← All posts The Computer Vision Metrics Handbook · Chapter 3 of 6 View all →

Object Detection Metrics: IoU, AP, and the mAP Everyone Quotes

How detection is scored: IoU to decide a hit, greedy matching into TP/FP/FN, per-class Average Precision as the area under the precision–recall curve, and the COCO mAP@[.5:.95] you see in every paper, all on one running parking-lot example, with torchmetrics.

Luis Condados · · Updated August 9, 2026
Object Detection Metrics: IoU, AP, and the mAP Everyone Quotes

“mAP 0.54” is three nested ideas stacked into one number: a geometric overlap test, a precision–recall curve, and an average over IoU thresholds and classes. We’ll unpack all three on one running scene, a parking-lot camera counting cars, with numbers you can check and boxes you can drag.

Detection adds localization on top of classification [4], so it reuses precision, recall, and AP from the foundations, but first it needs a way to decide whether a predicted box even counts as a hit.

The running example: LotCam, the parking-lot camera

LotCam watches a parking lot and draws a box around every car. In one frame there are 5 real cars (ground truth). LotCam returns 7 boxes, each with a confidence: 5 of them land on real cars (at various confidence), and 2 are mistakes: a trash bin it thinks is a car (0.60) and a shadow (0.30). We’ll score this exact frame every way detection knows how.

Step 1. IoU: is this box close enough?

Intuition: a prediction only counts if it lands on the object, and IoU measures how much the predicted box and the true box overlap, as a fraction of the total area they cover together.

IoU(A,B)=ABAB.\text{IoU}(A, B) = \frac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}.

It’s 0 for disjoint boxes, 1 for a perfect match.

Drag either box and watch overlap, IoU, and Dice react, and find where IoU crosses 0.5 (the box “becomes” a hit) and how forgiving Dice is by comparison:

ground truth prediction
Drag either box.

Step 2. Match predictions to truth (TP / FP / FN)

Intuition: hand the highest-confidence box first pick of the true cars, then the next, and so on, like seating guests by priority.

Within a class, sort predictions by confidence and greedily assign each to the best unused ground-truth box above the IoU threshold:

  • matches an unused GT box → True Positive
  • matches nothing, or a GT already taken → False Positive (duplicate boxes become FPs, which is why NMS matters)
  • a GT box left with no match → False Negative

There are no true negatives in detection, because the background isn’t enumerable, which is exactly why detection lives on precision/recall, not accuracy.

Step 3. Average Precision per class

Intuition: don’t pick one confidence threshold. Sweep them all and measure the area under the resulting precision–recall curve. A detector that keeps precision high as recall climbs scores near 1.

AP=n(RnRn1)P(Rn).\text{AP} = \sum_{n} (R_n - R_{n-1})\, P(R_n).

Modern (COCO) evaluation interpolates, replacing P(Rn)P(R_n) with the best precision at any recall ≥ the current one, to iron out wiggles; on a clean curve it barely moves the number. (Pascal VOC used a coarser 11-point version; Padilla et al. [3] work through both interpolations side by side.)

precisionrecall1.01.0AP ≈ 0.93
LotCam’s PR curve: precision holds at 1.0 through recall 0.6, dips when the trash bin enters, and the shaded area is AP ≈ 0.93.

Now you try

Each dot is one of LotCam’s 7 boxes (green = a real car, amber = a false box), ranked by confidence. Drag the threshold: watch precision and recall trade off and the live dot ride the PR curve. Confirm that no threshold gets you both perfect, because the trash bin at 0.60 always costs precision unless you also drop the lower-confidence real cars.

Step 4. Mean over classes (and over IoU): mAP

Intuition: compute that AP for every class and average. Then, because “hit” depends on the IoU bar, average over a range of IoU thresholds too.

mAP=1Cc=1CAPc.\text{mAP} = \frac{1}{C} \sum_{c=1}^{C} \text{AP}_c.

The other catch is the IoU threshold from Step 1, the single biggest source of “why don’t these numbers match?”:

  • Pascal VOC mAP / AP50 uses a single IoU threshold of 0.5. Lenient on localization.
  • COCO mAP is the headline, averaged over ten IoU thresholds 0.50→0.95:
mAPCOCO=110t{0.50,,0.95}mAP@t.\text{mAP}_{\text{COCO}} = \frac{1}{10} \sum_{t \in \{0.50, \dots, 0.95\}} \text{mAP@}t.
0.93car0.71truckmAP 0.82
mAP is just the average of the per-class AP bars, and the gap between them is what the single number throws away.

COCO also reports AP75 (strict) and AP-small / -medium / -large by object area, which is invaluable for spotting a detector that’s fine on big objects and quietly failing on small ones.

In code (torchmetrics)

The reference implementation is COCO’s pycocotools; torchmetrics wraps the same protocol with a friendlier API [5]:

import torch
from torchmetrics.detection import MeanAveragePrecision

# boxes are [x_min, y_min, x_max, y_max]
preds = [dict(
    boxes=torch.tensor([[10., 10., 50., 50.], [60., 60., 90., 90.]]),
    scores=torch.tensor([0.90, 0.45]),
    labels=torch.tensor([0, 1]),
)]
target = [dict(
    boxes=torch.tensor([[12., 12., 48., 52.], [62., 58., 92., 90.]]),
    labels=torch.tensor([0, 1]),
)]

metric = MeanAveragePrecision(iou_type="bbox", class_metrics=True)
metric.update(preds, target)
result = metric.compute()
print(result["map"])       # COCO mAP@[.5:.95]
print(result["map_50"])    # AP50 (VOC-style)
print(result["map_75"])    # AP75 (strict localization)
print(result["map_small"], result["map_medium"], result["map_large"])

The same iou_type="segm" switches the whole protocol to mask IoU for instance segmentation, where the metric machinery is identical and only the overlap definition changes (see segmentation metrics).

Caveats & common pitfalls

  • The IoU convention is the number-one reason scores don’t match. AP50 (VOC) and COCO mAP@[.5:.95] are different metrics: LotCam’s “hit” at IoU 0.68 is a TP at 0.5 and a miss at 0.75. Always state which, plus the dataset and split [1][2].
  • The confidence threshold silently breaks mAP. Evaluating at a high visualization threshold (e.g. 0.5) truncates the PR curve, losing exactly the car E we lost above, and collapses the score. Evaluate near 0.0–0.05 (this bites in practice, see the edge benchmark).
  • mAP averages away where you fail. The 0.82 hid trucks at 0.71; read AP-small/medium/large and per-class AP [2].
  • Greedy one-to-one matching has edge cases (duplicate boxes, crowd regions, ignore flags). Use the reference pycocotools protocol, or your numbers won’t be comparable.

Takeaways

  • IoU decides a hit; the threshold you choose is half the story (0.68 = hit at 0.5, miss at 0.75).
  • Detection has no true negatives → precision/recall/AP, never accuracy.
  • AP is area under the per-class PR curve (LotCam’s car = 0.93); mAP averages it over classes (0.82 with trucks) and over IoU thresholds.
  • Always state the IoU convention: AP50 (VOC) ≠ COCO mAP@[.5:.95].
  • Read AP-small/medium/large. It tells you where the detector fails.

Next: when the prediction is a mask, not a box → Segmentation metrics: IoU, Dice, and Panoptic Quality.

Further reading

References

[1] Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., & Zisserman, A. (2010). The PASCAL Visual Object Classes (VOC) Challenge. International Journal of Computer Vision, 88(2), 303–338. DOI:10.1007/s11263-009-0275-4.

[2] Lin, T.-Y., et al. (2014). Microsoft COCO: Common Objects in Context. ECCV. arXiv:1405.0312. (Defines the COCO mAP@[.5:.95] protocol.)

[3] Padilla, R., Netto, S. L., & da Silva, E. A. B. (2020). A Survey on Performance Metrics for Object-Detection Algorithms. International Conference on Systems, Signals and Image Processing (IWSSIP).

[4] Szeliski, R. (2022). Computer Vision: Algorithms and Applications (2nd ed.), §6.3 (object detection). Springer. Free PDF.

[5] TorchMetrics. Mean-Average-Precision (mAP). Docs. Reference protocol: pycocotools.