The Computer Vision Metrics Handbook
Every score in computer vision, from the ground up. One confusion matrix grows into precision and recall; those scale to many classes, stretch over boxes and masks, follow identities through time, and finally measure images with no ground truth at all. Six chapters, with the equations, the intuition, and runnable code.
- 01Classification
Classification Metrics, From the Ground Up: Precision, Recall, F1, ROC and AUC
The confusion matrix and everything that falls out of it: precision, recall, F1, specificity, the precision/recall trade-off, PR and ROC curves, and AUC, all built on one running toy example (a backyard cat-cam), with the equations, the intuition, and scikit-learn.
Confusion matrixPrecisionRecallF1SpecificityPR & ROC curvesAUC - 02Multiclass & multilabel
Multiclass and Multilabel Metrics: Micro, Macro, Weighted, and the Difference That Matters
How precision/recall/F1 scale past two classes: one-vs-rest, and the micro/macro/weighted averaging that decides whether rare classes count. Plus multilabel-specific metrics: subset accuracy, Hamming loss, and top-k, all on one running shelter-camera example, with scikit-learn.
One-vs-restMicroMacroWeightedSubset accuracyHamming lossTop-k - 03Object detection
Object Detection Metrics: IoU, AP, and the mAP Everyone Quotes
How detection is scored: IoU to decide a hit, greedy matching into TP/FP/FN, per-class Average Precision as the area under the precision–recall curve, and the COCO mAP@[.5:.95] you see in every paper, all on one running parking-lot example, with torchmetrics.
IoUAPPR curvemAP - 04Segmentation
Segmentation Metrics: IoU, Dice, mIoU, and Panoptic Quality
Scoring masks instead of boxes: pixel IoU/Jaccard and the Dice coefficient (and why they're related), mIoU for semantic segmentation, mask-AP for instance segmentation, and Panoptic Quality for the unified task, on one running cat-mask example, with the equations and code.
IoU / JaccardDicemIoUMask APPanoptic Quality - 05Object tracking
Object Tracking Metrics: MOTA, MOTP, IDF1, and HOTA
Tracking adds identity over time, so its metrics score two things at once: detecting objects and keeping their IDs consistent. MOTA, MOTP, ID switches, IDF1, and the modern HOTA that balances detection against association, on one running door-cam example, with the equations and motmetrics.
MOTAMOTPID switchesIDF1HOTA - 06Image generation
Image Generation Metrics: FID, IS, KID, LPIPS, and CLIPScore
How do you score images with no ground truth? You compare distributions and use learned perceptual models: Inception Score, Fréchet Inception Distance, KID, LPIPS, and CLIPScore for prompt fidelity, on one running cat-generator example, with the equations, the catches, and torchmetrics.
Inception ScoreFIDKIDLPIPSCLIPScore
Get new chapters by email
The handbook grows over time. Subscribe for new chapters and explainers as they publish, plus short notes on computer vision.