Handbook · 6 chapters

The Computer Vision Metrics Handbook

Every score in computer vision, from the ground up. One confusion matrix grows into precision and recall; those scale to many classes, stretch over boxes and masks, follow identities through time, and finally measure images with no ground truth at all. Six chapters, with the equations, the intuition, and runnable code.

Start with Chapter 1 → Free · no signup
  1. 01
    Classification

    Classification Metrics, From the Ground Up: Precision, Recall, F1, ROC and AUC

    The confusion matrix and everything that falls out of it: precision, recall, F1, specificity, the precision/recall trade-off, PR and ROC curves, and AUC, all built on one running toy example (a backyard cat-cam), with the equations, the intuition, and scikit-learn.

    Confusion matrixPrecisionRecallF1SpecificityPR & ROC curvesAUC
  2. 02
    Multiclass & multilabel

    Multiclass and Multilabel Metrics: Micro, Macro, Weighted, and the Difference That Matters

    How precision/recall/F1 scale past two classes: one-vs-rest, and the micro/macro/weighted averaging that decides whether rare classes count. Plus multilabel-specific metrics: subset accuracy, Hamming loss, and top-k, all on one running shelter-camera example, with scikit-learn.

    One-vs-restMicroMacroWeightedSubset accuracyHamming lossTop-k
  3. 03
    Object detection

    Object Detection Metrics: IoU, AP, and the mAP Everyone Quotes

    How detection is scored: IoU to decide a hit, greedy matching into TP/FP/FN, per-class Average Precision as the area under the precision–recall curve, and the COCO mAP@[.5:.95] you see in every paper, all on one running parking-lot example, with torchmetrics.

    IoUAPPR curvemAP
  4. 04
    Segmentation

    Segmentation Metrics: IoU, Dice, mIoU, and Panoptic Quality

    Scoring masks instead of boxes: pixel IoU/Jaccard and the Dice coefficient (and why they're related), mIoU for semantic segmentation, mask-AP for instance segmentation, and Panoptic Quality for the unified task, on one running cat-mask example, with the equations and code.

    IoU / JaccardDicemIoUMask APPanoptic Quality
  5. 05
    Object tracking

    Object Tracking Metrics: MOTA, MOTP, IDF1, and HOTA

    Tracking adds identity over time, so its metrics score two things at once: detecting objects and keeping their IDs consistent. MOTA, MOTP, ID switches, IDF1, and the modern HOTA that balances detection against association, on one running door-cam example, with the equations and motmetrics.

    MOTAMOTPID switchesIDF1HOTA
  6. 06
    Image generation

    Image Generation Metrics: FID, IS, KID, LPIPS, and CLIPScore

    How do you score images with no ground truth? You compare distributions and use learned perceptual models: Inception Score, Fréchet Inception Distance, KID, LPIPS, and CLIPScore for prompt fidelity, on one running cat-generator example, with the equations, the catches, and torchmetrics.

    Inception ScoreFIDKIDLPIPSCLIPScore

Get new chapters by email

The handbook grows over time. Subscribe for new chapters and explainers as they publish, plus short notes on computer vision.