Before this unit: Homogeneous coordinates

Affine vs Perspective Transform

Lesson 2 of Image Alignment. Written on (x, y, 1), a 3×3 matrix can translate (affine, six numbers) and, once its last row is allowed to vary, make parallel lines meet (projective, eight numbers). On a pickleball court, an affine map fixed by three corners misses the fourth by 3.50 m; a homography through four lands the held-out centre-line corner 0.47 cm from the rulebook.

Luis Condados ·
The court model pushed through an affine map fixed by three corners (red) and through a homography fixed by four (green). Only one of them lands on the paint.
The court model pushed through an affine map fixed by three corners (red) and through a homography fixed by four (green). Only one of them lands on the paint.

An affine transform can move, rotate, stretch and slant a picture, but lines that start parallel stay parallel. A perspective transform, or homography, can also make them converge, which is what a camera does to anything it sees at an angle.

A court diagram whose opposite sides are parallel; as one number in the bottom row of its matrix grows, the far end shrinks and the sidelines start to converge
The same court diagram as one number in the bottom row of the 3×3 grows from zero. At zero the map is affine and the sidelines stay parallel; any other value makes it a perspective map.

Which one you need depends on the camera:

  • Perspective: a document photographed at an angle, a projector aimed up at a wall (keystone correction), a court or a road seen from a camera beside it. Anything the camera sees obliquely.
  • Affine is enough: a scanner or a microscope, where the view is straight on and only shift, rotation and scale change; satellite images taken looking straight down; two frames of video a moment apart, when the camera has barely moved.
A printed card lying at an angle on a wooden bench; its four corners are marked, then the whole image is warped until the card is an upright rectangle
A document scanner needs the perspective map: the card’s far edge is shorter than its near edge in the photo, and no affine map can undo that. “Stationery on a bench (Unsplash).jpg” by Brigitte Tohm (Wikimedia Commons), CC0.

Write each point as (x,y,1)(x, y, 1) and a 3×3 matrix can do the two things a 2×2 could not. Its last column translates, which gives the six-number affine map. Letting its last row vary divides every output by a number that changes across the image, and that makes parallel lines meet: the eight-number projective map, or homography. On this court the difference is not subtle. An affine map fixed by three corners puts the fourth 3.50 m from where the paint is; a homography through four puts a corner it never saw 0.47 cm from the rulebook position.

Where we are

Part of image alignment. Lesson 1 showed that a 2×2 matrix cannot translate and cannot make parallel lines meet, and that on frame 45000 the best one misses every corner by 60 px. Both failures have the same fix, which is to use the extra coordinate from homogeneous coordinates.

First the last column: affine

Put the 2×2 in the top-left corner of a 3×3, a translation (tx,ty)(t_x, t_y) in the last column, and (0,0,1)(0, 0, 1) underneath:

(x′y′1)=(abtxcdty001)(xy1)\begin{pmatrix} x' \\ y' \\ 1 \end{pmatrix} = \begin{pmatrix} a & b & t_x \\ c & d & t_y \\ 0 & 0 & 1 \end{pmatrix}\begin{pmatrix} x \\ y \\ 1 \end{pmatrix}

The 1 at the end of the input picks up txt_x and tyt_y, so translation is now part of the multiplication. Six numbers, so three point pairs fix it. It still keeps parallel lines parallel, since the bottom row leaves the last coordinate at 1 and the top two rows are the old 2×2 plus a shift [1].

Then the last row: projective

Now let the bottom row be anything, (g,h,1)(g, h, 1):

x′=ax+by+txgx+hy+1y′=cx+dy+tygx+hy+1x' = \frac{a x + b y + t_x}{g x + h y + 1} \qquad y' = \frac{c x + d y + t_y}{g x + h y + 1}

The output is divided by a number that depends on where you are. Far from the camera the divisor is large and things shrink; near it they grow. That is perspective, and it is also why parallel lines can meet: the divisor is zero along one line of the input, and that line goes to infinity in the output, or the other way round. Eight numbers (the ninth is the overall scale, which does not matter for the same reason (2,4,2)(2, 4, 2) and (1,2,1)(1, 2, 1) are the same point), so four point pairs fix it [2].

Hartley and Zisserman put these in a ladder: a similarity keeps shape and has four numbers, an affine map keeps parallelism and has six, a projective map keeps only straight lines and has eight [1].

the court, from aboveaffine: 6 numbersopposite sides stay parallelprojective: 8 numbersonly straightness survives
The same square through the two maps. The affine image is a parallelogram whatever its six numbers; the projective one can be any four-sided shape, which is what a court looks like from behind a corner.

Why a plane is exactly projective

This is not a lucky fit. The camera models post writes a pinhole camera as a 3×4 matrix P=K[ r1 r2 r3 t ]P = K[\,r_1\ r_2\ r_3\ t\,] acting on (X,Y,Z,1)(X, Y, Z, 1). Every point of the court has Z=0Z = 0, so the third column never contributes, and what is left, K[ r1 r2 t ]K[\,r_1\ r_2\ t\,], is a 3×3 acting on (X,Y,1)(X, Y, 1). The map from a plane to a pinhole image is a homography, with no approximation beyond the pinhole itself [3].

The same comparison on every piece of held-out evidence the unit uses, in centimetres:

NBCNKCnear centre linefar right sidelinefar baseline
affine, 3 corners91.6188.5125.6397.0308.8
affine, least squares on 4122.434.178.6105.1551.6
homography, 4 corners0.475.632.665.9538.4

Giving the affine map all four corners and a least-squares fit does not rescue it. It spreads its error over the four corners it was given (42 to 114 cm at each) and gets worse on the far baseline, 5.52 m, because the far court is where perspective matters most and an affine map has no way to express it.

import cv2
import numpy as np

court = np.float32([[0.0254, 0.0254], [0.0254, 6.0746], [4.6004, 0.0254], [4.6004, 6.0746]])
image = np.float32([[-18.52, 587.39], [987.25, 959.91], [669.68, 500.09], [1555.24, 663.22]])

A_img2court = cv2.getAffineTransform(image[:3], court[:3])      # 2x3, six numbers
H_img2court = cv2.getPerspectiveTransform(image, court)         # 3x3, eight numbers
nbc = np.float32([[[332.06, 717.24]]])
print(cv2.transform(nbc, A_img2court))                          # [[[0.0252 2.1339]]]
print(cv2.perspectiveTransform(nbc, H_img2court))               # [[[0.0254 3.0547]]]
#include <opencv2/imgproc.hpp>
#include <iostream>

int main() {
    std::vector<cv::Point2f> court{{0.0254f, 0.0254f}, {0.0254f, 6.0746f}, {4.6004f, 0.0254f}, {4.6004f, 6.0746f}};
    std::vector<cv::Point2f> image{{-18.52f, 587.39f}, {987.25f, 959.91f}, {669.68f, 500.09f}, {1555.24f, 663.22f}};
    cv::Mat A = cv::getAffineTransform(image.data(), court.data());          // first three
    cv::Mat H = cv::getPerspectiveTransform(image, court);
    std::vector<cv::Point2f> nbc{{332.06f, 717.24f}}, a, h;
    cv::transform(nbc, a, A);
    cv::perspectiveTransform(nbc, h, H);
    std::cout << a[0] << " " << h[0] << "\n";
}
The court seen from behind one corner, with two overlays: a red affine court model that agrees at three circled corners and drifts badly everywhere else, and a green homography court model that lies on the painted lines, including the far half
Red: the whole court model through the affine map fixed by the three circled corners (the fourth is off the left edge). Green: through the homography fixed by all four. The green far baseline and far service lines were never fitted and still land on the paint behind the net. Frame from “2026.07.25 WD Open … (Gold Medal match)” by pickleball4you, CC BY 3.0.

Now you try

Start on Affine (3 pts). Drag the three white corners and watch the red model and the ghost of the fourth corner; the readout scores NBC and NKC in centimetres. Then switch to Homography (4 pts) and nudge one corner by a few pixels: the held-out error moves by centimetres near the camera, and the far-corner scale in the last line tells you why it moves by much more at the back.

Try it: drag the corners of the court

Frame 45000 of "2026.07.25 WD Open … (Gold Medal match)" by pickleball4you, CC BY 3.0.

In the wild

A different court, outdoors, from a different camera, where both baseline corners are outside the frame and were found as line intersections [4]. The same two maps, the same held-out checks:

outdoor courtNBCNKCnear centre line
affine, 3 corners111.3 cm115.5 cm47.3 cm
homography, 4 corners0.77 cm11.3 cm2.66 cm
An outdoor pickleball court seen from beside and behind the near baseline, with the full court model drawn in green over the painted lines, and two circled held-out corners
The outdoor court with the homography’s court model drawn over it. Two of the four corners it was fitted to are outside the frame; the circled corners were not used. Median of 31 frames of “2024.08.30 MS4.0 Philip Wong vs Tristan Clark (Round Robin, match 6)” by pickleball4you, CC BY 3.0.

The homography beats the affine map by 145× at NBC, 10× at NKC and 18× along the centre line; indoors the same factors were 195×, 33× and 47×. The size of the gap depends on the camera and the geometry of the evidence. The ordering, affine metres against projective centimetres, does not change.

Where this breaks

A homography through four points reproduces those four points exactly, whatever they are. If one of the four was found 3 px off, the fit bends the whole court to pass through the wrong pixel without complaint, and the only sign is in the held-out numbers. Look at the far baseline in the table: 38.4 cm, the largest error the homography makes, coming from four corners that are each good to a fraction of a pixel.

Four points leave nothing over to average against, and nothing to detect a bad one with. What is needed is a way to use more than four, and to say how an error in each one turns into an error on the court.

Next

Lesson 3 writes the homography as eight linear equations in nine unknowns, solves them with the SVD for any number of point pairs, and follows the error from the pixel to the far baseline.

Run it: every code block on this page has a cell in the unit’s notebook, open it in Colab.

References

[1] Hartley, R., & Zisserman, A. (2004). Multiple View Geometry in Computer Vision (2nd ed.), §2.4 “A hierarchy of transformations”, p. 37. Cambridge University Press.

[2] Hartley, R., & Zisserman, A. (2004). Multiple View Geometry in Computer Vision (2nd ed.), §2.3 “Projective transformations”, p. 32. Cambridge University Press.

[3] Hartley, R., & Zisserman, A. (2004). Multiple View Geometry in Computer Vision (2nd ed.), ch. 13 “Scene planes and homographies”, p. 325. Cambridge University Press.

[4] pickleball4you (2024). 2024.08.30 MS4.0 Philip Wong vs Tristan Clark (Round Robin, match 6). YouTube, CC BY 3.0. youtube.com/watch?v=K0qrASvix3Y