Before this unit: Homogeneous coordinates

Bird's-Eye View and Panoramas: Warping with a Homography

Lesson 5 of Image Alignment. Pushing each pixel forward through H leaves holes in 56% of the far court's bird's-eye view; pulling each output pixel back fills every one exactly once, and matches OpenCV's warpPerspective to one grey level. Interpolation decides what the pulled-back value is. Then a median of 31 warped frames removes the players, who cover 7% of the court in one frame.

Luis Condados ·
The court from directly above, resampled from 31 frames and blended by median: the players, who were never on the plane, are gone.
The court from directly above, resampled from 31 frames and blended by median: the players, who were never on the plane, are gone.

With a homography in hand you can redraw a whole photograph as if it had been taken from somewhere else: straight down onto a court, flat onto a page, or joined onto the next photograph of a panorama.

Two photographs of Barcelona harbour: the second is pulled onto the first by a homography, then the seam between them is blended away
Warp, then blend: the second photo pulled onto the first, a hard seam, then a multiband blend. “BarcelonaHarbour1.jpg” and “BarcelonaHarbour2.jpg” by Pap3rinik (Wikimedia Commons), public domain.

The same two steps, move the pixels and then combine them, are behind:

  • Bird’s-eye views of a court, a road or a car park, from a camera that sees them at an angle.
  • Document scanning, where the page is warped to a rectangle.
  • Panorama modes, where each photo is warped into a shared frame and the overlaps blended [3].
  • Graphics on the ground in sport, where the graphic is warped onto the pitch and blended under the players:
The same pickleball rally with a condados.ai banner lying flat on the near court, staying fixed to the floor while the players move over it
A banner warped onto the court and drawn under anything that differs from the empty court, so the players pass over it. Frames from “2026.07.25 WD Open … (Gold Medal match)” by pickleball4you, CC BY 3.0.

With a homography in hand, there are two ways to move an image through it, and only one works. Pushing every source pixel to where HH sends it leaves holes wherever the output is stretched: 56% of the far half of this court’s bird’s-eye view is never written. Pulling every output pixel back through H−1H^{-1} writes each one exactly once. The value you pull lands between four source pixels, and interpolation decides how to mix them. Once several images share one output, blending decides how to combine them: a median of 31 frames removes the players, who cover 7.0% of the court in any single one.

Where we are

Part of image alignment. Lessons 2 and 3 produced a homography from the court’s plane to frame 45000, and RANSAC is how you get one when some of the point pairs are wrong. This lesson uses it: every pixel of a bird’s-eye view of the court, at 2 cm per pixel, is filled from the photograph.

Push or pull

The obvious way is forward: take each pixel of the frame, send it through HH, and write it where it lands. The trouble is that HH stretches the far court and squeezes the near one. Where it stretches, neighbouring source pixels land two or three output pixels apart and the gaps between them are never written. Where it squeezes, several land on the same output pixel and all but the last are thrown away.

Backward mapping turns the loop around. For each output pixel, apply H−1H^{-1} to find where it came from in the source, and read the source there. Every output pixel is visited once and gets exactly one value. That is why warpPerspective and every resampler like it work this way [1].

forward: source → output5 of 9 never writtenbackward: output → sourcelands between four:interpolateevery outputpixel, once
Forward mapping writes where the source pixels happen to land. Backward mapping asks, for each output pixel, where it came from, and pays for that with a position that is almost never a whole pixel.

On the court, measured over the 205,569 top-view pixels the camera sees:

forward mappingnear halffar halfwhole court
pixels never written (holes)12.2%55.8%34.0%
pixels written more than once78.1%

The near half is compressed, most of its output pixels hit several times, and still leaves 12% unwritten; the far half is too sparse. The backward warp written out below and cv2.warpPerspective agree to a maximum difference of 1 grey level, over every visible pixel.

The bird's-eye view of the court filled by forward mapping, with most of the far half in magenta where no source pixel landed
Forward mapping. Magenta: output pixels no source pixel landed on. The far court, stretched the most, is mostly holes.

What the pulled-back value is

A pulled-back position is almost never a whole pixel. Nearest neighbour takes the closest one. Bilinear takes the four around it and mixes them by how close each is, first along xx, then along yy [1].

import cv2
import numpy as np

frame = cv2.imread("frame_045000.png")
H_court2img = np.array([[214.196342, 90.15869, -26.259561],
                        [28.479672, -12.463675, 587.254072],
                        [0.095124, -0.077166, 1.0]])
s, m = 50, 1.5                                   # 50 px per metre, 1.5 m margin
H_court2top = np.array([[0, s, m * s], [-s, 0, (13.41 + m) * s], [0, 0, 1.0]])
H_img2top = H_court2top @ np.linalg.inv(H_court2img)

top = cv2.warpPerspective(frame, H_img2top, (455, 820), flags=cv2.INTER_LINEAR)
cv2.imwrite("topview.png", top)                  # backward mapping inside
#include <opencv2/imgcodecs.hpp>
#include <opencv2/imgproc.hpp>

int main() {
    cv::Mat frame = cv::imread("frame_045000.png");
    cv::Matx33d H_court2img(214.196342, 90.15869, -26.259561,
                            28.479672, -12.463675, 587.254072,
                            0.095124, -0.077166, 1.0);
    const double s = 50, m = 1.5;
    cv::Matx33d H_court2top(0, s, m * s, -s, 0, (13.41 + m) * s, 0, 0, 1);
    cv::Matx33d H_img2top = H_court2top * H_court2img.inv();
    cv::Mat top;
    cv::warpPerspective(frame, top, H_img2top, {455, 820}, cv::INTER_LINEAR);
    cv::imwrite("topview.png", top);
}
Two enlarged crops of the far court in the bird's-eye view: nearest neighbour on the left with blocky, stepped lines, bilinear on the right with smooth lines
The far service line and centre line in the bird’s-eye view, enlarged four times. Left: nearest neighbour, each block one source pixel repeated. Right: bilinear. The diagonal texture in both is the net’s mesh, which is in front of this part of the court.

Blending: several images, one output

Stitching a panorama means several images land on the same output, and some rule has to combine them where they overlap. The simplest is to average, which ghosts anything that moved; feathering weights each image down towards its own edge; multiband blending does that separately for coarse and fine detail, so seams vanish without blurring [2], and it is what Brown and Lowe’s stitcher uses [3]. Szeliski covers the whole family under compositing [4].

A fixed camera gives the cleanest case to measure. The 31 frames either side of frame 45000 all share one homography, and a per-pixel median keeps whatever most of them agree on: the court, not the players. In frame 45000, 7.0% of the court’s bird’s-eye pixels differ from the median by more than 40 levels in some colour channel, and those pixels are the players and their shadows.

The court from above with no players on it: clean white lines on blue, the grey non-volley zones, and the net
The median of 31 frames, warped to the court’s plane. The players are gone because each one stands on a given patch of court in fewer than half of the frames. Frames from “2026.07.25 WD Open … (Gold Medal match)” by pickleball4you, CC BY 3.0.

The order does not matter here, and in a panorama it does. Taking the median of the 31 frames first and warping once, or warping all 31 and taking the median, differs by a mean of 0.55 grey levels (99% of pixels within 2), which is interpolation rounding. It works because every frame has the same HH. In a panorama each image has its own, and the only place the images agree about where a pixel is is the output, so there you have to warp first and blend second.

Now you try

The Warp tab renders the bird’s-eye view live, pulling every output pixel back through the homography the four corners define. Drag a corner a few pixels and watch the far court swing while the near court barely moves: that is the scale map from lesson 3, seen as an image. Toggle Bilinear off to see nearest neighbour.

Try it: drag the corners of the court

Frame 45000 of "2026.07.25 WD Open … (Gold Medal match)" by pickleball4you, CC BY 3.0.

In the wild

A panorama is where blending earns its place. Two photographs of Barcelona harbour taken from one spot on Montjuïc, panning right, published as source frames for a stitched panorama [5]. Everything in this unit, in order, at 1600 px wide: SIFT finds 6,788 and 8,197 keypoints, the ratio test keeps 1,011 matches, RANSAC keeps 906 of them (89.6%), and the DLT through those 906 reprojects them with a median error of 0.80 px. The second image is warped into the first’s frame by backward mapping.

Two photographs of a container port stitched into one wide panorama, the second warped into the frame of the first
The two photographs stitched through one homography and blended with a multiband blend. “BarcelonaHarbour1.jpg” and “BarcelonaHarbour2.jpg” by Pap3rinik (Wikimedia Commons), public domain.

The second photograph is exposed brighter: over the overlap, the first’s mean brightness is 0.76 of the second’s. That is what a seam shows, and the three blends handle it very differently. Measured as the average brightness step across the seam line, against 4.9 grey levels for the same measure a few columns away where there is no seam:

blendstep across the seam
hard cut40.5 grey levels
feather4.6
multiband, 5 levels5.5
Three crops of the same seam region: a hard cut with a visible vertical brightness step, a feathered blend and a multiband blend with no visible step
The same 400 px around the seam. Left: hard cut, the exposure step is a vertical line. Middle: feather. Right: multiband. Both blends bring the step down to the level of ordinary image texture.

Feather and multiband both erase the step here, and feather scores slightly better on this one number. Where they differ is away from the seam line: feathering spreads the exposure difference across the whole overlap as a slow gradient, visible in the sky of the panorama above as a lighter band, while multiband blends the low frequencies over a wide zone and the fine detail over a narrow one [2]. The step metric only looks at the seam itself, which is one of this comparison’s limits.

Where this breaks

Only the plane is mapped correctly. A homography is exact for points on the court and wrong for everything else. The player in frame 45000 comes out of the warp smeared into a long streak away from the camera, because the warp treats her head as a point on the floor far behind her feet. The net is the same, which is why its mesh is printed across the far court above. In a panorama the same effect is called parallax: step sideways between shots and anything off the dominant plane ghosts, whatever the blending.

The lens is not a pinhole, and the warp carries its bend into the output. The unit hub measures how much: the near baseline bows by 7.7 px, and a one-parameter correction cuts the far-sideline error from 5.95 cm to 1.96.

Next

That is the end of the unit: a map from a plane in one image to a plane anywhere else, computed from point pairs, made robust, and used. When the scene is not a plane, one 3×3 no longer describes it, and the tools are the ones in structure from motion. The unit’s hub collects every number and its limits in one place.

Run it: every code block on this page has a cell in the unit’s notebook, open it in Colab.

References

[1] OpenCV (5.x, commit a0cda08). warpPerspective, documentation comment in modules/imgproc/include/opencv2/imgproc.hpp, which defines each output pixel as the source sampled at M applied to (x, y), with M inverted first unless WARP_INVERSE_MAP is set. github.com/opencv/opencv

[2] Burt, P. J., & Adelson, E. H. (1983). A Multiresolution Spline with Application to Image Mosaics. ACM Transactions on Graphics, 2(4), 217–236. doi:10.1145/245.247

[3] Brown, M., & Lowe, D. G. (2007). Automatic Panoramic Image Stitching using Invariant Features. International Journal of Computer Vision, 74(1), 59–73. doi:10.1007/s11263-006-0002-3

[4] Szeliski, R. (2022). Computer Vision: Algorithms and Applications (2nd ed.), §8.4 “Compositing”, including §8.4.2 “Pixel selection and weighting (deghosting)” and §8.4.4 “Blending”. Springer. Free PDF

[5] Pap3rinik (2006). BarcelonaHarbour1.jpg and BarcelonaHarbour2.jpg. Wikimedia Commons, public domain. commons.wikimedia.org