Remove a white background without a green screen

A chroma key needs colour to separate on, and a white wall has none. This is what happens instead: the keyer gives up on colour and measures brightness, which works — until the wall stops being flat.

Why a chroma key has nothing to hold on to

A chroma key does not separate subject from backdrop by brightness. It converts each pixel to YCbCr, discards the brightness, and measures the distance from the key colour in the two colour-difference channels. Discarding brightness is deliberate: a shadow on a green screen is the same paint with less light, and in colour terms it has not moved.

A white wall breaks that. In key-presets.csv the white, black and grey backdrops all carry Cb 0.0000 and Cr 0.0000. Every pixel on a neutral backdrop — wall, subject, shadow, highlight — lands on the same point in the colour plane however light or dark it is. A distance in that plane is zero everywhere, and a threshold on zero sorts nothing.

That is not a tuning problem. baseline-comparison.csv runs a naive green-dominance keyer on the neutral fixture with the best threshold available: IoU 0.000000. Not weak — empty. There was never a signal.

What runs instead: a luminance key

The keyer checks the key colour's chroma magnitude first. Under 7, which is true for white, black and grey, it stops measuring colour and measures brightness:

d = |luma(pixel) − luma(key)|

Pixels inside a radius of the wall's brightness go transparent; everything further out stays. One number per pixel — the only lever that can tell a white wall from a white shirt, and what the tool falls back to it for.

One other thing changes, and it matters if you move the slider. The radius is lo = (tolerance / 100) × scale, and the scale is 118 for a colour key but 150 for a neutral one. At the default 25 that is 29.5 on green and 37.5 here: the same slider position, covering more ground. The tolerance page has the full sweep.

It works, right up to the point where it doesn't

The neutral-white fixture is a flat #f2f2f2 field with a solid #9a9a9a rectangle in front of it, 320×240. At the default tolerance the score is IoU 1.000000: 49,579 backdrop pixels keyed out, no subject pixels lost. Auto-detection reads the wall off the frame border and returns #f2f2f2 exactly — max channel delta 0.

The two greys are 88 units of luma apart: the wall is 242, the prop is 154. One number, one gap. That makes the sweep a step, not a curve:

The arithmetic explains the cliff. 0.60 × 150 = 90, and 90 is past 88. Both surfaces are flat, so every subject pixel sits at the same distance and all cross together — and raising tolerance only pushes the radius outward, never pulling them back.

Against a green screen the difference is the whole story. On shadow-green the score climbs gradually — 0.617718 at 25, 0.671108 at 30, 0.770641 at 35, 1.000000 at 40 — because a shadowed screen spreads the backdrop across a range of distances. A flat neutral backdrop has no range: perfect or nothing, with one abrupt edge between.

The margin you are actually working with

88 units out of 255 is about a third of the range, and the default radius of 37.5 uses less than half of it. On a flat synthetic wall that is comfortable. A real wall is not flat — it falls off towards the corners, takes the subject's shadow, and picks up whatever the room bounces onto it. Each of those eats into the same 88 units, and the subject sits at the near end of them.

A black backdrop runs the same code with the distances flipped: key-presets.csv lists black-wall at luma 0.0 and grey-card at 128.0, both neutral, both with the 37.5 default. Which one you want stops being a colour decision — it is whichever end of the brightness range your subject sits further from.

How to shoot so the key has something to measure

Where these numbers come from

Every figure comes from data/key-presets.csv, data/baseline-comparison.csv, data/auto-key-accuracy.csv and data/tolerance-sweep.csv in the chroma-key-reference repository. The method is written up at DOI 10.5281/zenodo.22916767. The fixtures are synthetic: 320×240 flat colour fields with a rectangle or ellipse subject in front of them, generated deterministically so the sweep reproduces exactly. They are not real photographs, and a real room has uneven light, sensor grain, compression and a subject made of dozens of colours. Read these as the shape of the behaviour, not as settings to copy.

When this does not work