Green screen for YouTube Shorts: shoot vertical footage that keys clean

A Short is a tall frame, shot close, usually on a phone. All of that makes the key harder. Here is what goes wrong, what the measurements say, and how to export a clip that survives the upload.

What the tall frame does to the key

A Short is a 9:16 vertical frame — 1080×1920 is the usual target — and Google's own help page puts the length ceiling at three minutes. That shape changes the shot before it changes anything technical: to fill a tall frame, the camera comes in close, which puts the subject near the backdrop. Distance is what keeps their shadow and the screen's bounce off them, and the vertical frame spends it.

It also leaves less backdrop on screen: in a wide shot the green fills the sides, in a tight vertical shot it is a narrow strip the keyer must read the colour from. The tool samples the frame border to pick the key colour. When that strip is partly shadowed, the sample describes the shadow rather than the screen — on shadow-green, auto-detection returns #00a23b against a true #00b140, a max channel delta of 15. Not a misread: the border really is darker than the middle, so the key colour fits the darkest part of the screen.

The shadow is the whole problem

A chroma key works in the two colour-difference channels and discards brightness, which is why a shadow on the screen usually still keys out. "Usually" has a limit. Dimmed to 0.65 of its lit brightness, the screen sits 29.3412 from the key and still comes out transparent at the default tolerance; at 0.60 the distance is 33.5337 and the pixel flips to fully opaque. The default radius is 29.5. Past roughly two-thirds brightness, an unevenly lit screen starts coming back as solid green.

Tolerance is the fix people reach for, and it works: shadow-green goes from IoU 0.617718 at tolerance 25 to 1.000000 at 40. On a screen that was already clean, the same slider does the opposite — flat-green at tolerance 90 keys out all 21,341 subject pixels and falls to 0.722122. It widens one radius around one colour and cannot tell a shadowed screen from a green shirt. Light the screen flat and use tolerance as a small correction: the lighting guide and the tolerance page cover both.

Spill is more visible here, not more likely

The other artefact is green light bouncing off the screen onto skin, hair and light clothing. A vertical Short shows it off: the subject is large in the frame, so a rim two pixels wide in a wide shot is a visible halo at phone size.

What the measurements say is narrower than most advice. On spill-green, the shipped defaults score IoU 1.000000 — spill does not make the subject transparent. It is a colour problem, and a coverage metric cannot see it. Scrubbing the rim with the edge control does cost pixels: at edge −100, 6,902 of 21,341 subject pixels go and the score drops to 0.889322 — on a close-up, a bite out of a chin or a hand. The spill page covers what that control separates.

How to shoot it

Exporting: the frame ceiling and the clock

Two limits decide which export you can use, and both bite on longer clips.

A PNG sequence is held in memory before it is zipped, so it is capped at 1,200 frames. The frame-rate control sets how many frames get pulled out of your clip — it has nothing to do with what the camera shot. At 30 fps, a 60-second Short wants 1,800 frames, over the ceiling. At 20 fps it is exactly 1,200; the default 12 fps gives 720 frames of the same minute. The tool warns you and suggests a rate that fits rather than failing at the end.

A transparent WebM has no frame ceiling, but it records in real time: the keyed canvas is captured at 30 fps into VP9 at 8,000,000 bits per second, dimensions rounded down to even numbers. A one-minute clip takes about a minute; a backgrounded tab can stall it.

For a Short that usually means WebM past about 40 seconds, the sequence for anything short you want lossless. The sequence-versus-WebM page goes through the trade-off, and the video page walks through the export.

Where these numbers come from

Every figure comes from data/shadow-distance.csv, data/tolerance-sweep.csv, data/edge-sweep.csv, data/auto-key-accuracy.csv and data/key-presets.csv in the chroma-key-reference repository. The method is written up at DOI 10.5281/zenodo.22916767. The fixtures are synthetic: 320×240 flat colour fields with a rectangle or ellipse subject in front of them, generated deterministically so every sweep reproduces exactly. They are not real phone footage — a real room has uneven light, rolling shutter, sensor noise, heavy compression and a subject made of dozens of colours. Read these as the shape of the behaviour, not as settings to copy.

When this does not work