Chroma key benchmark
We measured the keyer behind this site against two naive baselines on six synthetic fixtures. Every number below is regenerated by one command from a seeded generator, and the full write-up is archived with a DOI. We claim no accuracy advantage over a naive RGB threshold — and this page explains why.
Setup
- Six synthetic fixtures, 320×240: flat green, green with a shadow gradient, green with spill on the subject, green with both, flat blue, and a near-white neutral backdrop.
- Metric: intersection over union (IoU) between the produced alpha and the known ground-truth matte.
- Baselines are given oracle thresholds — each is evaluated at its own best value, which is the fairest possible comparison. Comparing against a badly tuned baseline would be meaningless.
- Reproduce with
node bench/run.mjsandnode bench/baseline.mjsin the reference repository.
Results — mean IoU over the six fixtures
| Method | Threshold | Mean IoU |
|---|---|---|
| Naive RGB distance (baseline A) | oracle | 0.995662 |
| Cb/Cr distance keyer (this site) | best tolerance | 0.994378 |
| Cb/Cr distance keyer (this site) | 25 — shipped default | 0.873398 |
| Green-dominance rule (baseline B) | oracle | 0.657181 |
What the table actually says
A naive RGB threshold matches us to within 0.002 mean IoU. That is the honest reading, and it is why we do not advertise the keyer as more accurate. On flat, evenly lit, synthetic backdrops, any colour-distance rule saturates — there are only two kinds of pixel, and any threshold separates them. A benchmark like this has almost no discriminating power.
What does separate the methods is the case where the assumption breaks:
- The green-dominance rule collapses on non-green backdrops. It scores 0.000000 on the flat blue and near-white fixtures. That is the argument for making the key colour a parameter rather than hard-coding green.
- Our shipped default scores 0.873398, well below our own best tolerance of 0.994378. Two fixtures — the shadowed ones — drag it down to roughly 0.62. A single global threshold cannot cover both a lit backdrop and a shadowed one in the same frame; you either keep the shadow or lose part of the subject. This is a structural property of the method, not a tuning bug, and it is why the tolerance control exists.
Limits you should hold us to
- Everything here is synthetic. No real footage, no real camera noise, no compression artefacts. A number that looks good on a gradient does not promise the same on a phone video.
- IoU rewards area, not edge quality. It says almost nothing about hair, fur, motion blur or fine detail — the places where keyers actually differ.
- The baselines are weak. A modern learned matting model would beat every row in this table. We compared against naive rules because they are what a dependency-free browser implementation competes with.
Full write-up
The preprint with the method, the fixture construction and the complete tables is archived at doi.org/10.5281/zenodo.22916767 (CC BY 4.0). The source and the CSVs are in the chroma-key-reference repository.