What Color Rendering Actually Measures
Why this matters
A rendering index is the number a customer points at when they say the new lights make everything look wrong, and it is almost always the number that says the new lights are fine. That is not because the customer is imagining it. The index is an average of eight comparisons against a reference that changes depending on which lamp you are testing, and both of those design choices throw away exactly the information a complaint about red merchandise or skin tone depends on. A shop that can say what the number contains can specify around the gap for the price of one extra line on a purchase order.
The metric compares a source to a moving reference
The general color rendering index, written Ra or loosely called CRI, works like this. Take eight standard test color samples, moderate saturation, no strong reds among them. Compute the color each sample would appear under a reference illuminant, compute the color it appears under the source being tested, and score the shift. Do that eight times and average the eight scores. A source that shifts nothing scores 100.
The reference is the part people miss. It is chosen to have the same correlated color temperature as the source under test. Below 5000 K the reference is a blackbody radiator at that temperature; at or above 5000 K it is a phase of daylight at that temperature. So a 2700K lamp is judged against a 2700K blackbody, and a 5000K lamp is judged against 5000K daylight, and neither is judged against the other.
The immediate consequence: a rendering index is not comparable across color temperatures. Two products both posting Ra 85, one at 2700K and one at 5000K, are equally close to two different targets. Neither number tells you which one will make a red sweater look like a red sweater, and comparing the two as though 85 equals 85 is the single most common misuse of the metric.
What the averaging hides
An arithmetic mean of eight numbers dilutes any one of them by a factor of eight. A single sample sitting 40 points below its peers moves the average by 40 / 8 = 5.0 points. So a five-point difference in Ra between two products can be one catastrophic sample or eight mildly imperfect ones, and the index cannot tell you which.
This is not theoretical. Two products, both real-shaped profiles:
| Ra (mean of eight) | Lowest single sample | Mean of the other seven | |
|---|---|---|---|
| Candidate M | 90 | 71 | 92.7 |
| Candidate N | 85 | 79 | 85.9 |
M's headline is 5 points better and M's worst sample is 8 points worse. Whichever object in the space happens to fall near that sample is the object the customer will complain about, and the headline told you the opposite of what to expect. Ask for the individual sample values, not the average. Every test report has them.
R9 is missing on purpose, and it is the one people notice
The eight samples in the average are moderately saturated pastels. Saturated red is not among them. It is a supplementary sample, reported separately as R9, and it is excluded from Ra entirely.
That matters because saturated red is where the human eye is least forgiving and where the merchandise usually is: raw meat, produce, cosmetics, paint chips, wood finishes, and skin. A source can post Ra 90 and R9 in the low teens, and every one of those numbers is correctly reported.
So R9 gets specified separately, always, on any job where color decisions get made under the light. It costs one clause.
Fidelity is not preference
Ra and its more recent successor metrics all measure fidelity: closeness to the reference. High fidelity is the right target when the job is to judge a color accurately. It is not automatically the right target when the job is for the space to look good, because a source that mildly increases saturation often reads as more vivid and more pleasant while scoring lower on fidelity for exactly that reason.
That is why the current two-number system pairs a fidelity index with a gamut index, one saying how close the colors are to the reference and the other saying whether the source enlarges or shrinks the range of colors relative to it, with 100 meaning the same area. A gamut figure above 100 means saturation is being increased.
Two cautions in the same clause as the numbers. First, the newer fidelity index is computed from a much larger sample set on a different scale, so it may not be compared numerically to Ra; a 2-point difference in one is not a 2-point difference in the other. Second, its reference also follows the test source's chromaticity, so it inherits the same prohibition on cross-color-temperature comparison.
The specification line, built field by field
This is what actually goes on the order. Each field exists because a preceding section showed what breaks without it.
- Application, and which colors carry decisions. Paint retail: saturated reds and skin tone. Warehouse pick face: printed label legibility. This field decides how hard the rest of the line has to be.
- Color temperature the rendering figures are quoted at. Required, because the reference moves with it.
- Fidelity floor. A minimum Ra, at that stated color temperature.
- Saturated red floor. A minimum R9, stated separately, because Ra does not contain it.
- Worst-sample floor, where the application is color-critical. A minimum for the lowest individual sample, not just the mean.
- Gamut range, where appearance rather than accuracy is the goal.
- The document that proves it. The electrical and photometric test report for the exact ordering code, since rendering figures move with color temperature and drive current within a family.
- Unit-to-unit consistency, which is a chromaticity tolerance question rather than a rendering one, and belongs to the card on why two lamps at the same color temperature look different.
Running two candidates through it
A paint retailer's sales floor. Color decisions are made under the light, on saturated chips and on skin. Both candidates quoted at 3000K, which satisfies the second field and makes the rest comparable.
| Candidate M | Candidate N | |
|---|---|---|
| Ra | 90 | 85 |
| Lowest individual sample | 71 | 79 |
| R9 | 12 | 62 |
| Gamut index | 96 | 104 |
On the headline, M wins by 5 points. On this application, N wins outright, and the fields say why. R9 is 62 against 12, a factor of 5.2, on a floor that the application makes non-negotiable. The worst individual sample is 79 against 71, so M's spread is wider as well as lower at the bottom. The gamut figures split around 100, with N at 104 slightly increasing saturation and M at 96 slightly reducing it, which on a sales floor favors N again and on a color-matching bench would favor M.
What flips the recommendation. Move this to a paint mixing station where the job is to judge whether a match is correct rather than to sell the chip, and the gamut preference reverses: you want the source that does not enlarge the color range, because it will not flatter a mismatch. The R9 and worst-sample floors do not move, because accuracy on red still matters. So the same two candidates, twenty feet apart in the same store, resolve differently, and only the field-by-field line makes that visible.
A rendering figure describes one operating point
Every index above is measured on a stabilized sample at one drive current, one thermal condition and one point in its life. Three things move it afterwards, and none of them appear on the sheet.
Dimming. Some drivers and modules shift chromaticity and spectral output as they dim, and where chromaticity moves, the reference the index is measured against moves with it. A product characterized at full output has not been characterized at 20 percent, and if the space will run dimmed most of the time, that is the condition to ask about.
Age. Phosphor systems change with operating hours, and the change is not required to be uniform across wavelengths. Whether the family holds its rendering over life is a color maintenance question with its own published data.
Thermal condition in the housing. The same module at a hotter junction is not spectrally identical to the one on the test bench.
Which is why the last step on any color-critical job is not arithmetic. Put the actual merchandise under the actual candidate at the actual illuminance and look at it, with the people who will make decisions under it, and do it at the level the space will really run rather than at full output on a bench. A mock-up settles in ten minutes what four indices only bound, and it is the one test that captures the source, the surfaces, the level and the observer at once.
Checking your own figures
- Dilution arithmetic printed. 40 / 8 = 5.0 points of movement in the mean from one bad sample, which is why a 5-point Ra gap between M and N carries no information about the worst case.
- The two profiles are internally consistent. M: mean 90, worst 71, so the other seven average (8 x 90 - 71) / 7 = 92.7. N: mean 85, worst 79, so the other seven average (8 x 85 - 79) / 7 = 85.9. Both recomputed rather than asserted.
- Cross-CCT comparison refused, then checked. Both candidates quoted at 3000K, so their Ra values are against the same reference and may be compared. Had one been at 4000K, the pair would have been struck.
- Metrics not blended. Ra values were compared only to Ra values, and the gamut figures only to each other and to 100. No Ra figure was set against a newer-scale fidelity figure.
- Every field in the line traces to a mechanism above it. The color temperature field to the moving reference, the worst-sample field to the eight-sample dilution, the R9 field to the excluded sample, the gamut field to the fidelity-versus-preference split.
- R9 gap stated with its base. 62 against 12 on a scale where 100 is the reference, a factor of 5.2, not a percentage of anything.
References
- CIE test-color method for the general color rendering index, including the reference illuminant selection at and above 5000 K, in the edition your specification names
- IES TM-30, the fidelity and gamut evaluation method, in the edition your specification names, which binds through that specification rather than on its own
- Manufacturer electrical and photometric test report for the exact ordering code, which is the only source for individual sample values and R9
- See related: What Color Temperature Does and Does Not Tell You; Why Two Lamps at the Same Color Temperature Look Different; How to Read a Luminaire Cut Sheet