Ingredients · measured

We built a skincare rating.
Then we killed it. Twice.

Every app you have used puts a number on your moisturiser. We wanted one too, for an ordinary reason: a star next to a search result gets clicked. So we built it, and measured it, and it told us something we did not expect about our own shelf.

Test one: a safety score across 107 products from the brands we cover
0.0-1.01.0-2.012.0-3.013.0-4.0134.0-5.0925.0

A rating that never says no is not a rating.

The short version

Rating apps are accurate about what they measure. What they measure is narrower than the star implies, and the gap is where the trouble lives.

What they actually score
Ingredient hazard, not product quality. Yuka takes the single worst ingredient; EWG scores hazard without exposure.
What no label can tell anyone
Concentration. Order is only meaningful down to 1%, and nothing below that line is ranked.
What we found on our own shelf
92 of 107 products had no declarable fragrance allergen at all. There was nothing to measure.
What we do instead
Decode the label, publish how much of it we could read, and check a product against what you already own.

How the two big scores are built

Neither app hides its method, and both are worth reading in the original. The mechanics below are theirs, described from their own documentation.

Yuka: the worst ingredient sets the band

78< 25out of 100
One red ingredient forces the whole product under 25, whatever else is in it.Yuka, on how penalties are calculated

EWG: a thin file is scored as a risk

Long recordlow hazardThin recordlow hazardpenalisedsame evidence of harm: none
The data-gap penalty makes newer molecules score worse for being less studied.EWG, on understanding Skin Deep ratings

Four ways an ingredient score goes wrong

  1. One ingredient decides everything

    Yuka sets a product’s band from its single worst ingredient. One red ingredient forces the score under 25 out of 100 whatever else is in the bottle. It guarantees a wide spread, which is exactly what our own flat rubric lacked, and it means a formula is judged by its least flattering line.

    The rule we took from itWeight by position. Never take the maximum.

  2. Hazard is not risk

    A substance can be hazardous in a beaker and irrelevant in a bottle at the amount actually used. Cosmetic chemists make this point about both apps: the scores flag presence, not exposure. Salicylic acid, titanium dioxide and phenoxyethanol all score badly on apps that ignore how much is there.

    The rule we took from itAn ingredient list gives order, never percentage. Say so.

  3. Following the law costs you points

    European rules require an essential oil to declare its allergens separately by name. Apps that count each declared name as another additive therefore penalise the brands that complied and reward the ones that did not. It is the cleanest example of a metric measuring paperwork rather than product.

    The rule we took from itA declared substance must never cost more than an undeclared one.

  4. Not knowing is treated as bad news

    EWG applies a penalty for missing data, so a newer, well-studied-in-a-different-way molecule scores worse than an old one simply for having a thinner file. The rating ends up measuring how long something has been around.

    The rule we took from itUnknown is not bad. Exclude it from the score and publish the coverage instead.

Test one

We scored safety, and everything came out the same

The obvious rating is a safety score. Take irritation, take pore-clogging, take how well studied the actives are, average them, print stars. We built it and ran it across everything we publish. Fifteen of nineteen products came out at four stars or more, inside a band of about a third of a star.

The reason turned out to be more interesting than the failure. Our own ingredient library says low pore-clogging risk for 89%of the 211 ingredients in it, and low irritation risk for70%. You cannot separate products on a measure where almost everything scores the same.

So we tested the one list that is a matter of law rather than opinion: the twenty-six fragrance allergens European regulations require to be named on a label. Across 107products from the brands we cover, 92 had nothing to declare at all. Standard deviation across the whole set: 0.36, against the0.5 we had written down in advance as the pass mark.

Balm0.00n=12
Cleanser0.00n=7
Exfoliant0.00n=2
Moisturizer0.22n=27
Other0.48n=50
Serum0.00n=3
Sunscreen0.30n=6

Standard deviation by category. Every one of them flat.

We could have fixed the spread in an afternoon by covering the perfumed drugstore products where these numbers vary enormously. That would have meant choosing what to write about in order to make our own rating look clever. The rating was the thing we were prepared to lose.

Test two

We scored value, and it passed everything except the point

The second idea was better. Forget whether a product is good, ask how much formula you get for the money. Price is the one thing our shelf is not pre-selected for: the brands we cover run from about five euros to about thirty-five, and what goes in the bottle at those prices genuinely differs.

It passed on the numbers, comfortably, on both sets we ran. Our own products: standard deviation 1.19. A wider set of 92 labels from the same brands: 1.91. Against a pass mark of 0.5, with both ends of the scale occupied and every category clearing the bar on its own.

Then we read the list.

These two products were the moment it fell apart. Both real, both scored by the rubric we had committed. Move the constant and watch them change places.

0.30the value we committed
  1. #3Cleanser

    CeraVe Hydrating Cleanser

    ~€12

    5.0 / 5

  2. #16Serum

    The Ordinary Retinol 0.5% in Squalane

    ~€8

    2.5 / 5

Wherever you set the constant, you decided the ranking. That is the problem.

They change places at 2.80, which is 9.3 times the value we committed. That distance is not reassuring. It means the ordering we published was a long way from the one a different reasonable person would have chosen.

See all 18 products at this setting
  1. 5.0CeraVe Daily Moisturizing LotionMoisturizer
  2. 5.0CeraVe Foaming Facial CleanserCleanser
  3. 5.0CeraVe Hydrating CleanserCleanser
  4. 5.0CeraVe Moisturizing CreamMoisturizer
  5. 5.0CeraVe PM Facial Moisturizing LotionMoisturizer
  6. 5.0CeraVe SA Smoothing CleanserCleanser
  7. 5.0Cetaphil Gentle Skin CleanserCleanser
  8. 5.0Eucerin Aquaphor Healing OintmentBalm
  9. 5.0The Ordinary AHA 30% + BHA 2% Peeling SolutionExfoliant
  10. 5.0The Ordinary Glycolic Acid 7% Toning SolutionExfoliant
  11. 5.0The Ordinary Vitamin C Suspension 23% + HA Spheres 2%Serum
  12. 4.1The Ordinary Niacinamide 10% + Zinc 1%Serum
  13. 3.5The Ordinary Natural Moisturizing Factors + HAMoisturizer
  14. 3.1Cetaphil Moisturizing CreamMoisturizer
  15. 2.9The Ordinary Hyaluronic Acid 2% + B5Serum
  16. 2.5The Ordinary Retinol 0.5% in SqualaneSerum
  17. 2.3La Roche-Posay Cicaplast Baume B5+Moisturizer
  18. 1.3Bioderma Sébium Gel MoussantCleanser

An eight-euro retinol serum with retinol in third position, below a twelve-euro cleanser. There is no honest sentence that explains it, and explaining every extreme in one sentence was a condition we had set ourselves before running anything.

The cause was in our own method. To compare a cleanser with a cleanser rather than with a serum, the rubric needs a benchmark per category, and we had picked those numbers ourselves. Ours made it very easy for a cleanser to reach full marks and very hard for a serum to.11 of 18 products sat on the ceiling. The score was not reporting a relationship between formula and price. It was reporting a constant we chose.

The wider set failed the same condition for a second reason:14 of its 92 labels scored zero, every one of them because we could not decode a single active ingredient on them. That is a fact about our library, not about the products.

The fix would have taken ten minutes. Adjust the constants until the retinol serum scores well, then publish. Not doing that is the entire point of writing the pass mark down first.

The number nobody publishes about themselves

Every score above presents itself as equally confident whether it read the whole label or half of it. Here is ours. Across 106 labels from the brands we cover, our library currently decodes a median of 43% of the ingredients on a page, and only 7 of them pass the eighty-percent bar we would want before rating anything.

0%25%50%75%100%80% after ~190 more0 added658 added
Median share of a label our library can decode, as ingredients are added most-frequent first. The curve is steep at the start because formulation reuses the same few hundred structural ingredients.

We would rather print that than imply we read everything. It is also the reason the value test could not be rescued: a distribution partly produced by what we cannot read is not a measurement of what we can.

What honest looks like instead

One idea survived the second test and may come back in a narrower form. The value axis separated bad value reliably and good value badly, which suggests a flag rather than a score: something that says a product is expensive for what is in it, and stays quiet otherwise. If it ships, it will ship with its own pass mark written down first.

Questions people ask about rating apps

Is Yuka accurate for skincare?

Yuka is accurate about what it actually measures, which is narrower than most people assume. Its cosmetic score is built from the health risk associated with each ingredient, and the product's band is set by its single worst one: a high-risk ingredient forces the score below 25 out of 100 regardless of the rest of the formula. What it does not include is concentration. Cosmetic chemists' main objection is that an ingredient can be flagged while present in an amount that is legal, tiny and inert. It is a useful screen for whether a formula contains anything on a watchlist. It is not a measure of whether a product is good, and it was never built to be.

Is EWG Skin Deep reliable?

Skin Deep rates ingredient hazard from 1 to 10 on a weight-of-evidence basis, and it is transparent about doing exactly that. The criticism from toxicologists is the hazard-versus-risk distinction: a hazard score describes what a substance could do, without the exposure that decides whether it will. Skin Deep also applies a penalty where data is thin, which means a newer molecule with a shorter research record scores worse than an older one for reasons that have nothing to do with the molecule. Read it as a flag to go and look something up, not as a verdict.

What is the best skincare rating app?

There is no honest single answer, because the apps measure different things and none of them measures product quality. If you want to know whether a formula contains anything on a hazard watchlist, Yuka and EWG Skin Deep both answer that. If you want per-ingredient acne and irritation numbers to interpret yourself, CosDNA publishes them and leaves blanks where there is no data, which is more honest than filling them in. If you want a score personal to your skin rather than a verdict on the product, SkinSort's match score is the model that avoids the trap entirely. And INCIDecoder, the biggest ingredient site there is, deliberately gives no score at all.

Is there a Yuka alternative for skincare?

For ingredient decoding, INCIDecoder is the deepest and gives no score. CosDNA gives per-ingredient numbers rather than a product verdict. SkinSort scores how well a product matches your own skin type and preferences instead of ranking it. We are a fourth kind: we decode the label, tell you what each ingredient does and how strong the evidence is, publish what share of the label we could read, and check a product against what is already on your shelf. What we do not do is put a number on it, and this page is the record of the two attempts that convinced us not to.

Why does Glow Simplified not give products a score?

We tried twice and pre-registered the pass mark both times, before running anything. The first attempt scored safety and came out flat, because almost every product we cover is already fragrance-free and low-irritation: 92 of 107 labels had no declarable fragrance allergen at all. The second scored value, formula against price, and passed every numeric test before failing on a condition we had set ourselves: we could not explain the ordering. An eight-euro retinol serum ranked below a twelve-euro cleanser, and the reason was a constant we had chosen, not anything about the products. Fixing it would have taken ten minutes of retuning, which is the thing the pre-registration existed to stop.

Sources for the mechanics described above, in the original:Yuka on evaluating cosmetics,EWG on Skin Deep ratings,CosDNA's FAQ,SkinSort on its match score. Our own figures come from the two tests described above. They are recomputed from the same records on every build, so nothing on this page is a number somebody typed.

Tests, tooling and this page engineered by Sasu Dan Liviu.