Most skin-analysis APIs are black boxes — they never tell you how accurate they are, or whether they work on every skin tone. We do. Here is exactly how DermIQ performs, measured on 219 held-out faces, with nothing hidden.
Skin AI's best-documented failure is bias — most models degrade on darker skin. We measured our error across the full Fitzpatrick scale. It varies by less than half a point. This is the number our competitors don't publish.
Lower error on darker skin (Type V: 5.51) than the overall average — the opposite of the industry's usual bias.
We don't pretend all 23 metrics are equally strong. Each is graded by how well it tracks the reference on held-out faces. 17 are high-confidence or good. The directional ones are labelled as such, and one is still in development and never shown in production. That honesty is the point.
What we measured. Agreement with a leading commercial skin-analysis engine across 219 held-out faces, using 5-fold cross-validation. Numbers describe agreement with that benchmark, not a clinical ground truth — we state that plainly rather than implying more.
MAE is the average absolute difference on a 0–100 scale (lower = closer). ρ is rank correlation (how well the metric orders faces from best to worst).
Percentile scoring. In the product, every score also comes with a percentile versus a reference population, so an "88" reads as "better than 62% of people" — informative, not a compressed number that means little.
Fairness. Skin tone is estimated per face via ITA and grouped by Fitzpatrick type. Type VI (dark) has the fewest samples (n=16), and we're actively expanding it — disclosed here rather than buried.
10 free analyses. Transparent pricing. The only skin API that shows its work.