Accuracy & Fairness, published

Most skin-analysis APIs are black boxes — they never tell you how accurate they are, or whether they work on every skin tone. We do. Here is exactly how DermIQ performs, measured on 219 held-out faces, with nothing hidden.

Consistent across every skin tone

Skin AI's best-documented failure is bias — most models degrade on darker skin. We measured our error across the full Fitzpatrick scale. It varies by less than half a point. This is the number our competitors don't publish.

5.59
Type I–II
Light
n=32
5.52
Type III
Light-medium
n=27
5.62
Type IV
Medium
n=52
5.51
Type V
Medium-dark
n=92
5.96
Type VI
Dark
n=16
5.58
Mean error (0–100 scale)
0.45 pts
Max gap across all skin tones
I–VI
Fitzpatrick range validated

Lower error on darker skin (Type V: 5.51) than the overall average — the opposite of the industry's usual bias.

Every metric, graded honestly

We don't pretend all 23 metrics are equally strong. Each is graded by how well it tracks the reference on held-out faces. 17 are high-confidence or good. The directional ones are labelled as such, and one is still in development and never shown in production. That honesty is the point.

MetricCorrelation (ρ)Confidence
  • Nasolabial folds0.82High confidence
  • Overall wrinkles0.81High confidence
  • Crow's feet0.77High confidence
  • Radiance0.75High confidence
  • Firmness0.75High confidence
  • Moisture0.72High confidence
  • Eye bags0.72High confidence
  • Pigmentation0.68Good
  • Periocular wrinkles0.68Good
  • Lower eyelid0.65Good
  • Pores (overall)0.63Good
  • Redness0.61Good
  • Oiliness0.61Good
  • Upper eyelid0.61Good
  • Forehead wrinkles0.59Good
  • Dark circles0.57Good
  • Pores (forehead)0.57Good
  • Pores (cheek)0.54Directional
  • Pores (nose)0.53Directional
  • Texture0.53Directional
  • Glabellar lines0.52Directional
  • Acne0.47Directional
  • Marionette lines0.41In development — not surfaced

How to read this

What we measured. Agreement with a leading commercial skin-analysis engine across 219 held-out faces, using 5-fold cross-validation. Numbers describe agreement with that benchmark, not a clinical ground truth — we state that plainly rather than implying more.

MAE is the average absolute difference on a 0–100 scale (lower = closer). ρ is rank correlation (how well the metric orders faces from best to worst).

Percentile scoring. In the product, every score also comes with a percentile versus a reference population, so an "88" reads as "better than 62% of people" — informative, not a compressed number that means little.

Fairness. Skin tone is estimated per face via ITA and grouped by Fitzpatrick type. Type VI (dark) has the fewest samples (n=16), and we're actively expanding it — disclosed here rather than buried.

Build on numbers you can verify

10 free analyses. Transparent pricing. The only skin API that shows its work.