baselit.app / blog / how skin scores work

[ Explainer · skin scoring ]

How skin scores work, and what they don't

Published 26 July 20269 min readWritten by the Baselit team

Short version: a 0 to 100 skin score is a cosmetic snapshot of how your skin looks in one photograph, graded against the app's own private scale. There is no shared standard behind these numbers, which is why the same face gets a different score in every app and why comparing across apps is pointless. The only property that makes a score useful is repeatability: does the same photo give the same number, so that a change over weeks is your skin rather than your bathroom light. You can test that in any app in about two minutes, and the method is below.

We build a skin-scoring app, so treat everything here as informed and interested at the same time. What follows is the part we would want a friend to know before they paid for any scanner, ours included.

What does a 0 to 100 skin score actually measure?

It measures how your skin looks in one photograph, scored against the app's own reference scale. Concretely, a model looks at visible surface appearance: how even the tone reads, how visible the texture is, how much shine or dullness there is, how clear the surface looks. Then it maps that onto a number. The number is not a physical unit. Nobody agreed that 70 means anything, the way everyone agreed what a gram is. It reads nothing below the surface, it is not a measurement of skin health, and it does not identify or rule out anything medical.

That sounds like a demolition, but it is not. A tape measure around your waist is also not a health measurement, and it is still the most useful thing you can do to track a change over months, precisely because it is repeatable and you already know what it does not tell you. A skin score works the same way. Its value is entirely in the comparison, not in the absolute number.

Why does the same person get a different score in every app?

Because there is no shared standard behind these numbers. Each app defines its own axes, its own anchor points and its own idea of what a 70 looks like, so the scales are not convertible into one another any more than one restaurant's five stars converts into another's. A 62 in one app and an 81 in another can both be internally consistent and still disagree completely. This is the single most common source of confusion in the category, and it is not a sign that one of the apps is broken.

The practical rule that falls out of it: never compare your score across apps, and be suspicious of any app that invites you to. Compare your score in one app across weeks, captured the same way each time. That is the only comparison that carries information.

Why did my score change when my skin did not?

Almost always the photo changed, not the skin. Most of the perceived inconsistency in skin scanners comes from different photos rather than from the model: different light, different distance, different angle, different time of day. Warm bathroom light shifts the colour cast towards red and can read as unevenness. Side light manufactures texture out of shadow. Holding the phone closer changes how much of the frame your face fills, and with it how large every pore looks in pixels. An app that does not correct for these is reading the room as if it were your face.

This is why capture matters more than model choice. A scanner that lets you shoot in any conditions and then reports whatever comes back will produce a number that bounces, and the bouncing will look like your skin misbehaving when it is actually your lamp.

What can an app do about it?

Roughly four things, and it is worth knowing which ones a scanner does before you trust its numbers. This is the pipeline Baselit uses, described plainly rather than as a feature list.

01
Refuse the bad photo

Check distance, angle, brightness, evenness of light and sharpness before the shutter unlocks. A photo that is too dark or lit from one side cannot be rescued afterwards, because blown highlights and crushed shadows contain no information to recover. Spending five extra seconds lining up beats scoring a photo that was never usable.

02
Normalize what got through

Rotate until the eye line is horizontal, scale to a fixed eye distance so the face occupies the same pixels every time, correct the colour cast from warm or cool light, even out the brightness, then mask out hair, eyes, lips and background so only skin is graded. This is the step that turns two photos taken in different rooms into two comparable inputs.

03
Make identical input give identical output

Fingerprint the normalized image and cache the result against it. Feed in the exact same photo and you get the exact same number back, every time, by construction rather than by luck. This closes the most embarrassing failure in the category: the same picture scoring 64 and then 71.

04
Average out the wobble on near-identical photos

A photo taken thirty seconds later is not byte-identical, so caching cannot help. Scoring it several times and taking the median absorbs individual outliers. It reduces the wobble. It does not abolish it, and anyone claiming it does is selling something.

How can you test whether an app's score is consistent?

Here is a two-minute test that works on any scanner, including ours. Take two scans five minutes apart, in the same spot, in the same light, at the same distance, without washing or touching your face in between. Your skin cannot meaningfully change in five minutes. Whatever gap you see between the two numbers is therefore not skin. It is the app's noise floor.

Then use that number as your reading glasses. If two scans five minutes apart differ by eight points, then a weekly change of six points tells you nothing, and every "you improved" message below that threshold in that app is decoration. If the two scans land within a point or two, you can trust smaller movements. Run this before you subscribe to anything, and run it on ours too.

What Baselit does about the same test: the same photo always returns the same number, because the result is cached against the image fingerprint. Two different photos taken minutes apart will still differ slightly, because there is genuine residual variance in any vision model even after normalization. We do not claim clinical reproducibility and we would not believe an app that did. We claim that the number moves when your skin moves, and not when your lighting does.

Why split the score into axes at all?

Because a single number cannot tell you what to change. If your overall score goes from 68 to 71, you have learned that something got better, which is pleasant and useless. If you can see that texture went up four points while evenness stayed flat, you have learned which part of your routine is earning its place. Baselit reads five separate axes, and the routine is built from the ones with the most room in them rather than from the headline number.

The five axes Baselit reads from one selfie
AxisWhat it describesHow honest the camera can be about it
ClarityHow clear the surface looks, including visible blemishes and congested-looking areasGood. This is surface appearance, which is exactly what a camera sees
EvennessHow uniform the tone reads across the face, including darker-looking marks and redder-looking areasGood, with a caveat: colour depends heavily on white balance, which is why the light correction step matters most here
TextureHow visible the surface irregularity is, including pores and roughnessGood, but sensitive to angle and sharpness. Side light invents texture that is not there
HydrationHow dry or comfortable the surface looksWeakest of the five. This is an impression from visible cues, not a moisture reading. See below
RadianceHow luminous or dull the surface looksModerate. Closely tied to lighting, so it is the axis that benefits most from consistent capture conditions

Can a phone camera measure hydration?

Not the way a clinic instrument does, and this is the honest weak point of every selfie scanner on the market. A corneometer reads the electrical properties of the outer skin layer. A phone camera reads red, green and blue values in a photograph. What a selfie can genuinely pick up is the visible impression of dryness: dullness, flaking, how light sits on the surface. That is a real cue and it is worth tracking under a consistent method. It is not a moisture measurement, and an app that presents it as one has picked the exact claim that makes the whole category look like a scam.

How big does a change have to be before it means anything?

Bigger than most apps let on. Every scoring system has a noise floor, and movements inside it are indistinguishable from measurement wobble. Baselit treats changes of about three points or less as noise rather than progress, and says so instead of throwing confetti at them. Two habits are worth more than staring at a single delta: look at the direction across three or more scans rather than the gap between two, and look at which axis moved rather than at the headline number.

There is a related design decision worth naming, because it explains something users notice. In Baselit, a single dip never rewrites your routine. A change in an axis has to hold across three separate scans on three different days before it changes what the app recommends. Skin genuinely fluctuates day to day, and a routine that reshuffles itself every time you sleep badly is worse than useless.

How long before a new routine shows up in the score?

It depends entirely on which axis you are watching, and this is where most people quit too early. Rough, honest ranges for visible cosmetic change:

Judging a routine at week two mostly measures your patience. This is also the honest answer to why a scanner is worth having at all: without a repeatable number, twelve weeks of a routine feels like nothing happened, whether or not it did.

What a skin score is not

Worth stating plainly, because the category earns its scepticism. A skin score is not a diagnosis and does not identify or exclude anything medical. It is not a health metric. It is not comparable between apps, or between you and anyone else, and any leaderboard framing is entertainment. It cannot see below the surface. And a good score is never a reason to ignore something that is changing, itching, bleeding or spreading, which belongs with a doctor and not with a phone.

What it is: a consistent, repeatable way to make a slow cosmetic change visible while it is still too small to notice in the mirror. That is a narrow claim, and it is the one worth paying for.

Frequently asked questions

What does a 0 to 100 skin score actually measure?
It measures how your skin looks in one photograph, scored against the app's own reference scale. It is a cosmetic snapshot of visible surface appearance: how even the tone looks, how visible the texture is, how much shine or dullness there is. It is not a physical unit, it does not read anything below the surface, and it is not a measurement of skin health. Two apps can both be internally consistent and still disagree by twenty points on the same face, because each one invented its own scale.
Why does the same person get a different score in every skin app?
Because there is no shared standard behind these numbers. Each app defines its own axes, its own anchor points and its own idea of what a 70 looks like, so the scales are not convertible into one another any more than one restaurant's five stars converts into another's. Comparing your score across apps tells you nothing. Comparing your score in one app across weeks, taken the same way each time, is the only reading that carries information.
Why did my score change when my skin did not?
Almost always the photo changed, not the skin. Most of the perceived inconsistency in skin scanners comes from different photos rather than from model variance: different light, different distance, different angle, different time of day. Warm bathroom light shifts the colour cast, side light manufactures texture out of shadow, and holding the phone closer changes how much of the frame your face fills. An app that does not correct for these reads the room as if it were your face.
How can I test whether a skin app's score is consistent?
Take two scans five minutes apart, in the same spot, same light, same distance, without touching your face in between. Your skin cannot meaningfully change in five minutes, so any gap between the two numbers is the app's noise floor. If the two scores differ by more than a couple of points, then any weekly change smaller than that gap is unreadable in that app. This test takes two minutes and works on any scanner, including ours.
Can a phone camera measure skin hydration?
Not the way a clinic instrument does. A corneometer reads the electrical properties of the outer skin layer; a phone camera reads red, green and blue values in a photograph. What a selfie can pick up is the visible impression of dryness: dullness, flaking, tightness in how light sits on the surface. That is a useful cue to track over time under a consistent method, and it is not a moisture measurement. Any app claiming otherwise is overselling its camera.
Is a 3 point change in my skin score meaningful?
Usually not on its own. Every scoring system has a noise floor, and small movements inside it are indistinguishable from measurement wobble. Baselit treats changes of about three points or less as noise rather than progress and says so rather than celebrating them. Two useful habits: look at the direction across three or more scans instead of the gap between two, and look at which axis moved rather than at the headline number.
How long before a new routine shows up in a skin score?
It depends entirely on which axis you are watching, and this is where most people give up too early. A hydration-type change tends to be visible within days to about two weeks. Radiance moves in roughly one to four weeks. Surface texture takes about four to eight weeks. Blemish-prone clarity takes four to six weeks before a first change, and up to twelve for the full picture, with a rougher stretch possible in weeks three to six while skin adjusts. Uneven tone is the slowest of the common goals at roughly eight to twelve weeks. Judging a routine at week two mostly measures your patience.
Are skin scores accurate?
Repeatability and accuracy are different things, and skin scores are far better at the first than the second. There is no external ground truth that says your face is objectively a 71, so accuracy is not really the question. Repeatability is: does the same input return the same output, and does the number move only when your skin does. That is the property that makes a score worth tracking, and it is the one worth checking before you subscribe to anything.

Baselit is a cosmetic skin-analysis app, not medical advice. The score describes the visible look of skin in a photo. It is not a diagnosis, not a health measurement, and no substitute for a doctor or dermatologist. Timeline ranges describe when a visible cosmetic change typically becomes noticeable and are not a promise about any individual result.

[ your skin, scored ]

Run the five-minute test
on us.

Two scans, five minutes apart, same light. Then decide whether the number is worth tracking. That is the only demo we would trust either.

Download on the App Store