baselit.app / blog / how skin scores work
[ Explainer · skin scoring ]
How skin scores work, and what they don't
Short version: a 0 to 100 skin score is a cosmetic snapshot of how your skin looks in one photograph, graded against the app's own private scale. There is no shared standard behind these numbers, which is why the same face gets a different score in every app and why comparing across apps is pointless. The only property that makes a score useful is repeatability: does the same photo give the same number, so that a change over weeks is your skin rather than your bathroom light. You can test that in any app in about two minutes, and the method is below.
We build a skin-scoring app, so treat everything here as informed and interested at the same time. What follows is the part we would want a friend to know before they paid for any scanner, ours included.
What does a 0 to 100 skin score actually measure?
It measures how your skin looks in one photograph, scored against the app's own reference scale. Concretely, a model looks at visible surface appearance: how even the tone reads, how visible the texture is, how much shine or dullness there is, how clear the surface looks. Then it maps that onto a number. The number is not a physical unit. Nobody agreed that 70 means anything, the way everyone agreed what a gram is. It reads nothing below the surface, it is not a measurement of skin health, and it does not identify or rule out anything medical.
That sounds like a demolition, but it is not. A tape measure around your waist is also not a health measurement, and it is still the most useful thing you can do to track a change over months, precisely because it is repeatable and you already know what it does not tell you. A skin score works the same way. Its value is entirely in the comparison, not in the absolute number.
Why does the same person get a different score in every app?
Because there is no shared standard behind these numbers. Each app defines its own axes, its own anchor points and its own idea of what a 70 looks like, so the scales are not convertible into one another any more than one restaurant's five stars converts into another's. A 62 in one app and an 81 in another can both be internally consistent and still disagree completely. This is the single most common source of confusion in the category, and it is not a sign that one of the apps is broken.
The practical rule that falls out of it: never compare your score across apps, and be suspicious of any app that invites you to. Compare your score in one app across weeks, captured the same way each time. That is the only comparison that carries information.
Why did my score change when my skin did not?
Almost always the photo changed, not the skin. Most of the perceived inconsistency in skin scanners comes from different photos rather than from the model: different light, different distance, different angle, different time of day. Warm bathroom light shifts the colour cast towards red and can read as unevenness. Side light manufactures texture out of shadow. Holding the phone closer changes how much of the frame your face fills, and with it how large every pore looks in pixels. An app that does not correct for these is reading the room as if it were your face.
This is why capture matters more than model choice. A scanner that lets you shoot in any conditions and then reports whatever comes back will produce a number that bounces, and the bouncing will look like your skin misbehaving when it is actually your lamp.
What can an app do about it?
Roughly four things, and it is worth knowing which ones a scanner does before you trust its numbers. This is the pipeline Baselit uses, described plainly rather than as a feature list.
Check distance, angle, brightness, evenness of light and sharpness before the shutter unlocks. A photo that is too dark or lit from one side cannot be rescued afterwards, because blown highlights and crushed shadows contain no information to recover. Spending five extra seconds lining up beats scoring a photo that was never usable.
Rotate until the eye line is horizontal, scale to a fixed eye distance so the face occupies the same pixels every time, correct the colour cast from warm or cool light, even out the brightness, then mask out hair, eyes, lips and background so only skin is graded. This is the step that turns two photos taken in different rooms into two comparable inputs.
Fingerprint the normalized image and cache the result against it. Feed in the exact same photo and you get the exact same number back, every time, by construction rather than by luck. This closes the most embarrassing failure in the category: the same picture scoring 64 and then 71.
A photo taken thirty seconds later is not byte-identical, so caching cannot help. Scoring it several times and taking the median absorbs individual outliers. It reduces the wobble. It does not abolish it, and anyone claiming it does is selling something.
How can you test whether an app's score is consistent?
Here is a two-minute test that works on any scanner, including ours. Take two scans five minutes apart, in the same spot, in the same light, at the same distance, without washing or touching your face in between. Your skin cannot meaningfully change in five minutes. Whatever gap you see between the two numbers is therefore not skin. It is the app's noise floor.
Then use that number as your reading glasses. If two scans five minutes apart differ by eight points, then a weekly change of six points tells you nothing, and every "you improved" message below that threshold in that app is decoration. If the two scans land within a point or two, you can trust smaller movements. Run this before you subscribe to anything, and run it on ours too.
What Baselit does about the same test: the same photo always returns the same number, because the result is cached against the image fingerprint. Two different photos taken minutes apart will still differ slightly, because there is genuine residual variance in any vision model even after normalization. We do not claim clinical reproducibility and we would not believe an app that did. We claim that the number moves when your skin moves, and not when your lighting does.
Why split the score into axes at all?
Because a single number cannot tell you what to change. If your overall score goes from 68 to 71, you have learned that something got better, which is pleasant and useless. If you can see that texture went up four points while evenness stayed flat, you have learned which part of your routine is earning its place. Baselit reads five separate axes, and the routine is built from the ones with the most room in them rather than from the headline number.
| Axis | What it describes | How honest the camera can be about it |
|---|---|---|
| Clarity | How clear the surface looks, including visible blemishes and congested-looking areas | Good. This is surface appearance, which is exactly what a camera sees |
| Evenness | How uniform the tone reads across the face, including darker-looking marks and redder-looking areas | Good, with a caveat: colour depends heavily on white balance, which is why the light correction step matters most here |
| Texture | How visible the surface irregularity is, including pores and roughness | Good, but sensitive to angle and sharpness. Side light invents texture that is not there |
| Hydration | How dry or comfortable the surface looks | Weakest of the five. This is an impression from visible cues, not a moisture reading. See below |
| Radiance | How luminous or dull the surface looks | Moderate. Closely tied to lighting, so it is the axis that benefits most from consistent capture conditions |
Can a phone camera measure hydration?
Not the way a clinic instrument does, and this is the honest weak point of every selfie scanner on the market. A corneometer reads the electrical properties of the outer skin layer. A phone camera reads red, green and blue values in a photograph. What a selfie can genuinely pick up is the visible impression of dryness: dullness, flaking, how light sits on the surface. That is a real cue and it is worth tracking under a consistent method. It is not a moisture measurement, and an app that presents it as one has picked the exact claim that makes the whole category look like a scam.
How big does a change have to be before it means anything?
Bigger than most apps let on. Every scoring system has a noise floor, and movements inside it are indistinguishable from measurement wobble. Baselit treats changes of about three points or less as noise rather than progress, and says so instead of throwing confetti at them. Two habits are worth more than staring at a single delta: look at the direction across three or more scans rather than the gap between two, and look at which axis moved rather than at the headline number.
There is a related design decision worth naming, because it explains something users notice. In Baselit, a single dip never rewrites your routine. A change in an axis has to hold across three separate scans on three different days before it changes what the app recommends. Skin genuinely fluctuates day to day, and a routine that reshuffles itself every time you sleep badly is worse than useless.
How long before a new routine shows up in the score?
It depends entirely on which axis you are watching, and this is where most people quit too early. Rough, honest ranges for visible cosmetic change:
- Hydration-type change: days to about two weeks. The fastest thing to see.
- Radiance: roughly one to four weeks.
- Surface texture: about four to eight weeks. Deeper structural change is a matter of months.
- Clarity on blemish-prone skin: four to six weeks for a first change, up to twelve for the full picture, and a rougher stretch is common around weeks three to six while skin adjusts to a new active.
- Uneven tone: the slowest common goal, roughly eight to twelve weeks, and daily sunscreen decides whether it holds.
Judging a routine at week two mostly measures your patience. This is also the honest answer to why a scanner is worth having at all: without a repeatable number, twelve weeks of a routine feels like nothing happened, whether or not it did.
What a skin score is not
Worth stating plainly, because the category earns its scepticism. A skin score is not a diagnosis and does not identify or exclude anything medical. It is not a health metric. It is not comparable between apps, or between you and anyone else, and any leaderboard framing is entertainment. It cannot see below the surface. And a good score is never a reason to ignore something that is changing, itching, bleeding or spreading, which belongs with a doctor and not with a phone.
What it is: a consistent, repeatable way to make a slow cosmetic change visible while it is still too small to notice in the mirror. That is a narrow claim, and it is the one worth paying for.
Frequently asked questions
What does a 0 to 100 skin score actually measure?
Why does the same person get a different score in every skin app?
Why did my score change when my skin did not?
How can I test whether a skin app's score is consistent?
Can a phone camera measure skin hydration?
Is a 3 point change in my skin score meaningful?
How long before a new routine shows up in a skin score?
Are skin scores accurate?
Baselit is a cosmetic skin-analysis app, not medical advice. The score describes the visible look of skin in a photo. It is not a diagnosis, not a health measurement, and no substitute for a doctor or dermatologist. Timeline ranges describe when a visible cosmetic change typically becomes noticeable and are not a promise about any individual result.