Scored by evidence, not by vibes.
The Physique Score reads a progress photo and rates seven muscle groups on a 0-999 scale, against a published band ladder, with the same standards for what the AI must refuse to score. This page documents the pipeline, exactly where your photo goes — and the bias our own test bench caught before launch.
Written and maintained by Alex Tirim, founder of RepStack. This documents the system that ships in the app — the same rubric, gates, and data path, not a marketing version. Published July 9, 2026.
The answer
Seven muscle groups, one photo, one honest mean.
The Physique Score is a number from 0 to 999 that summarizes how developed your visible muscle groups look in a standardized progress photo. An AI vision model scores seven groups — shoulders, chest, arms, back, abs, quads, calves — against a fixed band ladder, and your overall is the plain mean of the groups it could actually see. A score of 674 means the visible groups average deep into the unmistakably-gym-built range.
It is a training-analysis metric, not an attractiveness rating and not a body-fat estimate. It exists to answer one coaching question — which groups are carrying your physique, and which are lagging — and every judgment it makes is pinned to something visible in the photo you gave it. We publish the method for the same reason we published the Strength Score formula: a score you cannot inspect is a score you cannot argue with.
Developing
0–499
Trained
500–599
Solid
600–699
Advanced
700–799
Elite
800–999
The five physique tiers. Deliberately not the Strength Score tier names — a 500 here and a 500 there measure different things.
The pipeline
From a progress photo to a score, in four steps.
-
1
The photo is standardized on your phone
You line up a front pose against an on-screen silhouette (back and side are optional), and the app crops the shot to a canonical 3:4 frame on the device. Standardizing before scoring matters: the AI should be reading your body, not your camera angle.
-
2
One scoring call reads all poses together
The normalized photos are sent, once, to our scoring service, which passes them to a frontier vision model with the temperature set to 0 — the least-creative setting, chosen for repeatability. The model must answer against a fixed, versioned rubric: score each visible muscle group on a published band ladder, or return null for anything it cannot actually see.
-
3
The response is validated, not trusted
The model returns structured data — group scores, callouts, a confidence value — and our code checks all of it: callouts must come from a fixed claim list, a “strong” verdict requires that group to score at least 450, one callout per muscle group per pose, and a group score for anything the captured poses cannot show is rejected outright. A response that fails validation never reaches your screen.
-
4
The overall score is arithmetic, not opinion
Your overall is the plain mean of the non-null group scores, rounded — computed by our code, never by the model. The scan is then stored on your device with the rubric version stamped on it, and the uploaded copies are discarded. A score you got in July stays comparable to one from October because each records exactly which ruler measured it.
The band ladder, summarized
Every group score must anchor to a band. When the model is torn between two adjacent bands, the rubric tells it to take the lower one.
| Band | What it means for a muscle group |
|---|---|
| 0–149 | No discernible muscular development; the group reads as frame only. |
| 150–299 | Untrained. No distinct muscle shape — a typical untrained adult lands here. |
| 300–399 | First training effect. Slight fullness, no defined shape or separation yet. |
| 400–549 | Trained. Clear shape for the group — a stranger might notice this person lifts. |
| 550–749 | Unmistakably gym-built through advanced: obvious mass, shape, and separation. |
| 750–999 | Standout to exceptional: near-competition and stage-level development. Requires strong evidence. |
Sex calibration
Judged against the best of your own sex.
Competitive physique sport has never used one ruler for everyone: the IFBB publishes separate rulebooks for men's and women's divisions because the standards genuinely differ. Women carry more essential body fat and express elite development through shape, hardness, and separation at far lower absolute mass than men. A rubric that quietly means male mass when it says mass will file a stage-lean woman under mediocre forever.
Ours did exactly that, once. Our first calibration sweep scored a synthetic stage-lean female competitor 478 — mid-scale, for a physique that would place in a show. The rubric now instructs the model to infer the subject's sex from the photo and judge against elite standards for that sex: for men, elite requires genuine mass with separation; for women, hard conditioning, shape, and separation — not male-scale size. Same ladder, same bands, honest anchors. The full incident, with numbers, is in the test-bench section below.
Worked example
Seven group scores, one mean, nothing else.
The athletic test subject pictured in the ladder below returned these group scores in our July 2026 calibration sweep — front pose, studio light:
| Group | Shoulders | Chest | Arms | Back | Abs | Quads | Calves |
|---|---|---|---|---|---|---|---|
| Score | 742 | 704 | 726 | 676 | 612 | 709 | 548 |
Sum the seven, divide by seven: 4,717 ÷ 7 = 673.9, rounded to 674 — Solid, 26 points under the Advanced line. The useful part is not the headline number: it is calves at 548 sitting 194 points under the shoulders. That gap arrives as a tappable “work on” callout pinned to the photo, and the app turns it into a concrete change — which muscle group to add weekly sets for, and which day of your program fits them.
The gates
What the scanner refuses to score.
Before any scoring happens, hard gates run: no visible single adult subject, multiple people, an apparent minor, explicit content, not a physique photo, or a subject too obscured to read. Any of these produces a structured refusal — no score, no stored scan, and the local photo copies are deleted. The app adds its own gates before upload: scanning requires confirming you are an adult photographing yourself, and an account with a known under-18 age cannot start a scan at all.
A second tier asks for a retake instead of guessing: too dark, too far, a pose that hides proportions, clothing covering most groups, or model confidence below 0.55. A retake request is the system saying I will not pretend to see what I cannot — we count that as the score working, not failing.
Your photo's path
Where the photo goes — and where it never does.
The honest version is one sentence: your photo leaves your phone exactly once, for the scoring call, and what comes back and gets stored is numbers and text. We did not build a photo cloud, so there is no photo cloud to breach. The same honesty cuts the other way — if you delete the app, your scans go with it, because we hold no copy to restore.
Stays on your phone
The original photo. The normalized 3:4 crop. The saved scan — scores, callouts, archetype — in a device-local database that is excluded from iCloud sync on purpose.
Leaves once, comes back as numbers
The normalized photos, in a single scoring request. The response is JSON — scores and text. The uploaded copies are processed and discarded: no server-side photo storage, no image bytes in logs, no training on your photos.
Never exists anywhere
A server-side gallery of user photos. Photo backups we hold. Analytics containing your image — scan events record metadata like pose count and score, never pixels, and session replay masks all images by policy.
The scoring service enforces daily per-device and per-network caps as abuse throttles. Those caps count requests; they do not build profiles, and they hold no images. This design is what lets the feature stay inside the low-risk general-wellness boundary: training analysis for healthy adults, no diagnosis, no health measurement, nothing retained to repurpose later.
The test bench
How we test a body scorer without using anyone's body.
You cannot calibrate a physique scorer on a promise, and we were not going to build a test set out of real people's progress photos. So the bench is synthetic: we generate photorealistic test subjects with an image model (Google's Gemini image model, in July 2026) across the full range the scorer must handle — overweight, untrained, average, lean, athletic, and competition-level, men and women, in studio light and in deliberately bad conditions — plus trap cases that must be refused. Every image is scored through the production endpoint, the same code path your photo takes.
The four subjects below are from the current sweep, with the exact scores the system returned. None of them is a real person.
Synthetic subject
207 Developing
Synthetic subject
476 Developing
Synthetic subject
674 Solid
Synthetic subject
809 Elite
AI-generated calibration subjects with their actual returned scores — generated exactly so that no real person's photo has to be in a test set.
What the first sweep caught
The point of a bench is to fail before users can. The first full sweep, July 2026, failed usefully in four places:
| Test case | First sweep | After recalibration | What changed |
|---|---|---|---|
| Stage-lean female competitor | 478 — mid-scale | 809 — Elite | The v1 rubric read “mass” on an implicitly male scale. We rewrote it to judge each physique against elite standards for that sex. |
| Massive male competitor | 573 | 816 | Instructions tuned for caution compressed the top of the scale. Elite development is now allowed to reach the elite bands. |
| Athletic man, harsh overhead light | 594 — beat his studio score by 86 | Discount instructions added | Hard shadows fake definition. The rubric now says to judge stable development, not shadow theater — and the honest answer is that lighting still moves scores, which is why the app nags you to match conditions. |
| Dim, underexposed photo | Hard error — malformed response | “Retake: too dark” | A photo the model cannot read now degrades into a retake request instead of a failure. |
The trap cases, all refused
| What we sent | What came back |
|---|---|
| A fully clothed street snapshot in a winter coat | Refused — not a physique photo |
| A mountain landscape with no person in it | Refused — not a physique photo |
| Two people standing side by side in a gym | Refused — multiple people |
| A person hidden inside an oversized puffer coat | Refused — subject obscured |
Methodology
Both sweeps: 23 cases each — 12 body-development × sex combinations in studio light, 7 bad-condition variants, 4 adversarial traps — every image a single synthetic front pose generated in July 2026 and scored once through the production scoring endpoint. The unit of analysis is one generated image; nothing here is user data, and no real photos were used. Limits: single pose per case, one scoring run per image (the repeat-scan spread bar of ±15 on five runs of the same photo is tested separately, on real photos, before launch claims), and generated bodies differ between sweeps, so cross-sweep deltas are directional. Scores shown are the system's actual output on the stated date, unedited. Corrections: [email protected].
When the number lies
What the score cannot tell you.
It is not a body-fat percentage
Estimating body fat from one photo is guesswork wearing a decimal point, so the score refuses to do it. Nothing in the output is a body-fat number, a weight judgment, or a health measurement — by design, and by the wellness boundary we build inside.
It reads the photo, not the person
A pump, a flex, a flattering angle, or dramatic lighting all move the number even though your body did not change. The rubric instructs the model to discount these; our bench shows the discount is partial, not total. Matching conditions between scans — same time of day, same light, before training — protects the trend more than any prompt we can write.
The bench is synthetic
Our test subjects are generated images, chosen so we can probe extremes — including bodies at both ends of the scale — without putting anyone's real photos in a test set. The trade is realism: a synthetic sweep calibrates the ruler, it does not certify accuracy on your bathroom lighting. Our repeat-scan acceptance bar (same photo, five scores within 15 points) runs against real photos before we call the calibration settled.
Calibration is versioned and still moving
The mid-scale looks strict to us right now — the athletic female test case at 476 reads about one band low. When the rubric changes, the version stamped on every scan changes with it, this page gets a visible update note, and old scans keep the ruler that measured them.
Null means null, not zero
A group the photos cannot show — your back in a front-only scan — returns null and is left out of the mean entirely. A front-only scan is a front-only score. Adding back and side poses widens what the score is allowed to claim.
It cannot verify who is in the photo
The gates refuse multiple people, apparent minors, and photos that are not physique photos, and scanning requires confirming you are an adult photographing yourself. What no gate can prove is that the adult in the frame is you. The score trusts you to point it at your own body.
One more boundary worth stating in plain terms: this is a wellness feature for adults who train, built to the standard that health claims must be truthful and supported. It makes none. If a scan number ever makes you feel worse about training rather than clearer about it, close the feature — the log and the lifts were always the point.
In the app
The score ends where your program begins.
In RepStack, a scan is not a verdict screen — every “work on” callout links to the training bridge: which muscle group is lagging, how many weekly sets your current program gives it, and which program day is the natural place to add volume. The physique reading and your progression engine point at the same next session. Alongside your Strength Score, it also answers a question numbers alone cannot: whether you are built like you lift, lift like you are built, or both.
On iOS. Your first full scan is free; weekly rescans are part of RepStack Pro.
The bottom line
Trust the gaps, not the headline.
Read the Physique Score the way you would read a coach's clipboard note, not a lab result. The headline number moves with lighting and conditioning; the gaps between groups are the durable signal — a 194-point calf deficit does not vanish when the light changes. Scan under matched conditions, let the work-on callouts pick where next month's sets go, and judge the feature by whether your lagging groups stop lagging. That is the test we hold it to as well.
Sources
What this page is based on.
The pipeline, gates, band ladder, and data path above are documented from RepStack's shipped code — product methodology, not research findings. The test-bench numbers are our own system's output, with limits stated in the methodology box. The references below anchor the claims about physique judging standards and the wellness/claims boundary.
- IFBB official rulebooks — men's and women's divisions are judged as separate categories with different criteria
- FDA General Wellness guidance — the low-risk wellness boundary this feature is designed to stay inside
- FTC Health Products Compliance Guidance — the claim standard for anything health-adjacent
Spotted an error? Email [email protected]. Factual errors are corrected within 48 hours, with a visible note on this page.