The Health Score is GearCheck's single-number summary of your panel. It runs from 0 to 100, and it is deliberately calibrated for an athletic context rather than a general-population one. But the number is only half the information — the tier it falls in tells you what kind of reading it is.
Think of it like a car dashboard: the Health Score is a check-engine system, but one that knows you drive a race car. It does not just say something is off — it says how much of that is expected for what you are doing, and what is worth raising with a doctor or coach.
This guide breaks down the three tiers, what puts you in each one, and — more usefully — why the direction your score is moving matters more than where it currently sits.
What changed in July 2026
The scoring model was rebuilt around findings rather than scraped report text, and the tier system was simplified from five bands to three. Scores dropped across the board by design — the median moved from roughly 85 to 76 — because the old model was scoring the prose it had just written rather than the underlying data. If you have reports from before that date, do not compare their scores directly to newer ones. Compare within each era instead.
Markers sit inside their athletic-adjusted bands, and where they do not, the deviations carry context that explains them. Suppressed LH on exogenous testosterone, a raised hematocrit with normal blood pressure and normal platelets, an AST elevation alongside a high CK in a lifter — these are recognised as expected rather than penalised as damage. A STRONG score means the panel does not show a system under load. It does not mean nothing is happening; it means what is happening has a plausible, non-damaging explanation.
A 30-year-old on TRT at 120 mg/week. Lipids in range, hematocrit 0.51 with a 118/74 blood pressure, liver enzymes normal, total testosterone at the top of the TRT band. Two markers are technically outside standard lab ranges and both are contextualised.
What this is for: Confirming that your current setup is not producing measurable strain, and giving you a stable baseline to compare future panels against.
Several markers are drifting, or one is drifting meaningfully, and the context does not fully account for it. This is the most common tier among people running anything beyond replacement doses, and it is not an alarm. It is the band where the panel starts giving you something specific to talk about: which system is carrying the load, whether the drift is isolated or corroborated by neighbouring markers, and whether it is moving.
A 35-year-old on a 500 mg blast. HDL at 0.82 mmol/L, ApoB mildly raised, ALT 68 U/L with a normal GGT, hematocrit 0.55 with blood pressure at 132/84. The lipid picture has corroboration from two markers; the liver picture does not.
What this is for: Identifying which one or two systems are actually carrying load, so a consultation is about something concrete rather than a whole panel.
Multiple systems are showing load at the same time, or one marker sits far enough outside its band that context cannot explain it away. The model reserves this tier for patterns where markers corroborate each other — a lipid drift confirmed by ApoB and non-HDL together, a kidney signal confirmed by both eGFR and cystatin C, a hematocrit with blood pressure and platelets moving with it. A single dramatic number does not put you here. A pattern does.
A 32-year-old on a high-dose cycle. Hematocrit 0.58 with platelets at 480 and blood pressure at 148/94, HDL 0.46 mmol/L with ApoB clearly raised, ALT 120 U/L with GGT 65 U/L. Three separate systems, each with internal corroboration.
What this is for: Making the case, with specifics, that this panel deserves a conversation with a doctor sooner rather than at the next routine interval.
These tiers are not instructions
The score aggregates per-marker statuses. Those statuses were renamed in 2026 to remove clinical-directive language, so if you have older reports the vocabulary will differ. The current set:
The label and the score now come from the same scale. They used to be computed separately, which let them contradict each other: an eGFR of 67.5 could read Review in the marker table while contributing almost nothing to the score. Wherever a marker has a clinical severity curve, its label is read off that curve, so a harsher word always means a heavier cost.
Inside the athletic-adjusted band. Contributes full points.
Just outside the band. A small, early drift.
Measurably outside the band. Worth following across panels.
Substantially outside the band, or escalated by a contextual rule. A discussion topic.
Outside the band, but a rule explains it — muscle-origin AST, an eGFR artefact, expected androgen effects. Damped in the score rather than removed.
Informative rather than risk-bearing — suppressed LH on exogenous androgens, for instance. Also excluded from scoring.
How Context markers affect your score
A marker classified as Context is damped, not deleted. An eGFR of 72 in a strength athlete with a clean cystatin C is largely a creatinine artefact, not reduced kidney function — scoring it as impairment would be scoring the measurement error. But an explanation is not a guarantee, so the marker keeps a fraction of its weight, and how large a fraction depends on how strong the explanation is: strong evidence leaves 30% of the severity standing, medium 45%, weak 65%. The same eGFR without any supporting context stays in the score at full weight as Monitor or Review. This is where most of the difference between GearCheck and a standard lab readout comes from.
Removing them outright, which is what the model used to do, quietly turned "we have set this evidence aside" into "there is no evidence": a domain whose every flagged marker was explained scored a perfect 100. A Context marker also still counts fully towards your panel coverage. It was measured; we simply understand why it reads the way it does.
Some markers move the score more than others, because the model weights by how irreversible the damage is and how well the domain predicts hard outcomes. Three worth understanding:
Hematocrit
ApoB
Cystatin C
📈A stable 65 is better than a dropping 80. The single most important number is not the score itself — it is the direction the score is moving over time.
— GearCheck Health Score System
A single Health Score is a snapshot. It tells you where you are right now. The real value comes from tracking it across consecutive reports, which tells you where you are going. Here is how to read the movement:
What you are doing is holding. The value of the next panel is confirmation, not correction. Keep the monitoring interval you have.
An early signal. Something shifted — a compound change, a dose change, a diet change, a stretch of poor sleep. Worth identifying before it compounds.
Check the draw conditions first. Rest days, hydration and timing relative to a hard session move several markers at once and can produce a drop that is not real. Then look at what actually changed.
Usually cycle phases. Compare on-cycle scores to previous on-cycle scores rather than to off-cycle baselines. The trend within each phase is the signal.
Score Mechanics
The score starts at 100 and subtracts for each deviation, weighted by severity and by how much the affected domain predicts hard outcomes. Deviations ramp with distance from the band rather than stepping at a threshold, so a marker just outside its range costs far less than one well outside it. Corroboration multiplies: several markers in one organ system moving together cost more than the same number of isolated deviations, because concurrent signals mean more than scattered ones. Contextual markers are damped rather than deleted, and descriptive rows — including readings a body cannot physically produce, which are reporting errors rather than results — stay out of the calculation altogether. The result is that a 70 with one corroborated cardiovascular pattern is a very different report from a 70 with four scattered mild drifts — which is exactly why the tier alone is not the whole story.
Ten domains carry weight, in this order: kidneys and cardiovascular heaviest, then blood, liver and metabolic, then hormones, inflammation, thyroid, iron and vitamins. Your blood pressure is scored too, on a curve that starts at 120/80 and climbs steeply. And a panel that covers too little — fewer than three domains, fewer than eight markers, or too little depth in the domains it did touch — is given no number at all rather than a flattering one.
Two markers are graded on their own ladder rather than by distance from a reference range, because their diagnostic thresholds sit only a few percent above it. HbA1c is the strictest: below 5.5% the metabolic domain scores full marks, and every further 0.1% costs ten points, reaching zero at 6.4%. It caps the domain rather than averaging into it — a clean fasting glucose is one morning, HbA1c is the last three months, and only one of those is hard to have a good day on.
Score Is Not a Diagnosis
Test Consistently
Same point in the cycle, similar conditions, similar rest days. Comparing a mid-blast score to a cruise score tells you less than comparing two mid-blast scores.
Watch Trends, Not Snapshots
One drop might be noise. Three consecutive drops is a signal. Look for the pattern across 2—3 panels before drawing a conclusion.
Read the Findings, Not Just the Number
The score is a summary of the findings underneath it. Those findings say which system is carrying load and what would confirm it — that is the part worth taking to a consultation.
