How every grade is produced
Methodology
This page describes exactly how a grade on this site is produced. It is written so that a reader can disagree with us specifically rather than generally — if you think a grade is wrong, this tells you which input to argue about.
The central rule
We grade claims, not interventions. A compound can have decent evidence that it shifts a blood marker and no evidence at all that it helps anyone live longer or better. Those are separate claims, they get separate grades, and we never average them into a single score.
This is not a stylistic preference. Presenting “NAD+ levels increased” as evidence for “lives longer” is the single most common error in longevity marketing. Our software checks for it: where a claim is about lifespan or healthspan and no linked human study measured that outcome, an automatic notice appears on the page.
What we record about each study
Bibliographic details — title, journal, year, authors, DOI — are retrieved from PubMed and never typed by hand, so a citation on this site cannot drift from its source. On top of that, an editor records the appraisal facts:
- Subject. Human, animal, cell culture or modelling.
- Design. From meta-analysis of randomised trials down to case series and mechanistic work.
- Sample size and duration. How many participants, followed for how long.
- Outcome type. Whether a hard outcome or a surrogate marker was measured.
- Direction. Whether the result supports, contradicts or finds no effect for this specific claim.
- Risk of bias. Assessed against standard domains; recorded as low, some concerns or high.
- Funding and conflicts. Whether an interested party funded or authored the work.
- Retraction status. Retracted work is excluded from grading and shown as excluded.
How a grade is derived
Those recorded facts are run through a fixed rubric that produces a suggested grade. The same inputs always produce the same suggestion — it is code, not impression, and the reasoning is printed on every claim under “why this grade”. The rules that do most of the work:
- 1.No human study, no human grade. If nothing has been tested in people, the best available grade is preclinical only, whatever the animal results show.
- 2.Surrogate markers cannot reach “strong”. If every human study measured a laboratory marker rather than a clinical or functional outcome, the grade is capped below strong however large the trials were.
- 3.Replication is required for the top grade. A single impressive trial is early evidence, not strong evidence.
- 4.Credible disagreement produces “mixed”. Where good trials point both ways, we say so rather than picking the flattering one.
- 5.Absence of evidence is labelled as such. “Insufficient evidence” means nobody has properly looked. It is not a judgement that something does not work — that is a separate grade, evidence against, and it requires good trials finding no effect.
An editor may override the rubric, but the override, the original suggestion and a written rationale are all published on the claim. You can always see where human judgement departed from the rules.
The grades
- Strong human evidence
- Consistent findings from multiple well-conducted human trials, or meta-analysis of randomised controlled trials, with adequate sample sizes and a clinically meaningful endpoint.
- Moderate human evidence
- More than one human trial pointing the same way, but limited by sample size, duration, risk of bias, or reliance on surrogate endpoints.
- Early human evidence
- One or a small number of small, short or preliminary human studies. Directionally interesting, not yet dependable.
- Mixed evidence
- Human studies disagree, with credible trials on both sides, or results reverse under better methodology.
- Preclinical only
- Evidence comes from animals, cell cultures or modelling. No human trial has tested this claim. Animal lifespan results do not establish human benefit.
- Insufficient evidence
- Too little credible research exists to judge the claim either way.
- Evidence against the claim
- Well-conducted human research indicates the claimed effect does not occur, or is too small to matter.
Outcome types
What a study measured determines what it can support. We record this explicitly so a biomarker result is never quietly upgraded into a health claim.
- Lifespan
- Death from any cause was measured.
- Healthspan
- Years lived free of major disease or disability were measured.
- Disease outcome
- A diagnosed condition or clinical event was measured.
- Physical function
- A measured capability such as walking speed, grip strength or VO₂ max.
- Surrogate biomarkerSurrogate
- A laboratory marker measured as a stand-in for health. A change here does not by itself demonstrate a health benefit.
- Safety
- Adverse events, tolerability or harm were measured.
Study designs and their weight
| Design | Weight | Human |
|---|---|---|
| Meta-analysis of RCTs | 10 | Yes |
| Systematic review | 8 | Yes |
| Randomised controlled trial | 7 | Yes |
| Non-randomised trial | 5 | Yes |
| Prospective cohort | 4 | Yes |
| Retrospective cohort | 3 | Yes |
| Case-control | 3 | Yes |
| Cross-sectional | 2 | Yes |
| Case series | 1 | Yes |
| Mechanistic study | 1 | No |
| Animal study | 1 | No |
| Cell / in vitro study | 0 | No |
| Modelling study | 0 | No |
What we will not do
- Publish a numeric score such as “83/100 effective”. The underlying evidence does not support that precision, and it invites false confidence.
- Let commercial availability influence a grade. Whether a product is sold by our owner has no input into the rubric.
- Publish an imported or machine-drafted summary without a named human reviewer approving it.
- Describe a page as medically reviewed unless a named reviewer has genuinely reviewed it.
- Adjust publication dates to appear fresher than we are.
Found something wrong? Our corrections policy explains how we handle it, and every grade change is recorded publicly.