Methodology

How we grade the evidence.

Every claim in every gem carries one chip. Here’s what each one means in plain English — and why we built the system the way we did.

Most health content treats every study as equal. We don’t.

A 12-person pilot, a single mouse experiment, and a 90,000-person meta-analysis all show up the same way in your inbox: as “studies.” That’s how the noise gets in. Before any claim leaves a Distilled Gems edition, it gets one of four chips — Strong, Moderate, Emerging, or Mechanistic-only. The chip tells you, at a glance, how confident the science actually is. No chip means we won’t say it.

  • Not all studies count the same. A meta-analysis of 50 randomized trials is a different kind of evidence than a 12-person pilot. The grade tells you which.
  • Mechanism matters — but isn’t enough. “Plausible in a petri dish” is not “proven in humans.” We chip those differently on purpose.
  • One chip per claim, every time. No claim ships without one. If we can’t grade it, we don’t say it.

The four tiers.

What each chip means in plain English — with one worked example each.

Strong

Tier 1 — Strong

A clear physiological or statistical mechanism, multiple converging human studies — including at least one meta-analysis or foundational controlled human trial — and expert consensus across credible voices.

Worked example

Morning sunlight within 30 minutes of waking improves sleep quality. Backed by Berson, Dunn & Takao (Science, 2002), multiple human RCTs on phase-response curves, and consensus across Walker, Panda, Czeisler, and Foster.

Moderate

Tier 2 — Moderate

A strong mechanism is established, and at least one high-quality human RCT or large prospective cohort confirms the effect — but no full meta-analysis yet, or some well-designed studies disagree on effect size.

Worked example

Eight thousand steps a day captures most of the longevity benefit of higher counts. Paluch et al., The Lancet Public Health, 2022 — large prospective cohort, replication still in progress.

Emerging

Tier 3 — Emerging

Two or three solid human studies support the effect, no meta-analysis yet, mechanism is plausible but not fully mapped. Worth testing — still being refined.

Worked example

Resonance-frequency breathing at roughly six breaths per minute raises heart rate variability. Several small human studies, biologically plausible, no meta-analysis yet.

Mechanistic only

Tier 4 — Mechanistic only

The biological, physiological, or behavioral mechanism is documented in animal models, lab studies, or small human pilots — but the clinical effect in healthy adults has not been demonstrated at scale. Use with appropriate caveat.

Worked example

Compound X extends lifespan in mice by 30%. Robust mechanism in animals, intriguing — but no evidence yet that it does the same thing in humans living human lives.

How a claim earns its tier.

Three checks. Every claim passes them in this order before it gets a chip.

1

Source-grade check.

What kind of study is this? A meta-analysis of RCTs sits at the top. A single rodent experiment sits at the bottom. Every source is mapped against the upstream evidence hierarchies.

2

Convergence test.

Do other credible voices and other study designs land in the same place? Three independent RCTs plus a mechanism plus expert consensus is the bar for Strong. Disagreement drops the chip a tier.

3

Honesty rule.

If we don’t yet know — we say so. Every gem on a contested protocol carries a “what we don’t know yet” line. Replication uncertainty gets named, not buried.

No chip means no claim. If we can’t grade it, we won’t tell you to do it.

Where these grades come from.

Distilled Gems didn’t invent evidence grading. The medical and research community has been refining it for thirty years. Our four tiers are a plain-English compression of the four most-used systems — translated for a reader who has five minutes on a Sunday morning, not five hours in a journal.

GRADE
Grading of Recommendations Assessment, Development and Evaluation. The framework most clinical guidelines now use. Source of our “strength of recommendation” logic.
Cochrane Reviews
The gold standard for systematic reviews and meta-analyses in medicine. Source of our convergence test.
OCEBM Levels of Evidence
Oxford Centre for Evidence-Based Medicine. Source of our study-type hierarchy weighting.
USPSTF Recommendation Grades
U.S. Preventive Services Task Force. Source of our calibration on lifestyle and screening claims.
Four upstream systems, one reader-facing chip.

The ladder of evidence, from strongest to weakest.

Those four systems lean on the same simple idea: some study designs are more reliable than others. Here are the seven main types in plain English, strongest at the top. Next to each is the chip it usually earns in a gem.

  • 1

    Reviews that combine many studies

    Strong
    Systematic reviews and meta-analyses

    This is a study of studies. Researchers gather every solid study on one question, then pool the results into a single big-picture answer. Combining thousands of people washes out the flukes of any one study. This is the gold standard.

  • 2

    The fair, randomized test

    Strong
    Randomized controlled trials

    People are split into groups by chance. One group gets the thing being tested. The other gets a dummy pill or the usual care. Because chance sets the groups, they start out even, so a difference at the end points to the treatment, not luck. It is the strongest kind of single human study. One small trial on its own usually earns a Moderate chip.

  • 3

    Studies that follow a group forward

    Moderate
    Cohort studies

    Researchers track a large group of people for months or years and record what happens. They watch. They do not assign anything. This is how we learn the long-term effect of a habit we cannot test on purpose, like how you sleep or eat. It is weaker than a trial, because other parts of life can creep in and muddy the result.

  • 4

    Studies that look backward

    Emerging
    Case-control studies

    Start with people who already have a condition, plus a similar group who do not. Then look back to find what was different between them. This is faster and cheaper than following people forward. It is also less reliable, because old memories and records have gaps.

  • 5

    Snapshots and surveys

    Emerging
    Cross-sectional studies

    Measure a group of people once, at a single moment. These are good for spotting a link, like “people who sit more tend to weigh more.” But a snapshot cannot tell you which thing caused the other, or whether a third thing caused both.

  • 6

    A single story, or a handful

    Emerging
    Case reports and case series

    A close write-up of one person, or a few. They are useful for flagging something new or rare. But with no comparison group, you cannot tell if the result came from the treatment or from chance.

  • 7

    Lab and animal studies

    Mechanistic
    Test-tube and animal studies

    Done in petri dishes, in cells, or in animals like mice. They show that something is possible and hint at how it might work. But humans are not mice. A promising result here is a starting point, not proof in people.

This ladder is a guide, not a law. A large, careful study can outrank a small, sloppy one a rung above it. That is why we never lean on a single study. We look at where many studies land together, the convergence test from the steps above.

What we won’t do, no matter how strong the evidence.

A high evidence grade is not a license to give medical advice. These four lines are non-negotiable.

  • No specific medical advice. Educational content only. Every health gem is for thinking with your doctor, not instead of one.
  • No specific financial or tax advice. Principles, behavior, and mindset only. No tax math, no withdrawal-order math, no Roth-conversion strategy.
  • No prescription drug interpretation. GLP-1, statins, sleep aids, HRT, TRT — we lead with natural levers and present pharmaceutical options honestly. We never tell a reader to start or stop a medication.
  • No supplement-stack culture. We cover individual vitamins and supplements when they meet the Strong evidence bar — but never as ten-pill morning routines, never as proprietary blends, and always after the free, food-form, or lifestyle lever has been named first.

Educational content only — never medical or financial advice.

What this means when you read a gem.

When you open a Sunday gem and see a teal Strong chip next to a claim, you can act on it this week with high confidence the science is settled. An amber Moderate chip means the effect is real and worth doing, with a little more replication still to come. A rust Emerging chip means early but promising — worth testing, with the open questions named in plain sight; if it also carries a ↑ Rising tag, the supporting research is recent and building. A graphite Mechanistic-only chip means we’re showing it to you because the mechanism is interesting, not because the human evidence is in. Either way, the chip tells you the truth before the headline does.

The journals do the proving. We do the translating.