Quick answer

Numbers are manufactured by design choices, and reading the choices is the whole skill: endpoint, arithmetic, population, duration, analysis, hierarchy, registration. Learn the checklist once and every citation on this site becomes auditable — which is the point — and most GLP-1 marketing becomes transparent — which is the bonus.

Every number this market throws at you was manufactured by choices — endpoint, population, duration, analysis — and the choices decide what the number can honestly claim. This file teaches the reading skill once, so every trial citation on this site (and every marketing distortion off it) becomes checkable in minutes.

Start with the endpoint zoo

GLP-1 trials report three families of results that answer different questions. Continuous weight endpoints — mean percent change — summarize the average experience (STEP 1’s −14.9%). Categorical endpoints — the share losing ≥5/10/15/20% — reveal the spread the mean hides, and they’re where individual expectations should actually be set. Event endpoints — heart attacks in SELECT, kidney events in FLOW, apnea scores in SURMOUNT-OSA — measure health itself rather than its proxies. A claim quoting one family to answer another family’s question (“lose 20%!” …in the categorical tail, not the mean) is the first spin pattern to catch.

Absolute vs placebo-adjusted: the two honest numbers

Every placebo-controlled result has two true versions: the absolute change in the drug arm (−14.9% in STEP 1) and the placebo-subtracted difference (−12.5 points, since placebo lost 2.4%). Both are legitimate; mixing them across comparisons is not — and marketers reliably quote whichever runs larger. The site convention, worth adopting: lead with absolute (it’s what a patient experiences), keep the placebo arm in view (it prices the lifestyle-plus-attention floor every trial provides), and never compare one trial’s absolute against another’s adjusted.

Who was actually in the room

Trial populations are engineered: inclusion criteria (BMI floors, comorbidity mixes), exclusions (recent cardiovascular events, significant GI disease, prior bariatric surgery in many programs), and sometimes run-in periods that filter for tolerance or adherence before randomization. None of this is scandal — it’s how clean questions get asked — but it defines who the result describes. The practical reading move: check the baseline table (mean BMI, age, sex split, diabetes status) against yourself before borrowing the number, and notice that real-world persistence already tells you the trial’s supported, free-drug environment flattered adherence.

Duration decides what a number can mean

Weight-loss curves bend: rapid early loss, then plateau — so a 12-week result is a slope, a 68–72-week result is a destination, and comparing them is comparing a sprinter’s split to a marathoner’s finish. This is why the pivotal programs standardized around 68–72 weeks, why plateaus are physiology rather than failure, and why any ad quoting week-12 numbers is selling the steep part of a curve everyone eventually exits.

The analysis fine print: estimands, briefly

Modern trials pre-specify how dropouts and off-drug stretches count. The treatment-policy approach answers “what happened to everyone randomized, regardless of adherence” — conservative, real-world-flavored; the on-treatment / trial-product approach answers “what did the drug do while actually taken” — larger numbers, cleaner pharmacology. Both are published for the big programs, and the gap between them (often a few points) is exactly the adherence tax. You don’t need the statistics degree; you need the reflex of asking which question a quoted number answered — because marketing always picks the bigger answer without saying so.

Withdrawal designs: the class’s most underrated evidence

STEP 4 and SURMOUNT-4 used the field’s sharpest trick: treat everyone, then randomize responders to continue or switch to placebo. Continuers kept losing; switchers regained most of it — the same verdict twice, in the strongest design available for the question. Reading skill: withdrawal trials answer “what maintains results,” not “what starts them,” and their regain arms are the evidentiary core of every honest maintenance discussion — while their existence exposes any provider’s “easy off-ramp” pitch as unsupported.

Surrogates vs outcomes: the hierarchy

Weight, A1c, and blood pressure are surrogates — excellent, validated, still proxies. Heart attacks, kidney failure, and death are outcomes. The class’s historic turn is that it climbed the hierarchy: SELECT and FLOW measured the real things and won, which is why those rows in the index reorganized insurance in ways no scale number could. Reading rule: surrogate results justify expectations; outcome results justify indications — and claims that blur the tiers (“prevents heart disease!” from a weight-only trial) are reaching past their evidence.

Reading the tolerability tables

Adverse-event tables list everything reported, which makes them look terrifying and mean little raw. The three numbers that actually inform: placebo-subtracted rates (nausea at 44% means less when placebo shows 17%), severity splits (mild-moderate vs severe), and — most decision-relevant — discontinuation-due-to-adverse-events, the single figure that says “how often was this bad enough to quit.” The site’s tables file applies exactly this triage.

Sponsorship, registered endpoints, and honest paranoia

Nearly every trial in this field is industry-funded — that’s who can afford them — so the safeguard isn’t purity, it’s pre-registration and publication: primary endpoints declared in advance (checkable on ClinicalTrials.gov), results in peer-reviewed journals, both arms’ data visible. Calibrated skepticism reads funded-but-registered trials seriously, discounts press-release toplines until papers land (our ≈ convention), and reserves real suspicion for “studies” that were never registered anywhere — the compounded market’s entire evidence genre.

The ten-point checklist

Run any GLP-1 claim through: (1) Which trial, by name? (2) Which endpoint family? (3) Absolute or placebo-adjusted? (4) Mean or categorical tail? (5) What population — and am I in it? (6) How long? (7) Which estimand flavor? (8) Surrogate or outcome? (9) Peer-reviewed paper or topline? (10) Does the index row match the quote? Two minutes; catches essentially everything this market throws.

The bottom line

Numbers are manufactured by design choices, and reading the choices is the whole skill: endpoint, arithmetic, population, duration, analysis, hierarchy, registration. Learn the checklist once and every citation on this site becomes auditable — which is the point — and most GLP-1 marketing becomes transparent — which is the bonus.

Spin patterns, field guide edition

The recurring distortions, named for fast recognition: tail-as-mean (“patients lost up to 26%!” — a categorical tail sold as typical); generation borrowing (semaglutide sellers quoting SURMOUNT numbers, compounders quoting anyone’s); slope-selling (week-8 curves projected forever); arm laundering (quoting the drug arm’s absolute against a competitor’s placebo-adjusted); surrogate inflation (“heart-healthy!” from weight data alone); and estimand shopping (the on-treatment number, uncredited). Each has appeared in this market’s advertising; each dies on contact with one checklist pass — which is why the checklist, not outrage, is this file’s deliverable.

Practice round: three claims, graded

“Clinically proven: 20% weight loss.” Checklist items 1–4 fail at once — no trial named, and 20% is SURMOUNT-1’s top-dose mean, so the claim is honest only for tirzepatide 15 mg, 72 weeks, said plainly. “Our members lose 3× more than diet alone.” Item 9 fails — member data isn’t a registered trial — and “diet alone” smuggles a placebo-arm comparison the program never ran. “Semaglutide reduced cardiovascular events by 20% (SELECT, NEJM 2023).” Passes: named, outcome-tier, published — the sentence every claim in this field should be dressed like, and the standard our rubric holds providers to when they borrow the science for their storefronts.

Where meta-analyses and network comparisons fit

One tier up from single trials sit the syntheses: meta-analyses pooling similar trials for precision, and network meta-analyses that compare drugs never trialed head-to-head by routing through shared placebo arms. Useful — and bounded: pooling inherits every input trial’s population and duration choices (the garbage-in principle), and network comparisons rest on a similarity assumption that this field’s varied designs strain. Read them as sophisticated context that can rank probabilistically, never as substitutes for the same-trial rows — which is why the index marks head-to-heads specially, and why our own cross-trial ladder ships wrapped in caveats.

Teaching it forward

The checklist compounds when shared: it’s the answer to a relative’s viral screenshot, the pre-read for any appointment where a study will be discussed, and the two-minute audit for every provider claim this site’s rubric scores. Internally it’s already load-bearing — the ≈ convention for toplines, the methodology corners in comparison files, the index pairing — so consider this page the site’s epistemics, published: the same test we invite readers to run on us.

Final calibration: this skill isn’t cynicism — the field’s core trials are genuinely excellent science, which is exactly why learning to read them properly pays. The checklist exists to separate that excellence from its imitators, and after a few practice rounds the separation takes minutes: named trial, right endpoint, honest arithmetic, matching row — pass; anything else — pending. Reading well is the one intervention in this entire library with zero side effects and permanent duration.

Sources

Design and analysis details per the primary publications indexed at the trial index; estimand framing per the programs’ statistical analysis plans; registration verification via ClinicalTrials.gov. Links at sources.