Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.
[2026 update] Reading a supplementary appendix and finding the interesting part posts 31–45
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
This follows post #29 rather than contradicting it.
Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.
Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.
Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.
Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.
I disagree with the reply above, and I think the disagreement is substantive rather than terminological.
The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.
On post #33 — agreed on the reasoning, with one qualification.
The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.
Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.
Two things before anyone answers the substance.
First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.
Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.
Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.
Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.
For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.
On post #40 — agreed on the reasoning, with one qualification.
Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.
This topic was referenced in
- Comparators chosen for regulatory reasons rather than clinical onesEvidence › Trials · 2 replies
- Run-in periods and the population they select — the long versionEvidence › Trials · 2 replies
- Follow-up: Open-label extensions: what survives and what does notEvidence › Trials · 26 replies
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Open-label extensions: what survives and what does not
Posting this under the heading it deserves: Open-label extensions: what survives and what does not Everything below is what sits behind that. Comparing STEP 8 ( JAMA , 2022) with PIONEER 6 ( N Engl J Med ,…
|
4 | 22k | 3mo | |
|
[2026 update] Composite endpoints and the component doing the work
On the subject in the title: Composite endpoints and the component doing the work Working notes rather than a conclusion. Comparing SURMOUNT-1 ( N Engl J Med , 2022) with SURPASS-4 ( Lancet , 2021) and…
|
+76 | 87 | 994 | 3mo |
|
Trial registration and comparing the protocol with the paper
Trial registration and comparing the protocol with the paper — setting out what I have, and where I think it stops being reliable. I have seen LEADER ( N Engl J Med , 2016) cited in support of a claim I do…
|
+38 | 44 | 4.8k | 10mo |
|
Follow-up: Open-label extensions: what survives and what does not
Open-label extensions: what survives and what does not — setting out what I have, and where I think it stops being reliable. Comparing SURMOUNT-2 ( Lancet , 2023) with SURMOUNT-1 ( N Engl J Med , 2022) and…
|
+22 | 26 | 397 | 16mo |
|
What a statistical analysis plan adds that the paper does not
What a statistical analysis plan adds that the paper does not — that is the question, and I have not found it answered plainly anywhere I have looked. I have seen SCALE ( N Engl J Med , 2015) cited in support…
|
2 | 6.6k | 10h |
Related topics — sharing the tags estimand, intention to treat, open label
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Follow-up: Journal club: orforglipron phase 2 and the non-peptide question
Posting this under the heading it deserves: Journal club: orforglipron phase 2 and the non-peptide question Everything below is what sits behind that. Session topic: SURMOUNT-4 ( JAMA , 2024). Please read it…
|
+80 | 91 | 10k | 11mo |
|
What TRIUMPH is designed to answer, and why we should not pre-empt it
What TRIUMPH is designed to answer, and why we should not pre-empt it I have a specific reason for asking rather than idle curiosity, and the context is below. I have seen SURPASS-4 ( Lancet , 2021) cited in…
|
+4 | 8 | 142 | 6h |
|
Hepatic effects of glucagon receptor agonism: the mechanistic worry
Posting this under the heading it deserves: Hepatic effects of glucagon receptor agonism: the mechanistic worry Everything below is what sits behind that. Session topic: SURMOUNT-4 ( JAMA , 2024). Please read…
|
+51 | 56 | 700 | 22mo |
|
Journal club: semaglutide in MASH, and surrogate endpoints — does this still hold?
Journal club: semaglutide in MASH, and surrogate endpoints — does this still hold? I have a specific reason for asking rather than idle curiosity, and the context is below. Comparing STEP 1 ( N Engl J Med ,…
|
3 | 28k | 6mo | |
|
Trial registration and comparing the protocol with the paper
Trial registration and comparing the protocol with the paper — setting out what I have, and where I think it stops being reliable. I have seen LEADER ( N Engl J Med , 2016) cited in support of a claim I do…
|
+38 | 44 | 4.8k | 10mo |