The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Trials

Subgroup analyses: pre-specified versus discovered — a second dataset

ID
isotonic_driftTL1Member2 Mar 2026#1

On the subject in the title: Subgroup analyses: pre-specified versus discovered — a second dataset Working notes rather than a conclusion.

Comparing SURMOUNT-1 (N Engl J Med, 2022) with SURPASS-4 (Lancet, 2021) and finding the comparison harder than it looks.

Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.

Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?

27 likes 5mo
AM
a.mwangiTL2 Moderator5 Mar 2026#2

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

29 likes 5mo
CI
c.inglethorpeTL3Regular6 Mar 2026#3

post #2 answers the question as asked. The question underneath it is different.

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

0 likes 5mo
LD
l.dialloTL2 Moderator8 Mar 2026#4
c.inglethorpe, post #3: post #2 answers the question as asked. The question underneath it is different. Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting. Go to post

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

2 likes in reply to #3 5mo
C
CSagredoTL3Regular9 Mar 2026#5
c.inglethorpe, post #3: post #2 answers the question as asked. The question underneath it is different. Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting. Go to post

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

10 likes in reply to #3 5mo
HB
h.bhattacharyaTL2 Moderator10 Mar 2026 · edited#6

I read post #4 twice before replying, because I had assumed the opposite.

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

22 likes 5mo
OA
o.abrahamsenTL3Regular12 Mar 2026#7

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

0 likes 5mo
RN
r.novakTL2 Moderator13 Mar 2026#8

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

1 like 5mo
EK
e.kjeldsenTL2Member14 Mar 2026#9
r.novak, post #8: Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting. Go to post

Picking up post #6: that is the part I would want checked first.

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

28 likes in reply to #8 4mo
MN
m.nwosuTL2 Moderator15 Mar 2026#10

Coming back to post #8, because the follow-up matters more than the original answer.

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

0 likes 4mo
EL
endpoint_lineTL3Regular16 Mar 2026#11

On post #7 — agreed on the reasoning, with one qualification.

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

26 likes 4mo
ZA
z.adeyemiTL2 Moderator17 Mar 2026#12
endpoint_line, post #11: On post #7 — agreed on the reasoning, with one qualification. Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at… Go to post

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

12 likes in reply to #11 4mo
HN
h.nicolaidesTL3Regular18 Mar 2026#13

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

2 likes 4mo
IG
in.guerreroTL2 Moderator19 Mar 2026 · edited#14

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

0 likes 4mo
EF
erratum_fileTL3Regular20 Mar 2026#15
isotonic_drift, post #1: On the subject in the title: Subgroup analyses: pre-specified versus discovered — a second dataset Working notes rather than a conclusion. Comparing SURMOUNT-1 ( N Engl J Med , 2022) with SURPASS-4 ( Lancet , 2021) and finding the comparison harder than it looks. Different populations, different durations, different endpoints defined… Go to post

Worth separating two things that post #11 runs together.

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

0 likes in reply to #1 4mo
AC
a.cabreraTL2 Moderator21 Mar 2026#16
h.nicolaides, post #13: Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive. Go to post

post #15 is right about the mechanism and I think understates the practical bit.

Thank you for the correction. I have edited my earlier post with a note rather than silently, so the thread still makes sense to read. The error was mine and it was the kind that comes from remembering a figure instead of looking it up.

18 likes in reply to #13 4mo
R
RidgewayTL3Regular22 Mar 2026#17

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

4 likes 4mo
IG
i.grimaldiTL2 Moderator23 Mar 2026#18

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

0 likes 4mo
DB
d.bramleyTL3Regular24 Mar 2026#19
i.grimaldi, post #18: Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting. Go to post

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

0 likes in reply to #18 4mo
VB
v.bruunTL2 Moderator25 Mar 2026#20

post #19 answers the question as asked. The question underneath it is different.

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

25 likes 4mo
IB
i.beaulieuTL2 Moderator26 Mar 2026#21
isotonic_drift, post #1: On the subject in the title: Subgroup analyses: pre-specified versus discovered — a second dataset Working notes rather than a conclusion. Comparing SURMOUNT-1 ( N Engl J Med , 2022) with SURPASS-4 ( Lancet , 2021) and finding the comparison harder than it looks. Different populations, different durations, different endpoints defined… Go to post

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

9 likes in reply to #1 4mo
I
IMainwaringTL3Regular27 Mar 2026#22

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

21 likes 4mo
LL
l.lundgrenTL2 Moderator28 Mar 2026#23

This follows post #20 rather than contradicting it.

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

0 likes 4mo
T
TamburelloTL2Member28 Mar 2026#24
Ridgeway, post #17: Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot. Go to post

I read post #22 twice before replying, because I had assumed the opposite.

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

1 like in reply to #17 4mo
NK
n.kaufmannTL229 Mar 2026#25
VS
vial_slopeTL3Regular30 Mar 2026#26

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

15 likes 4mo
MA
m.amankwahTL2 Moderator31 Mar 2026 · edited#27

Picking up post #24: that is the part I would want checked first.

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

30 likes 4mo
N
NLoughranTL3Regular1 Apr 2026#28
l.lundgren, post #23: This follows post #20 rather than contradicting it. Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of… Go to post

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

0 likes in reply to #23 4mo
RM
r.molnarTL2 Moderator2 Apr 2026#29
Tamburello, post #24: I read post #22 twice before replying, because I had assumed the opposite. Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit. Go to post

post #28 is right about the mechanism and I think understates the practical bit.

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

3 likes in reply to #24 4mo
CI
c.inglethorpeTL3Regular3 Apr 2026#30

Worth separating two things that post #26 runs together.

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

10 likes 4mo