The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Trials

Non-inferiority margins: how they are chosen and how they are abused

NT
nl_translatorTL2Translator · NL16 Sep 2025#1

Posting this under the heading it deserves: Non-inferiority margins: how they are chosen and how they are abused Everything below is what sits behind that.

Comparing SURPASS-2 (N Engl J Med, 2021) with LEADER (N Engl J Med, 2016) and finding the comparison harder than it looks.

Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.

Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?

6 likes 10mo
K
KLindqvistTL4 Moderator24 Sep 2025#2
Staff post. Actions described here are recorded in the public moderation log and may be challenged in Meta.

Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing.

The correction was fair and I had been repeating something I had not checked carefully enough.

10 likes 10mo
CT
c.tullochTL2 Moderator29 Sep 2025 · edited#3

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

30 likes 10mo
DO
d.oyelaranTL3Pharmacist4 Oct 2025#4

I read post #2 twice before replying, because I had assumed the opposite.

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

0 likes 10mo
PA
p.amankwahTL2 Moderator9 Oct 2025#5
KLindqvist, post #2: Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing. The correction was fair and I had been repeating something I had not checked carefully enough. Go to post

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

6 likes in reply to #2 10mo
NR
n.rahimiTL2 Moderator13 Oct 2025#6
d.oyelaran, post #4: I read post #2 twice before replying, because I had assumed the opposite. Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than… Go to post

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

15 likes in reply to #4 9mo
LD
l.dziedzicTL2 Moderator17 Oct 2025#7

Picking up post #4: that is the part I would want checked first.

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

0 likes 9mo
NL
n.lehtinenTL2 Moderator21 Oct 2025#8

Coming back to post #6, because the follow-up matters more than the original answer.

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

1 like 9mo
K
KnowltonTL3Regular25 Oct 2025#9

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

9 likes 9mo
EK
e.kimaniTL2 Moderator29 Oct 2025#10

Worth separating two things that post #6 runs together.

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

21 likes 9mo
BS
buffer_sheetTL3Regular1 Nov 2025#11

I read post #9 twice before replying, because I had assumed the opposite.

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

3 likes 9mo
AK
a.kravchenkoTL2 Moderator5 Nov 2025#12
Knowlton, post #9: Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit. Go to post

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

0 likes in reply to #9 9mo
B
BirkelandTL3Regular8 Nov 2025#13
l.dziedzic, post #7: Picking up post #4: that is the part I would want checked first. For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use. Go to post

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

30 likes in reply to #7 9mo
PD
p.dialloTL2 Moderator12 Nov 2025#14

post #13 is right about the mechanism and I think understates the practical bit.

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

15 likes 8mo
BJ
b.jankowiakTL3Regular15 Nov 2025#15

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

6 likes 8mo
RM
r.mensahTL2 Moderator19 Nov 2025#16
p.diallo, post #14: post #13 is right about the mechanism and I think understates the practical bit. Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why… Go to post

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

1 like in reply to #14 8mo
YM
y.mensahTL3Wiki editor22 Nov 2025#17

On post #13 — agreed on the reasoning, with one qualification.

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

0 likes 8mo
CC
c.castellanosTL2 Moderator25 Nov 2025 · edited#18

post #17 answers the question as asked. The question underneath it is different.

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

21 likes 8mo
M
MJayawardenaTL3Regular28 Nov 2025 · edited#19

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

9 likes 8mo
NZ
n.zielinskiTL2 Moderator1 Dec 2025#20
a.kravchenko, post #12: Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot. Go to post

This follows post #17 rather than contradicting it.

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

2 likes in reply to #12 8mo
HE
h.eriksenTL2 Moderator5 Dec 2025#21

This follows post #18 rather than contradicting it.

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

18 likes 8mo
CL
c.lundgrenTL2 Moderator8 Dec 2025#22
e.kimani, post #10: Worth separating two things that post #6 runs together. Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost… Go to post

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

0 likes in reply to #10 8mo
DY
d.yilmazTL2 Moderator11 Dec 2025#23
c.castellanos, post #18: post #17 answers the question as asked. The question underneath it is different. For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use. Go to post

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

1 like in reply to #18 8mo
ZL
z.laurentTL2 Moderator14 Dec 2025 · edited#24

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

7 likes 7mo
CC
c.cardosoTL2 Moderator17 Dec 2025#25

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

25 likes 7mo
CG
c.grimaldiTL2 Moderator20 Dec 2025#26

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

0 likes 7mo
NN
n.nybergTL223 Dec 2025#27
ML
m.lehtinenTL2 Moderator25 Dec 2025#28

On post #24 — agreed on the reasoning, with one qualification.

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

12 likes 7mo
AC
a.coelhoTL2 Moderator28 Dec 2025 · edited#29

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

8 likes 7mo
CW
cohort_watchTL2Member31 Dec 2025#30
n.zielinski, post #20: This follows post #17 rather than contradicting it. For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use. Go to post

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

19 likes in reply to #20 7mo