Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.
Reading a trial's population section before its results posts 91–110
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.
Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.
Picking up post #91: that is the part I would want checked first.
Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.
Worth separating two things that post #91 runs together.
Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.
Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.
Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.
Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.
On post #95 — agreed on the reasoning, with one qualification.
The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.
post #99 answers the question as asked. The question underneath it is different.
Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.
Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.
Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing.
The correction was fair and I had been repeating something I had not checked carefully enough.
post #103 answers the question as asked. The question underneath it is different.
Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.
I read post #103 twice before replying, because I had assumed the opposite.
Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.
Collapsed as off-topic by two members at trust level 3 or above
I disagree with the reply above, and I think the disagreement is substantive rather than terminological.
The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.
Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.
Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.
Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.
Picking up post #107: that is the part I would want checked first.
Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Revisiting: Primary endpoint hierarchies and why order matters
Posting this under the heading it deserves: Revisiting: Primary endpoint hierarchies and why order matters Everything below is what sits behind that. Comparing SURMOUNT-4 ( JAMA , 2024) with SURPASS-2 ( N…
|
+21 | 25 | 46k | 3mo |
|
Comparators chosen for regulatory reasons rather than clinical ones
On the subject in the title: Comparators chosen for regulatory reasons rather than clinical ones Working notes rather than a conclusion. Session topic: SURMOUNT-1 ( N Engl J Med , 2022). Please read it before…
|
2 | 27k | 14mo | |
|
Follow-up: Open-label extensions: what survives and what does not
Open-label extensions: what survives and what does not — setting out what I have, and where I think it stops being reliable. Comparing SURMOUNT-2 ( Lancet , 2023) with SURMOUNT-1 ( N Engl J Med , 2022) and…
|
+22 | 26 | 397 | 16mo |
|
Subgroup analyses: pre-specified versus discovered
Subgroup analyses: pre-specified versus discovered — setting out what I have, and where I think it stops being reliable. Comparing SURMOUNT-2 ( Lancet , 2023) with SURPASS-4 ( Lancet , 2021) and finding the…
|
2 | 27k | 1d | |
|
Reading a supplementary appendix and finding the interesting part
Reading a supplementary appendix and finding the interesting part — setting out what I have, and where I think it stops being reliable. I have seen FLOW ( N Engl J Med , 2024) cited in support of a claim I do…
|
+87 | 92 | 1.1k | 17mo |
Related topics — sharing the tags estimand, surrogate endpoints, discontinuation & dropout
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Journal club: semaglutide in MASH, and surrogate endpoints — does this still hold?
Journal club: semaglutide in MASH, and surrogate endpoints — does this still hold? I have a specific reason for asking rather than idle curiosity, and the context is below. Comparing STEP 1 ( N Engl J Med ,…
|
3 | 28k | 6mo | |
|
Journal club: STEP 8 and the fairness of the comparator dose — a second dataset
Posting this under the heading it deserves: Journal club: STEP 8 and the fairness of the comparator dose — a second dataset Everything below is what sits behind that. Session topic: SURPASS-2 ( N Engl J Med ,…
|
+18 | 22 | 656 | 13h |
|
Non-inferiority margins: how they are chosen and how they are abused — a second dataset
Non-inferiority margins: how they are chosen and how they are abused — a second dataset — setting out what I have, and where I think it stops being reliable. Comparing FLOW ( N Engl J Med , 2024) with…
|
+32 | 37 | 574 | 5mo |
|
Journal club: the CagriSema phase 2 combination paper
On the subject in the title: Journal club: the CagriSema phase 2 combination paper Working notes rather than a conclusion. I have seen SURMOUNT-1 ( N Engl J Med , 2022) cited in support of a claim I do not…
|
2 | 329 | 21h | |
|
Does dual agonism explain the effect size, or is it dose?
Does dual agonism explain the effect size, or is it dose? I have a specific reason for asking rather than idle curiosity, and the context is below. Session topic: SURPASS-4 ( Lancet , 2021). Please read it…
|
+62 | 68 | 2.6k | 29d |