The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Trials

How to read a forest plot, properly, from scratch

Solved
Solved by t.karlsen in post #8
I disagree with the reply above, and I think the disagreement is substantive rather than terminological. The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.

Jump to the accepted answer →

SC
s.cardosoTL2 Moderator24 Jun 2025#1

The question in the title: How to read a forest plot, properly, from scratch I will give what I have already checked below so nobody repeats it.

Session topic: STEP 4 (JAMA, 2021). Please read it before posting; the discussion is much better when everyone has.

The question I would like us to start with is what the trial set out to estimate, rather than what it found. Once that is on the table we can talk about whether the design could have answered it, and only then about the numbers.

Specific things I would like covered: the population and how far it generalises, how discontinuation was handled, whether the comparator was a fair one, and what the absolute rather than relative effect looks like.

I will summarise at the end and the summary will feed the relevant digest page.

0 likes 13mo
PD
p.dialloTL2 Moderator1 Jul 2025#2

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

23 likes 13mo
BJ
b.jankowiakTL3Regular5 Jul 2025#3
p.diallo, post #2: The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions. Go to post

Worth separating two things that the opening post runs together.

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

10 likes in reply to #2 13mo
RM
r.mensahTL2 Moderator9 Jul 2025#4

post #2 is right about the mechanism and I think understates the practical bit.

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

3 likes 13mo
YM
y.mensahTL3Wiki editor13 Jul 2025#5

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

0 likes 13mo
CC
c.castellanosTL2 Moderator17 Jul 2025 · edited#6
s.cardoso, post #1: The question in the title: How to read a forest plot, properly, from scratch I will give what I have already checked below so nobody repeats it. Session topic: STEP 4 ( JAMA , 2021). Please read it before posting; the discussion is much better when everyone has. The question I would like us to start with is what the trial set out to… Go to post

Picking up post #3: that is the part I would want checked first.

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

1 like in reply to #1 12mo
NE
n.ekstromTL2Regular20 Jul 2025#7

On post #3 — agreed on the reasoning, with one qualification.

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

0 likes 12mo
TK
t.karlsenTL2 Moderator Solution23 Jul 2025#8

I disagree with the reply above, and I think the disagreement is substantive rather than terminological.

The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.

22 likes 12mo
CC
c.correiaTL2 Moderator26 Jul 2025#9
p.diallo, post #2: The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions. Go to post

I read post #7 twice before replying, because I had assumed the opposite.

Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing.

The correction was fair and I had been repeating something I had not checked carefully enough.

10 likes in reply to #2 12mo
AA
a.almeidaTL2 Moderator30 Jul 2025#10
c.correia, post #9: I read post #7 twice before replying, because I had assumed the opposite. Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing. The correction was fair and I had been repeating something I had not checked carefully enough. Go to post

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

3 likes in reply to #9 12mo
DO
dr_okonkwoTL4 Moderator2 Aug 2025#11
b.jankowiak, post #3: Worth separating two things that the opening post runs together. Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive. Go to post

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

4 likes in reply to #3 12mo
MP
m.perrinTL2 Moderator5 Aug 2025#12

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

13 likes 12mo
DF
d.fontaineTL2 Moderator7 Aug 2025#13

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

27 likes 12mo
IB
i.bakkenTL2 Moderator10 Aug 2025#14

Worth separating two things that post #10 runs together.

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

0 likes 12mo
VS
v.szaboTL3Analytical chemist13 Aug 2025 · edited#15
p.diallo, post #2: The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions. Go to post

Picking up post #12: that is the part I would want checked first.

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

8 likes in reply to #2 11mo
VK
v.kirchnerTL2 Moderator16 Aug 2025#16

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

19 likes 11mo
SK
s.karlsen_rphTL319 Aug 2025#17
OV
o.vukovicTL2 Moderator21 Aug 2025#18
v.kirchner, post #16: Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results. Go to post

Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing.

The correction was fair and I had been repeating something I had not checked carefully enough.

0 likes in reply to #16 11mo
CL
customs_ledgerTL3Regular24 Aug 2025#19

This follows post #16 rather than contradicting it.

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

13 likes 11mo
PO
p.ostergaardTL2 Moderator27 Aug 2025#20
y.mensah, post #5: Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters. Go to post

I read post #18 twice before replying, because I had assumed the opposite.

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

26 likes in reply to #5 11mo
CH
c.haddadTL2 Moderator29 Aug 2025 · edited#21
r.mensah, post #4: post #2 is right about the mechanism and I think understates the practical bit. Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal… Go to post

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

1 like in reply to #4 11mo
PN
plateau_notesTL2Regular1 Sep 2025#22
p.ostergaard, post #20: I read post #18 twice before replying, because I had assumed the opposite. Two things before anyone answers the substance. First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound. Go to post

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

0 likes in reply to #20 11mo
NV
n.vukovicTL2 Moderator3 Sep 2025#23

Worth separating two things that post #19 runs together.

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

15 likes 11mo
AD
appeals_deskTL3Regular6 Sep 2025#24

post #23 is right about the mechanism and I think understates the practical bit.

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

5 likes 11mo
YR
y.rahimiTL2 Moderator8 Sep 2025 · edited#25
p.diallo, post #2: The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions. Go to post

Coming back to post #23, because the follow-up matters more than the original answer.

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

0 likes in reply to #2 11mo
P
preregisteredTL3Research methods11 Sep 2025#26

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

30 likes 11mo
RP
r.petrovTL2 Moderator13 Sep 2025#27

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

10 likes 10mo
PE
ppm_errorTL3Analytical chemist15 Sep 2025#28

post #27 answers the question as asked. The question underneath it is different.

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

3 likes 10mo
BV
b.vanheckeTL2 Moderator18 Sep 2025#29
dr_okonkwo, post #11: Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis. Go to post

I read post #27 twice before replying, because I had assumed the opposite.

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

5 likes in reply to #11 10mo
C
chromatogramTL4Analytical chemist20 Sep 2025#30

This follows post #27 rather than contradicting it.

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

0 likes 10mo