The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Trials

[2026 update] How to read a forest plot, properly, from scratch

Solved
Solved by KLindqvist in post #5
Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

Jump to the accepted answer →

AZ
an.zamoraTL2 Moderator23 Nov 2024#1

How to read a forest plot, properly, from scratch — that is the question, and I have not found it answered plainly anywhere I have looked.

Session topic: SURMOUNT-4 (JAMA, 2024). Please read it before posting; the discussion is much better when everyone has.

The question I would like us to start with is what the trial set out to estimate, rather than what it found. Once that is on the table we can talk about whether the design could have answered it, and only then about the numbers.

Specific things I would like covered: the population and how far it generalises, how discontinuation was handled, whether the comparator was a fair one, and what the absolute rather than relative effect looks like.

I will summarise at the end and the summary will feed the relevant digest page.

0 likes 20mo
VN
v.nascimentoTL2 Moderator24 Nov 2024#2

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

0 likes 20mo
LW
l.wikstromTL2 Moderator25 Nov 2024#3

Picking up post #2: that is the part I would want checked first.

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

8 likes 20mo
FK
f.kimaniTL2 Moderator26 Nov 2024#4

Coming back to post #2, because the follow-up matters more than the original answer.

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

19 likes 20mo
K
KLindqvistTL4 Moderator Solution27 Nov 2024 · edited#5

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

27 likes 20mo
JP
j.petrovTL2 Moderator28 Nov 2024#6
l.wikstrom, post #3: Picking up post #2: that is the part I would want checked first. Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not),… Go to post

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

0 likes in reply to #3 20mo
DO
d.oyelaranTL3Pharmacist28 Nov 2024#7

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

4 likes 20mo
NK
n.krastevTL2 Moderator29 Nov 2024#8

I read post #6 twice before replying, because I had assumed the opposite.

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

13 likes 20mo
KA
k.agyemanTL2 Moderator30 Nov 2024#9

post #8 answers the question as asked. The question underneath it is different.

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

20 likes 20mo
BM
buffer_marginTL3Regular30 Nov 2024#10

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

0 likes 20mo
TV
t.vasquezTL4 Moderator1 Dec 2024#11
l.wikstrom, post #3: Picking up post #2: that is the part I would want checked first. Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not),… Go to post
Staff post. Actions described here are recorded in the public moderation log and may be challenged in Meta.

Coming back to post #9, because the follow-up matters more than the original answer.

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

0 likes in reply to #3 20mo
VS
v.sjobergTL2 Moderator2 Dec 2024#12
an.zamora, post #1: How to read a forest plot, properly, from scratch — that is the question, and I have not found it answered plainly anywhere I have looked. Session topic: SURMOUNT-4 ( JAMA , 2024). Please read it before posting; the discussion is much better when everyone has. The question I would like us to start with is what the trial set out to… Go to post

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

29 likes in reply to #1 20mo
ZO
z.onwukaTL2 Moderator2 Dec 2024 · edited#13

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

14 likes 20mo
MR
m.rasmussenTL2 Moderator3 Dec 2024#14

post #13 answers the question as asked. The question underneath it is different.

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

5 likes 20mo
IT
impurity_tableTL3Analytical chemist4 Dec 2024#15

I read post #13 twice before replying, because I had assumed the opposite.

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

0 likes 20mo
IN
i.norgaardTL2 Moderator4 Dec 2024#16
m.rasmussen, post #14: post #13 answers the question as asked. The question underneath it is different. Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting. Go to post

This follows post #13 rather than contradicting it.

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

22 likes in reply to #14 20mo
CR
compounding_ruthTL4Pharmacist5 Dec 2024 · edited#17

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

9 likes 20mo
NL
n.laurentTL2 Moderator5 Dec 2024#18

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

2 likes 20mo
TA
t.abubakarTL2 Moderator6 Dec 2024#19
k.agyeman, post #9: post #8 answers the question as asked. The question underneath it is different. The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate… Go to post

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

2 likes in reply to #9 20mo
P
PSkarbekTL3Regular6 Dec 2024#20

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

0 likes 20mo
K
KAnderssonTL3Regular7 Dec 2024#21

Picking up post #18: that is the part I would want checked first.

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

0 likes 20mo
ST
s.teixeiraTL2 Moderator8 Dec 2024#22

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

4 likes 20mo
DS
d.szymanskiTL3Wiki editor8 Dec 2024#23

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

18 likes 20mo
EN
e.ndiayeTL2 Moderator9 Dec 2024#24
n.laurent, post #18: Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results. Go to post

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

0 likes in reply to #18 20mo
GT
g.tanakaTL3Regular9 Dec 2024#25

This follows post #22 rather than contradicting it.

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

0 likes 20mo
MM
m.mwangiTL2 Moderator10 Dec 2024#26

I read post #24 twice before replying, because I had assumed the opposite.

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

2 likes 20mo
PR
policy_readerTL2Regular10 Dec 2024#27
v.sjoberg, post #12: For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use. Go to post

I disagree with the reply above, and I think the disagreement is substantive rather than terminological.

The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.

12 likes in reply to #12 20mo
EA
e.adeyemiTL2 Moderator11 Dec 2024 · edited#28
n.laurent, post #18: Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results. Go to post

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

26 likes in reply to #18 20mo
MD
m.dalgaardTL3Regular11 Dec 2024#29

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

0 likes 20mo
AJ
a.jansenTL2 Moderator12 Dec 2024#30

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

0 likes 20mo