The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Research Methods · Statistics

Second pass at: Regression to the mean in progress reports

SG
s.grimaldiTL2 Moderator12 Dec 2024#1

Second pass at: Regression to the mean in progress reports Writing it up because I had to work it out twice and would rather nobody else did.

Comparing SUSTAIN 6 (N Engl J Med, 2016) with STEP 4 (JAMA, 2021) and finding the comparison harder than it looks.

Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.

Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?

1 like 20mo
CK
c.kuuselaTL2 Moderator15 Dec 2024#2

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

0 likes 19mo
I
IbrahimoviTL2Member17 Dec 2024#3

Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.

14 likes 19mo
AM
a.molnarTL2 Moderator18 Dec 2024#4

This follows post #3 rather than contradicting it.

Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.

5 likes 19mo
DM
d.magalhesTL2Member20 Dec 2024 · edited#5
c.kuusela, post #2: For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use. Go to post

On the opening post — agreed on the reasoning, with one qualification.

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

0 likes in reply to #2 19mo
AW
am.wikstromTL2 Moderator22 Dec 2024#6

post #5 answers the question as asked. The question underneath it is different.

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

29 likes 19mo
M
MSaarinenTL3Regular23 Dec 2024#7

Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.

9 likes 19mo
BB
b.brandtTL2 Moderator25 Dec 2024#8
d.magalhes, post #5: On the opening post — agreed on the reasoning, with one qualification. Two things before anyone answers the substance. First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound. Go to post

Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.

2 likes in reply to #5 19mo
JD
j.delacroixTL3Regular26 Dec 2024#9

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

5 likes 19mo
RM
ra.mensaTL228 Dec 2024#10
ZA
z.adeyemiTL2 Moderator29 Dec 2024#11

post #10 answers the question as asked. The question underneath it is different.

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

11 likes 19mo
CN
cohort_notesTL2Member30 Dec 2024 · edited#12

On post #8 — agreed on the reasoning, with one qualification.

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

23 likes 19mo
SK
s.kravchenkoTL2 Moderator1 Jan 2025#13
am.wikstrom, post #6: post #5 answers the question as asked. The question underneath it is different. Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude. Go to post

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

0 likes in reply to #6 19mo
GF
gradient_fileTL2Member2 Jan 2025#14

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

3 likes 19mo
FE
f.espinozaTL2 Moderator3 Jan 2025#15

post #14 is right about the mechanism and I think understates the practical bit.

I disagree with the reply above, and I think the disagreement is substantive rather than terminological.

The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.

6 likes 19mo
W
WoodhouseTL2Member4 Jan 2025#16

Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.

17 likes 19mo
ZV
z.vogelTL2 Moderator6 Jan 2025#17
ra.mensa, post #10: post #9 is right about the mechanism and I think understates the practical bit. Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics. Go to post

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

0 likes in reply to #10 19mo
M
MakinenTL2Member7 Jan 2025#18
f.espinoza, post #15: post #14 is right about the mechanism and I think understates the practical bit. I disagree with the reply above, and I think the disagreement is substantive rather than terminological. The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this,… Go to post

I read post #16 twice before replying, because I had assumed the opposite.

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

1 like in reply to #15 19mo
ON
o.nybergTL2 Moderator8 Jan 2025#19

Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.

22 likes 19mo
DI
diluent_indexTL1Member9 Jan 2025#20
ra.mensa, post #10: post #9 is right about the mechanism and I think understates the practical bit. Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics. Go to post

Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.

0 likes in reply to #10 19mo
CC
c.castellanosTL2 Moderator10 Jan 2025#21

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

3 likes 19mo
NE
n.ekstromTL2Regular11 Jan 2025#22
c.castellanos, post #21: Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics. Go to post

post #21 is right about the mechanism and I think understates the practical bit.

I disagree with the reply above, and I think the disagreement is substantive rather than terminological.

The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.

0 likes in reply to #21 18mo
RM
r.mensahTL2 Moderator13 Jan 2025#23

I read post #21 twice before replying, because I had assumed the opposite.

Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.

31 likes 18mo
YM
y.mensahTL3Wiki editor14 Jan 2025#24

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

16 likes 18mo
PD
p.dialloTL2 Moderator15 Jan 2025#25

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

1 like 18mo
BJ
b.jankowiakTL3Regular16 Jan 2025#26
n.ekstrom, post #22: post #21 is right about the mechanism and I think understates the practical bit. I disagree with the reply above, and I think the disagreement is substantive rather than terminological. The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this,… Go to post

Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.

0 likes in reply to #22 18mo
AK
a.kravchenkoTL2 Moderator17 Jan 2025#27

Coming back to post #25, because the follow-up matters more than the original answer.

Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.

23 likes 18mo
B
BirkelandTL3Regular18 Jan 2025 · edited#28

Picking up post #25: that is the part I would want checked first.

Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.

11 likes 18mo
PM
p.mwangiTL2 Moderator19 Jan 2025 · edited#29
z.adeyemi, post #11: post #10 answers the question as asked. The question underneath it is different. Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p Go to post

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

0 likes in reply to #11 18mo
NA
n.abernathyTL3Analytical chemist20 Jan 2025#30
d.magalhes, post #5: On the opening post — agreed on the reasoning, with one qualification. Two things before anyone answers the substance. First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound. Go to post

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

32 likes in reply to #5 18mo