Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.
Second pass at: Absence of evidence and evidence of absence posts 31–56
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.
Worth separating two things that post #29 runs together.
Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing.
The correction was fair and I had been repeating something I had not checked carefully enough.
post #33 is right about the mechanism and I think understates the practical bit.
Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.
Coming back to post #33, because the follow-up matters more than the original answer.
Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.
Collapsed as off-topic by two members at trust level 3 or above
Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.
post #37 answers the question as asked. The question underneath it is different.
Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.
I read post #37 twice before replying, because I had assumed the opposite.
Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.
This follows post #37 rather than contradicting it.
Thank you for the correction. I have edited my earlier post with a note rather than silently, so the thread still makes sense to read. The error was mine and it was the kind that comes from remembering a figure instead of looking it up.
Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.
Coming back to post #40, because the follow-up matters more than the original answer.
Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing.
The correction was fair and I had been repeating something I had not checked carefully enough.
post #42 answers the question as asked. The question underneath it is different.
Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.
Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.
This follows post #42 rather than contradicting it.
P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".
I read post #44 twice before replying, because I had assumed the opposite.
Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.
I disagree with the reply above, and I think the disagreement is substantive rather than terminological.
The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.
Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.
Picking up post #46: that is the part I would want checked first.
Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.
Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.
On post #47 — agreed on the reasoning, with one qualification.
Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.
Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.
Worth separating two things that post #51 runs together.
Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.
post #55 is right about the mechanism and I think understates the practical bit.
Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Measurement error in home scales, with a worked standard deviation — a second dataset
Measurement error in home scales, with a worked standard deviation — a second dataset — setting out what I have, and where I think it stops being reliable. Comparing FLOW ( N Engl J Med , 2024) with SURPASS-2…
|
2 | 13k | 7mo | |
|
Regression to the mean in progress reports
On the subject in the title: Regression to the mean in progress reports Working notes rather than a conclusion. Session topic: PIONEER 6 ( N Engl J Med , 2019). Please read it before posting; the discussion…
|
+30 | 34 | 49k | 12mo |
|
Sample size intuition for a personal experiment
On the subject in the title: Sample size intuition for a personal experiment Working notes rather than a conclusion. Session topic: STEP 2 ( Lancet , 2021). Please read it before posting; the discussion is…
|
4 | 62k | 18mo | |
|
Measurement error in home scales, with a worked standard deviation — the long version
On the subject in the title: Measurement error in home scales, with a worked standard deviation — the long version Working notes rather than a conclusion. Session topic: SURMOUNT-1 ( N Engl J Med , 2022).…
|
+96 | 105 | 2k | 2mo |
|
What a confidence interval means, from scratch — the long version
The question in the title: What a confidence interval means, from scratch — the long version I will give what I have already checked below so nobody repeats it. I have seen SURMOUNT-2 ( Lancet , 2023) cited…
|
+16 | 20 | 2.3k | 8mo |
Related topics — sharing the tags worked example, confounding, number needed to treat
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Why "add 2 mL" is not an instruction — a second dataset
Why "add 2 mL" is not an instruction — a second dataset — that is the question, and I have not found it answered plainly anywhere I have looked. Reporting something rather than asking about it, in case the…
|
2 | 13k | 11mo | |
|
Second pass at: Multiplicity when you track fifteen variables
Posting this under the heading it deserves: Second pass at: Multiplicity when you track fifteen variables Everything below is what sits behind that. I have seen SCALE ( N Engl J Med , 2015) cited in support…
|
2 | 12k | 18mo | |
|
Journal club: STEP 3 and the intensive behavioural therapy floor — does this still hold?
The question in the title: Journal club: STEP 3 and the intensive behavioural therapy floor — does this still hold? I will give what I have already checked below so nobody repeats it. Session topic: STEP 4 (…
|
2 | 15k | 10mo | |
|
How much context is too much context in a first post?
How much context is too much context in a first post? I have a specific reason for asking rather than idle curiosity, and the context is below. Question in the title. Context below, and I have tried to…
|
+70 | 74 | 935 | 18d |
|
Regression to the mean in progress reports
On the subject in the title: Regression to the mean in progress reports Working notes rather than a conclusion. Session topic: PIONEER 6 ( N Engl J Med , 2019). Please read it before posting; the discussion…
|
+30 | 34 | 49k | 12mo |