The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Research Methods · Statistics

Sample size intuition for a personal experiment — does this still hold?

Solved
Solved by c.ostergaard in post #7
Two things before anyone answers the substance. First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

Jump to the accepted answer →

AW
a.westergaardTL3Regular3 Feb 2026#1

Sample size intuition for a personal experiment — does this still hold? I have a specific reason for asking rather than idle curiosity, and the context is below.

Comparing STEP 1 (N Engl J Med, 2021) with STEP 2 (Lancet, 2021) and finding the comparison harder than it looks.

Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.

Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?

15 likes 6mo
MS
m.strand_rphTL3Pharmacist12 Feb 2026#2

the opening post answers the question as asked. The question underneath it is different.

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

2 likes 5mo
NV
n.villalobosTL2 Moderator18 Feb 2026#3

Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.

0 likes 5mo
BI
blank_injectionTL2Analytical chemist24 Feb 2026#4

Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.

19 likes 5mo
FA
f.amankwahTL2 Moderator1 Mar 2026#5

Worth separating two things that the opening post runs together.

Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.

12 likes 5mo
EP
e.piresTL26 Mar 2026#6
CO
c.ostergaardTL2 Moderator Solution11 Mar 2026 · edited#7
m.strand_rph, post #2: the opening post answers the question as asked. The question underneath it is different. Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude. Go to post

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

7 likes in reply to #2 5mo
HS
hana.satoTL4 Moderator15 Mar 2026#8
Staff post. Actions described here are recorded in the public moderation log and may be challenged in Meta.

This follows post #5 rather than contradicting it.

Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.

26 likes 4mo
ZC
z.cardosoTL2 Moderator20 Mar 2026#9

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

2 likes 4mo
DM
d.moreauTL2Regular24 Mar 2026 · edited#10

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

0 likes 4mo
LV
l.vermeulenTL2 Moderator28 Mar 2026#11

post #10 is right about the mechanism and I think understates the practical bit.

Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.

7 likes 4mo
CE
crossover_entryTL3Regular2 Apr 2026#12
m.strand_rph, post #2: the opening post answers the question as asked. The question underneath it is different. Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude. Go to post

Worth separating two things that post #8 runs together.

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

18 likes in reply to #2 4mo
PL
p.lindqvistTL2 Moderator6 Apr 2026#13

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

0 likes 4mo
EK
e.kjeldsenTL2Member10 Apr 2026#14

Having read the exchange above, I think I was wrong earlier in this topic and I want to say so plainly rather than quietly editing.

The correction was fair and I had been repeating something I had not checked carefully enough.

1 like 4mo
KO
k.ogunleyeTL2 Moderator13 Apr 2026 · edited#15
crossover_entry, post #12: Worth separating two things that post #8 runs together. Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p Go to post

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

11 likes in reply to #12 3mo
AK
a.kwiatkowskiTL2Member17 Apr 2026#16
m.strand_rph, post #2: the opening post answers the question as asked. The question underneath it is different. Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude. Go to post

On post #12 — agreed on the reasoning, with one qualification.

Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.

24 likes in reply to #2 3mo
TV
t.vargaTL2 Moderator21 Apr 2026#17

Thank you for the correction. I have edited my earlier post with a note rather than silently, so the thread still makes sense to read. The error was mine and it was the kind that comes from remembering a figure instead of looking it up.

0 likes 3mo
N
NardoneTL2Member25 Apr 2026#18

Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.

3 likes 3mo
EB
e.bakkenTL2 Moderator29 Apr 2026#19

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

17 likes 3mo
MM
methods_marginTL3Regular2 May 2026#20

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

33 likes 3mo
T
TamburelloTL2Member6 May 2026#21

On post #17 — agreed on the reasoning, with one qualification.

Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.

1 like 3mo
LL
l.lundgrenTL2 Moderator9 May 2026#22

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

0 likes 3mo
F
FFaulknerTL3Regular13 May 2026#23
f.amankwah, post #5: Worth separating two things that the opening post runs together. Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed. Go to post

Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.

24 likes in reply to #5 3mo
VO
v.okonkwoTL2 Moderator16 May 2026 · edited#24
FFaulkner, post #23: Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction. Go to post

Thank you for the correction. I have edited my earlier post with a note rather than silently, so the thread still makes sense to read. The error was mine and it was the kind that comes from remembering a figure instead of looking it up.

11 likes in reply to #23 2mo
N
NLoughranTL3Regular20 May 2026#25

Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.

3 likes 2mo
MA
m.amankwahTL2 Moderator23 May 2026#26

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

0 likes 2mo
I
IMainwaringTL3Regular27 May 2026#27

Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.

32 likes 2mo
IB
i.beaulieuTL2 Moderator30 May 2026#28
n.villalobos, post #3: Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters. Go to post

This follows post #25 rather than contradicting it.

Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.

17 likes in reply to #3 2mo
T
ThibodeauTL3Regular3 Jun 2026#29
a.kwiatkowski, post #16: On post #12 — agreed on the reasoning, with one qualification. Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed. Go to post

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

0 likes in reply to #16 2mo
FC
f.chowdhuryTL2 Moderator6 Jun 2026#30

Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.

25 likes 2mo