We simulated 10,000 studies where there was nothing to find — two groups drawn from the same population. Analyzed honestly, 5.2% produced a false positive, right at the promised 5%. Given four ordinary-looking analysis choices — a second outcome measure, adding data when results were almost there, trying outlier exclusions, checking subgroups — 68.1% found a "significant" effect. No fraud, no fabricated data. Build your own p-hack below and watch it happen.
Every simulated study compares two groups drawn from the same population— there is nothing to find. Pick the "reasonable" analysis flexibilities you would allow yourself and see how often nothing becomes a significant finding.
Scenario:two groups of 20 (growing to 50 if you add data), outcome measured on a continuous scale, Welch's t-test at α = 0.05. 4of 4 flexibilities selected — a result counts as a "finding" if any allowed analysis path reaches p < 0.05.
Every simulated study compared two groups of 20 subjects (growing to 50 per group when data was added) drawn from the same normal population, analyzed with Welch's t-test. A strategy "finds an effect" if any analysis path it allows reaches significance.
| Analysis strategy | False-positive rate (α = 0.05) | False-positive rate (α = 0.01) |
|---|---|---|
| Honest analysis (prespecified: one outcome, everyone, n = 20/group) | 5.2% | 1.0% |
| Report either of 2 correlated outcomes (ρ = 0.5) | 9.3% | 1.9% |
| Add data if not significant (test at n = 20/30/40/50 per group) | 11.9% | 2.7% |
| Try outlier exclusions (none / |z| > 2.5 / |z| > 2, keep best) | 10.9% | 3.3% |
| Test subgroups (everyone / group X / group Y) | 12.4% | 2.8% |
| All four combined | 68.1% | 31.4% |
10,000 seeded, reproducible simulated studies per figure. This is the same exercise Simmons, Nelson & Simonsohn ran in their 2011 paper "False-Positive Psychology", where four combined flexibilities produced a 60.7% false-positive rate — our four choices land in the same territory.
A 5% significance level promises a 5% false-positive rate for one prespecified analysis. Every flexibility below multiplies the number of analyses hiding inside "the" analysis — and only the best one gets reported.
Individually each trick adds a few points. Together they compound, because a study only needs one of its many analysis paths to cross the threshold.
Nobody in these simulations fabricated data. Each choice — "this outcome is more sensitive", "let's collect a bit more data", "that participant clearly wasn't paying attention", "the effect is probably stronger in experienced users" — is defensible in isolation. That is exactly the danger: p-hacking is what happens when honest people make data-dependent choices and only notice the ones that worked. Gelman and Loken call this the garden of forking paths — the flexibility does the damage even when no single choice feels like cheating.
The practical consequence: a literature can fill up with "p < .05" findings of effects that do not exist, each one produced by a researcher who never intended to deceive anyone.