A Short Note on P-Value Hacking by Nassim Nicholas Taleb
The paper works out the exact probability distribution of p-values across repeated trials of the same phenomenon, and shows that an observed p-value of .02 can correspond to a "true" p-value above .1.
P-value hacking is modeled as picking the minimum p-value among m independent tests, which can land considerably lower than the true p-value even in a single trial, owing to the extreme skewness of the underlying meta-distribution. Taleb derives this metadistribution exactly, for small samples (2 < n ≤ 30) and for the large-sample limit, and shows p-values stay highly skewed and volatile regardless of sample size: about 75% of realizations of a true p-value of .05 read below .05, and 60% of a true p-value of .12 read below .05 as well.
Topics include:
• Metadistribution of p-values — exact PDF for small samples and the large-n limit
• P-value hacking — distribution of the minimum p-value across m trials
• Inverse power of a test — how reliable power estimates are as a measure
• Application — what this means for the 5% cutoff and for replication studies
Four pages, free on arXiv, first posted 2015 and revised through January 2018.
Link: Paper
Navigational hashtags: #armknowledgesharing #armarticles
General hashtags: #statistics #math
@data_science_weekly
Post #246
217

- 👍 4