Free lesson · Research & robust tuning
Multiple trials, false discoveries and selection diagnostics
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Searching many noisy strategies creates impressive winners even when no candidate has a real advantage. The winning backtest must be judged in light of the entire search, not as if it were the only experiment ever run.
Symbols, units & horizon
- α: per-test false-positive probability in the first identity, family target in the Bonferroni rule
- M: number of tests
- Independence: required for the exact first identity
- α_Bonf: per-test threshold controlling family-wise error by a union bound without independence
- False positive: rejection of a true null under valid tests
When and why to use this
Use multiplicity reasoning whenever a claimed edge emerged from searching many alternatives. It explains why a spectacular winner can be weak evidence.
Count model families, feature sets, thresholds, universes, holding horizons and discretionary revisions. Correlated trials are not independent, so neither a raw trial count nor an independence assumption fully describes the selection process. Retain the return matrix for all candidates when possible.
The probability of backtest overfitting framework compares in-sample selection with out-of-sample ranking across combinatorial splits. Deflated Sharpe diagnostics address selection and non-normality under additional assumptions. Neither is a posterior probability that your strategy is profitable, and neither replaces a future deployment-style evaluation.
For confirmatory hypotheses, family-wise or false-discovery-rate procedures require a defined family and valid underlying tests. Sequentially inventing a new family after each disappointing result defeats the correction. Report uncertainty and failed hypotheses instead of decorating one selected Sharpe with a naive p-value.
Multiple trials, false discoveries and selection diagnostics
- If M null tests are independent, the probability that none rejects is (1−α)^M. Take the complement.
- Without independence, the probability of any rejection is at most the sum of individual error probabilities.
- Setting each error probability to α/M makes that sum α. This conservative bound does not make an invalid test valid.
Twenty independent null tests at .05 produce probability 1−.95^20≈.6415 of at least one false positive. For family error .05 with 20 tests, the Bonferroni threshold is .0025.
Apply it in a strategy
- Retain all candidate outcomes and distinguish exploratory searches from fixed confirmatory hypotheses.
- Choose a justified selection diagnostic and preserve its assumptions, including trial dependence and serial dependence.
- Reserve a chronological external evaluation and report performance degradation after selection.
Research deliverable
Add a search-history appendix listing the full candidate family, selection rule and the uncertainty method applied after selection.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def multiple_testing(alpha,trials):
if not 0<alpha<1 or not isinstance(trials,int) or trials<1: raise ValueError("Valid alpha and trial count required")
return 1-(1-alpha)**trials,alpha/trials
print(multiple_testing(.05,20))Continue learning
Strategy Research, Backtesting & Robust Optimisation — all lessons- Write the experiment before the strategy
- Walk-forward validation, overlapping labels and purging
- Hyperparameter optimisation without an unrestricted search
- Multiple trials, false discoveries and selection diagnostics
- Dependent returns, block bootstrap and realistic stress tests
- Fine-tuning, retraining and the research-to-production decision
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations