Trading Dev AcademyFree quant education

Free lesson · Research & robust tuning

Multiple trials, false discoveries and selection diagnostics

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Searching many noisy strategies creates impressive winners even when no candidate has a real advantage. The winning backtest must be judged in light of the entire search, not as if it were the only experiment ever run.

Symbols, units & horizon
  • α: per-test false-positive probability in the first identity, family target in the Bonferroni rule
  • M: number of tests
  • Independence: required for the exact first identity
  • α_Bonf: per-test threshold controlling family-wise error by a union bound without independence
  • False positive: rejection of a true null under valid tests

When and why to use this

Use multiplicity reasoning whenever a claimed edge emerged from searching many alternatives. It explains why a spectacular winner can be weak evidence.

Count model families, feature sets, thresholds, universes, holding horizons and discretionary revisions. Correlated trials are not independent, so neither a raw trial count nor an independence assumption fully describes the selection process. Retain the return matrix for all candidates when possible.

The probability of backtest overfitting framework compares in-sample selection with out-of-sample ranking across combinatorial splits. Deflated Sharpe diagnostics address selection and non-normality under additional assumptions. Neither is a posterior probability that your strategy is profitable, and neither replaces a future deployment-style evaluation.

For confirmatory hypotheses, family-wise or false-discovery-rate procedures require a defined family and valid underlying tests. Sequentially inventing a new family after each disappointing result defeats the correction. Report uncertainty and failed hypotheses instead of decorating one selected Sharpe with a naive p-value.

P(at least one false positive)=1−(1−α)M,αBonf=αM
Model assumptions, derivation and arithmetic

Multiple trials, false discoveries and selection diagnostics

  1. If M null tests are independent, the probability that none rejects is (1−α)^M. Take the complement.
  2. Without independence, the probability of any rejection is at most the sum of individual error probabilities.
  3. Setting each error probability to α/M makes that sum α. This conservative bound does not make an invalid test valid.
Work it by hand

Twenty independent null tests at .05 produce probability 1−.95^20≈.6415 of at least one false positive. For family error .05 with 20 tests, the Bonferroni threshold is .0025.

Apply it in a strategy

  • Retain all candidate outcomes and distinguish exploratory searches from fixed confirmatory hypotheses.
  • Choose a justified selection diagnostic and preserve its assumptions, including trial dependence and serial dependence.
  • Reserve a chronological external evaluation and report performance degradation after selection.

Research deliverable

Add a search-history appendix listing the full candidate family, selection rule and the uncertainty method applied after selection.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def multiple_testing(alpha,trials):
    if not 0<alpha<1 or not isinstance(trials,int) or trials<1: raise ValueError("Valid alpha and trial count required")
    return 1-(1-alpha)**trials,alpha/trials

print(multiple_testing(.05,20))

Continue learning

Strategy Research, Backtesting & Robust Optimisation — all lessons
  1. Write the experiment before the strategy
  2. Walk-forward validation, overlapping labels and purging
  3. Hyperparameter optimisation without an unrestricted search
  4. Multiple trials, false discoveries and selection diagnostics
  5. Dependent returns, block bootstrap and realistic stress tests
  6. Fine-tuning, retraining and the research-to-production decision

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations