Free lesson · Research & backtests
Account for dependence and multiple experiments
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Overlapping positions share shocks, so the number of rows in a dataset can overstate the amount of independent information. Searching many candidate rules also increases the chance that one looks impressive purely by chance.
Symbols, units & horizon
- R̄: average of n periodic returns
- γ₀: return variance
- γ_k: covariance of returns k periods apart
- k: positive lag
- α: per-test false-positive probability
- M: independent tests under the null for the displayed family probability
- Var: variance
- Σ: sum over lags
- 1−k/n: fraction of covariance pairs at lag k
- ∪: union, at least one event occurs
- Eᵢ: false rejection event for test i
- α/M: Bonferroni per-test significance threshold
When and why to use this
Use dependence-aware uncertainty for persistent returns and record the complete experiment family. These are central to deciding whether a pattern’s apparent advantage deserves further testing.
The familiar t-statistic assumes a suitable sampling model. Overlapping trades, persistent positions, changing volatility, and shared market exposure can all invalidate an independent-observation standard error. More trades do not necessarily mean proportionally more information.
Count covariance pairs in an average
- . There are n diagonal terms γ₀.
- At positive lag k, there are n−k pairs in each direction. Add , then factor 1/n to obtain the displayed expression. Setting all nonzero-lag covariances to zero recovers the IID result.
For n=4, γ₀=1, γ₁=.2 and other lags zero: variance of the mean=[4+2(3)(.2)]/16=.325 rather than .25.
For a covariance-stationary series, is lag-k covariance. Positive serial dependence raises the variance of the sample mean. HAC estimators truncate and weight the covariance sum; block bootstrap methods resample contiguous observations. Both need explicit bandwidth or block-length choices.
Take the complement of no false positives
- One null test avoids a false rejection with probability 1−α. Under independence, all M avoid one with probability . Subtract from 1.
- Bonferroni instead uses the union bound . Testing each at α/M controls the family error at α without requiring independence.
At α=.05,M=100, the independent probability is 1−.95¹⁰⁰≈99.408%. Bonferroni’s individual threshold is .0005.
With 100 independent null tests at 5%, the chance of at least one false positive is about 99.4%. Correlated tests change this calculation, but do not eliminate selection bias. Bonferroni tests each hypothesis at α/M to control family-wise error; false-discovery-rate procedures answer a different question about the expected proportion of false discoveries.
Further reading: NBER: Backtesting Strategies Based on Multiple Signals ↗
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def mean_variance_from_autocovariance(n, gamma):
"""gamma[k] is the model covariance at lag k, for k=0,...,n-1."""
if n < 1 or len(gamma) < n:
raise ValueError("Need n covariance entries")
return (gamma[0]+2*sum((1-k/n)*gamma[k] for k in range(1,n)))/n
def independent_family_false_positive(alpha, tests):
if not 0 <= alpha <= 1 or tests < 0:
raise ValueError("Invalid probability or test count")
return 1-(1-alpha)**tests
print(independent_family_false_positive(.05, 20))
def bonferroni_threshold(family_alpha, tests):
if tests < 1 or not 0 < family_alpha < 1:
raise ValueError("Positive test count and alpha in (0,1) required")
return family_alpha/tests
Continue learning
Research & Backtest Design — all lessons- Build a point-in-time dataset
- Separate model selection from evaluation
- Account for dependence and multiple experiments
- Make the accounting identity your first test
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations