Trading Dev AcademyFree quant education

Free lesson · Research & backtests

Account for dependence and multiple experiments

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Overlapping positions share shocks, so the number of rows in a dataset can overstate the amount of independent information. Searching many candidate rules also increases the chance that one looks impressive purely by chance.

Symbols, units & horizon
  • R̄: average of n periodic returns
  • γ₀: return variance
  • γ_k: covariance of returns k periods apart
  • k: positive lag
  • α: per-test false-positive probability
  • M: independent tests under the null for the displayed family probability
  • Var: variance
  • Σ: sum over lags
  • 1−k/n: fraction of covariance pairs at lag k
  • ∪: union, at least one event occurs
  • Eᵢ: false rejection event for test i
  • α/M: Bonferroni per-test significance threshold

When and why to use this

Use dependence-aware uncertainty for persistent returns and record the complete experiment family. These are central to deciding whether a pattern’s apparent advantage deserves further testing.

The familiar t-statistic assumes a suitable sampling model. Overlapping trades, persistent positions, changing volatility, and shared market exposure can all invalidate an independent-observation standard error. More trades do not necessarily mean proportionally more information.

Var⁡(R‾)=1n[γ0+2∑k=1n−1(1−kn)γk]
Algebra and arithmetic

Count covariance pairs in an average

  1. Var(R‾)=n−2∑i∑jCov(Ri,Rj). There are n diagonal terms γ₀.
  2. At positive lag k, there are n−k pairs in each direction. Add 2(n−k)γk, then factor 1/n to obtain the displayed expression. Setting all nonzero-lag covariances to zero recovers the IID result.
Work it by hand

For n=4, γ₀=1, γ₁=.2 and other lags zero: variance of the mean=[4+2(3)(.2)]/16=.325 rather than .25.

For a covariance-stationary series, γk is lag-k covariance. Positive serial dependence raises the variance of the sample mean. HAC estimators truncate and weight the covariance sum; block bootstrap methods resample contiguous observations. Both need explicit bandwidth or block-length choices.

P(at least one false positive)=1−(1−α)M
Algebra and arithmetic

Take the complement of no false positives

  1. One null test avoids a false rejection with probability 1−α. Under independence, all M avoid one with probability (1−α)M. Subtract from 1.
  2. Bonferroni instead uses the union bound P(∪Ei)≤∑P(Ei). Testing each at α/M controls the family error at α without requiring independence.
Work it by hand

At α=.05,M=100, the independent probability is 1−.95¹⁰⁰≈99.408%. Bonferroni’s individual threshold is .0005.

With 100 independent null tests at 5%, the chance of at least one false positive is about 99.4%. Correlated tests change this calculation, but do not eliminate selection bias. Bonferroni tests each hypothesis at α/M to control family-wise error; false-discovery-rate procedures answer a different question about the expected proportion of false discoveries.

Further reading: NBER: Backtesting Strategies Based on Multiple Signals ↗

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def mean_variance_from_autocovariance(n, gamma):
    """gamma[k] is the model covariance at lag k, for k=0,...,n-1."""
    if n < 1 or len(gamma) < n:
        raise ValueError("Need n covariance entries")
    return (gamma[0]+2*sum((1-k/n)*gamma[k] for k in range(1,n)))/n

def independent_family_false_positive(alpha, tests):
    if not 0 <= alpha <= 1 or tests < 0:
        raise ValueError("Invalid probability or test count")
    return 1-(1-alpha)**tests

print(independent_family_false_positive(.05, 20))
def bonferroni_threshold(family_alpha, tests):
    if tests < 1 or not 0 < family_alpha < 1:
        raise ValueError("Positive test count and alpha in (0,1) required")
    return family_alpha/tests

Continue learning

Research & Backtest Design — all lessons
  1. Build a point-in-time dataset
  2. Separate model selection from evaluation
  3. Account for dependence and multiple experiments
  4. Make the accounting identity your first test

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations