Trading Dev AcademyFree quant education

Free lesson · Research & robust tuning

Walk-forward validation, overlapping labels and purging

Open interactive lessonPractice calculationsExplore labs

Start with the idea

A chronological split preserves the direction in which information arrives. Overlapping label intervals can nevertheless transfer future outcomes into training, even when feature rows are ordered by date.

Symbols, units & horizon
  • i: candidate training observation
  • s_i: sample information/start time
  • e_i: label end/availability time
  • v₀,v₁: validation interval boundaries
  • g: extra pre-validation time gap in the same units
  • T: eligible chronological training set
  • ∩: interval overlap
  • ∅: empty set
  • Strict inequality: conservative boundary convention

When and why to use this

Use interval-aware folds for forward returns, barrier labels, overlapping holdings and execution outcomes. It makes data availability auditable.

Store each sample’s feature availability time and the end of its label interval. For a model fitted before validation begins, a training sample is usable only if its label was fully observed by that fit time. For event labels, the end is the event’s actual stopping time, not a fixed guessed lag.

Purging removes training labels that overlap the protected validation interval. An additional gap can address dependencies induced by the data construction. In schemes that train on both sides of a test interval, embargo excludes observations immediately after the test. A fixed percentage is not universally sufficient.

Fit scaling, PCA, imputation, feature selection, calibration and any secondary models inside each training fold. Splitting only after preprocessing leaks information. Walk-forward performance evaluates a retraining policy; a model fit once and held forever answers a different question.

𝒯={i:ei<v0−g},[si,ei]∩[v0,v1]=⌀
Model assumptions, derivation and arithmetic

Walk-forward validation, overlapping labels and purging

  1. List each training label’s actual observation interval, rather than only its row timestamp.
  2. For an expanding window, retain samples whose label ends strictly before the fit cutoff v₀−g. Such intervals cannot overlap the future validation interval.
  3. For other fold layouts, check interval intersection explicitly: overlap occurs when s_i≤v₁ and e_i≥v₀. Remove those samples and apply the chosen post-test embargo if training uses later observations.
Work it by hand

Validation begins at day 100 with gap 2. Labels ending at days 96,98,101 retain only day 96 under the strict e<98 rule. A row dated day 95 with label ending 101 must be removed.

Apply it in a strategy

  • Attach availability and label-end timestamps to every sample before building folds.
  • Run all learned preprocessing inside the fold and persist its training-only parameters.
  • Evaluate the whole retraining schedule on outer future blocks, reporting gaps and excluded observations.

Research deliverable

Draw a sample timeline and export the exact train/validation indices for every fold, including purged and embargoed rows.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def past_training_indices(label_ends,validation_start,gap=0):
    if gap<0: raise ValueError("Gap must be nonnegative")
    return [i for i,end in enumerate(label_ends) if end<validation_start-gap]

def overlaps(start,end,test_start,test_end):
    return start<=test_end and end>=test_start

print(past_training_indices([96,98,101],100,2))

Continue learning

Strategy Research, Backtesting & Robust Optimisation — all lessons
  1. Write the experiment before the strategy
  2. Walk-forward validation, overlapping labels and purging
  3. Hyperparameter optimisation without an unrestricted search
  4. Multiple trials, false discoveries and selection diagnostics
  5. Dependent returns, block bootstrap and realistic stress tests
  6. Fine-tuning, retraining and the research-to-production decision

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations