Free lesson · Research & backtests
Separate model selection from evaluation
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Training learns numerical parameters, validation chooses among alternatives and a final holdout evaluates the chosen process. Every time feedback changes a choice, that feedback becomes part of the training history.
Symbols, units & horizon
- Xₜ: predictors known at time t
- yₛ: realised training target for sample s
- R̂ₜ₊₁: predicted next-period return
- f_θ: prediction model with parameters θ
- θ̂ₜ: parameters fitted using only permitted training rows
- 𝒯ₜ: training index set with mature labels available by t
- ℓ: loss function
- argmin: parameter value giving the smallest objective
- s∈𝒯ₜ: loop over the training set
When and why to use this
Use chronological walk-forward evaluation for rules that will be retrained over time. Purge label overlap so a training target does not contain the future return being tested.
Training estimates parameters; validation selects model choices; a final holdout evaluates the frozen process. Repeatedly consulting the holdout turns it into another training set. A chronological split is a starting point, not a universal 60/40 recipe.
Solve a simple training objective
- The argmin means “the parameter value giving the smallest loss.” With a constant forecast θ and squared loss, minimise .
- Differentiate: , so . Freeze that estimate before forecasting the next observation.
- With a linear forecast a+bx, the same operation gives the normal equations in the regression lesson. Change training dates without accessing the next test target.
Training targets 1%,2%,3% imply a constant forecast 2%. The next realised return must not be included in that mean until it is genuinely available.
The training set must finish before the test observation, including the label's full horizon. For a five-day forward-return label, removing only the last one-day observation does not remove overlap. Purge overlapping label intervals and use a gap justified by the information structure.
- Walk forward: train on an expanding or rolling past window, freeze, evaluate the next block, then advance.
- Report results by period, market regime, asset, turnover, and trade-size bucket; one aggregate Sharpe can hide concentration.
- Retain an experiment log including failed variants. Freeze the final model before paper trading.
- Paper trading tests data timing, order lifecycle, and reconciliations; simulated fills still do not demonstrate live capacity.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
from statistics import mean
def fit_constant_squared_loss(training_targets):
"""The minimiser of sum((y-theta)**2) is the training mean."""
return mean(training_targets)
def training_rows(rows, decision_time):
"""Both feature and label must have been available before fitting."""
return [r for r in rows if r["feature_time"] < decision_time
and r["label_available_at"] < decision_time]
print(fit_constant_squared_loss([.01, -.02, .04])) # .01Continue learning
Research & Backtest Design — all lessons- Build a point-in-time dataset
- Separate model selection from evaluation
- Account for dependence and multiple experiments
- Make the accounting identity your first test
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations