Free lesson · Quant strategy development
03 / Test predictive information before a complex model
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A predictive score should distinguish better from worse opportunities on data excluded from its fitting. Ranking quality and return calibration are different: a signal can order assets correctly but exaggerate the size of the expected payoff.
Symbols, units & horizon
- ICₜ: cross-sectional correlation at decision time t
- i: instrument index
- sᵢ,ₜ: frozen signal score
- Rᵢ,ₜ:ₜ₊ₕ: subsequent realised return over h periods
- T: number of evaluation dates
- IC̄: mean IC
- μ̂(s): calibrated expected return for score s
- â,b̂: training intercept and slope
- k: score bucket
- R̄_k: average bucket return
- ĉ_k: estimated matching bucket cost
When and why to use this
Use information coefficients for ranking diagnostics, and bucket returns or calibration slopes to convert scores into return units for sizing.
Start by asking whether the feature orders future returns. Plot training-sample buckets and evaluate the same predeclared buckets on untouched periods. A monotonic relationship that persists after costs and across samples is more informative than a single optimised threshold.
Calculate an information coefficient
- At a single date, centre scores and subsequent returns across eligible assets. Divide their cross-product sum by the square root of the two squared-deviation sums.
- Spearman IC first replaces each variable by its ranks, using average ranks for ties. Average date-level ICs to form the stated mean; this is not the same as pooling all asset-date observations.
Scores 1,2,3 and future returns 3%,1%,2% have centred cross-product −.01 and denominator .02, so Pearson IC=−.5.
A cross-sectional information coefficient correlates scores with future returns across eligible instruments at each date. Spearman IC uses ranks. Avoid treating every instrument-date as independent: common shocks, overlapping horizons, and persistent signals reduce effective sample size. Report uncertainty with a method suited to that dependence.
Calibrate a score and net out bucket costs
- Fit a straight line by minimising squared training residuals. Slope ; intercept . Predict .
- For a held-out score bucket, average its observed returns and subtract costs calculated for its actual turnover and size. Do not subtract a per-trade rate directly from a multi-trade portfolio return.
a=2 bp, b=3 bp per score unit, s=2 gives an 8 bp forecast. If that bucket costs 5 bp to trade over the horizon, estimated net is 3 bp.
The linear calibration is a baseline, not a claim that score-to-return behaviour is linear. Fit it only on training data. A useful rank ordering can still have poor absolute return calibration; sizing requires an estimate in return units, not just a rank.
- Run ablations: remove one feature family at a time, refit on the training set, and evaluate its incremental benefit under the same selection procedure.
- Use placebo features and time shifts as diagnostics for accidental leakage or an overly flexible pipeline.
- Compare neighbouring lookbacks and thresholds. A sharp isolated optimum is less convincing than a broad stable region.
- Record how many models, assets, horizons and thresholds were tried. Reuse of a test set is part of the search process.
- Check tail attribution: identify the fraction of P&L contributed by the best five trades, one asset, or one regime.
Further reading: NBER: how searching many signals distorts backtest evidence ↗
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
from statistics import mean, correlation, linear_regression
def information_coefficient(scores, future_returns):
"""Evaluation only: future_returns must never enter the score function."""
return correlation(scores, future_returns)
def calibrate(training_scores, training_returns):
slope, intercept = linear_regression(training_scores, training_returns)
return intercept, slope
def forecast(score, intercept, slope):
return intercept+slope*score
def net_bucket(realized_returns, estimated_cost):
return mean(realized_returns)-estimated_cost
print(calibrate([-1,0,1], [-.01,0,.01]))Continue learning
Quant Strategy Development — all lessons- 01 / Start with a source of return
- 02 / The variables that actually enter the decision
- 03 / Test predictive information before a complex model
- 04 / Momentum and trend: information that persists
- 05 / Mean reversion and relative value
- 06 / Carry, events, and liquidity provision
- 07 / Convert a forecast into a trade decision
- 08 / Build a bot that preserves the experiment
- 09 / Decide whether the edge is real enough to continue
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations