Free lesson · Financial machine learning
Calibration, meta-labels and conditional payoff estimation
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A score should be judged by the decision it supports. Calibration asks whether predicted probabilities match observed frequencies; meta-labelling asks whether an existing candidate trade should be accepted or sized differently.
Symbols, units & horizon
- BS: Brier score, lower is better
- p_i: probability forecast
- y_i: binary realised outcome
- n: evaluated examples
- v̂_i: predicted net value in currency or bps
- G_i,L_i: positive conditional gain and loss magnitudes in matching units
- C_i: cost not already incorporated in those payoffs
When and why to use this
Use calibration and meta-labelling to improve how an existing strategy acts, while preserving the primary signal’s out-of-sample lineage.
A primary strategy generates candidate direction and timing. A secondary model can estimate whether those candidates are likely to work after costs. Its training examples must come from out-of-fold primary predictions; otherwise the primary model’s in-sample optimism leaks into the secondary labels.
Separate calibration from ranking. Reliability bins compare predicted probability with observed frequency, while Brier score measures squared probability error. Calibrate on a separate past split and evaluate on later data. Small bins and dependent trades make apparent calibration noisy.
Combine probability with conditional gains/losses, not a fixed 50% threshold. Include opportunity cost and portfolio constraints. If a meta-model rejects most trades, compare it with a simple lower-turnover baseline to identify whether the benefit is selection or merely less trading.
Calibration, meta-labels and conditional payoff estimation
- A binary outcome is 0 or 1; subtract it from the forecast probability and square the error. Average across examples for Brier score.
- For decision value, average the gain and negative loss using p and 1−p, then subtract cost.
- Use a threshold based on this estimated value and uncertainty. Brier score does not by itself specify a position size.
Forecasts [.8,.4] for outcomes [1,0] have Brier (.04+.16)/2=.10. With p=.6, G=12, L=8 and C=2 bps, predicted net value is 7.2−3.2−2=2 bps.
Apply it in a strategy
- Generate candidate trades and scores from outer-safe, out-of-fold primary models.
- Fit the secondary acceptance model and calibration only on permitted historical candidates.
- Compare filtered and unfiltered strategies at matched turnover, exposure and costs, including rejected-trade outcomes.
Research deliverable
Provide a primary/secondary split diagram, reliability report and incremental net-value analysis for the filter.
Start with probability, price and a two-outcome payoff →
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def brier(probabilities,outcomes):
if not probabilities or len(probabilities)!=len(outcomes): raise ValueError("Aligned nonempty inputs required")
if any(not 0<=p<=1 for p in probabilities) or any(y not in (0,1) for y in outcomes): raise ValueError("Invalid probability or label")
return sum((p-y)**2 for p,y in zip(probabilities,outcomes))/len(outcomes)
def conditional_value(p,gain,loss,cost): return p*gain-(1-p)*loss-cost
print(brier([.8,.4],[1,0]),conditional_value(.6,12,8,2))Continue learning
Machine Learning for Quantitative Strategy Development — all lessons- Choose the model’s job: targets, horizons and decision layers
- Feature engineering, missingness and training-only transformations
- Regularised regression: an interpretable alpha baseline
- Logistic classification and cost-aware entry thresholds
- Trees and boosting: nonlinear interactions with controlled complexity
- Calibration, meta-labels and conditional payoff estimation
- Unsupervised learning, clusters and latent risk structure
- From model forecasts to a constrained strategy
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations