Trading Dev AcademyFree quant education

Free lesson · Financial machine learning

Calibration, meta-labels and conditional payoff estimation

Open interactive lessonPractice calculationsExplore labs

Start with the idea

A score should be judged by the decision it supports. Calibration asks whether predicted probabilities match observed frequencies; meta-labelling asks whether an existing candidate trade should be accepted or sized differently.

Symbols, units & horizon
  • BS: Brier score, lower is better
  • p_i: probability forecast
  • y_i: binary realised outcome
  • n: evaluated examples
  • v̂_i: predicted net value in currency or bps
  • G_i,L_i: positive conditional gain and loss magnitudes in matching units
  • C_i: cost not already incorporated in those payoffs

When and why to use this

Use calibration and meta-labelling to improve how an existing strategy acts, while preserving the primary signal’s out-of-sample lineage.

A primary strategy generates candidate direction and timing. A secondary model can estimate whether those candidates are likely to work after costs. Its training examples must come from out-of-fold primary predictions; otherwise the primary model’s in-sample optimism leaks into the secondary labels.

Separate calibration from ranking. Reliability bins compare predicted probability with observed frequency, while Brier score measures squared probability error. Calibrate on a separate past split and evaluate on later data. Small bins and dependent trades make apparent calibration noisy.

Combine probability with conditional gains/losses, not a fixed 50% threshold. Include opportunity cost and portfolio constraints. If a meta-model rejects most trades, compare it with a simple lower-turnover baseline to identify whether the benefit is selection or merely less trading.

BS=1n∑i(pi−yi)2,v^i=piGi−(1−pi)Li−Ci
Model assumptions, derivation and arithmetic

Calibration, meta-labels and conditional payoff estimation

  1. A binary outcome is 0 or 1; subtract it from the forecast probability and square the error. Average across examples for Brier score.
  2. For decision value, average the gain and negative loss using p and 1−p, then subtract cost.
  3. Use a threshold based on this estimated value and uncertainty. Brier score does not by itself specify a position size.
Work it by hand

Forecasts [.8,.4] for outcomes [1,0] have Brier (.04+.16)/2=.10. With p=.6, G=12, L=8 and C=2 bps, predicted net value is 7.2−3.2−2=2 bps.

Apply it in a strategy

  • Generate candidate trades and scores from outer-safe, out-of-fold primary models.
  • Fit the secondary acceptance model and calibration only on permitted historical candidates.
  • Compare filtered and unfiltered strategies at matched turnover, exposure and costs, including rejected-trade outcomes.

Research deliverable

Provide a primary/secondary split diagram, reliability report and incremental net-value analysis for the filter.

Start with probability, price and a two-outcome payoff →

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def brier(probabilities,outcomes):
    if not probabilities or len(probabilities)!=len(outcomes): raise ValueError("Aligned nonempty inputs required")
    if any(not 0<=p<=1 for p in probabilities) or any(y not in (0,1) for y in outcomes): raise ValueError("Invalid probability or label")
    return sum((p-y)**2 for p,y in zip(probabilities,outcomes))/len(outcomes)

def conditional_value(p,gain,loss,cost): return p*gain-(1-p)*loss-cost

print(brier([.8,.4],[1,0]),conditional_value(.6,12,8,2))

Continue learning

Machine Learning for Quantitative Strategy Development — all lessons
  1. Choose the model’s job: targets, horizons and decision layers
  2. Feature engineering, missingness and training-only transformations
  3. Regularised regression: an interpretable alpha baseline
  4. Logistic classification and cost-aware entry thresholds
  5. Trees and boosting: nonlinear interactions with controlled complexity
  6. Calibration, meta-labels and conditional payoff estimation
  7. Unsupervised learning, clusters and latent risk structure
  8. From model forecasts to a constrained strategy

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations