Free lesson · Financial machine learning
Trees and boosting: nonlinear interactions with controlled complexity
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A decision tree partitions feature space into regions with different predictions. Boosting adds small trees to correct the current model’s errors. These methods can capture interactions without manually specifying every product of inputs.
Symbols, units & horizon
- y_i: training target
- F_(m−1): current prediction ensemble
- e_i: residual for sample i
- h_m: new weak learner fitted to residuals for squared loss
- ν: learning rate between 0 and 1 in this example
- m: boosting iteration
- x: feature vector
- F_m: updated prediction
When and why to use this
Use tree ensembles for structured tabular features with plausible interactions and limited data, benchmarking them before sequence networks.
A tree might distinguish a momentum signal in liquid markets from the same signal when spreads are wide. Splits are learned from training data, so a plausible-looking rule can still fit noise. Deep trees and small leaves isolate unusual historical cases instead of stable effects.
For squared-error boosting, fit each weak learner to current residuals and add only a fraction of its prediction. Learning rate, depth, leaf size, number of trees and subsampling interact; tune them inside the chronological inner loop. For financial panels, group by date and preserve horizon purging.
Compare against ridge and a simple hand-specified interaction baseline. Feature importance is an attribution inside the fitted model, not causal evidence. Correlated features share or substitute for importance, and permutation of a time series may create unrealistic inputs. Use grouped or block-aware ablations.
Trees and boosting: nonlinear interactions with controlled complexity
- For loss (y−F)²/2, the negative derivative with respect to F is y−F. This gives the residual target.
- Fit a weak learner to approximate that residual using training data only.
- Add ν times its output to the existing forecast. A smaller ν limits each step but usually requires more rounds, which are selected by validation.
Current forecasts [1,1], targets [2,0] produce residuals [1,−1]. If a learner matches them and ν=.1, new forecasts are [1.1,.9]. Training error decreases; generalisation remains untested.
Apply it in a strategy
- Start with shallow trees, minimum leaf sizes and an explicit round budget.
- Select complexity with chronological validation and fixed feature preprocessing; audit stability across folds.
- Ablate whole feature families and compare net decision value at matched turnover.
Research deliverable
Create a model comparison including baseline, boosting, feature-family ablations and regime-level errors, with every tuning trial retained.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def boosting_step(target,current,weak_prediction,learning_rate=.1):
if not len(target)==len(current)==len(weak_prediction) or not 0<learning_rate<=1: raise ValueError("Aligned inputs required")
residual=[y-f for y,f in zip(target,current)]
updated=[f+learning_rate*h for f,h in zip(current,weak_prediction)]
return residual,updated
print(boosting_step([2,0],[1,1],[1,-1]))Continue learning
Machine Learning for Quantitative Strategy Development — all lessons- Choose the model’s job: targets, horizons and decision layers
- Feature engineering, missingness and training-only transformations
- Regularised regression: an interpretable alpha baseline
- Logistic classification and cost-aware entry thresholds
- Trees and boosting: nonlinear interactions with controlled complexity
- Calibration, meta-labels and conditional payoff estimation
- Unsupervised learning, clusters and latent risk structure
- From model forecasts to a constrained strategy
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations