Trading Dev AcademyFree quant education

Free lesson · Trading algorithms

Linear models, random forests and gradient boosting

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Predictive algorithms differ in the patterns they can represent and the variance they introduce. Select them by the structure and size of the dataset, then compare the decisions their predictions produce.

Symbols, units & horizon
  • f_m: individual model prediction
  • M: number of equally weighted models
  • σ²: equal prediction-error variance in the illustrative ensemble model
  • ρ: common pairwise error correlation
  • f̂: ensemble prediction
  • Variance formula: assumes equal variances and common correlations, not arbitrary forests

When and why to use this

Use this comparison to select interpretable linear, bagged-tree or boosted-tree baselines and assess whether model diversity adds useful stability.

Linear and ridge models offer transparent additive relationships. Random forests average many randomised trees, reducing variance when trees are not perfectly correlated. Gradient boosting adds learners sequentially to reduce the chosen loss. XGBoost and LightGBM are implementations of boosted-tree ideas with engineering and regularisation choices.

Tabular features such as lagged returns, carry, liquidity and events often justify trying trees before a large sequence network. Forests and boosting can learn nonlinear thresholds and interactions, but may extrapolate poorly beyond training ranges. Their probabilities may require calibration.

Use grouped chronological splits, minimum leaf sizes and a fixed tuning budget. Compare feature-family ablations rather than reading a single importance plot as economic truth. Choose the smallest model that shows a reliable incremental decision benefit at the required latency.

f^(x)=1M∑m=1Mfm(x),Var⁡(f^)=σ2[ρ+1−ρM]
Model assumptions, derivation and arithmetic

Linear models, random forests and gradient boosting

  1. Expand variance of the average as 1/M² times the sum of all variances and covariances.
  2. There are M diagonal terms σ² and M(M−1) off-diagonal terms ρσ².
  3. Simplify to σ²/M+ρσ²(M−1)/M. As M grows, correlated error remains as a floor.
Work it by hand

With σ²=1, ρ=.5 and M=10, ensemble error variance is .5+.5/10=.55. Adding more nearly identical models cannot remove the shared error.

Apply it in a strategy

  • Fit ridge, a constrained forest and shallow boosting under the same feature and label contract.
  • Calibrate outputs where probabilities drive thresholds, and use the same portfolio/execution layer.
  • Compare error dependence, net value, compute cost and stability across future periods.

Research deliverable

Build an algorithm selection table listing the input type, baseline, tuning budget, latency and measured incremental benefit.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def ensemble_variance(single_variance,correlation,models):
    if single_variance<0 or not 0<=correlation<=1 or not isinstance(models,int) or models<1: raise ValueError("Valid illustrative ensemble inputs required")
    return single_variance*(correlation+(1-correlation)/models)

print(ensemble_variance(1,.5,10))

Continue learning

Trading Algorithms: A Practical Selection Guide — all lessons
  1. Moving averages, EWMA and momentum/reversion rules
  2. Kalman filtering: combine a prediction with a noisy observation
  3. ARIMA for conditional means and GARCH for conditional variance
  4. Linear models, random forests and gradient boosting
  5. PCA, clustering, risk parity and quadratic programming
  6. TWAP, VWAP and percentage-of-volume execution
  7. Optimal execution: impact versus waiting risk

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations