Trading Dev AcademyFree quant education

Free lesson · Financial machine learning

Regularised regression: an interpretable alpha baseline

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Linear regression estimates how features combine into a conditional mean. Ridge regression discourages very large coefficients, stabilising a fit when predictors are noisy or correlated.

Symbols, units & horizon
  • X: n×p centred training feature matrix
  • y: n-vector of centred return labels
  • β: p-vector of coefficients
  • || ||₂²: sum of squared components
  • λ: nonnegative ridge penalty for this unnormalised loss
  • I: p×p identity matrix
  • T: transpose
  • β̂: fitted coefficients
  • No intercept: assumed in this derivation

When and why to use this

Use ridge as a stable benchmark for return forecasting, factor residual modelling and cost estimation before adopting more flexible models.

Start with a small, economically motivated feature set and a forward residual-return label. A linear baseline is easy to inspect and often reveals data leakage, unit errors or nonlinear complexity that adds little. Standardise features inside each fold so the penalty has a comparable meaning across coefficients.

Ridge trades some in-sample fit for coefficient stability. The penalty is a hyperparameter chosen in inner chronological validation, not on the final test. An intercept is typically unpenalised or removed by training-only centring; the simplified calculation below assumes centred data and no intercept.

Inspect coefficient signs over time, out-of-fold residuals and performance by regime. Collinearity can make individual coefficients unstable even if predictions are similar. Interpretability of a fitted association does not establish causality or guarantee the coefficient survives transaction costs.

L(β)=‖y−Xβ‖22+λ‖β‖22,(X𝖳X+λI)β^=X𝖳y
Model assumptions, derivation and arithmetic

Regularised regression: an interpretable alpha baseline

  1. Expand the squared residual loss and differentiate with respect to β: −2Xᵀ(y−Xβ)+2λβ.
  2. Set the gradient to zero and collect terms to obtain (XᵀX+λI)β=Xᵀy.
  3. Solve the linear system. With one feature this reduces to β=Σxy/(Σx²+λ), making the shrinkage explicit.
Work it by hand

For x=[−1,1], y=[−2,2] and λ=2, Σxy=4 and Σx²=2, so β=4/(2+2)=1. Unpenalised fitting gives β=2.

Apply it in a strategy

  • Fit a training-only centred/scaled baseline and choose λ in inner chronological folds.
  • Convert out-of-fold predictions into the same constrained decision rule used by competitors.
  • Compare coefficient stability, net returns and turnover with an unregularised or constant baseline.

Research deliverable

Report the ridge path, out-of-fold forecast errors and net portfolio results under a fixed allocation rule.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

import numpy as np

def ridge_fit(x,y,penalty):
    x=np.asarray(x,dtype=float); y=np.asarray(y,dtype=float)
    if x.ndim!=2 or len(y)!=len(x) or penalty<=0: raise ValueError("Aligned matrix and positive penalty required")
    return np.linalg.solve(x.T@x+penalty*np.eye(x.shape[1]),x.T@y)

print(ridge_fit([[-1],[1]],[-2,2],2))

Continue learning

Machine Learning for Quantitative Strategy Development — all lessons
  1. Choose the model’s job: targets, horizons and decision layers
  2. Feature engineering, missingness and training-only transformations
  3. Regularised regression: an interpretable alpha baseline
  4. Logistic classification and cost-aware entry thresholds
  5. Trees and boosting: nonlinear interactions with controlled complexity
  6. Calibration, meta-labels and conditional payoff estimation
  7. Unsupervised learning, clusters and latent risk structure
  8. From model forecasts to a constrained strategy

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations