Free lesson · Financial machine learning
Regularised regression: an interpretable alpha baseline
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Linear regression estimates how features combine into a conditional mean. Ridge regression discourages very large coefficients, stabilising a fit when predictors are noisy or correlated.
Symbols, units & horizon
- X: n×p centred training feature matrix
- y: n-vector of centred return labels
- β: p-vector of coefficients
- || ||₂²: sum of squared components
- λ: nonnegative ridge penalty for this unnormalised loss
- I: p×p identity matrix
- T: transpose
- β̂: fitted coefficients
- No intercept: assumed in this derivation
When and why to use this
Use ridge as a stable benchmark for return forecasting, factor residual modelling and cost estimation before adopting more flexible models.
Start with a small, economically motivated feature set and a forward residual-return label. A linear baseline is easy to inspect and often reveals data leakage, unit errors or nonlinear complexity that adds little. Standardise features inside each fold so the penalty has a comparable meaning across coefficients.
Ridge trades some in-sample fit for coefficient stability. The penalty is a hyperparameter chosen in inner chronological validation, not on the final test. An intercept is typically unpenalised or removed by training-only centring; the simplified calculation below assumes centred data and no intercept.
Inspect coefficient signs over time, out-of-fold residuals and performance by regime. Collinearity can make individual coefficients unstable even if predictions are similar. Interpretability of a fitted association does not establish causality or guarantee the coefficient survives transaction costs.
Regularised regression: an interpretable alpha baseline
- Expand the squared residual loss and differentiate with respect to β: −2Xᵀ(y−Xβ)+2λβ.
- Set the gradient to zero and collect terms to obtain (XᵀX+λI)β=Xᵀy.
- Solve the linear system. With one feature this reduces to β=Σxy/(Σx²+λ), making the shrinkage explicit.
For x=[−1,1], y=[−2,2] and λ=2, Σxy=4 and Σx²=2, so β=4/(2+2)=1. Unpenalised fitting gives β=2.
Apply it in a strategy
- Fit a training-only centred/scaled baseline and choose λ in inner chronological folds.
- Convert out-of-fold predictions into the same constrained decision rule used by competitors.
- Compare coefficient stability, net returns and turnover with an unregularised or constant baseline.
Research deliverable
Report the ridge path, out-of-fold forecast errors and net portfolio results under a fixed allocation rule.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
import numpy as np
def ridge_fit(x,y,penalty):
x=np.asarray(x,dtype=float); y=np.asarray(y,dtype=float)
if x.ndim!=2 or len(y)!=len(x) or penalty<=0: raise ValueError("Aligned matrix and positive penalty required")
return np.linalg.solve(x.T@x+penalty*np.eye(x.shape[1]),x.T@y)
print(ridge_fit([[-1],[1]],[-2,2],2))Continue learning
Machine Learning for Quantitative Strategy Development — all lessons- Choose the model’s job: targets, horizons and decision layers
- Feature engineering, missingness and training-only transformations
- Regularised regression: an interpretable alpha baseline
- Logistic classification and cost-aware entry thresholds
- Trees and boosting: nonlinear interactions with controlled complexity
- Calibration, meta-labels and conditional payoff estimation
- Unsupervised learning, clusters and latent risk structure
- From model forecasts to a constrained strategy
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations