Free lesson · Linear algebra
Regression, conditioning, and regularisation
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Least squares selects coefficients that minimise squared residuals. If predictors nearly duplicate one another, many coefficient combinations fit almost equally well; regularisation makes extreme combinations less attractive.
Symbols, units & horizon
- y: n-vector of observed targets
- X: n-by-k design matrix of known predictors
- β: k-vector of coefficients
- ε (epsilon): residual vector y−Xβ
- β̂: estimated coefficients
- T: transpose
- −1: matrix inverse when it exists
- λ: nonnegative ridge penalty, not an eigenvalue here
- I: identity matrix
- || ||²: sum of squared vector entries
- ∇: gradient, the vector of partial derivatives
When and why to use this
Use regression for hedge ratios, factor exposure and simple forecasting baselines. Use ridge when correlated predictors make coefficients unstable.
Obtain the normal equations
- Minimise . Expand: .
- Differentiate with respect to β: , so . If full rank, solve to obtain the displayed formula; numerical code should solve rather than form an inverse.
- For one predictor with an intercept, centring gives , then .
x=1,2,3 and y=2,3,5 give b=3/2=1.5 and a=1/3. Prediction at x=4 is 6⅓.
Rows of X are observations, columns are features; include an intercept column if the model requires one. This closed form assumes full column rank. Numerical software should generally use QR or SVD rather than explicitly invert the matrix.
Add a ridge penalty
- Add to squared-error loss. Its gradient adds , giving .
- For a centred single predictor, . Usually replace the intercept penalty entry by zero.
With Sxy=3, Sxx=2 and λ=1, ridge slope=1 instead of 1.5. Predictors need consistent scaling before λ has a meaningful comparison.
Ridge penalises squared coefficient size and stabilises collinear features. Standardise predictors using training-only statistics and usually leave the intercept unpenalised. λ is chosen using validation, not the final test. A smaller training residual is not the same as a better forecast.
Research sources, review dates and limitations
Extend the research question
Represent each traded position by its payoff in each scenario. Check feasibility and hedge units before solving for a portfolio.
Continue with the connected research module →
Connect the ideas: Dependence and diversification
Retrieve: Joint behavior matters when combining uncertain outcomes.
Check the change: Correlation, cointegration, covariance and event dependence answer different questions.
Statistics → Time series → Portfolio construction → Portfolio management theory → Prediction foundations
Self-assessed. Write your explanation before opening this comparison. Correlation measures co-movement under a chosen sample and horizon. It does not establish a stationary combination or contractual convergence; those require separate definitions, tests and implementation checks.Explain it yourself: Why does a highly correlated pair not automatically provide a converging spread?
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
import numpy as np # dependency: numpy
def ols(features, targets):
"""Include a column of ones in X if an intercept is wanted."""
x, y = np.asarray(features, float), np.asarray(targets, float)
return np.linalg.lstsq(x, y, rcond=None)[0] # avoid explicit inverse
def ridge(features, targets, penalty):
"""This version penalises ALL columns, including any intercept."""
x, y = np.asarray(features, float), np.asarray(targets, float)
if penalty <= 0:
raise ValueError("Use OLS for zero penalty")
return np.linalg.solve(x.T@x + penalty*np.eye(x.shape[1]), x.T@y)
print(ols([[1,0],[1,1],[1,2]], [1,3,5]))Continue learning
Linear Algebra — all lessons- Start with lists: addition, scaling and a dot product
- Your portfolio is a vector; your risk is a matrix
- Eigenvalues: where the risk actually lives
- Regression, conditioning, and regularisation
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations