Free lesson · Financial machine learning
Unsupervised learning, clusters and latent risk structure
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Unsupervised models describe patterns without a future-return label. They can compress correlated features or identify similar instruments, but similarity is not an expected-return forecast.
Symbols, units & horizon
- X_c: training-centred n×p feature matrix
- U,Σ,V: singular-value decomposition
- V_k: first k orthonormal loading columns
- Z: n×k component scores
- X̂: rank-k reconstruction of centred data
- T: transpose
- k: selected representation dimension
- Original units: restored by adding the training mean
When and why to use this
Use learned representations to reduce noise or organise a universe, then test their incremental effect on the portfolio task.
PCA finds directions of large training variation; clustering groups instruments according to a chosen distance. These tools can support factor removal, pair-candidate selection, risk grouping or feature compression. Their economic meaning depends on preprocessing, universe and distance metric.
Standardise training returns before correlation-oriented clustering, and distinguish that task from covariance-oriented risk modelling. A high-variance component may describe a common risk you want to hedge rather than profitable information. Refit on a causal schedule and track changing group membership.
Autoencoders learn nonlinear compression, while anomaly scores measure reconstruction difficulty. Neither automatically identifies mispricing: an outlier can be a genuine structural break. Use a supervised or economic evaluation of how the representation changes the strategy, with the representation itself fitted inside the fold.
Unsupervised learning, clusters and latent risk structure
- Centre columns using training means and compute a singular-value decomposition.
- Project each row onto the first k loading directions by multiplying by V_k.
- Reconstruct by mapping scores back through V_kᵀ. The difference X_c−X̂ is information omitted by this linear representation, not automatically tradable alpha.
For centred rows [1,1] and [−1,−1], one loading direction is (1,1)/√2. One component reconstructs both rows exactly. Its sign can flip without changing the represented subspace.
Apply it in a strategy
- Define whether similarity means correlated returns, fundamentals, liquidity or another economic relation.
- Fit the representation on formation data and apply it unchanged to the next evaluation block.
- Compare candidate selection or risk estimates with and without the representation, accounting for changing membership and turnover.
Research deliverable
Document loadings or clusters over time and evaluate the strategy outcome they are intended to improve.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
import numpy as np
def pca_fit_transform(train,future,k):
x=np.asarray(train,dtype=float); future=np.asarray(future,dtype=float)
if x.ndim!=2 or not 1<=k<=min(x.shape): raise ValueError("Invalid component count")
center=x.mean(axis=0)
_,_,vt=np.linalg.svd(x-center,full_matrices=False)
loading=vt[:k].T
scores=(future-center)@loading
return scores,scores@loading.T+center
print(pca_fit_transform([[1,1],[-1,-1]],[[2,2]],1)[1])Continue learning
Machine Learning for Quantitative Strategy Development — all lessons- Choose the model’s job: targets, horizons and decision layers
- Feature engineering, missingness and training-only transformations
- Regularised regression: an interpretable alpha baseline
- Logistic classification and cost-aware entry thresholds
- Trees and boosting: nonlinear interactions with controlled complexity
- Calibration, meta-labels and conditional payoff estimation
- Unsupervised learning, clusters and latent risk structure
- From model forecasts to a constrained strategy
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations