Trading Dev AcademyFree quant education

Free lesson · Financial machine learning

Unsupervised learning, clusters and latent risk structure

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Unsupervised models describe patterns without a future-return label. They can compress correlated features or identify similar instruments, but similarity is not an expected-return forecast.

Symbols, units & horizon
  • X_c: training-centred n×p feature matrix
  • U,Σ,V: singular-value decomposition
  • V_k: first k orthonormal loading columns
  • Z: n×k component scores
  • X̂: rank-k reconstruction of centred data
  • T: transpose
  • k: selected representation dimension
  • Original units: restored by adding the training mean

When and why to use this

Use learned representations to reduce noise or organise a universe, then test their incremental effect on the portfolio task.

PCA finds directions of large training variation; clustering groups instruments according to a chosen distance. These tools can support factor removal, pair-candidate selection, risk grouping or feature compression. Their economic meaning depends on preprocessing, universe and distance metric.

Standardise training returns before correlation-oriented clustering, and distinguish that task from covariance-oriented risk modelling. A high-variance component may describe a common risk you want to hedge rather than profitable information. Refit on a causal schedule and track changing group membership.

Autoencoders learn nonlinear compression, while anomaly scores measure reconstruction difficulty. Neither automatically identifies mispricing: an outlier can be a genuine structural break. Use a supervised or economic evaluation of how the representation changes the strategy, with the representation itself fitted inside the fold.

Xc=UΣV𝖳,Z=XcVk,X^=ZVk𝖳
Model assumptions, derivation and arithmetic

Unsupervised learning, clusters and latent risk structure

  1. Centre columns using training means and compute a singular-value decomposition.
  2. Project each row onto the first k loading directions by multiplying by V_k.
  3. Reconstruct by mapping scores back through V_kᵀ. The difference X_c−X̂ is information omitted by this linear representation, not automatically tradable alpha.
Work it by hand

For centred rows [1,1] and [−1,−1], one loading direction is (1,1)/√2. One component reconstructs both rows exactly. Its sign can flip without changing the represented subspace.

Apply it in a strategy

  • Define whether similarity means correlated returns, fundamentals, liquidity or another economic relation.
  • Fit the representation on formation data and apply it unchanged to the next evaluation block.
  • Compare candidate selection or risk estimates with and without the representation, accounting for changing membership and turnover.

Research deliverable

Document loadings or clusters over time and evaluate the strategy outcome they are intended to improve.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

import numpy as np

def pca_fit_transform(train,future,k):
    x=np.asarray(train,dtype=float); future=np.asarray(future,dtype=float)
    if x.ndim!=2 or not 1<=k<=min(x.shape): raise ValueError("Invalid component count")
    center=x.mean(axis=0)
    _,_,vt=np.linalg.svd(x-center,full_matrices=False)
    loading=vt[:k].T
    scores=(future-center)@loading
    return scores,scores@loading.T+center

print(pca_fit_transform([[1,1],[-1,-1]],[[2,2]],1)[1])

Continue learning

Machine Learning for Quantitative Strategy Development — all lessons
  1. Choose the model’s job: targets, horizons and decision layers
  2. Feature engineering, missingness and training-only transformations
  3. Regularised regression: an interpretable alpha baseline
  4. Logistic classification and cost-aware entry thresholds
  5. Trees and boosting: nonlinear interactions with controlled complexity
  6. Calibration, meta-labels and conditional payoff estimation
  7. Unsupervised learning, clusters and latent risk structure
  8. From model forecasts to a constrained strategy

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations