Trading Dev AcademyFree quant education

Free lesson · Linear algebra

Eigenvalues: where the risk actually lives

Open interactive lessonPractice calculationsExplore labs

Start with the idea

An eigenvector is a direction whose orientation a covariance matrix preserves; the eigenvalue measures variance along that unit direction. Principal components rotate correlated risks into uncorrelated directions in the fitted sample.

Symbols, units & horizon
  • Σ: symmetric covariance matrix
  • v_k: eigenvector for component k, a direction in asset space
  • λ_k (lambda): its eigenvalue in squared return units
  • tr: trace, sum of diagonal entries
  • k: component index
  • I: identity matrix
  • det: determinant
  • a,b: diagonal entries and c: off-diagonal entry of the two-by-two matrix
  • Q: matrix of unit eigenvectors
  • Λ: diagonal matrix of eigenvalues
  • ||v||: vector length

When and why to use this

Use PCA to diagnose a common market driver, compress correlated features or reveal concentrated risk hidden behind many holdings.

A covariance matrix describes an ellipsoid of risk. Its eigenvectors are the axes of that ellipsoid — the independent directions the portfolio can move in — and the eigenvalues are how much variance sits along each axis.

Σ𝐯k=λk𝐯k,∑kλk=tr⁡(Σ)=total variance
Algebra and arithmetic

Solve a two-by-two eigenproblem

  1. Nonzero v in (Σ−λI)v=0 requires det⁡(Σ−λI)=0. For Σ=(accb), expand to λ2−(a+b)λ+ab−c2=0.
  2. The quadratic formula gives λ±=[a+b±(a−b)2+4c2]2. Their sum is a+b, the trace. Solve one row for each eigenvector and normalise its length to 1.
  3. For symmetric Σ, Σ=QΛQT. A unit component v has variance vTΣv=λ. More generally portfolio variance is ∑kλk(wTvk)2.
Work it by hand

For Σ=(.04.02.02.04), eigenvalues .06,.02 sum to .08. Unit eigenvectors are (1,1)/√2 and (1,−1)/√2; the first accounts for 75% of total variance.

In a 500-stock universe you might expect 500 independent sources of risk. In practice the first eigenvector (roughly "everything goes up together" — the market) carries 30–50% of total variance, the next few are sectors and styles, and the remaining hundreds are idiosyncratic noise. A "diversified" long book of 500 names is very often one bet, sized 500 times.

PCA and SVD: compressing 50 indicators into 5 factors

Principal component analysis is eigen-decomposition applied to the correlation matrix of your data. Order the eigenvalues; keep the eigenvectors whose eigenvalues together explain, say, 90% of variance; project the data onto them. Fifty correlated indicators — RSI, three moving averages, MACD, five volatility measures — collapse into a handful of uncorrelated factors, which is what a model can actually learn from.

The singular value decomposition X=USV⊤ does the same job directly on the data matrix without forming Σ first; it is numerically better behaved and it is what libraries call under the hood.

Eigenvalues of an 8-asset correlation matrix: 5.2, 1.1, 0.6, 0.4, 0.3, 0.2, 0.1, 0.1. Fraction of variance in PC1?

They sum to 8 (the trace of a correlation matrix is N). 5.28=0.65. Two-thirds of the movement in eight "different" assets is one factor.

Why does reducing 50 indicators to 5 principal components usually improve a predictive model rather than hurt it?

Because the 50 indicators are largely redundant re-expressions of the same few underlying quantities (trend, volatility, volume). Feeding all 50 to a model gives it 45 dimensions of noise to overfit and makes coefficients unstable (multicollinearity). PCA keeps the directions with signal variance and discards the rest, so the model sees fewer, cleaner, uncorrelated inputs.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

import numpy as np  # dependency: numpy

def principal_components(covariance):
    """Symmetric covariance; columns of vectors match descending eigenvalues."""
    cov = np.asarray(covariance, dtype=float)
    if not np.allclose(cov, cov.T):
        raise ValueError("Symmetric matrix required")
    values, vectors = np.linalg.eigh(cov)
    order = np.argsort(values)[::-1]
    return values[order], vectors[:, order], np.trace(cov)

print(principal_components([[.04,.02],[.02,.04]]))

Continue learning

Linear Algebra — all lessons
  1. Start with lists: addition, scaling and a dot product
  2. Your portfolio is a vector; your risk is a matrix
  3. Eigenvalues: where the risk actually lives
  4. Regression, conditioning, and regularisation

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations