Free lesson · Statistics
Correlation: how many strategies do you really have?
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Covariance measures whether deviations tend to share a sign; correlation removes their measurement scale. It describes linear co-movement in a sample. Diversification reduces variance through the cross terms, but can disappear when dependence changes during stress.
Symbols, units & horizon
- xᵢ, yᵢ: paired observations
- x̄, ȳ: sample means
- n: pair count
- Cov: sample covariance
- sₓ, sᵧ: sample standard deviations
- ρ (rho): dimensionless correlation between −1 and 1
- Sxy: sum of cross-products of centred observations
- Sxx, Syy: sums of squared centred observations
- σ: common asset volatility in return units
- σ_b: equal-weight blend volatility
- Var: variance
- X,Y: paired random return variables, with σ_X and σ_Y denoting their SDs (estimated by sₓ,sᵧ in sample calculations)
- a,b,ε in y=a+bx+ε: intercept, slope and residual
- N: number of equally weighted, equal-volatility streams in the diversification example
When and why to use this
Use correlation for an initial redundancy check, then use covariance with actual weights for allocation. Compare normal periods with stressed co-movement before calling a second strategy a hedge.
Correlation runs from to and measures how much two return streams move together, after removing their scale.
Cancel the sample denominators in correlation
- Write , , .
- Dividing covariance by cancels n−1 and gives . Zero variance makes correlation undefined.
x=(1,2,3), y=(2,4,6) give Sxy=4, Sxx=2, Syy=8; ρ=4/√16=1. Reversing y gives −1.
When you run several strategies, their combined volatility is not the average of their volatilities. For two equal-weight streams with the same :
Expand the equal-weight blend
- For equal vol σ and weights ½, .
- Factor out : variance is . Take the square root. To solve for correlation from blend vol, rearrange to .
σ=20%, ρ=.5 gives blend vol = .2√.75=17.32%, not 10%.
| ρ | blend vol / single vol | what it means |
|---|---|---|
| +1.0 | 100% | one strategy, two names |
| +0.9 | 97% | one strategy, two names, plus paperwork |
| +0.5 | 87% | some diversification |
| 0.0 | 71% | uncorrelated — volatility falls by √2 |
| −0.5 | 50% | a hedge that still earns |
You add a second, independent () strategy with identical EV and , 50/50 weight. Compared with running one, EV per unit of risk goes…
EV of the blend is the same; drops to . EV/σ rises by . The Sharpe ratio scales with independent bets — this is the whole argument for diversification, and why it fails when the bets are not independent.
Regression: correlation with a slope
Linear regression fits . In trading the slope shows up as a hedge ratio (how much of X to hold against Y), as beta (sensitivity to the market), and as the core of mean-reversion signals (regress spread on time or on its own lag). Logistic regression does the same for a yes/no target — will the next bar close up? — producing a probability instead of a level, which plugs straight into the EV formula.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
from statistics import mean
from math import sqrt
def correlation(x, y):
if len(x) != len(y) or len(x) < 2:
raise ValueError("Matching samples of at least two required")
dx, dy = [v-mean(x) for v in x], [v-mean(y) for v in y]
denominator = sqrt(sum(v*v for v in dx)*sum(v*v for v in dy))
if denominator == 0:
raise ValueError("Correlation undefined for constant inputs")
return sum(a*b for a,b in zip(dx,dy))/denominator
def equal_weight_volatility(sigma, rho):
return sigma*sqrt((1+rho)/2)
print(correlation([1,2,3], [2,4,6]), equal_weight_volatility(.2, .5))
from statistics import linear_regression
def hedge_regression(x_returns, y_returns):
slope, intercept = linear_regression(x_returns, y_returns)
return intercept, slope
Continue learning
Statistics & Probability — all lessons- Start with counts, probabilities and averages
- Random outcomes, sample averages and the limits of the bell curve
- Conditional probability: update a belief with evidence
- Prediction markets: probability, price and net expected value
- Expected value: measure the payoff before choosing the risk
- Correlation: how many strategies do you really have?
- Conditional probability and Bayes: where the edge actually lives
- The central limit theorem: sampling means under explicit assumptions
- Estimation uncertainty and Bayesian updating
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations