Free lesson · Start here
The six numbers your first backtest must produce
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A trade list answers a different question from an investor equity curve. Trade statistics describe realised payoffs; daily portfolio returns describe the capital exposed through time. Reconcile the two before comparing strategies. A strategy can have positive average trade P&L and still lose money at the portfolio level if financing, overlapping exposure or missing trades are ignored.
Symbols, units & horizon
- xᵢ: currency P&L of trade i
- n: trade count
- x̄, median: average and middle ordered trade P&L
- s: sample standard deviation of trade P&L
- SE: standard error of the mean under IID sampling
- t: mean divided by SE, dimensionless test statistic
- μ: population mean
- σ: population standard deviation
- E[X]: probability-weighted average payoff
- p: win probability
- W, L: positive average win and loss magnitudes
- SR: Sharpe ratio of periodic excess returns
- V, peak: wealth and its running high
- DD: peak-relative fractional drawdown
- √: square root
- IID: independent and identically distributed
When and why to use this
Use this checklist at the first research review. It tells you whether to investigate the mean, the typical trade, sampling uncertainty or the path of losses next. Keep the starting capital and cash-flow conventions visible.
The curriculum connects market foundations, quantitative methods, strategy development and fund operations. Begin with these performance measures, then use the later modules to understand their assumptions. No single statistic is a pass/fail rule for a fund.
| # | number | formula | what it decides |
|---|---|---|---|
| 1 | mean trade (EV) | is there anything here at all | |
| 2 | median trade | middle of the sorted P&Ls | what a typical trade feels like; skew |
| 3 | standard deviation | size, and how noisy the EV estimate is | |
| 4 | t-statistic | mean relative to sampling uncertainty | |
| 5 | Sharpe ratio | comparable quality across strategies | |
| 6 | max drawdown | largest peak-to-trough fall | whether you would survive trading it |
Then one relationship: the correlation of this strategy's daily returns with anything else you run, or with the market. High correlation can limit diversification, but no fixed threshold determines value. Evaluate marginal expected return, factor exposure, tail dependence, capacity and costs.
A minimal, honest backtest loop
- Define the rule before looking at the data you will test it on. Write it down.
- Choose chronological training, validation and final test periods appropriate to the data and holding horizon. Fit only on training data and keep the final test independent.
- Subtract realistic costs: spread + commission + slippage on every fill. For most retail strategies this is the step that ends the project.
- Produce the six numbers on the out-of-sample set.
- Report observed drawdown and simulated drawdown distributions. Preserve serial dependence where relevant, and state the model limitations.
- Size with a fraction of Kelly or a vol target such that that drawdown is survivable.
Your 80-trade out-of-sample test shows EV +$25, σ $200. Is the t-stat above 2?
SE = 200/√80 = 22.4; t = 25/22.4 = 1.1. Under the IID approximation with unchanged estimates, about 256 trades would give t ≈ 2. This is a calculation, not a stopping rule or evidence that future trades will have the same mean.
Why is out-of-sample evaluation non-negotiable, in one paragraph you could say to a friend?
Any rule with a few adjustable parameters can be tuned until it fits the past — that is curve fitting, not forecasting. The only way to know whether the rule captured something real is to test it on data it never saw. If it holds up there, the edge might be real; if it doesn't, you have learned that cheaply. Reporting in-sample results is reporting how well you memorised the answer key.
Research sources, review dates and limitations
Extend the research question
Choose one narrow deliverable: a reproducible calculation, an evidence critique or an implementation-feasibility memo. A reasoned rejection is a valid outcome.
Continue with the connected research module →
Six statistics, one hand-calculation sheet
- For trades +30, −10, +20, the sum is 40 and mean is . Grouping wins and losses gives the same value: . Sort to get median 20.
- Deviations are 16.667, −23.333, 6.667. Their squared sum is 866.667. Sample variance is , so .
- Under IID sampling, ; divide the mean by SE to obtain . Rearranging gives , only if those estimates and assumptions stay fixed.
- For daily excess returns, divide their mean by their sample SD and multiply by under the IID annualisation approximation. Trade P&Ls alone cannot determine daily Sharpe.
- Start capital at 100: the cumulative path is 100,130,120,140. Running peaks are 100,130,130,140. Maximum percentage drawdown is .
A 7.692% drawdown and a $10 drawdown describe the same path in different units. Never divide a dollar trade mean by a percentage-return SD.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
from statistics import mean, median, stdev
from math import sqrt
def trade_summary(pnl):
n = len(pnl)
sd = stdev(pnl)
se = sd / sqrt(n)
return {"mean": mean(pnl), "median": median(pnl), "sd": sd,
"se_iid": se, "t_iid": mean(pnl)/se if se else None}
def sharpe(excess_returns, periods=252):
return mean(excess_returns) / stdev(excess_returns) * sqrt(periods)
def max_drawdown(wealth):
peak, worst = wealth[0], 0
if peak <= 0:
raise ValueError("Positive starting wealth required")
for value in wealth:
peak = max(peak, value)
worst = max(worst, (peak-value)/peak)
return worst
print(trade_summary([30, -10, 20]))
print(max_drawdown([100, 130, 120, 140]))Continue learning
Fund Lab Orientation — all lessonsQuantitative finance and development glossary · Python resources and libraries · Research sources and limitations