Trading Dev AcademyFree quant education

Free lesson · Start here

The six numbers your first backtest must produce

Open interactive lessonPractice calculationsExplore labs

Start with the idea

A trade list answers a different question from an investor equity curve. Trade statistics describe realised payoffs; daily portfolio returns describe the capital exposed through time. Reconcile the two before comparing strategies. A strategy can have positive average trade P&L and still lose money at the portfolio level if financing, overlapping exposure or missing trades are ignored.

Symbols, units & horizon
  • xᵢ: currency P&L of trade i
  • n: trade count
  • x̄, median: average and middle ordered trade P&L
  • s: sample standard deviation of trade P&L
  • SE: standard error of the mean under IID sampling
  • t: mean divided by SE, dimensionless test statistic
  • μ: population mean
  • σ: population standard deviation
  • E[X]: probability-weighted average payoff
  • p: win probability
  • W, L: positive average win and loss magnitudes
  • SR: Sharpe ratio of periodic excess returns
  • V, peak: wealth and its running high
  • DD: peak-relative fractional drawdown
  • √: square root
  • IID: independent and identically distributed

When and why to use this

Use this checklist at the first research review. It tells you whether to investigate the mean, the typical trade, sampling uncertainty or the path of losses next. Keep the starting capital and cash-flow conventions visible.

The curriculum connects market foundations, quantitative methods, strategy development and fund operations. Begin with these performance measures, then use the later modules to understand their assumptions. No single statistic is a pass/fail rule for a fund.

#numberformulawhat it decides
1mean trade (EV)pW−qLis there anything here at all
2median trademiddle of the sorted P&Lswhat a typical trade feels like; skew
3standard deviation∑(xi−x‾)2(n−1)size, and how noisy the EV estimate is
4t-statisticx‾(σn)mean relative to sampling uncertainty
5Sharpe ratioR‾σR×252comparable quality across strategies
6max drawdownlargest peak-to-trough fallwhether you would survive trading it

Then one relationship: the correlation of this strategy's daily returns with anything else you run, or with the market. High correlation can limit diversification, but no fixed threshold determines value. Evaluate marginal expected return, factor exposure, tail dependence, capacity and costs.

A minimal, honest backtest loop

  1. Define the rule before looking at the data you will test it on. Write it down.
  2. Choose chronological training, validation and final test periods appropriate to the data and holding horizon. Fit only on training data and keep the final test independent.
  3. Subtract realistic costs: spread + commission + slippage on every fill. For most retail strategies this is the step that ends the project.
  4. Produce the six numbers on the out-of-sample set.
  5. Report observed drawdown and simulated drawdown distributions. Preserve serial dependence where relevant, and state the model limitations.
  6. Size with a fraction of Kelly or a vol target such that that drawdown is survivable.
Your 80-trade out-of-sample test shows EV +$25, σ $200. Is the t-stat above 2?

SE = 200/√80 = 22.4; t = 25/22.4 = 1.1. Under the IID approximation with unchanged estimates, about 256 trades would give t ≈ 2. This is a calculation, not a stopping rule or evidence that future trades will have the same mean.

Why is out-of-sample evaluation non-negotiable, in one paragraph you could say to a friend?

Any rule with a few adjustable parameters can be tuned until it fits the past — that is curve fitting, not forecasting. The only way to know whether the rule captured something real is to test it on data it never saw. If it holds up there, the edge might be real; if it doesn't, you have learned that cheaply. Reporting in-sample results is reporting how well you memorised the answer key.

Research sources, review dates and limitations

Extend the research question

Choose one narrow deliverable: a reproducible calculation, an evidence critique or an implementation-feasibility memo. A reasoned rejection is a valid outcome.

Continue with the connected research module →

Algebra and arithmetic

Six statistics, one hand-calculation sheet

  1. For trades +30, −10, +20, the sum is 40 and mean is x‾=403=13.333. Grouping wins and losses gives the same value: (23)25−(13)10=13.333. Sort to get median 20.
  2. Deviations are 16.667, −23.333, 6.667. Their squared sum is 866.667. Sample variance is 866.667(3−1)=433.333, so s=20.817.
  3. Under IID sampling, SE=s3=12.019; divide the mean by SE to obtain t=1.109. Rearranging t=x‾ns gives n=(tsx‾)2, only if those estimates and assumptions stay fixed.
  4. For daily excess returns, divide their mean by their sample SD and multiply by 252 under the IID annualisation approximation. Trade P&Ls alone cannot determine daily Sharpe.
  5. Start capital at 100: the cumulative path is 100,130,120,140. Running peaks are 100,130,130,140. Maximum percentage drawdown is (130−120)130=7.692%.
Work it by hand

A 7.692% drawdown and a $10 drawdown describe the same path in different units. Never divide a dollar trade mean by a percentage-return SD.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

from statistics import mean, median, stdev
from math import sqrt

def trade_summary(pnl):
    n = len(pnl)
    sd = stdev(pnl)
    se = sd / sqrt(n)
    return {"mean": mean(pnl), "median": median(pnl), "sd": sd,
            "se_iid": se, "t_iid": mean(pnl)/se if se else None}

def sharpe(excess_returns, periods=252):
    return mean(excess_returns) / stdev(excess_returns) * sqrt(periods)

def max_drawdown(wealth):
    peak, worst = wealth[0], 0
    if peak <= 0:
        raise ValueError("Positive starting wealth required")
    for value in wealth:
        peak = max(peak, value)
        worst = max(worst, (peak-value)/peak)
    return worst

print(trade_summary([30, -10, 20]))
print(max_drawdown([100, 130, 120, 140]))

Continue learning

Fund Lab Orientation — all lessons
  1. The six numbers your first backtest must produce

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations