Trading Dev AcademyFree quant education

Python for quant trading: resources & libraries

Build a workflow you can explain, test and reproduce. Python 3.10+; NumPy for the first two recipes. All examples use synthetic inputs to teach mechanics.

Open the interactive developer reference · Download the Python recipes

Libraries by responsibility

ResponsibilityLibraryUseImplementation check
Arrays & numerical workNumPyVectorized returns, exposures and covariance arrays.Check shapes, finite values, precision and whether broadcasting is intentional.
Time-indexed datapandasAlign bars, events, corporate actions and feature tables.Record availability timestamps; a backward join alone cannot repair revised data.
Statistical modelsstatsmodelsRegression, time-series benchmarks and diagnostics.Check residual dependence and parameter stability before interpreting uncertainty.
Validation & predictionscikit-learnPipelines and chronological train/test splits with a gap.Fit transformations within training folds. A fixed gap is not a general label-overlap purge.
Constrained portfoliosCVXPYExpress convex objectives, exposure bounds and turnover penalties.Check solver status, feasibility and estimated-input sensitivity.
Instrument valuationQuantLibCash-flow schedules, curves and derivative pricing.Calendars, day counts and calibration conventions are part of the model.
Event-driven simulationBacktraderModel bars, orders, fills and a broker ledger.Verify fill timing and fee behavior with a hand-reconciled fixture.
Exchange integrationCCXTA shared interface for supported exchange market data and trading APIs.Inspect venue-specific precision, limits and order semantics; start with recorded data or a sandbox.
Concurrent I/OasyncioCoordinate streams, bounded queues and asynchronous I/O.Blocking calculations still block the event loop; define timeouts and backpressure.

1. Reconcile a causal signal and its costs

A forecast becomes a position only after the information exists. This small ledger earns each return with the preceding position and charges each exposure change once. It is a return approximation with constant capital, not a self-financing multi-asset execution engine.

t: daily interval index; rₜ: asset simple return over interval t, decimal/day; sₜ: target exposure decided at the end of t, dimensionless; wₜ=sₜ₋₁: exposure held during t; c: cost per unit of absolute exposure change, decimal; nₜ: net strategy return, decimal/day. Initial exposure is zero.

wt=st−1,nt=wtrt−c|wt−wt−1|
Derivation under the stated assumptions

Work it by hand

  1. Start with notional exposure divided by capital: wt. Multiplying by the interval return gives gross return gt=wtrt.
  2. Turnover counts purchases and sales positively: τt=|wt−wt−1|. Under the proportional-cost assumption the return cost is cτt.
  3. Subtract cost from gross return: nt=gt−cτt. Lag the decision before this multiplication; do not lag the market return.
Work it by hand

Use returns [0.01, −0.02, 0.03], end-of-day signals [1, 0, 1] and cost 10 basis points = 0.001. Held exposures are [0, 1, 0]. Net returns are [0, −0.021, −0.001]: the second interval loses 2% plus 0.1% entry cost, and the third pays 0.1% exit cost.

Useful for independently checking a backtester’s timing and fee ledger. Assumes immediate fills at the interval boundary, no financing, no market impact and no intraperiod rebalancing. Real delayed fills can invalidate the result.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

import numpy as np

def causal_net_returns(returns, signals, cost_bps=10.0):
    """1-D daily decimal returns; end-of-day unit exposures; one-way bps."""
    r, s = np.asarray(returns, float), np.asarray(signals, float)
    if r.ndim != 1 or r.size == 0 or r.shape != s.shape:
        raise ValueError("Use equally sized nonempty 1-D arrays")
    if not (np.isfinite(r).all() and np.isfinite(s).all()):
        raise ValueError("Inputs must be finite")
    if not np.isfinite(cost_bps) or cost_bps < 0:
        raise ValueError("Costs must be finite and nonnegative")
    held = np.r_[0.0, s[:-1]]
    turnover = np.abs(np.diff(np.r_[0.0, held]))
    return held * r - cost_bps / 10_000 * turnover

result = causal_net_returns([.01, -.02, .03], [1, 0, 1])
np.testing.assert_allclose(result, [0, -.021, -.001])
print(result)  # [ 0.    -0.021 -0.001]

2. Calculate portfolio risk without hiding the covariance

Portfolio risk depends on how returns move together. Two individually volatile holdings can diversify one another, but only while the estimated dependence remains relevant.

w: vector of capital weights, dimensionless; Σ: symmetric positive semidefinite covariance matrix of daily decimal returns, units decimal²/day; v: portfolio daily return variance; σ: daily return standard deviation. Negative weights mean short exposures.

v=w⊤Σw=∑i∑jwiwjΣij,σ=v
Derivation under the stated assumptions

Work it by hand

  1. The portfolio return is the weighted sum rp=∑iwiri. Subtract its mean and square the sum.
  2. Expanding the square gives (rp−E[rp])2=∑i∑jwiwj(ri−E[ri])(rj−E[rj]).
  3. Take expectations. Each expected product is Σij, giving v=w⊤Σw. Take the nonnegative square root to return to decimal-return units.
Work it by hand

For w=[0.5,0.5] and Σ=[[0.0004,0.0001],[0.0001,0.0009]], variance is 0.25×0.0004 + 2×0.25×0.0001 + 0.25×0.0009 = 0.000375. Daily volatility is √0.000375 ≈ 0.019365, or 1.9365%.

Useful for exposure checks and risk budgets. This identity is exact for fixed weights and a valid covariance matrix; an estimated covariance is uncertain. A historical estimate can fail when correlations change.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

import numpy as np

def portfolio_daily_risk(weights, covariance):
    """Return (daily variance, daily volatility); daily decimal covariance."""
    w, cov = np.asarray(weights, float), np.asarray(covariance, float)
    if w.ndim != 1 or not w.size or cov.shape != (w.size, w.size):
        raise ValueError("Covariance must match a nonempty weight vector")
    if not (np.isfinite(w).all() and np.isfinite(cov).all()):
        raise ValueError("Inputs must be finite")
    if not np.allclose(cov, cov.T, rtol=0, atol=1e-12):
        raise ValueError("Covariance must be symmetric")
    if np.linalg.eigvalsh(cov).min() < -1e-12:
        raise ValueError("Covariance must be positive semidefinite")
    variance = max(0.0, float(w @ cov @ w))
    return variance, variance ** .5

v, sigma = portfolio_daily_risk([.5, .5], [[.0004, .0001], [.0001, .0009]])
assert abs(v - .000375) < 1e-12
print(round(sigma, 6))  # 0.019365

3. Generate chronological validation windows

A walk-forward split trains on earlier observations and tests on later ones. An explicit gap separates the two. Choose the gap from the feature and label availability timeline, rather than copying a default.

Useful for reproducible model comparisons. Equal row spacing is assumed when comparing windows by duration. This example does not purge variable-length overlapping labels; implement that separately from their actual end timestamps.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def walk_forward(n_rows, min_train, test_size, gap=0):
    """Yield expanding train and fixed test index ranges, measured in rows."""
    values = (n_rows, min_train, test_size, gap)
    if any(type(x) is not int for x in values):
        raise ValueError("All sizes must be integers")
    if min(n_rows, min_train, test_size) < 1 or gap < 0:
        raise ValueError("Positive sizes and a nonnegative gap are required")
    for start in range(min_train + gap, n_rows - test_size + 1, test_size):
        yield range(start - gap), range(start, start + test_size)

folds = [(list(train), list(test)) for train, test in walk_forward(10, 4, 2, 1)]
assert folds == [([0, 1, 2, 3], [5, 6]), ([0, 1, 2, 3, 4, 5], [7, 8])]
print(folds)
# Fit scalers and models on each train range only.
# The incomplete final test window is deliberately excluded.

4. Make fill replay idempotent

A stream may deliver the same fill twice after a reconnect. Deduplicate by an immutable venue fill identifier before updating inventory. Keep each execution’s signed quantity; an order may have several partial fills.

Useful as an offline replay fixture for automation. Assumes fill identifiers are globally unique in this input; production keys often need venue and account as well. Persist deduplication and inventory atomically, and reconcile against broker records. This is a runnable in-memory exercise, not a live order service.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

from decimal import Decimal

def replay_inventory(fills):
    """(unique fill id, signed quantity string) pairs; output in base units."""
    seen, inventory = set(), Decimal("0")
    for fill_id, signed_quantity in fills:
        if not isinstance(fill_id, str) or not fill_id:
            raise ValueError("Each fill needs a nonempty string id")
        quantity = Decimal(signed_quantity)
        if not quantity.is_finite():
            raise ValueError("Fill quantity must be finite")
        if fill_id not in seen:
            inventory += quantity
            seen.add(fill_id)
    return inventory

fills = [("fill-a", "0.2"), ("fill-a", "0.2"), ("fill-b", "-0.05")]
assert replay_inventory(fills) == Decimal("0.15")
print(replay_inventory(fills))  # 0.15 base units

From research to automation

Record availability timestamps, test a causal baseline, validate chronologically, simulate partial fills and costs, and reconcile positions and cash. Profile representative workloads before vectorizing calculations or batching I/O.

Algorithm design · Strategy validation · Execution models

Research checkpoint

Reviewed 8 October 2026. Yin et al., Implementation Risk in Portfolio Backtesting: arXiv working paper submitted to Financial Innovation; abstract-only review. US equities, 180 S&P 500 stocks; historical data dates unspecified in the abstract. Engine and cost conventions affect comparisons. No independent replication here. Recency or numerical accuracy does not establish profitable implementation.