Free lesson · Research & backtests
Build a point-in-time dataset
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A reproducible backtest is an information-timing experiment. “Data from March” does not mean “available in March.” Reconstruct the exact sequence of publication, computation and executable prices before evaluating a feature.
Symbols, units & horizon
- t: decision time
- 𝓕ₜ: information available by t, including publication and processing delays
- ⊆: contained in, no future information allowed
- featuresₜ: inputs built only from available data
- f: specified decision rule
- positionₜ₊₁: next-period exposure based on the decision at t
- aⱼ: availability timestamp of record j
- {j:aⱼ≤t}: set of record indices available by decision time t
When and why to use this
Use an availability timestamp in every join and audit it when interpreting unusually strong predictions. This is especially useful for earnings, macro releases, revisions and index changes.
Each record needs two clocks: when an event happened and when your strategy could have known it. A financial statement dated in March may not have been published until May. Joining on the statement period leaks future information into the simulated decision.
Audit the information-to-position mapping
- Let aⱼ be the availability time of record j. The input set at decision time t is . A feature must be a function only of that set.
- If the rule uses the completed bar at t, its next-period position is f(featuresₜ). P&L then pairs that position with the return after its executable entry, not the return that produced the feature.
A release becomes public at 10:00:00, arrives at 10:00:01 and needs 0.2 seconds to process. A simulated fill at 09:59:59 violates the information constraint.
is the information available at decision time. This timing convention assumes a signal observed at t is implemented for the next holding period; actual implementation must specify the first executable quote after computation and order latency.
- Keep delisted securities and historical index membership; a universe of today’s survivors overstates opportunity.
- Version raw data, adjustment rules, calendars, time zones, and missing-value decisions.
- Align stock total returns, borrow fees, futures rolls, and currency conversion with the economic position.
- Fit scalers, feature selection, and missing-value models on training data only. Freeze the pipeline for evaluation.
Research sources, review dates and limitations
Capstone checkpoint 3 / Audit the information
Synthetic exercise · self-assessed. Use labels available at times 8, 10 and 12, a fitting cutoff of 10, decision time 11 and outcome time 12. Identify allowed labels, then describe what changes if the time-10 label arrives late.
Save in your practice notes: A timeline and an allowed-input list. Open practice studio →
The first two labels qualify only if both have actually arrived by the cutoff. An observation timestamp alone is insufficient.Check your reasoning
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def lag_positions(signals, initial_position=0):
"""Signals become exposure next row. Rows must follow executable timestamps."""
return [initial_position] + list(signals[:-1]) if signals else []
def available_records(records, decision_time):
"""Each record includes its actual available_at timestamp."""
return [row for row in records if row["available_at"] <= decision_time]
print(lag_positions([1, -1, 0])) # [0, 1, -1]Continue learning
Research & Backtest Design — all lessons- Build a point-in-time dataset
- Separate model selection from evaluation
- Account for dependence and multiple experiments
- Make the accounting identity your first test
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations