Free lesson · Putting it all together
2 / Make a point-in-time data contract
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A row of data is usable only when it has become available to the strategy. Record both what period a measurement describes and when the program could actually know it.
Symbols, units & horizon
- a_i: availability time of training label i
- t_fit: model-fitting cutoff
- t_decision: time the model is used
- t_outcome: end time of the future outcome being predicted
- times: one consistent ordered time unit
- inequality: required chronology, not a profitability formula
When and why to use this
Use availability rules to connect market data, feature generation, supervised targets and chronological validation.
For the ETF example, daily close t becomes a feature only after that session has completed and the data has arrived. An order based on that feature belongs to a later executable event. The target is the next session’s open-to-close return; it cannot enter training until that close is available.
Store instrument identity, venue session, timezone, observation timestamp, availability timestamp, price fields, corporate-action conventions and data quality status. Preserve raw records and version transformations. Adjusted close is useful for some return calculations, but it is not necessarily a price at which an order could fill.
Use a chronological split with outcome availability checked at the boundary. Purge labels that overlap evaluation intervals, and fit scaling, imputation and feature selection only inside permitted training windows. A shuffled row split does not respect this contract.
Missing data needs a written rule: skip a decision, stop trading or use a past value only if that choice is valid for the feature. Backfilling a missing price from the future is an information leak. Build a small hand-audited dataset before handling millions of rows.
2 / Make a point-in-time data contract
- Record each training label’s actual availability time, including publication or processing delay.
- Keep only labels available by the fitting cutoff.
- Require the fitting cutoff to precede the decision and the predicted outcome to occur after the decision. Add interval purging when labels overlap test windows.
Labels become available at times [8,10,12]. A model fitted at time 10 can use the first two, but not the third. It can then make a decision at 11 about an outcome ending at 12. Equality at the cutoff assumes the data has actually arrived.
Apply it in a strategy
- Write the time and adjustment schema before fitting a model.
- Audit ten rows by hand, including a missing value and a boundary label.
- Apply the purging and chronological-split workflow to every fitting step.
Research deliverable
Create a data dictionary plus an availability audit that identifies exactly which labels each model version may use.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def available_labels(availability,fit_cutoff,decision_time,outcome_time):
if not fit_cutoff<decision_time<outcome_time: raise ValueError("Fit, decision and outcome must be ordered")
return [i for i,t in enumerate(availability) if t<=fit_cutoff]
print(available_labels([8,10,12],10,11,12))Continue learning
Putting It All Together: Build a Complete Trading Research System — all lessons- 1 / Define the job and a small research contract
- 2 / Make a point-in-time data contract
- 3 / Turn an idea into a causal feature and a baseline
- 4 / Replay decisions into fills and net returns
- 5 / Add ML, regime models or RL only at a defined interface
- 6 / Convert forecasts into constrained portfolio positions
- 7 / Turn targets into idempotent orders and handle partial fills
- 8 / Reconcile fills, costs, cash and performance
- 9 / Run the miniature system and define the promotion decision
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations