Free lesson · Putting it all together
5 / Add ML, regime models or RL only at a defined interface
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A model should replace one well-defined estimate or decision rule. Keep the surrounding experiment fixed so you can tell whether that replacement actually helped.
Symbols, units & horizon
- r_candidate,t: net account return of candidate in interval t
- r_baseline,t: matched baseline return in the same interval
- n: paired evaluation intervals
- Δr̄: average incremental net account return
- both series: same capital, calendar, costs and risk convention
When and why to use this
Use paired ablations to decide whether a specific model improves the system at its chosen interface.
A supervised classifier could replace the assumed win probability in the example. Its target, horizon and cost interpretation must match the payoff model. A volatility model could replace the risk estimate used for sizing. A hidden-state filter could condition either estimate. These are separate experiments, not reasons to change all components together.
Deep learning is a function-approximation choice inside one of those roles. RL can choose sequential execution or inventory actions when today’s action changes future possibilities. For this one-session fixed-entry/fixed-exit baseline, adding RL is not automatically justified.
For every candidate, fit on past data, tune on separate chronological validation windows, and evaluate the complete selection rule on untouched later data. Record trials, random seeds, compute budget and features discarded. A parameter plateau is more persuasive than an isolated optimal point.
Compare incremental net results at matched risk and costs, and report uncertainty with methods appropriate to dependence. A better loss score can produce worse portfolio returns. A robustness score is evidence about the test process, not a probability of future profit.
5 / Add ML, regime models or RL only at a defined interface
- Align candidate and baseline returns on the same evaluation intervals.
- Subtract baseline from candidate in each interval to remove common variation from the comparison.
- Average the paired differences. This is an effect estimate; uncertainty, dependence and search correction require separate analysis.
Candidate returns [.004,−.002,.003] and baseline [.003,−.001,.001] produce differences [.001,−.001,.002]. Mean increment=.002/3≈.000667, about 6.67 basis points per interval. Three intervals are not meaningful evidence of an edge.
Apply it in a strategy
- Choose one interface: probability, volatility, regime conditioning or sequential action.
- Use the nested selection process and retain simple baselines.
- Compare net paired effects, uncertainty and failure scenarios before accepting added complexity.
Research deliverable
Produce an ablation table with matched baselines and an experiment log including unsuccessful trials.
Mechanics & research · reviewed 12 September 2026
Abstract and metadata reviewed 12 September 2026. The proposed robustness grade uses a reported 359,062-record calibration and synthetic tests, while its stated unseen real-market test finds no significant forward relationship. Full samples, code and claims were not independently verified. Use this as a reminder to distinguish statistical support from a forecast of profit, not as a scoring system adopted by the lab.
Further reading: Santoni et al. · Equity Strategy Backtesting: Luck or Edge? · August 2026 preprint ↗
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def paired_increment(candidate,baseline):
if not candidate or len(candidate)!=len(baseline): raise ValueError("Aligned nonempty returns required")
differences=[a-b for a,b in zip(candidate,baseline)]
return differences,sum(differences)/len(differences)
print(paired_increment([.004,-.002,.003],[.003,-.001,.001]))Continue learning
Putting It All Together: Build a Complete Trading Research System — all lessons- 1 / Define the job and a small research contract
- 2 / Make a point-in-time data contract
- 3 / Turn an idea into a causal feature and a baseline
- 4 / Replay decisions into fills and net returns
- 5 / Add ML, regime models or RL only at a defined interface
- 6 / Convert forecasts into constrained portfolio positions
- 7 / Turn targets into idempotent orders and handle partial fills
- 8 / Reconcile fills, costs, cash and performance
- 9 / Run the miniature system and define the promotion decision
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations