Free lesson · Reinforcement learning
Simulator transfer, regime randomisation and RL research promotion
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A policy learns the environment it is given. If simulated fills, impact or other traders are unrealistic, the policy may become excellent at a market that does not exist.
Symbols, units & horizon
- k: scenario index
- K: number of scenarios
- p_k: declared scenario weights summing to one
- G: episode return
- E[G|k]: mean return in scenario k
- Var_k: weighted dispersion of scenario means
- λ: penalty scaling
- J_robust: illustrative robustness objective, not a worst-case guarantee
When and why to use this
Use scenario-based evaluation to identify simulator assumptions that drive policy success and to compare robustness at matched economic constraints.
Randomise plausible volatility, arrival intensity, spread, cancellation, latency and regime duration during training. Reserve different combinations and more severe but defensible scenarios for evaluation. Randomisation should reflect a documented uncertainty range, not an arbitrary distribution selected to flatter the policy.
Use walk-forward market periods for historical components and independent simulator seeds/scenarios for generated dynamics. Compare with fixed and adaptive non-RL baselines at equal inventory, participation and risk limits. Report tail outcomes and constraint violations, not only mean reward.
Fine-tuning on weak scenarios can improve robustness, but those scenarios become training information. Hold out additional conditions to evaluate the adaptation procedure. Recent market-making research provides useful examples of inventory saturation and regime-aware policies, while remaining conditional on its simulated market assumptions.
Simulator transfer, regime randomisation and RL research promotion
- Estimate the policy’s mean episode return separately in each scenario using multiple random seeds.
- Weight those means by a prespecified probability or evaluation mixture.
- Subtract λ times their weighted squared deviations around the mixture mean. This discourages uneven scenario results but can still tolerate a severe low-probability loss.
Scenario means [2,−1] with equal weights have mean .5 and variance 2.25. With λ=.2, score=.5−.45=.05. A constant .4 in both scenarios scores .4 under this illustrative preference.
Apply it in a strategy
- Create an environment uncertainty register and separate training randomisation from held-out stresses.
- Evaluate matched-risk baselines across seeds, regimes and terminal conditions; inspect worst episodes.
- Move to shadow decisions with independent limits and measured simulator-to-observation discrepancies before considering broader operation.
Research deliverable
Submit an RL promotion dossier with environment specification, baseline comparison, support limits, scenario tails and reproducible worst-case trajectories.
Research checkpoint · reviewed 11 September 2026
These sources inform the questions to test. A result is conditional on its data, simulator and evaluation design. The examples in this module are teaching calculations, not reproductions of the reported experiments.
Richer action spaces. This preprint describes market/limit allocations across price levels, set-based order representations and potential-based shaping. Abstract and metadata reviewed. Results are illustrated in three simulated environments with noise, tactical and strategic traders; no historical deployment sample is stated. Its relevance is to action-space and environment design. It does not establish real-market profitability or resolve unsupported offline actions.
Further reading: Cheridito & Weiss · Multi-Level Market Making with Reinforcement Learning · 18 August 2026 ↗
The HFT section reviews a 10 September 2026 preprint on regime-aware market-making adaptation and records its simulator restrictions. Compare its training stresses with held-out regimes; once a scenario is used for fine-tuning, it is part of the development information set.
Further reading: Regime-switching market-making research checkpoint ↗
Research sources, review dates and limitations
Extend the research question
State the simulator’s action, fill and terminal-reward rules. Identify which actions have inadequate historical support and which constraints must be enforced outside the learned policy.
Continue with the connected research module →
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def scenario_score(means,probabilities,penalty):
if not means or len(means)!=len(probabilities) or any(p<0 for p in probabilities) or abs(sum(probabilities)-1)>1e-9 or penalty<0: raise ValueError("Valid scenario mixture required")
mean=sum(p*x for p,x in zip(probabilities,means))
variance=sum(p*(x-mean)**2 for p,x in zip(probabilities,means))
return mean,variance,mean-penalty*variance
print(scenario_score([2,-1],[.5,.5],.2))Continue learning
Reinforcement Learning for Trading & Execution — all lessons- Before RL: one action, one outcome and learning an average
- Choose an RL task: states, actions and the environment
- Bellman recursion and Q-learning
- From a value table to linear features and LSTD
- Policy gradients, actor–critic and PPO clipping
- Reward design, inventory penalties and terminal accounting
- Offline RL, logged actions and support limits
- Simulator transfer, regime randomisation and RL research promotion
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations