Free module · Machine learning & adaptive decisions
Reinforcement Learning for Trading & Execution
Model decisions through time, design the environment carefully, and test the policy’s assumptions.
MDPs · Bellman equations · Q-learning · policy gradients · offline evaluation
The building blocks
Start with a small choice: trade now or wait. An agent chooses an action, observes its consequence, and updates a decision rule. Later actions may depend on the inventory left by earlier ones.
- Compare fixed actions and average observed rewards
- Add state, future consequences and value updates
- Study function approximation and policy optimization
Lessons in this module
- Before RL: one action, one outcome and learning an average
- Choose an RL task: states, actions and the environment
- Bellman recursion and Q-learning
- From a value table to linear features and LSTD
- Policy gradients, actor–critic and PPO clipping
- Reward design, inventory penalties and terminal accounting
- Offline RL, logged actions and support limits
- Simulator transfer, regime randomisation and RL research promotion
Practice and apply
- Discount an episode — Rewards 1,2,−1; γ=.9.
- Update one action value — Q=2, reward=.5, γ=.9, best next Q=3, learning rate .1.
- Update an action average — Three rewards average 2. The fourth reward is −2. Find the new average.
- Solve a scalar value model — Current features [1,1], next [1,0], rewards [1,2], γ=.5. Solve θ=b/A.
Work through the practice exercises · Quant development tools