Trading Dev AcademyFree quant education

Free module · Machine learning & adaptive decisions

Reinforcement Learning for Trading & Execution

Model decisions through time, design the environment carefully, and test the policy’s assumptions.

MDPs · Bellman equations · Q-learning · policy gradients · offline evaluation

The building blocks

Start with a small choice: trade now or wait. An agent chooses an action, observes its consequence, and updates a decision rule. Later actions may depend on the inventory left by earlier ones.

  • Compare fixed actions and average observed rewards
  • Add state, future consequences and value updates
  • Study function approximation and policy optimization

Lessons in this module

  1. Before RL: one action, one outcome and learning an average
  2. Choose an RL task: states, actions and the environment
  3. Bellman recursion and Q-learning
  4. From a value table to linear features and LSTD
  5. Policy gradients, actor–critic and PPO clipping
  6. Reward design, inventory penalties and terminal accounting
  7. Offline RL, logged actions and support limits
  8. Simulator transfer, regime randomisation and RL research promotion

Open the interactive module

Practice and apply

  • Discount an episode — Rewards 1,2,−1; γ=.9.
  • Update one action value — Q=2, reward=.5, γ=.9, best next Q=3, learning rate .1.
  • Update an action average — Three rewards average 2. The fourth reward is −2. Find the new average.
  • Solve a scalar value model — Current features [1,1], next [1,0], rewards [1,2], γ=.5. Solve θ=b/A.

Work through the practice exercises · Quant development tools