Trading Dev AcademyFree quant education

Free lesson · Reinforcement learning

Bellman recursion and Q-learning

Open interactive lessonPractice calculationsExplore labs

Start with the idea

A value function estimates the reward still available from a state. A Q-function estimates the reward from taking one particular action and then following a policy or an optimal continuation.

Symbols, units & horizon
  • Q(s,a): estimated action value
  • s′: next state
  • a′: candidate next action
  • r: observed immediate reward
  • γ: discount factor
  • η: learning step size in (0,1]
  • max: largest feasible next-action value
  • Terminal transition: continuation term is zero
  • Bracket: temporal-difference error

When and why to use this

Use a small tabular environment to understand sequential cost trade-offs before scaling to DQN or a larger state representation.

The Bellman idea splits an episode into immediate reward plus the discounted value of the next state. This recursion allows long-horizon decisions to be learned from shorter transitions. Terminal states have no continuation reward under the stated episode definition.

Tabular Q-learning moves a state-action estimate toward a target built from observed reward and the best next action value. Deep Q-networks replace the table with a neural approximation and introduce additional stability issues, including moving targets and extrapolation to poorly observed actions.

For a discrete execution example, actions might be wait, place a passive order or cross the spread. The value of waiting depends on remaining quantity, deadline and price risk. A high next-state Q estimate learned from an unrealistic fill model propagates the simulator error backward into earlier decisions.

Q(s,a)←Q(s,a)+η[r+γmaxa′⁡Q(s′,a′)−Q(s,a)]
Model assumptions, derivation and arithmetic

Bellman recursion and Q-learning

  1. Form the target as immediate reward plus discounted best feasible next value, or immediate reward alone at termination.
  2. Subtract the current Q estimate to obtain the temporal-difference error.
  3. Move the estimate η of the way toward the target. This arithmetic is an update rule; convergence requires additional conditions and is not guaranteed for arbitrary neural approximations.
Work it by hand

Current Q=2, reward=.5, γ=.9, best next Q=3 and η=.1 give target 3.2, error 1.2 and updated Q=2.12. If terminal, target=.5 and update=1.85.

Apply it in a strategy

  • Construct a tiny finite execution environment with a known or exhaustively computed solution.
  • Verify terminal targets and feasible-action masks, then compare learned Q values with the benchmark.
  • Add function approximation only after the state, reward and fill model are validated.

Research deliverable

Show a transition table with rewards, continuation values and manual Q updates, including terminal and invalid-action cases.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def q_update(current,reward,next_values,gamma,learning_rate,terminal=False):
    if not 0<=gamma<=1 or not 0<learning_rate<=1: raise ValueError("Invalid learning parameters")
    if not terminal and not next_values: raise ValueError("Feasible next actions required")
    target=reward if terminal else reward+gamma*max(next_values)
    return current+learning_rate*(target-current),target

print(q_update(2,.5,[1,3],.9,.1),q_update(2,.5,[],.9,.1,True))

Continue learning

Reinforcement Learning for Trading & Execution — all lessons
  1. Before RL: one action, one outcome and learning an average
  2. Choose an RL task: states, actions and the environment
  3. Bellman recursion and Q-learning
  4. From a value table to linear features and LSTD
  5. Policy gradients, actor–critic and PPO clipping
  6. Reward design, inventory penalties and terminal accounting
  7. Offline RL, logged actions and support limits
  8. Simulator transfer, regime randomisation and RL research promotion

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations