Free lesson · Reinforcement learning
Bellman recursion and Q-learning
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A value function estimates the reward still available from a state. A Q-function estimates the reward from taking one particular action and then following a policy or an optimal continuation.
Symbols, units & horizon
- Q(s,a): estimated action value
- s′: next state
- a′: candidate next action
- r: observed immediate reward
- γ: discount factor
- η: learning step size in (0,1]
- max: largest feasible next-action value
- Terminal transition: continuation term is zero
- Bracket: temporal-difference error
When and why to use this
Use a small tabular environment to understand sequential cost trade-offs before scaling to DQN or a larger state representation.
The Bellman idea splits an episode into immediate reward plus the discounted value of the next state. This recursion allows long-horizon decisions to be learned from shorter transitions. Terminal states have no continuation reward under the stated episode definition.
Tabular Q-learning moves a state-action estimate toward a target built from observed reward and the best next action value. Deep Q-networks replace the table with a neural approximation and introduce additional stability issues, including moving targets and extrapolation to poorly observed actions.
For a discrete execution example, actions might be wait, place a passive order or cross the spread. The value of waiting depends on remaining quantity, deadline and price risk. A high next-state Q estimate learned from an unrealistic fill model propagates the simulator error backward into earlier decisions.
Bellman recursion and Q-learning
- Form the target as immediate reward plus discounted best feasible next value, or immediate reward alone at termination.
- Subtract the current Q estimate to obtain the temporal-difference error.
- Move the estimate η of the way toward the target. This arithmetic is an update rule; convergence requires additional conditions and is not guaranteed for arbitrary neural approximations.
Current Q=2, reward=.5, γ=.9, best next Q=3 and η=.1 give target 3.2, error 1.2 and updated Q=2.12. If terminal, target=.5 and update=1.85.
Apply it in a strategy
- Construct a tiny finite execution environment with a known or exhaustively computed solution.
- Verify terminal targets and feasible-action masks, then compare learned Q values with the benchmark.
- Add function approximation only after the state, reward and fill model are validated.
Research deliverable
Show a transition table with rewards, continuation values and manual Q updates, including terminal and invalid-action cases.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def q_update(current,reward,next_values,gamma,learning_rate,terminal=False):
if not 0<=gamma<=1 or not 0<learning_rate<=1: raise ValueError("Invalid learning parameters")
if not terminal and not next_values: raise ValueError("Feasible next actions required")
target=reward if terminal else reward+gamma*max(next_values)
return current+learning_rate*(target-current),target
print(q_update(2,.5,[1,3],.9,.1),q_update(2,.5,[],.9,.1,True))Continue learning
Reinforcement Learning for Trading & Execution — all lessons- Before RL: one action, one outcome and learning an average
- Choose an RL task: states, actions and the environment
- Bellman recursion and Q-learning
- From a value table to linear features and LSTD
- Policy gradients, actor–critic and PPO clipping
- Reward design, inventory penalties and terminal accounting
- Offline RL, logged actions and support limits
- Simulator transfer, regime randomisation and RL research promotion
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations