Free lesson · Reinforcement learning
Choose an RL task: states, actions and the environment
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Reinforcement learning learns a policy from the consequences of actions over time. It is most compelling when today’s action changes tomorrow’s feasible choices or costs, such as remaining execution quantity or inventory.
Symbols, units & horizon
- G_t: discounted future reward from t
- r_(t+k+1): one-step reward, not necessarily an asset return
- γ: discount factor in [0,1]
- T: terminal decision index
- π: policy distribution over actions
- s,a: state and action
- S_t,A_t: random state/action variables
- Reward units: must remain consistent over the episode
When and why to use this
Use RL when actions create meaningful sequential trade-offs, and keep a fixed schedule or supervised decision baseline to justify the extra complexity.
A Markov decision process specifies states, actions, transition probabilities and rewards. In trading, the true market state is only partially observed, so your input is usually an observation or a learned belief state. Calling a feature vector “state” does not make the Markov assumption true.
Useful bounded tasks include slicing a parent order, choosing passive versus aggressive execution, controlling market-maker inventory or rebalancing under costs. If actions do not materially change future opportunities, a supervised model or contextual bandit can be a simpler baseline.
State should include relevant holdings, cash, outstanding orders, time remaining and market features available now. Actions need feasible quantities and venue constraints. Specify terminal handling: an execution policy must account for any unfilled quantity, and a trading policy must mark or liquidate residual inventory consistently.
Choose an RL task: states, actions and the environment
- List each future reward after an action, preserving its event order.
- Multiply a reward k steps away by γ^k and sum until termination. γ expresses the objective’s weighting, not automatically a risk-free discount rate.
- A policy maps the available state to action probabilities. Learning seeks a policy with high expected G under the environment’s transition rules.
Rewards [1,2,−1] and γ=.9 give return 1+.9×2+.81×(−1)=1.99. For a finite execution task, γ=1 may be appropriate when all costs belong to the same episode objective.
Apply it in a strategy
- Choose one bounded task and list all state, action, reward and terminal conventions.
- Document what market dynamics your actions can change and which remain exogenous.
- Compare against TWAP, a simple inventory rule or a contextual policy at matched costs and constraints.
Research deliverable
Write an environment specification with state dimensions, action feasibility, episode boundaries and a manually checked trajectory.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def discounted_return(rewards,gamma):
if not 0<=gamma<=1: raise ValueError("Discount must lie in [0,1]")
return sum(gamma**k*r for k,r in enumerate(rewards))
print(discounted_return([1,2,-1],.9))Continue learning
Reinforcement Learning for Trading & Execution — all lessons- Before RL: one action, one outcome and learning an average
- Choose an RL task: states, actions and the environment
- Bellman recursion and Q-learning
- From a value table to linear features and LSTD
- Policy gradients, actor–critic and PPO clipping
- Reward design, inventory penalties and terminal accounting
- Offline RL, logged actions and support limits
- Simulator transfer, regime randomisation and RL research promotion
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations