Trading Dev AcademyFree quant education

Free lesson · Reinforcement learning

Choose an RL task: states, actions and the environment

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Reinforcement learning learns a policy from the consequences of actions over time. It is most compelling when today’s action changes tomorrow’s feasible choices or costs, such as remaining execution quantity or inventory.

Symbols, units & horizon
  • G_t: discounted future reward from t
  • r_(t+k+1): one-step reward, not necessarily an asset return
  • γ: discount factor in [0,1]
  • T: terminal decision index
  • π: policy distribution over actions
  • s,a: state and action
  • S_t,A_t: random state/action variables
  • Reward units: must remain consistent over the episode

When and why to use this

Use RL when actions create meaningful sequential trade-offs, and keep a fixed schedule or supervised decision baseline to justify the extra complexity.

A Markov decision process specifies states, actions, transition probabilities and rewards. In trading, the true market state is only partially observed, so your input is usually an observation or a learned belief state. Calling a feature vector “state” does not make the Markov assumption true.

Useful bounded tasks include slicing a parent order, choosing passive versus aggressive execution, controlling market-maker inventory or rebalancing under costs. If actions do not materially change future opportunities, a supervised model or contextual bandit can be a simpler baseline.

State should include relevant holdings, cash, outstanding orders, time remaining and market features available now. Actions need feasible quantities and venue constraints. Specify terminal handling: an execution policy must account for any unfilled quantity, and a trading policy must mark or liquidate residual inventory consistently.

Gt=∑k=0T−t−1γkrt+k+1,π(a|s)=P(At=a|St=s)
Model assumptions, derivation and arithmetic

Choose an RL task: states, actions and the environment

  1. List each future reward after an action, preserving its event order.
  2. Multiply a reward k steps away by γ^k and sum until termination. γ expresses the objective’s weighting, not automatically a risk-free discount rate.
  3. A policy maps the available state to action probabilities. Learning seeks a policy with high expected G under the environment’s transition rules.
Work it by hand

Rewards [1,2,−1] and γ=.9 give return 1+.9×2+.81×(−1)=1.99. For a finite execution task, γ=1 may be appropriate when all costs belong to the same episode objective.

Apply it in a strategy

  • Choose one bounded task and list all state, action, reward and terminal conventions.
  • Document what market dynamics your actions can change and which remain exogenous.
  • Compare against TWAP, a simple inventory rule or a contextual policy at matched costs and constraints.

Research deliverable

Write an environment specification with state dimensions, action feasibility, episode boundaries and a manually checked trajectory.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def discounted_return(rewards,gamma):
    if not 0<=gamma<=1: raise ValueError("Discount must lie in [0,1]")
    return sum(gamma**k*r for k,r in enumerate(rewards))

print(discounted_return([1,2,-1],.9))

Continue learning

Reinforcement Learning for Trading & Execution — all lessons
  1. Before RL: one action, one outcome and learning an average
  2. Choose an RL task: states, actions and the environment
  3. Bellman recursion and Q-learning
  4. From a value table to linear features and LSTD
  5. Policy gradients, actor–critic and PPO clipping
  6. Reward design, inventory penalties and terminal accounting
  7. Offline RL, logged actions and support limits
  8. Simulator transfer, regime randomisation and RL research promotion

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations