Trading Dev AcademyFree quant education

Free lesson · Reinforcement learning

Policy gradients, actor–critic and PPO clipping

Open interactive lessonPractice calculationsExplore labs

Start with the idea

A policy-gradient method changes the probabilities of actions directly. Actions associated with better-than-expected outcomes become more likely, while a critic estimates the baseline used to judge those outcomes.

Symbols, units & horizon
  • π_θ,π_old: new and behaviour policy action probabilities/densities
  • ρ_t: probability ratio, not market correlation
  • A_t: estimated advantage in reward units
  • ε: positive clipping width
  • clip: bound a number to the interval
  • min: pessimistic surrogate branch
  • L_clip: objective maximised during PPO policy training

When and why to use this

Use actor–critic methods for sequential control with a justified action space, after a simple baseline and environment audit are in place.

The actor maps observations to an action distribution; the critic estimates expected future reward. An advantage compares an action’s outcome with a baseline, reducing noise in the update. Continuous actions can represent target weights or order allocations, but must be mapped to feasible discrete orders when deployed.

PPO compares the new policy probability of a logged action with the old policy probability and clips the improvement incentive outside a small ratio range. This limits one surrogate objective, not the policy’s actual market risk or every possible distribution shift. Inventory bounds still belong in the environment and execution layer.

For trading, preserve episode structure and realistic rewards during policy updates. Report seed dispersion and interaction budget. Comparing a heavily tuned actor–critic with an untuned fixed schedule is not a fair algorithm comparison.

ρt(θ)=πθ(at|st)πold(at|st),Lclip=E[min⁡(ρtAt,clip⁡(ρt,1−ϵ,1+ϵ)At)]
Model assumptions, derivation and arithmetic

Policy gradients, actor–critic and PPO clipping

  1. Evaluate the probability of the same observed action under the new and old policies and divide.
  2. Multiply the ratio by its estimated advantage. Also form a version with the ratio clipped to [1−ε,1+ε].
  3. Take the smaller surrogate contribution. For positive advantage, ratios above the upper bound stop increasing the objective; for negative advantage, clipping behaves asymmetrically through the minimum.
Work it by hand

Old probability .2, new .3 imply ratio 1.5. Advantage 2 and ε=.2 give min(3,2.4)=2.4. With advantage −2, min(−3,−2.4)=−3.

Apply it in a strategy

  • Define action probabilities and feasible-order mapping, including rounding and remaining-quantity constraints.
  • Check advantage, probability ratio and clipping on a small stored trajectory.
  • Evaluate policy updates on disjoint episodes and stress reward/transition misspecification.

Research deliverable

Keep an actor–critic training report with episode budget, seeds, clipped ratios, constraint violations and benchmark comparisons.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def ppo_term(old_probability,new_probability,advantage,epsilon=.2):
    if old_probability<=0 or new_probability<0 or not 0<epsilon<1: raise ValueError("Invalid policy inputs")
    ratio=new_probability/old_probability
    clipped=max(1-epsilon,min(1+epsilon,ratio))
    return ratio,min(ratio*advantage,clipped*advantage)

print(ppo_term(.2,.3,2),ppo_term(.2,.3,-2))

Continue learning

Reinforcement Learning for Trading & Execution — all lessons
  1. Before RL: one action, one outcome and learning an average
  2. Choose an RL task: states, actions and the environment
  3. Bellman recursion and Q-learning
  4. From a value table to linear features and LSTD
  5. Policy gradients, actor–critic and PPO clipping
  6. Reward design, inventory penalties and terminal accounting
  7. Offline RL, logged actions and support limits
  8. Simulator transfer, regime randomisation and RL research promotion

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations