Free lesson · Reinforcement learning
Reward design, inventory penalties and terminal accounting
Open interactive lessonPractice calculationsExplore labs
Start with the idea
The reward tells the agent what to optimise. Any omitted cost, unmarked position or inconsistent deadline becomes an incentive to exploit the simulator rather than solve the intended trading problem.
Symbols, units & horizon
- ΔW_t: marked-wealth change including executed costs
- q_t: inventory
- λ: penalty scaled to reward per inventory² per time
- Δt: elapsed time
- r_t: economic reward plus preference penalty
- Φ(s): state potential
- γ: same discount used in episode return
- r′: shaped reward
- Terminal potential: must be handled consistently
When and why to use this
Use reward design to encode a clearly specified execution or risk preference while retaining a separate, auditable cash P&L metric.
For market making, use changes in marked wealth including fees, then explicitly add any risk preference such as inventory penalties. For execution, measure implementation shortfall relative to a fixed arrival benchmark and account for all remaining quantity at the deadline. For rebalancing, include turnover, financing and the chosen capital base.
Reward shaping can improve learning by providing intermediate feedback. Potential-based shaping has a telescoping structure under the same discount factor; policy-invariance claims require correct terminal handling and other assumptions. Arbitrary bonuses for fills can instead encourage excessive turnover.
Compare the training reward with the final economic metric on every episode. Separate preference penalties from actual cash costs so the reported P&L remains interpretable. A penalty may steer average behaviour, but an independent feasible-action constraint is needed for a hard inventory limit.
Reward design, inventory penalties and terminal accounting
- Compute actual marked-wealth change from the ledger, then subtract the chosen inventory-time penalty.
- Add a potential difference to shape intermediate feedback. In the discounted sum, adjacent potential terms cancel.
- The remaining change is −Φ(s_0)+γ^TΦ(s_T). A fixed initial state and zero or appropriately fixed terminal potential avoid changing policy ranking through this boundary term.
Wealth change .5, inventory 2, λ=.01 and Δt=1 give reward .46. With γ=.9, current potential 1 and next potential 1.2, shaping adds .08, producing .54.
Apply it in a strategy
- Reconcile each reward term with fills, marks, inventory and elapsed time on a hand-worked episode.
- Test for reward exploits: never liquidating, excessive churning, delaying to episode end or exploiting free cancellations.
- Report both shaped reward and unshaped economic results, with terminal costs and independent constraints.
Research deliverable
Write a reward specification with units, terminal conditions and adversarial toy episodes demonstrating that unwanted behaviour is not rewarded.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def inventory_reward(wealth_change,inventory,penalty,elapsed):
if penalty<0 or elapsed<0: raise ValueError("Nonnegative penalty and time required")
return wealth_change-penalty*inventory**2*elapsed
def shape_reward(reward,current_potential,next_potential,gamma):
return reward+gamma*next_potential-current_potential
r=inventory_reward(.5,2,.01,1)
print(r,shape_reward(r,1,1.2,.9))Continue learning
Reinforcement Learning for Trading & Execution — all lessons- Before RL: one action, one outcome and learning an average
- Choose an RL task: states, actions and the environment
- Bellman recursion and Q-learning
- From a value table to linear features and LSTD
- Policy gradients, actor–critic and PPO clipping
- Reward design, inventory penalties and terminal accounting
- Offline RL, logged actions and support limits
- Simulator transfer, regime randomisation and RL research promotion
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations