Trading Dev AcademyFree quant education

Free lesson · Reinforcement learning

Offline RL, logged actions and support limits

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Offline RL learns from a fixed history generated by another policy. Historical market prices alone do not reveal what would have happened after every possible action. Reliable evaluation needs support for the proposed decisions.

Symbols, units & horizon
  • V̂_IPS: inverse-propensity estimate for a one-step contextual-bandit task
  • π: target policy
  • b: behaviour/logging policy
  • a_i,s_i,r_i: logged action, context and reward
  • n: logged observations
  • Support: b must be positive wherever π acts
  • Sequential RL: this one-step estimator is not sufficient

When and why to use this

Use support diagnostics before claiming an offline policy improves execution or allocation based on historical logs.

Logged execution data should include state, available actions, chosen action, behaviour probability when known, reward and next state. A new policy that proposes actions never observed in similar states requires extrapolation. In markets, its action may also change fills and future prices.

Importance weighting corrects a simple contextual-bandit estimate by the ratio of target-policy probability to logging-policy probability. Sequential trajectories require products or more advanced estimators and can have enormous variance. Unknown or zero logging probabilities make the simple estimator unusable for those actions.

Offline RL methods often regularise against unsupported actions or use conservative value estimates. That addresses one failure mode under assumptions, not all market counterfactuals. Treat simulator evaluation, logged-policy evaluation and prospective paper/live evidence as different sources with distinct limitations.

V^IPS=1n∑iπ(ai|si)b(ai|si)ri
Model assumptions, derivation and arithmetic

Offline RL, logged actions and support limits

  1. For an action observed under behaviour b, weight its reward by target probability divided by behaviour probability.
  2. In expectation, summing b(a|s) times π(a|s)/b(a|s) times reward replaces the behaviour action distribution with the target distribution, if support and logging assumptions hold.
  3. Average the weighted observations. Large ratios increase variance; clipping ratios introduces bias and should be reported.
Work it by hand

Two logged rewards [1,0], behaviour probabilities [.5,.5] and target probabilities [.75,.25] give weights [1.5,.5] and IPS estimate .75. Two observations provide almost no precision.

Apply it in a strategy

  • Audit logging propensities, action feasibility and coverage of proposed target actions.
  • Use an estimator appropriate to the actual one-step or sequential problem and report uncertainty and weight concentration.
  • Test conservative baselines and obtain prospective evidence for actions outside reliable historical support.

Research deliverable

Provide an action-support report and clearly separate logged-data estimates from simulator and prospective outcomes.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def bandit_ips(rewards,target_probabilities,behaviour_probabilities):
    if not rewards or not len(rewards)==len(target_probabilities)==len(behaviour_probabilities): raise ValueError("Aligned logs required")
    if any(not 0<=p<=1 for p in target_probabilities) or any(not 0<p<=1 for p in behaviour_probabilities): raise ValueError("Valid positive logging support required")
    weights=[p/b for p,b in zip(target_probabilities,behaviour_probabilities)]
    return sum(w*r for w,r in zip(weights,rewards))/len(rewards),weights

print(bandit_ips([1,0],[.75,.25],[.5,.5]))

Continue learning

Reinforcement Learning for Trading & Execution — all lessons
  1. Before RL: one action, one outcome and learning an average
  2. Choose an RL task: states, actions and the environment
  3. Bellman recursion and Q-learning
  4. From a value table to linear features and LSTD
  5. Policy gradients, actor–critic and PPO clipping
  6. Reward design, inventory penalties and terminal accounting
  7. Offline RL, logged actions and support limits
  8. Simulator transfer, regime randomisation and RL research promotion

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations