Trading Dev AcademyFree quant education

Free lesson · Reinforcement learning

Before RL: one action, one outcome and learning an average

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Start with two buttons and a score after each choice. Keep an average score for each button. This tiny problem explains what learning from actions means before adding states, trajectories or neural networks.

Symbols, units & horizon
  • n: count of observations for this action after the new reward
  • r_n: new observed reward in fixed reward units
  • r̄_(n−1): previous sample mean for this action
  • r̄_n: updated mean
  • 1/n: averaging step size
  • reward convention: for execution it may be negative cost, so larger is better

When and why to use this

Use an explicit action/reward ledger and a simple fixed or bandit baseline before attempting sequential execution or inventory control.

A one-step bandit chooses an action and observes its reward; there is no modeled future state changed by that choice. For an execution analogy, compare two ways to place a small order and measure cost relative to the same arrival benchmark. Treat the example as a simplified experiment, since actual execution can have sequential consequences.

You only observe the consequence of the action you took. An unchosen alternative does not come with a known counterfactual fill. Exploration collects information about alternatives; exploitation uses the current preferred action. Historical observations can be selected by an earlier policy, so naive comparisons may be biased.

A sample average can be updated without storing every reward: move the old average partway toward the new observation. A step size of one divided by the new count gives the exact average. A fixed step size instead forgets old data gradually and is a different estimator.

Now add one remaining order and a deadline. Waiting today changes tomorrow’s remaining quantity and risk. The objective must include future consequences, not just the current reward. That is the reason to introduce an MDP and Bellman recursion in the following lessons. Learn the small table first; a larger network does not fix an incomplete reward ledger.

r‾n=(n−1)r‾n−1+rnn=r‾n−1+rn−r‾n−1n
Model assumptions, derivation and arithmetic

Before RL: one action, one outcome and learning an average

  1. The previous n−1 observations have total reward (n−1) times their average. Add the new reward.
  2. Divide by n to form the new mean. Expand (n−1)/n as 1−1/n.
  3. Collect the old mean and the correction: old mean plus (new reward−old mean)/n. The n=1 case initializes the estimate directly from the first reward.
Work it by hand

An action has three observations averaging 2 reward units. A fourth reward is −2. Updated mean=(3×2−2)/4=1, equivalently 2+(−2−2)/4=1. One bad observation changes the estimate but does not identify its long-run quality.

Apply it in a strategy

  • Define a bounded action and one measurable reward with consistent costs.
  • Keep per-action counts, means and context, including how actions were chosen.
  • Only add sequential state when actions change future choices or costs; then specify terminal accounting.

Research deliverable

Update two action averages by hand and explain what unobserved counterfactuals prevent you from concluding.

Sources & evidence · reviewed 12 September 2026

The linear-value lesson records the supplied 2015 thesis; the final RL checkpoint records 2026 research. Start with this original arithmetic example before comparing their more complex methods.

Further reading: Historical RL study and current extensions ↗

Research sources, review dates and limitations

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

from math import isfinite

def update_reward_mean(old_mean,new_reward,new_count):
    if not isinstance(new_count,int) or new_count<1 or not all(isfinite(x) for x in [old_mean,new_reward]): raise ValueError("Finite rewards and positive integer count required")
    return new_reward if new_count==1 else old_mean+(new_reward-old_mean)/new_count

print(update_reward_mean(2,-2,4))

Continue learning

Reinforcement Learning for Trading & Execution — all lessons
  1. Before RL: one action, one outcome and learning an average
  2. Choose an RL task: states, actions and the environment
  3. Bellman recursion and Q-learning
  4. From a value table to linear features and LSTD
  5. Policy gradients, actor–critic and PPO clipping
  6. Reward design, inventory penalties and terminal accounting
  7. Offline RL, logged actions and support limits
  8. Simulator transfer, regime randomisation and RL research promotion

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations