Trading Dev AcademyFree quant education

Free lesson · Reinforcement learning

From a value table to linear features and LSTD

Open interactive lessonPractice calculationsExplore labs

Start with the idea

A table needs a separate value for every state. A linear value model shares information by expressing each state as a short list of features and taking their weighted sum.

Symbols, units & horizon
  • V̂(s): estimated discounted value of state s in reward units
  • φ(s): column feature vector, dimensionless in this example
  • θ: feature coefficient vector in reward units
  • superscript ⊤: transpose
  • γ: discount factor in [0,1]
  • r_(t+1): observed immediate reward
  • t: transition index
  • φ_(t+1): zero vector for a true terminal transition
  • sums: empirical moments from a fixed policy’s permitted transitions

When and why to use this

Use small linear value approximations to understand feature design, Bellman projection and numerical conditioning before deep value networks.

A feature might be remaining inventory divided by its starting value, time remaining divided by the deadline, or a constant intercept. A value estimate is the dot product of these features with learned coefficients. Begin with one feature and one coefficient before using a matrix.

For a fixed policy, the temporal-difference residual is immediate reward plus discounted next-state estimated value minus the current estimate. Least-squares temporal-difference learning, or LSTD, asks for an empirical feature-weighted residual of zero. This produces a linear system rather than a sequence of small gradient updates.

This is a projected Bellman moment equation, not ordinary regression against observed full future returns and not direct minimization of squared TD residuals. Its matrix need not be symmetric. A solve can be poorly conditioned or singular; adding a diagonal regularizer changes the estimate and does not establish a profitable policy.

At a true terminal transition, set next-state features to zero after including all terminal costs in the reward. Policy evaluation estimates the value of a specified policy. Policy improvement is a separate operation; least-squares policy iteration alternates evaluation and improvement with additional action features and data-support requirements. The historical RL thesis motivates this progression, not an immediate jump to a modern deep agent.

V^(s)=ϕ(s)⊤θ,[∑tϕt(ϕt−γϕt+1)⊤]θ=∑tϕtrt+1
Model assumptions, derivation and arithmetic

From a value table to linear features and LSTD

  1. Write TD residual δ_t=r_(t+1)+γφ_(t+1)ᵀθ−φ_tᵀθ. Multiply by current feature vector φ_t.
  2. Set the sum of feature-weighted residuals to zero. Move the coefficient terms to the other side.
  3. Factor out θ to obtain Aθ=b, where A is the sum of outer products φ_t(φ_t−γφ_(t+1))ᵀ and b is the sum of φ_t r_(t+1).
  4. Solve the linear system without explicitly inverting A. In one dimension this is ordinary division when the denominator is nonzero.
Work it by hand

One-feature transitions have current φ=[1,1], next φ=[1,0], rewards [1,2] and γ=.5. A=1(1−.5)+1(1−0)=1.5; b=1×1+1×2=3; θ=3/1.5=2. A constant feature shares one value across states, so it may be misspecified.

Apply it in a strategy

  • Solve a tiny fixed-policy environment and compare with an exact table.
  • Construct feature and reward matrices, setting terminal successor features to zero.
  • Check conditioning, representation error and held-out policy behavior before trying policy improvement.

Research deliverable

Show the individual outer-product contributions to A and reward contributions to b, then solve and interpret the coefficient.

Sources & evidence · reviewed 12 September 2026

Reviewed 12 September 2026: Imperial College MEng thesis, 18 June 2015. Background, data table, LSTD/LSPI formulation and evaluation discussion reviewed in full text. The data table lists eleven FX pairs from 2 January–19 December 2014. These are one-minute bid candles; the thesis defers explicit/implicit transaction-cost analysis. Its proof-of-concept and limited reported profitability motivate explicit state/reward design and simple value approximation; they do not validate a contemporary bot. The lesson’s two-transition calculation is original, and the historical backtest was not reproduced.

Further reading: James Cumming · An Investigation into the Use of Reinforcement Learning Techniques within the Algorithmic Trading Domain ↗

Use the module’s current preprint notes to compare more recent environment and action-space choices after understanding fixed-policy evaluation.

Further reading: Recent RL research checkpoint ↗

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

import numpy as np

def lstd_value(features,next_features,rewards,gamma,regularization=0):
    x=np.asarray(features,dtype=float); nxt=np.asarray(next_features,dtype=float); r=np.asarray(rewards,dtype=float)
    if x.ndim!=2 or x.shape[0]==0 or x.shape[1]==0 or nxt.shape!=x.shape or r.shape!=(len(x),): raise ValueError("Transition dimensions must match")
    if not all(np.isfinite(a).all() for a in [x,nxt,r]) or not 0<=gamma<=1 or not np.isfinite(regularization) or regularization<0: raise ValueError("Invalid numeric inputs")
    a=x.T@(x-gamma*nxt); b=x.T@r
    return np.linalg.solve(a+regularization*np.eye(x.shape[1]),b)

print(lstd_value([[1],[1]],[[1],[0]],[1,2],.5))

Continue learning

Reinforcement Learning for Trading & Execution — all lessons
  1. Before RL: one action, one outcome and learning an average
  2. Choose an RL task: states, actions and the environment
  3. Bellman recursion and Q-learning
  4. From a value table to linear features and LSTD
  5. Policy gradients, actor–critic and PPO clipping
  6. Reward design, inventory penalties and terminal accounting
  7. Offline RL, logged actions and support limits
  8. Simulator transfer, regime randomisation and RL research promotion

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations