Free lesson · Reinforcement learning
From a value table to linear features and LSTD
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A table needs a separate value for every state. A linear value model shares information by expressing each state as a short list of features and taking their weighted sum.
Symbols, units & horizon
- V̂(s): estimated discounted value of state s in reward units
- φ(s): column feature vector, dimensionless in this example
- θ: feature coefficient vector in reward units
- superscript ⊤: transpose
- γ: discount factor in [0,1]
- r_(t+1): observed immediate reward
- t: transition index
- φ_(t+1): zero vector for a true terminal transition
- sums: empirical moments from a fixed policy’s permitted transitions
When and why to use this
Use small linear value approximations to understand feature design, Bellman projection and numerical conditioning before deep value networks.
A feature might be remaining inventory divided by its starting value, time remaining divided by the deadline, or a constant intercept. A value estimate is the dot product of these features with learned coefficients. Begin with one feature and one coefficient before using a matrix.
For a fixed policy, the temporal-difference residual is immediate reward plus discounted next-state estimated value minus the current estimate. Least-squares temporal-difference learning, or LSTD, asks for an empirical feature-weighted residual of zero. This produces a linear system rather than a sequence of small gradient updates.
This is a projected Bellman moment equation, not ordinary regression against observed full future returns and not direct minimization of squared TD residuals. Its matrix need not be symmetric. A solve can be poorly conditioned or singular; adding a diagonal regularizer changes the estimate and does not establish a profitable policy.
At a true terminal transition, set next-state features to zero after including all terminal costs in the reward. Policy evaluation estimates the value of a specified policy. Policy improvement is a separate operation; least-squares policy iteration alternates evaluation and improvement with additional action features and data-support requirements. The historical RL thesis motivates this progression, not an immediate jump to a modern deep agent.
From a value table to linear features and LSTD
- Write TD residual δ_t=r_(t+1)+γφ_(t+1)ᵀθ−φ_tᵀθ. Multiply by current feature vector φ_t.
- Set the sum of feature-weighted residuals to zero. Move the coefficient terms to the other side.
- Factor out θ to obtain Aθ=b, where A is the sum of outer products φ_t(φ_t−γφ_(t+1))ᵀ and b is the sum of φ_t r_(t+1).
- Solve the linear system without explicitly inverting A. In one dimension this is ordinary division when the denominator is nonzero.
One-feature transitions have current φ=[1,1], next φ=[1,0], rewards [1,2] and γ=.5. A=1(1−.5)+1(1−0)=1.5; b=1×1+1×2=3; θ=3/1.5=2. A constant feature shares one value across states, so it may be misspecified.
Apply it in a strategy
- Solve a tiny fixed-policy environment and compare with an exact table.
- Construct feature and reward matrices, setting terminal successor features to zero.
- Check conditioning, representation error and held-out policy behavior before trying policy improvement.
Research deliverable
Show the individual outer-product contributions to A and reward contributions to b, then solve and interpret the coefficient.
Sources & evidence · reviewed 12 September 2026
Reviewed 12 September 2026: Imperial College MEng thesis, 18 June 2015. Background, data table, LSTD/LSPI formulation and evaluation discussion reviewed in full text. The data table lists eleven FX pairs from 2 January–19 December 2014. These are one-minute bid candles; the thesis defers explicit/implicit transaction-cost analysis. Its proof-of-concept and limited reported profitability motivate explicit state/reward design and simple value approximation; they do not validate a contemporary bot. The lesson’s two-transition calculation is original, and the historical backtest was not reproduced.
Further reading: James Cumming · An Investigation into the Use of Reinforcement Learning Techniques within the Algorithmic Trading Domain ↗
Use the module’s current preprint notes to compare more recent environment and action-space choices after understanding fixed-policy evaluation.
Further reading: Recent RL research checkpoint ↗
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
import numpy as np
def lstd_value(features,next_features,rewards,gamma,regularization=0):
x=np.asarray(features,dtype=float); nxt=np.asarray(next_features,dtype=float); r=np.asarray(rewards,dtype=float)
if x.ndim!=2 or x.shape[0]==0 or x.shape[1]==0 or nxt.shape!=x.shape or r.shape!=(len(x),): raise ValueError("Transition dimensions must match")
if not all(np.isfinite(a).all() for a in [x,nxt,r]) or not 0<=gamma<=1 or not np.isfinite(regularization) or regularization<0: raise ValueError("Invalid numeric inputs")
a=x.T@(x-gamma*nxt); b=x.T@r
return np.linalg.solve(a+regularization*np.eye(x.shape[1]),b)
print(lstd_value([[1],[1]],[[1],[0]],[1,2],.5))Continue learning
Reinforcement Learning for Trading & Execution — all lessons- Before RL: one action, one outcome and learning an average
- Choose an RL task: states, actions and the environment
- Bellman recursion and Q-learning
- From a value table to linear features and LSTD
- Policy gradients, actor–critic and PPO clipping
- Reward design, inventory penalties and terminal accounting
- Offline RL, logged actions and support limits
- Simulator transfer, regime randomisation and RL research promotion
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations