Trading Dev AcademyFree quant education

Free lesson · Deep learning for finance

Recurrent networks and LSTM gates

Open interactive lessonPractice calculationsExplore labs

Start with the idea

A recurrent model carries a state from one step to the next. An LSTM uses gates to control what is retained, added and exposed. This lets it represent dependencies whose useful duration changes with context.

Symbols, units & horizon
  • c_t: LSTM cell memory
  • h_t: exposed hidden state
  • f_t,i_t,o_t: forget, input and output gate values between 0 and 1 from learned sigmoid transforms
  • g_t: candidate memory, commonly tanh-transformed
  • t−1: previous state
  • Vector form: products are elementwise
  • This equation: LSTM architecture definition, not an optimal finance law

When and why to use this

Use recurrent memory when a decision benefits from path context, and test whether it improves upon simpler causal summaries at a feasible latency.

A plain recurrent state can struggle to retain long information through repeated nonlinear updates. The LSTM cell has an additive memory path: a forget gate scales old memory and an input gate scales new content. An output gate decides how much transformed memory reaches the hidden state.

Potential financial uses include evolving trend strength, volatility state or order-flow context. The architecture does not know which state is economically meaningful; that must emerge from training and be evaluated out of sample. Compare with simple EWMA and autoregressive states.

Reset or carry hidden state according to the deployed schedule. Do not carry a training sequence’s future-conditioned state into an earlier validation period. Truncated backpropagation changes the learning horizon; sequence length, batching and state reset conventions belong in the model specification.

ct=ftct−1+itgt,ht=ottanh⁡(ct)
Model assumptions, derivation and arithmetic

Recurrent networks and LSTM gates

  1. Compute gates and candidate content from current inputs and the previous hidden state using learned transformations.
  2. Multiply old memory by the forget gate and new candidate by the input gate, then add them.
  3. Apply tanh to the new cell and multiply by the output gate. Repeating this update defines the recurrent computation; gradients flow through both paths.
Work it by hand

Previous cell=.8, forget=.75, input=.2 and candidate=.5 give new cell=.6+.1=.7. With output gate .9, hidden state=.9tanh(.7)≈.543931.

Apply it in a strategy

  • Choose sequence boundaries, state reset policy and target horizon before tuning memory size.
  • Benchmark an LSTM against lagged linear features and EWMA state under equal validation splits.
  • Test state carryover, gaps, session boundaries and inference cost, reporting dispersion across seeds.

Research deliverable

Document gate/state dimensions and a replay test showing that training and inference use identical reset and update rules.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

from math import tanh

def lstm_cell(previous,forget,input_gate,candidate,output_gate):
    if any(not 0<=g<=1 for g in (forget,input_gate,output_gate)): raise ValueError("Gate values must lie in [0,1]")
    cell=forget*previous+input_gate*candidate
    return cell,output_gate*tanh(cell)

print(lstm_cell(.8,.75,.2,.5,.9))

Continue learning

Deep Learning: Sequences, Representations & Financial Decisions — all lessons
  1. Neural networks and backpropagation from first principles
  2. Causal windows, temporal convolutions and order-book tensors
  3. Recurrent networks and LSTM gates
  4. Attention and transformers: which history can the model use?
  5. Forecast loss versus trading loss, turnover and differentiable decisions
  6. Fine-tuning, financial text and foundation-model contamination

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations