Trading Dev AcademyFree quant education

Free lesson · Deep learning for finance

Attention and transformers: which history can the model use?

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Attention forms a weighted combination of representations based on learned relevance scores. It can connect distant observations, but its mask and training data determine whether those connections are legitimate for a trading decision.

Symbols, units & horizon
  • Q,K: query and key matrices
  • V: value matrix, not option value here
  • d_k: key/query feature dimension
  • T: transpose
  • M: additive mask, zero for allowed positions and −∞ for forbidden positions
  • softmax: rowwise positive weights summing to one
  • A: attention weights
  • H: weighted output representations

When and why to use this

Use attention when long or cross-feature context plausibly improves the target, with explicit masking and a simpler baseline.

Queries ask what information is useful, keys determine matching scores, and values carry the information to combine. A causal mask blocks attention to future positions for per-time predictions. A final-window forecast may attend to the entire already-observed window, but must not use later target data.

Transformers can process multivariate sequences, patches of returns or text representations. Positional information tells the model where observations belong. Irregular market events may need elapsed-time features rather than assuming equally spaced steps. Larger context is not automatically helpful when old regimes are irrelevant.

Attention weights are not causal explanations or reliable feature importance by themselves. Compare temporal architectures at matched data, tuning and compute budgets. For HFT, measure end-to-end inference latency and stale-feature handling; a better offline loss can be unusable within the edge lifetime.

A=softmax⁡(QK𝖳dk+M),H=AV
Model assumptions, derivation and arithmetic

Attention and transformers: which history can the model use?

  1. Take query–key dot products and divide by √d_k to control score scale.
  2. Set forbidden future scores to negative infinity before rowwise softmax; their exponential weights become zero.
  3. Normalise allowed exponential scores by their sum, then use those weights to average the value vectors.
Work it by hand

For two allowed scores both zero, weights are [.5,.5]. Values [2,6] produce output 4. If the second position is future-masked, weights become [1,0] and output 2.

Apply it in a strategy

  • Specify whether predictions occur at each timestamp or only at the window end, then set the corresponding attention mask.
  • Verify prefix invariance where required and prevent future-aware pretraining or normalisation.
  • Compare net decision quality, latency and compute against convolutional and recurrent baselines.

Research deliverable

Include an attention-mask diagram, training-cutoff record and same-budget architecture comparison.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

import numpy as np

def causal_attention(q,k,v):
    q=np.asarray(q,float); k=np.asarray(k,float); v=np.asarray(v,float)
    if q.ndim!=2 or q.shape!=k.shape or len(v)!=len(q): raise ValueError("Aligned sequence matrices required")
    scores=q@k.T/np.sqrt(q.shape[1])
    scores[np.triu_indices(len(q),1)]=-np.inf
    weights=np.exp(scores-scores.max(axis=1,keepdims=True))
    weights/=weights.sum(axis=1,keepdims=True)
    return weights@v,weights

print(causal_attention([[0],[0]],[[0],[0]],[[2],[6]])[0])

Continue learning

Deep Learning: Sequences, Representations & Financial Decisions — all lessons
  1. Neural networks and backpropagation from first principles
  2. Causal windows, temporal convolutions and order-book tensors
  3. Recurrent networks and LSTM gates
  4. Attention and transformers: which history can the model use?
  5. Forecast loss versus trading loss, turnover and differentiable decisions
  6. Fine-tuning, financial text and foundation-model contamination

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations