Trading Dev AcademyFree quant education

Free lesson · Deep learning for finance

Causal windows, temporal convolutions and order-book tensors

Open interactive lessonPractice calculationsExplore labs

Start with the idea

A sequence model receives a history, not just one feature row. Causality requires that every output use only inputs that would already be available when that output is acted upon.

Symbols, units & horizon
  • x_t: observed input at time t
  • a_j: convolution weight at lag index j
  • K: kernel length
  • d: positive integer dilation
  • b: bias
  • h_t: pre-activation output
  • R: receptive-field span of one layer in input steps
  • Negative indices: require a stated padding or warmup convention

When and why to use this

Use causal convolutions for local patterns in returns or order-flow events where the relevant dependency is a bounded history.

For daily signals, a sample might contain the preceding 60 days of returns, volatility, carry and volume. For an order-book model, dimensions may represent event time, book level, side and channel. State the tensor ordering explicitly; flattening or transposing incorrectly can silently swap time with features.

A causal convolution combines only current and earlier inputs. Dilation skips evenly spaced historical lags to extend the receptive field without using future observations. Symmetric padding or a centred moving average can expose future data when training per-timestamp outputs.

Normalisation across the whole sequence may leak if it uses values later than a particular output time. Bidirectional encoders are valid when the entire input window is past and only the final-window prediction is traded; they are invalid for outputs that pretend to have been available inside that window.

ht=∑j=0K−1ajxt−dj+b,R=1+(K−1)d
Model assumptions, derivation and arithmetic

Causal windows, temporal convolutions and order-book tensors

  1. For j=0 use x_t; increasing j moves back by d steps, never forward.
  2. The earliest input is at t−d(K−1). Counting both endpoints gives span 1+d(K−1).
  3. Multiply each historical input by its kernel weight and sum. Stacking layers expands the receptive field, but that expansion must be computed for the actual architecture.
Work it by hand

Kernel [.5,.3,.2], d=1, inputs x_t=4,x_(t−1)=2,x_(t−2)=1 produce h=2+.6+.2=2.8 with zero bias. A three-point kernel at d=2 spans five timestamps.

Apply it in a strategy

  • Draw the exact tensor dimensions and the latest observable timestamp for each output.
  • Run a prefix-invariance test: changing future inputs must not alter earlier causal outputs.
  • Compare a small convolution with simple lagged features under the same horizon, costs and training budget.

Research deliverable

Provide the window definition, receptive field, padding policy and a prefix-causality test for the model.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def causal_convolution(values,kernel,dilation=1):
    if dilation<1 or not isinstance(dilation,int) or not kernel: raise ValueError("Valid kernel and dilation required")
    warmup=(len(kernel)-1)*dilation
    out=[None]*min(warmup,len(values))
    for t in range(warmup,len(values)):
        out.append(sum(a*values[t-j*dilation] for j,a in enumerate(kernel)))
    return out

print(causal_convolution([1,2,4],[.5,.3,.2]))

Continue learning

Deep Learning: Sequences, Representations & Financial Decisions — all lessons
  1. Neural networks and backpropagation from first principles
  2. Causal windows, temporal convolutions and order-book tensors
  3. Recurrent networks and LSTM gates
  4. Attention and transformers: which history can the model use?
  5. Forecast loss versus trading loss, turnover and differentiable decisions
  6. Fine-tuning, financial text and foundation-model contamination

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations