Free lesson · Deep learning for finance
Fine-tuning, financial text and foundation-model contamination
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Transfer learning starts from a representation learned elsewhere. It may reduce task-specific data needs, but only if the earlier training information and the new evaluation are compatible. Historical finance tests are especially sensitive to unknown pretraining cutoffs.
Symbols, units & horizon
- W: frozen d_out×d_in weight matrix
- A: r×d_in learned low-rank factor
- B: d_out×r learned factor
- W′: adapted effective weight
- r: adapter rank, not return
- N_adapter: trained parameter count excluding optional biases/scales
- BA: low-rank update with rank at most r
When and why to use this
Use transfer learning for representations or document-derived features when provenance, availability and prospective validation can be established.
For numerical sequences, compare frozen features with a small trained head, partial-layer fine-tuning and full retraining. For financial text, start with timestamped document extraction: publication time, retrieval time, company mapping, event type, magnitude and confidence. A language model’s fluent explanation is not an alpha estimate.
Adapters such as low-rank updates modify a small parameter subspace. Freezing most weights can reduce training cost and overfitting capacity, but adapter rank, data selection, prompts and checkpoints remain hyperparameters. Use disjoint chronological validation and version the base model as well as the adapter.
A foundation model may already know later events from its pretraining corpus. If the cutoff or data provenance is unknown, do not present a historical backtest as clean out-of-sample evidence. Use prospective timestamped evaluation for the incremental feature and compare against simple text dictionaries, linear embeddings or structured-data baselines.
Fine-tuning, financial text and foundation-model contamination
- Represent a weight update as a product of a narrow and a wide matrix: BA has the same dimensions as W.
- Count r×d_in parameters in A and d_out×r in B; add them to get r(d_in+d_out).
- During inference apply Wx+B(Ax). This is algebraically equivalent to (W+BA)x; a scaling factor can be included but is omitted in this teaching convention.
For a 100×100 matrix and rank 4, the adapter trains 4(100+100)=800 parameters instead of 10,000 for a full update. Fewer parameters do not guarantee better generalisation.
Apply it in a strategy
- Record base-model version, known training cutoff, document timestamps and the exact fine-tuning corpus.
- Compare frozen-head, adapter and simpler baselines under a fixed chronological experiment and compute budget.
- Validate extracted facts against source documents and measure incremental net strategy value prospectively when historical contamination cannot be ruled out.
Research deliverable
Create a model-and-data lineage sheet and a prospective evaluation plan for one timestamped financial feature.
Research checkpoint · reviewed 11 September 2026
These sources inform the questions to test. A result is conditional on its data, simulator and evaluation design. The examples in this module are teaching calculations, not reproductions of the reported experiments.
Architecture evidence and its boundaries. The companion ML section records the reviewed 2026 financial sequence benchmark, including its gross-return optimisation, sample-label discrepancy and validation-based seed selection. Use that source to motivate an experiment, then compare architectures under your own causal data, cost and compute constraints. The neural operations taught here are standard mathematical building blocks, not a reproduction of a paper’s strategy.
Further reading: Read the financial ML research checkpoint ↗
Research sources, review dates and limitations
Extend the research question
Freeze the volatility target, availability rule and evaluation dates before comparing architectures. Reconcile sample labels and distinguish gross metrics from cost-adjusted results.
Continue with the connected research module →
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
import numpy as np
def adapted_linear(weight,a,b,x):
w=np.asarray(weight,float); a=np.asarray(a,float); b=np.asarray(b,float); x=np.asarray(x,float)
if (b@a).shape!=w.shape: raise ValueError("Adapter dimensions must match weight")
return w@x+b@(a@x)
def adapter_parameters(input_dim,output_dim,rank):
if min(input_dim,output_dim,rank)<1: raise ValueError("Positive dimensions required")
return rank*(input_dim+output_dim)
print(adapter_parameters(100,100,4),adapted_linear([[1,0],[0,1]],[[1,0]],[[.1],[.2]],[2,3]))Continue learning
Deep Learning: Sequences, Representations & Financial Decisions — all lessons- Neural networks and backpropagation from first principles
- Causal windows, temporal convolutions and order-book tensors
- Recurrent networks and LSTM gates
- Attention and transformers: which history can the model use?
- Forecast loss versus trading loss, turnover and differentiable decisions
- Fine-tuning, financial text and foundation-model contamination
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations