Free lesson · Deep learning for finance
Neural networks and backpropagation from first principles
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A neural network composes simple transformations into a flexible function. Training changes its parameters so that predictions reduce a chosen loss. Backpropagation is the chain rule applied to those composed operations.
Symbols, units & horizon
- x: scalar feature in the teaching example
- w,b: weight and bias
- z: pre-activation
- ŷ: predicted scalar
- y: target
- tanh: bounded hyperbolic tangent
- L: squared-error loss
- ∂L/∂w: parameter gradient
- η: learning rate used in the update w_new=w−η gradient
When and why to use this
Use this derivation to understand how financial forecast errors change parameters, and to debug losses or custom differentiable trading components.
A dense layer combines features with weights and a bias, then applies a nonlinear activation. Multiple layers can express interactions that a linear model cannot. More capacity also allows more ways to fit noise, so the first comparison should be a linear or tree baseline on the same dataset.
Think in tensors: a batch has samples, features and possibly time. Track dimensions at every layer. Fit input normalisation on training data; a network cannot repair a target that leaks future prices. Treat learning rate, width, depth, dropout and early stopping as part of the recorded search.
The example below uses one tanh neuron and squared error so the entire gradient is visible. Real deep-learning frameworks automate this bookkeeping, but you should verify a small finite-difference gradient and distinguish training mode from inference mode before trusting a larger model.
Neural networks and backpropagation from first principles
- Differentiate loss with respect to the prediction: ∂L/∂ŷ=ŷ−y.
- Differentiate tanh: ∂ŷ/∂z=1−ŷ². The affine derivative is ∂z/∂w=x.
- Multiply along the dependency chain. A gradient-descent step subtracts η times this gradient; the bias gradient uses the same expression without x.
At x=2, w=0, b=0, y=1, prediction=0 and weight gradient=(−1)×1×2=−2. With η=.1 the updated weight is .2; the updated bias would be .1 if trained simultaneously.
Apply it in a strategy
- Implement and verify one small network on synthetic data with a known relationship.
- Move to causal financial inputs and compare identical out-of-fold decisions with ridge or boosting.
- Inspect learning curves and seed dispersion, keeping the validation-selection budget visible.
Research deliverable
Trace one observation through forward prediction, loss, gradient and update before reporting a larger network’s results.
Research sources, review dates and limitations
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
from math import tanh
def neuron_gradient(x,w,b,target):
prediction=tanh(w*x+b)
chain=(prediction-target)*(1-prediction**2)
return prediction,.5*(prediction-target)**2,chain*x,chain
print(neuron_gradient(2,0,0,1))Continue learning
Deep Learning: Sequences, Representations & Financial Decisions — all lessons- Neural networks and backpropagation from first principles
- Causal windows, temporal convolutions and order-book tensors
- Recurrent networks and LSTM gates
- Attention and transformers: which history can the model use?
- Forecast loss versus trading loss, turnover and differentiable decisions
- Fine-tuning, financial text and foundation-model contamination
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations