Trading Dev AcademyFree quant education

Free lesson · Deep learning for finance

Forecast loss versus trading loss, turnover and differentiable decisions

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Training on return prediction and training directly on portfolio outcomes are different choices. A decision-aware objective can reflect costs and risk, but it can also exploit unrealistic simulator assumptions more directly.

Symbols, units & horizon
  • z_t: network position score
  • w_t: bounded signed exposure between −1 and 1
  • r_(t+1): future return after decision
  • c: cost per absolute exposure change
  • R: net strategy return on stated capital
  • L: loss to minimise
  • λ: risk-penalty scale for the return units
  • Var: population variance over the training sequence in this example

When and why to use this

Use a decision-aware loss only when the simulator and capital constraints are sufficiently explicit to make the objective economically meaningful.

Squared forecast error treats all observations according to magnitude; log loss targets probabilities; ranking losses target ordering. A Sharpe-based loss couples many observations through estimated mean and volatility, and can be unstable in small or low-variance batches. Never compare objectives using the same period repeatedly until one wins.

A direct position network can output a bounded weight and optimise net returns with a risk penalty. Include turnover on changes from previous holdings, preserving sequence order. Randomly shuffling individual returns can make transaction-cost computation meaningless. Smooth approximations to absolute turnover are optimisation conveniences, not actual fee schedules.

Compare forecast-then-allocate with direct allocation using the same outer evaluation, capital and costs. Track gross return, net return, turnover, tail losses and constraint violations. An attractive differentiable objective does not establish executable fills or adequate uncertainty estimates.

wt=tanh⁡zt,Rt+1=wtrt+1−c|wt−wt−1|,L=−R‾+λVar⁡(R)
Model assumptions, derivation and arithmetic

Forecast loss versus trading loss, turnover and differentiable decisions

  1. Map the unconstrained score through tanh to enforce a simple exposure bound.
  2. Compute future held return and deduct the cost of changing from previous exposure.
  3. Average net returns and calculate mean squared deviations around that average. Negate the mean and add λ times variance. The absolute cost has a kink at zero turnover; implementations need a documented subgradient or approximation.
Work it by hand

Fixed weights [.5,.5], returns [.01,−.002], cost .001 and initial weight 0 give net [.0045,−.001]. Mean=.00175 and variance=.0000075625. With λ=10, loss=−.001674375.

Apply it in a strategy

  • Choose a target loss for a specific decision and include realistic cost accounting during evaluation.
  • Train on ordered sequences with fixed initial-state conventions and independently enforced bounds.
  • Compare direct and two-stage policies through outer future blocks and cost/latency stresses.

Research deliverable

Report a loss-to-ledger reconciliation showing exactly how each training output becomes a held position and a net return.

A staged volatility-forecasting project

Begin with the variance and GARCH walkthrough. Define one future volatility target and when it becomes observable. Keep rolling variance and GARCH as baselines. Train a small sequence model on the same permitted data, with preprocessing fitted inside each training window.

Compare forecast error first, then feed each forecast into the same constrained risk-sizing rule. Hold the expected-return signal, costs and exposure limits fixed to isolate the volatility model’s contribution. Record turnover and tail outcomes as well as predictive fit. Only add cross-asset graphs or hybrid architectures when an ablation demonstrates useful incremental information.

Supplied review and recent empirical check · 12 September 2026

Ogunruku’s Advanced deep learning approaches for forecasting financial market volatility (GSC Advanced Research and Reviews, June 2025) is a narrative review. The publisher PDF returned HTTP 403; the author-uploaded full-text sections on architectures, applications and limitations were reviewed instead. It motivates a staged volatility project but supplies no single reproducible market/sample/cost specification adopted here. Its broad performance assertions are not treated as verified trading results.

Further reading: Requested GSC review · publisher DOI ↗

Further reading: Author-uploaded full text reviewed ↗

For a recent primary empirical comparison, see the 2026 volatility-forecast checkpoint. Predictive accuracy and portfolio benefit are distinct evaluation targets.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def decision_loss(weights,returns,cost,risk_penalty,initial=0):
    if not weights or len(weights)!=len(returns) or min(cost,risk_penalty)<0: raise ValueError("Aligned nonempty inputs required")
    net=[]; previous=initial
    for w,r in zip(weights,returns):
        if abs(w)>1: raise ValueError("Exposure bound exceeded")
        net.append(w*r-cost*abs(w-previous)); previous=w
    mean=sum(net)/len(net)
    variance=sum((x-mean)**2 for x in net)/len(net)
    return net,-mean+risk_penalty*variance

print(decision_loss([.5,.5],[.01,-.002],.001,10))

Continue learning

Deep Learning: Sequences, Representations & Financial Decisions — all lessons
  1. Neural networks and backpropagation from first principles
  2. Causal windows, temporal convolutions and order-book tensors
  3. Recurrent networks and LSTM gates
  4. Attention and transformers: which history can the model use?
  5. Forecast loss versus trading loss, turnover and differentiable decisions
  6. Fine-tuning, financial text and foundation-model contamination

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations