Trading Dev AcademyFree quant education

Free lesson · Time series

Hidden Markov models: predict, observe, update

Open interactive lessonPractice calculationsExplore labs

Start with the idea

You can observe a thermometer without observing the underlying condition directly. A hidden Markov model combines a changing hidden state with observations that are more or less likely in each state.

Symbols, units & horizon
  • q_(t−1,i): previous filtered probability of state i
  • A_ij: one-step transition probability from i to j
  • y_t: observation available at time t
  • b_j(y_t): likelihood of this observation under state j
  • u_j: unnormalized evidence weight
  • k: index over all destination states in the normalizer
  • q_(t,j): updated filtered probability
  • discrete likelihoods: dimensionless, continuous densities: units cancel in normalization

When and why to use this

Use filtered regime probabilities as uncertain risk or feature inputs. Compare against observable volatility rules before using them to change allocations or execution behavior.

A hidden state is a model category you cannot directly read from the data. An emission is the observation distribution conditional on that state. For a first example, use the observed category “large move” instead of a complicated multivariate Gaussian density. Neither the category nor the state should be called profitable by definition.

Perform two operations in order. First use the transition model to predict the state probabilities before seeing today’s observation. Then weight each predicted state probability by the likelihood of the observed evidence under that state, and normalize the weights to sum to one. This is a Markov prediction followed by Bayes’ rule.

Filtering estimates the current state using observations available through now. Smoothing revises past state probabilities using later observations. Viterbi finds a most likely whole path under a fitted model, which differs from choosing the most probable state separately at every date. Full-sample smoothing or decoding is useful for retrospective interpretation but cannot be a live feature for an earlier decision.

Learning an HMM commonly alternates estimates of hidden-state responsibility with parameter updates, called expectation maximization or Baum–Welch. Keep that fitting inside each training window. During refits the numeric state labels may swap: state 0 today need not describe state 0 next month. Align economic meaning using only prior information and measure the turnover induced by refits.

uj=bj(yt)∑iqt−1,iAij,qt,j=uj∑kuk
Model assumptions, derivation and arithmetic

Hidden Markov models: predict, observe, update

  1. Predict today’s state weights by multiplying yesterday’s probability row by A. For the example these are [.78,.22].
  2. Multiply the weights by the large-move likelihoods [.1,.6], giving [.078,.132].
  3. Add the unnormalized weights to obtain .21 and divide each by .21. The posterior is [.371429,.628571].
  4. Use that posterior only after this observation is available. Repeating prediction and normalization implements a scaled forward filter without using future observations.
Work it by hand

Yesterday’s calm/stress probabilities [.8,.2] and transitions [[.9,.1],[.3,.7]] imply prior stress=.22. Seeing a large move with likelihood .1 in calm and .6 in stress raises posterior stress to .132/.21≈.628571.

Apply it in a strategy

  • Start with the two-state discrete example and verify every probability.
  • Fit transition and emission models on past data, then filter sequentially.
  • Test probability stability, label alignment, turnover and net outcomes against a simple regime baseline.

Research deliverable

Produce a table with prior, likelihood, unnormalized weight and posterior for each state at each step.

Sources & evidence · reviewed 12 September 2026

Reviewed 12 September 2026: the supplied URL served the draft dated 19 August 2026. Markov chains, emissions, forward inference, decoding and parameter-learning sections reviewed. This is a textbook treatment in language processing, not a finance backtest. Our trading analogy uses causal filtering; the simple discrete filter is not a full Gaussian-HMM fitter.

Further reading: Jurafsky & Martin · Speech and Language Processing, Appendix A: Hidden Markov Models ↗

The allocation lesson records the supplied 2026 regime-aware investing preprint and explains its research limitations.

Further reading: Regime-aware investing: implementation checkpoint ↗

Research sources, review dates and limitations

Extend the research question

Annotate observation, release, revision and decision timestamps. A lagged series can still contain hindsight information if its values were revised later.

Continue with the connected research module →

Connect the ideas: Information and decision time

Retrieve: Use only information available when the decision is made.

Check the change: Observation dates, release delays, revisions and label maturity require different availability checks.

Statistics → Research & backtests → Point-in-time data → Research & robust tuning → Financial machine learning → Putting it all together

Explain it yourself: Does shifting a feature by one row guarantee that it was available?

Self-assessed. Write your explanation before opening this comparison.

No. A revised value or delayed release may still contain unavailable information. Audit actual availability timestamps and fit preprocessing inside each training window.

Connect the ideas: Dependence and diversification

Retrieve: Joint behavior matters when combining uncertain outcomes.

Check the change: Correlation, cointegration, covariance and event dependence answer different questions.

Statistics → Linear algebra → Portfolio construction → Portfolio management theory → Prediction foundations

Explain it yourself: Why does a highly correlated pair not automatically provide a converging spread?

Self-assessed. Write your explanation before opening this comparison.

Correlation measures co-movement under a chosen sample and horizon. It does not establish a stationary combination or contractual convergence; those require separate definitions, tests and implementation checks.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

from math import isfinite

def hmm_update(previous,transition,likelihood):
    n=len(previous)
    if not n or len(likelihood)!=n or len(transition)!=n or any(len(row)!=n for row in transition): raise ValueError("State dimensions must match")
    for row in [previous]+list(transition):
        if not all(isfinite(x) and 0<=x<=1 for x in row) or abs(sum(row)-1)>1e-9: raise ValueError("Invalid probability row")
    if not all(isfinite(x) and x>=0 for x in likelihood): raise ValueError("Nonnegative finite likelihoods required")
    prior=[sum(previous[i]*transition[i][j] for i in range(n)) for j in range(n)]
    # Rescale likelihoods first; relative weights and posterior are unchanged.
    scale=max(likelihood)
    if scale==0: raise ValueError("Impossible evidence")
    weights=[p*(b/scale) for p,b in zip(prior,likelihood)]
    total=sum(weights)
    if total==0: raise ValueError("Evidence has zero probability under the prior")
    return prior,[x/total for x in weights]

print(hmm_update([.8,.2],[[.9,.1],[.3,.7]],[.1,.6]))

Continue learning

Time Series Analysis — all lessons
  1. Start with time order, lags and differences
  2. Before GARCH: mean, shocks and changing variance
  3. Stationarity: the assumption every test makes and every market breaks
  4. Autocorrelation: momentum, mean reversion, or coin flips
  5. GARCH: volatility clusters, and you can model the cluster
  6. Cointegration: a stationary relationship to test
  7. Forecast horizons, EWMA, and model diagnostics
  8. Markov chains: a two-state model you can calculate by hand
  9. Hidden Markov models: predict, observe, update

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations