Trading Dev AcademyFree quant education

Free lesson · Charts & patterns

Pattern recognition: rules, features, shapelets and image models

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Recognition is a measurement problem first and a prediction problem second. Make the label and distance unambiguous before training a model to optimise them.

Symbols, units & horizon
  • xⱼ: sequence observation at position j
  • sⱼ: template observation, not SD here
  • μ_train,s_train: mean and positive SD from the permitted training sample
  • Tilde: normalised value
  • m: common sequence length
  • d(x,s): root-mean-square aligned distance
  • TP,FP,FN: true positives, false positives and false negatives
  • p: probability of a winning executed trade
  • W̄,L̄: positive mean win and loss amounts
  • c̄: mean trade cost in matching currency
  • ê: estimated net currency expectancy

When and why to use this

Use simple detectors and explicit OHLCV features as baselines for more complex shapelet or image research. Choose performance measures that match the intended trading decision.

A rule detector is an auditable baseline: use body fractions, wick fractions, gaps, relative volume, lagged volatility and distance to a prior boundary. Continuous features often retain more information than a yes/no label. Add context only if it is observable at the decision time; scale using training data or a trailing window, never the complete sample.

x~j=xj−μtrainstrain,d(x,s)=1m∑j=1m(x~j−s~j)2
Algebra and arithmetic

Compute an aligned template distance

  1. Subtract the training mean from each value, then divide by the positive training SD. Apply the intended, documented normalisation to the template too.
  2. For normalised sequences (0,1,0) and (0,.5,0), differences are (0,.5,0). Square, sum, divide by m=3 and take the square root.
Work it by hand

Distance=√(.25/3)=.288675. This simple metric uses fixed time alignment; it is not DTW.

A shapelet is a short subsequence used as a pattern template. The displayed distance compares aligned, equally long sequences after a specified normalisation. Dynamic time warping (DTW) instead permits constrained alignments between sequences; unconstrained warping may match economically unrelated timing. Kernel smoothing and centred filters can also use future values, so replace them with causal versions or delay their availability in a trading test.

Image models learn from rendered charts. A chart image is a transformation of the underlying OHLCV window, so it contains no additional market observations beyond those inputs. Compare image models with tabular OHLCV features under matched data, horizon, tuning budget and capacity. Check sensitivity to axis scaling, colours, padding and chart-window length; these rendering choices should not accidentally encode the target.

precision=TPTP+FP,recall=TPTP+FN,e^=pW‾−(1−p)L‾−c‾
Algebra and arithmetic

Count correct signals and price their payoffs

  1. Of all predicted positives, TP+FP, the correct fraction is TP/(TP+FP). Of all actual positives, TP+FN, the detected fraction is TP/(TP+FN). Zero denominators mean undefined metrics.
  2. For trade win probability p, average win W and positive average loss L, net expectancy is pW−(1−p)L−c. This adds payoff size and costs that classification metrics omit.
Work it by hand

TP=12,FP=8,FN=18 gives precision=.60 and recall=.40. With W=$10,L=$20,c=$1, a .60 win rate still gives EV=6−8−1=−$3.

Define the target before fitting: for example, next-horizon return above a cost-aware threshold, or a fully specified trade outcome. Use chronological train/validation/test periods; remove label overlap at boundaries; include delisted assets and a point-in-time universe. Report precision, recall, base rate, conditional payoff, turnover, drawdown and net portfolio returns. A detector can identify a shape perfectly while predicting future returns no better than a baseline.

Evaluate context variables with ablations: compare the same baseline with and without volatility, trend, volume, spread, time of day and events. Save every tried specification, use dependence-aware uncertainty, and assess stability across assets and market conditions. Select a simple model unless the extra complexity produces repeatable economic value.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

from math import sqrt

def normalized_distance(sequence, template, training_mean, training_sd):
    if len(sequence)!=len(template) or not sequence or training_sd<=0:
        raise ValueError("Matching nonempty sequences and positive training SD required")
    # Same fixed training scale for both sequences in this version.
    x=[(v-training_mean)/training_sd for v in sequence]
    s=[(v-training_mean)/training_sd for v in template]
    return sqrt(sum((a-b)**2 for a,b in zip(x,s))/len(x))

def detection_metrics(tp,fp,fn):
    if min(tp,fp,fn)<0: raise ValueError("Counts must be nonnegative")
    return (None if tp+fp==0 else tp/(tp+fp),
            None if tp+fn==0 else tp/(tp+fn))

def net_expectancy(win_probability, mean_win, mean_loss, mean_cost):
    return win_probability*mean_win-(1-win_probability)*mean_loss-mean_cost

print(normalized_distance([0,1,0],[0,.5,0],0,1))
print(detection_metrics(12,8,18), net_expectancy(.6,10,20,1))

Continue learning

Candles, Structures & Pattern Research — all lessons
  1. First principles: what a candle actually records
  2. Doji, hammer, shooting star and long-body bars
  3. Engulfing, inside bars and multi-candle sequences
  4. Trends, ranges, breakouts and chart structures
  5. Indicators as arithmetic: ATR, moving averages, RSI and bands
  6. Known strategy families: from chart idea to complete rules
  7. Pattern recognition: rules, features, shapelets and image models
  8. Quantitative recognition I: build a causal candle feature table
  9. Quantitative recognition II: shapelets and constrained dynamic time warping
  10. Quantitative recognition III: causal encoders and contrastive learning
  11. Quantitative recognition IV: GAF images, CNNs, transformers and visual-model audits
  12. Quantitative recognition V: calibrate, abstain and test the complete strategy
  13. Research review: what the evidence does and does not establish

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations