Trading Dev AcademyFree quant education

Free lesson · Charts & patterns

Quantitative recognition III: causal encoders and contrastive learning

Open interactive lessonPractice calculationsExplore labs

Start with the idea

An encoder turns a window into a small vector. Train nearby vectors to represent a chosen kind of similarity; define that similarity before choosing a neural network.

Symbols, units & horizon
  • z_a,z_p,z_n: anchor, positive and negative encoder vectors with matching coordinates
  • || ||₂²: sum of squared coordinate differences
  • γ: nonnegative margin in squared embedding units
  • L: nonnegative triplet loss
  • embeddings: dimensionless in this teaching example

When and why to use this

Use embeddings for motif retrieval, similarity features or a downstream classifier. Inspect nearest training neighbors to understand what the representation actually groups.

  • Begin with a feature vector and a linear model. Add a 1D CNN only when local ordered combinations are useful beyond that baseline.
  • A causal convolution uses the current and earlier positions. Dilation spaces its inputs farther apart; a larger receptive field sees longer context.
  • A completed-window encoder may inspect every position inside that past window. It must never consume the target continuation. Token-level forecasting requires an appropriate causal mask.
  • Contrastive training uses an anchor, a related positive and a contrasting negative. A margin asks the negative to be farther away than the positive.
  • Define positives from a valid augmentation or training-only label. Small permitted scale changes may preserve a task; time reversal, sign reversal and changing candle bodies often change the meaning.
  • Use hard negatives with similar trend and volatility but different local structure. Do not choose negatives by examining held-out future returns.
  • A shapelet-distance vector can feed logistic regression or an SVM. A pooled causal-CNN embedding can feed the same classifier; compare both under matched tuning budgets.
  • Pretraining, clustering and hard-negative mining must be restricted to allowed history. Self-supervised learning can still leak future regimes through its training corpus.
  • Ablate the learned branch: raw OHLCV only, template distances only, encoder only, then fused context. Report seed variability and calendar-period stability.
ℒ=max⁡(0,‖za−zp‖22−‖za−zn‖22+γ)
Model assumptions, derivation and arithmetic

Quantitative recognition III: causal encoders and contrastive learning

  1. Anchor=[1,0], positive=[.8,.2]. Differences=[.2,−.2], so squared distance=.04+.04=.08.
  2. Negative=[0,1]. Differences=[1,−1], so squared distance=2.
  3. With margin .2, the expression is .08−2+.2=−1.72. Taking max with zero gives loss 0.
  4. Swap the positive and negative: 2−.08+.2=2.12. Training would penalize this reversed ordering.
Work it by hand

Loss zero means this triplet satisfies its margin. It neither proves all classes separate nor estimates the probability of a profitable trade.

Use the rule

  • Fix a valid positive/negative construction.
  • Train on past windows and fit a simple downstream head.
  • Stress renderer, trend and regime changes; retain complexity only after an incremental held-out benefit.

Before moving on

Document the encoder receptive field, training dates, augmentations and false-neighbor examples.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

def triplet_loss(anchor,positive,negative,margin=.2):
    if not anchor or len(anchor)!=len(positive) or len(anchor)!=len(negative) or margin<0:
        raise ValueError("Matching vectors and nonnegative margin required")
    squared=lambda a,b: sum((x-y)**2 for x,y in zip(a,b))
    return max(0,squared(anchor,positive)-squared(anchor,negative)+margin)

print(triplet_loss([1,0],[.8,.2],[0,1]))

Continue learning

Candles, Structures & Pattern Research — all lessons
  1. First principles: what a candle actually records
  2. Doji, hammer, shooting star and long-body bars
  3. Engulfing, inside bars and multi-candle sequences
  4. Trends, ranges, breakouts and chart structures
  5. Indicators as arithmetic: ATR, moving averages, RSI and bands
  6. Known strategy families: from chart idea to complete rules
  7. Pattern recognition: rules, features, shapelets and image models
  8. Quantitative recognition I: build a causal candle feature table
  9. Quantitative recognition II: shapelets and constrained dynamic time warping
  10. Quantitative recognition III: causal encoders and contrastive learning
  11. Quantitative recognition IV: GAF images, CNNs, transformers and visual-model audits
  12. Quantitative recognition V: calibrate, abstain and test the complete strategy
  13. Research review: what the evidence does and does not establish

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations