Free lesson · Charts & patterns
Quantitative recognition III: causal encoders and contrastive learning
Open interactive lessonPractice calculationsExplore labs
Start with the idea
An encoder turns a window into a small vector. Train nearby vectors to represent a chosen kind of similarity; define that similarity before choosing a neural network.
Symbols, units & horizon
- z_a,z_p,z_n: anchor, positive and negative encoder vectors with matching coordinates
- || ||₂²: sum of squared coordinate differences
- γ: nonnegative margin in squared embedding units
- L: nonnegative triplet loss
- embeddings: dimensionless in this teaching example
When and why to use this
Use embeddings for motif retrieval, similarity features or a downstream classifier. Inspect nearest training neighbors to understand what the representation actually groups.
- Begin with a feature vector and a linear model. Add a 1D CNN only when local ordered combinations are useful beyond that baseline.
- A causal convolution uses the current and earlier positions. Dilation spaces its inputs farther apart; a larger receptive field sees longer context.
- A completed-window encoder may inspect every position inside that past window. It must never consume the target continuation. Token-level forecasting requires an appropriate causal mask.
- Contrastive training uses an anchor, a related positive and a contrasting negative. A margin asks the negative to be farther away than the positive.
- Define positives from a valid augmentation or training-only label. Small permitted scale changes may preserve a task; time reversal, sign reversal and changing candle bodies often change the meaning.
- Use hard negatives with similar trend and volatility but different local structure. Do not choose negatives by examining held-out future returns.
- A shapelet-distance vector can feed logistic regression or an SVM. A pooled causal-CNN embedding can feed the same classifier; compare both under matched tuning budgets.
- Pretraining, clustering and hard-negative mining must be restricted to allowed history. Self-supervised learning can still leak future regimes through its training corpus.
- Ablate the learned branch: raw OHLCV only, template distances only, encoder only, then fused context. Report seed variability and calendar-period stability.
Quantitative recognition III: causal encoders and contrastive learning
- Anchor=[1,0], positive=[.8,.2]. Differences=[.2,−.2], so squared distance=.04+.04=.08.
- Negative=[0,1]. Differences=[1,−1], so squared distance=2.
- With margin .2, the expression is .08−2+.2=−1.72. Taking max with zero gives loss 0.
- Swap the positive and negative: 2−.08+.2=2.12. Training would penalize this reversed ordering.
Loss zero means this triplet satisfies its margin. It neither proves all classes separate nor estimates the probability of a profitable trade.
Use the rule
- Fix a valid positive/negative construction.
- Train on past windows and fit a simple downstream head.
- Stress renderer, trend and regime changes; retain complexity only after an incremental held-out benefit.
Before moving on
Document the encoder receptive field, training dates, augmentations and false-neighbor examples.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def triplet_loss(anchor,positive,negative,margin=.2):
if not anchor or len(anchor)!=len(positive) or len(anchor)!=len(negative) or margin<0:
raise ValueError("Matching vectors and nonnegative margin required")
squared=lambda a,b: sum((x-y)**2 for x,y in zip(a,b))
return max(0,squared(anchor,positive)-squared(anchor,negative)+margin)
print(triplet_loss([1,0],[.8,.2],[0,1]))Continue learning
Candles, Structures & Pattern Research — all lessons- First principles: what a candle actually records
- Doji, hammer, shooting star and long-body bars
- Engulfing, inside bars and multi-candle sequences
- Trends, ranges, breakouts and chart structures
- Indicators as arithmetic: ATR, moving averages, RSI and bands
- Known strategy families: from chart idea to complete rules
- Pattern recognition: rules, features, shapelets and image models
- Quantitative recognition I: build a causal candle feature table
- Quantitative recognition II: shapelets and constrained dynamic time warping
- Quantitative recognition III: causal encoders and contrastive learning
- Quantitative recognition IV: GAF images, CNNs, transformers and visual-model audits
- Quantitative recognition V: calibrate, abstain and test the complete strategy
- Research review: what the evidence does and does not establish
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations