Free lesson · Financial machine learning
Feature engineering, missingness and training-only transformations
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Features translate a hypothesis into measurable inputs. A robust feature carries a defined unit, window and availability timestamp. Preprocessing is part of the fitted model and must obey the same training boundary.
Symbols, units & horizon
- x_ij: feature j for sample i
- μ_j_train: mean estimated only from training observations
- s_j_train: positive training feature scale
- z_ij: dimensionless standardised feature
- n: training count
- Zero-variance column: remove or explicitly map to zero
- Sample versus population SD: choose and document consistently
When and why to use this
Use training-only preprocessing to make signal features comparable without allowing future distributions to influence the model fit.
Useful inputs can include lagged residual returns, volatility-normalised momentum, spread/depth, volume surprises, carry, valuation, earnings revisions and event age. The mechanism and horizon should justify the input. Hundreds of transformations of the same series are many correlated research trials, not hundreds of independent signals.
Standardisation subtracts a training mean and divides by a training scale. Apply those frozen values to validation and test. Missing-value imputation must be fitted on training data too; add a missingness indicator when absence may be informative, while checking whether it proxies for future survival or vendor coverage.
Cross-sectional ranks use the universe actually available at that timestamp. Fundamental reports have announcement delays and revisions; macro series need historical vintages. Fit winsorisation bounds, PCA and feature-selection rules inside the fold. Retain source timestamps and version IDs alongside the feature matrix.
Feature engineering, missingness and training-only transformations
- Compute the column mean using the training rows. Compute its scale on the same rows.
- Subtract that training mean from each future value and divide by the frozen scale.
- Do not recalculate moments on validation data unless a causal, predeclared online update is part of the deployed procedure.
Training values [1,2,3] have mean 2 and sample SD 1. A later value 4 maps to z=2. Recentring on a batch containing 4 would answer a different, future-informed question.
Apply it in a strategy
- Write a feature dictionary with units, mechanism, data source, window and last required timestamp.
- Fit imputation, scaling and feature selection inside each chronological training fold.
- Stress missing values and stale feeds; compare feature families by out-of-fold ablation.
Research deliverable
Save the fitted preprocessing object with each model version and demonstrate identical transformations in training and inference.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
from statistics import mean,stdev
def fit_scale(values):
if len(values)<2: raise ValueError("At least two training values required")
return mean(values),stdev(values)
def transform(values,center,scale):
if scale<=0: raise ValueError("Remove or handle constant columns explicitly")
return [(v-center)/scale for v in values]
center,scale=fit_scale([1,2,3])
print(transform([4],center,scale))Continue learning
Machine Learning for Quantitative Strategy Development — all lessons- Choose the model’s job: targets, horizons and decision layers
- Feature engineering, missingness and training-only transformations
- Regularised regression: an interpretable alpha baseline
- Logistic classification and cost-aware entry thresholds
- Trees and boosting: nonlinear interactions with controlled complexity
- Calibration, meta-labels and conditional payoff estimation
- Unsupervised learning, clusters and latent risk structure
- From model forecasts to a constrained strategy
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations