Advanced C · Systematic and quantitative trading
From market questions to reproducible code, controlling visibility, selection bias, and transaction costs.
- Chapter 22 · SignalHow do you perform Feature Engineering and combine multiple signals?
- Chapter 23 · Market regimeHow can a Regime Model adjust or pause strategies in different conditions?
- Chapter 25 · BacktestWhat traps differ between event-driven and vectorized backtests?
- Chapter 26 · Data biasHow can you construct point-in-time datasets without survivorship bias?
- Chapter 27 · In-sample and out-of-sampleHow do Walk-forward and Purged Validation prevent information leakage?
- Chapter 29 · RobustnessHow can you test Parameter Stability, and distinguish plateaus from isolated peaks?
Market Question before the model
Study whether past information improves next-period net-return decisions, not which curve looks best. The research workbook supplies a complete standard-library Python experiment: fixed-seed inputs, historical features, chronological splits, training selection, costs, final OOS, block bootstrap, and path drawdowns. It demonstrates discipline without asserting synthetic-data market edge.
Save raw inputs and definitions before derived features. Predictions, positions, and fills differ. Correct predictions may have no net returns because of costs, capacity, or unavailable execution.
Time Series: clocks are the first constraint
Preserve observation order. Rolling windows read only data published before decisions. Record label starts/ends; remove training labels crossing test boundaries. Temporally overlapping trade results are not independent experiments.
Delay one quote's publication a period and check signal movement. Unchanged original-time execution leaks the future. Require event_time, available_at, decision_time, and label_end per row, with visibility before decisions. Missing quotes require stops/skips, not future interpolation.
Factor Research: explain expected compensation first
Factors are shared explanatory variables cross-sectionally or over time. Size, liquidity, and momentum may proxy risks; correlation is not causation. Define mechanism, direction, horizon, and costs before ranks or regressions.
Compare a past-return candidate with market-only holding; report group returns, information coefficients, and turnover. Fix group counts before validation. If only illiquid assets work, inspect unachievable prices. Retain delistings, factor versions, and risk controls. Gross correlation is not Alpha.
Signal Combination: weights overfit too
Signals from one historical window need not add independent information when averaged. Align clocks/units, then compare equal, advance-fixed, and training-fitted weights. Weights are parameters and enter attempts/OOS.
Ablate momentum and flow individually. Keep complexity only when combined signals add stable improvement on identical test, risk, and costs. Submit correlation matrices, weight versions, and individual/combined comparisons. Missing inputs follow frozen rules, not hindsight reweighting toward winners.
Feature Engineering: every transform needs a training boundary
Returns, volatility, rolling quantiles, and volume ratios are derived features. Normalization, clipping, and imputation contain parameters. Fit only on training, apply to validation/testing, and never use full-sample means to anticipate future distributions.
Add a large test move and verify unchanged training normalization. Inject missing input and inspect predefined substitution or no signal. Retain transform order, fit windows, units, and exception rules. Clipping real extremes alters tails; retain raw risk series too.
Cross-sectional Strategy: compare the original universe
Ranking contemporaneous tradable assets requires rebuilding then-current universes. Selecting similarly priced but riskier assets may merely raise Beta. Equal long/short notional is not automatic factor neutrality.
Construct members period by period from teaching listings, delistings, and missing data, excluding future membership. Output ranks, exclusions, and notional. Check liquidity floors, borrow availability, asset caps, and industry concentration; permit flat positions when constraints fail.
Alternative Data: labels and revisions are data too
On-chain labels, news, flows, and social text include delays, revisions, expanding coverage, and licensing constraints. Event time is not tradable time. Today's final revisions cannot automatically be historical point-in-time data.
Save original/revised event versions, generate both signals, and quantify changes. Without historical vintages, limit conclusions to exploration. Submit sources, first visibility, versions, licensing boundaries, and missing rates. Code must require no account keys, customer data, or live-trading permission.
Statistical Arbitrage: test relationship stability
Statistical arbitrage uses predictable spread/residual repair while bearing model/financing risk. Return correlation is not price cointegration; even training stability may fail as institutions, supply, and participants change.
Estimate a teaching pair in training, then inspect test residual means, variances, and recovery times. Add a permanent level jump and inspect stops. Submit economics, break detection, both-leg costs, and holding limits. Reestimating until failures disappear cannot establish uninterrupted validity.
ML Signal: complexity requires incremental OOS evidence
Machine learning fits nonlinear relations and noise. Begin with constants, linear models, or simple-rule baselines. Feature selection, model classes, and hyperparameters belong within inner training/validation.
Freeze labels/costs and compare a baseline with one candidate. Report errors, directional calibration, turnover, and net profit; none replaces another. Sparse samples may justify retaining the baseline. Include training seeds, software versions, features, and failed attempts, not only a model file.
Regime Model: labels must be known then
Past volatility, spread, or liquidity can bound applicability. Hindsight labeling losers “crisis” and filtering leaks outcomes. Hidden-state models must filter with past data, not pass full-sequence smoothed probabilities as real-time estimates.
Add a pause based only on prior volatility, comparing net returns and reduced-trading opportunity costs. Report transition evidence, hysteresis/cooldowns, and mistakes. State defaults with inadequate data; never select the retrospectively best branch.
Optimization: objectives cannot define reality
Maximizing training returns may concentrate noise, thin liquidity, or extreme leverage. Include costs, positions, turnover, concentration, and capital constraints. Omitted tails may make “optimal” mean efficiently bearing unpriced risk.
Restrict grids to a small preregistered set and plot neighborhoods, not only peaks. Select stable regions before final-test lock. Report searches, binding constraints, failed solutions, and simple baselines. Post-test tuning starts new research; the test becomes used data.
Walk-forward: move research through time
Each fold trains on past windows, validates later, then tests later still. Predefine rolling/expanding windows and refit transforms/parameters each fold. Join genuinely tested returns to approximate operating order.
Draw three fold boundaries with readable data and model versions. Tried window choices still require a final independent test. Never double-count overlapping test-day returns or insert validation performance into the final curve.
Purged Validation: isolate overlapping labels
Remove training labels whose time intervals overlap validation/testing. Cross-validation using future training samples also needs an embargo after validation to limit neighboring feature/label information transfer. Length needs a dependence rationale, not a universal percentage.
The workbook trains on past and tests future, with two-period labels and removed boundary crossings; it does not pretend to be random K-fold. Lengthen holding and verify more purged samples. Submit split indices and leakage assertions; the word “purged” is not completion.
Project: reproducible research code
Run the workbook code, retaining inputs, parameters, seeds, commands, and complete outputs. Replace teaching inputs with sourced, availability-timed data and rerun the pipeline. Without tradable inputs, conclusions remain engineering exercises without promotion.
Completion means another reader can run code from an empty directory; changed costs affect net returns; changed test data leaves training selection fixed; overlaps purge; and missing, insufficient, or nonfinite inputs explicitly reject. Submit Memo, code, inputs, and Backtest Report to the institutional pipeline.