Advanced A · Probability and statistics
Turn trading judgments into refutable statistical conclusions through distributions, updating, testing, and simulation.
- Chapter 3 · Candles and volumeWhat do return distributions look like, and why do fat tails distort pattern statistics?
- Chapter 7 · ProbabilityHow do conditional probability and Bayesian updates revise judgments with new evidence?
- Chapter 8 · Expected valueHow can historical samples estimate expectancy, and how large is estimation error?
- Chapter 10 · Variance and drawdownHow can Monte Carlo simulate a strategy's possible drawdown distribution?
Fix the question and observation unit first
Researching “whether next-period net returns after a signal are positive” requires defining a period as a day or trade, signal visibility, costs, and overlapping positions. Ten consecutive signal days during one holding are not ten independent experiments. All returns, probabilities, and thresholds below are teaching assumptions, not market estimates or recommendations.
Work through one Research Memo: define hypotheses and data, calculate, then permit “insufficient evidence” or “rejected.” The research workbook provides copyable templates and runnable code; the chapter map connects Foundation nodes.
Probability Distribution: inspect distributions before averages
Let net return be random variable X. A distribution describes probabilities of outcomes. Discrete outcomes fit a table; continuous price changes fit histograms or empirical quantiles. Identical means may conceal different left tails. Volatility measures dispersion about a mean, not all gap, liquidity, or insolvency risk.
Use four equally probable returns: −3%, −1%, 1%, 3%. Mean is 0%, population variance 0.0005, and population standard deviation approximately 2.2361%. If these estimate an unknown distribution, denominator n−1 gives sample variance 0.0006667. Specify finite-population description versus unknown-population estimation; do not switch denominators for prettier results.
Replace the worst return with −15% and recalculate mean, deviation, worst value, and quantiles. This breaks the small symmetric-movement assumption first. Fitting a normal distribution to sparse data does not remove unseen tails. Submit a distribution plot, sample size, and maximum loss—not a Sharpe-only page.
Conditional Probability: the denominator defines the question
P(A|B) uses occurrences of B as its denominator. In 100 nonoverlapping observations, suppose 40 signals include 24 next-period rises, while 60 nonsignals include 30 rises. Conditional rise frequency is 24/40=60%; overall it is 54/100=54%. This is a sample difference, not proof of causation.
If signals appear only during high volatility, compare signals and nonsignals within the same regime. Otherwise regime differences may masquerade as signal contribution. Rebuild two contingency tables with advance-known high/low-volatility labels, recording missing labels. Acceptance requires denominators, timestamps, and grouping rules; never define “trend days” after future returns.
Bayesian Thinking: update beliefs rather than chase winning streaks
Define H as “the strategy has a positive net edge,” and E as “the preregistered validation result occurs.” Updating P(H|E) requires prior P(H), the likelihood of E under H, and false-positive likelihood under not-H.
P(H|E) = P(E|H) × P(H) /
[P(E|H) × P(H) + P(E|not H) × P(not H)]With P(H)=20%, P(E|H)=60%, P(E|not H)=10%, posterior is 0.12/(0.12+0.08)=60%. Passing validation is not certainty. For a selected winner, the 10% false-positive rate usually cannot simply remain unchanged: model selection too.
Keep likelihoods fixed and change the prior to 5%; posterior is about 24%. Explain the prior change and evidence that would lower it further. List prior sources and sensitivity intervals. Without reasonable likelihoods, keep scenario ranges rather than disguising subjective confidence as precise probabilities.
Expected Value and Variance: estimates have errors too
Expectation weights returns by probabilities; historical averages are estimates. A teaching strategy wins 2R with 40% probability and loses 1R with 60%, paying fixed 0.1R costs: gross expectancy 0.2R, net 0.1R. Net outcomes are 1.9R and −1.1R; variance is 2.16R². Define R beforehand, never from realized losses afterwards.
At costs of 0.3R, net expectancy becomes −0.1R. Fixed costs shift distributions without changing this variance. Real impact relates to volatility, so costs need not be constant. Compare fixed costs with doubled worst-regime costs and explain left-tail changes. Submit net series, cost fields, and uncertainty, not gross win rate alone.
Covariance and Correlation: co-movement is not a shared mechanism
Covariance averages simultaneous products of deviations from means. Correlation divides by both standard deviations, removing units. If either variance is zero, correlation is undefined, not zero. Rising price levels can create spurious correlation; justify returns, spreads, or residuals.
For X=[−1%, 1%] and Y=[−2%, 2%], population covariance is 0.0002 and correlation 1. Different assets provide no diversification in this example. Change Y to [2%, −2%] and inspect the sign. Two points illustrate formulas, not allocation evidence.
Real research synchronizes clocks, explains missing data, and reports rolling/stress windows. Do not treat stale prices in nontrading periods as new observations. List windows, effective overlap, and stress correlations before sizing in Portfolio and risk.
Regression: conditional relationships are what it explains
The linear model r_strategy = α + β × r_market + ε separates market co-movement from the remainder. β is sample sensitivity; α is the intercept after controlling for that factor. Omitted factors, nonlinear option exposure, and wrong clocks distort interpretations. Positive α is not automatically tradable edge.
For X=[−1%, 0%, 1%] and Y=[−1%, 1%, 3%], the line is Y=1%+2X. Three perfectly fitted points verify arithmetic, not future 1% excess returns. Add an outlier to the last Y and compare slopes, then discuss retaining it versus evidence of bad data.
Submit residual plots, train/test β, factor definitions, and fee treatment. Autocorrelated residuals invalidate ordinary IID standard errors. Use time-block resampling or an explicitly appropriate robust method and explain window selection rather than simply switching software options.
Hypothesis Testing: define refutation in advance
Write H0 “mean net return is not positive,” then one-/two-sidedness, statistic, cutoff date, and attempt count. A p-value is the probability of at least such an extreme statistic under the null and model assumptions, not the probability the strategy is false.
Shuffle signal labels while retaining returns and costs to build a no-predictive-signal control. With temporal dependence, pointwise shuffling destroys it; permute time blocks without crossing holdout boundaries and explain exchangeability. Reporting only the minimum p-value after many parameters exaggerates evidence.
Submit attempt registration including failures. Fix a primary hypothesis in advance, mark others exploratory, and move the final candidate into unseen testing. Changing hypotheses while viewing tests does not remain OOS.
Confidence Interval: express estimation uncertainty
With IID observations, sufficient sample size, and an approximately normal mean, intervals use sample mean ± critical value × standard error; standard error is sample deviation divided by square root of sample size. Sparse samples, heavy tails, and continuous holdings break simple approximations.
A 95% confidence interval describes repeated-sampling coverage of the method, not a 95% probability that this fixed interval contains the fixed truth. Economic relevance matters too: a slightly positive interval may not cover unmodeled impact.
Compare trade-by-trade with contiguous-block bootstrap mean intervals on one series. Record block length, seed, and repetitions. Explain retaining “insufficient evidence” when intervals cross zero, effective samples are small, or block choices matter. More simulation repetitions cannot conceal inadequate original data.
Monte Carlo: expose path risk
Reordering identical returns leaves the mean unchanged but changes drawdown. Choose the generating mechanism first: independent sampling models an ideal independent setting, blocks retain local losing streaks, and regime transitions allow persistent stress. Update each step using beginning equity times return, then calculate drawdown from previous peaks.
The research workbook's fixed-seed code prints mean intervals and path maximum-drawdown quantiles. Generated teaching data is not historical validation. Increase block length, costs, or insert a stress loss separately, retaining all results.
No insolvency in finite simulations does not imply impossibility. Historical resampling cannot invent never-observed exchange freezes; use separate stress scenarios. List omitted mechanisms, shutdown thresholds, and restart conditions for institutional governance.
Project: Research Memo
Use the copyable template: question, mechanism, availability, samples/groups, attempt registry, assumptions, net distribution, intervals, counterexamples, next decision. Attach runnable code and inputs, not pictures alone.
Acceptance recalculates three sensitivities: remove an extreme sample, raise costs, and fully isolate testing. Original results remain. A sound submission may reject a strategy; it cannot hide adverse results or present teaching assumptions as external facts.