Trader OS
Advanced

Advanced research workbook

Copyable research, strategy, risk, execution, and review templates, with a runnable chronological experiment.

How to use it

Complete A · Statistics, B · Strategy, C · Code, D · Risk, E · Execution, and F · Governance in order. Copy templates and reference actual evidence; blank templates are not submissions. Use teaching or public research data only, without trading accounts or orders.

Research Memo

Research ID / author / date / version:
Primary question and mechanism: who pays, and what risk is borne?
Observation / Mechanism / Hypothesis:
Null H0 / refutation / minimum economic significance:
Observation unit / holding period / overlap treatment:
Source / retrieval / first availability / version / license:
Universe / delistings / missing values / label revisions:
Training, validation, final-test boundaries:
All attempts and failures / parameter-selection rule:
Baselines / ablation / cost stress:
Net distribution / count / intervals / tails:
Time dependence / multiple testing / omitted mechanisms:
Conclusion: support, reject, insufficient evidence; cite outputs:
Next step: research, Testing, pause, retire; owner:
Input files / code version / command / output checks:

Select a signal row at random and trace it to then-visible inputs. Even after removing the most favorable explanation, remaining evidence should explain why conclusions hold or fail.

Strategy Object

These 23 fields match the research system. Copy separately per strategy; “same as above” must not hide different legs/costs. Values such as “strict risk control” or “machine learning” alone are not inspectable.

Strategy ID: Stable unique identifier
Name: Readable name and version
Market: Market, venues, timezone
Instrument: Assets, contract specifications, currencies, each leg
Horizon: Signal frequency, maximum holding, label interval
Category: Return mechanism; application context separately
Hypothesis: Refutable claim and outcome metric
Mechanism: Participants, constraints, why compensation exists
Data: Source, availability, version, missing/revision treatment
Features: Historical windows, transforms, units
Signal: Explicit logic, thresholds, invalid-input behavior
Entry: Submission timing, executable prices, unfilled handling
Exit: Normal/time/invalidation exits and unknown orders
Sizing: Target notional, risk scaling, hard caps
Risk: Strategy, portfolio, venue, financing, shutdown limits
Execution: Order types, children, deadlines, latency, remainders
Cost: Fee, Slippage, financing, others; no double counting
Backtest: Inputs, code, baselines, costs, complete outputs
OOS: Isolated boundaries, leakage checks, failures
Regime: Then-known states, applicability, defaults
Failure: Contrary mechanisms, pause triggers
Capacity: Normal/stress capacity, capital, exits
Status: Lifecycle, approver, date, evidence

Backtest Report

Object ID / version / linked Memo:
Data manifest and checksums / clocks / visibility:
Historical universe / missing values / exits:
Rules, parameters, seeds, total attempts:
Fill/funding assumptions / initial capital / costs:
Training, validation, testing / Purge / Embargo rationale:
Gross/net returns, maximum drawdown, turnover, exposure:
Cash/holding baselines / single-signal ablation:
Parameter neighborhoods / regimes / costs / capacity / gap stress:
Period equity / fills / remainders / rejections / cash ledger:
Statistical intervals / dependence assumptions / unsupported risks:
Reproduction steps / expected results / stop conditions:
Conclusion and promotion decision with independent review:

Risk Report

Report timestamp / data cutoff / strategy versions / currency:
Total capital / allocated / cash / available margin:
Strategy net/gross exposure, risk budgets, venues:
Factors / Beta / other factors / nonlinear exposure:
Covariance window / correlations / contributions and sums:
Volatility target / realized volatility / leverage/turnover caps:
Stress inputs, losses, funding gaps, executable exits:
Drawdown / triggered tiers / residual orders and positions:
Concentration / counterparties / shared data/execution dependencies:
Gaps / owners / immediate actions / next review:
Approve/pause decision / restart conditions / independent approver:

Execution Analysis

Parent ID / side / quantity / decision time/price:
Arrival time/benchmark / deadline / terminal mark:
Three execution rules and shared inputs, frozen beforehand:
Every child: time, venue, quantity, fills, remainder, fees:
Total fills / remainder / average / arrival and decision deviations:
Opportunity cost / duration / adverse selection / in-flight risk:
Single-variable price, depth, fee, latency comparisons:
Queue, impact, routing assumptions / unproven items:
Applicability / evidence for retaining or replacing rules:

Strategy Review and quarterly review

Quarter/strategy ID / opening plan / changes:
Nine-component PnL Attribution / independent ledger / Residual investigation:
Signal, cost, capacity, correlation, factor changes:
Mechanism still valid / counterexamples / failed trials:
Risk/discipline events / shutdown and restart records:
Lifecycle decision: retain, downgrade, pause, retire:
Next-quarter capital / risk budgets / research priorities:
Market-structure changes and testable evidence:
Proposer / risk objections / approver / effective time:

Reproducible research code

Download advanced-research.py, or save the code below under that filename and run python3 advanced-research.py. Requires Python 3.10 or newer, with no third-party packages, network, or accounts. It checks inputs, boundaries, future-data isolation, costs, and fixed seeds before JSON output. Failures stop rather than print false success.

Inputs are 300 independent synthetic-period returns generated with a fixed seed—not a market. Decisions occur every two periods and hold for two. Next open equals current close, with no opening gaps; each side costs 10 bp. This simplified cost experiment cannot replace executable-quote backtests. Three Walk-forward folds diagnose development. Final parameters are selected only from validation 180–220; after 220 is reserved for final testing.

always_long_round_trip_block_mean holds long for every two-period interval and pays both sides each interval. It is not long-term buy-and-hold trading only at the entire sample's endpoints.

Each label spans two periods; purge labels crossing training/validation ends. Splits train only on past data, not random K-fold. Future-training cross-validation requires redesigned purge/embargo. Bootstrap resamples contiguous runs of three nonoverlapping test holding blocks to show local-dependence sensitivity; it cannot invent unseen incidents.

Default final testing has 39 observations and 16 held positions. Mean net period return is about −0.42796%, maximum drawdown about 16.26065%. Raising each-side costs to 30 bp gives mean approximately −0.59206%. These are teaching outputs, not history or forecasts. The script records input SHA-256; check seeds/checksums before comparing within reasonable floating-point tolerance.

Change costs, block length, and synthetic seed once each, retaining originals. Do not keep selecting seeds until profitable. Even making final testing rally strongly must leave frozen parameters unchanged. Changing holding periods requires split/nonoverlap changes and invalidates prior validation conclusions.

"""Synthetic teaching research, Python 3.10+, standard library only. No orders or network."""
import hashlib
import json
import math
import random
import statistics

SEED = 20260930
HORIZON = 2
LOOKBACK = 10
CANDIDATES = (0.0, 0.003, 0.006)


def validate(returns):
    if len(returns) < 300 or any(not math.isfinite(x) or x <= -1 for x in returns):
        raise ValueError("need >=300 finite simple returns above -1")


def data(seed=SEED):
    rng = random.Random(seed)
    # Independent synthetic returns: this does not assert market predictability.
    return [rng.gauss(0, 0.015) for _ in range(300)]


def rows(returns):
    validate(returns)
    return [dict(t=t, end=t + HORIZON,
                 feature=sum(returns[t-LOOKBACK+1:t+1]),
                 label=math.prod(1+x for x in returns[t+1:t+HORIZON+1])-1)
            for t in range(LOOKBACK-1, len(returns)-HORIZON, HORIZON)]


def segment(samples, start, stop):
    # Purge labels crossing the right boundary; features use only history.
    return [r for r in samples if start <= r["t"] < stop and r["end"] < stop]


def evaluate(samples, threshold, cost_bps=10):
    if not samples or not math.isfinite(cost_bps) or cost_bps < 0:
        raise ValueError("nonempty samples and nonnegative finite cost required")
    if not math.isfinite(threshold) or threshold < 0:
        raise ValueError("threshold must be finite and nonnegative")
    pnl = []
    traded = 0
    for r in samples:
        # Synthetic periods have no opening gap: next open equals t close.
        # Decision after t close, enter next open, then hold two returns.
        position = int(r["feature"] > threshold)
        # Two-way proportional cost; no leverage/financing, no overlapping trades.
        pnl.append(position * r["label"] - 2 * cost_bps / 10000 * position)
        traded += position
    return dict(mean=statistics.mean(pnl), traded=traded, returns=pnl)


def choose(samples):
    # Deterministic tie break: first registered candidate. No test input.
    return max(CANDIDATES, key=lambda p: evaluate(samples, p)["mean"])


def max_drawdown(returns):
    equity = peak = 1.0
    worst = 0.0
    for r in returns:
        if not math.isfinite(r) or r <= -1:
            raise ValueError("invalid path return")
        equity *= 1+r
        peak = max(peak, equity)
        worst = max(worst, 1-equity/peak)
    return worst


def block_bootstrap(returns, block=3, repetitions=1000, seed=SEED):
    if not returns or not isinstance(block, int) or not 1 <= block <= len(returns):
        raise ValueError("invalid block")
    if not isinstance(repetitions, int) or repetitions < 20:
        raise ValueError("need at least 20 repetitions")
    rng = random.Random(seed)
    means, drawdowns = [], []
    for _ in range(repetitions):
        resample = []
        while len(resample) < len(returns):
            start = rng.randrange(len(returns)-block+1)
            resample.extend(returns[start:start+block])
        resample = resample[:len(returns)]
        means.append(statistics.mean(resample))
        drawdowns.append(max_drawdown(resample))
    means.sort()
    drawdowns.sort()
    # Empirical nearest-index percentile, not an analytic confidence guarantee.
    quantile = lambda xs, p: xs[round((len(xs)-1)*p)]
    return dict(mean_interval_95=[quantile(means, .025), quantile(means, .975)],
                drawdown_p50=quantile(drawdowns, .5),
                drawdown_p95=quantile(drawdowns, .95), block=block,
                repetitions=repetitions, seed=seed)


def run(returns):
    samples = rows(returns)
    folds = []
    for boundary in (100, 140, 180):
        training = segment(samples, 0, boundary)
        test = segment(samples, boundary, boundary+40)
        threshold = choose(training)
        folds.append(dict(train_end=boundary, test_end=boundary+40,
                          threshold=threshold, test_mean=evaluate(test, threshold)["mean"]))
    # Development ends at 220. Choose on 180..220 validation; never inspect final test.
    validation = segment(samples, 180, 220)
    selected = choose(validation)
    final = segment(samples, 220, len(returns))
    base = evaluate(final, selected)
    stressed = evaluate(final, selected, cost_bps=30)
    return dict(seed=SEED, input_sha256=hashlib.sha256(json.dumps(returns).encode()).hexdigest(),
                input_rows=len(returns), horizon=HORIZON, attempts=list(CANDIDATES),
                folds=folds, final_start=220, selected=selected, observations=len(final),
                traded=base["traded"], test_net_mean=base["mean"],
                cash_mean=0.0, always_long_round_trip_block_mean=statistics.mean(r["label"] for r in final)-0.002,
                high_cost_mean=stressed["mean"], max_drawdown=max_drawdown(base["returns"]),
                uncertainty=block_bootstrap(base["returns"]))


def self_test():
    original = data()
    samples = rows(original)
    for boundary in (100, 140, 180, 220):
        training = segment(samples, 0, boundary)
        assert all(r["end"] < boundary for r in training)
        assert any(r["t"] < boundary <= r["end"] for r in samples)
    changed = original[:220] + [0.1] * (len(original)-220)
    assert segment(rows(changed), 0, 220) == segment(samples, 0, 220)
    assert run(changed)["selected"] == run(original)["selected"]
    assert all(a["end"] <= b["t"] for a, b in zip(samples, samples[1:]))
    result = run(original)
    assert result == run(original)
    assert result["high_cost_mean"] <= result["test_net_mean"]
    assert abs(max_drawdown([.1, -.2, .1]) - .2) < 1e-12
    for invalid in ([], [0.0]*299, [math.nan]*300, [-1.0]*300):
        try:
            rows(invalid)
        except ValueError:
            pass
        else:
            raise AssertionError("invalid data accepted")
    for block in (0, 1000):
        try:
            block_bootstrap([.01, -.01], block=block)
        except ValueError:
            pass
        else:
            raise AssertionError("invalid block accepted")


if __name__ == "__main__":
    self_test()
    print(json.dumps(run(data()), ensure_ascii=False, indent=2))

On this page