Skip to content
Strata
Methodology

We don’t try to prove your strategy works. We try to kill it.

Most backtests are optimism with a chart. A single historical run tells you what happened once, on one ordering of trades, on data the strategy was built to fit. Strata treats every strategy as guilty until proven robust — three independent attempts to break it, each scored, none optional.

01

The null hypothesis: beat random, or it isn't edge

Any market with a drift will make a lot of random strategies look smart. So the first question isn't "did it make money?" — it's "did it make more money than strategies with no idea behind them at all?"

For every scored run, the engine composes a pool of random strategies — same instrument, same data, same period, same trade frequency, no idea behind them — and banks their P&L. Your strategy gets a percentile rank against that distribution. A strategy at the 95th percentile cleared almost the entire random pile; a strategy at the 60th is statistically hard to tell apart from luck.

This is the cheapest, most brutal filter we have, and it's worth 10% of the Strata Score. Plenty of nice-looking equity curves don't survive it.

02

Monte Carlo reshuffle: remove the luck of the draw

A backtest hands you one ordering of trades — the historical one. But max drawdown is mostly a property of trade order, and order is partly luck. The same trades in a crueler sequence can blow a drawdown limit the original backtest sailed past.

So the engine resamples your strategy's trade P&Ls with replacement across 200 paths and rebuilds the equity curve for each one. Out of that distribution we report the share of paths that stayed profitable, the drawdown envelope, and the 95th-percentile max drawdown — the crueler sequences, not just the lucky one that actually occurred.

The 95th-percentile drawdown is the number that matters if you trade a drawdown-limited account: it tells you what the same trades can do in a worse order. This component carries 25% of the Strata Score.

03

Out-of-sample: the only result that counts is on data it never saw

A strategy tuned on six years of data will look brilliant on those six years. That's not insight, that's memorization. The honest test is performance on windows the strategy had no part in choosing.

Strata holds out the most recent 30% of the history — the strategy is scored in-sample on the first 70%, then judged on the untouched tail. Both halves get the full metric suite, and the gate demands real trade counts in each: a strategy that only worked in the past, or barely trades in the holdout, doesn't pass.

Out-of-sample stability is worth 20% of the Strata Score, and it's where most curve-fit strategies go to die.

04

What we deliberately don’t model

A methodology page that only lists strengths is marketing. These are the limits:

  • Order-book depth, queue position, and partial fills. No retail backtesting tool has honest access to that data, so we don't simulate a fiction of it. Fills are next-bar-open with commission and slippage costed on every trade — deliberately conservative.
  • Sub-bar microstructure. The engine builds its bars from CME 1-minute source data (via Databento), and the smallest tradable timeframe is 5 minutes. If your edge lives inside the bar, Strata is the wrong tool and we'd rather say so.
  • The future. Every test on this page is a statement about how a strategy behaved under historical stress, reshuffled and held out in every way we know how. None of it is a forecast. Backtest results do not guarantee future returns.

All of the above rolls up into one number. How the Strata Score is built →

Futures and derivatives carry substantial risk. Backtest results do not guarantee future returns. Strata is software — not investment advice.