QuantZ
Checking account

Overfitting auditor

A backtest is not evidence until you say how many you ran

Paste an equity curve and the number of variants you tried, and the auditor says what the result is worth given that search. Forty-five no-skill variants over 5 years hand the luckiest a Sharpe of about 1.0, so a Sharpe that size found that way carries no information.

This reads your equity curve, not your code. It cannot see look-ahead, survivorship bias in your universe, fills you would never have got, or costs you left out — those live in the backtest, not in its output. It also has to take your trial count on trust. What it can do is tell you how much the number you already have is worth given the search you say produced it.

Paste a column straight out of a spreadsheet. Several columns — one per configuration you compared — unlock the split test. Nothing here is uploaded; the arithmetic runs in this tab.

The parser follows TradingView's documented List of Trades format and picks the return column by its name. If your export is refused, send us the first ten rows, header included — the format is matched from the documentation rather than from a sample file.

Every parameter set, indicator swap and threshold you looked at before settling here — including the ones you discarded. There is no default, because a number chosen after the result is known is the problem this measures.

Sign in and the auditor reads N from your research ledger directly; until then the count is the one you type, and it is labelled that way.

Observation frequency

How the auditor works

What noise alone will hand you

The first column is how many configurations were tried. The second is the Sharpe ratio the luckiest of them is expected to show over 5 years when none of them has any edge. The third is how much history you would need before a search that size could be told apart from skill at a Sharpe of 1.0.

Variants triedBest Sharpe from luck aloneHistory needed
10.000.0 yr
50.531.4 yr
200.853.6 yr
451.005.0 yr
1001.136.4 yr
5001.379.3 yr
2,0001.5411.9 yr

Computed at render by the same functions the auditor above runs on, over 5 years of daily data. The independence assumption behind them is conservative: real searches are correlated, which makes the true no-skill maximum lower than these figures.

Where the numbers come from

The Deflated Sharpe Ratio and the no-skill maximum are Bailey and López de Prado (2014); the minimum backtest length is Bailey, Borwein, López de Prado and Zhu (2014); the track-record length is Bailey and López de Prado (2012); the split test is the Combinatorially Symmetric Cross-Validation of Bailey, Borwein, López de Prado and Zhu (2015); the serial-correlation correction is Lo (2002), and the reason it matters for smoothed returns is Getmansky, Lo and Makarov (2004).

None of this is ours. What is ours is that it runs on your numbers, for free. Our own Momentum Score is published as inputs, method and cadence, and is not tested or offered as a forecast, see how the score is computed.

What the auditor reads

This audit reads an equity curve, which is an output. Most of the ways a backtest misleads are properties of the code that produced it: look-ahead, where a signal quietly used tomorrow's close; survivorship bias, where the universe excludes the companies that went bankrupt; fills at prices nobody could have got; costs left out. Those live in the backtest, not in its output, so a clean result on this page is a statement about the curve alone.

The trial count is on trust unless your ledger measured it. The whole calculation turns on how many variants you actually tried, which is why the result names the source of N every time — self-reported by the author, or measured by the research ledger that recorded the search.