Leaderboard

Ranked by what the evidence is worth

Sorting strategies by their headline number ranks whoever searched hardest, because the best of many no-skill attempts looks impressive by construction. This board orders on the Deflated Sharpe Ratio instead, the trial count is printed beside every row, and an entry that will not state one is listed but never ordered.

The board

There are no entries yet. Nothing below this line is a strategy anyone submitted, and no placeholder rows have been invented to fill the space — on a page about whether numbers can be believed, that would be a strange place to start.

An entry needs three things: a return series, the number of configurations it was selected from, and — to be ranked in the upper band — a prediction sealed in the pre-registration vault before the run that tests it. The first two are what make a result comparable; the third is what makes it a claim rather than an observation.

The ordering rule, in full

Entries are ordered by their Deflated Sharpe Ratio — the probability the result reflects skill rather than the luckiest of however many configurations were tried. Entries carrying a prediction sealed before the run are ranked above those that do not, and among them the ordering multiplies that probability by how close the prediction came. Ties go to the entry that searched less. Anything that cannot say how many configurations it was selected from is not ranked at all.

Searching forty-five variants of a strategy over five years of daily data, with no skill whatsoever, produces a best-of Sharpe ratio of about 1.0. A leaderboard that sorts on the headline number without knowing how many variants produced it is therefore close to a ranking of who searched hardest. That is why the count is required here and printed beside every row, and why an entry without one is listed but never ordered.

What the rule does that a normal ranking does not

The four rows below are invented for this illustration — they are not entries and nobody submitted them. They exist to show the ordering doing something a return-sorted board cannot: the entry that delivered less ranks first, because it said in advance what it would do and was close.

#Illustrative entrySharpe · DSR · NPredicted → deliveredScore
1Predicted 8%, delivered 7.6%· sealed prediction0.88 · DSR 0.336 · N 128% → 7.6%0.320
2Predicted 40%, delivered 12%· sealed prediction0.89 · DSR 0.342 · N 1240% → 12.0%0.201
3Same result, found on the 900th attempt· no prediction0.89 · DSR 0.024 · N 9000.024

Not ranked — Will not say how many variants were tried

This entry does not say how many configurations it was selected from, so its result cannot be deflated and cannot be compared with anything. It is not ranked last — it is not ranked, because a rank would assert a comparison nobody measured.

Computed at render by rankBoard() — the same function the real board calls — over synthetic return series generated from a fixed seed. A hand-typed illustration of a ranking rule goes stale the first time the rule changes, and then the page is teaching something false.

What this board still cannot check

It takes the trial count on trust unless it came from an account's own trials ledger, and every row says which of the two it is. It reads a return series, so it inherits every limit the auditor has: look-ahead, survivorship and unrealistic fills all live in the code that produced the curve, not in the curve. And a sealed prediction proves the ORDER of a person's own records, not a notarised time.

What it does do is remove the one thing every other ranking of this kind rewards by accident: searching until something looks good, and then reporting only the something.