Backtesting & results
The Edge Report
The one card that asks whether your backtest describes an edge or just describes the past — eight checks, what each one tries to break, and how to read a card that has no verdicts yet.
Updated 2026-09-07·Rev. 2026.09
A backtest tells you what already happened. The Edge Report asks whether it happens again.
It sits under your metrics as one card, and it holds eight checks. Each one attacks the result from a different angle and reports whether the result broke. They are not eight opinions — they are eight attempts to falsify the same claim, and the card is the argument they add up to.
Why the card exists
The seductive failure in strategy building is not a losing backtest. It's a winning one that only wins on the exact window, the exact settings, and the exact trade order it was measured in. Every metric on the results panel — Sharpe, drawdown, win rate — is a description of one history. None of them can tell you whether that history was representative or flattering.
That is the entire job of this card: to separate "we checked and it held" from "we didn't check" from "we checked and it fell apart".
An absence is never a pass
If a check produced no report, the card says so. It never fills the gap with a reassuring green. A missing verdict is a missing verdict.
The eight checks
They are grouped by what they're for, and each group heading is the question its checks answer.
Select any row below to see the exact thresholds behind its tick or cross, and one worked run that lands on them. Each row carries its own scale — the coloured strip is that check's pass, caveat and fail bands, with a mark where the worked run fell — so the eight are comparable at a glance even before you open one.
Grade gate — six of the eight can hold your letter down. Every strip below a row is that check's own scale, marked where the worked run landed.
Out-of-sample check
Is the edge real? · Deep run · grade gateThe window is cut in two by time. The strategy is scored again on the later slice alone, and the two risk-adjusted returns are divided: held-out Sharpe ÷ in-sample Sharpe.
- Holds up — Holds. — Retention ≥ 0.70×.
- Degraded — Caveat. — Retention between 0.30× and 0.70×.
- Fitted — Broke. — Retention below 0.30×, negatives included.
- Not enough evidence — No verdict. — Fewer than 20 closed trades in the held-out slice, a slice under 10% of the window, or an in-sample Sharpe at or below 0.1.
180-day run, 25% held out. In-sample Sharpe 1.80; the held-out 45 days score 1.53 over 34 trades.
1.53 ÷ 1.80 = 0.85×✓ Holds up
Above the 0.70× line, on a slice with enough trades to mean something. The strongest single statement the card can make — and still only about the past.
Walk-forward
Is the edge real? · Deep run · grade gateThe window is cut into 4, 6 or 12 consecutive periods, each scored on its own with the same graph and settings. A period counts only with at least 10 closed trades and 30 bars.
Clearing this line is necessary, not sufficient: a single period below -0.5 Sharpe reads Concentrated however many others were positive.
- Consistent — Holds. — At least 60% of scored periods profitable AND the worst period's Sharpe no lower than -0.5.
- Concentrated — Broke. — Under 60% profitable, or one period below -0.5 Sharpe. The whole-window result rests on a favourable stretch.
- Not enough windows — No verdict. — Fewer than 3 periods could be scored — usually 12 windows over a range too short to fill them.
- Account lost — Broke. — The account was wiped out inside the window. That is the outcome, not a missing reading.
6 periods of ~30 days. Five profitable; the losing one scores −0.2 Sharpe.
83% positive, worst period −0.2✓ Consistent
Clears both halves of the rule. Had that one period scored −0.8, the same 5-of-6 would read Concentrated: a strategy that disappears for a whole period at a time is not consistent, however good the majority looks.
vs Random entries
Is the edge real? · Edge Check run · grade gateThe strategy is re-run at least 400 times with the same trade count, the same long/short mix and the same holding periods — but random entry timing. Your return is ranked against that field.
- Beats random — Holds. — Your return sits at the 80th percentile or above.
- Same as random — Caveat. — Between the 20th and 80th. The entry rules are decoration — the return came from exposure, not from the signal.
- Worse than random — Broke. — At the 20th percentile or below.
- Not enough trades — No verdict. — Fewer than 30 closed trades, so a percentile would be ranking noise.
112 closed trades. 400 randomly-timed re-runs at the same exposure; 28 of them beat your return.
93rd percentile✓ Beats random
Being in the market at all was not the story. Note what this check holds fixed: same number of trades, same holding time — so it isolates timing and nothing else.
vs Shuffled market
Is the edge real? · Edge Check run · grade gateThe same candles' moves are dealt in a random order, so trends and reversals are destroyed while volatility and costs survive. The p-value is the share of shuffled markets that did at least as well as the real one.
- Beats shuffled market — Holds. — p ≤ 0.05. The only route to the top grade.
- Barely beats shuffled — Caveat. — p between 0.05 and 0.2.
- Same as shuffled — Broke. — p above 0.2 — the strategy was being paid for the shape of its bets, not for reading the market.
- Not measured — No verdict. — Fewer than 20 closed trades, or under 100 bars to shuffle.
500 shuffled markets built from the same candles. 60 of them matched or beat the real run.
p 0.12! Barely beats shuffled
Not a failure — one run in eight of pure noise would have looked this good. It is the row to fix before claiming an edge, and the one blocking the top grade.
Parameter sensitivity
Is the edge real? · Edge Check run · grade gateThe two most influential numeric settings are swept to neighbouring values and the strategy re-run at each. Two figures come out: how much of the Sharpe the neighbours retain, and what share of them still made money.
Second axis, not drawn: the share of neighbours that stayed profitable. Under 50% of them is knife-edge on its own; a plateau needs 80%.
- Plateau — Holds. — Neighbours retain ≥ 0.70× of the Sharpe AND at least 80% of them made money. Both halves, because either alone can be satisfied by a ridge.
- Slope — Caveat. — Anything between the two verdicts below and the plateau.
- Knife-edge — Broke. — Retention under 0.25× OR under 50% of the neighbours profitable. Either one on its own says the result belongs to one exact value.
- Not measured — No verdict. — Fewer than 4 neighbouring settings could be scored — a cell needs 20 closed trades of its own.
EMA length 20 → Sharpe 2.4. Swept to 18 and 22: Sharpe 0.5 and 0.3; 3 of 8 neighbours profitable.
0.18× retained, 38% profitable✕ Knife-edge
The 2.4 exists at one value and nowhere near it. A higher return at one exact setting is a worse finding than a lower one across a range — you cannot trade a value you only found by looking.
Track record
Can you trust these numbers? · Quick run · grade gateA small edge takes more history to separate from luck than a large one. The minimum track record length for your Sharpe at 95% confidence is computed, and your run's observations are compared against it.
- Long enough — Holds. — Observations meet or exceed the requirement.
- Too short — Caveat. — Below the requirement. The strategy is not bad — the number is unproven.
- Unprovable — Broke. — The requirement exceeds 20 years. An edge this small cannot be demonstrated with any window you can run.
- Not measurable — No verdict. — No usable Sharpe to size a requirement from.
Sharpe 0.9 on 90 days of daily observations. That Sharpe needs about 163 days at 95% confidence.
90 ÷ 163 = 0.55×! Too short
The fix is more history, not a different strategy. Note the direction: a higher Sharpe needs LESS time — so this row gets easier exactly when the others get harder.
Trade-order test
Can you trust these numbers? · Quick run · describes onlyYour closed trades are resampled into 1,000+ different orders, and a random 20% of them dropped in a second variant. Your equity curve is one path out of that spread.
Second axis, not drawn: how often dropping a random 20% of the trades turns the run negative. At 15% or more the verdict is Fragile wherever the drawdown ranked.
- Order-robust — Holds. — Your drawdown is a typical one for these trades, and dropping a random 20% of them rarely turns the run negative.
- Lucky ordering — Caveat. — Your drawdown ranks in the mildest 10%. Plan for a deeper one — the same trades in a different order would have hurt more.
- Fragile — Broke. — 15% or more of the runs go negative when a random 20% of the trades is dropped. The result rests on a handful of trades.
- Not enough trades — No verdict. — Fewer than 30 closed trades to reorder.
84 closed trades, 1,000 reorderings. Your −14% drawdown ranks 34th out of 100; dropping a random fifth turns 4% of runs negative.
34th percentile, 4% negative✓ Order-robust
Both signals clear. This row never caps your grade — it changes the drawdown you budget for, which is a planning fact rather than a verdict on the edge.
Across market conditions
Where did the return come from? · Quick run · describes onlyTrades are bucketed by the market they were opened into — up-trend, range, down-trend, crossed with high and low volatility. A bucket is scored with at least 30 bars and 10 trades in it.
Crossing 80% only flags concentration when that condition also covers 25% or less of the bars. Consistent is decided elsewhere — by the share of scored conditions that made money.
- Consistent — Holds. — At least 75% of the scored conditions were profitable.
- Edge concentrated — Caveat. — One condition produced 80% or more of the net return while covering 25% or less of the bars.
- Mixed — No verdict. — Neither reading holds. Information, not a verdict.
- Not enough evidence — No verdict. — Fewer than 2 conditions had the bars and trades to be scored.
Up-trend / high-vol covers 21% of the bars and produces 86% of the net return.
86% of return from 21% of bars! Edge concentrated
Not a failure and never a grade gate — it is a description of what has to keep being true. This strategy is a bet on trending, volatile markets, and it should be deployed as one.
Is the edge real?
Five ways of asking whether the result survives outside the window, timing and settings it was measured in. These are the checks that can hold your grade down.
| Check | The question it answers |
|---|---|
| Out-of-sample check | Does the edge survive on data the strategy never saw? |
| Walk-forward | Was it there in every period, or only in one? |
| vs Random entries | Did your entry rules do anything, or was being in the market the whole story? |
| vs Shuffled market | Do your rules read this market's order, or would they profit from any noise? |
| Parameter sensitivity | Is this a region you can stand on, or one lucky setting? |
The order is deliberate. "Does it survive unseen data" comes first; "was it there in every period" is that same question asked repeatedly. The two null comparisons follow, because they ask whether there was ever anything there to survive — vs Random entries replaces your timing with random timing at the same trade count and holding period, and vs Shuffled market deals the same candles' moves in a random order so trends are destroyed while volatility and costs are not. Sensitivity closes the group, because it asks about the settings that produced all four answers above it.
A higher return at one exact setting is a worse finding, not a better one
Sensitivity re-runs the strategy at neighbouring values of its most influential settings. A result that survives its neighbours is a plateau you can stand on. One that vanishes a single step away was drawn, not discovered.
Can you trust these numbers?
| Check | The question it answers |
|---|---|
| Track record | Is this window even long enough to prove a result this size? |
| Trade-order test | Is the drawdown on screen the drawdown to plan for? |
A small edge needs more history to separate from luck than a large one does. Track record says how much calendar time a Sharpe this size actually requires, and whether your run has it. A window that's too short doesn't make the strategy bad — it makes the number unprovable.
Trade-order test resamples your trades into thousands of different orders. Your equity curve is one path out of many the same trades could have produced; this shows the whole spread and where your path sat inside it. A drawdown that ranks among the mildest of them means you were shown a flattering curve.
Where did the return come from?
| Check | The question it answers |
|---|---|
| Across market conditions | When did it work — everywhere, or in one kind of market? |
Trades are sorted by the market they were opened into: trending, ranging, calm, volatile. A return earned entirely in one condition stops the day that condition does. This one is a description, not a gate — it never changes your grade.
The headline
The card's collapsed header carries one phrase for the whole argument.
| Headline | What it means |
|---|---|
| Edge holds up | Every check that ran returned a clean result. The strongest statement this panel can make — still not a statement about the future. |
| Edge with caveats | Everything that could score did, and at least one returned a warning. Nothing here says the strategy is broken; it says what to look at first. |
| Edge questioned | At least one check found the result does not hold. Open the card — the failing row names what broke. |
| Checks pending | Checks are still waiting to be asked for, or the run is too thin for them to score. |
| Not measurable | No check could produce a verdict at all. That absence is not a finding about the strategy. |
Why a clean sweep can still say “Checks pending”
The headline refuses to claim an edge holds up while checks remain un-run. A card reading "Edge holds up" with three of its five hardest checks never asked for is exactly the over-claim this design avoids.
Reading the chips: "Not run" vs "No verdict"
Beside the headline sit counts. The two grey ones look alike and mean opposite things, so this is the distinction worth learning:
- Not run — the check hasn't been asked for yet. There's an action behind it. The three deferred checks (vs Random entries, vs Shuffled market, Parameter sensitivity) are released together by the Edge Check button inside the card. One run releases all three.
- No verdict — the check did run and honestly couldn't answer, or it needs a deep run and this was a quick one. Usually too few closed trades or too short a window.
The first has a button. The second needs a different run: a longer date range, a lower timeframe, or a strategy that trades more often. Hovering any chip on the collapsed card spells this out.
Grades and ceilings
Six of the eight checks are grade gates — they can hold your letter grade down no matter how good the raw metrics look. Those six are the five in Is the edge real? plus Track record. The trade-order test and the market-conditions breakdown describe the result; they never ceiling it.
When a gate is what's capping your grade, the card names it on the row responsible, so "why is my grade a B" has a visible home rather than living in a popover somewhere else.
Tip
Clearing vs Shuffled market is the only route to the top grade. A strategy that still makes money on shuffled data was being paid for the shape of its bets, not for reading the market.
Which checks run when
| Quick run | Deep run | Deep + Edge Check | |
|---|---|---|---|
| Track record, Trade-order test, Across market conditions | ✓ | ✓ | ✓ |
| Out-of-sample check, Walk-forward | — | ✓ | ✓ |
| vs Random entries, vs Shuffled market, Parameter sensitivity | — | — | ✓ |
A quick run covers only the candles currently on your chart — too short for a split, and re-running it hundreds of times would cost more than the run itself. On a quick run the panel names those checks in one locked strip with a Set up deep run link — it switches you to the Deep tab and opens the window controls.
Next: out-of-sample and walk-forward, the two settings that decide how the window gets cut.