Honest Strategy Lab

Study · October 2026

Why most backtests lie: overfitting, costs and luck

A backtest is a story the past tells about a rule. Three habits make that story far better than the truth. Here they are, with numbers from our own studies, and the questions that catch them.

1. Overfitting: the past, fitted

Try enough settings and one of them will look brilliant on the years you tried them on. That best result is partly skill and partly the noise of those years. The more variants you compare, the more of it is noise.

In our golden cross study the chosen pair earned +1.93 R per trade on the years it was chosen on. Measured year by year, always choosing on the three years before, it swung from −0.52 R in 2022 to +4.07 R in 2023. One good stretch of years can make any rule look like a discovery.

The fix: decide the rule and what counts as success before looking, then judge it on years and stocks that played no part in choosing it.

2. Costs: small per trade, huge per rule

We tested a popular intraday approach built on fair value gaps and liquidity on QQQ and SPY, minute by minute, from 2019 to 2026: about 3,000 trades over three separate tests. Before costs, its average trade was about zero (between +0.00 and +0.03 R). Costs were only 3.6 basis points a trade, but its stops were so close, about 0.2% of the price, that those costs came to about 0.21 R per trade. After costs it lost about 0.2 R per trade in every test.

The fix: always charge fees and slippage, and look at them in R. The tighter the stop, the more a basis point costs.

3. Luck: would random dates do as well?

A rule can make money simply because the market rose while it was in it. To separate the signal from the market, take the same trades, with the same lengths, on random dates. If the rule does no better than its random twin, the signal adds nothing.

The intraday approach above did exactly as well as its random twin: the difference was +0.02 R per trade, with an interval from −0.07 to +0.12 R. Its win rate with a 1 R target was 51%, and it still lost money after costs. A high win rate is not evidence.

Five questions before you trust a backtest

  1. Was the rule fixed before the results were seen, and how many variants were tried?
  2. Does it hold on years, and on markets, that played no part in building it?
  3. Are fees and slippage charged on every trade, and what are they worth in R?
  4. Does it beat simply holding the same thing, on risk-adjusted return?
  5. Does it beat the same trades taken on random dates?

The fair test asks all five of every rule you build. Most fail, including ours. Past prices; not advice.

Join the waiting list How the test works