Why your backtest looks better than your live trading results
Overfitting, look-ahead bias, survivorship bias and missing costs explain why most backtests disappoint in live trading. Here's how to catch each one.
Most people who build a trading strategy hit the same moment. The backtest shows a smooth equity curve, small drawdowns and a Sharpe ratio worth showing someone. Then the strategy trades real money and the results look nothing like it.
That gap is rarely bad luck. Usually the backtest was measuring something other than what you thought. These are the most common causes, and what you can do about each.
Overfitting: the strategy learned the past, not the market#
Every parameter you tune, whether it's a moving-average length, a stop-loss distance or an entry threshold, gives the strategy another way to fit the historical data. Try enough combinations and one of them will look good by chance alone.
The usual safeguard is to hold back some data for a final test. It protects you less than you'd hope. In The Probability of Backtest Overfitting, Bailey, Borwein, López de Prado and Zhu argue that standard hold-out validation tends to be unreliable for investment backtests. Part of the problem is human. After each round of changes it's tempting to check the hold-out set again, and once you've looked at it a few times it isn't really unseen. The authors propose a way to estimate the probability that a chosen strategy is overfit, based on how its ranking holds up across many different splits of the same data.
You can use the lesson without implementing their method. Keep a log of every variant you test, not only the winner, because ten variants and two hundred variants are very different situations even when the best result looks the same. Prefer fewer parameters: a strategy with three settings that works across a range of values is more believable than one with ten settings that only works at one exact combination. And check the neighbourhood. If nudging a parameter slightly wrecks performance, you've probably found noise.
Multiple testing: the bar for "significant" is higher than you think#
Professional researchers fall into the same trap. Harvey, Liu and Zhu looked at the hundreds of return "factors" proposed in published finance research. They argued that the usual cut-off for significance, a t-statistic above 2.0, is too lenient once you account for how much searching has gone on, and estimated that a newly proposed factor should clear a t-statistic above 3.0 instead. They also concluded that many claimed findings in financial economics are likely false.
For an individual trader the takeaway is simple. If you tested a lot of ideas to find one that works, ask for much stronger evidence from it than you would from an idea you only tested once.
Look-ahead bias: decisions made with information you didn't have yet#
Look-ahead bias creeps in when the backtest makes a decision using data that wouldn't have existed at that moment. Some common ways it happens:
- using a day's closing price to decide a trade that's supposed to execute at that day's open;
- using a company's quarterly figures as of the quarter's end, when they were actually published weeks later and sometimes revised after that;
- calculating an indicator over the whole dataset, such as normalising by the full-period average, and then trading on it day by day.
The same leak shows up when machine-learning models are tested with ordinary shuffled cross-validation. The scikit-learn documentation for TimeSeriesSplit explains that standard cross-validation is inappropriate for time-ordered data because it "would lead to training on future data and evaluating on past data". TimeSeriesSplit only trains on earlier observations and tests on later ones. Its gap parameter leaves a buffer between the two, which helps when your features use rolling windows that would otherwise overlap the test period.
One habit prevents most of this. Timestamp every input with the moment it became known, not the moment it describes, and only let the backtest read data that was known before each decision.
Survivorship bias: the failures vanished from your data#
If your historical universe only contains companies, funds or coins that still exist today, you've quietly removed the ones that failed. What's left looks healthier than the real market did at the time.
Fund researchers have studied this for decades. The paper Survivor Bias and Mutual Fund Performance, published in The Review of Financial Studies in 1996, starts from exactly this problem: funds that disappear tend to do so because they performed poorly, so studying only the survivors makes the group look better than it was.
For stock strategies, the fix is a point-in-time universe, meaning the list of instruments you could actually have traded on each date, including ones that were later delisted. Many cheap or free datasets leave delisted instruments out, so check before you trust a long backtest.
Costs and slippage#
A backtest that buys and sells at the exact price on the chart with no fees describes a market nobody trades in. In real trading you pay commission, you cross the bid–ask spread, and large orders move the price against you.
These costs hurt most for strategies that trade often or in thinly traded instruments. If the average profit per trade is smaller than a realistic round-trip cost, there's no edge left. Three checks are worth doing every time. Model fees and spreads explicitly, using your broker's actual fee schedule. Re-run the backtest with costs doubled, because an edge that disappears was thin to begin with. And compare your position sizes with typical trading volume to see whether your own orders would have moved the price.
A validation routine#
None of this needs special tools. A routine like this catches most problems:
- Write down the hypothesis before testing it: what inefficiency you're exploiting, and why it should persist.
- Use point-in-time data that includes delisted instruments.
- Split by time, not at random. Develop on an early period and evaluate on later data you haven't touched, walking forward through several windows instead of relying on one split.
- Log every variant you test, and raise your bar for evidence as the count grows.
- Include realistic costs, then stress-test them.
- Paper trade before going live, and compare the live fills with what the backtest assumed.
If you're automating a strategy#
Automation helps you run a strategy consistently, but it also makes it easy to deploy an overfit one quickly. When clients ask us at TheAILAB to automate their strategies, most of the engineering goes into these checks (clean point-in-time data, walk-forward testing and hard risk limits) rather than the entry signal, which is theirs. That's where most of the risk sits.
Treat a great-looking backtest as a hypothesis about the past, not a forecast, and live trading is much less likely to surprise you.
This article is for education only and isn't financial advice. Trading involves risk of loss.
Sources
- The Probability of Backtest Overfitting — Bailey, Borwein, López de Prado & Zhu — Journal of Computational Finance (via SSRN) (checked 2026-10-10)
- ... and the Cross-Section of Expected Returns (NBER Working Paper 20592) — Harvey, Liu & Zhu — National Bureau of Economic Research (checked 2026-10-10)
- TimeSeriesSplit — scikit-learn documentation — scikit-learn (checked 2026-10-10)
- Survivor Bias and Mutual Fund Performance — The Review of Financial Studies, Vol. 9, No. 4 (1996) (checked 2026-10-10)
Trading a strategy by hand?
If you have rules you already follow and want them automated, tell us about them. You define the strategy; we help automate it.