Trading Strategies
How to Backtest a Strategy Without Fooling Yourself
A backtest measures your assumptions first and the strategy second. Six of those assumptions decide whether the curve on your screen has anything to do with an account.
Every backtest returns a number, and the number is a measurement of your assumptions. The strategy only enters into it if those assumptions happen to be right. Six failure modes account for most of the distance between a curve on a screen and a statement at the end of the year, and each one begins as a decision somebody made without noticing they were making a decision at all.
The universe has already been filtered
Pull today’s index constituents, run ten years of history against them, and every company that went bankrupt, got acquired or fell out of the index is absent from the file. All of them. The remaining names are there because they lasted, which is both the definition of the bias and the reason it is invisible while you work.
Arithmetic on an invented universe shows the size of it. Two hundred stocks. Over the test window 180 of them return 9 percent on average and 20 are wiped out completely. The honest average across all two hundred is (180 x 9 + 20 x -100) / 200 = (1,620 - 2,000) / 200 = -1.9%. Test on the survivors alone and you measure +9%. Nothing about the strategy changed between those two figures. The file changed.
The damage lands hardest on rules that hunt for weakness, since a reversion screen buys whatever has fallen the furthest, and the names that fell furthest and then kept falling to zero are precisely the ones a survivor universe has deleted, so the simulation is never offered the trades that would have destroyed it. The honest version of that method, filters included, is in mean reversion trading.
Two things fix it. Point in time constituent lists, so the universe on any test date is the universe as it stood on that date. And delisted securities carried through to their final prices, including the ones that went to nothing. Most free data has neither. If you cannot obtain it, write that limitation next to the result and keep it there, because a caveat you remember in March is a caveat you have forgotten by June.
Look ahead bias is usually one line
The rule is simple to state. A simulated decision may use only what existed at the moment of that decision, which sounds like a low bar until you notice how many data sources arrive carrying today’s version of yesterday’s facts, with no memory at all of what yesterday’s version said. Violating it takes almost no effort.
The standard offender fills at the closing price of the bar that produced the signal. You computed the signal from that close, so the fill requires knowing the close before it printed, and the correction is to fill at the next open and measure what the change costs you, because that number was being hidden from you by the old assumption.
Fundamental data carries the same trap in a quieter form. A quarter ending in March gets filed in April or May, so a screen stamped with the quarter end has you trading on numbers nobody had yet. Restated figures are worse still, since the database holds the corrected version and the market at the time held the wrong one. Index membership has an effective date and an announcement date, and using the effective date backwards puts you in the stock before anyone knew it was joining.
Overfitting has an arithmetic
Count the search. Six lookback windows, five entry thresholds, four stop multiples and three liquidity filters give 6 x 5 x 4 x 3 = 360 variants, and reporting the best of them is what most people mean when they say they backtested something.
Now suppose every one of those 360 variants is worthless. Results still scatter, so roughly 360 x 0.01 = 3.6 of them will clear a bar that only one in a hundred should clear. You will find those three or four. Finding them is exactly what the search was built to do, and the equity curve attached to the winner will look convincing, because that is the property you selected on.
Every parameter you add multiplies the search space and spends sample at the same time. Four parameters against ninety trades is 90 / 4 = 22.5 trades for each degree of freedom, which is not a measurement. It is a drawing.
There is a cheap diagnostic. Look at the neighbours of your best setting. A lookback of 20 that works while 18 and 22 both fail is noise you have carved a shape out of. A real effect shows up as a plateau, with performance degrading smoothly as you move away from the centre, and the sensible choice is the middle of the plateau even when a spike somewhere else tested better.
A strategy can die on three basis points
Take an average trade that gains 0.14 percent gross, in a stock around $100 quoted six cents wide. Crossing half that spread costs $0.03 / $100 = 0.030% on each side, so 0.060% for the round trip. Commission at half a cent a share adds $0.005 / $100 = 0.005% a side, or 0.010%. Correcting the fill from the signal close to the next open took another 0.040% off in the measurement above.
| Line | Per round trip | Across 180 round trips |
|---|---|---|
| Gross average trade | 0.140% | 25.2% |
| Spread on a six cent quote | minus 0.060% | minus 10.8% |
| Commission at half a cent a share | minus 0.010% | minus 1.8% |
| Filling at the next open | minus 0.040% | minus 7.2% |
| Net | 0.030% | 5.4% |
A 25 percent gross year is a 5.4 percent net year. Now run the identical rule on smaller companies, where the quote is twelve cents wide, and the extra 0.060% a round trip turns the net trade into 0.030% - 0.060% = -0.030%, or 180 x -0.0003 = -5.4% across the year. Three basis points a side did all of that. Why the spread exists and who collects it is set out in market makers and liquidity.
Four more costs belong in the model. Slippage on stop orders, which fill as market orders into whatever the book offers. Market impact, once your size is a real fraction of the volume. Borrow fees and availability on the short side. Financing on anything carried with margin. None of the five appear anywhere in a price series, which is why a simulation built from prices alone will always report a better strategy than the one you can run.
Walk forward, and who has seen what
Let time run in one direction. Twelve years of data, fitted on the first three, then traded unchanged through year four, after which both windows roll forward by a year and the whole thing repeats until the data runs out. What you end up with is years four through twelve, nine stretches of performance where the parameters were chosen without any knowledge of the period they were applied to.
The accounting rule underneath it is stricter than most people apply. Data you have looked at is in sample from the moment you look. Check the out of sample result, dislike it, change a threshold and run again, and the out of sample window has become a slow and very expensive training set. Do that four times and you have overfitted the holdout while believing you validated on it.
So keep a final block sealed. The last two years, or a different market entirely, opened once when you believe the work is finished. If it disagrees with everything before it, the finding is that you have no strategy, and that finding cost you two years of data and saved you an account. Momentum rules are the usual place this goes wrong, because they look superb across a single long trend and fall apart at the turn, which momentum trading describes in detail.
How many trades a result needs
The standard error of an average shrinks with the square root of the sample, so standard error = standard deviation of trade returns / square root of number of trades. Take the 0.14 percent average trade above and assume the individual trades scatter with a standard deviation of 1.9 percent, a placeholder you replace with whatever figure your own record produces.
| Trades | Standard error | Average divided by standard error |
|---|---|---|
| 60 | 0.245% | 0.57 |
| 150 | 0.155% | 0.90 |
| 400 | 0.095% | 1.47 |
| 1,500 | 0.049% | 2.85 |
Sixty trades tells you nothing. At 400 the average is still well inside the noise, and only somewhere past a thousand does the edge start to separate from zero, which is an uncomfortable requirement for a rule that fires twice a month. Quadrupling the sample only halves the error bar. That is the arithmetic and no amount of clever reporting improves it.
Independence matters as much as the count. Forty positions opened on the same morning, in the same sector, on the same signal, are one observation wearing forty hats, and a backtest that reports 1,200 trades taken in 40 clusters has closer to 40 observations in it. Count the clusters. The same correlation problem sizes your live positions too, which is the portfolio heat discussion in risk management for traders.
Regimes are the other axis. A sample drawn entirely from one long rising market has tested a single environment over and over, whatever the trade count says, so look for a falling market, a flat one and at least one volatility shock inside the window, and if the history you hold contains none of them then the honest finding is that the strategy is untested in the conditions it has never met, which is a very different statement from saying it failed and a much more useful one to carry into a live account. Return per unit of risk across those stretches, which the Sharpe ratio calculator computes, tells you more about survivability than total return does.
Paper trading is the last check
A simulation cannot see your broker. Paper trading can, and it catches a specific list of failures: limit orders that never fill at the price the file assumed, shares you cannot borrow at any price, a data vendor whose bars disagree with the tape, and the discovery that placing thirty orders before the open takes longer than the open gives you.
It also flatters you. Fills are generous, your size never moves the book, and no part of you is frightened, which removes the single largest source of live underperformance. Treat the paper result as a ceiling. Then run small real size for a stretch before you run full size, log every deviation from the rules as it happens, and compare the live average trade with the backtested one in blocks of fifty.
One account of exactly this going wrong, with the costs added at the end and the edge disappearing, is in the backtest that died when I added commissions.
Frequently asked questions
What is survivorship bias in backtesting?
Testing a rule against a universe built from the companies that still exist today, which quietly removes every firm that went bankrupt, was acquired or dropped out of the index. The survivors are in the file because they survived. Strategies that buy weakness suffer worst, because the names they would have bought and lost on are the exact ones missing.
What is look ahead bias?
Using information in a simulated decision that was not available at the moment of that decision. The common version fills the order at the closing price of the bar that generated the signal, which requires knowing the close before the close. Fundamental data stamped with the quarter it describes, when the filing landed weeks later, does the same thing more quietly.
How many trades does a backtest need?
Enough that the average trade is large compared with the standard error around it, and the standard error falls only with the square root of the trade count. Quadrupling the sample halves the error bar. A few hundred trades is usually the bare minimum for a small per trade edge, and the trades have to be reasonably independent to count at all.
What is walk forward testing?
Fitting parameters on one block of history, trading them unchanged on the block that follows, then rolling both windows forward and repeating. The result is the stitched together record of the untouched blocks. It is the closest a simulation gets to the order in which you would have actually made decisions.
Does paper trading prove a strategy works?
No, and it is still worth doing. Paper trading catches the operational failures a backtest cannot see, such as orders that never fill, shares you cannot borrow and data that disagrees with the tape. It flatters fills and removes the emotional load of real money, so treat the result as an upper bound on live performance.