Why Picking Your Best Backtest Result Is Already a Form of Overfitting
The act of sorting your training results and keeping the top performers is itself a source of bias, separate from and in addition to overfitting the rules.
You didn’t overfit by writing bad rules. You overfit the moment you sorted the results table and kept the top thirty.
That’s the part of the process that never gets treated as a modeling decision, because it doesn’t look like one. It looks like housekeeping. sorted_train = sorted(res_train.items(), key=lambda x: x[1]['pnl'], reverse=True) reads like a formatting step, something you’d do to make a printout readable. It isn’t. It’s a filter, and filters that select on the exact metric you’re trying to validate always introduce bias, even when every underlying pattern in the list is completely legitimate.
The moment the bias enters
Here’s the mechanism. Imagine, generously, that a chunk of those three hundred patterns represent genuine, if modest, edges, somewhere in the 53% to 58% true win rate range. On any given training window, each of those patterns will land somewhere around its true rate, but not exactly on it, because twenty or forty trades is a small sample and variance pushes the observed number around in both directions.
Sorting by PnL and keeping the top performers doesn’t just surface the patterns with the best true edge. It surfaces the patterns that got the most favorable draw of variance on this specific dataset, and a pattern with a mediocre true edge and lucky variance will consistently outrank a pattern with a better true edge and unlucky variance. That’s not a bug in the ranking. That’s what ranking by a noisy metric always does. It’s the same effect that makes rookie-season standouts across every sport regress toward average the next year, not because they got worse, but because the metric that got them noticed already had luck baked into it.
Why test performance should be expected to fall, even for real patterns
This is the part that trips people up emotionally more than statistically. A pattern shows a 65% win rate in training and a 54% win rate on the out-of-sample test set, and the instinct is to read that as failure, as the pattern breaking the moment it hit new data. Sometimes that’s exactly what happened. But often it isn’t. It’s regression to the mean doing exactly what it’s supposed to do: the training number was inflated by the same selection process that put the pattern in the top 30 to begin with, and the test number, closer to the pattern’s actual long-run rate, was never going to match it.
This is one of the reasons a validated win rate landing in the 52% to 62% range shouldn’t read as disappointing. That range is closer to what real, durable edges in liquid markets like XAUUSD actually look like once the training-set inflation gets stripped away by a genuinely unseen window. A test result that looks unglamorous next to its training counterpart is frequently the more honest number in the room, not a strategy in decline.
| Stage | What the number reflects |
|---|---|
| Training win rate | True edge + selection bias + this-window luck |
| Test win rate | Closer to true edge, minus whatever luck didn’t repeat |
| Live win rate | True edge, plus real slippage and cost drag the backtest underweighted |
One split isn’t enough insurance when you searched this hard
The train/test split is still the right foundation, and still the only step in this pipeline that checks a pattern against data it never influenced. But it was built to validate a small number of deliberate choices, not to absorb a selection process that already ran the training metric through a ranking filter before the test set ever got touched. Every pattern in that top 30 got there partly because of favorable training-window luck. The test set corrects for some of that. It wasn’t designed to correct for all of it, especially not from a single pass.
The trap sitting right next to this one is subtler and easy to walk into with good intentions. If a pattern disappoints on the test set, the natural next move is to go back, adjust the combination, rerun, and check again. Do that enough times and the test set stops being a holdout and quietly becomes a second training set, one you’re now optimizing against just as surely as the first, just with an extra step in between. At that point you need a third, genuinely untouched window before you can trust anything, and most pipelines, this one included, don’t have one built in by default.
What happens after you’ve done it right
Say a pattern survives all of that cleanly. It goes into a JSON config with its start_hour, end_hour, lot_size, sl, and tp, and it starts trading live. The next test isn’t statistical anymore. It’s psychological.
A validated system is going to lose money on a meaningful fraction of its trades, by design, even when it’s working exactly as intended. The temptation after a losing stretch is to intervene, tighten the filter, nudge the risk, second-guess the config, using the same instinct that built the original top-30 list in the first place. That instinct is precisely what needs to sit still here. A system validated at a 55% win rate over hundreds of test trades is expected to have losing weeks. That’s not new information. That’s the variance you already signed up for when you validated it.
The harder judgment call is telling that expected variance apart from an actual regime shift, the kind where the pattern’s real edge has genuinely decayed because the market structure it was built on has moved on. Both look identical in the moment, a losing streak is a losing streak either way. The only real way to tell them apart is a review cadence set in advance, against the original validated baseline, rather than a reaction triggered by however the last ten trades happened to go. One of those is a process. The other is just re-running the same selection bias that got you here, live, with real money, one trade at a time.