Your Backtest Doesn't Know What Regime It's In, and That's the Real Problem
A backtest run over three years of data doesn't validate a strategy against three years of market conditions, it validates it against an average of conditions that never coexisted.
Run a backtest over three years of EUR/USD data and the report will hand you one number: total return, one win rate, one expectancy per trade. It reads like a single coherent verdict on the strategy. It isn’t. Those three years almost certainly contain a trending stretch, a ranging stretch, a low-volatility grind, and at least one violent macro-driven move that behaved nothing like the rest of the dataset. The backtest doesn’t know the difference. It blends all of it into one average, and the strategy you think you validated is really a strategy validated against a regime that never actually existed as a single market condition.
This is the part of backtesting almost nobody checks, because the report doesn’t surface it. You get one number, it looks stable, and stability reads as validation. But an average built across regimes that behave differently isn’t evidence of robustness. It’s often the opposite: a strategy that made money in a trend and lost money in a range can produce the exact same blended win rate as a strategy that performs consistently in both. The report can’t tell you which one you’re looking at.
What “average edge” actually hides
Take a pattern that wins 71% of the time during a clean directional trend and 41% of the time during a choppy range. If your three-year window happens to contain roughly equal amounts of both, the blended win rate lands somewhere in the mid-50s, comfortably inside the range most people are taught to trust as evidence of a real, non-overfit edge. The number looks healthy. It’s actually two completely different strategies wearing the same statistic, one of which is currently profitable and one of which is currently bleeding, and the average can’t tell you which regime the market is in right now.
This is why a strategy can pass every sanity check on paper and still start losing the moment it goes live, without anything about the strategy itself changing. The market didn’t break the pattern. The market simply moved into the half of the blend that was never actually working, and the backtest never separated the two halves in the first place.
Why the train/test split alone doesn’t catch this
The train/test split is still the right foundation, it remains the only validation that tells you whether a pattern generalizes at all rather than just fitting noise in the training window. But splitting chronologically, train on the first eighteen months and test on the last six, doesn’t fix the blending problem. It just moves it. If both halves still contain a mix of regimes, and they usually do, you’ve validated that the blended average holds up out of sample, not that the underlying edge is regime-stable. A strategy can pass a clean train/test split and still be entirely dependent on one regime showing up in both halves in similar proportion, purely by chance.
The fix isn’t to abandon the split, it’s to add a second axis to it. Train and test on data segmented by regime, not just by time, so you can see whether the edge holds inside a single regime and not just across an arbitrary calendar boundary that happens to contain a lucky mix.
Where the blending hides in plain sight
The place this shows up most often in a forex bot’s backtest is session structure. A pattern tested across the full 24-hour cycle is implicitly blending Asian session behavior (00:00–08:00 UTC), London session behavior (08:00–16:00 UTC), New York session behavior (13:00–21:00 UTC), and the London/New York overlap (13:00–16:00 UTC) into one number, even though these windows have different liquidity profiles, different volatility character, and frequently different dominant regimes entirely. A pattern that’s trend-following in nature might be genuinely edge-positive during the overlap, where directional volume concentrates, and edge-negative during the low-liquidity Asian session, where price often chops without conviction. Blend those two and you get a mediocre-looking average that hides a strategy which is actually strong, just not everywhere, all the time.
This is exactly the kind of thing a JSON config with start_hour and end_hour fields is built to handle, but only if those values were chosen after checking whether the edge survives being isolated by regime and session rather than set once across the whole day because the blended backtest looked acceptable.
Segmenting by regime instead of by calendar
Practically, this means running the backtest three times instead of once: once filtered for trending conditions, once filtered for ranging conditions, once filtered for high-volatility transition periods, however you choose to define those boundaries for your instrument. If expectancy holds up reasonably consistently across all three, you have something closer to a genuine structural edge. If the entire result is being carried by one regime while the other two are flat or negative, the blended number was always going to be misleading, and you now know exactly which conditions the strategy actually depends on to be profitable.
This also gives you something the blended backtest never could: an early warning system. If you know a pattern’s edge lives almost entirely in trending conditions, you have a concrete reason to reduce exposure or pause the bot the moment the market shows structural signs of transitioning into a range, rather than waiting for the blended equity curve to confirm it three weeks later after the losses have already accumulated.
What to do once you know which regime is carrying the strategy
The uncomfortable part of this exercise is that most patterns, once segmented honestly, turn out to be regime-dependent rather than regime-agnostic. That’s not a failure of the strategy, it’s just the truth the blended backtest was hiding. The mistake is reacting to that discovery by trying to patch the strategy so it performs everywhere, tightening stops here, adding a filter there, until it’s been reshaped entirely around the specific historical mix of regimes in your dataset. That’s how a legitimately regime-dependent edge turns into an overfit one.
The better response is to let the strategy be what it is. Keep the regime dependency, but make it explicit and monitored rather than hidden inside a blended average. A pattern that only works in trending conditions is still a useful, tradeable pattern, as long as you know that’s what it is before the market tells you the hard way.