How Many Random Patterns Look Profitable by Chance Alone

At the win rate thresholds most backtests use, a meaningful share of 'profitable' patterns are exactly what pure chance would produce anyway.


Flip a coin twenty-five times. Get fifteen heads. You now have a 60% win rate.

Nobody would call that a system. But run twenty-five trades through a fixed 10-point TP and 10-point SL, land fifteen winners, and the exact same arithmetic gets treated as evidence. The math doesn’t know the difference between a coin and a candle. Only the framing changes.

The coin flip nobody thinks they’re flipping

Look at the actual thresholds a pattern has to clear here to earn a place in the results. MIN_COMPLETED_TRADES is set to 20. MIN_WIN_RATE is set to 40%. Neither of those is a hard bar. Twenty trades is a small sample by any statistical standard, and a 40% win rate on a 1:1 risk-reward setup is only barely above breakeven once you factor in the trade cost. So the actual gate a pattern needs to pass is closer to “get lucky over about twenty to thirty attempts,” which is a much lower bar than it sounds like when it shows up as a clean percentage in a printed table.

Here’s what that looks like under a true null. Assume a pattern has no real edge at all, a genuine 50% win rate, and gets 25 completed trades in the test window. Using a normal approximation to the binomial distribution, the odds of that pattern showing a 60% win rate or higher, purely from variance, land somewhere around one in five. Not one in a thousand. Not a rare fluke. Roughly a fifth of truly worthless patterns will look like a solid 60% edge over that sample size, just from noise.

Now scale that up. Run three hundred semi-independent patterns through the same window, the way a combinatorial event miner does by default, and the expected number of pure-noise patterns clearing 60%+ isn’t one or two. It’s dozens. Fewer than the naive multiplication suggests, because many of those three hundred share overlapping conditions and aren’t truly independent draws, but still enough that a “top 30 by PnL” list is, at minimum, partly built out of exactly this effect.

Why the bar has to move as the haystack grows

This is the part that doesn’t show up anywhere in the printed output: the threshold for what counts as a real signal isn’t fixed, it depends on how many candidates you searched to find it. A 65% win rate found by testing one hand-picked setup means something. A 65% win rate found as the best of three hundred candidates means something much weaker, because you already know, from the math above, that a nontrivial number of zero-edge patterns were going to clear that same bar regardless of what the market actually did.

That’s part of why validated, durable win rates on real systems tend to sit in a narrower and less impressive-looking band, something closer to 52% to 62%, rather than the 70%+ numbers that show up near the top of an unfiltered training sort. A 70% win rate surviving a large combinatorial search isn’t a red flag because high win rates are inherently suspicious. It’s a red flag because it’s exactly the range where noise-driven results cluster once you’ve given noise enough tries to produce one.

The costs backtest math conveniently forgets

There’s a second filter that separates real edges from lucky ones, and it’s not statistical, it’s structural. This script charges TRADING_COST_POINTS of 0.03 per completed trade. On XAUUSD, real-world spread alone during normal London or New York session hours typically runs well above that, before slippage on entries and exits gets added, and before commission if the broker charges it separately. A backtest that undercharges for cost doesn’t just shave a little off the PnL column. It systematically favors exactly the noise-driven patterns discussed above, because those are the ones with the thinnest real margin to begin with. A pattern with a genuine edge tends to survive a more realistic cost model with room to spare. A pattern that only ever had a 50% coin-flip win rate dressed up as 60% by variance has no such cushion, and a cost model close to reality is often the fastest way to tell the two apart without waiting for a second dataset.

What actually distinguishes signal from this noise floor

Sample size is the honest answer, and it’s a less satisfying one than most people want. Twenty or thirty trades is not enough to separate a real 55% edge from a lucky 50% coin flip with any real confidence, no matter how the win rate prints. The gap closes as trade count grows into the hundreds, which is exactly why the out-of-sample test window matters more than the training sort ever will: it’s not just unseen data, it’s a second, independent sample that a noise-driven pattern has to get lucky in twice, and the odds of that compound in your favor fast.

If a pattern clears 60% on 25 training trades and then clears something in the low-to-mid 50s on a larger out-of-sample set, that’s not a strategy falling apart. That’s closer to what a real edge is supposed to look like once the coin-flip luck of a small sample gets averaged out by a bigger one. The uncomfortable version of that lesson is that the buzzy-looking training numbers were probably never real in the first place, and the plainer test numbers are the actual result.

Knowing how many patterns were noise-shaped by design is one problem. The next one is what happens the moment you look at that list and decide, based on nothing but a gut sense of which entries feel right, which thirty of three hundred deserve a second look.