Two independent ways a backtest overstates itself, both invisible from inside the backtest, both measured here on our own research.
Clustering. A t-statistic assumes independent draws. Trading signals arrive in bursts — one regime, one news day, one correlated basket — so forty trades are nowhere near forty independent observations. Recomputing by clustering on date and counting each date once, every t-statistic in our research shrank by a factor of about 1.9. One of ours went from 5.98 to 1.29.
Selection. Search a thousand variants, keep the winner, quote its solo p-value: that number is meaningless, because you did not run one test. Pure noise hands you a best-of-1000 t-statistic around 3.1 for free. We shifted our signals against price so the true answer was known to be nothing, ran the same search, and it reported a 39% false-discovery rate where the method advertised 10%.
Neither error is visible from inside the result. Both are arithmetic once you know to look, and both are why we no longer believe a backtest that has not been through them.
Run this on your own data.
Money Mind sells the check, not a strategy — we have no proven edge and do not sell signals. See the /multipletest spec (free) or browse the shop.