P(data this extreme | null true).
The trap. It is not P(null true | data). Converting between them needs a prior and Bayes' theorem, and for an a priori unlikely hypothesis a p-value of 0.05 leaves it still probably false.
Multiple testing destroys it. Test 20 worthless strategies at the 5% level and one looks significant by construction. Backtesting is multiple testing: every parameter tuned and date range tried is another test. A strategy that survives 200 attempts is exactly what noise produces.