What Works on Wall Street Ch. 5: Does It Still Work?
阅读中文版Data mining, survivorship bias, post-publication factor decay, and what happened to O'Shaughnessy's own funds.
🔊 Listen to Article (Chinese Audio)
What Works on Wall Street Ch. 5: Does It Still Work?
"Test enough strategies against the same data and you will certainly find one that looks superb — even if it means nothing." — the central caution of quantitative research
Investment Context
The first four chapters present impressive backtested numbers. This one asks the question that must be answered: how much of that can be trusted?
This is not nitpicking. A systematic gap separates backtests from live results, and understanding where that gap comes from matters more than memorising any strategy's annualised return.
The Wall Street Translation
1. The Data-Mining Problem
O'Shaughnessy tested hundreds of factor combinations. When you test enough combinations against the same history, the best performer necessarily owes part of its result to luck rather than genuine regularity.
Statisticians call this the multiple comparisons problem: test 100 worthless strategies and roughly five will appear "significant" at 95% confidence by chance alone. The strategy with the best backtest is precisely the one most likely to contain the most luck.
2. Survivorship Bias and Other Data Traps
| Trap | Meaning | Consequence |
|---|---|---|
| Survivorship bias | Database excludes delisted companies | Overstates historical returns |
| Look-ahead bias | Uses financials not yet published at the time | Overstates achievability |
| Ignored trading costs | Backtests fill at closing prices | Omits spreads and market impact |
| Small-cap liquidity | Best results often come from micro-caps | Large money cannot actually execute |
The last matters most: many factors concentrate their excess return in the smallest, least liquid stocks — exactly the ones real money finds hardest to buy.
3. Post-Publication Decay
Academic work repeatedly finds the same pattern: after a factor is published, its excess return decays on average by a third to a half. Arbitrage capital arrives, and the luck embedded in the original result fails to repeat.
O'Shaughnessy's own funds provide an honest test. He founded an asset management firm on this research and launched public funds, but their live results did not replicate the book's backtests, and several products were eventually closed or folded into other strategies. This is not a criticism of him personally — it is the common fate of nearly every published quantitative strategy.
4. So Is It Worthless?
No. Value and momentum are the two most robust factors in the literature, supported across markets, periods, and asset classes. The sound conclusion is not that factors do not work but that the excess return actually available is far smaller than backtests show, and comes with the risk of failing for long stretches.
Actionable Trading Rules
- Halve every backtest: Cut a published strategy's backtested excess return in half as a realistic forward expectation, and execute only if it still justifies the effort.
- Distrust over-engineered strategies: The more parameters and rules, the more likely the result is data mining. Simple robust factors are more credible than elaborate combinations.
- Check liquidity feasibility: If a strategy's returns come mainly from micro-caps, it may work for small personal sums — but recognise you cannot verify it still holds at your trade size.
Relevance to a Retirement Portfolio
This chapter's value to retirees extends well past quantitative investing: it teaches you to evaluate any financial product sold on historical performance.
Every fund brochure shows an excellent past track record, because products that performed badly are never marketed. Every attractive backtest you are shown has already passed through a survivorship filter. For anyone in the withdrawal phase who cannot absorb a serious mistake, that scepticism is essential protection.