Here is a trade I would have taken a year ago without thinking about it.
A stock has a violent week. Realised volatility triples. The obvious inference is that the next few weeks will also be wild, so you widen the stop, cut the size, maybe sell some premium into it. Volatility clusters. Everybody knows this. It is one of the oldest documented facts in finance.
The question I could not answer was a narrower one: how much of that is the moment, and how much of it is just this stock? Some names are permanently jumpy. If you rank a thousand stocks by trailing volatility and the same names sit near the top every month for twenty years, then your “volatility forecast” is not forecasting anything. It is a roster.
That distinction turned out to matter far more than I expected, and it is the reason this project exists. Before testing whether a strategy works, we wanted to know whether the market behaviour underneath it is real, and what specifically it is made of.
Seven modules have now run. Six are closed. The seventh is deliberately held behind an external review gate, and the eighth is blocked behind it until that clears.
The Test That Broke Almost Everything
The device that did the most damage is embarrassingly simple, and I would recommend it to anyone testing a cross-sectional signal.
Take your panel. Shuffle each stock's history within that stock - keep every observation the name ever produced, destroy the order they arrived in. Now re-run the exact same test. Any predictive power that survives cannot be about timing, because you deleted time. It can only be the fact that some names are permanently different from other names.
We call that the personality floor. Whatever your signal scores above it is the part that is genuinely about now.
The results were not close.
Read the top bar as a practical instruction. A twenty-day volatility screen is roughly 70% a list of names that are always volatile and 30% a statement about current conditions. Hold it out to a year and the state component nearly disappears.
This does not make volatility clustering false. It survived every null we threw at it. What it changes is what you should do with it. Sizing a position off a stock's own volatility level is sound, because the level is stable and that stability is the whole point. Rotating capital between stocks on trailing volatility is a much weaker idea than the raw statistic suggests, because you are mostly re-buying the same roster and paying spread to do it.
There was one clean exception, and it is the reason I trust the method rather than suspecting it just destroys everything. Liquidity state genuinely persists. When a name gets harder to trade, it stays harder to trade, and that survives the shuffle with room to spare. The tool is capable of returning a positive answer. It just did not return many.
The most unsettling result in the whole module was a time-series one. Between 2005-2009 and 2023-2026, the genuine state component of volatility clustering collapsed by roughly four times, while the headline statistic everybody quotes went slightly up. Twenty years of structural change in how volatility behaves, completely invisible in the number on the dashboard.
Variance Is Predictable. Direction Is Not.
If you take one thing away, take this one. It is the most replicated finding in the program - three modules, built by different people at different times, sharing nothing but the underlying panel.
Look at those two grey slivers next to the teal bars. That gap is why the parts of trading that work tend to be the parts that do not require you to be right about direction.
Concretely, this is the difference between two things that look superficially similar on a chart:
- Sizing off a 20-day ATR. You are forecasting something with a rank correlation around 0.62. That is a real, usable, stable relationship, and it is why volatility targeting and risk parity survive contact with live markets.
- Filtering entries on 20-day momentum. You are forecasting something with a rank correlation of roughly negative 0.016. That is not a weak edge. That is the absence of one.
The same asymmetry showed up in the shock module and the volume module independently. Unusual volume predicts future volatility with overwhelming statistical strength. The same variable predicts the future sign with something roughly an order of magnitude weaker - real, but nothing you would build on alone.
Most retail strategy design spends its effort on the half of the problem where there is nothing to find.
The Same Test, Two Answers: p = 0.000 and p = 0.888
The seventh module went after the statistic the industry quotes more than any other regime variable: average pairwise correlation, and the share of market variance explained by its first principal component. When correlation rises, the story goes, diversification stops working and the regime has changed.
We tested 155 separate claims that these measures persist - that a high reading today tells you something about the reading tomorrow. Under conventional inference, they sail through.
Then we corrected for one thing. These measures are computed on rolling windows. A 120-day rolling correlation observed today and the same measure observed tomorrow share 119 of their 120 days. They are not two observations. They are one observation, reported twice, with a rounding error between them. Do that across a decade and your five thousand data points carry roughly the information content of forty.
That turn matters, and I want to be careful about it. The failure is specifically about forecasting the level. Testing whether a high-correlation state behaves differently from a low-correlation state is a different question, and those tests overwhelmingly survive. Correlation describes where you are. It does not tell you where the reading goes next.
Which is a genuinely useful thing to know before you build a regime filter on it. If your model treats a correlation reading as a forecast, the number underneath it is doing far less work than your backtest suggests.
Module seven is not closed. It sits at an external review gate that returned a hold, and every claim in it carries a flag saying no causal proof was established. I would rather publish it in that state than round it up.
Sixty Percent of the Crash Was the Companies That Never Came Back
The shock module started with a clean-looking result: after an unusually large negative move, stocks underperform for weeks. A tidy, tradeable-sounding decline.
Then we rebuilt the universe properly. Our corpus carries 36,684 securities, 20,911 of which are delisted - the ones that merged, went private, or went to zero, each present with the dates they actually traded and absent afterwards. Most datasets you can buy quietly drop them.
Sixty percent of the effect was survivorship, running the other way from how people usually describe it. And a separate headline result in the same module - a “structural change in market behaviour since 2005” - evaporated completely once the universe was handled correctly. It had never been a change in the market. It was a change in who was in the sample.
The practical version: if you are testing a buy-the-crash rule against today's index members, the companies that crashed and never recovered are not in your test. You have pre-selected for recovery and then measured recovery.
We Broke Our Own Tools Eighteen Times
Every module has a ledger of defects we found in our own machinery. It currently holds eighteen entries. Ten of them would have produced a confident wrong answer rather than a visible error - no crash, no warning, just a number that looked fine.
The pattern we learned to fear is not “obviously wrong.” It is “suspiciously clean.” A result with no ragged edges, no unexplained outliers and no awkward subperiod is usually not a discovery. It is a leak.
A few of them are worth naming because they generalise well beyond this project.
A guard nobody has watched fail is not known to be a guard. We had a look-ahead check that had passed on every run since it was written. When we finally forced a violation through it deliberately, it did not fire. It had been passing because it was never reachable.
An exclusion window has to point backwards. One control excluded the days within a fortnight of an event on both sides. Excluding days after the event means the control set is chosen using information from the future. Subtle, silent, and it inflates the result.
Two effects that come from the same helper function will always agree. We had two independent-looking detectors corroborating each other. They shared a routine, so their agreement was arithmetic rather than evidence.
The self-consistency checker that audits the ledger has itself been rewritten three times, because it kept passing runs it should have failed.
What I Am Not Claiming
This is the section I would want to read first if someone handed me this article, so it is going here rather than in a footnote.
None of this is a strategy. Nothing in the program was backtested with costs, sizing, capacity or execution. No artifact across seven modules carries a return claim, and none of these numbers convert into one. A rank correlation of 0.62 is not 62% of anything.
One of our strongest findings does not generalise. The first module found something clean: stock-specific moves reverse hardest in the short run and continue hardest in the medium run, and both effects strengthen when you isolate the most idiosyncratic names. The sixth module was built specifically to test that across a full specification grid, and it survives in 15 of 3,240 specifications. That tension is unresolved, and we are publishing it as unresolved rather than picking the module we prefer.
Some results were withdrawn. Findings that had already been written up were pulled after the machinery underneath them turned out to be defective. Those are on the record too. A research programme that only publishes its survivors is exactly the thing this one exists to argue against.
What This Changed About How I Test Anything
Five rules came out of this that I now apply before I look at a single performance number.
-
Shuffle within the name before you believe a cross-sectional signal.
If the effect survives having time deleted, it is a roster, not a forecast. This one test would have killed several ideas I was previously fond of, in about ten minutes each.
-
Count your independent observations, not your rows.
Overlapping windows are the most common way a research process convinces itself. If consecutive observations share most of their input, your sample is a fraction of what your spreadsheet says, and every t-statistic is inflated by roughly the square root of the overlap.
-
Forecast size, not sign, wherever you have the choice.
The evidence for variance is between one and two orders of magnitude stronger than the evidence for direction. Build the parts of the system that can run on that.
-
Test on a universe that includes the dead.
If your symbol list came from today's index, your backtest has already been told which companies survived.
-
Write down every defect, especially the silent ones.
Ten of our eighteen would have shipped a confident wrong answer. The ledger is the only reason we know that.
Where This Goes Next
Three phenomena remain unexecuted, and the eighth is blocked until the seventh clears its review gate. That is deliberate. The whole point of the sequence is that a module whose inference is under question cannot be used as a foundation for the next one.
What we have now is not a strategy and was never meant to be. It is a set of things we can say about US equity behaviour with the corrections applied, a considerably longer set of things we have stopped saying, and a tested toolkit for telling those two apart. When a strategy idea arrives, the first question is no longer whether it backtests well. It is which of these seven behaviours it is claiming to exploit, and whether that behaviour survived its own null.
Most ideas do not get past that question. That is the point.
The tools that separate a real market effect from a well-dressed artifact - shuffle nulls, overlap-aware standard errors, survivorship-complete universes - have mostly lived inside hedge funds and academic departments. The person trading their own money almost never gets to touch them.
That is who we built July Backtester for, and it is why this write-up leads with what failed. Publishing only the survivors is how a research process fools itself and everyone reading it. If you want to learn to tell a genuine effect from a coin flip, the engine is open-source and free. July Backtester is free to download and use →
Views are my own and those of Light Water Capital, provided for information only; nothing here is investment advice or a recommendation, and no result described above was tested as a tradeable strategy. Figures are from an internal US equity research panel covering 2004-2026, built on a delisting-complete security corpus. Module seven remains under external review and no causal claim is made for it.
Leave a comment - join the conversation
Which signal in your own process would survive having time shuffled out of it? Tell us the one you would put through the test first.
Comment on LinkedIn