Summary: Instead of testing strategies, we spent six days testing the seven market behaviours that strategies are built on top of - momentum, volatility clustering, shocks, volume, resiliency, commonality and correlation - across US equities back to 2004. Two corrections did most of the damage. Subtracting each stock's permanent character removed 70% to 87% of what looked like volatility memory. Correcting for overlapping windows took 155 correlation-persistence results down to zero. Along the way we broke our own measuring tools eighteen times, ten of them silently. This is what was left standing.

Here is a trade I would have taken a year ago without thinking about it.

A stock has a violent week. Realised volatility triples. The obvious inference is that the next few weeks will also be wild, so you widen the stop, cut the size, maybe sell some premium into it. Volatility clusters. Everybody knows this. It is one of the oldest documented facts in finance.

The question I could not answer was a narrower one: how much of that is the moment, and how much of it is just this stock? Some names are permanently jumpy. If you rank a thousand stocks by trailing volatility and the same names sit near the top every month for twenty years, then your “volatility forecast” is not forecasting anything. It is a roster.

That distinction turned out to matter far more than I expected, and it is the reason this project exists. Before testing whether a strategy works, we wanted to know whether the market behaviour underneath it is real, and what specifically it is made of.

Seven modules have now run. Six are closed. The seventh is deliberately held behind an external review gate, and the eighth is blocked behind it until that clears.

The Test That Broke Almost Everything

The device that did the most damage is embarrassingly simple, and I would recommend it to anyone testing a cross-sectional signal.

Take your panel. Shuffle each stock's history within that stock - keep every observation the name ever produced, destroy the order they arrived in. Now re-run the exact same test. Any predictive power that survives cannot be about timing, because you deleted time. It can only be the fact that some names are permanently different from other names.

We call that the personality floor. Whatever your signal scores above it is the part that is genuinely about now.

The results were not close.

What is left after you subtract the stock itself Share of measured persistence that survives a within-name shuffle. US equities, 2004-2026. Volatility persistence, 20 days 70% 30% Volatility persistence, one year 87% 13% Underperformance after a shock 63% 37% The stock's permanent character Genuine state - the part about now The exception: liquidity state clears the same floor comfortably. That one really does move. Source: Light Water Capital research - US equity panel, 2004-2026 (modules V03, V04, S01)
Rank a thousand stocks on trailing volatility and hold the list for a year, and roughly seven-eighths of what you have measured is a permanent roster rather than a forecast.

Read the top bar as a practical instruction. A twenty-day volatility screen is roughly 70% a list of names that are always volatile and 30% a statement about current conditions. Hold it out to a year and the state component nearly disappears.

This does not make volatility clustering false. It survived every null we threw at it. What it changes is what you should do with it. Sizing a position off a stock's own volatility level is sound, because the level is stable and that stability is the whole point. Rotating capital between stocks on trailing volatility is a much weaker idea than the raw statistic suggests, because you are mostly re-buying the same roster and paying spread to do it.

There was one clean exception, and it is the reason I trust the method rather than suspecting it just destroys everything. Liquidity state genuinely persists. When a name gets harder to trade, it stays harder to trade, and that survives the shuffle with room to spare. The tool is capable of returning a positive answer. It just did not return many.

The most unsettling result in the whole module was a time-series one. Between 2005-2009 and 2023-2026, the genuine state component of volatility clustering collapsed by roughly four times, while the headline statistic everybody quotes went slightly up. Twenty years of structural change in how volatility behaves, completely invisible in the number on the dashboard.

Variance Is Predictable. Direction Is Not.

If you take one thing away, take this one. It is the most replicated finding in the program - three modules, built by different people at different times, sharing nothing but the underlying panel.

The market tells you how much, not which way Rank correlation between a stock's past and its future. Same panel, same windows, same test. Future variance, 5 days 0.77 Future variance, 20 days 0.62 Future direction, 5 days +0.034 Future direction, 20 days −0.016 Sorted into buckets, variance ranked in the right order 9 times out of 9. Direction managed 4 of 9 - a coin. Source: Light Water Capital research - US equity panel, 2004-2026 (module V02, corroborated by S16 and L01/L03)
Depending on the horizon, variance is between 23 and 354 times more predictable than direction. The bars for direction are not small because the effect is subtle. They are small because there is almost nothing there.

Look at those two grey slivers next to the teal bars. That gap is why the parts of trading that work tend to be the parts that do not require you to be right about direction.

Concretely, this is the difference between two things that look superficially similar on a chart:

  • Sizing off a 20-day ATR. You are forecasting something with a rank correlation around 0.62. That is a real, usable, stable relationship, and it is why volatility targeting and risk parity survive contact with live markets.
  • Filtering entries on 20-day momentum. You are forecasting something with a rank correlation of roughly negative 0.016. That is not a weak edge. That is the absence of one.

The same asymmetry showed up in the shock module and the volume module independently. Unusual volume predicts future volatility with overwhelming statistical strength. The same variable predicts the future sign with something roughly an order of magnitude weaker - real, but nothing you would build on alone.

Most retail strategy design spends its effort on the half of the problem where there is nothing to find.

The Same Test, Two Answers: p = 0.000 and p = 0.888

The seventh module went after the statistic the industry quotes more than any other regime variable: average pairwise correlation, and the share of market variance explained by its first principal component. When correlation rises, the story goes, diversification stops working and the regime has changed.

We tested 155 separate claims that these measures persist - that a high reading today tells you something about the reading tomorrow. Under conventional inference, they sail through.

Then we corrected for one thing. These measures are computed on rolling windows. A 120-day rolling correlation observed today and the same measure observed tomorrow share 119 of their 120 days. They are not two observations. They are one observation, reported twice, with a rounding error between them. Do that across a decade and your five thousand data points carry roughly the information content of forty.

0.000 → 0.888
The same p-value, on the same statistic, on the same window - computed conventionally, then computed with an overlap-aware standard error. It went from the most significant result in the module to indistinguishable from noise.
You can describe the regime. You cannot forecast its level. Claims about correlation and dispersion surviving overlap-aware inference, out of 1,130 tested Does the level persist? 0 of 155 survived Does the state differ from other states? 925 of 975 The turn: correlation is not meaningless. It is a thermometer, not a forecast. Source: Light Water Capital research - US equity panel, 2004-2026 (module 7, held at external review)
Zero of 155 persistence claims survived. But 925 of 975 contrast claims did - so the finding is not that correlation is useless. It is that the thing it is useful for is not the thing it gets used for.

That turn matters, and I want to be careful about it. The failure is specifically about forecasting the level. Testing whether a high-correlation state behaves differently from a low-correlation state is a different question, and those tests overwhelmingly survive. Correlation describes where you are. It does not tell you where the reading goes next.

Which is a genuinely useful thing to know before you build a regime filter on it. If your model treats a correlation reading as a forecast, the number underneath it is doing far less work than your backtest suggests.

Module seven is not closed. It sits at an external review gate that returned a hold, and every claim in it carries a flag saying no causal proof was established. I would rather publish it in that state than round it up.

Sixty Percent of the Crash Was the Companies That Never Came Back

The shock module started with a clean-looking result: after an unusually large negative move, stocks underperform for weeks. A tidy, tradeable-sounding decline.

Then we rebuilt the universe properly. Our corpus carries 36,684 securities, 20,911 of which are delisted - the ones that merged, went private, or went to zero, each present with the dates they actually traded and absent afterwards. Most datasets you can buy quietly drop them.

60%
Share of the measured post-shock decline that was not a market effect at all, but the universe changing composition underneath the measurement window as names dropped out.

Sixty percent of the effect was survivorship, running the other way from how people usually describe it. And a separate headline result in the same module - a “structural change in market behaviour since 2005” - evaporated completely once the universe was handled correctly. It had never been a change in the market. It was a change in who was in the sample.

The practical version: if you are testing a buy-the-crash rule against today's index members, the companies that crashed and never recovered are not in your test. You have pre-selected for recovery and then measured recovery.

We Broke Our Own Tools Eighteen Times

Every module has a ledger of defects we found in our own machinery. It currently holds eighteen entries. Ten of them would have produced a confident wrong answer rather than a visible error - no crash, no warning, just a number that looked fine.

The pattern we learned to fear is not “obviously wrong.” It is “suspiciously clean.” A result with no ragged edges, no unexplained outliers and no awkward subperiod is usually not a discovery. It is a leak.

A few of them are worth naming because they generalise well beyond this project.

A guard nobody has watched fail is not known to be a guard. We had a look-ahead check that had passed on every run since it was written. When we finally forced a violation through it deliberately, it did not fire. It had been passing because it was never reachable.

An exclusion window has to point backwards. One control excluded the days within a fortnight of an event on both sides. Excluding days after the event means the control set is chosen using information from the future. Subtle, silent, and it inflates the result.

Two effects that come from the same helper function will always agree. We had two independent-looking detectors corroborating each other. They shared a routine, so their agreement was arithmetic rather than evidence.

The self-consistency checker that audits the ledger has itself been rewritten three times, because it kept passing runs it should have failed.

What I Am Not Claiming

This is the section I would want to read first if someone handed me this article, so it is going here rather than in a footnote.

None of this is a strategy. Nothing in the program was backtested with costs, sizing, capacity or execution. No artifact across seven modules carries a return claim, and none of these numbers convert into one. A rank correlation of 0.62 is not 62% of anything.

One of our strongest findings does not generalise. The first module found something clean: stock-specific moves reverse hardest in the short run and continue hardest in the medium run, and both effects strengthen when you isolate the most idiosyncratic names. The sixth module was built specifically to test that across a full specification grid, and it survives in 15 of 3,240 specifications. That tension is unresolved, and we are publishing it as unresolved rather than picking the module we prefer.

Some results were withdrawn. Findings that had already been written up were pulled after the machinery underneath them turned out to be defective. Those are on the record too. A research programme that only publishes its survivors is exactly the thing this one exists to argue against.

What This Changed About How I Test Anything

Five rules came out of this that I now apply before I look at a single performance number.

  1. Shuffle within the name before you believe a cross-sectional signal.

    If the effect survives having time deleted, it is a roster, not a forecast. This one test would have killed several ideas I was previously fond of, in about ten minutes each.

  2. Count your independent observations, not your rows.

    Overlapping windows are the most common way a research process convinces itself. If consecutive observations share most of their input, your sample is a fraction of what your spreadsheet says, and every t-statistic is inflated by roughly the square root of the overlap.

  3. Forecast size, not sign, wherever you have the choice.

    The evidence for variance is between one and two orders of magnitude stronger than the evidence for direction. Build the parts of the system that can run on that.

  4. Test on a universe that includes the dead.

    If your symbol list came from today's index, your backtest has already been told which companies survived.

  5. Write down every defect, especially the silent ones.

    Ten of our eighteen would have shipped a confident wrong answer. The ledger is the only reason we know that.

Where This Goes Next

Three phenomena remain unexecuted, and the eighth is blocked until the seventh clears its review gate. That is deliberate. The whole point of the sequence is that a module whose inference is under question cannot be used as a foundation for the next one.

What we have now is not a strategy and was never meant to be. It is a set of things we can say about US equity behaviour with the corrections applied, a considerably longer set of things we have stopped saying, and a tested toolkit for telling those two apart. When a strategy idea arrives, the first question is no longer whether it backtests well. It is which of these seven behaviours it is claiming to exploit, and whether that behaviour survived its own null.

Most ideas do not get past that question. That is the point.

Why we publish the negatives

The tools that separate a real market effect from a well-dressed artifact - shuffle nulls, overlap-aware standard errors, survivorship-complete universes - have mostly lived inside hedge funds and academic departments. The person trading their own money almost never gets to touch them.

That is who we built July Backtester for, and it is why this write-up leads with what failed. Publishing only the survivors is how a research process fools itself and everyone reading it. If you want to learn to tell a genuine effect from a coin flip, the engine is open-source and free. July Backtester is free to download and use →


Views are my own and those of Light Water Capital, provided for information only; nothing here is investment advice or a recommendation, and no result described above was tested as a tradeable strategy. Figures are from an internal US equity research panel covering 2004-2026, built on a delisting-complete security corpus. Module seven remains under external review and no causal claim is made for it.

Light Water Capital  ·  September 2026