Summary: We tested a popular retail setup — buy the 50-EMA breakout during the New York morning, stop below the breakout candle, take profit at 1:2 — on EURUSD, GBPUSD and XAUUSD across 16.5 years of hourly data. It lost 63%. Then we spent three more rounds trying to rescue it: higher-timeframe trend filters, volume confirmation, a different entry trigger, both 1H and 4H. Roughly 60 configurations, none with a real edge. Along the way a single trade lost 168× its intended risk and drove the account negative. This is the full teardown — including the test we should have run first.

There's a version of this article where we find a tweak that works, and you learn a new setup. This isn't that one. What we found instead is more useful: a way to prove, in a single chart, that an entry signal contains no information — so you can stop tuning it and go do something else.

The rules came in exactly as a trader would describe them:

The setup. Instruments: EURUSD, GBPUSD, XAUUSD. Timeframe: 1 hour. Enter long when price breaks above the 50 EMA, with the candle closing between 9:30am and 12:30pm New York. Stop loss below the candle that broke the EMA. Take profit at 1:2 risk-reward. Never hold two positions at once.

Every one of those clauses is doing work, and two of them turned out to matter far more than anyone expected.

Three Decisions Before a Single Trade

Before you can test a rule you have to make it unambiguous, and that's where quiet assumptions get baked in. Three came up here.

The session window doesn't fit the bars. Hourly candles close on the hour, so "closing between 9:30 and 12:30" leaves exactly three eligible closes per day: 10:00, 11:00 and 12:00 New York. There is no 9:30 or 12:30 hourly close. We also read "EST" as New York local time rather than a fixed UTC−5, since that's what a trader means — which shifts the eligible bars by an hour across daylight saving.

Gold is a forex ticker. On our data provider, XAUUSD lives on the currency endpoint as C:XAUUSD, not the crypto-style X: prefix. Small detail, and the run silently returns nothing if you get it wrong.

Costs are not optional. We charged a round-trip spread on every trade — 1.0 pip on EURUSD, 1.4 on GBPUSD, 30 cents on gold — split half on entry and half on exit. Hold that number in mind. It ends up being the entire story.

Round One: The Honest Baseline

Sixteen and a half years, 1,411 trades after the one-position-at-a-time rule, and the result was unambiguous.

VariantTradesWin rateTotal returnMax drawdownProfit factor
One trade at a time1,41132.9%−63.2%73.7%0.92
No concurrency limit2,04532.6%−79.8%88.4%0.93

Now the number that matters. At 1:2 risk-reward, you need to win 33.33% of the time to break even — one win of 2R pays for two losses of 1R. The strategy won 32.88%.

That's not "close to profitable." That's indistinguishable from a random entry. So we re-ran the identical test with the spread set to zero, to separate the signal from the cost:

Win rateExpectancy per trade
Gross — zero transaction cost33.83%+0.0015 R
Net — real spread32.88%−0.0606 R
Breakeven requirement33.33%0.000 R

Gross expectancy was +0.0015 R with a standard error of 0.038 — a t-statistic of 0.04, and a 95% confidence interval running from −0.074 to +0.077. In plain terms: after 1,411 trades and 16 years, we could not distinguish this entry from flipping a coin.

The spread, meanwhile, is not uncertain at all. It costs 6.2% of your risk on every single trade, deterministically. A zero-edge signal paying 6% of R in friction, 1,411 times, compounds to −63%. That's the whole mechanism.

A coin flip with a 2:1 payout is a fair game. A coin flip with a 2:1 payout and a toll booth is a slow, statistically inevitable loss. Most retail setups fail here — not because the idea is crazy, but because it's neutral, and neutral doesn't survive costs.

Round Two: The Filters Made It Worse

The obvious response is that a bare EMA cross is too naive — it needs context. So we added the two filters everyone reaches for: a higher-timeframe trend filter (only go long if price is above the 4H or daily EMA) and volume confirmation (only take breakouts on above-average volume). Twelve combinations.

All twelve lost money. And the trend filter — the one people trust most — made the per-trade edge worse in every variant:

Higher-timeframe filterTradesExpectancy (R)vs. no filter
None (baseline)1,411−0.061
Above 4H EMA50864−0.083worse
Above 1D EMA50789−0.093worse
Above 1D EMA200804−0.078worse

In hindsight it's intuitive. A 50-EMA cross is already a momentum condition. Demanding that price also sit above a higher-timeframe average keeps the crosses that fire late in an established move — exactly the ones most likely to be exhaustion, where a stop at the breakout candle's low gets taken out on the first pullback.

Volume confirmation looked better — total return improved from −63% to −45% — but that improvement is a mirage of a different kind, and it's worth naming because it recurs everywhere in this project:

The "fewer trades looks better" trap. When per-trade expectancy is negative, any filter that removes trades improves total return. It isn't finding good trades; it's taking fewer bad ones. The tell: as the volume filter tightened, the t-statistic moved toward zero (−1.58 to −0.99). Less evidence, not more edge.

One methodology note that cost us some certainty here. Forex is an over-the-counter market with no consolidated tape, so there is no true volume figure. What our provider reports for currency pairs is a quote-update tick count — a reasonable activity proxy, but not traded size. A genuine volume filter isn't testable on spot FX at all. If volume confirmation matters to you, the honest venue is a futures contract, where exchange volume is real.

Round Three: Remove the EMA — and Break the Account

Next we threw out the EMA entirely and replaced it with pure candle strength: enter when a candle's return exceeds the average return of all prior candles. Tested on 1H and 4H, two definitions of "return," three strictness levels. Twenty-four more cells.

The first thing we measured killed the idea in its literal form. The mean hourly return of a currency pair is 0.00000017 — indistinguishable from zero. So "return greater than average return" reduces to "the candle is green," and it fired on 49.7% of all eligible bars. That isn't a filter; it's a coin, and it produced the worst cells in the entire study — around −99% with t-statistics near −5. Those were the only statistically significant results we found all day, and they were significantly bad.

Then the results table showed something impossible: drawdowns of −101%, −113%, −118%, and a CAGR of NaN. You cannot lose 118% of your money. Equity had gone negative.

−168 R
The worst single trade in the study — a loss 168 times its intended risk, on a strategy where every trade was supposed to risk exactly 1%. Six such trades drove the account to −$18,574.

The cause is a genuine flaw in the rule as written, not just a coding slip. "Stop below the breakout candle" says nothing about how far below. When the signal candle is a doji — a tiny bar where open and close nearly touch — the stop sits a hair beneath the entry. Risk per unit approaches zero. And position size is risk budget divided by risk per unit, so as that denominator collapses, size explodes toward infinity. Then a weekend gap jumps straight through the hair-thin stop, and a "1% risk" trade takes 168%.

No broker would extend that leverage, so the unguarded numbers were fiction. We added two guards that any real implementation needs: a 30:1 leverage cap and a minimum stop distance of 0.03%, below which the trade is simply skipped. The blowups vanished; the 4H cells went from −92% and −114% to −73% and −69%. The conclusion didn't change — 4 of 24 cells finished with positive expectancy, none of them significant — but the numbers became real.

If you ever place a stop at a structural level — a candle's low, a swing point, yesterday's range — you need a minimum-distance rule. Without one, the quietest bar on your chart is the one that ends your account.

Round Four: The Test We Should Have Run First

By now one possibility remained. Maybe the entry did have signal, and the 1:2 target was simply the wrong shape for it — targets too far away, winners never reaching them. There's a clean way to settle that.

For any fixed stop with a target at R:R times the risk, the breakeven win rate is exactly:

1 / (1 + RR)
1:1 needs 50%. 1:2 needs 33.3%. 1:3 needs 25%. 1:5 needs 16.7%. If an entry carries no directional information, its observed win rate will track that curve at every risk-reward ratio. If it carries real information, some ratio will beat its own breakeven by more than noise.

This test is immune to the "fewer trades looks better" trap that made every earlier filter appear to help — because it compares each configuration against its own breakeven, not against total return. We swept six risk-reward ratios across four different entries, including the original specification, and measured everything gross of costs so we were testing the entry rather than the spread.

EntryR:RTradesGross win rateBreakevenEdgez
1H EMA50 cross0.51,69964.6%66.7%−2.1 pp−1.84
1H EMA50 cross1.01,58048.4%50.0%−1.6 pp−1.27
1H EMA50 cross2.01,41033.8%33.3%+0.4 pp0.32
1H EMA50 cross5.01,07916.9%16.7%+0.2 pp0.18
1H candle strength2.01,03334.6%33.3%+1.3 pp0.89
4H candle strength0.555068.2%66.7%+1.5 pp0.75
4H close-to-close2.031236.0%33.3%+2.7 pp1.01
4H close-to-close5.018518.5%16.7%+1.8 pp0.66

Eight of the 24 cells shown; the full sweep is in the tearsheet below. "pp" = percentage points. z = the edge divided by its standard error.

Zero of 24 cells reached statistical significance. Every observed win rate sat within one standard error of its theoretical breakeven. The mean edge across the entire sweep was +0.70 percentage points, against standard errors of 1.1 to 2.7 points.

That is the signature of an entry with no information. The 1:2 target was never cutting winners short — there were no winners to cut short. And because the result holds at every ratio from 1:0.5 to 1:5, no exit rule can rescue it. That single chart closed a line of investigation we could otherwise have tuned for weeks.

Two Things Worth Stealing

The original setup is slightly worse than random. The 50-EMA cross showed negative gross edge at tight targets — −2.1 points at 1:0.5, −1.6 at 1:1 — and it was the only entry with a negative mean across the sweep. Not significant at two standard errors, but it leans the same wrong way at all three of the tightest targets. The mild implication: price tends to pull back right after a 50-EMA cross in the New York morning. If anything, that cross is a marginally better fade than a follow.

The apparent improvement was a sample-size illusion. Across our four entries, mean gross edge climbed steadily — −0.55, +0.83, +1.07, +1.47 points — which looks exactly like filters doing their job. But every step also cut the sample, from 1,699 trades down to 312. The standard error grew right alongside the measured edge, and the significance never moved: 1.84, 0.89, 0.75, 1.01. The number got bigger and the evidence got weaker.

This is the most common way a research process fools itself. Each filter you add makes the headline number prettier and the sample smaller. Somewhere around the fourth filter you have a beautiful equity curve built on 40 trades and no statistical power whatsoever — and you'll believe it, because you watched it improve at every step.

What We Learned

Compare win rate to breakeven, not to zero. "32.9% win rate" is meaningless on its own. Against the 33.33% that 1:2 demands, it's the whole verdict in one line. Every risk-reward ratio has its own bar; a strategy is only interesting if it clears its own.

Measure gross before net. Separating the zero-cost result from the after-cost result tells you which problem you have. Gross positive and net negative is a cost problem — tradable on a better venue, a cheaper broker, a longer timeframe. Gross flat, as here, is a signal problem, and no amount of execution improvement will touch it.

Distrust every improvement that removes trades. If expectancy is negative, cutting trades always flatters the total return. Ask what happened to the t-statistic. If it drifted toward zero, you bought a nicer-looking chart with statistical power.

Put a floor under structural stops. A stop at a candle's low with no minimum distance is an unbounded position-size bug waiting for a quiet bar and a weekend gap. Ours cost 168R in a system designed to risk 1%.

Know what your data actually measures. We nearly reported a volume filter on an instrument class that has no volume. The field was populated, plausible, and meant something else entirely.

Where the Idea Goes

Nowhere, in this form — and that's a result worth having in an afternoon rather than a quarter. Four rounds, roughly 60 configurations, 16.5 years of hourly data across three instruments, and the premise that failed is a specific one: a single candle's direction inside a fixed morning window does not predict the next several hours in EURUSD, GBPUSD or XAUUSD. Both trigger families said so, and all six risk-reward ratios agreed.

If we picked this thread up again, we wouldn't add a seventh filter. We'd change the premise — test it where exchange volume is real, so volume confirmation becomes a genuine variable rather than a proxy, or replace the entry with something that has a documented economic reason to work rather than a shape on a chart. The tooling stays; only the hypothesis changes. That's the point of building the harness once.

None of this makes for a triumphant write-up. But the alternative — publishing the one cell out of 60 that finished green, at a t-statistic of 1.01, and calling it a strategy — is how most retail setups get born. We'd rather show you the graveyard and the shovel.

Who we built this for

Institutional-grade backtesting — expectancy in R, breakeven-adjusted win rates, significance testing, cost modelling — has mostly lived behind the walls of hedge funds and prop desks. The everyday trader almost never gets to touch it. That's exactly who we built July for.

It's open-source and completely free, because we believe the person risking their own money deserves the same tools — and the same honesty about what those tools really show — that the professionals use. This teardown isn't a strategy we're selling you; it's a worked example of how to kill a bad idea quickly, mistakes and blown-up accounts included, so anyone can build real research instincts at no cost. If you want to learn to tell a genuine edge from a coin flip, that's what we're here to help with. July is free to download and use →

Download the full tearsheets

Every round from the session, exported as a full PDF report — complete result grids, equity curves, expectancy charts and significance tests. Every configuration we ran is in here, including the ones that lost 99%.

Round 2 — trend & volume filter grid (0 of 12 profitable)The 50-EMA baseline plus every higher-timeframe and volume combination. PDF Round 3 — candle strength, 1H and 4H (24 cells)The EMA removed entirely — and the leverage bug that drove equity negative. PDF Round 4 — the risk-reward sweep (the decisive test)Win rate against theoretical breakeven at six ratios — every series lands on the line. PDF
Light Water Capital  ·  July 2026