Does FVG entry beat random entry? We're testing it — in public
Search for "do fair value gaps work" and you'll find two camps shouting past each other: believers with annotated winners, and skeptics pointing out that nobody shows their losers. Both are right about the other. What's missing from the argument is the boring thing that settles it: a test specified before the data arrives. This post documents ours, in enough detail that you can hold us to it. No results are presented here, because the honest answer today is: we don't know yet.
The claim under test
Not "FVGs work" — that's untestable. The precise claim: entries at the retest of the confirming leg's last FVG, taken only after the full sequence (H4 bias → H1 zone → sweep → displacement MSS), produce a pooled net result that beats duration-matched random entries after modeled costs. Every threshold in that sentence is published and frozen.
Why "beats random" is the bar
In a trending market, almost any long entry "works". In chop, almost nothing does. Quoting raw results without a null benchmark is how this industry manufactures conviction. So the engine's pooled net result is compared against thousands of simulated entries at random times on the same symbols, held for comparable horizons, paying the same modeled costs. If random does about as well, FVG entries — as this engine defines them — carry no information worth paying for.
The design, frozen before data
- Sample trigger: the test runs at 60 closed trades, or at 42 days with at least 30 — whichever comes first. Not when the numbers look good.
- Data: the live paper ledger only. No backtest, anywhere — historical data invites exactly the freedom pre-registration removes.
- Costs: taker fees, slippage and funding from one versioned model, identical for engine trades and random entries.
- Counting: a win is a target-hit on the stored bracket, pessimistic tie rule, costs included. No first-take-profit tricks, no maximum-favorable-movement accounting.
- Robustness: a split-half check and a second framing of the null test run alongside; if the two framings disagree, we publish the disagreement rather than the friendlier number.
What would make us wrong
Being falsifiable cuts both ways, so let's be explicit. If the pooled result clears the benchmark, that is one live sample clearing one null — it buys four more weeks of forward confirmation, not a profit claim. If it doesn't clear, then this mechanization of FVG entry — whatever its charting appeal — did not beat randomness in its live test, and the product will say so on its front page. A sibling research project of ours ran ~170 pre-registered experiments across 9 signal families and watched every single one die in this kind of test (internal and unpublished — our claim about ourselves, not independent evidence). We built this engine expecting the same fate and hoping otherwise. That's the right posture for a falsifiable product.
How to follow along
The sample accumulates on the open ledger in real time — losses included, small-sample labels on. The trigger state is on the evaluation page. When the verdict lands, it will be appended to this post, to the evaluation page, and to the changelog, whichever way it goes. If you joined the waitlist, you'll get it by e-mail — that mail goes out on pass and on fail alike.