Research

Do Fair Value Gaps Work? A 23,899-Trade Backtest With Costs

Evidence category: Historical backtest

At a glance

In our fair value gap (FVG) retest backtest on NQ, ES, EURUSD and XAUUSD, 23,899 completed historical events averaged +0.020R before costs and −0.240R after the stated cost assumptions, where 1R is the width of the gap zone. Two of the twelve market-and-timeframe datasets, NQ 15-minute and NQ 1-hour, were positive after costs; the other ten were negative. This is a mechanical event study with overlapping trades, not a portfolio return or a verdict on every FVG strategy.

On this page
The RoboXpert robot examines a short row of candlestick blocks through a magnifying glass, holding a blank clipboard. Headline: “Do fair value gaps work?”
AI-generated illustration with a headline added by RoboXpert. The charts below are from the study’s own results.

A fair value gap is easy to mark after the chart has unfolded. A test needs more: a definition available at the time, a specific entry, an exit, a cost model and a comparison.

We tested a deliberately simple FVG retest and inverted FVG retest across NQ, ES, EURUSD and XAUUSD on three timeframes. The study was first run on 29 September 2026. For this article, we inspected its implementation and reran all twelve original datasets on 1 October; the results matched the original output apart from the generation timestamp.

The main finding is narrow: this pooled mechanical retest model was negative after the stated costs. It is not proof that all FVG-based strategies fail. A different entry, filter, exit or execution assumption is a different test.

Our follow-up trading cost sensitivity analysis freezes these baseline retest events and varies the cost allowance. It shows the zero-mean thresholds and cost distribution without changing entries or exits.

This event study does not model an account equity curve. A sum of overlapping event outcomes cannot establish equity drawdown; that requires an account path with sizing and simultaneous exposure.

The FVG and inverted FVG retest rules we tested

At the close of bar t, a bullish FVG exists when low[t] > high[t−2]. Its zone lies between the earlier high and the current low. The bearish definition reverses those relationships. This describes a three-candle pattern; it does not by itself prove institutional orders, missing liquidity or future price direction.

Scroll horizontally if needed

RuleFVG retest AInverted FVG retest B
Minimum zone size0.25 × Wilder ATR(14)Same original zone
Eligible fromThe bar after formationThe bar after a close through the far edge
EntryFirst retest at the near edge, in the gap’s directionFirst retest of the broken far edge, against the original gap
Stop and targetOpposite edge; target 2 zone widths from the entry levelOther edge; target 2 zone widths from the entry level
Expiry100 chart bars after formationBreak must occur within 100 bars; retest within 100 bars of the break
Time exitClose after 200 chart bars from the fillSame

Here 1R means the original zone width in price units. It is not a percentage of an account. Orders become eligible only after the relevant confirming bar. The session-gap filter rejects a formation when the first-to-third-bar time distance exceeds three chart intervals.

For a hypothetical bullish zone from 100 to 102, the standard A entry level is 102, the stop level is 100 and the target is 106. Those numbers illustrate the rule; they are not a trade from the market sample. A gap through an entry or exit level can alter the realized model payoff.

The engine permits overlapping events from multiple zones. It does not allocate capital, impose margin limits or combine exposures into a tradable portfolio. Do not turn the summed event results into an account-return claim.

Pooled FVG retest results before and after costs

Study chart comparing mean R before and after stated costs for FVG, IFVG and control events; all five pooled after-cost values are negative.
Original chart from RoboXpert FVG lab v1.00, run 29 September 2026. Historical event study; fixed cost assumptions, unequal periods and overlapping events.

The main setting produced 23,899 completed FVG retest events (model A) and 19,320 completed inverted FVG retest events (model B); each event is one simulated retest trade. The table reports trade-weighted means across all twelve datasets, in R (1R = the zone width of that gap). All five rows, including the opposite-direction and random-entry controls, have a negative mean after the stated cost assumptions.

Scroll horizontally if needed

ConstructionCompleted eventsPositive after costsMean gross RMean net RNet profit factor
FVG retest A23,89933.34%+0.0201−0.24030.714
Same FVG entries, opposite direction23,90232.78%+0.0020−0.25850.694
Random entries, ten seeds combined238,94132.76%−0.0014−0.26190.691
IFVG retest B19,32032.13%−0.0112−0.28510.670
Same IFVG entries, opposite direction19,32134.21%+0.0448−0.22910.726

Positive after costs means net R greater than zero. It is not the same metric as reaching the 2R target. The FVG retest model’s (A) gross-positive share was 33.80%. The familiar 33.33% gross break-even rate assumes every trade ends at exactly −1R or +2R; time exits, gaps and costs mean that shortcut does not describe the net results above.

The opposite-direction controls reuse the model entry opportunities. Counts differ slightly because the engine omits positions that remain unfinished at the dataset’s end. The random control chooses chart bars and directions with fixed seeds 0–9 and draws risk sizes from completed A events in the same dataset. Ten seeds provide repeated simulations on the same history—not ten independent market histories.

Results by market and timeframe

Some cells were positive. Selecting only those cells would tell a different story from the pooled result, but would also select winners after seeing the results.

Scroll horizontally if needed

DatasetIntrabar dataSample starts, UTCFVG retest (A) eventsMean gross RMean net R
NQ M51m18 Jun 20262,030−0.0191−0.0636
NQ M153m21 Nov 20251,877+0.0502+0.0215
NQ H115m30 Apr 20231,834+0.0913+0.0735
ES M51m18 Jun 20262,073+0.0131−0.2124
ES M153m21 Nov 20251,859+0.0287−0.0836
ES H115m30 Apr 20231,815+0.0109−0.0514
EURUSD M51m23 Jun 20262,534−0.0556−1.2599
EURUSD M153m8 Dec 20252,115+0.0575−0.4627
EURUSD H115m2 Jul 20231,853+0.0484−0.1652
XAUUSD M51m18 Jun 20262,168+0.0321−0.1391
XAUUSD M153m21 Nov 20251,851+0.0402−0.0398
XAUUSD H115m30 Apr 20231,890−0.0302−0.1315

All samples end on 29 September 2026, between 10:00 and 10:30 UTC; exact first/last timestamps and file hashes are in the download. Each retained sample contains approximately 20,000 chart bars. The histories overlap and have unequal lengths, so this is not an equal-period tournament between timeframes.

Mean net R across the twelve FVG datasets: NQ 15-minute and hourly are positive; the other ten are negative, with EURUSD 5-minute the lowest.
All twelve FVG cells, including the positive ones. Different start dates and costs limit direct comparisons; the best observed cell is not an independently validated selection.

Futures (TradingView symbols NQ1! and ES1!) came from TradingView’s delayed CME_MINI_DL continuous-contract feeds; EURUSD and XAUUSD came from OANDA. The export did not establish the futures back-adjustment state conclusively. Roll-period jumps remain in the data. We have not verified these results against individual futures contracts or another data provider.

Why small gaps can be expensive to trade

The engine subtracts a fixed round-trip cost in price units, divided by each event’s zone width:

cost in R = assumed round-trip cost / zone width
net R = model gross R − cost in R

Scroll horizontally if needed

MarketAssumed round-trip cost in price units
NQ0.50 index points
ES0.35 index points
EURUSD0.00012, equivalent to 1.2 pips
XAUUSD0.35 USD per ounce in quoted price

These are study assumptions, not current broker quotes or a complete execution model. They stand in for friction; actual spread, commissions, slippage, financing and fills depend on the instrument and venue.

For a hypothetical EURUSD gap of 0.00012, the assumed 0.00012 round-trip cost is 1R. That explains how a narrow stop can make apparently small price costs large relative to the model risk. In the actual EURUSD M5 sample, the median cost was 1R and the mean gross/net difference was about 1.204R. A median is not the average deduction.

Changing the minimum gap size also changed the selected events. With no ATR minimum, the FVG retest model’s (A) mean net result was −0.9454R across 45,660 events; with a 0.5 ATR minimum, it was −0.1358R across 11,812. Both remained negative under the same market cost assumptions. These are the two predefined sensitivity settings, not an optimized threshold recommendation.

Later-period and direction checks

The engine divides each dataset at 70% of its retained bars and assigns events by their fill time. No parameter is fitted on the first segment.

Scroll horizontally if needed

ConstructionEarlier 70%: events / mean net RLater 30%: events / mean net R
FVG A16,714 / −0.2290R7,185 / −0.2666R
IFVG B13,502 / −0.2800R5,818 / −0.2970R

This is a historical stability check. The result file calls the later segment “out of sample”, but the split does not purge trades crossing the boundary, and the random-control risk distribution uses completed events across the full dataset. It is not an untouched prospective trial, a demo-forward test or live evidence. See backtests versus forward and live results.

A separate symmetric ±1R test asks which side is reached first from the entry level. The expected-direction share was 49.49% for FVG and 47.49% for IFVG, among resolved bracket outcomes. This is a different question from winning at a 2R target after costs.

The original output contains 95% Wilson intervals. We retain them for transparency but do not use them to claim statistical significance here: overlapping events and repeated market histories are not independent trials. Dependence-aware inference needs an appropriate design; ordinary interval overlap also does not establish equivalence. Zhang and Cheng’s research on dependent time-series inference explains why dependence must be addressed, but its method has not been implemented in this study.

What the simulation can still get wrong

Intrabar order matters. Entries and exits are evaluated in lower-timeframe bars, each following an assumed OHLC path. The chart/intrabar mapping is 5m→1m, 15m→3m, 1h→15m. We did not resolve every trade with one-minute data or market ticks. The path is based on the OHLC convention described in TradingView’s broker-emulator documentation; the study itself runs in Python.

A touched limit is assumed filled. There is no order queue, partial fill or bid/ask reconstruction. An added cost allowance cannot demonstrate that an order would have filled. Six retained chart bars have an intrabar range mismatch. The final exported bar may be incomplete, and ATR is seeded on the retained history. These are reasons to investigate sensitivity further, not to describe the output as exact execution.

The validation is bounded. The engine’s synthetic random-walk checks compared modeled outcomes with known tick paths, with agreement between 99.29% and 100% for those checks. That validates particular code paths; it does not establish that much accuracy on real market data.

The most attractive cell is selected after inspection. NQ H1 had the highest mean net result here. We cannot infer that it is a repeatable trading edge—or prove that luck caused it—from this table alone. The selection experiment in our Expert Advisor guide illustrates the same selection problem in a simplified probability model. Bailey and co-authors’ backtest-overfitting paper addresses the broader selection problem. We did not compute their overfitting estimator.

Download, inspect and reproduce

Reader resource · ZIP

Inspect the FVG study and reproduce the calculation

Python engine, Pine exporter, original protocol, aggregate results, file hashes and publication review. Raw market exports are not included.

Download the code and results

The archive contains the original Python v1.00 engine, Pine data exporter, dated local protocol, aggregate JSON results, manifests and the 1 October publication review. The review qualifies earlier shorthand such as “coin flip”, the interpretation of confidence intervals and the use of “out of sample”. The original experiment files are preserved so the qualification is visible.

The local protocol records that rules were fixed before the first run. We have not established an independently timestamped public preregistration. The raw TradingView market exports are not redistributed: you need permitted data in the documented format to rerun the market calculation. You can run the synthetic checks without market data:

python3 analyze.py --verify
python3 analyze.py datasets.json --out my-results.json

For exact replication, compare the raw-file hashes, settings and retained time ranges. A new export can contain a different history or revised data. The code and original documentation were developed with AI assistance; this article reports an implementation review and repeated calculation, not independent peer review.

Common questions

Does this show that FVG trading never works?

No. It shows the outcomes of the stated mechanical rules, samples and cost assumptions. Additional filters or discretionary selection require their own defined test. They cannot borrow a profitable result from a different setup or dismiss the losing baseline without showing what changes.

Is a high gap-fill rate enough?

No. Returning to a zone is an entry or path event. Profit depends on what happens after the chosen entry, the exits and costs. This study evaluates specified retest trades; it does not report a universal percentage of all gaps eventually filled.

What would improve the evidence?

Independent data, confirmed contract handling, execution sensitivity, dependence-aware statistics and a prospectively frozen test would strengthen it. A portfolio version would also need explicit capital and exposure rules. The next useful question is whether a precisely defined improvement survives those checks, not whether one attractive chart is persuasive.

Friendly robot illustration representing the RoboXpert pen name.
AI-generated avatar

About the author

RoboXpert

Pen name

The person behind RoboXpert writes about expert advisors, trading evidence and programming, and develops their own trading software. They report 10 years of experience in these areas; this is self-reported, not an independently verified qualification or performance record.

Sources & further reading

Your next question

Continue reading

Testing & evidence

Trading Cost Sensitivity: How Much Cost Can a Strategy Survive?

Break-even trading costs for a backtest: in our 23,899-event sample, the mean reached zero at 7.7% of the assumed costs. Method, 12 datasets and code.

Read the guide →
Testing & evidence

100% a Month With a Trading Indicator? A Reality Check

Understand why a recognizable chart pattern is not enough to establish usable trading returns.

Read the guide →