XAUUSD 4,151Live gold map →
Academy · Chapter 12 of 12

Honest Research: Testing Ideas Without Fooling Yourself

≈25 min read · Real gold futures examples with XAUUSD equivalents · Free

In this chapter: You will learn to judge any claim that "this strategy works", including your own claims and ours, using the standards of honest research. We cover the classic ways a backtest lies (look-ahead bias, survivorship and data bias), why choosing the best of many tests produces brilliant-looking results from pure luck, and how the Deflated Sharpe Ratio and the Probability of Backtest Overfitting correct for it. We then look at the silent killers of order-flow strategies (impossible fills, slippage, commissions and fake volume), lay out a complete testing protocol (pre-registration, a time-based split with a locked out-of-sample period, walk-forward analysis and realistic costs), and explain forward testing as the first truly honest evidence. Finally, we describe in general terms what failed in our own testing and why we tell you, and we give you a checklist for evaluating any trading educator, including us.

Why Most Backtests Lie: Look-Ahead & Survivorship

A backtest is the process of running a strategy's rules over historical data to see what would have happened "if we had traded this way back then". It is the most common tool in trading research and one of the easiest to get wrong. Small mistakes in how a backtest is built can produce a result that looks excellent and is completely unreal.

The worst thing about these mistakes is that they produce no error message. The code runs. The equity curve rises smoothly. Nothing on the screen tells you that the result is impossible.

Look-aheadusing data you would not have had at the timeOverfittingthe best of 100 settings fits the pastCosts & fillsspread, commission, slippage, impossible fillssolid: what the backtest shows · dashed: what you actually get
Figure 1. Three ways a backtest flatters an idea. Each one makes the past look better than the future will be.

Look-ahead bias: reading tomorrow's newspaper

Look-ahead bias means using information in a decision that was not yet available at the moment the decision was made. Imagine betting on yesterday's horse races while holding today's newspaper: you would look like a genius, and you would learn nothing about betting.

In order-flow and volume work, look-ahead hides in many places:

The defence is procedural. Every number a strategy uses should carry a timestamp of when it became knowable, not when it happened. Profiles, VWAP, highs and lows must be computed "up to that moment". Detections must be recorded at the moment they could really have been detected. Normalisation must use only past data. And it helps a great deal to have someone else review the code with these questions in mind.

A gold illustration

Illustrative, not a real result. Imagine two versions of the same simple rule on GC: "buy when price touches the day's VAL". Version A uses the completed day's VAL. Version B uses yesterday's VAL, or the developing VAL as it stood at that moment. The difference between them is one line of code. Version A will very often look far better, because the completed VAL is, by construction, a price where the day's trading later proved to be important. Version A has quietly looked at the answer.

Survivorship and data bias

Survivorship bias means your data is an unrealistic sample of the past because something has been removed. The classic stock-market example is testing on today's index members only: the companies that went bankrupt and dropped out are missing, so the past looks better than it was.

Gold futures do not have bankrupt companies, but they have their own versions of the problem:

Key idea: A backtest can be wrong without any visible error. Only process saves you: timestamps of when each piece of information was knowable, correct contract handling, and independent review.

Common mistake: "My backtest covers five years of data, so it's reliable." The length of the data does not cure look-ahead. Five years of tomorrow's newspaper is still tomorrow's newspaper.

A related claim, "my indicator doesn't repaint", is a popular claim, not proven, unless you have checked how the indicator calculates each value. An indicator "repaints" if its past values change after new data arrives, which is a form of look-ahead built into the chart itself.

Try it: Take one rule you believe in. Write down every input it uses. Next to each input, write the exact moment that input becomes knowable in real time. If any input becomes knowable after the decision moment, your test of that rule has look-ahead.

Overfitting & Multiple Testing: The "Best of 100" Trap

Overfitting: memorising noise

Overfitting means a strategy has been tuned so closely to the details of past data that it has memorised the noise rather than learned a pattern. It fits the past perfectly and the future poorly.

Order-flow strategies are especially exposed, because the space of choices is enormous: the delta threshold, the size that counts as a "big trade", the time window, the type of level, the time of day, the speed filter, the imbalance ratio, the bar type. Every combination of choices is, in effect, a separate experiment.

The multiple-testing problem

Here is the core problem. If you test 100 strategies that are completely random, with no edge at all, some of them will produce very good results by chance. If you report only the best one, you look like a genius. And that best one will almost always perform much worse in the future, because its "edge" was luck.

A coin example makes it concrete. Flip a fair coin ten times. The chance of getting nine or more heads is 11 in 1,024, about 1.1%. Now give 100 people a coin each and have them all flip ten times. The chance that at least one of them gets nine or more heads is about 1 − (1 − 0.011)^100, roughly two in three. The "winning" coin gets a crown, and on its next ten flips it is a fair coin again.

The same thing happens with Sharpe ratios. The Sharpe ratio is average return divided by the variability of returns, usually annualised: a measure of return per unit of risk. With two years of daily results, a strategy with zero true edge has a measured annual Sharpe that wanders around zero with a typical spread of about 0.7. The best of 100 such strategies will usually show a Sharpe around 1.7 or higher, purely from noise. On a chart, that looks like a very attractive strategy.

The key discipline is to count every experiment, including the ones you "just tried quickly" and the parameter tweaks you made in your head. The number of trials is part of the result.

Tools that correct for it

Deflated Sharpe Ratio (DSR). Developed by David Bailey and Marcos López de Prado. It compares the Sharpe ratio you observed with the highest Sharpe ratio you would expect to see by luck alone, given how many trials you ran, how spread out their results were, how skewed and fat-tailed the returns are, and how long the test period was. Its output is, roughly, the probability that the true Sharpe is above zero after all these corrections. A strategy that looks impressive after one trial can look ordinary after 300.

Probability of Backtest Overfitting (PBO). Developed by Bailey, Borwein, López de Prado and Zhu. The data is split into many pieces, which are recombined into many in-sample and out-of-sample pairs. For each pair, the procedure asks: "Did the configuration that was best in-sample end up below the median out-of-sample?" The fraction of pairs where the answer is yes is an estimate of how likely your selection process is to pick an overfit winner.

Statistical multiple-testing corrections. Methods such as Bonferroni and Holm, or controlling the false discovery rate with Benjamini–Hochberg, make the significance threshold stricter as the number of tests grows. In finance, Campbell Harvey, Yan Liu and Heqing Zhu argued that, given how many factors researchers have tried, a new "discovery" should need a t-statistic above about 3.0 rather than the traditional 2.0.

None of these tools is magic. Every one of them depends on an honest count of how many things were tried. If you tried 300 combinations and tell the formula you tried 3, the formula will happily give you a wrong answer.

Key idea: Flip 100 coins ten times each and one of them will look like a genius. That is what most backtests are: the luckiest coin, presented as skill.

Common mistake: "I only have three parameters, so I can't be overfitting." What matters is how many versions you tried, not just how many parameters the final version has. Three parameters with ten values each is 1,000 versions.

"A backtest Sharpe above 2 means the strategy is real" is a popular claim, not proven. Without knowing how many trials produced it, a Sharpe number on its own tells you very little.

Try it: Look back at the last idea you tested. Count honestly how many variations you tried: different thresholds, time windows, filters, markets, bar types. Write the number at the top of your notes. That number belongs in every report of the result.

Impossible Fills, Costs & Fake Volume

Order-flow strategies are usually short-term: small objectives measured in a few ticks, and many trades. In this world, costs and the quality of fills decide almost everything. A small error in how fills are modelled can be larger than the entire edge being measured.

1. Impossible fills

"My limit order filled when price touched it." On CME, orders at the same price are generally filled in time priority: first in, first out (FIFO). A touch is not a fill (Chapter 10). If price touches your limit price and then turns away, the orders ahead of you in the queue may have absorbed all the trading at that price. A conservative backtest only fills a limit order when price trades through it by at least one tick, or uses a model of queue position.

The unknown path inside the bar. A bar has an open, high, low and close, but many different paths can produce the same four numbers (Chapter 1). If both your stop and your objective sit inside one bar's range, the bar does not tell you which was hit first. A backtest that assumes the favourable one has invented a fill.

More volume than really traded. A backtest that fills 50 contracts at a price where only 12 traded has filled orders that the market could not have absorbed.

Zero latency. A signal that depends on a bar closing, or on data processing, is only visible after that happens. Your real reaction time, and your order's travel time, are never zero. A backtest that enters at the exact price of the signal moment assumes instant reaction.

2. Slippage

Market orders and stop-market orders fill worse than the expected price in fast markets, especially around news (Chapter 11). A conservative assumption is at least one tick of slippage per side for market and stop orders, and more around major releases.

3. Fixed costs

Commission, exchange and clearing fees, regulatory fees, and data and platform costs. Express them in ticks, because that makes them comparable with the size of your objective.

Illustrative arithmetic. Check your own broker's costs.

Suppose all-in commission and fees for a round trip on one MGC are about $2. On MGC a tick is $1, so that is 2 ticks per round trip. Add one tick of slippage on exits by stop.

Now take a simple strategy with a 6-tick objective and a 6-tick stop:

Nothing about the market changed. Only the accounting did. On GC, where a tick is worth $10, the same dollar commission is a much smaller fraction of a tick, which is one reason why costs per tick differ so much between the two contracts even though the price is the same.

Think of it as a waterfall: gross result, minus commission and fees, minus slippage, minus the fills you assumed but would not have received, equals the net result. For many short-term ideas, the waterfall drains most of the gross.

4. Tick volume versus real volume

In decentralised markets, such as spot forex or gold CFDs (XAUUSD), the "volume" on most charts is tick volume: the number of price changes in your broker's feed, not the number of contracts traded. A trade of 1,000 units and a trade of 1 unit each count as one tick, if they count at all. A backtest of "delta" or "footprint" patterns on such data does not measure what those words mean in futures (Chapter 2). That is why this book reads GC and MGC data: a central market with real volume and an aggressor side for every trade.

The recurring pattern

In published research and in our own testing, one pattern repeats: many order-flow signals show "something" before costs and then disappear after costs and realistic fills for a non-high-frequency trader. The edges that exist at the scale of milliseconds tend to belong to the fastest participants, whose costs and latencies are far below a retail trader's.

Key idea: For short-term strategies, costs are destiny. Always write the cost in ticks next to the objective in ticks.

Common mistake: "Commission is negligible." On MGC, with objectives of a few ticks, it is not.

"My platform runs backtests realistically" is a popular claim, not proven, unless you have checked its fill assumptions. And "a footprint on a gold CFD is the same as a footprint on gold futures" is simply false: the CFD footprint is built from estimated data from one broker's feed.

Try it: For any short-term idea you have, write three numbers: objective in ticks, stop in ticks, and round-trip cost plus expected slippage in ticks. Compute the break-even hit rate before and after costs. If the gap shocks you, that is the lesson.

How to Test an Order-Flow Idea Properly

Here is a testing protocol suitable for teaching and close to what we use in our own research. It is not the only valid protocol, but every step addresses one of the traps of the previous sections.

Step 1 — Pre-registration

Before you see any results, write a short document and date it. It contains:

Pre-registration stops you from moving the goalposts: deciding what counts as success after you have seen what you got.

Step 2 — Split the data by time

The split is always chronological (past → future), never random. Randomly assigning days to train and test lets information leak between neighbouring days, because markets have memory over hours and days. It is also good practice to leave a small gap, an embargo, between periods so that overlapping effects (for example, a trade that opens at the end of one period and closes in the next) cannot leak.

PeriodPurposeHow often you look
TrainBuild and tuneAs often as needed
EmbargoGap against leakageNever used
ValidationChoose among a few versionsA limited number of times
EmbargoGap against leakageNever used
Locked OOSFinal, honest testOnce

Step 3 — Walk-forward

Instead of one split, move a window forward through time: tune on period 1, test on period 2; tune on periods 1 + 2, test on period 3; and so on. Then put all the test periods side by side. Walk-forward shows whether the idea is stable across different market regimes: quiet and volatile, trending and balanced, different rate environments.

Step 4 — Execution realism from day one

Fills based on the queue, or on a conservative assumption (filled only when traded through). Slippage. Commission and fees. Latency. These are not refinements to add at the end; they belong in the first version, because an idea that only works without them is not an idea.

Step 5 — Correct statistics

Record the total number of trials. Use the Deflated Sharpe Ratio, PBO or a multiple-testing correction. Report the distribution of R, the maximum drawdown and the number of trades, not just the share of winning trades. A high share of winners with occasional large losers can still have negative expectancy.

Step 6 — A negative result is a result

If the idea fails, document it, and if you publish research, publish the failure too (see "What Failed in Our Testing" below). A literature that contains only successes is itself biased.

Key idea: Lock part of your data in a vault and open it once. If you peek and adjust, the vault is empty.

Common mistake: "Just one small look" at the out-of-sample results, followed by "just one small change". After that, the OOS is contaminated; it is effectively in-sample.

"A good walk-forward result guarantees the future" is a popular claim, not proven. Walk-forward provides stronger evidence than a single split; it does not provide certainty.

Try it: Write a pre-registration form for one idea using the five fields above: Hypothesis, Rules, Data, Success criteria, Maximum variants. Date it. Do not run anything until it is complete.

Forward Testing: The First Honest Evidence

Even the most careful backtest was run on data that the researcher knew something about, directly or indirectly, while building the idea. You may not have looked at the locked out-of-sample period, but you lived through those months; you remember how gold behaved. A forward test runs frozen rules on data that did not exist yet when the idea was built. It is the first evidence that is genuinely about the future.

Two forms

FormHow it worksStrengthWeakness
Paper / simulatedOrders go to a simulator fed by live dataCheap; no money at riskFills are optimistic: a simulated order is not really in the queue
Small liveReal orders at minimum size (for example, one MGC)Real fills, real slippage, real psychologyCosts real money, even if small

Paper testing is a useful first step; small live testing reveals what the simulator hides.

A protocol for forward testing

Honest limits

A forward test is also a limited sample, and a good period can be luck. But the combination of a strict backtest and a frozen forward test is far stronger evidence than either alone. The backtest says "this could work"; the forward test is the first time the market itself gets a vote.

Key idea: A backtest is a story about the past. A frozen forward test is the first time the market gets a vote.

Common mistake: Changing the rules after three bad trades in the forward test. That resets the test; the new rules have zero forward evidence.

"Two profitable weeks on paper means you are ready for a large account" is a popular claim, not proven. Two weeks is a tiny sample, and paper fills are optimistic.

Try it: For an idea you are working on, write an "expected band" card: average R range, trades per week range, worst expected drawdown in R. Then write the stop-and-review criteria. Pin both to your screen before the forward test begins.

What Failed in Our Testing (And Why We're Telling You)

This section is framed explicitly as "in our testing": results on our own data, with our own methods. They are not universal truths. They were obtained with a protocol like the one described above (pre-registered hypotheses, locked out-of-sample periods, realistic costs), and we describe them in general terms.

Example 1 — The side of an iceberg as a direction signal

The idea is attractive: "If a buy-side iceberg is detected, price will rise in the following minutes." In our testing, the side of a detected iceberg did not predict the direction of the next move in any useful way when run live. The likely reasons were covered in Chapter 10. First, an iceberg can only be detected after it has refilled several times, so by the time it is labelled, part of its story is over. Second, synthetic icebergs, built by a trading platform rather than the exchange, look just like ordinary orders and create ambiguity. Third, and most fundamentally, an iceberg is evidence of hidden size at a price, not evidence of its owner's overall direction. A large passive buyer at one price might be hedging, executing one leg of a spread, or completing an order that has nothing to do with the next five minutes.

Example 2 — Naive big-trade detection

The idea: "Large prints mean large players." In our testing, a naive detector that flagged large prints directly from the raw trade feed was wrong much more often than it was right: a large share of its flags were not single large orders at all. The reason is the one explained in Chapter 10: one real order often prints as many small fills, so the raw sequence of prints is a poor picture of the orders behind it. Grouping prints into the orders behind them improved the quality of detection considerably. But even when large orders were identified correctly, "a big trade happened" was not, on its own, a profitable signal after costs in our tests.

The broader pattern

Across the published literature and our own work, the same pattern appears again and again: many order-flow signals that show "something" before costs vanish after costs and realistic fills for a trader who is not operating at high frequency. Edges measured in seconds or less tend to belong to faster competitors.

The common cause in both examples is worth stating plainly: what we see on the screen is not the same as the real event behind it. Prints are not orders. A visible iceberg is not an intention. The data is a shadow of the decisions that created it.

So what is left?

Order flow as confirmation, timing and veto at levels chosen in advance, as described in Chapter 11. The veto role, declining trades that the flow contradicts, is especially valuable, because it reduces the number of trades and therefore the costs. This, too, is a hypothesis, and anyone using it should test it on their own data with the protocol of this chapter.

Why publish failures?

Because publishing only successes is survivorship bias at the level of content. If every educator shows only what worked, the overall picture the public sees is a highlight reel of the luckiest coins. An educator's credibility comes, in large part, from telling you what did not work.

Key idea: The right conclusion is not "order flow is useless". It is that order flow changes role: from signal generator to filter in context.

Common mistake: "These ideas failed for you because you didn't use them correctly." That is possible. That is exactly why we describe the method and its limits openly and encourage independent testing.

"Commercial tool X has solved this problem" is a popular claim, not proven, unless it is backed by an independent test with genuine out-of-sample evidence.

Try it: Pick one order-flow idea you find attractive. Write down the hidden assumption behind it (for example, "a large print is one large order" or "an iceberg's side is its owner's direction"). Then write how you would test whether that assumption is true before testing the trading idea itself.

How to Evaluate Any Trading Educator (Including Us)

The market for trading education is full of claims that cannot be checked. The goal of this section is not blind cynicism. It is to give you a standard set of questions, and to ask you to apply them to this book as well.

Official red flags

The US Commodity Futures Trading Commission (CFTC) publishes warnings about trading education and advice. Among the red flags they describe:

Simulated results have built-in limits

In the United States, CFTC Rule 4.41 requires that advisers it covers who present hypothetical or simulated performance must accompany it with a specific warning. In substance, the warning says that simulated results have inherent limitations, do not represent actual trading, and are designed with the benefit of hindsight. Whenever you see a "backtest" or "system results" in marketing, remember that warning, whether or not it is printed.

The question checklist

  1. Is the claim falsifiable: specific and precisely defined, so that it could be shown to be wrong?
  2. Are the results real and verifiable, or selected screenshots?
  3. How many ideas or parameter sets were tried to arrive at this result?
  4. Is there out-of-sample or forward evidence?
  5. Are costs and slippage included?
  6. Are failures and limits discussed?
  7. Where does the educator's income come from: courses, tools, broker or prop-firm commissions, or trading? Is there a conflict of interest?
  8. Are terms defined, or are they attractive labels ("smart money", "institutional") without a definition?
  9. Do the examples include ordinary and bad days, or only spectacular ones?
  10. Does the educator encourage you to test things yourself, or only to trust them?

A comparison exercise

Imagine two anonymous posts about gold (fictional, no real person or brand). The first says: "This GC order-flow setup wins 90% of the time!!" and shows one selected screenshot. The second gives a precise definition of the pattern, the test period, the number of variations tried, the cost assumptions, and one example where it failed. Score both with the checklist. The first scores close to zero; the second may not be right, but it gives you everything you need to find out.

About us

This book makes no profit or win-rate claims. It labels folklore as folklore ("a popular claim, not proven"). It describes what failed in our testing and explains the methods behind our conclusions. If you find a place where we break these principles, use this same checklist to criticise us.

Three myths to finish. "He drives an expensive car, so his strategy works": the income may come from selling courses. "A P&L screenshot is proof": screenshots can be selected, faked, and are not representative. "X years of experience means the method is right" is a popular claim, not proven: experience and evidence are different things.

Key idea: A good claim is specific, falsifiable and verifiable. Ask how many things were tried, whether there is out-of-sample evidence, and whether costs were included.

Common mistake: Treating confidence as evidence. The most certain-sounding educators are often the ones with the least verifiable evidence.

Try it: Score the next three trading posts you see on social media with the ten-question checklist. Then score this book. Write down where each one loses points.

Chapter summary

Checklist

Quiz

  1. At 10:00 ET you use today's final POC as a level in your backtest. What kind of error is this?
  2. You tried 300 parameter combinations and report the best one. What else must you report, and why?
  3. In spot forex or gold CFD charts, what does "volume" usually count?
  4. You peeked at your out-of-sample results and changed one rule. What is the status of your out-of-sample period now?
  5. In our testing, why did naive big-trade detection from raw prints perform poorly?

Quiz answers

  1. Look-ahead bias. Today's final POC only exists at the end of the session; at 10:00 ET the profile was still developing.
  2. The number of combinations tried (and ideally the spread of their results). Without it, nobody can judge how much of the best result is luck; tools like the Deflated Sharpe Ratio need that number.
  3. The number of price changes (ticks) in that broker's feed, not the number of contracts traded.
  4. It is contaminated. Because you adjusted the idea after seeing it, it has effectively become in-sample data, and it can no longer serve as an honest test.
  5. Because one real order often prints as many small fills, so individual prints are a poor picture of the orders behind them, and many of the flags did not correspond to single large orders. Even correctly identified big trades were not, on their own, a profitable signal after costs.
Prefer the PDF? This chapter is part of the free 301-page eBook Gold Order Flow: From Zero to Volume Trading.
Get the free PDFGlossary

Education only. Not financial advice. Trading involves substantial risk of loss.