How to use negative controls in a backtest
A negative control is a test designed to have no credible connection to your thesis while sharing the same data pipeline and much of the same backtest machinery. If the control produces the result you expected only from the real signal, something else may be doing the work. The culprit could be timing leakage, a broad market exposure, the universe definition, or the testing process itself.
Use negative controls beside ordinary robustness checks. Build one that breaks the proposed mechanism, one that challenges timing, and one that preserves the screen's broad exposures while removing its stock-selection claim. A failed control doesn't tell you which defect exists. It tells you the attractive curve hasn't earned its story yet.
What is a negative control?
The idea comes from observational research, where investigators can't always assign subjects randomly. Lipsitch, Tchetgen Tchetgen, and Cohen define negative controls as exposures or outcomes that can't plausibly cause the effect under study but share potential sources of bias with the main analysis. An association in the control warns that the main result may be spurious (their paper on negative controls).
A backtest has the same need. Historical market data arrive with time structure, missing records, changing universes, and common market forces. A screen then adds thresholds and portfolio rules. A conventional stress test asks whether the result survives a nearby assumption. A negative control asks a different question: does an intentionally irrelevant or impossible signal appear to work too?
Suppose your thesis says improving cash conversion identifies companies whose reported earnings are becoming more dependable. A useful negative control might shift the accounting signal to a period after the simulated decision date. That future value cannot have informed the trade. If it creates a large advantage, the result doesn't validate cash conversion. It exposes a timing path that can see information too early.
The control should share the machinery you want to test. Swapping the entire strategy for a coin flip won't diagnose a filing-date bug because the coin flip never touches a filing. Good controls preserve the suspect pipeline and remove the claimed economic mechanism.
Which defects can a control expose?
Negative controls are especially good at finding results that survived because every ordinary variant inherited the same mistake.
Timing leakage. A filing-driven thesis should act only after a filing became public. Move the same fundamental value backward so the simulated rule receives it before publication. That deliberately impossible version ought to look better than the valid version. If both versions behave identically, inspect whether the backtest is aligning on fiscal period ends or whether the reporting lag has any effect. Reporting lag is a testable assumption, and a timing control checks whether the assumption reaches the trades at all.
Broad exposures. A quality screen may beat a broad index because it favors a sector, company size, or valuation style that happened to lead. Construct a control that preserves those broad weights while randomizing selection within each bucket. If the control keeps most of the excess result, the named accounting rule contributed less than the exposure. Testing whether a screen is a sector bet is one specific version of this logic.
Universe artifacts. Apply a meaningless feature to the exact same eligible universe. A result that persists can point to an investable-universe effect, a delisting problem, or thin historical coverage. The feature may deserve no credit. The universe isn't neutral. It can carry a return pattern before the score awards a single point.
Implementation artifacts. Reverse a tie-breaker or replace a ranked selection with a within-bucket shuffle while keeping rebalance and exit rules fixed. If the result barely changes, the portfolio construction may dominate the ranking claim.
One control can flag several possibilities. It won't name the guilty line of code. Treat the failure as a reason to narrow the investigation, then audit dates, eligibility, and trades.
Start by writing the mechanism
You can't break a mechanism you haven't stated. Write the thesis as a short chain:
A deterioration in cash conversion appears in public filings, the screen observes it after publication, and subsequent returns differ because the deterioration reveals weak earnings quality.
Each link suggests a control.
- Public filing: use a field from a later filing to test for leakage.
- After publication: vary the lag and confirm that trades move with it.
- Cash conversion: replace it with a sham feature that shares coverage but has no economic role in the claim.
- Weak earnings quality: preserve industry and size exposure while scrambling the company ranking.
This exercise also exposes vague theses. If you can't say what should stop working when the signal is broken, the backtest has no distinctive claim to validate. Go back to the discipline of turning a hunch into measurable proxies before adding more tests.
Choose the control before reading its curve. A control invented after a disappointment can become another search for a favorable interpretation. Record what outcome would worry you and what follow-up you would run. The point is a diagnostic with a predeclared meaning.
How do you build a timing control?
Timing controls should be deliberately wrong and clearly labeled. Don't save an impossible thesis where someone could mistake it for a usable screen.
Begin with the valid specification. Fundamentals become visible on or after their filing date, the reporting lag passes, and the simulated entry occurs under the stated rule. Then create a research-only comparison that violates one timing boundary, such as exposing the later filing value at the earlier fiscal period end.
The expected ordering is simple. The impossible version has more information and may produce a stronger hypothetical result. The valid version should weaken once that information advantage is removed. If the impossible and valid versions produce the same decisions, check whether the feature changed between those dates and whether any qualifying event fell inside the gap. Identical output can be legitimate when there was nothing new to reveal.
Now add a plausible-delay control. Increase the reporting lag without changing the thesis, universe, benchmark, rebalance cadence, or exit logic. The shape below is illustrative, and the bundled templates plus the Quantery DSL documentation are the reference for exact fields.
# Valid baseline: illustrative
backtest:
rebalance: monthly
reporting_lag_days: 1
entry: next_close
exit: { max_hold_months: 12, on_gate_fail: true }
benchmarks: [SPY]
The proposed settings are test assumptions. The useful comparison changes only the lag. If a modest delay destroys the direction, the result may depend on entering before a real research process could digest the filing. If nothing changes, verify that the rebalance clock isn't masking the lag difference. A monthly evaluation can absorb a small timing change because both signals wait for the same evaluation date.
Every return remains hypothetical and excludes trading costs. The purpose of this pair is to inspect information timing, not to estimate an executable return down to the last basis point.
How do you control for sector and size?
A good exposure control keeps the parts of the selection problem that could create a disguised beta while removing the named stock-picking rule.
Take every company eligible on a simulated date and place it in coarse buckets defined before the run, such as sector and size range. Record how many names the real thesis selects from each bucket. Then form control portfolios by selecting different eligible names from the same buckets. Keep the rebalance dates, holding rules, and benchmark fixed.
This answers a concrete question: could a portfolio with the same broad shape have produced a similar result without the thesis score? If yes, the result belongs first to the exposure. The accounting rule still may improve selection within that exposure, but the original broad-index comparison overstated the evidence.
Don't choose buckets so narrowly that each contains the original company and no substitutes. Don't choose them so broadly that a tiny manufacturer and a global software company become interchangeable. The buckets should capture the obvious exposure you are trying to hold constant.
Run several shuffles because one random draw is an anecdote. Summarize the distribution of control outcomes, including the inconvenient draws. This isn't a new hunt for the best control. It is a check on how unusual the real selection was among comparable alternatives.
Choosing a matched backtest benchmark handles the investable baseline. The within-bucket control goes further by preserving the actual screen's changing exposure at each decision date. Use both when the economic claim is stock selection inside a style or sector.
What does a failed control mean?
A failed negative control identifies a problem class. It doesn't settle a conviction. A future-information signal that "works" may reveal leakage, but it also may show that the future filing genuinely contains predictive information. The impossible availability date is what makes that result a pipeline warning. Inspect when the values entered the simulation.
A shuffled portfolio matching the real result may reveal broad exposure. It may also occur by chance in one draw. Repeat the control according to the plan and compare the real result with the full set. Then inspect holdings and dates. Controls point. They don't finish the audit.
A passing control deserves restraint too. Lipsitch and coauthors warn that a null negative-control result can't prove the main analysis is free of bias. The control may fail to share the relevant defect. A sector-matched shuffle won't detect a filing-timing leak, and a timing test won't reveal that one industry carried the portfolio.
That is why controls come in a small, purposeful set. Each one should target a different failure mode. Adding dozens after seeing results recreates the multiple-testing problem the controls were meant to discipline.
Keep controls beside the thesis
For each control, record the mechanism it breaks, the defect it targets, the settings held fixed, and the result that would trigger investigation. Keep the associated run identifier and thesis version. Quantery persists versions and reports, so the control can remain attached to the exact baseline it challenged.
Read a negative control after the basic data checks, then beside the ordinary perturbations. Point-in-time inputs and survivorship-aware universes remain prerequisites, as the backtest integrity checklist explains. Nearby thresholds and alternate windows test fragility. Negative controls test whether the backtest can manufacture your story when the story has been removed.
That last question is severe on purpose. A plausible mechanism can make any attractive result feel explained after the fact. Break the mechanism. Keep the machinery. If the curve survives, investigate what the machinery is rewarding before you give the thesis credit. The rules are yours, and the placebo rules should be yours too.
Research tooling, not investment advice. Nothing here is a recommendation to buy, sell, or hold any security. Screens, scores, and backtests are informational only; backtested results are hypothetical, exclude costs such as commissions and slippage, and do not guarantee future results. Verify against primary filings and make your own decisions.
Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.
← All articles