How to use a holdout period in a backtest
A holdout period is a block of historical data you refuse to inspect while building a stock screen. You develop the rules on earlier dates, freeze the thesis, then run it once on the untouched period. The holdout asks whether the result travels beyond the data that helped create it.
That sounds easy. The hard part is preserving the blindness. If you check the holdout after every edit, dislike what you see, and edit again, it has joined the development sample. You can still learn from it, but you can't keep calling it an untouched test. A useful holdout is a process rule before it's a date range.
What problem does a holdout period solve?
Backtest overfitting happens when rules learn the accidents of a historical sample. It doesn't require fancy machine learning. A person moving a valuation cutoff, dropping an awkward criterion, and shifting the start date is fitting the sample by hand.
The danger grows with the number of alternatives tried. Bailey, Borwein, López de Prado, and Zhu show that strong simulated performance can be found after trying relatively few strategy configurations. Harvey, Liu, and Zhu make the same multiple-testing point across published return factors: after researchers have tested many candidates, the familiar significance threshold is too lenient.
A holdout separates two jobs that otherwise blur together:
- Development period: choose proxies, thresholds, gate structure, universe, benchmark, and rebalance cadence.
- Holdout period: evaluate the frozen choices on observations that didn't influence them.
The split won't make a weak thesis strong. It makes disappointment informative. Failure on untouched dates is evidence that the development result didn't travel. Without the split, another edit can always make the old chart look better.
This is stricter than an ordinary subperiod check. A subperiod checked after you already know its result is a robustness view. It may be useful, but it isn't independent evidence. The difference is whether the period was hidden from the decision process.
How should you choose the holdout dates?
Choose the boundary before tuning. Put the most recent completed block in the holdout and use the earlier history for development. Recent dates are useful because they resemble the conditions under which the finished rules would next operate, and because their results are harder to rationalize as ancient market plumbing.
Make the block long enough to contain several rebalance cycles and enough qualifying events to read. There is no universal number of months or events. A broad, frequently rebalanced screen generates evidence faster than a narrow deep-value thesis. Decide the minimum evidence before opening the result. If the holdout produces too few observations, conclude that the evidence is insufficient. Don't move the boundary until the sample cooperates.
Respect the information clock at both ends. A filing belongs to the test only after it became public and after your chosen reporting lag. A position opened during development can continue into the holdout, which muddies the question. The cleaner design begins measuring new decisions after the boundary while keeping entry and exit rules fixed.
Don't choose the split because it puts a crash, rally, or awkward year on the convenient side. That makes the boundary another tuned parameter. If you want to ask whether the thesis survived a named regime, run that as a separate stress test after preserving the original holdout.
Data coverage can also dictate a sensible boundary. Quantery's documentation explains that backtests use point-in-time EDGAR filings, and price depth depends on the user's market-data plan. Read the coverage note before declaring a long development period. A decade printed on the date selector isn't useful if the early eligible universe is a thin reconstruction. Check data coverage before the backtest, then record the actual span you can defend.
What must be frozen before the test?
Freeze every choice that can change membership or measured performance. That includes the feature definitions, thresholds, gate mode, required fields, universe, sector exclusions, benchmark, rebalance cadence, reporting lag, entry timing, and exit logic. Save the exact thesis version and write down the test dates.
Also freeze how you'll judge the result. Pick the primary comparison in advance. It might be median excess return at one event-study horizon, or a hypothetical portfolio's return relative to its benchmark with maximum drawdown considered alongside it. Backtest returns are hypothetical and exclude costs. Don't wait to see which panel looks best before deciding which panel mattered.
A short research note is enough:
# Holdout plan: illustrative
development:
from: 2013-01-01
to: 2023-01-01
holdout:
from: 2023-01-01
to: 2026-01-01
frozen:
thesis_version: quality-value-v7
rebalance: monthly
benchmark: SPY
reporting_lag_days: 1
primary_read:
event_horizon_months: 12
statistic: median_excess_return
failure_rule:
direction: non_positive
minimum_evidence: declared_before_run
The shape is illustrative. The bundled templates are the reference for exact thesis fields, and the Quantery thesis documentation covers the DSL. The dates and settings above are proposed research parameters. No actual run produced them.
Keep a search log beside the plan. List the thesis versions you tried during development and what changed in each. The winning version means less if it emerged from a large, unrecorded search. The log also stops you from remembering a tidy story that wasn't how the rules were built.
How do you run the test without contaminating it?
Finish development first. Use nearby thresholds, alternate windows, criterion ablations, and a score ladder inside the development period. Those checks tell you whether the thesis is stable enough to spend the holdout.
Then save the final version. Run the development dates one last time and preserve that report. Without changing the version, run the holdout dates with the same benchmark and mechanics. Read coverage and event counts before returns. A direction based on sparse events is weak evidence even if the line is attractive.
Now compare the claims without demanding matching figures. If development showed positive median excess return at the declared horizon, did the holdout retain that direction? Did the result come from a broad set of events, or from one cluster? Did drawdown change enough to make the implementation story different? Exact return equality would be surprising. Direction and economic shape matter more than cosmetic similarity.
Don't repair a failed holdout in place. If you loosen a threshold and rerun the same dates, those dates have taught you something. Move the revised thesis back into research status. The old holdout is now development evidence.
This one-use rule is why a holdout should come late. Spend it after the thesis has survived cheaper challenges. Perturb one dial at a time, inspect whether the screen is a hidden sector bet, and use negative controls to see whether broken mechanisms still seem to work. A holdout shouldn't be the first time you notice a sign error or duplicate signal.
What if the holdout fails?
First, don't call failure useless. It answered the question the design was built to ask. The frozen development conclusion didn't repeat on untouched dates.
Audit mechanics before economics. Confirm that the thesis version, universe, benchmark, timing, and coverage match the plan. Look for missing fundamentals, a changed effective universe, and too few qualifying events. A pipeline mismatch means the comparison wasn't the one you specified.
If the mechanics are sound, inspect the failure without immediately optimizing against it. The thesis may have depended on one market environment. A criterion may have been redundant. The development winner may have been the luckiest member of the search. Record which explanation you investigate and why.
You may revise the rules after learning from the failure. That's research. But the revised version needs fresh evidence. In live research, future observations become the next true holdout. If you keep slicing the same finite history, use nested walk-forward testing: each outer test block stays untouched while all selection and tuning happen inside earlier data. Even that doesn't create endless independent samples.
A successful holdout needs restraint too. One pass doesn't prove causality or guarantee another period will agree. It says the frozen result survived one genuinely unseen block. The formal literature goes further: Bailey and coauthors define the probability of backtest overfitting by repeatedly separating model selection from out-of-sample ranking across complementary splits. Their method is more elaborate, but its discipline is familiar: selection and evaluation can't use the same observations without paying a penalty.
When is walk-forward testing better?
A single holdout is clearest when you're finishing one thesis and can afford to reserve recent history. Walk-forward testing is better when the research process itself updates over time. At each simulated date, rules are chosen using only earlier observations, then evaluated on the next unseen block. The clock advances and the process repeats.
Walk-forward testing uses more of the history for sequential out-of-sample checks, but it is easier to implement badly. Every tuning decision inside each step must stay behind that step's information boundary. If a rule was designed after you saw the whole period, replaying it through rolling windows doesn't restore blindness.
Use both ideas at different levels. Develop and stress-test the rule on an early sample. Preserve a final holdout for the full research decision. If the strategy requires scheduled recalibration, make that recalibration part of the frozen walk-forward procedure. A sound backtest still needs point-in-time inputs and filing-date alignment; a holdout can't rescue data that leaked from the future.
The practical rule is blunt: unseen means unseen. Mark the boundary, freeze the version, declare the read, and run it once. If you change the rules afterward, say what the holdout taught you and stop borrowing its independence. Quantery preserves the thesis version and report, so keep both as the record. The next test should challenge the process you actually followed. A cleaner story written afterward doesn't count.
Research tooling, not investment advice. Nothing here is a recommendation to buy, sell, or hold any security. Screens, scores, and backtests are informational only; backtested results are hypothetical, exclude costs such as commissions and slippage, and do not guarantee future results. Verify against primary filings and make your own decisions.
Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.
← All articles