How to stress-test a backtest exit rule
Stress-test a backtest exit rule by holding the entry signal and every other setting fixed, then changing only when positions close. Compare a declared baseline with nearby maximum holds (a cap on how many months a simulated position can stay open) and with gate-fail exits switched on or off (closing a position when the company no longer passes the thesis gate at a later evaluation). If the conclusion survives, the entry thesis has support across plausible holding policies. If it flips, the exit rule is part of the thesis and needs its own explanation.
That distinction matters because an attractive portfolio curve can come from a good selection rule, a fortunate selling rule, or the interaction between them. A backtest blends all three. Pull the exit policy apart before giving the entry signal credit.
An exit rule makes an economic claim
An exit rule is the condition that closes a simulated position. Quantery theses can use a maximum holding period, an exit after the company fails the gate at a later evaluation, or both. The rebalance schedule supplies the dates when the engine can notice that failure. Rebalance cadence and holding period do different jobs, so keep them separate in your test plan.
The rule should follow the mechanism in the thesis. Suppose the claim is that a newly cheap, cash-generative business tends to be re-rated as the market digests its filings. A maximum hold says how long you expect that process to deserve room. A gate-fail exit says fresh evidence can cancel the original claim before the clock runs out.
Those are different beliefs. The clock-based rule assumes the signal decays with time. The gate-based rule assumes losing eligibility is useful exit information. Combining them is reasonable, but the combination can hide which belief produced the result.
Write the exit sentence before running anything:
Hold until the original condition fails at a scheduled evaluation, subject to a maximum research horizon.
Now write what would make that sentence wrong. Perhaps companies that dip below the gate recover before the next filing. Perhaps most gains arrive after the chosen horizon. Perhaps a price-sensitive ratio flickers around its cutoff and creates churn. Each possibility becomes a controlled rerun.
Hold the entry side still
An exit test teaches you little if the companies entering the portfolio also change. Preserve every entry-side field: universe, gate, features, criteria, reporting lag, entry timing, and rebalance cadence. Save and activate one new thesis version for each exit variant, changing only one exit setting in each version. Run every version over the same test window with the same benchmark.
Start from a baseline chosen for a reason. The shape below is illustrative, and the bundled templates are the reference for exact fields. The Quantery thesis documentation covers the workflow around the DSL.
# Baseline exit policy: illustrative
backtest:
rebalance: monthly
reporting_lag_days: 1
entry: next_close
exit:
max_hold_months: 12
on_gate_fail: true
benchmarks: [SPY]
The values are proposed test parameters. They aren't universal defaults. Save the baseline, then create controlled versions:
- Shorten
max_hold_monthswhile leavingon_gate_failunchanged. - Lengthen
max_hold_monthswith the same gate-fail behavior. - Switch
on_gate_failoff while preserving the baseline maximum hold.
Run the variants in the order you declared. Don't inspect the first result and invent the next rule needed to repair it. Bailey, Borwein, López de Prado, and Zhu show why trying more strategy configurations raises the chance of selecting an impressive backtest produced by overfitting (their paper on backtest overfitting). An exit grid is still a search. Preserve every run, including the awkward ones.
Maximum holds test signal decay
A maximum hold closes a position once it reaches a chosen age, unless another exit condition closes it first. It tests a claim about how long the signal remains useful.
Read neighboring holds as a shape rather than a contest. If a shorter hold and the baseline point in the same direction, the result doesn't require waiting for the last stretch. If the longer hold keeps adding relative performance, the original horizon may have cut the proposed mechanism short. If performance peaks at one exact hold and collapses on either side, the timer is carrying more of the result than the business rule.
Don't compare total return alone. Open the event study at matching forward horizons. An event study follows each fresh qualification from its signal date, independent of the portfolio's eventual exit. If the event pattern fades before the portfolio's maximum hold, the policy may be keeping positions after the entry signal has spent itself. If events keep improving after the portfolio closes them, test whether the longer horizon survives in another period.
The event study and portfolio weight history differently. Overlapping positions share market dates, while event observations begin whenever a company newly qualifies. Their disagreement isn't a broken report. It is evidence that holding policy matters. Read the event study and portfolio as separate questions before changing the rule.
A long hold can also turn a stock-selection test into an exposure test. Positions admitted during one market regime may remain in the simulated portfolio long after their original filing signal. The report doesn't expose holding-window composition or cohort attribution, so the curve can't tell you which group carried it. Split the date range into fixed subperiods and rerun the same thesis versions. Then test whether the screen is a sector bet with controlled universe variants as a separate sensitivity test.
Gate-fail exits test whether fresh evidence cancels the thesis
An exit on gate failure closes a held company after a later evaluation says it no longer qualifies. This feels disciplined, but it can mean several things.
A filing-driven gate failure may contain fresh business evidence: cash flow turned negative, leverage crossed a limit, or profitability weakened. A price-driven failure may say only that the stock became less cheap. A relative-rank failure can happen because peers changed even when the company did not. The same switch serves different economic claims.
Compare the baseline with gate-fail exits disabled. If results improve without the gate-fail rule, compare the aggregate portfolio statistics and equity curves, then use the go-event list to inspect companies that repeatedly qualified. The report doesn't provide a position ledger, so it won't identify every gate-fail sale for you. The original gate may be useful for entry while being poor for sale decisions. Cheapness is a common example: a rising price can make a company fail a value gate for the exact reason the entry thesis expected. Treating that failure as new bad evidence muddles valuation with deterioration.
If disabling gate-fail exits makes results worse, the report still can't tell you whether the rule removed deteriorating companies or just cut losers quickly in that sample. To investigate criterion families, create a separate thesis sensitivity test. Changing the gate changes entry eligibility too, so don't include that run in the exit-only comparison. A test that removes a valuation condition from the gate asks a different question from one that removes a cash-generation condition.
Boundary churn deserves its own check. A company near a hard cutoff can leave and later create another go event even though its business changed little. Compare the go-event lists for repeated qualifications by the same company. Treat that pattern as a warning. It isn't a complete trade count because the report has no position ledger or holding-duration table. Moving a gate threshold or replacing a strict gate with a score gate may reveal sensitivity to the boundary, but either change also alters entries. Label the result as a second-stage thesis test and keep it out of the exit grid.
Faster exits bring excluded costs closer
Every Quantery backtest is hypothetical and excludes trading costs. A gate-fail rule may create more trading, but the report doesn't expose transaction turnover or subtract spreads, slippage, taxes, or commissions.
The mechanism doesn't require a cost estimate to matter. FINRA notes that more buying and selling can raise a fund's transaction costs, which reduce returns (FINRA's explanation of portfolio transaction costs and turnover). Use repeated go events as a warning that a gate may be unstable. They don't measure trading frequency. The report provides no full evaluation history or position ledger, so transaction burden remains unknown. A gross advantage that appears only with gate-fail exits deserves less confidence for that reason.
Don't bolt an invented flat cost onto every trade and declare the problem solved. Spreads vary by company and date, while market impact depends on trade size. Use plausible cost assumptions when you have defensible inputs. Until then, keep the conclusion about the signal modest.
Read stability before the winning variant
A useful exit stress test produces a map of dependence:
- Stable direction: nearby maximum holds and either gate-fail choice preserve the broad conclusion. Keep investigating the entry signal.
- Horizon dependence: only one holding range works. The thesis now includes a timing claim that needs evidence across other periods.
- Gate dependence: results hinge on selling as soon as eligibility disappears. Use a separate thesis sensitivity test to study the gate, knowing that it changes entries too.
- Repeated qualifications: the go-event list shows the same companies qualifying again. Treat that as a warning about gate stability and unknown trading costs.
- Period dependence: the result appears in one fixed subperiod and disappears in another. Keep the exit policy fixed before testing universe changes separately.
None of those readings turns a historical association into causality. Coverage, benchmark choice, market regime, and outliers still matter. What makes a backtest trustworthy remains the floor beneath the exercise.
Stop once the declared variants are finished. Choose an exit policy whose mechanism you can state before seeing the curve, then take that fixed policy to an untouched period. If it fails there, don't return to the original sample for a more flattering timer. Record the failure and revise the claim.
The useful result may be that the entry gate works across several exits. It may be that a deterioration rule carries all the value. It may be that nothing survives. Quantery keeps the policy editable and each run tied to its settings, so those outcomes stay visible. The selling rule is yours. Make it earn its place.
Research tooling, not investment advice. Nothing here is a recommendation to buy, sell, or hold any security. Screens, scores, and backtests are informational only; backtested results are hypothetical, exclude costs such as commissions and slippage, and do not guarantee future results. Verify against primary filings and make your own decisions.
Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.
← All articles