← Quantery Blog

How to test a thesis across market regimes

September 24, 2026 · 8 min readbacktestingmarket-regimesrobustness

A market-regime test asks whether the same thesis behaves differently under a condition defined before you inspect its returns. Split one backtest by a rule such as falling versus rising policy rates, keep the stock-selection rules fixed, and compare the thesis with the same benchmark inside each bucket. If the result only holds in one bucket, you've found a conditional claim that needs another test.

That answer sounds simple. The trap is the word regime. It can mean a condition you could have observed at the time, or a story attached after the chart is complete. Only the first kind belongs in a clean test. "The inflation shock" is a historical label. "The latest published inflation reading is above its trailing average" is a rule a simulated investor could have followed.

Define a regime as a reproducible rule

In research, a market regime should be a reproducible classification of dates. Give two people the inputs and the rule, and they should put each date in the same bucket without seeing what your thesis returned afterward.

Useful regime variables include policy rates, inflation, credit conditions, broad-market trend, and volatility. Each one describes a different mechanism. A leveraged value thesis may plausibly depend on financing conditions. A quality thesis may respond to economic contraction. A short-horizon momentum thesis may care more about market trend. Start with the mechanism in your thesis, then choose the variable.

Write the classification in one sentence before running it:

At each monthly rebalance, classify the date by whether the published policy-rate series is above or below its level twelve months earlier.

The Federal Funds Effective Rate series is a public monthly series sourced from the Federal Reserve Board. It makes the rule auditable. Twelve months is an illustrative lookback, chosen in advance. Change that lookback later and you've run another specification, which belongs in the research log.

Resist labels such as easy money or panic. They sound precise because you know the history already. They aren't calculations. A regime rule needs a named input and threshold. It also needs an observation date and update schedule.

Hindsight can leak through the label

Recessions are the cleanest example of a label that arrives late. The National Bureau of Economic Research dates US business-cycle peaks and troughs after weighing several measures of activity, including employment and income. On June 8, 2020, its committee identified February 2020 as the monthly peak. On July 19, 2021, it identified April 2020 as the trough.

That chronology is useful for describing history. It isn't a tradable signal on those turning dates. A backtest that switches behavior in February because February was later called a peak knows something the simulated investor didn't know then. The FRED recession indicator records the official chronology as a binary series, but the underlying turning points are still retrospective.

There are two defensible ways to use such labels:

Don't slide from the first use to the second claim. "This thesis performed differently during recessions as later dated" is a historical observation. "This thesis could detect recessions" is a new claim and needs a signal available in real time.

The same problem appears in price-based labels. Calling an interval a bear market after measuring its eventual decline uses the endpoint to define the beginning. A mechanical trend rule based only on prices already observed at each rebalance avoids that leak, though it answers a narrower question.

Keep the thesis fixed while the calendar changes

A regime test isn't a license to maintain one thesis per environment. Keep the universe, features, gates, score bands, rebalance cadence, reporting lag, and benchmark fixed. Then partition the resulting decisions or returns by the predeclared regime rule.

That order matters. Suppose you loosen the debt gate during falling-rate months because the first result looked weak there. You no longer know whether the difference came from the environment or the edited thesis. First test the same machine in different weather. Only afterward should you propose a second machine, and then it needs its own out-of-sample test.

Use this sequence:

  1. State why the thesis should depend on the chosen condition.
  2. Freeze the regime rule and the thesis before reading bucket returns.
  3. Run the full period with one benchmark and one set of assumptions.
  4. Assign each rebalance date using information available then.
  5. Compare excess returns and drawdowns inside each bucket. Record turnover and event counts too.
  6. Move one regime boundary or lookback and run the partition again.

Excess return means the thesis return minus its benchmark return over the same dates. The benchmark must stay aligned because a bucket can contain a broad rally or selloff. If both the thesis and benchmark suffer in one environment, the relevant question is whether the thesis behaved differently from the alternative you could have held. Choosing the right backtest benchmark explains that baseline.

All backtest returns in this workflow are hypothetical and exclude trading costs. Turnover belongs beside performance because an apparent regime advantage can require much more trading precisely when spreads and slippage are least friendly.

Small buckets produce big stories

A full backtest is one historical path. Splitting it creates smaller samples, and adjacent months within one episode aren't independent experiments. A thesis with repeated signals during the same contraction may have dozens of rows but only one economic event behind them.

Read the bucket table with four questions:

The third question catches a common impostor. A rate-sensitive result may really be a sector result if one bucket selected far more banks, homebuilders, or utilities. Check sector mix inside each bucket. The same discipline used to ask whether a stock screen is just a sector bet applies here.

The fourth question is a perturbation test. If "rising rates" means higher than a year earlier, repeat the analysis with a nearby predeclared lookback. If one reasonable boundary reverses the conclusion, don't average the variants until the result looks respectable. Report that the classification controls the finding. A robust backtest result keeps its direction when a setting that shouldn't own the thesis gets nudged.

Avoid adding buckets to rescue a weak result. Two states are often enough for a first pass. Every extra split reduces the observations per bucket and creates another place for chance to look like a mechanism.

Put the economic claim before the code

Quantery theses define company selection, while the regime analysis classifies dates around the resulting backtest. The thesis itself should contain the business variables that express the proposed resilience. For a rate-sensitive quality test, that might mean cash generation and interest coverage. Use the policy-rate series as the external calendar partition. Don't smuggle it into the company score.

The shape below is illustrative, and the bundled templates are the reference for exact fields. The Quantery thesis documentation covers the full DSL.

# Financing resilience: illustrative
params:
  coverage_strong: 6
  coverage_floor: 3
  debt_assets_max: 0.45

features:
  ebit_ttm: ttm(operating_income)
  interest_ttm: ttm(interest_expense)
  interest_coverage: if(interest_ttm > 0, ebit_ttm / interest_ttm, null)
  debt: newest(total_debt)
  assets: newest(total_assets)
  debt_assets: if(assets > 0, debt / assets, null)
  fcf_ttm: ttm(free_cash_flow)

criteria:
  coverage:
    rules:
      - { when: "is_null(interest_coverage)", score: 0, flag: no_coverage }
      - { when: "interest_coverage >= $coverage_strong", score: 2 }
      - { when: "interest_coverage >= $coverage_floor", score: 1 }
      - { else: 0 }
  leverage:
    rules:
      - { when: "is_null(debt_assets)", score: 0, flag: no_leverage }
      - { when: "debt_assets <= $debt_assets_max", score: 2 }
      - { else: 0 }
  cash:
    rules:
      - { when: "fcf_ttm > 0", score: 2 }
      - { else: 0 }

gate:
  mode: strict

Now the claim is concrete: companies with stronger debt service, moderate balance-sheet leverage, and positive trailing cash flow may hold up better when rates are rising. Run that one thesis across the full period. Then compare its hypothetical excess returns in dates classified by your policy-rate rule.

A positive result would narrow the claim. It would say these rules were associated with different outcomes under this specified condition in the available history. It wouldn't establish that rates caused the difference, nor that a future rate cycle will resemble the observed ones. Point-in-time inputs and filing-date alignment still govern the company data, while release timing governs the macro series.

A conditional result can still be useful

Suppose the thesis beats its matched benchmark in rising-rate months and trails in falling-rate months. Don't promote the first bucket and bury the second. The pair is the result. Ask whether the business mechanism fits, whether several episodes contribute, and whether nearby definitions preserve the direction.

If those checks hold, you have a conditional hypothesis worth testing on another period. If they don't, you have learned where the story breaks. Both outcomes are better than one full-period average that mixes unlike conditions and explains none of them.

Save the regime definition beside the run notes. Save every variant too. The useful asset isn't a permanent label for the market. It's a reproducible question you can change, re-run, and try to disprove.

Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.

← All articles