← Quantery Blog

How to choose the right backtest benchmark

September 15, 2026 · 8 min readbacktestingbenchmarksmethod

The right benchmark is the return you could plausibly have earned from the same part of the market without running your thesis. A broad US large-company screen can reasonably start with SPY. A small-company screen needs a small-company comparison such as IWM. A screen built to favor cheap small companies should also be checked against a small-value proxy such as IWN. Match the alternative to the screen's investable universe before you look at the result.

This changes the question a backtest answers. Beating a broad market fund doesn't establish that your accounting rules found an edge if the screen spent the period loaded with small or value companies and those groups had a good run. First ask whether the thesis beat an investable alternative with similar built-in exposures. Then inspect whether the remaining gap survives different periods and assumptions.

A benchmark is the alternative you chose not to run

A benchmark isn't scenery behind an equity curve. It's the baseline claim. If a hypothetical screen returned more than its benchmark before costs, the difference is the screen's apparent excess return. Change the baseline and that difference changes too, even though the screen made exactly the same selections.

SPY is a useful first baseline because the fund seeks to track the S&P 500 before expenses, according to State Street's SPY fund page. It represents a simple alternative for a thesis operating among large US companies. It doesn't represent every corner of the US equity market.

The mismatch becomes obvious with a tiny-company balance-sheet thesis. Quantery's Graham Net-Net template permits much smaller companies than a large-company index contains. It also looks for deep discounts to net current asset value, which gives it an unmistakable value bias. Compare that thesis only with SPY and the reported gap mixes at least two things: the screen's rules and the behavior of small value stocks during the test.

IWM seeks to track the Russell 2000 and offers broad exposure to US small-cap stocks, as BlackRock describes on the fund page. IWN targets US small-cap value stocks and tracks the Russell 2000 Value Index, according to its corresponding fund page. Neither is a perfect twin for a net-net screen. Either can tell you more than SPY alone, and IWN asks the harder question: did the thesis add anything beyond holding the same broad style?

Match the benchmark to the universe first

Start with eligibility. Which companies could the rules select? Compare the screen's exchange and size constraints with the proposed benchmark's holdings policy. Check the sector exclusions and liquidity floor as well. This is why choosing the universe your screen runs on belongs in backtest design.

Size comes first because it changes the opportunity set. The bundled Buffett Quality Value template starts with larger companies, while Graham Net-Net keeps the floor low enough to admit microcaps. SPY is much closer to the first population. IWM is closer to the second. A benchmark needn't hold every eligible company, but it shouldn't describe a different market by construction.

Style comes next. A high free-cash-flow-yield rule has a value tilt before any result is produced. A high return-on-capital rule adds a quality tilt. A falling-leverage rule can favor a different financial profile again. These aren't accidental labels applied after the backtest. They're consequences of the equations.

Sector exclusions matter too. A thesis that removes banks, real estate, and utilities isn't taking the same bet as an all-market fund. The benchmark can rally because sectors the thesis forbids had a strong period. That doesn't make the thesis defective. It means the relative return contains an allocation decision that should be named.

Don't hunt for a benchmark that mirrors every rule. At that point you'd have rebuilt the thesis and learned nothing. Match the broad opportunity set and its obvious style, then use diagnostics to account for the rest.

Use a benchmark ladder instead of one perfect index

One comparison rarely separates all the moving parts. Use a short benchmark ladder, with each rung asking a stricter question.

  1. Broad market: Did the screen outperform a simple US equity alternative?
  2. Matched size: Did it outperform companies from a similar size range?
  3. Matched style: Did it outperform a size-and-style alternative that already owns much of the thesis's structural tilt?

For a small-company value screen, that ladder could be SPY, IWM, then IWN. For a broad large-company quality screen, the first comparison may already be close enough, with sector and quality exposure handled as diagnostics. Treat the tickers as examples of investable proxies. Choose your own defaults from the thesis.

Read the ladder from broad to demanding. If a hypothetical screen beats SPY but trails IWM, its apparent edge may be small-company exposure. If it beats IWM but trails IWN, value exposure may explain more of the result. If it beats all three, the stock-selection rules have earned further investigation. They haven't proved causality. A sector concentration, rebalance effect, or handful of events can still account for the gap.

Keep each comparison on identical dates. Funds start trading at different times, data coverage can thin in earlier periods, and a report can shift endpoints around available prices. Compare every screen return with the benchmark return recorded inside that same run. Don't carry a headline return from one report into another.

Factor returns explain the result but aren't portfolios

A factor is a long-short return series designed to isolate an exposure. The Fama-French size factor, SMB, subtracts the average return of large-company portfolios from small-company portfolios. The value factor, HML, subtracts growth portfolios from high book-to-market portfolios. The Kenneth French Data Library explains both constructions and publishes the underlying research returns.

Those series answer a diagnostic question: during the months when your thesis won, were small companies or value companies winning too? A strong relationship doesn't make the screen useless. It tells you what risk or market regime may be paying it.

Don't substitute a factor return for an investable benchmark. SMB and HML are constructed research portfolios with long and short legs. SPY, IWM, and IWN are funds a reader could identify as practical alternatives, with fees and tracking differences. Use an ETF comparison for the decision baseline. Use factors to explain why the baseline gap moved.

There is a simpler check when you don't want to run a regression: bucket the results. Compare performance in periods when SMB was positive with periods when it was negative. Do the same for HML. If the thesis only works when its favored style is already winning, call it a conditional result and test that condition across another window.

Put the benchmark choice in the thesis before the run

Benchmark selection can become another form of result shopping. Run against SPY, dislike the answer, switch to IWM, then keep whichever comparison flatters the thesis. The cure is dull and effective: write down the benchmark ladder before running the test.

The illustrative shape below shows where that decision belongs. The bundled templates are the reference for exact fields, and the Quantery thesis documentation covers the full DSL.

# Small-company value test: illustrative
universe:
  exchanges: [NYSE, NASDAQ, AMEX]
  min_market_cap: 25000000
  exclude_sectors: [Financial Services, Real Estate]

features:
  current_assets: newest(current_assets)
  liabilities: newest(total_liabilities)
  ncav: current_assets - liabilities
  price_to_ncav: if(ncav > 0, market_cap / ncav, null)

criteria:
  discount:
    rules:
      - { when: "is_null(price_to_ncav)", score: 0, flag: no_ncav }
      - { when: "price_to_ncav <= 0.66", score: 2 }
      - { when: "price_to_ncav <= 1.0", score: 1 }
      - { else: 0 }

gate:
  mode: strict

backtest:
  rebalance: monthly
  benchmarks: [IWN]

The benchmark does not alter which companies pass. It changes the comparison you promised to take seriously. Save another version with a different predeclared benchmark when you want the full ladder, and keep the screen rules and test dates fixed.

Every return from these runs is hypothetical and excludes trading costs. That warning bites harder when the thesis reaches smaller companies: the investable fund may trade cheaply while the selected names carry wider spreads and more slippage. A matched index solves the comparison problem. It doesn't solve implementation.

A fair benchmark makes failure useful

Suppose a screen beats SPY and loses to a matched value proxy. That's not an embarrassing result to hide. It says the screen may be a complicated way to obtain a familiar exposure. You can simplify it, change the claim, or demand evidence that the accounting rules improve on that exposure in another period.

Suppose it beats the matched proxy but fails when the start date moves. The benchmark choice was sound and the result was fragile. Use the perturbation method in how to tell if a backtest result is real: change one setting at a time and keep the conclusion only if it survives. The broader rules for point-in-time data, look-ahead, and excluded costs still apply, as what makes a backtest honest explains.

A useful report should leave you with a narrower claim than the one you started with. This small-value thesis beat a broad market proxy, but its size tilt explained the difference. Or it beat a matched alternative across several windows, while losing when liquidity assumptions tightened. Those are research findings you can act on by changing the test.

Pick the alternative first. Add a stricter rung when the thesis has an obvious style. Keep the dates fixed, read each run against its own baseline, and inspect what remains. The benchmark isn't there to make the curve look respectable. It's there to stop borrowed exposure from masquerading as a discovered edge.

Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.

← All articles