← Quantery Blog

Is your stock screen just a sector bet?

September 19, 2026 · 8 min readbacktestingsectorsmethod

A stock screen is a sector bet when its rules consistently favor the accounting shape of one industry, even though the thesis claims to find a trait that should travel across industries. Test for that by splitting every result by sector, comparing the signal within sectors, and running versions that remove the dominant sector. If the apparent edge vanishes, narrow the claim.

This matters because ratios don't start on level ground. A margin rule favors business models with high markups. An asset-turnover rule favors businesses that push a great deal of revenue through a small asset base. A leverage rule can reject capital-heavy industries before it measures what their debt funds. The screen may be doing the arithmetic correctly and still answer the wrong question.

Why can a sensible ratio become an industry classifier?

Financial ratios divide one statement value by another. The numerator and denominator reflect the economics of the business. That sounds obvious, but it means a ratio can identify how a company operates before it identifies how well it operates.

Take net margin, net income divided by revenue. A grocer can run an excellent operation on thin margins because its inventory turns fast. A software company can carry a much wider margin because distributing another copy costs little. Set one market-wide margin threshold and you've partly encoded a preference for the second model. Passing the rule isn't yet evidence that one company is better managed.

Return on assets has the same problem from another direction. Asset-light businesses need little property or inventory relative to earnings. Manufacturers, railroads, and utilities need a great deal. A universal hurdle mixes operating quality with the amount of physical capital the industry requires.

Even a relative rule can hide the issue. pct_rank() tells you where a company sits among the names in the declared universe. If that universe contains every sector, the top decile of a margin rank can fill with naturally high-margin industries. The calculation is valid. The comparison group is doing more work than the thesis admits.

Sector labels aren't natural law either. The MSCI description of GICS says the system classifies companies by primary business activity through a hierarchy from sector to sub-industry. A conglomerate still receives one classification, and companies near a boundary won't become economic twins because a label says so. Use sectors as a diagnostic partition, then inspect the businesses that make the conclusion move.

Start with the sector table before the equity curve

For each rebalance or qualification date, record the sector of every company that passed. Then summarize three things:

The second line supplies the base rate. If a sector supplies a large share of the eligible universe, a large share of selections isn't surprising. The useful comparison is selection rate: selected companies in the sector divided by eligible companies in that sector. A screen that selects one sector at several times the rate of the rest has a concentration worth explaining.

Contribution matters because counts can mislead. A sector may produce many uneventful selections while a different sector accounts for most of the excess return. Or one extreme company may carry a sector's average. Read the individual events behind any dominant contribution before telling a broad story.

Use sector composition at every date instead of a summary assembled from today's survivors. Historical membership changes. Delisted companies still belong in the test, and a company can change its principal activity. The same point-in-time discipline that prevents survivorship and look-ahead bias applies to classification data.

Industry portfolios offer a useful reference for this habit. The Kenneth French Data Library's industry construction assigns listed companies to industry portfolios from their SIC codes and reconstitutes those portfolios each June. You don't need to copy that exact convention. You do need a convention that was knowable at the simulated date and applied consistently.

Compare each company with the right peers

If your claim is relative, make the comparison relative. “High margin” across the market and “high margin for this kind of business” are different theses.

The clean design is a within-sector rank. First calculate the feature for every eligible company. Then rank each company only against peers in the same sector or industry. Combine those local ranks afterward if the broader thesis truly applies across business models. A retailer near the top of retail can then stand beside a software company near the top of software without pretending their raw margins mean the same thing.

Quantery's current pct_rank(feature_name) ranks across the thesis universe after exclusions. It doesn't partition the rank by sector. So don't write pct_rank(net_margin) and call the result sector-neutral. You have two practical choices today: narrow the universe to an economically coherent group and run separate thesis versions, or keep the broad screen and treat the sector breakout as a required diagnostic rather than a solved adjustment.

That limitation can improve the research. Separate versions force you to ask whether the same proxy belongs in every group. Banks and insurers can't be treated as low- or high-leverage industrial companies. Their liabilities are part of the operating model. Real estate businesses and utilities have their own accounting shapes as well. Choosing the universe your screen runs on explains why some sectors need different arithmetic instead of a more forgiving threshold.

Don't tune a custom threshold for every sector after seeing returns. That turns a diagnostic into a search over knobs. Define the peer groups and thresholds from business reasoning first, save the versions, and keep the failed runs.

Run the leave-one-sector-out test

A leave-one-sector-out test reruns the same hypothetical backtest after removing one sector. Nothing else changes: same feature definitions, thresholds, dates, cadence, benchmark, and reporting lag. Repeat for each sector that dominates selections or performance.

Read the outcomes in this order:

  1. Does the direction of excess return survive?
  2. Does the number of qualification events remain broad enough to interpret?
  3. Do drawdowns and turnover change for an economic reason?
  4. Does a different sector take over, suggesting the ratio is still sorting business models?

If removing one sector erases the result, the right conclusion isn't automatically that the screen failed. Perhaps the original claim genuinely belongs to that sector. Rewrite it that way and test it against an appropriate sector comparison. What failed was the broad claim.

If the direction survives every removal but the magnitude jumps around, keep the direction and distrust the headline size. If excluding any large sector leaves too few events, you don't have much evidence about portability. “Works everywhere” needs observations from more than one kind of business.

All of these backtests are hypothetical and exclude costs. Use the same benchmark inside each run, and remember that a broad benchmark can obscure a concentrated sector exposure. The benchmark ladder method handles the investable alternative. Sector removal handles the composition of the thesis. You need both checks because they answer different questions.

What should the thesis code say?

Start by making the scope visible. The shape below is illustrative, and the bundled templates are the reference for exact fields. The Quantery thesis documentation covers the full DSL.

# Broad quality test with sector scope made explicit: illustrative
meta:
  name: broad-operating-quality
  label: Broad Operating Quality
  description: Profitability and cash backing across non-financial businesses.

universe:
  exchanges: [NYSE, NASDAQ, AMEX]
  min_market_cap: 500000000
  exclude_sectors: [Financial Services, Real Estate]

quality:
  required:
    market_cap: no_market_cap
    revenue_ttm: no_revenue

params:
  margin_ok: 0.08
  conversion_ok: 0.70

features:
  revenue_ttm: ttm(revenue)
  income_ttm: ttm(net_income)
  fcf_ttm: ttm(free_cash_flow)
  net_margin: if(revenue_ttm > 0, income_ttm / revenue_ttm, null)
  cash_conversion: if(income_ttm > 0, fcf_ttm / income_ttm, null)

criteria:
  profitability:
    rules:
      - { when: "is_null(net_margin)", score: 0, flag: no_margin }
      - { when: "net_margin >= $margin_ok", score: 1 }
      - { else: 0 }
  cash_backing:
    rules:
      - { when: "is_null(cash_conversion)", score: 0, flag: no_conversion }
      - { when: "cash_conversion >= $conversion_ok", score: 1 }
      - { else: 0 }

gate:
  mode: strict

The code doesn't solve sector dependence. It states the market-wide claim clearly enough to test. Save that version before seeing the outcome. Then make copies that each target one coherent group, or copies that exclude a dominant group, without moving the criteria.

The net_margin rule is the obvious suspect. Cash conversion may also vary with working-capital cycles and capital intensity. Keep both visible in the result so you can tell which criterion changes the sector mix. This is another reason to avoid counting the same signal twice: a stack of related profitability ratios can give an industry's accounting shape several votes.

When should you narrow the claim?

Narrow it when the result depends on a sector and you can explain why the rule belongs there. “This margin-and-cash rule identified stronger companies within one industry during these samples” is a useful finding. “Quality wins” isn't, if quality meant membership in that industry.

Also narrow it when the raw metric has a different economic meaning across groups. Don't rescue a broken cross-sector leverage rule by giving every sector its own cutoff. Build the financial-company version from ratios that fit financial companies, and keep the industrial-company version separate. Different arithmetic deserves a different thesis.

A result that survives sector breakouts, peer comparisons, and leave-one-sector-out runs has earned a stronger next test. It still hasn't established causality, and a size tilt or one historical window may explain it. Shake those next, one change at a time.

The useful outcome isn't a certificate that a screen is sector-neutral. It's a claim with its hidden industry preference dragged into view. Keep the broad version if the evidence travels. Narrow it if it doesn't. The rules are yours, which means the comparison group has to be yours too.

Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.

← All articles