How to avoid counting the same signal twice
A thesis counts the same signal twice when several criteria react to the same underlying fact. Return on equity, net margin, and earnings growth can all improve because profit rose. Giving each one a full allotment of points makes that profit change look like several independent arguments. It isn't.
Fix it by naming the business claim behind every criterion, grouping criteria that share a driver, and limiting how much each group can contribute to the total score. The goal isn't a score with the most inputs. It's a score whose parts can disagree for useful reasons.
Why do overlapping criteria distort a score?
A composite score looks additive. A company earns points for quality, cash generation, balance-sheet strength, and valuation, then the points are summed. That structure invites a dangerous shortcut: keep adding ratios that sound sensible.
The trouble is that accounting ratios reuse the same statement lines. Consider a cluster built from these features:
- return on equity: net income divided by equity
- net margin: net income divided by revenue
- earnings growth: the change in net income across periods
- cash conversion: free cash flow divided by net income
Each ratio asks a different question. Still, net income runs through every one. A temporary gain can lift return on equity and net margin at once. A write-down can depress both. Cash conversion may then move in the opposite direction because its denominator changed. The score is reacting several times to one accounting event.
Some duplicates are exact. Price-to-earnings and earnings yield contain the same information because one is the reciprocal of the other. Adding a criterion for each doesn't broaden the thesis. It gives the same valuation observation another vote.
Other overlaps are economic. Return on equity and return on assets can both reward the same profitable operation, though leverage makes them diverge. Free cash flow yield and earnings yield can both reward a cheap, mature business, though working capital and capital spending separate them. These pairs aren't interchangeable. They also aren't independent by default.
That distinction matters because a high score feels like accumulated evidence. If most points trace back to one line in one filing, the confidence implied by the total is false precision.
Start with business claims before choosing ratios
Write the scorecard in plain language first. Each criterion family should answer a question that could fail while the others pass.
- Operating economics: Does the business earn an adequate return on the capital supporting it?
- Cash backing: Do reported profits turn into cash after the asset base is funded?
- Balance sheet: Could debt claims consume the cash the operation produces?
- Valuation: How much operating or equity cash flow does the current price offer?
Now assign a proxy to each claim. Return on equity might represent operating economics. Free cash flow divided by net income might represent cash backing. Net debt divided by EBITDA might represent balance-sheet pressure. Free cash flow yield might represent valuation.
The labels force a useful question: what new fact does this ratio add? If net margin sits beside return on equity, decide whether it captures a distinct part of the claim or repeats profitability with another denominator. If it adds context but doesn't deserve another vote, keep it as a feature for inspection and leave it out of the score.
A feature doesn't have to award points to be useful. It can explain a result, flag an edge case, or help you inspect a survivor. Scores should be scarce. Diagnostics don't need to be.
This is the same discipline used when turning a hunch into measurable proxies: state what each proxy stands for and write down the gap between the number and the idea. Here, go one step further and write down which other proxies share that gap.
Can the criteria disagree for a real reason?
Use counterexamples before using statistics. For every pair of criterion families, imagine a business that passes one and fails the other.
A high-return business can have weak cash backing when receivables swell or capital spending absorbs the operating cash. A cash-rich balance sheet can sit under a mediocre operation. A strong business can trade at a yield below your valuation threshold. A cheap business can carry enough debt to make the equity fragile.
Those disagreements show that the criteria observe different failure modes. If you can't describe a plausible disagreement, inspect the formulas. You may have renamed the same bet.
Don't demand zero correlation. Good businesses often look good across several measures, and a broad market shock can move many valuation ratios together. Statistical independence is too strong a standard for financial statements. What you want is decision independence: each family has a reason to change the verdict when the others hold steady.
Sector effects deserve the same test. A high-margin software company and a high-turnover retailer can both have sound operating economics, but a single margin threshold favors one business model before quality has been assessed. Compare a company with peers when the claim is relative, or narrow the universe when the arithmetic belongs to one sector. The universe is part of the thesis, including the sector assumptions built into its ratios.
How should a thesis limit duplicate votes?
The cleanest structure gives each claim one criterion and lets ordered rules grade the best proxy for that claim. Related measures can act as guards or annotations without creating another point source.
The shape below is illustrative. The bundled templates are the reference for exact fields, and the Quantery documentation covers the full thesis DSL.
# Separate claims, limited votes: illustrative
params:
roe_strong: 0.15
roe_ok: 0.10
conversion_strong: 1.0
conversion_ok: 0.70
leverage_max: 2.0
yield_strong: 0.06
yield_ok: 0.03
features:
ni_ttm: ttm(net_income)
fcf_ttm: ttm(free_cash_flow)
equity: latest(total_equity)
cash: latest(cash)
debt: latest(total_debt)
ebitda_ttm: ttm(ebitda)
roe: if(equity > 0, ni_ttm / equity, null)
cash_conversion: if(ni_ttm > 0, fcf_ttm / ni_ttm, null)
net_debt_ebitda: if(ebitda_ttm > 0,
(debt - cash) / ebitda_ttm, null)
fcf_yield: if(market_cap > 0, fcf_ttm / market_cap, null)
criteria:
operating_quality:
rules:
- { when: "is_null(roe)", score: 0, flag: no_roe }
- { when: "roe >= $roe_strong", score: 2 }
- { when: "roe >= $roe_ok", score: 1 }
- { else: 0 }
cash_backing:
rules:
- { when: "is_null(cash_conversion)", score: 0, flag: no_conversion }
- { when: "cash_conversion >= $conversion_strong", score: 2 }
- { when: "cash_conversion >= $conversion_ok", score: 1 }
- { else: 0 }
balance_sheet:
rules:
- { when: "is_null(net_debt_ebitda)", score: 0, flag: no_leverage }
- { when: "net_debt_ebitda <= 0", score: 2 }
- { when: "net_debt_ebitda <= $leverage_max", score: 1 }
- { else: 0 }
valuation:
rules:
- { when: "is_null(fcf_yield) or fcf_ttm <= 0",
score: 0, flag: no_positive_fcf }
- { when: "fcf_yield >= $yield_strong", score: 2 }
- { when: "fcf_yield >= $yield_ok", score: 1 }
- { else: 0 }
gate:
mode: score
min_score: 4
The proposed thresholds are starting assumptions. The structure is the point. No family can flood the total with a pile of near-duplicate ratios, and missing data earns no accidental credit. The score can still reveal combinations: strong operations with a stretched price, or modest operations supported by cash and low leverage.
You can choose a stricter arrangement. Make leverage a hard gate if the thesis claims debt is a dealbreaker. Combine return on equity and margin in one criterion if both are necessary to define operating quality. What you shouldn't do is let several formulas win separate points while all of them depend on the same claim being true.
Does the backtest reward the idea or the duplication?
Once the structure is explicit, remove one family and re-run the backtest. This is an ablation test: take away one component to see what work it was doing. If deleting a supposed source of independent evidence barely changes which companies qualify, that family may be redundant. If the result changes sharply, inspect the names and periods that moved before deciding the family was valuable.
Change one family at a time. A backtest is hypothetical and excludes trading costs, so its job here is comparison: did the conclusion survive a change that shouldn't destroy the underlying claim? The practical method for perturbing thresholds and time windows applies to criterion families too. For a broader treatment of what separates a trustworthy backtest from a fragile one, see what makes a backtest honest.
Keep a record of every version you tried. Bailey, Borwein, López de Prado, and Zhu call the selection of a winner from many tested configurations backtest overfitting, and show why the risk rises as more alternatives are tried (Bailey et al.'s paper on the probability of backtest overfitting). Harvey, Liu, and Zhu reach the related problem in published factor research: after extensive factor searching, the usual statistical hurdle admits too many false discoveries (Harvey, Liu, and Zhu's NBER paper on multiple testing and factor discovery).
A duplicate-heavy score creates more dials to search. Drop this ratio, add that one, change each threshold, keep the version with the best curve. Version history preserves the path you took, including the runs you'd prefer to forget. Use it as a research log.
What should survive the audit?
A finished thesis doesn't need every sensible ratio. It needs a small set of claims that can challenge one another.
For each scored criterion, write its claim, its raw inputs, and the failure mode it catches. Group shared inputs. Delete exact transformations. Demote useful-but-overlapping ratios to diagnostics. Then remove each remaining family in turn and see whether the thesis still behaves like the idea you wrote down.
Quantery makes those choices editable because there isn't a universal answer. A debt ratio may be a gate in one thesis and context in another. Cash conversion may carry the quality claim by itself, or it may guard a profitability score. The rules are yours, but so is every duplicate vote you leave in them.
Research tooling, not investment advice. Nothing here is a recommendation to buy, sell, or hold any security. Screens, scores, and backtests are informational only; backtested results are hypothetical, exclude costs such as commissions and slippage, and do not guarantee future results. Verify against primary filings and make your own decisions.
Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.
← All articles