How to test whether a stock score means anything
A stock score means something only if demanding more points produces a pattern you can explain and test. In Quantery, save a series of thesis versions with progressively stricter minimum scores, then backtest each version over the same dates. A higher gate should generally improve the evidence if another point really measures more of what your thesis claims to reward.
That sequence is a score ladder: the ordered relationship between the minimum score required and what happened afterward. Quantery reports only the companies that pass each gate, so this is a ladder of cumulative groups, not a report that places every observation into a separate score bucket. It doesn't need to rise perfectly at every step. Markets aren't that tidy. But if a stricter gate makes the result worse, or one narrow setting carries the whole result, the total score hasn't earned the confidence its precision suggests.
A total score is a claim about order
Suppose a thesis awards points for operating quality, cash backing, leverage, and valuation. A company can finish with a total near the bottom, middle, or top of the scale. The score makes two claims at once.
The first is a classification claim: companies above the gate are suitable for further research under the stated rules. The second is an ordering claim: a higher total represents more of the attributes the thesis is meant to capture. That second claim often slips by untested. We look at the companies that pass, compare their hypothetical return with a benchmark, and never ask what the lower scores did.
A gate can look successful even when the score underneath it is nonsense. Move enough cutoffs and one boundary may flatter the sample without showing that each added point carried useful information.
Read several gates instead. If outcomes improve as the minimum score rises, the score has evidence of order. If results jump only at one cutoff, treat that cutoff as a suspect parameter. The same applies when a score is used purely for ranking: the order matters more than the label attached to its upper tail.
This is why a composite score must expose its components. Two companies can earn the same total through different routes. Before treating the total as one continuous measure, make sure those routes produce a relationship worth compressing.
Build the gate ladder before looking at returns
Define the analysis before you open the performance report. Choose the outcome horizon and benchmark. Then write down the minimum scores you'll test.
For a compact integer score, create one thesis version per sensible gate. A four-point score might be tested with min_score set to 1, then 2, then 3, with everything else unchanged. Run each version over identical dates with the same rebalance cadence. Quantery records the thesis version behind every backtest, so the reports preserve which cutoff produced each result.
This doesn't isolate the exact-score groups. The run with a minimum of two contains companies scoring two or better, while the run with a minimum of three contains companies scoring three or better. The groups overlap by design. What you can test in the shipped app is whether demanding another point changes the portfolio and gate-entry event study in a consistent direction.
Keep the outcome identical across the ladder. A clean first pass might compare hypothetical excess return over the same benchmark after one fixed holding period. Excess return means a security's return minus the benchmark's return over matching dates. The backtest reports the event count beside the mean and median excess returns. Read the count. A striking result from a thin gate should prompt another test before any conclusion. These results exclude trading costs.
Don't use a short horizon for the loose gate and a long one for the strict gate. Don't switch benchmarks when one comparison looks awkward. The universe stays fixed too. A score of three among large industrial companies isn't comparable with the same total from a later test that admits banks and tiny listings. Every rank inherits its screening universe, so preserve the universe definition across versions.
Read shape before the winning endpoint
Across the reports, inspect the portfolio return and maximum drawdown (the largest peak-to-trough loss during the test period). Then read the event study's mean and median excess returns, hit rate, and event count at your chosen horizon. The median matters because one spectacular event can pull an average upward. The hit rate asks whether the apparent advantage was broad or concentrated.
Now draw the gate ladder. There are several useful shapes.
Broadly improving: stricter gates tend to improve the outcome without collapsing the event count. A wobble between neighbors is ordinary. The score deserves another test.
Flat: the reports look alike as the gate rises. The points may summarize accounting facts without separating future outcomes. That's still useful for organizing research, but the backtest doesn't support using the total as a stronger-versus-weaker prediction.
Peaked: one middle gate fares better than both looser and stricter versions. The strict run may contain expensive quality or a sector cluster. It may just have too few events. Inspect its go-event list before rewriting the threshold.
Cliff-shaped: outcomes change only at one boundary. This can happen when one criterion dominates the total or when the gate was tuned to the sample. Shift the boundary by one point. If the result disappears, the cliff was the thesis.
Reversed: stricter gates fare worse. Don't rescue the model by flipping its interpretation after seeing the result. First check sign errors and null handling. Then inspect sector composition and the economic story. A clean failure saves time.
None of these shapes proves causality. An improving ladder can be driven by size, industry, valuation, or one historical regime. It does show whether the score's claimed ordering exists in the sample. That's more demanding than asking whether one gate finished ahead.
Test the points inside the total
A gate ladder can improve for the wrong reason. If the same accounting fact earns several points, the total may just be a louder version of one signal. Counting the same signal twice explains how shared inputs create duplicate votes. Follow the ladder with removal tests.
Take away one criterion family and rerun the same gates. This is an ablation test: remove one component to see what work it was doing. If deleting leverage changes little, it may add description without adding separation. If deleting cash backing destroys the ladder, that family was carrying the result. Either way, you've learned what claim the composite is making.
Companies can reach the same total by different routes. One may get its points from quality and cash, another from valuation and leverage. Quantery's go-event list records each event's total score, while the thesis version preserves the component rules. Inspect the names around a gate that changes sharply. If one route dominates, turn its decisive family into a gate. You could instead cap duplicate families or keep the route visible as an annotation.
Missing data needs its own audit. A null must never earn a favorable point by falling through an unguarded rule. Rejecting every incomplete company can also change who reaches each gate. If the loosest version mostly admits missing-data cases, its poor outcome doesn't validate the economic thesis.
Write the score so the test stays legible
The illustrative shape below keeps criterion families separate and gives each one a limited vote. The bundled templates are the reference for exact fields, and the Quantery thesis documentation covers the full DSL.
# A score with four distinct claims: illustrative
params:
roe_floor: 0.12
conversion_floor: 0.80
leverage_max: 2.0
yield_floor: 0.04
features:
ni_ttm: ttm(net_income)
fcf_ttm: ttm(free_cash_flow)
equity_now: newest(total_equity)
debt_now: newest(total_debt)
cash_now: newest(cash)
ebitda_ttm: ttm(ebitda)
roe: if(equity_now > 0, ni_ttm / equity_now, null)
conversion: if(ni_ttm > 0, fcf_ttm / ni_ttm, null)
net_debt_ebitda: if(ebitda_ttm > 0,
(debt_now - cash_now) / ebitda_ttm, null)
fcf_yield: if(market_cap > 0, fcf_ttm / market_cap, null)
criteria:
quality:
rules:
- { when: "is_null(roe)", score: 0, flag: no_roe }
- { when: "roe >= $roe_floor", score: 1 }
- { else: 0 }
cash_backing:
rules:
- { when: "is_null(conversion)", score: 0, flag: no_conversion }
- { when: "conversion >= $conversion_floor", score: 1 }
- { else: 0 }
balance_sheet:
rules:
- { when: "is_null(net_debt_ebitda)", score: 0, flag: no_leverage }
- { when: "net_debt_ebitda <= $leverage_max", score: 1 }
- { else: 0 }
valuation:
rules:
- { when: "is_null(fcf_yield)", score: 0, flag: no_yield }
- { when: "fcf_yield >= $yield_floor", score: 1 }
- { else: 0 }
gate:
mode: score
min_score: 3
This structure produces an inspectable total without claiming that one ratio deserves three different names. The proposed thresholds are starting assumptions. Save copies with progressively stricter min_score values, then run the same historical test on each copy. Quantery's backtest will evaluate only the names that clear the active gate, which is exactly why each cutoff needs its own version and report.
Use those versions as a research log. First run the declared score. Then change one thing. Start by moving the gate for the initial ladder. A later ablation removes one family. Alter the test window only after you have preserved both earlier states. Perturb one dial at a time or you won't know what repaired or broke the result.
A good ladder can still be overfit
After enough configurations have been tried across cutoffs, test windows, and which criteria to include, somebody will find an improving sequence. That is backtest overfitting in another outfit. Bailey, Borwein, López de Prado, and Zhu show why trying multiple strategy configurations can produce impressive simulated performance that fails out of sample (their paper on backtest overfitting). Declare the test first and preserve every attempt. Once you reach the holdout period, stop changing the score.
Re-run the same gate ladder in another period. Change monthly rebalancing to weekly only if cadence is part of the robustness plan. Tighten the universe to test whether one size or sector group carried the order. Use a benchmark matched to the opportunity set. Every return remains hypothetical and excludes costs. A strict gate full of thinly traded names can lose its apparent advantage in implementation.
A score that survives doesn't become a forecast. It becomes a better research instrument.
That is enough. Keep the total only when progressively stricter gates support its ordering. If the sequence is flat, open the rules. If one duplicate vote carries it, change the structure. The score is your compressed claim, and each saved version gets to grade it.
Research tooling, not investment advice. Nothing here is a recommendation to buy, sell, or hold any security. Screens, scores, and backtests are informational only; backtested results are hypothetical, exclude costs such as commissions and slippage, and do not guarantee future results. Verify against primary filings and make your own decisions.
Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.
← All articles