How to check data coverage before a backtest
Before you backtest a thesis, check whether your local data can evaluate its rules across the dates you plan to test. In Quantery, open representative companies from a scan and read Data coverage: the price range and filing periods stored, plus gaps and collection attempts. Then refresh the missing inputs or choose a test window that fits the evidence.
This matters because a declared universe and an evaluable universe can be different. Quantery's backtest documentation says a company is included only on dates with visible fundamentals and fresh prices. If its price history begins halfway through your test, it enters halfway through. If a required filing value is absent, a feature can become null and the company can drop out. The calculation may be correct while the research question has narrowed without you noticing.
Data coverage is part of the test
A thesis declares which companies are eligible. Its exchange scope, market-cap floor, sector exclusions, and any named universe define the outer boundary. Coverage draws another boundary inside it: which eligible companies can be evaluated on each date.
Call that the effective universe. It changes when a company lacks a required input, when a filing is too stale for the thesis, or when the local price history doesn't reach the decision date. That isn't a clerical detail. The companies lost to missing data can share traits such as a recent listing, an unusual reporting pattern, or thin vendor coverage. Your result then describes the companies left behind.
This is the same problem covered in how to choose the universe your screen runs on, viewed from the data side. Writing “US equities” in the thesis doesn't make every US equity evaluable. The lake has to support each feature at each decision date.
Start with a current scan because it gives you concrete companies to inspect. Pick an expected survivor and a result that surprised you. If your thesis spans very different industries or company sizes, include examples from those edges. You aren't estimating a formal coverage rate from a few names. You're looking for reasons the planned test could be malformed before you spend time interpreting it.
Start with the inputs your thesis requires
Don't audit every dataset in the card. Audit the ones the thesis uses.
A filing-based quality thesis might need net income, operating cash flow, assets, debt, current assets, current liabilities, and share count. A valuation feature may add a current price or market capitalization. A momentum rule needs enough daily prices to cover its lookback. Short interest doesn't matter to that thesis if no rule references it.
Make a small dependency list from the thesis before opening any company. For each feature, write down the raw fields it reads, how far back it looks, and what happens when an input is null. This is where auditing a stock screen result helps: follow a feature down to its periods and sources instead of treating the finished ratio as an independent fact.
Pay special attention to comparisons. A current value may need only the latest filing. A change signal also needs its comparison period. A trailing calculation needs a run of usable flow data. A price drawdown needs the relevant high as well as today's close. One populated cell on a company page doesn't prove that the historical feature exists on every rebalance date.
Null policy belongs on the dependency list too. Does a missing feature fail the gate, earn no score, or make the entire result unevaluable? Those choices change who remains in the effective universe. If you can't state the policy, the thesis isn't ready for a backtest.
Read the collection status before the dates
The status sentence tells you what the app knows about an apparent blank. Read it before diagnosing the gap.
Last collected means the lake contains rows and records when collection last succeeded. That date is evidence about the pull. It isn't a claim that the dataset extends through that same date, so read the displayed range and freshness beside it.
No 10-K or 10-Q data found means the SEC fundamentals check completed without finding the filing types the pipeline uses. Last attempt failed means collection encountered an error. Not collected yet means the app hasn't tried. The data coverage documentation spells out those distinctions because each one calls for a different response.
A blank isn't one kind of blank. Retry a failed attempt after checking the connection or provider. Run the relevant collection job when nothing has been attempted. When the SEC check found no supported filings, inspect the issuer on SEC EDGAR before assuming repeated refreshes will fix it. The issuer may report under a different filing regime, or the identifier may need investigation.
The same restraint applies when a dataset has rows. “Collected” doesn't certify completeness. It says rows exist. The range tells you the span. Cadence and gaps tell you whether the feature can use it.
Check depth, gaps, and cadence separately
For prices, compare the first and last stored dates with the window your feature and backtest require. Then read freshness and any listed gaps. A series can be current at the right edge and still be too short at the left edge. It can also span the whole window with an interruption in the middle.
For fundamentals, compare the first and last fiscal periods, the latest reported date, and the missing-period list. Check whether cash-flow data is present when the thesis depends on it. A balance-sheet row alone can't supply operating cash flow. Annual cadence also needs different expectations from quarterly cadence; expected gaps shouldn't be treated as failed collection.
Next, open Data and filings for a suspicious company. The coverage card summarizes the shape of the local record. The underlying rows show the fiscal periods, filing dates, sources, and EDGAR links that the feature can use. Filing date matters because a backtest must wait until information was public. Period end and filing date answer different questions.
A missing period doesn't always make every feature unusable. It may break a year-over-year comparison while leaving a direct balance-sheet value intact. It may leave a trailing calculation with fewer usable inputs than you intended. Follow the exact feature. Don't convert “there is a gap” into “the whole company has no data.”
This preflight still can't certify the whole market by hand. After the run, read the report's coverage note and any warning. If the effective universe thins sharply in an early segment, the report is telling you that the attractive curve rests on a smaller opportunity set there. Stop interpreting returns until you've resolved that mismatch.
Fix the smallest missing slice
If collection hasn't run or a recent pull failed, use the app's data controls to run a targeted price or fundamentals job for the affected tickers. Keep the scope narrow at first. A small job lets you confirm the source and the resulting range before a broad refresh changes the lake underneath several experiments.
A refresh can retrieve available data. It can't create a filing that the issuer never submitted, and it can't exceed the historical depth supplied by your market-data plan. Price depth depends on that plan. If the required history still isn't available, shorten the test window, replace the feature with a defensible proxy, or make the missing-data rule explicit.
Don't “fix” coverage by filling unknown values with zero unless zero is economically true. Zero debt and unknown debt aren't interchangeable. Neither are flat revenue and missing revenue. A default that helps a company pass is evidence invented by the thesis itself.
Sometimes the right fix is a gate that refuses incomplete inputs. Sometimes it's a score that withholds a point while keeping the company visible. The choice depends on the claim. A solvency thesis may require the balance-sheet fields outright. A broad composite may tolerate one unavailable secondary signal if the output labels that limitation. Write the policy where you can inspect and retest it.
Rerun without moving the rules
After refreshing data or changing the test window, rerun the same saved thesis version with the same backtest settings. Record which dataset changed, which companies you refreshed, and whether the report's coverage warning changed. If you alter the rules at the same time, you won't know whether a different result came from better data or a different thesis.
Read coverage before performance. Then inspect membership and events. Only after that should you read returns, which remain hypothetical and exclude costs. What makes a backtest worth trusting depends on point-in-time alignment and the universe actually available, not the universe you meant to have.
Coverage work can make a backtest less impressive. Good. A shorter valid window beats a longer window padded with companies the engine couldn't evaluate. A strict null rule may remove a result you liked. Good again. The rules are yours, including the rule for what happens when the evidence runs out.
Research tooling, not investment advice. Nothing here is a recommendation to buy, sell, or hold any security. Screens, scores, and backtests are informational only; backtested results are hypothetical, exclude costs such as commissions and slippage, and do not guarantee future results. Verify against primary filings and make your own decisions.
Want to try this on your own rules? Quantery is free for 14 days: the full app, no card required.
← All articles