Finance AI agents have a basic reliability problem: language models can produce plausible answers, including plausible numbers. Plausible is not enough when a user asks what a company reported, when the information became public, or whether the figure was later restated.
One practical approach is a grounding layer built from dated financial facts.
For developers building retrieval-augmented generation systems, filing copilots, screening agents, research assistants, or automated company briefs, point-in-time fundamentals can provide a structured bridge between SEC filings and model-generated explanations. The model can reason over the data, but it should not invent the data.
Tradevo Data provides point-in-time US equity fundamentals sourced from SEC EDGAR. Its annual dataset covers 5,168 US companies, 632,466 point-in-time rows, 16 annual concepts, and up to 12 fiscal years. Its quarterly dataset contains 1,089,635 rows across 5,382 companies and seven quarterly concepts. Coverage varies by filer.
Why ordinary fundamentals can mislead an AI agent
A conventional fundamentals database may emphasize the latest known value. That can be useful for current analysis, but it may be inappropriate for a historical question.
Suppose an agent is asked:
What revenue and net income were publicly available for a company on June 15 of a past year?
If retrieval returns a figure restated months later, the answer contains information that was unavailable on June 15. The number may now be technically correct while still being historically invalid for that date.
That distinction matters for:
- Historical research and event reconstruction
- Backtests using fundamental signals
- Filing-aware RAG systems
- Agents answering “as of” questions
- Audit trails for generated financial summaries
- Comparing first-reported figures with subsequent revisions
This is lookahead bias at the data layer. A longer explanation is available at https://tradevodata.com/blog/lookahead-bias-fundamental-backtests?utm_content=blog-grounding-a-finance-llm-in-as-reported-fundamentals&utm_source=devto&utm_medium=syndication&utm_campaign=syndicate.
The fields an agent needs for defensible answers
A useful finance grounding record should do more than return a value. It should establish provenance.
Each Tradevo Data row includes:
-
first_filed: the date the value became public -
original_value: the first-reported, point-in-time-safe value -
latest_value: the current revision -
restated: a flag for a greater-than-0.5% change under the same XBRL tag, including amendments -
qa_status: a quality-assurance status
The dataset labels 37,804 restatements.
For an AI agent, those fields support a stronger response pattern:
- Retrieve facts available by the requested date.
- Use
original_valuefor an as-reported historical answer. - Compare it with
latest_valueif revisions matter. - Include
first_filedand the requested concept in the response. - State uncertainty or coverage gaps instead of filling them with generated numbers.
An XBRL tag is not a complete semantic guarantee: filers can use extensions, presentation varies, and accounting context still matters. But a dated record containing the original value, revised value, restatement flag, and QA status is more auditable than an unsupported number in model output.
Example query for a finance AI agent
Tradevo Data exposes one JSON fundamentals endpoint with server-side point-in-time filtering:
GET https://tradevodata.com/v1/fundamentals?ticker=AAPL&as_of=2020-06-15&concept=Revenue&period=annual&utm_content=blog-grounding-a-finance-llm-in-as-reported-fundamentals&utm_source=devto&utm_medium=syndication&utm_campaign=syndicate
The server applies:
first_filed <= as_of
The date is inclusive. Annual data is the default, while period=quarterly requests quarterly rows.
A basic agent tool contract might look like this:
{
"tool": "get_fundamentals",
"arguments": {
"ticker": "AAPL",
"as_of": "2020-06-15",
"concept": "Revenue",
"period": "annual"
}
}
The model should then compose an answer from returned records rather than relying on memorized parameters. A system instruction can require it to include the fiscal period, first_filed, original_value, and concept in every numerical citation.
For example:
Use only retrieved values. If no qualifying row exists by the requested
as-of date, say that the dataset returned no supported value. Do not infer
or interpolate a financial figure.
That final rule is important. Retrieval does not prevent hallucination unless the agent is explicitly required to abstain when evidence is missing.
A practical RAG architecture
A filing agent can separate structured facts from unstructured context:
- Use point-in-time fundamentals for supported numerical claims.
- Use filing text retrieval for management commentary, accounting policies, and risk disclosures.
- Store source metadata alongside every chunk or fact.
- Require citations in the final answer.
- Validate numerical statements against retrieved structured records before returning them.
Structured fundamentals are not a replacement for the filing. They can serve as a compact numerical layer alongside it.
For bulk workflows, the $29-per-month Pro plan includes /v1/download and /v1/snapshot?as_of, each supporting period=annual|quarterly. These endpoints can support a local retrieval store or reproducible historical snapshot. Parquet format is not included.
More background on the data model is available at https://tradevodata.com/blog/point-in-time-fundamentals-data?utm_content=blog-grounding-a-finance-llm-in-as-reported-fundamentals&utm_source=devto&utm_medium=syndication&utm_campaign=syndicate.
What Tradevo Data covers—and what it does not
The annual dataset contains these 16 concepts:
- Revenue
- NetIncome
- Assets
- StockholdersEquity
- OperatingCashFlow
- EPSDiluted
- DilutedShares
- GrossProfit
- OperatingIncome
- PretaxIncome
- IncomeTaxExpense
- CapitalExpenditures
- CashAndCashEquivalents
- CurrentAssets
- CurrentLiabilities
- NetPPE
Quarterly coverage includes seven concepts. Q4 is reported where tagged or derived and labelled where supported. Derived Q4 EPS and share figures are not provided.
The service does not provide TTM calculations, non-US coverage, delisted-company coverage, or Parquet output. It is not presented as a full filing-text corpus, market-data feed, estimates database, or accounting ontology.
These limits matter. An agent should not present absence as zero, assume every filer reports every concept, or silently substitute one accounting concept for another.
Fair comparison for finance-agent developers
| Option | Best fit | Potential strengths | Important considerations |
|---|---|---|---|
| Tradevo Data | Developers wanting a budget-tier API and bulk access for core US fundamentals | Server-side as_of, original and latest values, filing dates, restatement labels, and QA status |
Limited concept set; no delisted, TTM, non-US, or Parquet coverage |
| Raw SEC EDGAR/XBRL | Teams needing maximum control and direct primary-source processing | Public-domain source material and direct control over interpretation | Requires ingestion, taxonomy mapping, duplicate handling, amendment processing, QA, and point-in-time logic |
| Sharadar | Researchers evaluating a broader established data product | A credible alternative with its own coverage and data model | Confirm current scope and licensing; see their pricing page: https://data.nasdaq.com/databases/SF1 |
| Tiingo | Developers evaluating fundamentals alongside other financial APIs | A credible provider with a broader product context | Verify point-in-time semantics and coverage; see their pricing page: https://www.tiingo.com/pricing/overview |
| QuantConnect | Teams building inside an integrated research and execution platform | Data access within a broader algorithmic environment | Platform fit may matter more than a standalone fundamentals API; see their pricing page: https://www.quantconnect.com/pricing/ |
No provider should be chosen from a feature label alone. Test amendment handling, filing-date semantics, survivorship assumptions, missing values, concept normalization, redistribution rights, and historical coverage against your actual agent workflow.
When another option wins
Another provider can be the better choice when you need a broader security universe, delisted companies, non-US issuers, more accounting concepts, analyst estimates, TTM figures, market data, or a managed research platform. A specialist vendor may also win when your institution needs enterprise support, contractual service levels, or licensing terms beyond a developer-oriented product.
If your workflow already runs inside QuantConnect, its integrated environment may reduce engineering work. If you need a broader commercial dataset, Sharadar or Tiingo may fit better after you verify current coverage and point-in-time behavior.
Tradevo Data is positioned as an honest budget tier of research-grade point-in-time data, not as the only affordable option and not as a universal financial-data layer.
When to build it yourself
Build directly from SEC EDGAR when provenance control is more important than implementation speed, your required concepts fall outside the supported set, or your team has specialized accounting and data-engineering expertise.
Be prepared to handle:
- Filing and amendment chronology
- XBRL standard tags and company extensions
- Units, periods, contexts, and duplicate facts
- Fiscal-year differences
- Restatements and comparable-tag logic
- Quarterly versus year-to-date facts
- Q4 derivation rules
- Missing and malformed filings
- Reprocessing when extraction rules change
- Reproducible historical snapshots
EDGAR is public domain, so self-building can be rational. The practical burden is maintaining consistent interpretation and QA over time.
Inspect the evidence before connecting an agent
A public proof pack is available at https://github.com/christianpichichero-max/pit-fundamentals. It contains five companies, their latest three fiscal years, 225 rows, and the full methodology with no signup.
On reliable-filing rows in that five-company proof pack only, measured lookahead averaged 35.2 days and reached a maximum of 48 days. Those measurements describe the proof pack, not the full dataset.
For API evaluation, https://tradevodata.com/?utm_content=blog-grounding-a-finance-llm-in-as-reported-fundamentals&utm_source=devto&utm_medium=syndication&utm_campaign=syndicate offers a card-backed seven-day free trial covering 10 companies, a rolling three-year history, and 100 requests per day without bulk. Unless canceled, it renews at $29 per month. Pro provides complete available history, 5,000 requests per day, and bulk endpoints. Documentation is at https://tradevodata.com/docs?utm_content=blog-grounding-a-finance-llm-in-as-reported-fundamentals&utm_source=devto&utm_medium=syndication&utm_campaign=syndicate.
Start with the public sample, test the dates and concepts against source filings, and make abstention part of the agent design. The goal is not to make a model sound certain. It is to make supported financial numbers traceable.
Not investment advice.
Top comments (0)