DEV Community

Sravya Dangeti
Sravya Dangeti

Posted on Originally published at github.com

Firecrawl vs Tavily for an AI agent: 97% vs 82% fact recall, and why retrieval breadth decides grounding

I run a competitor-research agent: give it a company URL and it finds and verifies that company's competitors. It is framework-free (a raw tool-calling loop on Gemini), and it has a deterministic fabrication check: every page the agent cites must be in the ledger of pages my code actually fetched.

Retrieval was Tavily only. I added Firecrawl as a second backend behind a provider switch and benchmarked both. Here is what I found, including where Firecrawl lost.

How I measured

  • 10 companies, 20 pages, chosen for different site types: JS-heavy apps, docs-heavy developer tools, pricing pages, sparse early-stage startup sites.
  • An answer key of facts checked against each official page (pricing tiers, product names, taglines, YC batch). Every fact was audited against the raw HTML with a plain HTTP request, neither provider involved. Corrections are logged in the repo.
  • Two experiments:
    • A, retrieval only: fetch each page twice, no LLM involved, and score which facts made it into the content.
    • B, end to end: run the full agent with model, prompts and temperature held constant.

Experiment A: retrieval quality

Tavily (basic, as deployed) Firecrawl
Fetch success 100% 100%
Fact recall 82% (56/68) 97% (66/68)
Latency p50 250 ms 1,391 ms
Median content 11,735 chars 16,257 chars

Firecrawl found more of the facts that matter. It was about 5.6x slower, and by published pricing about 5x the credit cost per page.

Tavily's advanced extract tier did not help: it returned byte-identical content to basic on 19 of 20 pages, and on one page (docs.firecrawl.dev/billing) it returned a 404 where basic succeeded. I reproduced that with a raw API call outside my code.

The boilerplate filter has a real cost

Firecrawl's only_main_content=True strips navigation and footers. I measured it on and off:

Page Boilerplate (on / off) Facts found (on / off)
brickanta.com 12% / 19% 1/1 / 1/1
composio.dev 10% / 16% 2/2 / 2/2
composio.dev/pricing 4% / 18% 2/2 / 2/2
notion.com 13% / 45% 0/1 / 1/1

It was free on three pages. On notion.com it removed the product names in the navigation menu, which was the one fact I needed. Getting it back meant 45% boilerplate. Usually free, occasionally expensive.

Experiment B: retrieval breadth decides grounding

The surprise came from the end-to-end runs. Pooling all modes by how many sources the agent consulted:

Sources consulted Runs Citations rejected Per run
10 or fewer 9 48 5.3
More than 10 42 12 0.3

Every low-source run came from one configuration of mine: Firecrawl search with full-page scraping, capped at 3 results per search to save credits. With fewer sources, the agent searched more, hit its 9-call budget more often, and cited pages it had never retrieved.

The fabrication check caught every one of those citations. None reached a final report.

The lesson: starve an agent of sources and it cites pages it never read. That looks like a hallucination problem. It is a retrieval problem.

One caveat I want to be clear about: all 9 low-source runs used that same 3-result configuration, so breadth and configuration cannot be separated in this data. It shows an association, not isolated causation, and it is a finding about my settings, not about Firecrawl's quality.

Where Firecrawl lost

  • About 5.6x slower at p50.
  • More credits per page and per end-to-end run.
  • On brickanta.com it dropped the hero heading and paragraph while keeping YouTube embed chrome, reproducibly.
  • Worse boilerplate outliers (23% max vs 15%).
  • The Notion navigation trade-off above.

Also worth knowing

Running Tavily search with Firecrawl scraping (my "hybrid" mode) cost the same as pure Tavily and gave no measurable benefit.

What I would tell someone choosing

  • If fact coverage on messy real-world pages matters most, Firecrawl is worth the latency and cost.
  • If you need speed and low cost, Tavily basic is strong.
  • Either way, give your agent enough sources per search. Breadth mattered more for grounding than the vendor did.

Limitations

  • 10 companies, 20 pages: a small sample.
  • End-to-end runs span two days, so sites may have changed in between.
  • Some runs were excluded because of Gemini free-tier quota limits. The exclusions are counted per mode in the results.
  • Tavily's cost comparison uses published pricing, not measured usage.
  • Answer-key wording was corrected after an audit against raw HTML, before the final run. Every change is logged.
  • I scraped firecrawl.dev with Firecrawl itself.

Reproduce it

Method, raw results, answer key and change log are all in the repo:
https://github.com/sravya520/competitor-research-agent/blob/main/docs/firecrawl_vs_tavily.md

Feedback welcome, especially from anyone who has tuned Firecrawl or Tavily settings for agent workloads.

Top comments (0)