DEV Community

Cover image for Stop Paying for Search APIs: How I Ended Up Building Around Jiro
Adarsh Kushwah
Adarsh Kushwah

Posted on AI-assisted

Stop Paying for Search APIs: How I Ended Up Building Around Jiro

I have a graveyard of abandoned AI agent prototypes. The common thread? They all died at the search layer.

You know the drill. You build a beautiful RAG pipeline. You wire up the tools. You spend a weekend on prompt engineering. Then you hit a wall: your agent needs to know what happened on the web today, not just what was in a static PDF from 2023. You look at the usual suspects — SerpAPI, Tavily, Exa, Brave Search — and you realize that real-time grounding is either expensive at scale or a compliance nightmare. The alternative is rolling your own scraper, which lasts about forty-eight hours before Google starts serving you CAPTCHAs.

So I started looking for a middle ground. A self-hosted search layer that could actually survive contact with production. That is how I stumbled into Jiro.

What Jiro Actually Is (and Isn’t)

Let me be clear about the framing because the repository description does a lot of heavy lifting, and I almost scrolled past it.

Jiro is not a managed cloud service trying to lock you into a per-query billing model. It is an open-source, self-hosted search API and scraping platform that aggregates results from nine search engines and twelve social platforms, with a built-in MCP server and hybrid search ranking. You run it on your own infrastructure. It is MIT licensed. You can pip install jirosearch and have a local search endpoint running in under a minute.

The “free SerpAPI alternative” tagline is accurate, but it undersells what the project actually does. This is not just a proxy for Google results. It is a search intelligence layer designed specifically for AI agents and LLM tool use.

The Architecture That Made Me Stay

I spend a lot of time reading source code before I trust a tool in an agent loop. The first thing that caught my attention in Jiro’s repo is the directory structure. This is not a script someone threw together over a weekend.

The project separates concerns in a way that makes sense for a production search system:

search/ handles hybrid ranking, reranking, and multi-query expansion

scraping/ runs the nine engine adapters and the twelve social platform scrapers

ai/ manages LLM integration for answer synthesis

mcp.py exposes sixteen MCP tools out of the box

stealth.py handles TLS/JA3 fingerprint rotation and anti-bot bypass

You can see the full architecture breakdown in the repo. The short version is this: when my agent asks for a search, it does not get a flat list of links. It gets hybrid-ranked results that combine keyword signals with semantic freshness, along with query-aware snippet extraction and optional answer synthesis.

MCP Integration: The Real Unlock

If you are building with Claude Desktop, Cursor, or any other MCP-compatible client, this is where Jiro becomes genuinely useful. The project ships with a built-in MCP server. You drop a single configuration block into your client, and your agent immediately gains access to sixteen tools:

json
{
  "mcpServers": {
    "jiro": {
      "command": "jiro",
      "args": ["mcp"]
    }
  }
}

Enter fullscreen mode Exit fullscreen mode

The free tier alone gives you search, scrape, compare_engines, smart_classify, and monitor_status. The Pro tier unlocks ai_search (AI research with citations), search_hybrid, search_structured, and the full social scraping suite.

What this means in practice: I can ask Claude to “find the latest benchmarks for agentic RAG and summarize the key takeaways.” Claude calls jiro.ai_search, Jiro runs the hybrid search, scrapes the top results, synthesizes an answer, and returns it with citations — all locally, without my data leaving my machine.

The Social Scraping Layer Is Underrated

Most search APIs stop at web results. Jiro handles twelve social platforms: Reddit, Twitter/X, YouTube, LinkedIn, TikTok, Instagram, Facebook, Threads, Hacker News, Bluesky, Telegram, and Pinterest.

This matters more than it sounds. If you are building a monitoring agent for a product launch, or a research agent tracking sentiment around a topic, you need social data. The fact that it is baked into the same API surface as web search means you are not stitching together three different services and normalizing three different response schemas.

Code Walkthrough: Search in 30 Seconds

Here is the part I actually care about. Does it work without a PhD in infrastructure?

bash

Install

pip install jirosearch
Enter fullscreen mode Exit fullscreen mode

Start the server

jiro serve
Enter fullscreen mode Exit fullscreen mode

Search

curl -X POST http://localhost:8000/search \
  -H "Content-Type: application/json" \
  -d '{"q": "latest AI research", "engine": "google"}'
That is it. You now have a local search endpoint with caching, hybrid ranking, and structured extraction.

Enter fullscreen mode Exit fullscreen mode

The Python SDK is equally straightforward:

python

from jiro_sdk import JiroClient

client = JiroClient(api_key="your-key")
results = client.search("python web scraping")
content = client.scrape("https://example.com")
answer = client.ai_ask("What is Python?")
There is also a WebSocket streaming API if you need real-time results, and parallel search across multiple engines with a single call.
Enter fullscreen mode Exit fullscreen mode

The Self-Hosted Economics

Let me put the numbers on the table because this is the part that made me close the other tabs.

SerpAPI’s paid tier starts around $50/month for a few thousand searches. ScraperAPI gives you 5,000 requests. Bright Data is enterprise pricing. If you are running an agent that searches once per task, or a monitoring loop that polls every five minutes, you blow through those quotas fast.

Jiro’s self-hosted Pro tier is a one-time ₹4,999 purchase. It gives you 500 requests per minute, 100,000 requests per day, AI search, and commercial use rights. There is no per-query billing. There is no cloud lock-in. You are running the infrastructure, and you own the results.

The free self-hosted tier gives you 100 RPM and 10,000 requests per day, which is more than enough for personal projects and development environments.

The Honest Caveats
I am not going to pretend this is a perfect drop-in replacement for every use case. If you need a fully managed service with zero ops overhead, SerpAPI is still the easier choice. If you are scraping at a scale where even self-hosted rate limits are a bottleneck, you will need to think about horizontal scaling.

But if you are an AI engineer building agents that need real-time web grounding, social intelligence, and structured data extraction without handing your budget to a SaaS provider, Jiro is the most complete open-source option I have found this year.

The repository is at github.com/DevAnimecx/jiro. The documentation is at searchjiro.vercel.app/docs. The MIT license means you can fork it, audit it, and ship it in commercial products without asking for permission.

If you are tired of watching your agent’s search bill climb every time you test a new retrieval strategy, give it a look. Your infrastructure, your rules, your data. That is the pitch, and for once, it actually holds up.

Top comments (0)