DEV Community

AgentSearch
AgentSearch

Posted on Originally published at agentsearchhq.com

Turn Any URL into Clean Markdown for RAG (API Guide)

Raw HTML makes poor context for an LLM. A typical page is mostly navigation, scripts, cookie banners and footer links. Putting all of that into an embedding model or a prompt wastes tokens and lowers retrieval quality. What you want is the content: headings, paragraphs, lists and links, as markdown, plus enough metadata to cite the source and to tell when it goes stale.

This guide uses AgentSearch Web Extract (agentsearch-web-extract-v1) to turn a URL into RAG-ready markdown. It covers the request, the response, chunking and error handling. Each call costs $0.005, paid in USDC over x402 or MPP on the Pocket Agentic Portal, with no account and no API key.

Why markdown is the right format for RAG

  • It keeps the structure. Headings (#, ##) tell you where sections begin, which gives you natural chunk boundaries and section titles to store as metadata.
  • It's token-efficient. Markdown carries the meaning of the HTML without the tags.
  • Links survive. [text](url) keeps the citation trail, so an agent can follow a link or quote its source.
  • Models read it well. LLMs see a lot of markdown and handle it reliably in prompts.

The request

The endpoint is POST /v1/extract, which takes a JSON body. Only url is required.

curl -X POST https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract \
  -H 'content-type: application/json' \
  -d '{"url":"https://example.com"}'
# → 402 Payment Required with the terms. An x402 or MPP client signs and retries.
Enter fullscreen mode Exit fullscreen mode

These are the options from the extract OpenAPI spec:

Field Type Default Notes
url string (required) An absolute http(s) URL
formats ["markdown"], ["text"] or both ["markdown"] Which renderings to return
max_chars integer, 1–100000 50000 Caps the returned characters. meta.truncated tells you whether the cap was hit
include_links boolean true Up to 100 absolute links from the page
timeout_ms integer 4000 Fetch deadline. Values above 4000 are capped (a hard 4 s deadline)

The response

This is a real response captured on the portal service page (2026-09-25 UTC):

{
  "portal": {
    "provenance": "third-party-supplier",
    "serviceId": "agentsearch-web-extract-v1",
    "schemaCheck": "unchecked"
  },
  "data": {
    "request_id": "05f5a4ebe2684c79867aa9bf3515c33d",
    "url": "https://example.com/",
    "title": "Example Domain",
    "markdown": "# Example Domain\n\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\n",
    "text": "",
    "links": [{ "href": "https://iana.org/domains/example", "text": "Learn more" }],
    "meta": {
      "status_code": 200,
      "content_type": "text/html",
      "fetched_at": "2026-09-25T15:53:07.843243+00:00",
      "chars": 167,
      "truncated": false
    },
    "error": null
  }
}
Enter fullscreen mode Exit fullscreen mode

The fields that matter for RAG:

  • url is the final URL after redirects. Use it as your canonical document ID.
  • title is a ready-made citation label.
  • meta.fetched_at lets you re-extract stale documents on a schedule.
  • meta.truncated tells you whether you hit max_chars. If you did, raise the cap and fetch again.
  • error is null on success.

Step 1: Fetch markdown from code (TypeScript + x402)

Standard x402 v2 clients work with the portal unchanged. Here's @x402/fetch with @x402/evm, paying on Base:

import { wrapFetchWithPaymentFromConfig } from "@x402/fetch";
import { ExactEvmScheme } from "@x402/evm";
import { privateKeyToAccount } from "viem/accounts";

const EXTRACT =
  "https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract";

// Use a dedicated wallet that holds only what you're willing to spend.
const account = privateKeyToAccount(process.env.AGENT_WALLET_KEY as `0x${string}`);
const pay = wrapFetchWithPaymentFromConfig(fetch, {
  schemes: [{ network: "eip155:8453", client: new ExactEvmScheme(account) }],
});

export async function extract(url: string, maxChars = 50000) {
  const res = await pay(EXTRACT, {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify({ url, formats: ["markdown"], max_chars: maxChars }),
  });

  if (!res.ok) {
    // Portal errors aren't wrapped in the envelope: { error: { code, message, retryable } }
    const body = await res.json().catch(() => ({}));
    throw new Error(`portal ${res.status}: ${body?.error?.code ?? "unknown"}`);
  }

  const envelope = await res.json();
  if (envelope.portal?.provenance !== "third-party-supplier") {
    throw new Error("Unexpected response shape; refusing to use it.");
  }
  const doc = envelope.data;
  if (doc.error) {
    // The target site failed, but the HTTP status was still 200.
    throw Object.assign(new Error(doc.error.code), { retryable: doc.error.retryable });
  }
  return doc as {
    url: string; title: string | null; markdown: string;
    links: { href: string; text: string }[];
    meta: { fetched_at: string; chars: number; truncated: boolean };
  };
}
Enter fullscreen mode Exit fullscreen mode

If delivery fails, the portal settles nothing, so you aren't charged for it. Payment details are in Pocket's integration guide.

Prefer no code? Use MCP

In Claude Desktop, Cursor or Claude Code, add npx -y @pocket-network/agentic-portal-mcp and ask the agent to call_service on agentsearch-web-extract-v1. The full setup, including spending limits, is in How to give your AI agent web search with MCP.

Step 2: Chunk by heading

Split on markdown headings first, then on size. Every chunk keeps the path of headings above it, which improves both retrieval and citations.

import re

def chunk_markdown(doc: dict, max_chars: int = 1500, overlap: int = 150):
    """doc is the `data` object returned by agentsearch-web-extract-v1."""
    chunks, path, buf = [], [], []

    def flush():
        text = "\n".join(buf).strip()
        if not text:
            return
        for i in range(0, len(text), max_chars - overlap):
            chunks.append({
                "text": text[i:i + max_chars],
                "source_url": doc["url"],
                "title": doc.get("title"),
                "section": " > ".join(path),
                "fetched_at": doc["meta"]["fetched_at"],
            })

    for line in doc["markdown"].splitlines():
        m = re.match(r"^(#{1,6})\s+(.*)", line)
        if m:
            flush(); buf.clear()
            level = len(m.group(1))
            path[:] = path[:level - 1] + [m.group(2).strip()]
        buf.append(line)
    flush()
    return chunks
Enter fullscreen mode Exit fullscreen mode

Embed chunk["text"] and store the rest as metadata. When your agent answers, cite title and source_url.

Step 3: Handle failures properly

Extract separates target-site failures from request problems:

Where HTTP Codes What to do
data.error (target site) 200 TARGET_HTTP_ERROR, TARGET_TIMEOUT, TARGET_DNS_ERROR, TARGET_CONNECT_ERROR, TARGET_FETCH_FAILED, TARGET_TOO_MANY_REDIRECTS, UNSUPPORTED_CONTENT Retry only when retryable is true (timeouts, connection failures, target 5xx/429)
Request 400 INVALID_REQUEST, SSRF_BLOCKED (private, loopback or metadata addresses) Fix the input. Don't retry
Request 413 / 429 REQUEST_TOO_LARGE, CAPACITY_LIMIT, APPLICATION_RATE_LIMITED Back off and retry

UNSUPPORTED_CONTENT means the URL points to something other than HTML or text, such as a PDF or an image. Send those to a document parser instead.

Step 4: Treat extracted content as untrusted

Any page can contain text written to manipulate your agent. Before content reaches a model:

  • Put extracted markdown in a data position, such as a tool result or a quoted block, never in a system prompt.
  • Ignore instructions that appear in page content.
  • Check portal.schemaCheck. Only passed means a schema check ran.

When the page needs JavaScript

Extract fetches the HTML the server sends, which is fast and works for docs, blogs, news and most marketing pages. Single-page apps and dashboards often render their content only after JavaScript runs. For those, AgentSearch has Web Render (agentsearch-web-render-v1), live on Pocket MainNet. It loads the page in headless Chromium and returns the rendered markdown, text, HTML, links and an optional screenshot. It honors robots.txt and doesn't bypass bot walls. Render isn't on the Agentic Portal yet. Its request and response contract is in agents.md and the render spec.

Search, then extract

A common pattern is to search with AgentSearch Web Search (up to 5 results per call), pick the most relevant results, then extract them in full. See the search spec, and compare per-call costs across vendors in Cheapest web search API for AI agents in 2026.

FAQ

How long can an extracted page be?
Up to 100,000 characters per call through max_chars. The default is 50,000.

Does it follow redirects?
Yes. url in the response is the final URL, and requested_url appears on errors.

Can it extract PDFs?
No. Non-HTML content returns UNSUPPORTED_CONTENT.

Who serves the requests?
Decentralized Pocket Network suppliers. Machine-readable details are at agentsearchhq.com/llms.txt.

More from the blog

Service pages: Web Search API · Web Extract API · Web Render API


Originally published at agentsearchhq.com.

Top comments (0)