It's not just that I made a crawler, but it's how the docs actually look that makes me love this so much.
I made an MCP server this weekend, and I am kind of in awe of how well this thing works.
It's called pulpie-mcp, and the basic idea is pretty simple:
Give AI coding agents their own tool for pulling full web pages and documentation into clean Markdown.
Not summarized documentation. Not chopped-up search results. Not "here are the important parts I think you need."
The actual content.
Tables, links, images, code, headings, all turned into Markdown and saved somewhere the agent can use again later.
pulpie-mcp
An MCP server that gives AI agents their own tools for turning web pages into clean Markdown. It runs the Pulpie content extraction model locally, so tables, code blocks, links, and images come through intact instead of getting summarized away.
The agent can read a page inline, save it as a .md file, crawl a whole docs section into a folder, and check what it already saved before fetching again.
Why I made it
I already had a little local UI for Pulpie that I used to pull docs and articles into Markdown, and it works really well. But I was still the one doing the pulling. Most agent fetch tools run pages through a small model that summarizes or trims them, which is not what you want when the agent needs the actual API reference.
So this gives the agent the same thing I was using. It can…
The test that kind of blew my mind
I had Claude test crawl_docs against a documentation site with 60 pages.
The backend wasn't already running, so this included starting the server and loading the extraction model onto my GPU.
| Phase | Time |
|---|---|
| Backend cold start | ~10s |
| Crawling and saving 60 pages | ~37s |
| Total | ~49s |
So in under a minute, Claude went from having none of those docs locally to having 60 clean Markdown files it could search, read, and reference.
I just thought this was so awesome.
And once the model is already loaded, obviously that initial ~10-second startup cost disappears.
Why I built it
I've already been using Pulpie locally to turn web pages into Markdown.
It uses a small content extraction model rather than just grabbing the raw page or trying to summarize it, and I've been really impressed with the results.
But there was one annoying part:
I was still the one pulling the documentation.
If I wanted an agent to work with a library or framework, I might grab some docs, save them into the project, point the agent toward them, and then let it work.
Which is fine.
But also... why am I doing that?
The agent has MCP tools. It knows what documentation it needs. It knows what project it's working on.
So I built an MCP server that lets it do the whole thing itself.
What pulpie-mcp gives the agent
There are currently four tools:
| Tool | What it does |
|---|---|
fetch_markdown |
Fetch a page and return clean Markdown directly to the agent |
save_markdown |
Fetch a page and save it as a Markdown file |
crawl_docs |
Crawl an entire documentation section and save it locally |
list_library |
Check which documentation has already been saved |
The distinction between fetching and saving turned out to be especially useful.
Sometimes an agent just needs to quickly read one page.
Other times I'm working on a project where I want it to pull an entire section of documentation and keep it around.
For example, I can basically tell Claude:
Pull the uv docs into this project's DOCS/reference folder.
And it handles the rest.
Or:
Do we already have the FastAPI docs saved somewhere?
No elaborate prompts required.
It doesn't shove 60 pages into the context window
This part was important to me.
When crawl_docs saves documentation, it doesn't return the contents of every page to the agent.
That would be completely insane.
Instead, the docs are written to disk and the agent gets useful metadata back. Each file also includes frontmatter with the source URL, title, and fetch time.
So the agent can grab a whole documentation section once and then read only the files it actually needs.
That turns the local filesystem into a pretty nice little documentation library.
For project-specific docs, the agent can save them inside the project.
For general research, pulpie-mcp has a global library at:
~/.pulpie/library/
Crawling is actually fairly smart too
crawl_docs doesn't just blindly click every link it sees.
It looks for sitemap information first:
robots.txt- a sitemap under the documentation path
- the site's root sitemap
- link crawling as a fallback
It stays on the same host and underneath the documentation path you started with.
There are four concurrent crawling workers, so network requests can overlap while extraction is happening.
That's how it managed to tear through those 60 pages so quickly.
The model runs locally
This is probably my favorite part.
pulpie-mcp uses:
feyninc/pulpie-orange-small
It's only a 210M parameter encoder model, and on my machine it uses around 420 MB of VRAM.
The MCP process itself doesn't load PyTorch or the model.
Instead, there's a tiny MCP server and a separate local backend.
When the agent uses one of the tools:
Agent
↓
MCP server
↓
Local Pulpie backend
↓
Web page
↓
Clean Markdown
If the backend isn't running yet, the MCP server starts it automatically.
Once it's running, every agent session shares the same backend, so I don't end up with Claude, Codex, and whatever else I'm abusing that day each loading their own copy of the model.
After 30 minutes without a request, it shuts itself down and releases the memory.
That was one of those little implementation details that took this from "neat experiment" to something I can actually leave installed and forget about.
Adding it to Claude Code
Install it:
uv tool install git+https://github.com/pinkpixel-dev/pulpie-mcp
Then:
claude mcp add --scope user pulpie -- pulpie-mcp
That's basically it.
The first time it runs, the Pulpie model downloads from Hugging Face.
Codex
Add this to:
~/.codex/config.toml
[mcp_servers.pulpie]
command = "pulpie-mcp"
And other MCP clients can use a normal MCP server configuration.
There are still some limitations
I haven't come across any pages that haven't worked yet, but it's possible that some completely client-rendered pages could come back empty. And it doesn't handle PDFs.
I've only personally tested it on Linux with NVIDIA/CUDA so far.
There is also one license note: the pulpie-mcp code is Apache 2.0, but the pulpie-orange-small model is licensed CC BY-NC 4.0, so the model is for non-commercial use unless you arrange something different with its creator.
This is exactly what I want MCP to be used for
I think that's what has me so excited about this project.
MCP gets discussed a lot in terms of connecting agents to giant services and APIs.
But I really like this smaller category of MCP tools:
Give the agent one very specific capability that removes a repetitive step from my workflow.
I used to:
- Find the documentation.
- Pull the relevant pages.
- Clean them up.
- Save them somewhere.
- Tell the coding agent where they are.
- Finally start working.
Now I can basically say:
Grab the docs you need and figure this out.
And it does.
Seeing it pull 60 pages in about 49 seconds, save them as clean Markdown, build an index, and then immediately start using them was one of those wonderful little:
"Oh. Damn. This is actually really useful."
moments.
And those are my favorite kinds of projects.
Repo:
pulpie-mcp
Top comments (2)
The fetch/save split is the right cut — and disk-as-library with frontmatter (source URL, title, fetch time) is the part most crawlers skip. One thing worth splitting in your numbers: ~37s for 60 pages with four workers is ~600ms per page, and most of that is the extraction model doing inference per page. That's the real price of ML-based cleaning, and it buys fidelity on hostile markup — but the render path costs an order of magnitude less on the same job: a system WebView extracting the DOM natively lands around 10ms per documentation-style page, no model, no GPU. Two trades, not a ranking: pulpie wins when the HTML fights back, the native path wins when the site already has sane structure — which is most docs. The robots → sitemap → links politeness chain is right either way. Between this (building the library) and an MCP browser for the live, authenticated, screenshot-needing half — navette, which I wrote up last week — an agent's web layer is getting properly shaped.
The thirty-minute idle shutdown on the local backend is the detail that makes this practical day to day. Most local MCP helpers either try to run as permanent background daemons that hold onto VRAM indefinitely, or they pay the PyTorch import and weight-loading latency on every tool call. A watchdog timer that drops memory when the agent finishes gives you the speed on bursts without having to manage background services by hand.