Ran into this building an agent that needed to look at web pages.
Every screenshot API is SaaS, which is fine, until you notice the
page-understanding part runs on their model. So the text of every page you
capture gets sent to a third party and processed there. For a public marketing
page, who cares. For anything behind a login with real data on it, that's a
problem I couldn't get around.
Looked for a self-hostable one. The MCP screenshot servers that exist (Urlbox,
ScreenshotOne) are good but all point at the vendor's cloud. Didn't find one you
run yourself, so I built it.
Runs with docker compose. The extraction part runs on whatever model you point
at it — including Ollama on localhost, in which case the page text never leaves
the machine that rendered it. AGPL-3.0.
Honest about where it's weak: no security audit, extraction quality is decent
not amazing, and you can't capture your own private-IP hosts yet because the
SSRF guard blocks RFC1918 with no opt-out (allowlist is the next thing I'm
building). I also benchmarked it against two hosted APIs and the latency margins
weren't reproducible across runs, so I'm not claiming it's faster — the data and
the caveats are both in the repo.
github.com/route1-ai/shotbase
The thing I'm unsure about: is "run it yourself so page content stays in your
network" actually the useful part here, or do people mostly just want
screenshots and I've over-thought it? Genuinely asking — I built this for my own
use case and I don't know if it generalizes.
Top comments (0)