The crawler behind cl0q.com is now public. Anyone can read the code, run a node, and help index the open web.
https://codeberg.org/cl0qsearch/crawler-node
Why open source it?
Search shouldn't be a black box. The big engines don't tell you what they crawl, how they rank, or what they do with the data. We think that's wrong.
cl0q is an independent search engine with its own crawler and index. We've mapped 38.5M domains so far, working toward 100M. Now the crawler itself is open for anyone to inspect, run, or improve.
If you've ever wondered "what does a real web crawler look like under the hood," this is your answer. Every line is readable. No magic.
What it does
The node software discovers websites, fetches their homepages politely, and sends what it finds to cl0q's index. That's it. It doesn't scrape personal data, doesn't follow you around, doesn't build profiles.
Five services, one Postgres database, zero public ports. Everything runs locally on your machine.
It's a good bot (by design, not by promise)
Most crawlers say they respect robots.txt. Ours can't not respect it. The politeness rules are hardcoded:
- One honest User-Agent that names your node. It never pretends to be a browser.
- Public internet only. It physically cannot connect to private networks, localhost, or cloud metadata endpoints.
- robots.txt is checked first, every time. Crawl-delay is obeyed.
- At least 1 second between requests to any server. No hammering.
- Only fetches homepages, at most once every 120 days per site.
These aren't settings you can turn off. They're in the code.
Run your own node
You need Docker and a machine that's online. That's it.
git clone https://codeberg.org/cl0qsearch/crawler-node.git
cd crawler-node
cp .env.example .env
docker compose build
docker compose run --rm cl0q-node enroll --contact you@example.com
docker compose run --rm preflight
docker compose up -d
The enroll step registers your node with cl0q and gives you a token. The preflight check makes sure everything works before you start crawling.
What's next
This is v0.1.0. Coming up: more repos open sourced, a GitHub mirror, and performance work.
The code is MIT licensed. Fork it, break it, improve it, send PRs.
Run a node. Help index the open web.
Top comments (0)