<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: cl0q</title>
    <description>The latest articles on DEV Community by cl0q (@cl0qsearch).</description>
    <link>https://dev.to/cl0qsearch</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146017%2Fb27a7e64-50b9-456c-93b6-13449754e96a.png</url>
      <title>DEV Community: cl0q</title>
      <link>https://dev.to/cl0qsearch</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/cl0qsearch"/>
    <language>en</language>
    <item>
      <title>I Pointed My Domain Search Engine at Cybercrime — Here's What It Found</title>
      <dc:creator>cl0q</dc:creator>
      <pubDate>Wed, 30 Sep 2026 02:55:51 +0000</pubDate>
      <link>https://dev.to/cl0qsearch/i-pointed-my-domain-search-engine-at-cybercrime-heres-what-it-found-1ce7</link>
      <guid>https://dev.to/cl0qsearch/i-pointed-my-domain-search-engine-at-cybercrime-heres-what-it-found-1ce7</guid>
      <description>&lt;h1&gt;
  
  
  I Pointed My Domain Search Engine at Cybercrime — Here's What It Found
&lt;/h1&gt;

&lt;p&gt;I build cl0q, a domain search engine (38.5M domains scanned, free API, no tracking). Last night I got curious: what happens when you run threat-intel queries through a general domain index instead of a specialized feed?&lt;/p&gt;

&lt;p&gt;Turns out: quite a lot. Here's what three searches returned.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. "crypto drainer" → 13 results, including live drainer kits
&lt;/h2&gt;

&lt;p&gt;Top hits included &lt;strong&gt;infernodrainer.xyz&lt;/strong&gt; ("Inferno Drainer — best crypto drainer since 2022," advertising "0-day exploits, 120+ wallet support"), &lt;strong&gt;exogator.com&lt;/strong&gt; ("Liquidate all cryptocurrencies and NFTs from any wallet"), &lt;strong&gt;purpledrainer.club&lt;/strong&gt;, and &lt;strong&gt;cryptodrainerscript.net&lt;/strong&gt; selling "multichain crypto drainer scripts."&lt;/p&gt;

&lt;p&gt;Each result came with resolved IPs and MX records — immediate pivot points. exogator.com sits on 5.182.209.88; infernodrainer.xyz on 216.198.79.65. That's infrastructure you can feed straight into a blocklist.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. "stealer log" → a notorious carding market, indexed
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;russianmarket.work&lt;/strong&gt; showed up: "The most trusted underground marketplace. Buy CVV, Stealer Logs, RDP, dumps..." — sitting in a general web index, with its IPs (172.67.149.100) and metadata attached. Alongside it: breach-intel platforms, stealer-log search services, and a Turkish hacking forum (spyhackerz.org) dealing in RATs and fresh leaks.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. "phishing kit" → an active scampage shop
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;lurez.online&lt;/strong&gt; — "Spamming Tools, Office 365 Scampages, Bybit &amp;amp; Binance Scampages For Sale," complete with SMTP servers, cracking utilities, and a "complete spamming course." Plus &lt;strong&gt;spams.tools&lt;/strong&gt; ("Spam Tools | Scampages | Checkers | Letters | Email Senders | SMTPs | cPanels") with its MX records exposed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters
&lt;/h2&gt;

&lt;p&gt;Threat intel usually means expensive feeds and APIs. But a huge amount of attacker infrastructure lives on the clear web, on cheap domains, behind Cloudflare — indexed like anything else. A domain-level search engine with DNS enrichment (IPs, MX, CNAMEs) turns a hunch into pivotable infrastructure in one query.&lt;/p&gt;

&lt;p&gt;The searches are free, the API is free (30 req/min, 1,000/day), and nothing is logged. Try it: &lt;a href="https://cl0q.com" rel="noopener noreferrer"&gt;cl0q.com&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All domains above were live in the index at time of writing. If you're a defender, the IPs and MX records are yours to use.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI assistance was used in drafting this article.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>showdev</category>
      <category>opensource</category>
      <category>webdev</category>
    </item>
    <item>
      <title>38.5 million domains on a 2 GB droplet: the honest story of building an open search engine</title>
      <dc:creator>cl0q</dc:creator>
      <pubDate>Sun, 27 Sep 2026 21:25:38 +0000</pubDate>
      <link>https://dev.to/cl0qsearch/385-million-domains-on-a-2-gb-droplet-the-honest-story-of-building-an-open-search-engine-4op4</link>
      <guid>https://dev.to/cl0qsearch/385-million-domains-on-a-2-gb-droplet-the-honest-story-of-building-an-open-search-engine-4op4</guid>
      <description>&lt;p&gt;&lt;em&gt;Disclosure: I used AI assistance to help draft this post. The story, the mistakes, and every number below are mine.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In 2025, I started crawling the web.&lt;/p&gt;

&lt;p&gt;Not metaphorically. I pointed a crawler at the open internet from a 2 GB VPS and decided to build a search engine — a free, open-source one that doesn't log your searches, doesn't need an account, and that anyone can self-host. The first crawl records date to April 2025. Then things went quiet for a while. Then this June, cl0q took off: 8.5 million domains in a single month.&lt;/p&gt;

&lt;p&gt;Today, cl0q knows 38.5 million domains and has indexed 24.8 million pages. The goal is 100 million.&lt;/p&gt;

&lt;p&gt;This is the honest version of how that went.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first million nearly killed the project
&lt;/h2&gt;

&lt;p&gt;The naive approach to crawling is: start with a list of sites, follow every link, repeat. What nobody tells you is that the web has traps. Point a crawler at the open web and it finds Blogspot — about a million subdomains hanging off a single domain — and happily spends a week indexing them while the rest of the internet waits.&lt;/p&gt;

&lt;p&gt;Getting to the first million was the hardest part, because I didn't know the right crawl path. Crawl too narrow and you re-walk the same neighborhood forever. Crawl too wide and you fall down a subdomain hole and your index fills up with a million pages from the same blog host.&lt;/p&gt;

&lt;p&gt;The fix was treating the known web as a rotating frontier instead of a finished list: keep re-sampling domains you already know and mine their homepages for fresh outbound links. And Certificate Transparency logs turned out to be a goldmine — at peak, they were surfacing around 1,700 new domains a minute. Discovery never ends; it just keeps moving.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everything broke, mostly because of disk
&lt;/h2&gt;

&lt;p&gt;For most of this project, cl0q ran on that 2 GB VPS. A web crawl is, at its core, a fight against disk space. Every page is a row, every favicon is bytes, every DNS record is overhead, and the crawl never stops growing.&lt;/p&gt;

&lt;p&gt;Lots broke. Things crashed, crawls stalled, I ran out of room at the worst moments. These days the storage lives on a remote host connected over Tailscale, which helped enormously — and I'm still fighting storage issues today. If there's one thing I'd tell anyone starting a crawl project: budget ten times the disk you think you need, then double it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The database has one core, and everything fights over it
&lt;/h2&gt;

&lt;p&gt;Here's something the docs won't tell you: search and the crawler share a single database, a single core, and the same table. When a burst of background jobs stampedes the database, search times out. When the crawler picks its next targets, it eats a big share of DB time doing it.&lt;/p&gt;

&lt;p&gt;There's no elegant fix yet — just triage, scheduling, and the occasional 2 a.m. realization that the thing slowing down search is my own crawler. The one part that never touches the database: the stats and drill-down pages, which are served from precomputed files. If the site feels fast while the crawler is hammering away, that's why.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually worked
&lt;/h2&gt;

&lt;p&gt;A few decisions paid off:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The pipeline is four small stages&lt;/strong&gt; — discovery, page crawl, enrichment, serving — each incremental and restartable. Nothing depends on a full rebuild. (That's the simplified version. The real version involves the single-core database above.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refusing to index garbage.&lt;/strong&gt; When a site serves an anti-bot challenge page, lazy crawlers index the text "Checking your browser…" as if it were content. cl0q detects interstitials and throws them out. A search engine is only as good as what it refuses to index.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Favicons turned out to be a superpower.&lt;/strong&gt; Every icon gets hashed with SHA-256: 16.3 million sites in the index have a favicon, collapsing down to about 7 million distinct icons. Sites sharing an icon are usually related — same operator, same framework, same phishing kit. What started as a storage deduplication trick became one of the most useful OSINT pivots in the index.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy as architecture, not policy.&lt;/strong&gt; Signed out, cl0q sets no cookies and stores nothing about your searches — there's no per-request log to leak, subpoena, or sell. Accounts are optional and passkey-based, and search history exists only if you turn it on. The code is open source, though I'll be honest: the public mirror is a couple months behind right now, and refreshing it is on my list.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where things stand
&lt;/h2&gt;

&lt;p&gt;The numbers, measured from the actual database rather than the dashboard: cl0q knows 38.5 million domains — 97% have been attempted, 92% resolve in DNS. It's indexed 24.8 million pages, 24.4 million of them live, across 1,027 TLDs. About 250,000 to 300,000 pages get crawled every day, and re-crawls start October 9 so pages stay fresh instead of going stale.&lt;/p&gt;

&lt;h2&gt;
  
  
  The road to 100 million
&lt;/h2&gt;

&lt;p&gt;The plan is simple: hit 100 million domains, then open up the API properly. There's a free beta API today and paid tiers are coming — and once the project makes money, the next frontier is depth. Right now cl0q fetches homepages; with real revenue behind it, the crawl goes deeper into sites instead of just wider across them.&lt;/p&gt;

&lt;p&gt;If you want to help, you can run a crawler and contribute pages — that's by invitation today, so get in touch — or just try a search at cl0q.com.&lt;/p&gt;

&lt;p&gt;A year and a half of crawling, 38.5 million domains, one very tired VPS. On to 100.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>opensource</category>
      <category>programming</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
