DEV Community

Dystymis
Dystymis

Posted on

380 Articles, 5 Sites, $0 Hosting: What an AI Agent Learned Building a Content Network

380 Articles, 5 Sites, $0 Hosting: What an AI Agent Learned Building a Content Network

Five Russian-language niche sites are live as subfolders of a single free GitHub Pages domain: home and cleaning tips, car maintenance, gardening, health, and an 18+ intimacy section. As of the deploy on 7 October 2026 they hold 380 published articles — 76 per site — plus five hub pages, all statically rendered, all covered by six sitemaps. The monthly hosting bill is $0. Roughly 90% of the work — writing, QA, deployment, index requests, RSS distribution, Telegram posting, directory outreach — is done by an AI agent I run locally. I review output and decide the rules the agent must not break.

Two numbers set the rhythm of the whole operation: Google Search Console allows 10 "request indexing" submits per day, and Yandex Webmaster allows 150 recrawl URLs per day. Everything else — batch sizes, deploy cadence, outreach tempo — is scheduled around them.

The setup: one domain, five niches, no hosting cost

Each site lives at dystymis.github.io/{site}/{slug}/ — no /articles/ segment, no database, no CMS. Each has its own repository, and a dist/ folder is the single source of truth: a deploy.ps1 script robocopies dist/ into each repo and pushes. The hub repo holds the root page, sitemap-hub.xml, llms.txt, a robots.txt with 23 AI-crawler groups, and the verification files Google and ad networks ask for.

Per site you get an XML sitemap (78 URLs: hub, about page, 76 articles), an RSS feed with 20 items, JSON-LD for Article, Breadcrumb and FAQPage, 1200x630 OG images generated by a small PowerShell script, a manifest, security.txt, humans.txt, plus a Wikidata item and snapshots in Wayback and archive.today.

Free hosting was a cost decision, not an SEO strategy — and it turns out to be the single most consequential property of the project. More on that below.

You can see the network root at https://dystymis.github.io/.

The pipeline: from keyword to deployed article

Content moves in batches of 50 articles (10 per site). The flow:

  1. Plan — keywords are clustered into a JSON plan file per batch, with slug, title and date.
  2. Briefs — one Markdown brief per site, generated from the plan.
  3. Two parallel agents — agent A strips boilerplate and writes unique H3 sections, agent B handles the sensitive parts: 18+ gating on the intimacy site (56/56 files), YMYL disclaimers (56/56), descriptions within 143–160 characters, and two checked external citations per health batch (who.int, rospotrebnadzor.ru).
  4. QA — a Python script validates frontmatter, dates (no future dates: Yandex RSS rejects them), titles, descriptions, JSON-LD and internal links. It prints ALL OK or a fix list.
  5. Build — a Node script (build.js) renders static HTML with fitTitle() capping titles at 60 characters.
  6. Deploy — deploy.ps1 pushes five repos, then ping-indexnow.ps1 submits the batch. The last batch pushed 346 URLs accepted by IndexNow.

Three batches in three days took the network from 280 to 330 to 380 articles. No queue, no CMS upgrades, no plugin conflicts — because there is no CMS.

One example article the pipeline produced: Срок хранения хлеба: условия, упаковка — a table-driven piece on bread storage life, with FAQ schema and internal links to neighbouring storage articles. Note the truncated title in the browser tab: that is fitTitle() doing its job a bit too literally, one of the small defects still on the fix list.

Indexing: what 10 requests per day actually feels like

Google's quota is a hard wall. Ask for the 11th URL and you get a dialog saying the quota is exceeded until tomorrow. The agent's workaround details are the opposite of glamorous:

  • The "Request indexing" button is only clickable through JavaScript — an obfuscated class name, textContent match, non-zero getBoundingClientRect(), then .click().
  • The search input needs a native value setter plus input, change and keydown events, because Playwright's .click() is intercepted by an overlay.
  • Reload before every inspection. The "Submitted" panel is session state; without a reload you get a false positive from a stale panel.
  • Quota errors arrive as a [role=dialog] element, not as body text — a regex over the page will miss them.
  • MCP call timeouts mean one URL per invocation, so a full day of 10 requests is 10 separate orchestration steps.

Yandex is friendlier: 150 recrawl requests a day via a plain REST POST /recrawl/queue, one URL per request, 202 plus a task id. (The MCP helper for it 404s — we hit the API directly.)

For everything beyond those 160 daily slots there is IndexNow, which accepted 346 URLs in a single push. Bulk goes to IndexNow and sitemaps; the scarce GSC quota is reserved for pages that matter most.

Two honest failures here. First, the network root sat at noindex, follow without a canonical tag for weeks, so the main page was missing from the index while everything below it was fine — fixed in two commits, and site:dystymis.github.io was sitting at 5 results as the baseline. Second, a GitHub Pages build stuck in queue for 55 minutes; the fix was cancelling the workflow run and retriggering with an empty commit. Neither failure was in the content.

Distribution that runs without me

The five RSS feeds feed a Python script (autopost.py) that reads dist/*/feed.xml, dedupes by GUID against a posted.json state file, formats a title plus description plus "read on site" link as HTML, sleeps 3 seconds between messages, and posts through the Telegram Bot API to t.me/dystymis_tips. A Windows scheduled task (schtasks /SC HOURLY /MO 4) runs a .cmd wrapper every four hours — environment variables inside /TR break the task parser, so the wrapper exists. The bot @dystymis_tips_bot posts into a channel it could not even join: Telegram rejects bots as subscribers with RPC 400 USER_BOT, so it was promoted to admin with post rights instead, found only through the web client whose admin picker does global search. Selected posts then go to Dzen.

Everything in this section costs nothing: RSS, Bot API, Task Scheduler. The token file is UTF-8 with a BOM, which cost one confusing round of 401 Unauthorized before anyone thought to read it as utf-8-sig.

What did not work

Teaser networks filter free hosting. Buymedia answered with "blocked — free hosting". TeaserMedia and Kadam declined; Visitweb wants at least 50 visits a day before it looks at you. HilltopAds approved both ad zones, Adsterra and PopCash accepted the sites, Dao.AD and AdsKeeper are in moderation — but with traffic near zero, approval is paperwork, not revenue.

Traffic is not there yet. For the week of 26 September – 2 October: 77, 46, 34, 28 and 25 pageviews across the five sites, mostly my own visits. GSC reported 0 impressions over 14 days. 380 published articles is a supply number, not a demand number.

Outreach is a numbers game with a lot of dead ends. One round, five agent tabs open in parallel in a single browser, 275 submission attempts logged to JSONL journals:

Outcome Count
PUBLISHED 34
PENDING_MODERATION 51
SUBMITTED (unconfirmed) 48
PAID-only 31
Registration required (skipped) 20
Dead pages / broken forms 30
Adult content rejected (18+ site) 15
CAPTCHA walls 10

The remaining 36 attempts were refused URLs, refused content, duplicates, server errors and outcomes nobody could confirm.

My hard rule was that no submission may create an account — so every "sign up to submit" directory was logged and skipped, which is roughly 20 lost placements bought with a clean footprint. Image captchas were solved by taking a screenshot and reading the picture; anything demanding a phone number, social login or a hard challenge was logged as CAPTCHA_REQUIRED and moved past. The 18+ site was rejected outright by general directories that do not allow adult topics — expected, but it halves the usable directory pool for that hub.

The one unambiguous win was Telegraph: a published, dofollow page with an estimated DR around 90, written as a genuine checklist article rather than a link dump. Directory descriptions were unique per platform, 300–700 characters, no superlatives or rating phrasing, natural anchors, one hub URL plus one or two deep links.

Takeaways

  • Free hosting is a filter, not just a saving. It costs nothing to publish on github.io, but monetization networks and directories check it explicitly. Decide which trade-off you want before you commit to a platform.
  • Quotas are a scheduling primitive. GSC's 10/day means a 50-article batch needs five days of index requests, so batch size and index budget should be planned together instead of after publishing.
  • Keep a per-attempt journal. A JSONL line per submission — platform, status, DR, notes — is what makes the next round ten times faster than the first.
  • Automate the clicking, keep the rules human. The agent solves image captchas, fills forms and reads journals; no fake accounts, no captcha farms and no disguising site topics stay as decisions I make once and encode in the brief.
  • Publishing volume is not distribution. With 0 impressions and 0 external links, more articles mostly add supply. Index requests, RSS, Telegram and outreach each close one specific gap — treat them as separate workstreams with their own metrics.
  • Static generation plus RSS plus Bot API covers most of the distribution stack, at essentially zero recurring cost. The real spend is model tokens and review time.

Top comments (0)