The hardest part of running a public Telegram collection set is not scraping. The t.me/s/ preview layer is polite: a page fetch, a parse, done. The hard part is cadence - deciding how often to poll 200 channels without (a) drowning in requests, (b) missing the post that matters, or (c) waking up to a dead pipeline because one HTML change broke the parser at 04:00 UTC.
Here is the scheduling design that survived six months of unattended runs on GitHub Actions, zero paid infrastructure.
Tier the channels by alert value, poll accordingly. Warzone alert channels: every 10 minutes. Aggregators that summarize: hourly. Slow official channels (ministries, agencies): 6-hourly. This one decision cuts request volume ~70% versus uniform polling, and it puts the latency budget where the alerts live.
Poll in randomized order within each tier. Fixed-order polling at fixed intervals is a fingerprint; a datacenter IP hammering channels in the same sequence every 10 minutes is the easiest bot pattern to spot. Shuffle per cycle, add jitter of seconds. The preview endpoint has no auth, so politeness is what keeps your IP usable.
Parse defensively, store raw. The preview HTML changes without notice - a class rename once silently zeroed my ingestion for two days because the parser had one selector and no alarm. Store the raw HTML fragment per post in the database before parsing. When a parser breaks, you re-parse history; you don't lose it. Alert on parse rate, not on fetch rate: 100 fetches, 0 posts parsed = alarm.
Gap-fill on demand, not on schedule. Sequential post IDs (data-post) mean you can fetch a specific range when a channel shows a gap - a missing 30 posts during a network outage gets backfilled with 30 targeted fetches, in the next cycle, not in a special recovery mode. The public preview is stateless; treat history as a range you can walk at your own pace.
Watch the quiet channels loudest. A channel in your set that goes silent for 12 hours when its peers are posting at full volume is the most informative event in the corpus - deletions, admin arrests, platform pressure, and offline periods all look identical in the ingest logs and all matter. Silence detection is free; you already poll, absence is observable.
The failure modes that actually bit me, in order: parser selector breakage (fixed by raw-HTML storage), subscriber-count locale format changes (k vs. thousands separators), GitHub Actions scheduler drift during load spikes (add a watchdog ping outside Actions), and one memorable week where the preview layer served cached pages ~20 minutes stale for a whole tier - detected only because a mirror's post timestamp predated its origin in my store. Clock sanity checks on ingested timestamps caught it.
The collection cadence, parse-rate alarm, and gap-fill rules are part of the Telegram & Web OSINT Bundle ($5). Free sample brief shows the output format.
Runs free on GitHub Actions - no server, no paid APIs.
Top comments (0)