A lot of OSINT work dies at the first step: Telegram's official API only sees channels you join, and joining 200 channels to monitor 200 channels is not a plan. Here is the no-login method I run in production, in the order I actually use it.
1. The web preview endpoint
Every public channel has a web preview at https://t.me/s/<channel>. It renders the last ~20 messages as plain HTML — no account, no API key, no rate-limit token. That is the whole trick: you are reading the page a logged-out human sees.
curl -s https://t.me/s/durov | head -c 400
If you can read one, you can read all of them. Language doesn't matter; a Russian, Spanish, and English channel all expose the same markup.
2. Pagination is a cursor, not a page number
?after=<message_id> walks forward, ?before=<message_id> walks backward. Message IDs are dense integers per channel, so the loop is:
- Fetch page, extract the newest message ID you can see.
- Request
?after=that_id. - Stop when the page returns the same set (you've hit live edge) or a 429.
There is no offset, no page count, and no way to jump to an arbitrary date — you can only walk. A channel with 40 posts/day is ~40 requests/day to keep current. That's polite-crawlable.
3. What the preview hides from you
The ceiling is real and you must report it honestly:
- Views/reaction counts on the preview lag the app.
- Media previews are thumbnails; full-resolution needs a second fetch.
- Very old history is not walkable — archives older than the crawl window are simply not there.
In my 25-channel run the hard ceiling was 1,271 messages in 46 minutes (~4.5 msg/sec effective), and 30% of captured messages were forwards of each other. Dedupe by normalized text hash before you report any count.
4. Parsing without a browser
The message blocks sit under div.tgme_widget_message with text in div.tgme_widget_message_text and the timestamp in <time datetime=...>. One CSS-selector pass per page is enough; no JS execution, no headless browser. If someone sells you a "Telegram scraper" that launches Chrome for this, they haven't looked at the markup.
5. The polite-crawl rules that keep it running
- 1 request/sec per channel, exponential backoff on 429.
- Cache every page you've fetched (message IDs are immutable).
- Verify channel names against the page's own
<meta property="og:title">— usernames get squatted, and a typo in your config becomes data from someone else's channel.
The full method write-up is here, and the pipeline that turns raw captures into three clean CSVs is this one. The ready-to-run poller with this exact schema ships in the guide.
Top comments (0)