DEV Community

Yuhe He
Yuhe He

Posted on

How to Scrape Public Telegram Channels Without an API Key (Step by Step)

A lot of OSINT work dies at the first step: Telegram's official API only sees channels you join, and joining 200 channels to monitor 200 channels is not a plan. Here is the no-login method I run in production, in the order I actually use it.

1. The web preview endpoint

Every public channel has a web preview at https://t.me/s/<channel>. It renders the last ~20 messages as plain HTML — no account, no API key, no rate-limit token. That is the whole trick: you are reading the page a logged-out human sees.

curl -s https://t.me/s/durov | head -c 400
Enter fullscreen mode Exit fullscreen mode

If you can read one, you can read all of them. Language doesn't matter; a Russian, Spanish, and English channel all expose the same markup.

2. Pagination is a cursor, not a page number

?after=<message_id> walks forward, ?before=<message_id> walks backward. Message IDs are dense integers per channel, so the loop is:

  1. Fetch page, extract the newest message ID you can see.
  2. Request ?after=that_id.
  3. Stop when the page returns the same set (you've hit live edge) or a 429.

There is no offset, no page count, and no way to jump to an arbitrary date — you can only walk. A channel with 40 posts/day is ~40 requests/day to keep current. That's polite-crawlable.

3. What the preview hides from you

The ceiling is real and you must report it honestly:

  • Views/reaction counts on the preview lag the app.
  • Media previews are thumbnails; full-resolution needs a second fetch.
  • Very old history is not walkable — archives older than the crawl window are simply not there.

In my 25-channel run the hard ceiling was 1,271 messages in 46 minutes (~4.5 msg/sec effective), and 30% of captured messages were forwards of each other. Dedupe by normalized text hash before you report any count.

4. Parsing without a browser

The message blocks sit under div.tgme_widget_message with text in div.tgme_widget_message_text and the timestamp in <time datetime=...>. One CSS-selector pass per page is enough; no JS execution, no headless browser. If someone sells you a "Telegram scraper" that launches Chrome for this, they haven't looked at the markup.

5. The polite-crawl rules that keep it running

  • 1 request/sec per channel, exponential backoff on 429.
  • Cache every page you've fetched (message IDs are immutable).
  • Verify channel names against the page's own <meta property="og:title"> — usernames get squatted, and a typo in your config becomes data from someone else's channel.

The full method write-up is here, and the pipeline that turns raw captures into three clean CSVs is this one. The ready-to-run poller with this exact schema ships in the guide.

Top comments (0)