DEV Community

Yuhe He
Yuhe He

Posted on

A Telegram Scraper in 20 Lines of Python (No API Keys, No Login)

Every "telegram scraper" tutorial starts with an API ID, an API hash, and a phone number to log in with. Then it shows you Telethon, explains that Telegram's MTProto is not HTTP, and asks you to register an application on a developer portal that wants your identity. For monitoring public channels, this is a cathedral built on a parking lot.

There is a dumber option: https://t.me/s/<channel> returns plain HTML. You can curl it. You can requests.get it. No keys, no login, no phone number, no app registration.

The whole scraper

import requests, re
from html.parser import HTMLParser

def fetch(channel: str, before: int | None = None) -> list[dict]:
    url = f"https://t.me/s/{channel}"
    if before:
        url += f"/?before={before}"
    html = requests.get(url, timeout=15).text

    posts = []
    # each message block carries data-post="<name>-<id>"
    for m in re.finditer(r'data-post="[^"]*?-(\d+)"[^>]*>', html):
        posts.append({"id": int(m.group(1))})
    # text lives in .tgme_widget_message_text
    for p, blk in zip(posts, html.split('tgme_widget_message_text')):
        p["text"] = strip_html(blk[:4000])
    # subscriber count, once per page
    subs = re.search(r'subscriber_cont">\s*([\d,.]+)', html)
    return posts  # and subs.group(1) if you're polling counts

def strip_html(s: str) -> str:
    s = re.sub(r'<br[^>]*>', '\n', s)
    s = re.sub(r'<[^>]+>', '', s)
    return re.sub(r'\s+', ' ', s).strip()
Enter fullscreen mode Exit fullscreen mode

Not a framework. A regex and an HTML strip. That's the honest size of this problem for the public web-preview layer.

What this gets you

The last ~20 posts per call, in plain HTML: text, timestamps, view counts, forwards, media links. Pass ?before=<id> to walk backwards through history, and pass nothing to grab the live window. That's a full backfill (within the preview's coverage ceiling) plus a live tail, from one endpoint.

Where the dumber option stops being enough

  • Private channels. The preview only serves public ones.
  • Firehose speed. If you need sub-second latency on 500 channels, MTProto is your friend.
  • Media downloads at scale. The cdn4 links work, but they rot; store what you keep.

For monitoring, alerting, research, and market-signal work, none of those three apply. The preview window is the product.

The part nobody tells you

The scraper is 20 lines; the system is the part that pays. Post-ID cursors that survive restarts. Dedupe against forwarded copies. Deciding that a subscriber count is a time series, not a number. Flagging repost rings before they pollute your signal. Those are collection-layer decisions, and they transfer to every channel you'll ever scrape.

I wrote them down: the $5 collection playbook (with the code patterns above productionized) is at https://heyuhe.gumroad.com/l/poddr. The free pay-what-you-want sample brief — a real channel analysis, flags and timelines included — is at https://heyuhe.gumroad.com/l/ruldgm. $0 is a valid price; take the brief, skip the bundle, that's fine.

Top comments (0)