Every "telegram scraper" tutorial starts with an API ID, an API hash, and a phone number to log in with. Then it shows you Telethon, explains that Telegram's MTProto is not HTTP, and asks you to register an application on a developer portal that wants your identity. For monitoring public channels, this is a cathedral built on a parking lot.
There is a dumber option: https://t.me/s/<channel> returns plain HTML. You can curl it. You can requests.get it. No keys, no login, no phone number, no app registration.
The whole scraper
import requests, re
from html.parser import HTMLParser
def fetch(channel: str, before: int | None = None) -> list[dict]:
url = f"https://t.me/s/{channel}"
if before:
url += f"/?before={before}"
html = requests.get(url, timeout=15).text
posts = []
# each message block carries data-post="<name>-<id>"
for m in re.finditer(r'data-post="[^"]*?-(\d+)"[^>]*>', html):
posts.append({"id": int(m.group(1))})
# text lives in .tgme_widget_message_text
for p, blk in zip(posts, html.split('tgme_widget_message_text')):
p["text"] = strip_html(blk[:4000])
# subscriber count, once per page
subs = re.search(r'subscriber_cont">\s*([\d,.]+)', html)
return posts # and subs.group(1) if you're polling counts
def strip_html(s: str) -> str:
s = re.sub(r'<br[^>]*>', '\n', s)
s = re.sub(r'<[^>]+>', '', s)
return re.sub(r'\s+', ' ', s).strip()
Not a framework. A regex and an HTML strip. That's the honest size of this problem for the public web-preview layer.
What this gets you
The last ~20 posts per call, in plain HTML: text, timestamps, view counts, forwards, media links. Pass ?before=<id> to walk backwards through history, and pass nothing to grab the live window. That's a full backfill (within the preview's coverage ceiling) plus a live tail, from one endpoint.
Where the dumber option stops being enough
- Private channels. The preview only serves public ones.
- Firehose speed. If you need sub-second latency on 500 channels, MTProto is your friend.
- Media downloads at scale. The
cdn4links work, but they rot; store what you keep.
For monitoring, alerting, research, and market-signal work, none of those three apply. The preview window is the product.
The part nobody tells you
The scraper is 20 lines; the system is the part that pays. Post-ID cursors that survive restarts. Dedupe against forwarded copies. Deciding that a subscriber count is a time series, not a number. Flagging repost rings before they pollute your signal. Those are collection-layer decisions, and they transfer to every channel you'll ever scrape.
I wrote them down: the $5 collection playbook (with the code patterns above productionized) is at https://heyuhe.gumroad.com/l/poddr. The free pay-what-you-want sample brief — a real channel analysis, flags and timelines included — is at https://heyuhe.gumroad.com/l/ruldgm. $0 is a valid price; take the brief, skip the bundle, that's fine.
Top comments (0)