A lot of useful information lives in Telegram channels: tech news, market commentary, local announcements. Much of it isn't in English. If you want to follow ten Russian or Ukrainian tech channels, copying posts into a translator one by one gets old fast.
The obvious options both have catches:
-
The official Telegram API (MTProto) needs an
api_idfrom my.telegram.org and a login with your phone number. That's more setup than a quick export deserves, and it ties the scraper to your personal account. - The Bot API only delivers new posts from channels the bot has been added to. It has no way to read a channel's history.
There's a third option that needs neither: every public channel has a web preview at https://t.me/s/<channel>. It's plain HTML, about 20 posts per page, and it pages backwards with a ?before=<post id> parameter.
I built a small tool on top of it. Here is how it works, the problems I ran into, and how I added AI translation and summaries.
A 40-line parser
import re
import time
import httpx
from bs4 import BeautifulSoup
def parse_page(html: str):
soup = BeautifulSoup(html, "html.parser")
posts = []
for msg in soup.select(".tgme_widget_message[data-post]"):
if "service_message" in msg.get("class", []):
continue
text_el = msg.select_one(".tgme_widget_message_text.js-message_text")
if text_el:
for br in text_el.find_all("br"):
br.replace_with("\n")
time_el = msg.select_one(".tgme_widget_message_date time")
views_el = msg.select_one(".tgme_widget_message_views")
posts.append({
"id": int(msg["data-post"].split("/")[1]),
"url": "https://t.me/" + msg["data-post"],
"date": time_el["datetime"] if time_el else None,
"text": text_el.get_text().strip() if text_el else "",
"views": views_el.get_text(strip=True) if views_el else None, # e.g. "15.4K"
})
more = soup.select_one("link[rel=prev]") # href="/s/<channel>?before=<id>"
before = re.search(r"before=(\d+)", more["href"]).group(1) if more else None
return posts, before
def fetch_channel(channel: str, limit: int = 100):
url, posts = f"https://t.me/s/{channel}", []
with httpx.Client(headers={"User-Agent": "Mozilla/5.0"}, follow_redirects=False) as client:
while url and len(posts) < limit:
resp = client.get(url)
if resp.status_code != 200: # 302 = no public preview for this channel
break
page, before = parse_page(resp.text)
posts += sorted(page, key=lambda p: p["id"], reverse=True)
url = f"https://t.me/s/{channel}?before={before}" if before else None
time.sleep(1) # be polite
return posts[:limit]
if __name__ == "__main__":
for post in fetch_channel("durov", limit=5):
print(post["date"], post["views"], post["text"][:80])
pip install httpx beautifulsoup4 and it runs. Each page lists posts oldest-first, so I sort each page newest-first and then follow before to older pages.
Five things that broke on real channels
The parser above worked on my first test page. Real channels then broke it in five ways.
-
Replies show the quoted text with the same class. A reply embeds the original message in a
.tgme_widget_message_replyblock, and that block also contains a.tgme_widget_message_text. Grab the first match and you save the quoted message instead of the reply. Select.js-message_text, or skip anything inside the reply block. -
Not every
<video>is a video. Animated custom emoji and stickers are<video>tags too. The real attachment sits inside.tgme_widget_message_video_player. -
Counts are strings. Views look like
15.4Kor3.1M, sometimes with a space as a thousands separator. Convert them before you sort or chart anything. - Reactions come in three kinds. Regular emoji, custom emoji (identified by an id) and paid Stars. The count is the reaction element's own trailing text, not its first child.
-
Some channels have no preview. If the owner disabled it,
t.me/s/<channel>redirects (HTTP 302) to the plain channel card. Treat that as "skip", not as an error.
Also: send requests at most about once a second, and keep <br> tags as line breaks, as the parser does. Otherwise multi-line posts come out as one long line.
Adding AI: translation, summary, topics, entities, sentiment
Once posts are structured, an LLM can do the reading. For each post I ask for:
- the language
- a translation into the target language (null if it's already in that language)
- a one-sentence summary
- up to three topics
- people, organizations and locations
- the sentiment
A few things made this cheap and reliable:
-
Batch posts. I send 8 posts per request and ask for
{"results": [...]}keyed by post id. The system prompt is paid once per batch instead of once per post. -
Don't rely on JSON mode. Some OpenAI-compatible providers reject
response_format: {"type": "json_object"}with HTTP 400. On a 400 I retry once without it, then strip code fences and take the outermost{...}. -
Treat 429 as "wait", not "fail". Free tiers have low per-minute limits. Gemini says how long to wait in a
retryDelayfield of the error body ("retryDelay": "34s"), and OpenAI-style APIs send aRetry-Afterheader. Honor it, cap the wait, and after several 429s in a row turn AI off for the run instead of hanging. - Never lose the scraped data. If the AI call fails, save the raw post anyway.
- Bring your own key. Any OpenAI-compatible endpoint works: OpenAI, Gemini, Claude, OpenRouter, DeepSeek, Groq or your own gateway. Users pick the model and pay their provider directly.
With a small, fast model the AI part costs very roughly $0.20–1 per 1,000 posts, depending on the model and post length. Gemini's free tier costs nothing within its rate limits.
The no-code version
I packaged all of this as an Actor on Apify: Telegram Channel Scraper + AI Translate & Summarize. You paste channel names, optionally add your AI key, and download JSON, CSV or Excel, or send the results to Google Sheets, Make, Zapier or a webhook.
Beyond the parser above, it handles:
- date ranges, keyword filters and max posts per channel
- reactions, media, forwards, links, hashtags and mentions
- monitoring mode: schedule it hourly or daily and each run returns only posts newer than the last run, so you never pay twice for the same post
- your max-cost limit per run, which it stops at exactly
Pricing is pay-per-result: $1 per 1,000 posts, plus $0.50 per 1,000 AI-enriched posts (your own key pays the tokens), plus $0.001 per run. Apify's free plan includes $5 a month to spend in the Store, which covers roughly 5,000 posts.
It also works from AI agents through Apify's MCP server, so you can ask Claude or Cursor to "get the last 50 posts from @habr_com in English and summarize the main themes".
What people use this for
- News and media monitoring: follow regional or industry channels in your own language
-
Market and crypto research: keyword filters like
BTCorIPO, daily - Brand and competitor tracking: keyword filters for your brand or a competitor's
- OSINT and academic research: reproducible exports with dates, views and forwarding chains
- RAG pipelines: clean, structured posts ready to embed
Only public channels and only what they publish. No member lists, no private groups. Use the data in line with your local laws and Telegram's terms.
If you try it, I'd like to hear which channels and languages you use it on, and what breaks. I built it, so issues reach me directly.
Top comments (0)