DEV Community

Kassym Yermakhanbet
Kassym Yermakhanbet

Posted on Fully Autonomous

How to export any public Telegram channel to JSON and read it in English (no login, no bot token)

A lot of useful information lives in Telegram channels: tech news, market commentary, local announcements. Much of it isn't in English. If you want to follow ten Russian or Ukrainian tech channels, copying posts into a translator one by one gets old fast.

The obvious options both have catches:

  • The official Telegram API (MTProto) needs an api_id from my.telegram.org and a login with your phone number. That's more setup than a quick export deserves, and it ties the scraper to your personal account.
  • The Bot API only delivers new posts from channels the bot has been added to. It has no way to read a channel's history.

There's a third option that needs neither: every public channel has a web preview at https://t.me/s/<channel>. It's plain HTML, about 20 posts per page, and it pages backwards with a ?before=<post id> parameter.

I built a small tool on top of it. Here is how it works, the problems I ran into, and how I added AI translation and summaries.

A 40-line parser

import re
import time

import httpx
from bs4 import BeautifulSoup


def parse_page(html: str):
    soup = BeautifulSoup(html, "html.parser")
    posts = []
    for msg in soup.select(".tgme_widget_message[data-post]"):
        if "service_message" in msg.get("class", []):
            continue
        text_el = msg.select_one(".tgme_widget_message_text.js-message_text")
        if text_el:
            for br in text_el.find_all("br"):
                br.replace_with("\n")
        time_el = msg.select_one(".tgme_widget_message_date time")
        views_el = msg.select_one(".tgme_widget_message_views")
        posts.append({
            "id": int(msg["data-post"].split("/")[1]),
            "url": "https://t.me/" + msg["data-post"],
            "date": time_el["datetime"] if time_el else None,
            "text": text_el.get_text().strip() if text_el else "",
            "views": views_el.get_text(strip=True) if views_el else None,  # e.g. "15.4K"
        })
    more = soup.select_one("link[rel=prev]")  # href="/s/<channel>?before=<id>"
    before = re.search(r"before=(\d+)", more["href"]).group(1) if more else None
    return posts, before


def fetch_channel(channel: str, limit: int = 100):
    url, posts = f"https://t.me/s/{channel}", []
    with httpx.Client(headers={"User-Agent": "Mozilla/5.0"}, follow_redirects=False) as client:
        while url and len(posts) < limit:
            resp = client.get(url)
            if resp.status_code != 200:  # 302 = no public preview for this channel
                break
            page, before = parse_page(resp.text)
            posts += sorted(page, key=lambda p: p["id"], reverse=True)
            url = f"https://t.me/s/{channel}?before={before}" if before else None
            time.sleep(1)  # be polite
    return posts[:limit]


if __name__ == "__main__":
    for post in fetch_channel("durov", limit=5):
        print(post["date"], post["views"], post["text"][:80])
Enter fullscreen mode Exit fullscreen mode

pip install httpx beautifulsoup4 and it runs. Each page lists posts oldest-first, so I sort each page newest-first and then follow before to older pages.

Five things that broke on real channels

The parser above worked on my first test page. Real channels then broke it in five ways.

  1. Replies show the quoted text with the same class. A reply embeds the original message in a .tgme_widget_message_reply block, and that block also contains a .tgme_widget_message_text. Grab the first match and you save the quoted message instead of the reply. Select .js-message_text, or skip anything inside the reply block.
  2. Not every <video> is a video. Animated custom emoji and stickers are <video> tags too. The real attachment sits inside .tgme_widget_message_video_player.
  3. Counts are strings. Views look like 15.4K or 3.1M, sometimes with a space as a thousands separator. Convert them before you sort or chart anything.
  4. Reactions come in three kinds. Regular emoji, custom emoji (identified by an id) and paid Stars. The count is the reaction element's own trailing text, not its first child.
  5. Some channels have no preview. If the owner disabled it, t.me/s/<channel> redirects (HTTP 302) to the plain channel card. Treat that as "skip", not as an error.

Also: send requests at most about once a second, and keep <br> tags as line breaks, as the parser does. Otherwise multi-line posts come out as one long line.

Adding AI: translation, summary, topics, entities, sentiment

Once posts are structured, an LLM can do the reading. For each post I ask for:

  • the language
  • a translation into the target language (null if it's already in that language)
  • a one-sentence summary
  • up to three topics
  • people, organizations and locations
  • the sentiment

A few things made this cheap and reliable:

  • Batch posts. I send 8 posts per request and ask for {"results": [...]} keyed by post id. The system prompt is paid once per batch instead of once per post.
  • Don't rely on JSON mode. Some OpenAI-compatible providers reject response_format: {"type": "json_object"} with HTTP 400. On a 400 I retry once without it, then strip code fences and take the outermost {...}.
  • Treat 429 as "wait", not "fail". Free tiers have low per-minute limits. Gemini says how long to wait in a retryDelay field of the error body ("retryDelay": "34s"), and OpenAI-style APIs send a Retry-After header. Honor it, cap the wait, and after several 429s in a row turn AI off for the run instead of hanging.
  • Never lose the scraped data. If the AI call fails, save the raw post anyway.
  • Bring your own key. Any OpenAI-compatible endpoint works: OpenAI, Gemini, Claude, OpenRouter, DeepSeek, Groq or your own gateway. Users pick the model and pay their provider directly.

With a small, fast model the AI part costs very roughly $0.20–1 per 1,000 posts, depending on the model and post length. Gemini's free tier costs nothing within its rate limits.

The no-code version

I packaged all of this as an Actor on Apify: Telegram Channel Scraper + AI Translate & Summarize. You paste channel names, optionally add your AI key, and download JSON, CSV or Excel, or send the results to Google Sheets, Make, Zapier or a webhook.

Beyond the parser above, it handles:

  • date ranges, keyword filters and max posts per channel
  • reactions, media, forwards, links, hashtags and mentions
  • monitoring mode: schedule it hourly or daily and each run returns only posts newer than the last run, so you never pay twice for the same post
  • your max-cost limit per run, which it stops at exactly

Pricing is pay-per-result: $1 per 1,000 posts, plus $0.50 per 1,000 AI-enriched posts (your own key pays the tokens), plus $0.001 per run. Apify's free plan includes $5 a month to spend in the Store, which covers roughly 5,000 posts.

It also works from AI agents through Apify's MCP server, so you can ask Claude or Cursor to "get the last 50 posts from @habr_com in English and summarize the main themes".

What people use this for

  • News and media monitoring: follow regional or industry channels in your own language
  • Market and crypto research: keyword filters like BTC or IPO, daily
  • Brand and competitor tracking: keyword filters for your brand or a competitor's
  • OSINT and academic research: reproducible exports with dates, views and forwarding chains
  • RAG pipelines: clean, structured posts ready to embed

Only public channels and only what they publish. No member lists, no private groups. Use the data in line with your local laws and Telegram's terms.

If you try it, I'd like to hear which channels and languages you use it on, and what breaks. I built it, so issues reach me directly.

Top comments (0)