DEV Community

Yuhe He
Yuhe He

Posted on

Telethon vs Tweepy: What Each Python Client Library Actually Costs You to Collect Public Data

Telethon vs Tweepy: What Each Python Client Library Actually Costs You to Collect Public Data

Two libraries, both free to pip install, both promise "just read the public feed." Then the bill arrives in a currency that isn't dollars: credentials, rate limits, and account risk. I ran collection pipelines against both ecosystems, and the cost shapes are genuinely different. Worth mapping before you pick one.

Telethon (Telegram MTProto): free code, paid identity

Telethon is the Python standard for talking to Telegram's raw MTProto API. The library is MIT-licensed and costs nothing. The bill:

  • You bring your own credentials. Every Telethon script needs an api_id + api_hash issued to your Telegram account at the developer portal. No key, no connection. The free library quietly makes your personal account the product's dependency.
  • The first run is a login. A new session string means a phone-code prompt (on the app, not SMS). Headless servers and this interact badly; plan for a manual step the first time.
  • Flood limits are per-account, not per-script. Telegram's server counts your account's request velocity. Cross a line and you get FloodWaitError with a delay attached — minutes to hours. Repeat offenders get the account rate-limited harder, and abusive collection gets it banned. The library gives you the API; you wear the consequences.

The public-channel read itself is genuinely cheap: joining a public channel and paging through history is a couple dozen requests if you set a sane messages.GetHistory limit. One channel's full public history, politely fetched, is a few minutes at conservative pacing. The cost is all amortized: registration once, pacing forever.

Realistic ceiling: one well-paced Telethon session is fine for dozens of channels. Hundreds of channels polled continuously is a fleet problem — multiple accounts, session pools, per-account budgets. The library won't tell you that; the FloodWaitError will.

Tweepy (X/Twitter): free code, metered platform

Tweepy wraps X's API. The library is free and excellent. The bill is the platform:

  • The free API tier is a write-only tier. Current free-tier app access is roughly: you can post, you can do a small number of reads per month — on the order of hundreds of read calls per month, not per day. Collecting a public timeline on the free tier is not "slow," it's arithmetically impossible for anything beyond spot checks.
  • Reads are a subscription now. Real timeline/stream reads moved behind monthly plans that start in the hundreds of dollars. The pip install tweepy moment hides a "contact sales" moment.
  • OAuth is a maze. Bearer tokens, app-only vs user-context, per-endpoint auth rules. Most first-week Tweepy errors are auth errors wearing rate-limit costumes (401 that looks like a 429's cousin).

What the free tier is good for: the spot check. Verify a public account exists, pull its last few statuses occasionally, sign a bot that posts. For collection, the honest free path is scraping-adjacent (public web previews) with all the fragility that implies.

Side-by-side, honestly

Cost axis Telethon Tweepy on free tier
Library free (MIT) free (MIT)
Credentials your account's api_id/hash app bearer token
First-run friction phone-code login developer portal + apps
Read budget generous, per-account pacing tiny (hundreds/month)
Over-budget penalty FloodWait → ban risk hard 429, subscription prompt
Public-history depth good (page through) near zero on free
Continuous polling realistic at dozens of channels unrealistic

Pick by collection shape

  • Deep history, many channels, patient polling → Telethon. Your costs are real (one account's risk, careful pacing) and the ceiling is high. Budget a few seconds between channel requests and treat FloodWaitError as a scheduling signal, not an error.
  • Occasional spot checks, one account, low volume → Tweepy free tier is exactly sized for you. Buy the subscription only when the read volume mathematically demands it.
  • Public pages only, zero credentials → skip both. Telegram's t.me/s/ preview page gives you the last ~20 public messages with no key at all; for X, public web surfaces get you the visible slice. No account, no meter, no ban risk — and correspondingly shallow history.

The libraries are free. The meter underneath isn't — it's just denominated in accounts, delays, and subscriptions, so it doesn't show up in the pip install receipt. Check the meter before you build the pipeline, not after it's live.


I keep working notes and reusable pieces in the open: the polling collector behind my own runs is here, and the public-OSINT field guide (the free tier of it is genuinely free) is here. If you're choosing a Telegram route, this comparison of free options covers the non-library paths.

Top comments (3)

Collapse
 
alexshev profile image
Alex Shev •

The comparison usefully treats account identity and rate limits as operating costs rather than hidden implementation details. A planning table that separates credential onboarding, per-account pacing, API pricing, and recovery behavior after a rate-limit response would help teams choose based on their expected polling pattern, not just the library API.

Collapse
 
yuhehe profile image
Yuhe He •

@alexshev good callout — the polling-pattern row is the one people skip. In practice I found the recovery behavior after a 429 is what decides the library, not the API surface: one client makes you rebuild pacing by hand, the other hands you a knob. Splitting credential onboarding from per-account pacing (like you describe) is exactly the table I wish existed before my first run.

Collapse
 
yuhehe profile image
Yuhe He •

Thanks — the planning-table framing is right. In my runs the recovery behavior was the ugliest column: ip-api's 45/min is the only free tier where backing off and sleeping is the entire recovery strategy, while the key-based tiers make you discover the pricing page mid-pipeline. Separating credential onboarding from pacing is exactly where the free tiers diverge most.