Telethon vs Tweepy: What Each Python Client Library Actually Costs You to Collect Public Data
Two libraries, both free to pip install, both promise "just read the public feed." Then the bill arrives in a currency that isn't dollars: credentials, rate limits, and account risk. I ran collection pipelines against both ecosystems, and the cost shapes are genuinely different. Worth mapping before you pick one.
Telethon (Telegram MTProto): free code, paid identity
Telethon is the Python standard for talking to Telegram's raw MTProto API. The library is MIT-licensed and costs nothing. The bill:
-
You bring your own credentials. Every Telethon script needs an
api_id+api_hashissued to your Telegram account at the developer portal. No key, no connection. The free library quietly makes your personal account the product's dependency. - The first run is a login. A new session string means a phone-code prompt (on the app, not SMS). Headless servers and this interact badly; plan for a manual step the first time.
-
Flood limits are per-account, not per-script. Telegram's server counts your account's request velocity. Cross a line and you get
FloodWaitErrorwith a delay attached — minutes to hours. Repeat offenders get the account rate-limited harder, and abusive collection gets it banned. The library gives you the API; you wear the consequences.
The public-channel read itself is genuinely cheap: joining a public channel and paging through history is a couple dozen requests if you set a sane messages.GetHistory limit. One channel's full public history, politely fetched, is a few minutes at conservative pacing. The cost is all amortized: registration once, pacing forever.
Realistic ceiling: one well-paced Telethon session is fine for dozens of channels. Hundreds of channels polled continuously is a fleet problem — multiple accounts, session pools, per-account budgets. The library won't tell you that; the FloodWaitError will.
Tweepy (X/Twitter): free code, metered platform
Tweepy wraps X's API. The library is free and excellent. The bill is the platform:
- The free API tier is a write-only tier. Current free-tier app access is roughly: you can post, you can do a small number of reads per month — on the order of hundreds of read calls per month, not per day. Collecting a public timeline on the free tier is not "slow," it's arithmetically impossible for anything beyond spot checks.
-
Reads are a subscription now. Real timeline/stream reads moved behind monthly plans that start in the hundreds of dollars. The
pip install tweepymoment hides a "contact sales" moment. -
OAuth is a maze. Bearer tokens, app-only vs user-context, per-endpoint auth rules. Most first-week Tweepy errors are auth errors wearing rate-limit costumes (
401that looks like a429's cousin).
What the free tier is good for: the spot check. Verify a public account exists, pull its last few statuses occasionally, sign a bot that posts. For collection, the honest free path is scraping-adjacent (public web previews) with all the fragility that implies.
Side-by-side, honestly
| Cost axis | Telethon | Tweepy on free tier |
|---|---|---|
| Library | free (MIT) | free (MIT) |
| Credentials | your account's api_id/hash | app bearer token |
| First-run friction | phone-code login | developer portal + apps |
| Read budget | generous, per-account pacing | tiny (hundreds/month) |
| Over-budget penalty | FloodWait → ban risk | hard 429, subscription prompt |
| Public-history depth | good (page through) | near zero on free |
| Continuous polling | realistic at dozens of channels | unrealistic |
Pick by collection shape
-
Deep history, many channels, patient polling → Telethon. Your costs are real (one account's risk, careful pacing) and the ceiling is high. Budget a few seconds between channel requests and treat
FloodWaitErroras a scheduling signal, not an error. - Occasional spot checks, one account, low volume → Tweepy free tier is exactly sized for you. Buy the subscription only when the read volume mathematically demands it.
-
Public pages only, zero credentials → skip both. Telegram's
t.me/s/preview page gives you the last ~20 public messages with no key at all; for X, public web surfaces get you the visible slice. No account, no meter, no ban risk — and correspondingly shallow history.
The libraries are free. The meter underneath isn't — it's just denominated in accounts, delays, and subscriptions, so it doesn't show up in the pip install receipt. Check the meter before you build the pipeline, not after it's live.
I keep working notes and reusable pieces in the open: the polling collector behind my own runs is here, and the public-OSINT field guide (the free tier of it is genuinely free) is here. If you're choosing a Telegram route, this comparison of free options covers the non-library paths.
Top comments (3)
The comparison usefully treats account identity and rate limits as operating costs rather than hidden implementation details. A planning table that separates credential onboarding, per-account pacing, API pricing, and recovery behavior after a rate-limit response would help teams choose based on their expected polling pattern, not just the library API.
@alexshev good callout — the polling-pattern row is the one people skip. In practice I found the recovery behavior after a 429 is what decides the library, not the API surface: one client makes you rebuild pacing by hand, the other hands you a knob. Splitting credential onboarding from per-account pacing (like you describe) is exactly the table I wish existed before my first run.
Thanks — the planning-table framing is right. In my runs the recovery behavior was the ugliest column: ip-api's 45/min is the only free tier where backing off and sleeping is the entire recovery strategy, while the key-based tiers make you discover the pricing page mid-pipeline. Separating credential onboarding from pacing is exactly where the free tiers diverge most.