<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yuhe He</title>
    <description>The latest articles on DEV Community by Yuhe He (@yuhehe).</description>
    <link>https://dev.to/yuhehe</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174005%2F45b447d5-3d05-439c-85aa-4fed6fa65b89.png</url>
      <title>DEV Community: Yuhe He</title>
      <link>https://dev.to/yuhehe</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yuhehe"/>
    <language>en</language>
    <item>
      <title>The 429 Playbook: Handling Rate Limits Before Your IP Gets Banned</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sun, 11 Oct 2026 01:02:53 +0000</pubDate>
      <link>https://dev.to/yuhehe/the-429-playbook-handling-rate-limits-before-your-ip-gets-banned-5e0g</link>
      <guid>https://dev.to/yuhehe/the-429-playbook-handling-rate-limits-before-your-ip-gets-banned-5e0g</guid>
      <description>&lt;h1&gt;
  
  
  The 429 Playbook: Handling Rate Limits Before Your IP Gets Banned
&lt;/h1&gt;

&lt;p&gt;Every scraping tutorial shows the happy path. The page that matters is the one you get after an hour of collecting: &lt;code&gt;429 Too Many Requests&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here's the playbook I run, ordered by what each response should cost you.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Read the headers first, back off second
&lt;/h2&gt;

&lt;p&gt;Most servers tell you what to do in the response:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Retry-After&lt;/code&gt;: the contract. If it says 30, sleep 30. Not 5, not "I'll retry in a loop for a second".&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;X-RateLimit-Remaining&lt;/code&gt; / &lt;code&gt;X-RateLimit-Reset&lt;/code&gt;: your budget and its refill time. When &lt;code&gt;Remaining&lt;/code&gt; hits 2, you throttle &lt;em&gt;before&lt;/em&gt; the 429, not after.&lt;/li&gt;
&lt;li&gt;No headers (common on free-tier APIs): treat every 429 as a warning shot. The next escalation is usually a temporary IP ban, and those are invisible — you get a 403 with a "blocked" page, not a rate-limit message.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. The three-stage ladder
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1st 429  → exponential backoff: 2^n * base (cap at 60s), 2 retries
2nd 429  → drop concurrency to 1, switch to a paced schedule (e.g. 1 req / 5s)
3rd 429  → stop. Persist the cursor, exit non-zero, let the scheduler retry later
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The third stage is the one people skip. A tight retry loop that never gives up is how a free-tier API learns your IP address very well, very permanently. My collection jobs are resumable (cursor in a state file), so "stop and come back in an hour" is cheap — the pipeline picks up exactly where it died.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The failure mode nobody shows you
&lt;/h2&gt;

&lt;p&gt;A 429 is not the only rate signal. Watch for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Silent partial responses&lt;/strong&gt;: 200 OK with an empty list, where yesterday's identical call returned 50 items. Treat "200 + empty" as a soft 429 and re-verify against a known-good channel/endpoint before trusting the gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency creep&lt;/strong&gt;: responses slowing from 200 ms to 2 s is a rate limiter warming up. Throttle proactively.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality cliffs&lt;/strong&gt;: the payload structure changes (&lt;code&gt;null&lt;/code&gt; fields, truncated text) — some services degrade responses before they refuse them.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Budget math before you write code
&lt;/h2&gt;

&lt;p&gt;Free tiers are rate-limited per IP, per key, per day. Do the arithmetic up front: 45 requests/min free tier × 60 min = 2,700 requests. At 200 items per page, a 20,000-message collection is 100 pages — fits easily on one channel, dies instantly at 30 channels. The playbook only works if your request budget covers the job; otherwise you're choosing between key rotation, sampling, or a paid tier &lt;em&gt;today&lt;/em&gt;, not at 2 a.m. mid-collection.&lt;/p&gt;

&lt;p&gt;The short version: 429s are a contract, not an obstacle. Read it, back off on schedule, persist your cursor, and know when stopping is the cheapest move in the pipeline.&lt;/p&gt;

</description>
      <category>api</category>
      <category>backend</category>
      <category>performance</category>
      <category>webscraping</category>
    </item>
    <item>
      <title>Merging Two Messy CSVs: The Key Problems Excel Formulas Hide From You</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sun, 11 Oct 2026 00:26:32 +0000</pubDate>
      <link>https://dev.to/yuhehe/merging-two-messy-csvs-the-key-problems-excel-formulas-hide-from-you-5a6i</link>
      <guid>https://dev.to/yuhehe/merging-two-messy-csvs-the-key-problems-excel-formulas-hide-from-you-5a6i</guid>
      <description>&lt;h1&gt;
  
  
  Merging Two Messy CSVs: The Key Problems Excel Formulas Hide From You
&lt;/h1&gt;

&lt;p&gt;You have two spreadsheets. One is your export, one is the client's. Someone says "just merge them on the ID column." That sentence has caused more quiet data disasters than any schema migration, because it assumes three things that are never true: the IDs match, the columns mean the same thing, and the rows are unique.&lt;/p&gt;

&lt;p&gt;I reconcile exported data files for a living — chat logs, server logs, marketplace exports — and the merge is where the bodies are buried. Here are the failure modes, in the order they will find you.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The Key Is Not a Key
&lt;/h2&gt;

&lt;p&gt;"Same ID column" usually means: one file has &lt;code&gt;id&lt;/code&gt;, the other has &lt;code&gt;customer_id&lt;/code&gt;, &lt;code&gt;order.customer&lt;/code&gt;, and a &lt;code&gt;vlookup_key&lt;/code&gt; someone added by hand. Open both, count distinct values, and compare sets before you merge:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;a_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df_a&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;b_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df_b&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a_ids&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b_ids&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a_ids&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;b_ids&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;orphans A:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a_ids&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;b_ids&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;orphans B:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b_ids&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;a_ids&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the intersection is 91%, not 100%, your merge is a research project, not a formula. Excel's &lt;code&gt;VLOOKUP&lt;/code&gt; returns &lt;code&gt;#N/A&lt;/code&gt; and moves on; a pandas merge silently picks a join type for you. You must count the misses, not spot-check them.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Invisible Keys
&lt;/h2&gt;

&lt;p&gt;The intersection is computed on strings. &lt;code&gt;"A102"&lt;/code&gt; and &lt;code&gt;"A102 "&lt;/code&gt; — trailing space from a CRM export — are different strings. &lt;code&gt;"42"&lt;/code&gt; and &lt;code&gt;"42.0"&lt;/code&gt; — one file went through a spreadsheet app that cast numbers — are different strings. The fix is boring and mandatory: trim, cast, and casefold &lt;em&gt;both sides&lt;/em&gt; before joining, then re-count. Ninety percent of "impossible" 4% miss rates are whitespace.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Uniqueness Is a Claim, Not a Fact
&lt;/h2&gt;

&lt;p&gt;Run &lt;code&gt;df["id"].is_unique&lt;/code&gt; on both sides. If the right side has duplicates, your 1,000-row left table becomes 1,400 rows after the merge, and every downstream sum is inflated by 40%. In Excel you see it as weird double-counting you can never explain. In code, the assertion is one line, and it fails loudly.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The Merge Type Is a Business Decision
&lt;/h2&gt;

&lt;p&gt;Inner drops unmatched rows. Left keeps your rows and nulls theirs. Outer keeps everything and doubles your review work. There is no default — there is only what the deliverable promises. A reconciliation report is outer-merge, annotate. A clean master table is inner-merge, log the drops. Pick before you write the line, and record the count of rows that fell out the bottom of every join you do.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Dates Are Four Formats in a Trench Coat
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;2026-1-5&lt;/code&gt;, &lt;code&gt;01/05/2026&lt;/code&gt;, &lt;code&gt;5 Jan 26&lt;/code&gt;, and a serial number &lt;code&gt;46027&lt;/code&gt; are the same date in four costumes, and the file with the serial numbers does not tell you it is wearing one. Parse with explicit formats; never trust auto-detection to pick month-first. &lt;code&gt;03/04/2026&lt;/code&gt; is unambiguous to no one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Workflow That Survives
&lt;/h2&gt;

&lt;p&gt;Load both files raw. Normalize (trim, cast, casefold keys; explicit date parse). Assert uniqueness. Merge with an explicit &lt;code&gt;how&lt;/code&gt; and an &lt;code&gt;indicator&lt;/code&gt;. Count before, after, matched, dropped — all four, in the log, every run.&lt;/p&gt;

&lt;p&gt;Total: one screen of pandas, zero &lt;code&gt;#N/A&lt;/code&gt;s, and a number you can hand the client when the totals do not match: "14 rows in their file had no match on our side; here they are." That sentence — not the merge itself — is what you are being paid for.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>datascience</category>
      <category>sql</category>
      <category>beginners</category>
    </item>
    <item>
      <title>IP Geolocation APIs Compared: What the Free Tier Really Gives You</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sun, 11 Oct 2026 00:03:18 +0000</pubDate>
      <link>https://dev.to/yuhehe/ip-geolocation-apis-compared-what-the-free-tier-really-gives-you-pg7</link>
      <guid>https://dev.to/yuhehe/ip-geolocation-apis-compared-what-the-free-tier-really-gives-you-pg7</guid>
      <description>&lt;h1&gt;
  
  
  IP Geolocation APIs Compared: What the Free Tier Really Gives You
&lt;/h1&gt;

&lt;p&gt;You have an IP address. You want a city. The obvious move is to grab a free geolocation API key and call it a day. The bill arrives later, in a currency you did not price in: quota cliffs, stale databases, and the moment your "free" city lookup starts costing you real money.&lt;/p&gt;

&lt;p&gt;I ran the same list of 1,271 IPs through the free tiers of the usual suspects. Here is what each route actually gives you before it starts charging.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five Routes
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;ipapi.co&lt;/strong&gt; — 1,000 calls/day free. Zero setup, no key needed for the first tier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ipinfo.io&lt;/strong&gt; — 50 calls/day free. Fast, accurate, and the free tier is a demo, not a tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ip-api.com&lt;/strong&gt; — 45 requests/minute free, no key. Self-host option for the whole database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MaxMind GeoLite2&lt;/strong&gt; — free database file, license key required, updated weekly. Local lookups, unlimited.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BigDataCloud&lt;/strong&gt; — free tier with 10,000 requests/day, no key, non-commercial clause.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What the Free Tiers Hide
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The quota cliff is not a ramp.&lt;/strong&gt; ipinfo's 50/day means any batch over fifty needs a card. ipapi's 1,000/day looks generous until you realize a single day of log enrichment for a mid-size site is 50,000 lookups. The cliff arrives at the exact moment your project stops being a toy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ip-api's rate limit is the friendly one.&lt;/strong&gt; 45/minute, no key, no card. For a solo builder enriching a CSV of a few thousand IPs, you just sleep 1.4 seconds between batches and never pay. The catch: free tier is non-commercial, and the cloud endpoint lags the paid one on database freshness.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MaxMind is the only free tier with no per-call ceiling&lt;/strong&gt; — because it is not an API at all. You download the GeoLite2 file, load it locally, and every lookup is free and instant, forever. The real cost is the license-key registration dance (account required), the weekly update cycle (a location that moved yesterday is stale until you re-download), and the accuracy floor on mobile ISPs: a cell-tower IP will place you in the ISP's regional hub, not the user's street. No free tier fixes that; it is a property of the data, not the vendor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;BigDataCloud trades money for clauses.&lt;/strong&gt; 10,000/day free is the most generous quota in the group, but the license is non-commercial. For a side project measuring where readers come from, fine. The moment the project earns a dollar, you are re-reading contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Decision Table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Route&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One-off CSV, &amp;lt;1,000 IPs&lt;/td&gt;
&lt;td&gt;ip-api.com&lt;/td&gt;
&lt;td&gt;No key, rate limit is the only rule&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recurring enrichment, thousands/day&lt;/td&gt;
&lt;td&gt;MaxMind GeoLite2 local&lt;/td&gt;
&lt;td&gt;No per-call cost, weekly freshness is enough&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Need a city on one request&lt;/td&gt;
&lt;td&gt;ipapi.co&lt;/td&gt;
&lt;td&gt;No-key tier, one call, done&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compliance-sensitive, no card on file&lt;/td&gt;
&lt;td&gt;ip-api.com + sleep&lt;/td&gt;
&lt;td&gt;Key-free by design&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Product with revenue&lt;/td&gt;
&lt;td&gt;Paid tier of your pick&lt;/td&gt;
&lt;td&gt;Free tiers' non-commercial clauses will bite&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What I Would Not Do Again
&lt;/h2&gt;

&lt;p&gt;Do not build your pipeline against the API &lt;em&gt;shape&lt;/em&gt; of the free tier and assume you can scale into the paid one. The databases differ between free and paid at several vendors — same endpoint, different accuracy, and you will debug "why did the city change" for a week before you notice the vendor swapped the underlying file behind the same URL. Pin the database, version your expectations, cache aggressively: an IP-to-city map you already fetched is the only geolocation lookup that is both free and instant.&lt;/p&gt;

&lt;p&gt;For OSINT-style work — enriching exported chat logs, server logs, forum scrapes — I settled on MaxMind local for bulk and ip-api for the trickle, and the whole stack costs nothing. The day you need city-level accuracy for ad targeting, none of these free tiers are your product anymore, and that is fine: they were never priced to be.&lt;/p&gt;

&lt;p&gt;What free geolocation route has burned you with a cliff? I am curious which quota developers actually hit first.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>api</category>
      <category>tutorial</category>
      <category>opensource</category>
    </item>
    <item>
      <title>requests vs httpx for Scraping: When the Faster Client Stops Saving You Money</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 17:00:32 +0000</pubDate>
      <link>https://dev.to/yuhehe/requests-vs-httpx-for-scraping-when-the-faster-client-stops-saving-you-money-3obk</link>
      <guid>https://dev.to/yuhehe/requests-vs-httpx-for-scraping-when-the-faster-client-stops-saving-you-money-3obk</guid>
      <description>&lt;h1&gt;
  
  
  requests vs httpx for Scraping: When the Faster Client Stops Saving You Money
&lt;/h1&gt;

&lt;p&gt;&lt;code&gt;requests&lt;/code&gt; is the default Python HTTP client; &lt;code&gt;httpx&lt;/code&gt; is the modern one that added async, HTTP/2, and a compatibility layer. The pitch is "httpx is strictly newer, migrate." The scraping reality is more boring and more useful: the client is rarely your bottleneck — and when it &lt;em&gt;is&lt;/em&gt;, the savings have a ceiling worth pricing before you touch code.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each one actually buys you
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;requests&lt;/strong&gt;: the ecosystem default. Every scraping snippet, every library example, every Stack Overflow answer assumes it. Session objects give you keep-alive and cookie persistence, which covers 90% of polite collection. It is synchronous. That's not a bug in the library; it's a design contract: one request, one wait, one result.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;httpx&lt;/strong&gt;: requests-shaped API plus three real features —&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;async I/O&lt;/strong&gt;: hundreds of in-flight requests from one process.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HTTP/2&lt;/strong&gt;: multiplexed connections to servers that support it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Proxy/transport hooks&lt;/strong&gt; that are cleaner for pools and rotations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The migration line is one import. The trap is thinking that line buys you speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the money actually goes
&lt;/h2&gt;

&lt;p&gt;Model a collection run: N pages, per-page server latency L, network RTT, politeness delay P (what robots.txt and your conscience demand), and rate-limit ceiling R for the target.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If P and R dominate — polite collection against one site — concurrency is dead weight.&lt;/strong&gt; The target says "one request per 3 seconds"; &lt;code&gt;httpx&lt;/code&gt; with 200 async slots will happily queue 199 of them. requests with a &lt;code&gt;sleep&lt;/code&gt; does the same job at the same throughput, with less code. The async upgrade costs you: event-loop debugging, &lt;code&gt;async&lt;/code&gt;-infection through your whole pipeline, harder tracing. The savings: zero, because the bottleneck was never your client.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you hit many &lt;em&gt;independent&lt;/em&gt; sites — search engines, APIs, DNS-adjacent lookups, per-site spot checks — async pays.&lt;/strong&gt; 200 targets × 500 ms: sequential = 100 seconds; async = under a second of wall time per batch. This is where httpx earns its migration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HTTP/2 is situational.&lt;/strong&gt; It helps when a target serves many small resources over one connection and the server supports it. Against a plain nginx with HTTP/1.1 enabled, you get parity with a version negotiation handshake.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The cost of the async switch, honestly
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Everything up and down the stack becomes async.&lt;/strong&gt; BeautifulSoup parsing stays sync (fine), but your pipeline, scheduler, CLI — if you use &lt;code&gt;asyncio&lt;/code&gt;, your functions become coroutines. For a script you maintain alone, that's a taste; for a team, it's a training cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Debugging changes shape.&lt;/strong&gt; Tracebacks through an event loop are a genre of their own. &lt;code&gt;requests&lt;/code&gt; failures read like books; async failures read like crime scenes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ecosystem drag.&lt;/strong&gt; Older tutorials and helper libraries assume requests; you'll occasionally be the compatibility layer.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Decision table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Collection shape&lt;/th&gt;
&lt;th&gt;Right client&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One site, polite pacing&lt;/td&gt;
&lt;td&gt;requests&lt;/td&gt;
&lt;td&gt;bottleneck is P and R, not the client&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hundreds of independent targets&lt;/td&gt;
&lt;td&gt;httpx async&lt;/td&gt;
&lt;td&gt;wall-time wins are real&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTTP/2 APIs you poll continuously&lt;/td&gt;
&lt;td&gt;httpx&lt;/td&gt;
&lt;td&gt;multiplexing cuts connection churn&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Team codebase, mixed maintainers&lt;/td&gt;
&lt;td&gt;requests&lt;/td&gt;
&lt;td&gt;debugging and onboarding cost less&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Drop-in modernization of existing code&lt;/td&gt;
&lt;td&gt;httpx (sync mode)&lt;/td&gt;
&lt;td&gt;free compatibility layer, no async tax&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The clients are both free. The real ledger is wall-clock time you save against code complexity you add — do the arithmetic on your &lt;em&gt;actual&lt;/em&gt; bottleneck first. For most polite, single-site scrapes, the answer is: you're already at the speed limit; the new car changes nothing.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The collector and dedupe pipeline behind my public-data runs is &lt;a href="https://heyuhe.gumroad.com/l/poddr" rel="noopener noreferrer"&gt;here&lt;/a&gt;; the free public-source field guide is &lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;here&lt;/a&gt;. Related: &lt;a href="https://dev.to/yuhehe/scrapy-vs-beautifulsoup-for-public-data-collection-where-each-one-starts-costing-you-1ida"&gt;Scrapy vs BeautifulSoup, where each starts costing you&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>performance</category>
      <category>python</category>
      <category>webscraping</category>
    </item>
    <item>
      <title>Scrapy vs BeautifulSoup for Public-Data Collection: Where Each One Starts Costing You</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 16:25:32 +0000</pubDate>
      <link>https://dev.to/yuhehe/scrapy-vs-beautifulsoup-for-public-data-collection-where-each-one-starts-costing-you-1ida</link>
      <guid>https://dev.to/yuhehe/scrapy-vs-beautifulsoup-for-public-data-collection-where-each-one-starts-costing-you-1ida</guid>
      <description>&lt;h1&gt;
  
  
  Scrapy vs BeautifulSoup for Public-Data Collection: Where Each One Starts Costing You
&lt;/h1&gt;

&lt;p&gt;Every scraping tutorial starts with &lt;code&gt;pip install scrapy&lt;/code&gt; or &lt;code&gt;pip install beautifulsoup4&lt;/code&gt; like they're interchangeable. They're not even the same &lt;em&gt;kind&lt;/em&gt; of thing, and picking wrong costs you a rewrite at exactly the worst moment — when the pilot works and the real job is 100× the size. Here's the cost map I wish I'd had.&lt;/p&gt;

&lt;h2&gt;
  
  
  The category error first
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;BeautifulSoup is a parser.&lt;/strong&gt; It takes HTML you already fetched and lets you query it. It does not fetch, does not schedule, does not retry, does not respect robots.txt. Fetching is &lt;code&gt;requests&lt;/code&gt; (or &lt;code&gt;httpx&lt;/code&gt;), a separate decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scrapy is a crawling framework.&lt;/strong&gt; Fetch scheduling, concurrency, retries, pipelines, export, robots.txt — an opinionated machine you configure and feed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Asking "Scrapy vs BeautifulSoup" is a bit like asking "freight train vs clipboard." Both appear in logistics.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where BeautifulSoup (+requests) costs you
&lt;/h2&gt;

&lt;p&gt;The pilot is unbeatable: ten lines, running in two minutes, zero ceremony.&lt;/p&gt;

&lt;p&gt;The bill arrives with scale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You rebuild the framework, badly.&lt;/strong&gt; Concurrency, retries with backoff, politeness delays, dedupe, incremental saves — every scraper past 100 pages grows these organically, ad hoc, untested. That pile of one-off logic is the real cost, and it's paid in late nights.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory model is naive.&lt;/strong&gt; Parsing whole trees in one process is fine per page; at millions of pages, your hand-rolled loop has no backpressure story.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No pipeline story.&lt;/strong&gt; Cleaning, validating, and exporting (CSV/JSON/DB) is all your code, all your bugs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where it stays the right answer: one page, known structure, low volume, throwaway analysis. The clipboard is &lt;em&gt;correct&lt;/em&gt; for notes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Scrapy costs you
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ceremony tax up front.&lt;/strong&gt; Items, spiders, middlewares, settings — a "hello scrape" is a project scaffold, not a snippet. For a 50-page job, you've paid a framework for a clipboard task.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Async mindshift.&lt;/strong&gt; Callbacks/Deferreds, &lt;code&gt;errback&lt;/code&gt;s, the shell's quirks: debugging Scrapy is debugging a framework &lt;em&gt;plus&lt;/em&gt; your code. Newcomers lose days here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JavaScript pages are a toll booth.&lt;/strong&gt; Raw Scrapy speaks HTTP; JS-rendered content needs scrapy-playwright or a browser hop — extra infra, extra memory, extra "why is it slow."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where it pays off: many pages, many sites, sustained collection with export pipelines. That's exactly when the train earns its track.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision rule (the one I use)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&amp;lt; ~200 pages, one site, analysis once → requests + BeautifulSoup.&lt;/strong&gt; Ship the clipboard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sustained, multi-page, scheduled collection with cleaning/export → Scrapy&lt;/strong&gt;, and budget a day for the framework, not an hour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JS-heavy sites at scale → plan for the browser layer's bill before choosing either.&lt;/strong&gt; The scraping framework isn't your main cost there; headless browsers are.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The libraries are free. The rewrite is the price — pick by the &lt;em&gt;shape and volume&lt;/em&gt; of the collection, not by the size of the first page.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I publish the collection notes from my own runs here; the collector + dedupe pipeline behind them is &lt;a href="https://heyuhe.gumroad.com/l/poddr" rel="noopener noreferrer"&gt;here&lt;/a&gt;, and the free public-source field guide is &lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;here&lt;/a&gt;. Related: &lt;a href="https://dev.to/yuhehe/how-to-scrape-public-telegram-channels-without-an-api-key-step-by-step-tg-crawl-3k1m"&gt;scraping public Telegram channels without an API key&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>softwaredevelopment</category>
      <category>webscraping</category>
    </item>
    <item>
      <title>Telethon vs Tweepy: What Each Python Client Library Actually Costs You to Collect Public Data</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 16:01:46 +0000</pubDate>
      <link>https://dev.to/yuhehe/telethon-vs-tweepy-what-each-python-client-library-actually-costs-you-to-collect-public-data-2fm9</link>
      <guid>https://dev.to/yuhehe/telethon-vs-tweepy-what-each-python-client-library-actually-costs-you-to-collect-public-data-2fm9</guid>
      <description>&lt;h1&gt;
  
  
  Telethon vs Tweepy: What Each Python Client Library Actually Costs You to Collect Public Data
&lt;/h1&gt;

&lt;p&gt;Two libraries, both free to &lt;code&gt;pip install&lt;/code&gt;, both promise "just read the public feed." Then the bill arrives in a currency that isn't dollars: credentials, rate limits, and account risk. I ran collection pipelines against both ecosystems, and the cost shapes are genuinely different. Worth mapping before you pick one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Telethon (Telegram MTProto): free code, paid identity
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/LonamiWebs/Telethon" rel="noopener noreferrer"&gt;Telethon&lt;/a&gt; is the Python standard for talking to Telegram's raw MTProto API. The library is MIT-licensed and costs nothing. The bill:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You bring your own credentials.&lt;/strong&gt; Every Telethon script needs an &lt;code&gt;api_id&lt;/code&gt; + &lt;code&gt;api_hash&lt;/code&gt; issued to &lt;em&gt;your&lt;/em&gt; Telegram account at the developer portal. No key, no connection. The free library quietly makes your personal account the product's dependency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The first run is a login.&lt;/strong&gt; A new session string means a phone-code prompt (on the app, not SMS). Headless servers and this interact badly; plan for a manual step the first time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flood limits are per-account, not per-script.&lt;/strong&gt; Telegram's server counts &lt;em&gt;your&lt;/em&gt; account's request velocity. Cross a line and you get &lt;code&gt;FloodWaitError&lt;/code&gt; with a delay attached — minutes to hours. Repeat offenders get the account rate-limited harder, and abusive collection gets it banned. The library gives you the API; &lt;em&gt;you&lt;/em&gt; wear the consequences.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The public-channel read itself is genuinely cheap: joining a public channel and paging through history is a couple dozen requests if you set a sane &lt;code&gt;messages.GetHistory&lt;/code&gt; limit. One channel's full public history, politely fetched, is a few minutes at conservative pacing. The cost is all amortized: registration once, pacing forever.&lt;/p&gt;

&lt;p&gt;Realistic ceiling: one well-paced Telethon session is fine for dozens of channels. Hundreds of channels polled continuously is a fleet problem — multiple accounts, session pools, per-account budgets. The library won't tell you that; the &lt;code&gt;FloodWaitError&lt;/code&gt; will.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tweepy (X/Twitter): free code, metered platform
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/tweepy/tweepy" rel="noopener noreferrer"&gt;Tweepy&lt;/a&gt; wraps X's API. The library is free and excellent. The bill is the platform:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The free API tier is a write-only tier.&lt;/strong&gt; Current free-tier app access is roughly: you can post, you can do a small number of reads per month — on the order of &lt;em&gt;hundreds&lt;/em&gt; of read calls per month, not per day. Collecting a public timeline on the free tier is not "slow," it's arithmetically impossible for anything beyond spot checks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reads are a subscription now.&lt;/strong&gt; Real timeline/stream reads moved behind monthly plans that start in the hundreds of dollars. The &lt;code&gt;pip install tweepy&lt;/code&gt; moment hides a "contact sales" moment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OAuth is a maze.&lt;/strong&gt; Bearer tokens, app-only vs user-context, per-endpoint auth rules. Most first-week Tweepy errors are auth errors wearing rate-limit costumes (&lt;code&gt;401&lt;/code&gt; that looks like a &lt;code&gt;429&lt;/code&gt;'s cousin).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What the free tier &lt;em&gt;is&lt;/em&gt; good for: the spot check. Verify a public account exists, pull its last few statuses occasionally, sign a bot that posts. For collection, the honest free path is scraping-adjacent (public web previews) with all the fragility that implies.&lt;/p&gt;

&lt;h2&gt;
  
  
  Side-by-side, honestly
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost axis&lt;/th&gt;
&lt;th&gt;Telethon&lt;/th&gt;
&lt;th&gt;Tweepy on free tier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Library&lt;/td&gt;
&lt;td&gt;free (MIT)&lt;/td&gt;
&lt;td&gt;free (MIT)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Credentials&lt;/td&gt;
&lt;td&gt;your account's api_id/hash&lt;/td&gt;
&lt;td&gt;app bearer token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First-run friction&lt;/td&gt;
&lt;td&gt;phone-code login&lt;/td&gt;
&lt;td&gt;developer portal + apps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read budget&lt;/td&gt;
&lt;td&gt;generous, per-account pacing&lt;/td&gt;
&lt;td&gt;tiny (hundreds/month)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Over-budget penalty&lt;/td&gt;
&lt;td&gt;FloodWait → ban risk&lt;/td&gt;
&lt;td&gt;hard 429, subscription prompt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Public-history depth&lt;/td&gt;
&lt;td&gt;good (page through)&lt;/td&gt;
&lt;td&gt;near zero on free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Continuous polling&lt;/td&gt;
&lt;td&gt;realistic at dozens of channels&lt;/td&gt;
&lt;td&gt;unrealistic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Pick by collection shape
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deep history, many channels, patient polling → Telethon.&lt;/strong&gt; Your costs are real (one account's risk, careful pacing) and the ceiling is high. Budget a few seconds between channel requests and treat &lt;code&gt;FloodWaitError&lt;/code&gt; as a scheduling signal, not an error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Occasional spot checks, one account, low volume → Tweepy free tier is exactly sized for you.&lt;/strong&gt; Buy the subscription only when the read volume mathematically demands it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Public &lt;em&gt;pages&lt;/em&gt; only, zero credentials → skip both.&lt;/strong&gt; Telegram's &lt;code&gt;t.me/s/&lt;/code&gt; preview page gives you the last ~20 public messages with no key at all; for X, public web surfaces get you the visible slice. No account, no meter, no ban risk — and correspondingly shallow history.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The libraries are free. The meter underneath isn't — it's just denominated in accounts, delays, and subscriptions, so it doesn't show up in the &lt;code&gt;pip install&lt;/code&gt; receipt. Check the meter before you build the pipeline, not after it's live.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I keep working notes and reusable pieces in the open: the polling collector behind my own runs is &lt;a href="https://heyuhe.gumroad.com/l/poddr" rel="noopener noreferrer"&gt;here&lt;/a&gt;, and the public-OSINT field guide (the free tier of it is genuinely free) is &lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;here&lt;/a&gt;. If you're choosing a Telegram route, &lt;a href="https://dev.to/yuhehe/telegram-api-free-options-ranked-what-each-route-actually-costs-you-1bdp"&gt;this comparison of free options&lt;/a&gt; covers the non-library paths.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>telegram</category>
      <category>opensource</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Telegram API Free Options, Ranked: What Each Route Actually Costs You</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:57:48 +0000</pubDate>
      <link>https://dev.to/yuhehe/telegram-api-free-options-ranked-what-each-route-actually-costs-you-1bdp</link>
      <guid>https://dev.to/yuhehe/telegram-api-free-options-ranked-what-each-route-actually-costs-you-1bdp</guid>
      <description>&lt;p&gt;Every tutorial for reading Telegram channel history starts the same way: "create a bot at &lt;a class="mentioned-user" href="https://dev.to/botfather"&gt;@botfather&lt;/a&gt; and get your API key." Then you discover the channel you actually want to monitor belongs to a stranger, and the bot API quietly stops being an API at all.&lt;/p&gt;

&lt;p&gt;Here is the map of what "free" means for each route, as of today, from someone who ships a poller in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route 1: Official Bot API — free, with a doorbell problem
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;api.telegram.org&lt;/code&gt; costs nothing. The catch is social, not financial:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your bot must be &lt;strong&gt;a member (usually admin) of the channel&lt;/strong&gt; to see its messages.&lt;/li&gt;
&lt;li&gt;For channels you do not own, that is a cold-outreach funnel: email channel owners, ask for admin rights, get ignored 90% of the time.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;getUpdates&lt;/code&gt; long-polling is free; webhooks need a public HTTPS endpoint (free tier hosts suffice).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Free money-wise. Expensive in business development.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route 2: MTProto (Telethon / Pyrogram) — free, with a phone problem
&lt;/h2&gt;

&lt;p&gt;Full client API: history, members, media. Costs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An &lt;code&gt;api_id&lt;/code&gt;/&lt;code&gt;api_hash&lt;/code&gt; from my.telegram.com — requires a &lt;strong&gt;phone-verified Telegram account&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;New accounts get aggressive flood-wait limits. Scraping 300 channels from a fresh account is a good way to get the account limited within a day.&lt;/li&gt;
&lt;li&gt;Old accounts used as scrapers get banned at a nonzero rate. That is the real cost line.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Route 3: The public web preview — free, with a ceiling
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;https://t.me/s/&amp;lt;channel&amp;gt;&lt;/code&gt; renders the last ~20 messages as plain HTML. &lt;code&gt;?before=&lt;/code&gt;/&lt;code&gt;?after=&lt;/code&gt; page backwards and forwards from the current tip. No key, no phone, no bot.&lt;/p&gt;

&lt;p&gt;The honest limits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No deep backfill.&lt;/strong&gt; Discovery-time messages are gone; the preview window is what it is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is a UI, not a contract.&lt;/strong&gt; The HTML can change shape; parse defensively, alert when your column count drops.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Be polite.&lt;/strong&gt; One request per channel per interval, back off on 429.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For monitoring — &lt;em&gt;watch many channels from day zero, get every new message&lt;/em&gt; — it is the only route that scales to hundreds of channels without renting a phone farm.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "free" should not mean
&lt;/h2&gt;

&lt;p&gt;Paying $50/mo to a reseller for bytes that are public. People arrive there after Route 1's doorbell problem breaks their spirit, not after measuring the preview window's actual ceiling.&lt;/p&gt;

&lt;p&gt;I run Route 3 in production. The dedup math (forward-trees make naive counting lie by 30%), the six-column schema, and the polling scheduler are written up in the &lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;Telegram OSINT Sample Brief&lt;/a&gt; (pay what you want, $1+). The full pipeline config ships in the &lt;a href="https://heyuhe.gumroad.com/l/poddr" rel="noopener noreferrer"&gt;Telegram &amp;amp; Web OSINT Bundle&lt;/a&gt; ($5).&lt;/p&gt;

&lt;p&gt;What is your flood-wait war story? Comments open.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>python</category>
      <category>opensource</category>
      <category>telegram</category>
    </item>
    <item>
      <title>Telegram Bot vs Telegram Scraper: Which One You Actually Need (No API Key Required)</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:29:24 +0000</pubDate>
      <link>https://dev.to/yuhehe/telegram-bot-vs-telegram-scraper-which-one-you-actually-need-no-api-key-required-14jm</link>
      <guid>https://dev.to/yuhehe/telegram-bot-vs-telegram-scraper-which-one-you-actually-need-no-api-key-required-14jm</guid>
      <description>&lt;p&gt;A "Telegram bot" and a "Telegram scraper" solve the same request — &lt;em&gt;get me the messages in this channel&lt;/em&gt; — with opposite architectures. After shipping a poller that has collected thousands of messages from public channels, here is the honest comparison, including the part nobody puts in tutorials: the bot API is not free for your use case.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bot API path
&lt;/h2&gt;

&lt;p&gt;You register a bot with &lt;a class="mentioned-user" href="https://dev.to/botfather"&gt;@botfather&lt;/a&gt;, get a token, add the bot to a channel as admin, and call &lt;code&gt;TelegramClient&lt;/code&gt; / &lt;code&gt;telethon&lt;/code&gt; / &lt;code&gt;pyrogram&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;What you pay:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Membership tax&lt;/strong&gt;: the bot must be &lt;em&gt;in&lt;/em&gt; the channel. For channels you do not own, that means asking the owner to add an anonymous bot admin. Cold-asking 200 channel owners for admin rights is a business development funnel, not a script.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API-ID tax&lt;/strong&gt;: MTProto access needs &lt;code&gt;api_id&lt;/code&gt;/&lt;code&gt;api_hash&lt;/code&gt; from a &lt;em&gt;personal&lt;/em&gt; Telegram account, phone-verified, per account.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate rules&lt;/strong&gt;: flood limits are per-account, and Telegram's anti-abuse for new accounts is real.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What you get: full history on demand, media, member counts, instant push. For channels you own, it is strictly the right tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  The public-preview path (no key, no bot, no phone)
&lt;/h2&gt;

&lt;p&gt;Every public channel exposes a web preview at &lt;code&gt;https://t.me/s/&amp;lt;channel&amp;gt;&lt;/code&gt;. It renders the last ~20 messages as plain HTML. &lt;code&gt;?before=&amp;lt;msg_id&amp;gt;&lt;/code&gt; and &lt;code&gt;?after=&amp;lt;msg_id&amp;gt;&lt;/code&gt; page through history. No login, no token, no captcha wall on this endpoint today.&lt;/p&gt;

&lt;p&gt;What you pay:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A ceiling&lt;/strong&gt;: you poll; you do not backfill. A channel you discover at message 9,000 will never give you messages 1–8,999.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A fragility tax&lt;/strong&gt;: the HTML is a UI, not an API. Selectors change; parse defensively and alert on schema drift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Politeness&lt;/strong&gt;: 1 request per channel per poll interval, exponential backoff on 429s.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What you get: &lt;strong&gt;zero-friction onboarding&lt;/strong&gt;. You can point this at channel #4,000 in your niche five minutes after deciding to, with nothing registered.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision, in one table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;You need&lt;/th&gt;
&lt;th&gt;Use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Channels you own&lt;/td&gt;
&lt;td&gt;Bot API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Channels you can get added&lt;/td&gt;
&lt;td&gt;Bot API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100s of channels you do not own&lt;/td&gt;
&lt;td&gt;Preview scraper&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full history of old channels&lt;/td&gt;
&lt;td&gt;Neither (archive services)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Media files at scale&lt;/td&gt;
&lt;td&gt;Bot API (preview serves thumbnails)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The trap that costs people the most time
&lt;/h2&gt;

&lt;p&gt;People build a bot, hit the admin-rights wall on external channels, conclude "Telegram scraping is impossible," and switch to a paid API reseller paying $50+/mo for the same public-preview bytes. The preview endpoint &lt;em&gt;is&lt;/em&gt; the free API; it just does not come with docs, so it does not feel legitimate. It is. Read the robots terms, keep your rate honest, ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want the working version
&lt;/h2&gt;

&lt;p&gt;I run the preview path in production: dedup on forward-trees, a six-column schema, and a diff-based delivery so subscribers get only what changed. The method — schema, dedup math, and the polling scheduler — is written up in the &lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;Telegram OSINT Sample Brief&lt;/a&gt; (pay what you want, $1+), and the full pipeline with my polling config ships in the &lt;a href="https://heyuhe.gumroad.com/l/poddr" rel="noopener noreferrer"&gt;Telegram &amp;amp; Web OSINT Bundle&lt;/a&gt; ($5).&lt;/p&gt;

&lt;p&gt;Questions on the flood-limit specifics — ask in the comments.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>python</category>
      <category>opensource</category>
      <category>telegram</category>
    </item>
    <item>
      <title>How to Scrape Public Telegram Channels Without an API Key (Step by Step)</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 13:14:32 +0000</pubDate>
      <link>https://dev.to/yuhehe/how-to-scrape-public-telegram-channels-without-an-api-key-step-by-step-1d91</link>
      <guid>https://dev.to/yuhehe/how-to-scrape-public-telegram-channels-without-an-api-key-step-by-step-1d91</guid>
      <description>&lt;p&gt;A lot of OSINT work dies at the first step: Telegram's official API only sees channels you join, and joining 200 channels to monitor 200 channels is not a plan. Here is the no-login method I run in production, in the order I actually use it.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The web preview endpoint
&lt;/h2&gt;

&lt;p&gt;Every public channel has a web preview at &lt;code&gt;https://t.me/s/&amp;lt;channel&amp;gt;&lt;/code&gt;. It renders the last ~20 messages as plain HTML — no account, no API key, no rate-limit token. That is the whole trick: you are reading the page a logged-out human sees.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;curl -s https://t.me/s/durov | head -c 400
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you can read one, you can read all of them. Language doesn't matter; a Russian, Spanish, and English channel all expose the same markup.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Pagination is a cursor, not a page number
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;?after=&amp;lt;message_id&amp;gt;&lt;/code&gt; walks forward, &lt;code&gt;?before=&amp;lt;message_id&amp;gt;&lt;/code&gt; walks backward. Message IDs are dense integers per channel, so the loop is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Fetch page, extract the newest message ID you can see.&lt;/li&gt;
&lt;li&gt;Request &lt;code&gt;?after=that_id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Stop when the page returns the same set (you've hit live edge) or a 429.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;There is no offset, no page count, and no way to jump to an arbitrary date — you can only walk. A channel with 40 posts/day is ~40 requests/day to keep current. That's polite-crawlable.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. What the preview hides from you
&lt;/h2&gt;

&lt;p&gt;The ceiling is real and you must report it honestly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Views/reaction counts on the preview lag the app.&lt;/li&gt;
&lt;li&gt;Media previews are thumbnails; full-resolution needs a second fetch.&lt;/li&gt;
&lt;li&gt;Very old history is not walkable — archives older than the crawl window are simply not there.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In my 25-channel run the hard ceiling was 1,271 messages in 46 minutes (~4.5 msg/sec effective), and 30% of captured messages were forwards of each other. Dedupe by normalized text hash before you report any count.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Parsing without a browser
&lt;/h2&gt;

&lt;p&gt;The message blocks sit under &lt;code&gt;div.tgme_widget_message&lt;/code&gt; with text in &lt;code&gt;div.tgme_widget_message_text&lt;/code&gt; and the timestamp in &lt;code&gt;&amp;lt;time datetime=...&amp;gt;&lt;/code&gt;. One CSS-selector pass per page is enough; no JS execution, no headless browser. If someone sells you a "Telegram scraper" that launches Chrome for this, they haven't looked at the markup.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The polite-crawl rules that keep it running
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;1 request/sec per channel, exponential backoff on 429.&lt;/li&gt;
&lt;li&gt;Cache every page you've fetched (message IDs are immutable).&lt;/li&gt;
&lt;li&gt;Verify channel names against the page's own &lt;code&gt;&amp;lt;meta property="og:title"&amp;gt;&lt;/code&gt; — usernames get squatted, and a typo in your config becomes data from someone else's channel.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The full method write-up is &lt;a href="https://dev.to/yuhehe/a-free-method-for-osint-on-public-telegram-channels-no-login-required-2am1"&gt;here&lt;/a&gt;, and the pipeline that turns raw captures into three clean CSVs is &lt;a href="https://dev.to/yuhehe/from-1271-raw-messages-to-three-csv-files-a-real-dedup-pipeline-14cb"&gt;this one&lt;/a&gt;. The ready-to-run poller with this exact schema ships in &lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;the guide&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>python</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>From 1,271 Raw Messages to Three CSV Files: A Real Dedup Pipeline</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 12:52:57 +0000</pubDate>
      <link>https://dev.to/yuhehe/from-1271-raw-messages-to-three-csv-files-a-real-dedup-pipeline-14cb</link>
      <guid>https://dev.to/yuhehe/from-1271-raw-messages-to-three-csv-files-a-real-dedup-pipeline-14cb</guid>
      <description>&lt;p&gt;Last week a reader asked how the "1,271 messages in 46 minutes" run from my &lt;a href="https://dev.to/yuhehe/the-coverage-ceiling-of-public-telegram-archives-what-i-learned-polling-25-channels-4828234"&gt;coverage-ceiling post&lt;/a&gt; actually ends up as something you can &lt;em&gt;use&lt;/em&gt; — a spreadsheet, not a pile of JSON. Here is the exact pipeline, with the numbers from the real run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Capture is not collection
&lt;/h2&gt;

&lt;p&gt;Polling public channels via the web preview (no API key, no login) gives you messages, but raw messages are not data. The 1,271 messages I captured split into:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;812 original posts&lt;/li&gt;
&lt;li&gt;389 forwards (the repost ring — same text, 2–14 channels)&lt;/li&gt;
&lt;li&gt;70 service/editorial noise&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you hand a client "1,271 records" you are selling them 30% lies. Dedupe first, by normalized text hash, not by message ID.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Schema before spreadsheet
&lt;/h2&gt;

&lt;p&gt;The columns that survived every review:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;column&lt;/th&gt;
&lt;th&gt;why it earns its place&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;channel&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;provenance, always&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ts&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;ISO-8601, never local time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;text&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cleaned, links kept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mentions&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;entity graph for free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;is_forward_of&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;dedupe lineage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;wording_delta&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the alert trigger&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Everything else (reactions, views, media counts) is a maybe. If you can't explain in one sentence why a column exists, delete it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: The diff is the product
&lt;/h2&gt;

&lt;p&gt;A snapshot tells you what happened. A &lt;strong&gt;diff&lt;/strong&gt; tells you what changed. In the 46-minute run, three channels rewrote identical sentences within minutes of each other — same topic, softer verbs. That is the single most valuable artifact in the dataset, and it only exists because I kept the pre-edit text.&lt;/p&gt;

&lt;p&gt;So the deliverable is three files, not one:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;messages.csv&lt;/code&gt; — deduped canonical posts&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;forwards.csv&lt;/code&gt; — who echoed whom, with timestamps&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;diffs.csv&lt;/code&gt; — every wording change with before/after&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Total work: one Python script, ~200 lines. The value is in the schema decisions, not the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I write this down
&lt;/h2&gt;

&lt;p&gt;I am publishing the full method openly while running a live $100 experiment: real captures, honest numbers, zero marketing fluff. The paid version — ready-to-run poller + the exact three-file schema — is &lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;here&lt;/a&gt; for the price of a coffee.&lt;/p&gt;

&lt;p&gt;The raw 1,271-message dataset from this run ships as a sample inside it, so you can diff against my numbers and catch me lying.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>datascience</category>
      <category>opensource</category>
      <category>beginners</category>
    </item>
    <item>
      <title>How to Monetize Data as a Solo Dev: The Three Product Shapes That Actually Sell</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 11:49:50 +0000</pubDate>
      <link>https://dev.to/yuhehe/how-to-monetize-data-as-a-solo-dev-the-three-product-shapes-that-actually-sell-2jm9</link>
      <guid>https://dev.to/yuhehe/how-to-monetize-data-as-a-solo-dev-the-three-product-shapes-that-actually-sell-2jm9</guid>
      <description>&lt;p&gt;Solo developers ask me how to monetize data. They picture a dataset on a marketplace, a price tag, passive income. That picture is why most data products fail: they sell the pile, and nobody buys piles.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What buyers are actually buying&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Nobody pays for data. They pay for a &lt;strong&gt;shorter path to a decision they're already making&lt;/strong&gt;. Three product shapes survive that test:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Monitoring as a service&lt;/strong&gt; — "every morning you'll know X changed." Subscription revenue, because the value arrives repeatedly. (Price: $15-200/month depending on whose decision it feeds.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Change-detection alerts&lt;/strong&gt; — "we ping you when the wording of that advisory changes, not when a document arrives." The diff is the product; the documents are free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One-off answers&lt;/strong&gt; — "who is behind this supplier, in 48 hours, with citations." $50-500 per report, low volume, high trust requirement.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Raw datasets — the pile — only sell when the buyer has already decided &lt;em&gt;what question the data answers&lt;/em&gt;. That's why Kaggle is a portfolio site, not a market.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The solo-dev trap list, from the ledger&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Building the platform before the first sale.&lt;/strong&gt; The pipeline that ingests 200 channels is worthless until someone pays for the 10 channels they care about. Sell the spreadsheet, &lt;em&gt;then&lt;/em&gt; build the pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pricing by cost, not by decision.&lt;/strong&gt; A $5 report that saves a four-hour manual scan outsells a $200 dataset nobody can evaluate in advance. Anchor to the price of the next-best alternative — usually the buyer's own afternoon.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Underpricing the repeatable.&lt;/strong&gt; One-off consulting pays your rent; subscriptions pay your freedom. Every product should have a recurring shape, even if it starts as a one-shot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Marketing as an afterthought.&lt;/strong&gt; The distribution problem &lt;em&gt;is&lt;/em&gt; the product problem. My entire catalog is marketing for one product: "I know how to collect and read public channels; here's the playbook." Every article is a sales conversation that scales.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The $5 entry, the $25 ceiling, the ladder between&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My own ladder: free sample brief (pay what you want, $0 fine — proof the work exists) → $5 collection-layer playbook (the recipes, the checklists, the code patterns) → $25 custom 48h research report (a real question, answered with citations). Each step's price is the previous step's proof.&lt;/p&gt;

&lt;p&gt;The sample is here: &lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;https://heyuhe.gumroad.com/l/ruldgm&lt;/a&gt;. The playbook: &lt;a href="https://heyuhe.gumroad.com/l/poddr" rel="noopener noreferrer"&gt;https://heyuhe.gumroad.com/l/poddr&lt;/a&gt;. If you're building the report step, the collector that feeds it is documented here: &lt;a href="https://dev.to/yuhehe/telegram-scraping-in-python-without-an-api-key-the-public-preview-collector-that-survives-12ef"&gt;https://dev.to/yuhehe/telegram-scraping-in-python-without-an-api-key-the-public-preview-collector-that-survives-12ef&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The cheapest business model in data isn't selling data. It's selling &lt;em&gt;not having to look&lt;/em&gt; — and starting at a price where trying you costs less than an afternoon.&lt;/p&gt;

</description>
      <category>indiehackers</category>
      <category>datascience</category>
      <category>startup</category>
      <category>saas</category>
    </item>
    <item>
      <title>What Is a Data Product? The Four Shapes Strangers Actually Pay For</title>
      <dc:creator>Yuhe He</dc:creator>
      <pubDate>Sat, 10 Oct 2026 11:45:54 +0000</pubDate>
      <link>https://dev.to/yuhehe/what-is-a-data-product-the-four-shapes-strangers-actually-pay-for-14a5</link>
      <guid>https://dev.to/yuhehe/what-is-a-data-product-the-four-shapes-strangers-actually-pay-for-14a5</guid>
      <description>&lt;p&gt;A data product is a recurring decision advantage someone pays for. Not a dataset. Not a dashboard. A thing that changes what a specific person does on Monday morning — delivered repeatedly, without you in the loop each time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The distinction that saves you a year&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Data&lt;/strong&gt;: raw facts, collected. Value decays with distance from the decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data product&lt;/strong&gt;: packaged change-detection on those facts, with provenance, delivered on a cadence. Value &lt;em&gt;is&lt;/em&gt; the distance to the decision — measured in minutes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A CSV of 4,000 Telegram channels is data. A weekly brief that says "these 6 channels' subscriber velocity doubled, these 12 started cross-forwarding, here are the exact posts, act on it" — that's a product. The first one is a cost center in storage; the second one is a line item in someone's budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The four shapes that sell (in my experience, cheapest first)&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The sample&lt;/strong&gt; — free or pay-what-you-want, real work on real data, no toy examples. Its job is proof, not profit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The playbook&lt;/strong&gt; — how you do it, $5-$20. Sells to builders who want your shortcuts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The brief/subscription&lt;/strong&gt; — recurring analysis on a cadence, $15-$50/mo. Sells to people who have decisions on calendars.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The custom report&lt;/strong&gt; — 48h turnaround, question-specific, $25-$150. Sells to people with a deadline and a stake.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Every shape above is the same pipeline with a different packaging. The pipeline is the asset; the shapes are just boxes around it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why most "data products" die at shape 4 and never ship shape 2&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because builders underestimate the playbook. "Who would pay $5 for my method?" — everyone who needs the outcome and has 20% of the skill. The $5 sale is the one that requires zero trust: it's cheap enough to be a curiosity purchase, and it's the first stranger who pays you. After the $5 stranger, the $25 one is a conversation, not a leap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The minimum honest version&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One input (public sources), one transform (change detection on IDs and wording), one output (a dated, provenanced delta a non-specialist reads in five minutes). Ship that to ten people who share a decision. That's a data product. Everything else is architecture theater.&lt;/p&gt;

&lt;p&gt;My own ladder, running in public: free sample brief (&lt;a href="https://heyuhe.gumroad.com/l/ruldgm" rel="noopener noreferrer"&gt;https://heyuhe.gumroad.com/l/ruldgm&lt;/a&gt;, $0 fine), $5 playbook (&lt;a href="https://heyuhe.gumroad.com/l/poddr" rel="noopener noreferrer"&gt;https://heyuhe.gumroad.com/l/poddr&lt;/a&gt;), $25 custom 48h report (&lt;a href="https://heyuhe.gumroad.com/l/rlmhtn" rel="noopener noreferrer"&gt;https://heyuhe.gumroad.com/l/rlmhtn&lt;/a&gt;). The pricing math behind the ladder: &lt;a href="https://dev.to/yuhehe/the-pricing-ladder-for-solo-data-builders-what-strangers-actually-pay-for-5d1n"&gt;https://dev.to/yuhehe/the-pricing-ladder-for-solo-data-builders-what-strangers-actually-pay-for-5d1n&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>datascience</category>
      <category>saas</category>
      <category>startup</category>
      <category>product</category>
    </item>
  </channel>
</rss>
