DEV Community

Jefferson Valandro
Jefferson Valandro

Posted on Fully Autonomous

7 tricks and 3 traps from scraping Twitch, YouTube, TikTok, Kick, Bluesky and Patreon with no login and no browser

Over the last few weeks I built a family of scrapers for creator data: profiles, followers, social links and business contacts. My rule was no headless browser and no login, plain HTTP only. That makes them cheaper, faster and less fragile. Most of this isn't documented anywhere, so here is what worked, what didn't, and the numbers.

1. Twitch: the public GraphQL does almost everything, except pagination

The Twitch website talks to gql.twitch.tv/gql with a public Client-Id (the web player's). With it you get profiles, followers, live status, VODs, clips and even the channel panels, which is where streamers put their business email.

The trap: paginating with after: returns failed integrity check, because it needs an integrity token generated in the browser. The workaround is to not paginate. Each query accepts first: 100, so I run one query per language and take the 100 most-watched streams of each.

In practice, for VALORANT + Portuguese, 6 out of 10 streamers had a contact email in their panels.

2. Bluesky: search works without login, but only on one host

public.api.bsky.app serves profiles, feeds and followers, but searchPosts returns 403 there. On api.bsky.app the same search works unauthenticated. The cursor for page two, though, returns 403.

The workaround is to walk back in time: the next call uses until = the oldest indexedAt of the previous page. That gets thousands of posts per keyword, no login.

3. YouTube: the About tab is already in the HTML

/@channel/about ships an aboutChannelViewModel inside ytInitialData, with subscribers, views, country, join date, links and the description. It also has a signInForBusinessEmail field. The business email itself sits behind login and captcha, but that field tells you whether the channel has one.

Channel search uses the sp=EgIQAg== filter plus youtubei/v1/search continuations. Two gotchas:

  • the right continuation token is inside continuationItemRenderer, not the first continuationCommand you find;
  • in the new layout, the subscriber count comes in a field called videoCountText. Really.

4. TikTok: profiles are easy, video lists are not

The whole profile is in <script id="__UNIVERSAL_DATA_FOR_REHYDRATION__">: followers, likes, bio link and business category. So are the stats of a single video, from its URL. A profile's video list needs a signature (X-Bogus), and I skipped it.

Fun fact: from the cloud, residential proxies got an empty page, while datacenter IPs worked.

5. Kick, Linktree and Patreon

  • Kick: kick.com/api/v2/channels/<slug> returns the channel and its social links. web.kick.com/api/v1/livestreams lists live streams with a real cursor.
  • Linktree: everything is in __NEXT_DATA__, including country, plan and page creation date.
  • Patreon: there is a public JSON:API with search and membership tiers. Patron counts show up in search results even when the creator hides them on their page.

Trap 1: Cloudflare looks at your HTTP library

On Patreon, the same URL from the same cloud IP returned 200 with requests and 403 ("Just a moment...") with httpx. That's the TLS fingerprint. So "it worked in my test" means nothing if the test used a different library.

Trap 2: memory with 2.7 million rows

Another product of mine matches Google Maps places to Brazil's company registry. Searching São Paulo loaded the whole city (2.7M companies) and blew past 4 GB.

The fix was a two-pass filter in PyArrow:

  1. count how many companies in the city contain each searched word;
  2. keep only rows with the same phone, the same postal code, or rare words (name matching gives up on words with more than 400 occurrences anyway).

Before shipping, I diffed old against new on 210 places: zero differences. Memory went from 3.9 GB (crashing) to 350 MB.

Trap 3: marketplace search rewards whoever got there first

I publish these on the Apify Store, priced per result. Plenty of people open the pages, about 15% open the input form, and half of those run it. But in the store's own search, new products barely show up, because ranking weighs users and reviews. Outside posts brought the visitors.

The first scraper outside my Brazilian-company niche to get a paying customer was the Twitch one, on the day it went live.


If you want to try them, they're on Apify (the free plan includes $5 of monthly credit):

Top comments (0)