DEV Community

Alom Dev
Alom Dev

Posted on

Scraping YouTube comments without the API quota (and the bug where 300 comments meant one thread)

The YouTube Data API gives you 10,000 quota units a day by default. Reading comments costs 1 unit per page of up to
100 comments, which sounds generous until you need replies too, or more than a handful of videos, or a second
project running on the same key. People who analyse comments at any scale hit that wall pretty quickly.

I built a comments scraper that reads what youtube.com shows a logged-out visitor instead, so there is no key and no
quota. A few things I learned on the way, in case you are building something similar.

"300 comments" should mean 300 comments

My first version had one limit, maxComments, and replies counted toward it. Seemed reasonable. Then a test asked
for 300 comments with replies on a popular music video and got back 1 comment and its 299 replies.

Of course it did. Replies are delivered right after their parent, and the most liked comment on a big video has
thousands of them. The limit was used up by one thread.

Nobody who types "300" means that. The fix was two limits:

  • maxComments counts top-level comments only
  • maxRepliesPerComment caps the replies under each one (default 5)

The second one matters for cost too. If replies were unlimited, 200 comments could turn into 4,000 rows and a bill
nobody expected. With the default, 100 comments return at most 600 rows, and the input form says so.

Top comments vs newest

YouTube has two orders, Top and Newest. Top is what you want for "what do people think". Newest is what you want
for monitoring, for example comments on a launch video since yesterday. A date filter on top of Newest ("only the
last 30 days") lets the scraper stop early instead of paging through years of comments.

The timeouts that weren't network problems

This one took me a while. Runs with a short time limit kept ending as TIMED-OUT instead of stopping cleanly and
saving what they had. The code had a deadline check that should stop work 25 seconds before the limit. It didn't
fire in time.

The cause: on Apify, 256 MB of memory also means about 1/16 of a CPU core. Parsing a lot of YouTube pages at once
kept that tiny CPU so busy that the timer meant to stop the run fired late, sometimes by 20+ seconds. Raising the
minimum to 512 MB (twice the CPU) fixed it completely, and the extra cost is still a fraction of a cent per thousand
rows. If your scraper misses its own deadlines, check the CPU before you blame the network.

The tool

Disclosure: this is the YouTube Comments Scraper I published on
Apify. Paste video or Shorts URLs, pick Top or Newest, optionally replies and a date window. It costs $1.00 per
1,000 comments on Apify's free plan and $0.45 on the bigger plans.

{
  "startUrls": [{ "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ" }],
  "maxComments": 300,
  "sortCommentsBy": "TOP_COMMENTS",
  "includeReplies": true,
  "maxRepliesPerComment": 5
}
Enter fullscreen mode Exit fullscreen mode

There are also two siblings built on the same code: a channel scraper
that exports every video of a channel with its stats, and a Shorts scraper
for Shorts by channel or keyword.

If you'd rather stay on the official API, that's a perfectly good choice for small jobs. Just budget your quota for
replies, they add up faster than you think.

Top comments (0)