DEV Community

Alom Dev
Alom Dev

Posted on

Scraping Bilibili comments: the 3-comment limit was my own fault

If you've tried to pull comments from Bilibili (B站) without logging in, you've probably hit the same wall I did:
every video gives you 3 comments. Not 20, not "the first page". Three.

I believed this for a while. Guests get the top 3, log in for the rest. It sounded plausible, Chinese platforms are
strict about logged-out access, so I wrote it into my README as a known limitation and moved on.

Then, while working on something else, a request went out without my usual cookies and came back with 20 comments.

What was actually going on

To look like a normal browser, I was sending the visitor cookies Bilibili hands out on the first page load (buvid3
and friends). Most scraping guides tell you to do this, and for search and video details it's fine.

For the comment API, those cookies mark you as a guest with a session, and guests with a session get a preview: 3
top comments, then a prompt to log in. A request with no cookies at all is treated like an anonymous API call and
gets normal pages of 20.

I tested it properly to make sure it wasn't a fluke. Without cookies the paging just kept going: about 1,200 unique
comments across the videos I tried, no login anywhere.

The fix was literally to stop sending something. That stung a bit.

The other things that tripped me up

WBI signing. Most Bilibili endpoints want a w_rid signature built from two keys in the nav response, mixed
with a fixed shuffle table and a timestamp. Skip it and you get error codes instead of data. The keys rotate, so don't
hardcode them.

The 1,000-result search cap. Search stops at about 50 pages no matter what. If you need more videos for a keyword,
sort by publish date and walk backwards in date windows. Each window gets its own 1,000. Slow, but it works.

Danmaku (弹幕) pools are capped. The bullet comments you see flying over the video are a separate XML or protobuf
feed. For popular videos Bilibili only keeps a slice of them. In one test the video had 128,575 danmaku and the pool
held 3,600. If you're doing "danmaku per minute" analysis, keep that in mind, the counts are a sample, not the whole
thing.

Reply threads. The main comment list shows only a few replies under each comment. To get the full thread you page
a second endpoint per root comment. Easy to miss, and it's where the interesting arguments live.

If you just want the data

Disclosure: after all of this I packaged it as a Bilibili scraper on Apify,
so if you don't want to maintain the signing and paging yourself, this does it. No login, no cookies, and the field
names match the most common Bilibili scraper, so switching is easy.

Comments with full reply threads for a few videos:

{
  "mode": "video_comments",
  "videoUrls": ["https://www.bilibili.com/video/BV1LraH6qEr5", "av170001"],
  "includeReplies": true,
  "maxComments": 200,
  "sortComments": "hot"
}
Enter fullscreen mode Exit fullscreen mode

Search plus comments under each result:

{
  "mode": "search",
  "searchQuery": "人工智能",
  "maxResults": 100,
  "includeComments": true,
  "maxComments": 50
}
Enter fullscreen mode Exit fullscreen mode

b23.tv short links work as input too. Pricing is $4 per 1,000 rows on Apify's free plan and goes down to $2 on the
bigger plans. A danmaku profile (the per-minute reaction summary for a video) is an optional extra at $0.02 per video.

If you'd rather build it yourself

Totally doable. The short version of everything above:

  1. Don't send visitor cookies to the comment API.
  2. Sign requests with WBI and refresh the keys from nav regularly.
  3. Use date windows to get past the search cap.
  4. Page reply threads separately.
  5. Treat danmaku counts as a sample on popular videos.

That's most of the pain. Took me longer than I'd like to admit, mostly because of step 1.

Top comments (0)