Comments are the cheapest audience research you have, and the hardest to read. A video with a few thousand of them is too many to skim, and a single sentiment score from a library will mislead you. This article covers why the scores go wrong, what to do about it, and then a short pipeline: collect comments from the official API, classify them in Python, summarise per video.
Why comment sentiment is noisy
Most sentiment tools were built on reviews or tweets and then pointed at YouTube. Four things break them:
- Sarcasm. "Great, another ten-minute sponsor read" contains a positive word and a negative meaning. Lexicon tools score the words, not the intent.
- Emoji and slang. "This is sick" and a row of skull emoji are praise. Some tools cover part of this, many ignore it.
- Short text. "first" and "lol" carry almost no sentiment, but they still count in your percentages.
- Spam. Link drops, "check my channel" and copy-pasted replies are not audience opinion.
There is a fifth, quieter problem: context. A reply of "same" means nothing without the comment above it.
What to do about it
- Clean before you score. Drop the creator's own comments, anything with a link, very short comments and exact duplicates.
- Compare, do not read absolutes. A tool that is biased in a consistent way still shows whether this video is better or worse than your last five.
- Weight by likes. A negative comment with many likes speaks for more viewers than one with none.
- Check against a hand-labelled sample. Label 50 comments yourself and see how often the tool agrees. If it is poor, swap in a better model or an LLM with a fixed three-label prompt and run the same check.
- Read what is behind the number. Always look at the most-liked comments in each bucket.
Getting the comments
The official YouTube Data API v3 returns comments, but you handle paging, turning a channel into its recent videos, replies and quota yourself. Best Damn YouTube Comments Scraper is an Apify Actor that wraps the API and returns one flat record per comment. This input collects comments for sentiment work, from one video plus a channel's three most recent uploads:
{
"startUrls": ["https://www.youtube.com/watch?v=jNQXAC9IVRw", "https://www.youtube.com/@YouTube"],
"maxVideosPerChannel": 3,
"maxComments": 500,
"authorData": "hash"
}
Replace the channel with your own @handle. maxComments applies per video, maxVideosPerChannel sets how many recent uploads are read, and the default sort is top comments, which the Actor's docs recommend for sentiment work. authorData: "hash" replaces each commenter's channel ID with a stable hash and drops name and avatar, so you can still see when one account posts again and again without storing who it is.
Save that as input.json and call the API:
curl -X POST "https://api.apify.com/v2/acts/josh99smith~youtube-comments-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d @input.json -o comments.json
Each delivered record has text, likeCount, replyCount, isReply, parentCommentId, isByVideoOwner, publishedAt, videoId, videoTitle and commentUrl, plus success: true. isByVideoOwner still works with pseudonymised authors.
Classify and summarise
This uses VADER (pip install vaderSentiment), a lexicon tool that handles some slang and emoji. It is generic: swap sia.polarity_scores for any model.
import json
from collections import Counter, defaultdict
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer
rows = [r for r in json.load(open("comments.json")) if r["success"]]
seen, comments = set(), []
for r in rows: # clean
text = r["text"].strip()
if r["isByVideoOwner"] or "http" in text or len(text) < 4 or text.lower() in seen:
continue
seen.add(text.lower())
comments.append(r)
sia = SentimentIntensityAnalyzer()
for r in comments: # score: +/-0.05 is VADER's usual cut-off
c = sia.polarity_scores(r["text"])["compound"]
r["label"] = "positive" if c >= 0.05 else "negative" if c <= -0.05 else "neutral"
by_video = defaultdict(list)
for r in comments:
by_video[r["videoTitle"]].append(r)
for title, items in by_video.items(): # summarise
count, weight = Counter(), Counter()
for r in items:
count[r["label"]] += 1
weight[r["label"]] += 1 + r["likeCount"]
print(title, len(items), "comments")
for lab in ("positive", "neutral", "negative"):
print(f" {lab:9}{count[lab] / len(items):4.0%} of comments {weight[lab] / sum(weight.values()):4.0%} like-weighted")
negative = sorted((r for r in comments if r["label"] == "negative"), key=lambda r: r["likeCount"], reverse=True)
for r in negative[:5]: # read the evidence
print(r["likeCount"], r["text"][:100], r["commentUrl"])
The output is a three-line table per video and the five most-liked negative comments with links back to YouTube. Run it before and after a change, such as a new format or sponsor, and compare. I am not showing results here because they depend entirely on your channel.
Cost and limits
- Price: $0.40 per 1,000 comments, charged per delivered comment or reply. The input above reads at most 4 videos (one video plus 3 from the channel) at 500 comments each, so up to 2,000 comments, or $0.80. You can set a maximum cost per run.
-
Free: videos with comments off, private or deleted videos, invalid URLs and quota errors are not charged. A video with comments off comes back as
{"success": false, "errorType": "comments-disabled", ...}. Made-for-kids videos never have comments. -
Quota: by default the Actor uses a shared API key. Each key has a daily quota of 10,000 units, and one page of up to 100 comments costs one unit, shared with every other user. For regular or large runs, create your own free key in Google Cloud Console and paste it into the
apiKeyinput, which the docs put at roughly a million comments a day. When a quota runs out, the run stops cleanly and the remaining videos are returned as freerate-limitedrecords. It resets at midnight Pacific Time. -
Replies: off by default. Set
includeRepliestotrueto get them; they carryisReplyandparentCommentId, count towardmaxCommentsand are billed like any comment. - Coverage: public videos only, no login, so no members-only content. You may get fewer comments than the count YouTube shows, because that count includes replies and comments held for review or removed as spam.
-
Privacy: comments are personal data under GDPR and similar laws. Use
hashoromitforauthorDataif you do not need to know who wrote what.
If this saves you time, a review on the Store page helps other people find it, and a star on GitHub helps too.
Disclosure: I built this Actor. Source: github.com/josh99smith/youtube-comments-scraper. YouTube is a trademark of Google; the Actor is not affiliated with Google.
Top comments (0)