I needed the comment section of a YouTube video for a small project, without setting up a Google Cloud project. Here is the logged-out route that actually survived a week of testing, and the parts that did not.
Everything below reads content that is already publicly visible in a browser. If you need scale or guarantees, use the official Data API. This is for the case where you just need to read what people wrote.
Why not just use an Invidious mirror
Mirrors are the first thing everyone reaches for, and they are fine for a quick look. But they are operated by volunteers with rate limits, and their comment caches can lag days behind the video. I built a reading on one public instance and it did not survive a direct check against the live page: replies and hearts were missing, and the ordering was stale. If you use a mirror, treat the number as a hint, not a fact, and re-check against the page.
What worked: the innertube next endpoint
YouTube's web player loads comments through an internal "innertube" endpoint, not the Data API. The trick is finding the right continuation token in the watch page HTML.
- Fetch the watch page with a desktop user-agent.
- Pull the comment
continuationCommandtoken out of the embedded JSON. - Base64-decode it far enough to see what it points at. A token for the comments section contains the string
comments-section; a token for the engagement panel next to the player containsengagement-panel. It is easy to grab the wrong one and get zero comments back. - POST that token to the innertube
nextendpoint and read the comments out of the response.
Roughly:
import json, re, base64, urllib.request
UA = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 " \
"(KHTML, like Gecko) Chrome/126 Safari/537.36"
def watch_html(video_id):
req = urllib.request.Request(
f"https://www.youtube.com/watch?v={video_id}",
headers={"User-Agent": UA},
)
return urllib.request.urlopen(req).read().decode("utf-8", "replace")
def comment_token(html):
for tok in re.findall(r'"continuationCommand":\{"token":"(.*?)"', html):
pad = tok + "=" * (-len(tok) % 4)
try:
raw = base64.urlsafe_b64decode(pad)
except Exception:
continue
if b"comments-section" in raw:
return tok
return None
Then POST {"context": ..., "continuation": token} to https://www.youtube.com/youtubei/v1/next?key=.... The comments come back as commentThreadRenderer items.
The gotchas that cost me the most time
Bot-gating from data-center IPs. Requests from cloud or sandbox IPs frequently come back as LOGIN_REQUIRED ("Sign in to confirm you're not a bot"). A desktop user-agent on the watch-page fetch helps; the innertube call itself can still be gated from some IP ranges. Test from the network you will actually run in before you build anything on top.
The wrong token. As above: comments-section vs engagement-panel. If you get an empty comment list, your token is almost certainly the wrong one.
Control-validate every method. Before trusting a reading method, run it against a video you know has hundreds of comments. If it returns few or none, the method is broken, not the video.
What I could not get working logged out
- Live chat in a VOD replay. The replay does not expose the chat as a readable comment stream.
- Kick blocks headless browsers outright.
- Twitch live chat required an account for the endpoints I tried.
Takeaway
The innertube path works and needs no API key, but it is undocumented and can change without notice. Keep a mirror as a fallback, keep the official Data API as the serious option, and always compare your method against a known-good video before you trust a number.
Top comments (0)