I once collected about 22,000 comments from roughly 140 Korean YouTube videos about AI coding tools and classified them. (Quotes below are translated from Korean.)
I wanted to see what people were asking. What came out was something else.
What the videos teach
Put the titles and tags of those 140 videos in one pile and they say:
How to install. How to get started. How to build an app. Which tool is best.
All of it is "starting." Follow along, a result appears on the screen, the video ends.
What the comments say
The comments sweeping up the likes were telling a different story.
"Verifying AI mistakes takes so much time. Checking every answer for nonsense got so tiring I just do the work myself now." (👍598)
"Coding with AI makes me anxious. If one bug ships, I'm the one responsible. Checking and debugging everything one by one ends up being more work." (👍265)
"I pay every month and it lies about work matters like it's nothing." (👍72)
"Tokens burn too fast… added $50 and it was gone in half a day." (👍30)
It compresses into three complaints: expensive, can't trust it, can't fix it.
The videos teach the start. The people are dying right after the start.
The scariest comment
"Asked it for shampoo recommendations and it recommended one that doesn't exist. Slipped it in between real products — with the weight, the benefits, even a price." (👍49)
That comment is the essence of the problem. When AI is wrong, it doesn't look wrong. The fake sits among the real ones, wearing plausible numbers.
This is why "just write better prompts" is half an answer. Better prompts lower the odds of being wrong. They don't create a way to know when it's wrong. Drop the error rate from 10% to 3% and you still don't know where the 3% is hiding. If that 3% detonates inside payment logic, money leaves the building.
One more finding — where the real questions live
While collecting, I noticed the nature of comments changes with channel size.
multi-million-sub videos real questions/needs = 12% of comments — the rest is reactions and anxiety
10k–300k sub channels real needs = 26% — "I followed along and got stuck RIGHT HERE"
It's the difference between spectators and people actually doing the thing. Comments under the big videos ask "what happens to us in the AI era." Comments under mid-size channels ask "how do I fix this error."
The real questions were in the mid-size channels. And collecting reply threads, not just top-level comments, surfaced 470 more needs that the top level never showed.
So I picked a direction
Content that teaches starting is already everywhere. I decided to work on the stretch right after the start — the can't-trust-it, can't-fix-it stretch.
The earlier posts in this series are the first result of that decision: giving my LLM an exam, planting traps in it, grading by severity instead of pass/fail, why passing still wasn't enough to ship. Every one of them is about turning "I can't trust it" into "I verified it."
People don't die at the start. They die right after it.
But almost everyone is selling the start.
P.S. The classifier that produced the percentages above later sat an exam of its own — the same kind I give my LLMs — and failed three fatal-grade questions. That public correction is its own story, coming later in this series.
Next up: how to steal this exam and port it to your own pipeline, step by step.
P.P.S. If you'd rather get new experiments by email: ramses203.substack.com/subscribe
Top comments (9)
The strongest signal in that scrape is how much of the pain shows up after the demo has already "worked." Install tutorials make sense as acquisition content, but the comments are asking for recovery paths, bad-output examples, and a way to know when the model is bluffing.
I would probably split the exam set the same way. One bucket for capability, one for falsifiability. If the test only proves the model can do the happy path, it misses the part your commenters are actually afraid of.
The split you're naming is what my exam does implicitly — and naming it makes it countable. Of my 29 questions, the normal orders are the capability bucket; everything else — plausible non-orders, undecidable orders, mid-message reversals, after-learning cases — is the falsifiability bucket. That's about 90% of the exam, which is why an earlier post in this series is literally titled "A Good LLM Exam Is 90% Traps."
What your label adds: once the bucket is explicit, you can watch the ratio as the exam grows. New questions drift toward easy capability padding unless something pushes back — in my exam the accident list does that job, and your two-bucket label would make the drift visible as a number.
On your three reader-wants: "a way to know when the model is bluffing" is what this series has been about. The recovery part — "it's wrong, now what" — another commenter here just called the emptiest shelf in the store. Between the two of you, I think the next season of this series just got scoped.
Your split has a sibling on dev.to, and the numbers rhyme. I measured 1,350 top articles across 20 categories last week: the platform WRITES about agents and getting-started (agents alone: 14 % of articles), but the conversation lives elsewhere - failure/postmortem pieces draw the highest comment rate of any category (58 comments per 100 reactions), and when you count what people name as their pain, "can I trust my green check" beats "context/memory" roughly three to one. Two platforms, two languages, two formats - same two-conversations structure: the content teaches starting, the audience is stuck right after it. That's strong evidence your finding isn't a YouTube artifact or a Korea artifact; it's the shape of the market.
One hard-won addition to Vinh's audit suggestion (which is right): when you hand-label the stratified sample, resist the pull to only re-examine the comments your classifier got wrong-looking results on. We made that mistake in our own error analysis once - auditing only the lost cases - and the correction rate looked wildly different from the true one, because the found cases hid errors of the opposite sign. The per-bucket error rate only means something if the sample is drawn blind from the full bucket, wins included.
And a question for the next part: within the mid-size channels' 26 % real needs - how many are verification needs ("is this output right?") versus recovery needs ("it's wrong, now what?")? Your series covers the first beautifully; my scrape of my own niche says the second is almost unserved, and "recovery content" might be the emptiest shelf in the whole store.
Thank you for running the numbers on your side — 1,350 articles is a bigger sample than mine, and the same split showing up on a different platform, in a different language, is the strongest confirmation this finding has gotten. Your blind-sampling warning also lands at the right moment: Vinh, one comment over, just asked me to audit the 12%/26% with hand-labeled samples, and left to myself I would have re-checked only the wrong-looking cases. I'll draw blind from the full buckets, wins included.
On your question — I don't have the verification-vs-recovery split measured, and I'd rather not guess it in a comment. The audit sample can carry a second label ("is it right?" vs "it's wrong, now what?"), so the answer will come out as a real number instead.
One thing I can say now: my grade table already runs on recoverability — FATAL literally means "nobody can undo it" — but you're right that everything I've written teaches checking, not recovering. The empty shelf is real.
Drawing blind from the full buckets, wins included, is the whole audit - sampling only the wrong-looking cases just measures your own suspicion. And the second label means the audit answers the verification-vs-recovery question for free, as a real number. Looking forward to it.
When you publish the split, I'll run the same two labels over the audit sample on my side, so the platforms stay comparable - same definitions, different corpus. That's how your "the numbers rhyme" observation becomes a series instead of a coincidence.
The 12% versus 26% split is the number I would want audited before building a content direction on it, given your own P.S. that the classifier later failed three fatal-grade questions. That does not need a re-scrape: hand-label a stratified random sample from each bucket, a hundred or so per side, and report the classifier's error rate per bucket instead of overall. A comparison between two groups only survives if the error is roughly symmetric across them, and short reaction-type comments are exactly the sort of thing a classifier misreads at a different rate in the multi-million-sub pile than in the mid-size one, which is the case where the gap moves without a single underlying comment changing.
Accepted — and this one stings more than the last one, because the disclosure was already sitting in my own P.S. and I still hadn't done the arithmetic on what it means for the comparison. The plan, so you can check it: stratified random sample, about 100 comments per bucket, drawn blind from the full buckets including the ones the classifier got right, hand-labeled, then per-bucket error rates reported — symmetric or not. If the 12/26 gap doesn't survive, that post gets the same treatment the model comparison got. Labeling is human work, so this will arrive as its own post rather than a quick edit.
Great article, John! I really enjoyed your deep dive into this data—it’s such a clever way to figure out what engineers are actually struggling with versus what the hype is focused on.
Takeaway: Your observation that "spectators" flock to the massive channels while the actual builders solving real errors live in the mid-size channels is spot on. It completely validates why the industry needs to shift focus from "getting started" tutorials to solving the "can't trust it, can't fix it" reality.
Question: Since you're building out these rigorous LLM exams to catch those dangerous, plausible-sounding hallucinations, have you experimented with automating these tests directly inside a CI/CD pipeline (like GitHub Actions) so they act as a deterministic safety gate before production?
Keep up the awesome work on this series!
Feels relatable: a video about AI tools, and the comments drift off into memes and life hacks. The internet loves multi-track conversations—no hard feelings, just data.