DEV Community

Cover image for I Spent a Weekend Binge-Watching LLM Content So You Don't Have To (Okay, You Still Should, But Here's the Map)
Prayush Adhikari
Prayush Adhikari

Posted on

I Spent a Weekend Binge-Watching LLM Content So You Don't Have To (Okay, You Still Should, But Here's the Map)

Real talk: everyone and their uncle is "using AI" right now, but ask most people what a token actually is and you get a blank stare. I was in that camp too, not long ago. So this month I did something slightly unhinged — I sat down and went through three completely different explanations of how LLMs actually work, back to back, and it clicked in a way that no single YouTube video was ever going to give me.

Here's the thing: each of these sources attacks the problem from a totally different altitude. One gets its hands dirty with code and whiteboard sketches. One skips theory entirely and just shows you how to build something. And one is basically a full grad-school course crammed into a lecture series. Stack them together and you get something close to an actual curriculum, minus the tuition.

If you're on the same journey — curious about AI beyond "type prompt, get answer" — I'm breaking down what I learned from each one, and honestly, I'd watch them in the same order I did.

Big credit up front to the people who actually made this stuff:

Let's get into it.

Part 1: What's Actually Happening Inside the Box

If you've ever nodded along in a conversation about AI while quietly having zero clue what "embedding" means, start with Piyush Garg's video. It's not a theory lecture — it's a "let me rip the hood off this thing" video, and it works because it stays concrete the entire time.

The core message is refreshingly deflating in a good way: ChatGPT, Gemini, Claude — none of it is magic, none of it is sentient. It's math, code, and a lot of data. Even the name gives it away. GPT is basically a spec sheet:

  • Generative — it makes new stuff on the fly instead of retrieving and ranking existing pages like a search engine.
  • Pre-trained — it already chewed through a mountain of internet text, books, and conversations before you typed anything.
  • Transformer — the architecture underneath all of it, from Google's 2017 paper "Attention Is All You Need."

From there the video walks you through the actual pipeline your sentence goes through:

Tokenization. Before a computer can do anything with your text, it needs numbers. A tokenizer chops your sentence into chunks based on a vocabulary, and the size of that vocabulary decides whether you get clean whole-word tokens or things get sliced into subwords and stray characters. Garg demos this live with Hugging Face's transformers library and Google's Gemma tokenizer, and honestly that's the moment it stops being abstract and starts being "oh, that's literally what's happening."

Vector embeddings. Once you've got tokens, each one gets mapped to a dense vector — we're talking hundreds or thousands of dimensions — that's supposed to encode meaning. The analogy that actually landed for me: think direction and distance. The path from Cat to Milk looks a lot like the path from Dog to Pedigree. Dog to Animal mirrors Man to Human. The model isn't memorizing a dictionary — it's storing relationships between concepts, and you can literally go poke at this yourself through the OpenAI embeddings API.

Positional encoding. Here's a detail a lot of people skip over: Transformers don't read left to right the way you do. Without help, "the dog chased the cat" and "the cat chased the dog" would look identical to the attention mechanism, which — yeah, that's a problem. Positional encoding bolts sine and cosine functions onto the embeddings so word order (and the meaning that comes with it) actually survives.

Self-attention and multi-head attention. This is the star of the show, and it's the actual innovation from the 2017 paper everyone name-drops. Self-attention lets every token look at every other token and update itself based on context. That's how the model knows "bank" in "river bank" has nothing to do with "bank" in "ICICI bank" — something the older RNN and CNN architectures genuinely struggled with. Multi-head attention just runs a bunch of these attention processes in parallel, so the model can track multiple relationships at once — a dog, on a train, doing something, being a specific color — all simultaneously instead of one at a time.

Feed-forward layers, then output. These attention blocks stack on top of each other, refining the representation layer by layer, until a final linear layer spits out a probability distribution over the entire vocabulary for "what token comes next." Softmax turns those raw scores into actual probabilities, and temperature decides how safe or how spicy the model gets with its pick — low temperature plays it safe, high temperature gambles on something less likely (and usually more interesting, sometimes more unhinged).

The video also draws a clean line between training (the model sees known input-output pairs, calculates loss, updates weights via backpropagation) and inference (live deployment — the model just predicts one token at a time, autoregressively, until it hits an end-of-sequence tag, with zero learning happening). Sounds obvious once you say it out loud, but it clears up a lot of confusion about what's actually happening when you're chatting with one of these things.

What makes this video worth your time isn't that any of it is secret — it's that Garg keeps everything tied to code you can actually run: AutoTokenizer, the OpenAI embeddings endpoint, a from-scratch PyTorch generation loop with AutoModelForCausalLM. You walk away with a mental model you can go verify yourself instead of just trusting some guy on YouTube (me included).

Part 2: Cool, Now How Do I Actually Build Something?

Understanding attention mechanisms doesn't teach you how to ship a product. I learned that gap the hard way trying to wire an LLM into a side project with zero idea what I was doing. This is exactly where Dev Weekends' "Level 0 and Level 1" session comes in, and it's the most immediately useful of the three if your goal is writing software instead of passing an exam.

It starts with the unglamorous stuff that actually matters: you pay for tokens, full stop, and your bill is input tokens (your prompt) plus output tokens (the response). If you've ever gotten a surprise API bill, you already know why this matters more than the fancy theory.

From there it breaks down the three roles you'll use in basically every single LLM API call:

  • System message — sets the model's personality and ground rules, invisible to whoever's using your app.
  • User message — whatever the human actually typed.
  • Assistant message — the model's response, generated using both of the above as context.

A few engineering patterns get covered here that I think people skip more than they should:

  • Streaming (stream=True) — instead of waiting for the full response, you get it chunk by chunk, which is the entire reason ChatGPT feels like it's "typing" instead of freezing and dumping a wall of text on you.
  • Structured output / JSON schemas — if you're plugging an LLM into a real system, you cannot have it randomly deciding to wrap its JSON in a friendly little paragraph. Define a schema up front and you kill an entire category of parsing bugs before they happen.
  • Tool calling / function calling — this is the mechanism that lets a model say "I don't know today's weather, but here's a function signature for a weather API, go call it and hand me the result." The model itself never executes anything — it just recognizes when it needs help and passes the ball back to your app.

On the prompting side, it runs through the standard toolbox: zero-shot (just ask), few-shot (show it examples first), chain-of-thought (make it reason step-by-step before answering), and ReAct (interleaving reasoning with actual tool calls to chew through multi-step problems). I'd read about all of these separately before, but seeing them laid out as a progression — from "just ask" to "reason and act" — made a lot more sense than any individual blog post had.

The part I found genuinely most useful, though, was the fine-tuning vs. RAG breakdown. Fine-tuning means retraining chunks of the network on your own data — expensive, heavy on compute, and static, since any new information means retraining the whole thing over again. RAG (Retrieval-Augmented Generation) sidesteps all of that by keeping the model frozen and just hooking it up to an external, updatable vector database instead.

The RAG pipeline breaks down into five steps:

  1. Parse your source documents — PDFs, databases, scraped pages via BeautifulSoup, whatever.
  2. Chunk the text into fixed-size pieces (think 512 tokens) with some overlap (150 tokens) so you're not losing context at the seams.
  3. Embed and store those chunks in a vector database (ChromaDB, PgVector) along with metadata like source and timestamp.
  4. Search — turn the user's query into an embedding and run cosine similarity to pull the top-k most relevant chunks.
  5. Inject those chunks into the prompt so the model actually grounds its answer in your data instead of just guessing confidently.

And it doesn't stop at "ship it and pray" — it flags RAGAS as a framework for actually evaluating whether your retrieved context is faithful, your answers are relevant, and you're not quietly hallucinating despite technically having a "grounded" system. That last part is easy to forget. RAG lowers your hallucination risk. It does not erase it. Learn that the easy way, not the hard way.

Part 3: The Academic Backbone — Stanford CME295

If the first two are "how it works" and "how you build with it," Stanford's CME295 (taught by Afshine and Shervine Amidi) is the "why it works, rigorously, and where this is all headed" course. It's a full nine-lecture sequence, and it's the one I'd point to if your goal is going from "informed enthusiast" to "person who can actually reason about model design decisions in a room full of people who'd know."

The syllabus tracks a clean arc:

  1. Transformer — core attention mechanics and encoder-decoder structure, covered with academic precision instead of intuition-first analogies.
  2. Transformer-Based Models & Tricks — how BERT, GPT, and T5 diverge architecturally, different positional embedding schemes, and the tricks that make large-scale training actually converge instead of blowing up.
  3. Transformers & LLMs — the jump from small task-specific models to today's massive general-purpose ones, scaling laws, and emergent capabilities that only show up once you're big enough.
  4. LLM Training — pre-training objectives, dataset prep at scale, and the distributed-cluster/memory-efficiency tricks required to train something this large in the first place.
  5. LLM Tuning — Supervised Fine-Tuning, Parameter-Efficient Fine-Tuning (LoRA and friends), and alignment via RLHF and Direct Preference Optimization.
  6. LLM Reasoning — how post-training unlocks reasoning, and the increasingly hot idea of scaling test-time compute instead of just throwing more compute at training.
  7. Agentic LLMs — designing autonomous agents: task decomposition, memory architectures, tool use, multi-agent coordination.
  8. LLM Evaluation — systematic benchmarking, the "LLM-as-a-judge" approach, and safety/robustness red-teaming.
  9. Recap & Current Trends — tying it all together and pointing at where frontier research is actually headed.

What this course adds that the other two structurally can't is the why. Garg's video tells you self-attention solves the context problem. Stanford's course is where you'd actually get the math behind scaled dot-product attention, the reasoning behind specific positional embedding choices, and the real tradeoffs between RLHF and DPO for alignment. It's also the only one of the three treating agentic systems and formal evaluation as first-class topics instead of an afterthought bolted on at the end — which, not coincidentally, is exactly where the field's energy is right now.

Putting It All Together

Lined up side by side, these three cover roughly the same ground a CS curriculum would, just compressed into a weekend of watching instead of a semester of tuition:

Piyush Garg Dev Weekends Stanford CME295
Best for Building intuition for the internals Shipping real LLM-powered software Deep theoretical and research literacy
Format Single video, whiteboard + code Hands-on workshop, live API/RAG demos 9-lecture university course
Core focus Tokens → embeddings → attention → sampling API roles, streaming, JSON schemas, tool calling, RAG vs. fine-tuning Math foundations, scaling, alignment, reasoning, agents, evaluation
Tools shown Hugging Face transformers, PyTorch, OpenAI API OpenAI API, ChromaDB, LangChain, BeautifulSoup, RAGAS Academic frameworks, frontier benchmarks

Honest take: none of these alone gives you the full picture, and that's fine, because they're not trying to. Garg hands you the mechanical model of how a token turns into a prediction. Dev Weekends shows you what to actually do with that model once it's sitting behind an API and your credit card is on the line. Stanford gives you the rigor to understand why the field made the choices it did, and where it's probably going next — reasoning models, agentic systems, evaluation frameworks that can actually keep pace.

If you're early in this journey, that's the order I'd go: intuition, then application, then theory. Same path I took, minus the part where I procrastinated for two weeks before starting.

Now go tokenize something.

Let's Connect

I'm Prayush Adhikari, currently grinding through the 8th semester of computer engineering and running an IT club on the side while trying to actually understand the stuff I keep building with. If you've watched any of these three (or have a better resource I'm missing), drop it in the comments — I'm always collecting more.

Which one are you starting with: the internals video, the build-something workshop, or going straight for the Stanford deep end?

Top comments (0)