If you're prepping for an AI engineering interview, this question is coming: "How does an LLM generate a response?"
Most candidates get the first two beats right — tokens, embeddings — then start hand-waving. The ones who land the offer name all five beats without prompting, including the one most people skip.
There's a companion piece, "Your LLM Prompt Costs More Than You Think", that builds a token/cost calculator in Colab you can run in 10 minutes with zero API key. (Link it here once both are published.) This one is the pure Q&A version — no code, just the questions and the answers that separate a senior response from a junior one.
The core question, and the 5 beats interviewers listen for
"How does an LLM generate a response?"
A strong 90-second answer hits all five:
- Tokens — sub-word units; the model reads these, not words
- Embeddings — each token becomes a vector of numbers capturing meaning
- Attention — scores relevance across every token in context, for each new token generated
- Generation / sampling — next-token prediction, one token at a time, until done
- Hallucination root cause — plausibility ≠ truth; there's no truth-check step in the loop
Miss beat 5 and expect a direct follow-up. It's the most commonly skipped beat, and interviewers know it.
8 questions from this session, answered
1. What is a token, and why does it matter for your job?
A sub-word chunk — the model's basic reading unit. Not a word, not a character. "unbelievable" is 3 tokens. "ChatGPT" is 3 tokens. It matters because you pay per token, not per word, and the model has a hard limit on how many it can process per request.
2. What's the difference between "context" and the "context window"?
Context is everything the model can see in one conversation — your messages, documents, its own replies. The context window is the size limit on that everything: the maximum token count (input plus output) per request. It resets completely on every new call — not persistent memory.
3. What is an embedding, in plain terms?
A list of numbers representing a token's meaning and usage, not its spelling. Similar words land at nearby coordinates; unrelated words land far apart. The same word can carry different embeddings depending on the sentence it's in ("bank" the institution vs. "bank" the river).
4. What does attention actually do?
For every token about to be generated, attention scores every other token in context for relevance. High-relevance tokens get more influence over what comes next. This is also why a cluttered, noisy prompt produces worse output — irrelevant tokens compete for attention against the instruction that actually matters.
5. Why does an LLM hallucinate — what's the real root cause?
The model predicts statistically plausible text. There is no truth-check step anywhere in generation. Confident, authoritative writing was abundant in training data, so the model learned confident tone as a default style — a style that shows up whether or not the underlying claim is true. Plausibility is not truth.
6. What does temperature control, and when do you use low vs. high?
How peaked or flat the next-token probability distribution is. Low (0.0–0.3): the model almost always picks its top-probability token — use for fact extraction, structured output, anything that needs to be the same every call. High (0.7–1.0): more willing to pick lower-probability tokens — use for brainstorming or creative variation. Temperature never changes what the model knows, only how consistently it says it.
7. How does a chatbot "remember" earlier messages if the model resets every call?
It doesn't, on its own. The application layer replays prior turns back into the context window on every new request. The continuity a user experiences is maintained by the app, not the model. Strip out that replay logic and the model has never heard of the user before.
8. Why can't an LLM answer questions about your company's internal documents?
Training data is frozen at a cutoff date, and internal documents were never in it to begin with. Anything private or created after the cutoff has to be explicitly supplied in context — which is the entire premise behind retrieval-augmented generation.
What separates a strong answer from a weak one
Say: "Tokens, not words." "Embeddings capture meaning as vectors." "Attention scores relevance across the full context." "Plausibility, not truth — there's no truth-check step." "Context window is a fixed buffer that resets."
Avoid: "It's just autocomplete" (technically true, reads as surface-level). "It thinks like a human" (signals a real misconception). "It searches the internet" (only true with retrieval tools wired in). "It hallucinates because it lacks data" (names the wrong root cause — it hallucinates even with the data, because nothing checks the output against it).
Want the full 5-beat model answer, word for word?
The course session covers this question with a complete scored model answer, the do's/don'ts list interviewers are actually listening for, and the hands-on Colab notebook that makes tokens and cost tangible before you ever sit down for the interview:
https://confidentprep.com/courses/ai-ml-for-interview/1-how-llms-actually-work/
Say your answer out loud before your next interview. Reading it silently and being able to say it under pressure are two different skills.
Top comments (0)