Not a quiz for fun. A gut check.
These are three questions that show up in real AI engineering interviews — not "define RAG" trivia, but the scenario-style questions interviewers actually use to tell a candidate who's built things apart from one who's only read about them. Read each one, answer it out loud before you read the framework underneath, and be honest about how close you got.
🎯 TL;DR: Three real interview questions, no trivia. Answer each one out loud before reading the framework. If you reach for a definition instead of a process, that's the exact gap other candidates are closing right now — not by reading more, but by building.
🩹 1. "Your RAG chatbot just gave a customer a confident, completely wrong answer. Walk me through how you'd find out why — and stop it from happening again."
This is a debugging question disguised as a RAG question. Interviewers ask it because "explain how RAG works" tells them you read an article, but this tells them whether you've actually operated one.
How to answer it:
- Split the failure in two before you diagnose anything: was it a retrieval problem (the right chunk never made it into context) or a generation problem (the right chunk was there and the model ignored it anyway)? Say this split out loud first — it's the single biggest signal you know what you're doing.
- Describe how you'd check retrieval: log the retrieved chunks alongside the query, and eyeball whether the answer's source text was even in there.
- Describe how you'd check generation: if the source was there but the model still got it wrong, that's a prompt/grounding problem — talk about instructing the model to cite or refuse when the context doesn't support an answer.
- Close with prevention, not just the one-time fix: a small eval set of known-answer questions you can re-run after every prompt or retrieval change, so this doesn't quietly regress next sprint.
⚖️ 2. "We could fix this with a better prompt, or by fine-tuning the model. How do you decide which one?"
This question is a trap for candidates who only know one hammer. It's testing judgment, not a definition.
How to answer it:
- Lead with prompt engineering as the default: it's fast, cheap, needs no training data, and you can iterate on it in minutes.
- Name the specific conditions that push you toward fine-tuning instead: you have a large, stable, labeled dataset; you need consistent output formatting or tone at a scale prompting can't reliably hold; or your prompt is already maxed out on context budget and instructions are still being ignored.
- Mention the cost you're trading away: fine-tuning is slower to iterate and harder to reverse than a prompt change. Say that out loud — it shows you're weighing the decision, not just picking a technique.
🚨 3. "The feature demos perfectly, but real users are breaking it in ways you never saw in testing. What's your process?"
This is the question that separates "I built a demo" from "I've operated something in production."
How to answer it:
- Start with the gap between demo inputs and real inputs: real users send empty strings, five-paragraph pastes, other languages, and adversarial prompts your test set never covered.
- Talk about what you need logged before this happens, not after: full request/response pairs, retrieved context, any tool calls — without that, you're debugging blind.
- Mention a containment step: a feature flag or circuit breaker so you can turn the feature off or roll back the version for affected users while you dig in, instead of leaving it broken in front of customers while you investigate.
🪞 How did you do?
If you answered all three with that level of specificity, unprompted — that's rare, and it means you've actually operated systems like this before.
If you found yourself reaching for a textbook definition instead of a process, that's useful information. It means you understand the concepts but haven't yet built the judgment that only comes from shipping and breaking a few of these yourself. Other candidates in your interview pool are closing that exact gap right now — not by reading more, but by building.
That judgment isn't something an article can hand you — including this one. It comes from building the actual systems these questions are about.
The takeaway: knowing the definition doesn't get the offer. Having broken the thing once does.
Top comments (0)