DEV Community

PandeyC
PandeyC

Posted on

Walk Me Through How You'd Design a Production Prompt — 9 Interview Questions, Answered

There's one question that opens almost every AI engineering interview's prompt-design section:

"Walk me through how you'd design a production prompt for a structured extraction task."

It sounds open-ended. It isn't. Interviewers are listening for five specific beats, and the follow-up questions below are how they check for the ones you skip. This is the companion piece to the hands-on Colab notebook where we measured an 80%-to-0% format-drift swing from adding few-shot examples — if you want the receipts behind question 3 below, that's where they are.

1. What's the difference between a system prompt and a user message?

The system prompt is developer-controlled — role, rules, output format — and runs invisibly on every call. The user message is the runtime task or input, changing with each call. Miss this distinction and you'll bake per-request data into a prompt that should be static, or vice versa.

2. What does RCTF stand for, and what happens if you skip one letter?

Role, Context, Task, Format. Skip Format and the model chooses its own output shape — usually inconsistently. Skip Context and it fabricates facts it doesn't have (product names, plan tiers). Skip Role and tone drifts call to call. Skip Task and you get a summary when you wanted a classification. Interviewers who ask "what happens if you skip X" are testing whether you understand why each field exists, not just that it exists.

3. What temperature would you use for a support ticket classifier, and why?

Zero, or close to it. Classification feeds a parser, and parsers need identical output for identical input. Temperature above zero introduces sampling randomness — and that randomness doesn't just reword the answer, it can surface as a genuinely different output shape. Say the number if you have it: a naive zero-shot classifier can drift to a different JSON schema over 80% of the time at moderate temperature. Anchoring the same call with a few worked examples took that to zero.

4. When do you reach for few-shot instead of zero-shot?

Always start zero-shot — it's cheaper and easier to maintain. Move to few-shot only when zero-shot's content is right but the format is wrong or inconsistent. Few-shot examples teach output shape, not new facts. If the content itself is wrong, more examples won't fix it — that's a Context problem, not a technique problem.

5. What's the difference between few-shot prompting and fine-tuning?

Few-shot: examples live in the prompt, inference-time only, weights unchanged, cheap to update. Fine-tuning: the model is retrained on your dataset, weights change permanently, requires thousands of examples and real infrastructure. Try few-shot first. Fine-tuning is for stable, high-volume, narrowly-defined tasks where prompting has genuinely been exhausted — not a first instinct.

6. When does chain-of-thought help, and what does it cost you?

CoT helps on multi-condition or multi-step tasks where jumping straight to an answer causes errors — escalation decisions, trade-off analysis. The instruction "think step by step" forces intermediate reasoning tokens that improve the final answer. The cost: more tokens, more latency. Don't reach for it on simple single-condition tasks — that's needless spend.

7. What is prompt injection, and how do you actually prevent it?

User-supplied text overwrites developer instructions because, to the model, both are just tokens — there's no privileged execution layer separating "your rules" from "their input." The fix that holds up: XML-delimit user input (<user_input>...</user_input>) and instruct the model to treat that block as data only. The fix that doesn't hold up: asking nicely in the system prompt and hoping. A system prompt is a request, not a boundary.

8. Why does a model confidently answer with stale stock prices or outdated facts?

Knowledge cutoff. The model has no internet access and no live retrieval — it generates the statistically most plausible continuation from training data, which may be months or years stale, with the same confident tone as a correct answer. There's no "I don't know current data" reflex unless you build one. The fix is injecting current data into context directly (RAG), never trusting model memory for anything time-sensitive.

9. What does max_tokens actually control, and what breaks if you set it wrong?

The length of the response, not the input — that's the context window's job. Set too low, structured output truncates mid-JSON and your parser throws. Set too high, you're paying for and waiting on output nobody asked for. Rule of thumb: roughly 2x your expected output length, tighter (80-150) for constrained JSON classifiers.


The beat interviewers probe hardest: failure modes. A candidate who nails RCTF and parameters but never mentions injection, knowledge cutoff, or format validation unprompted gets asked "what could go wrong?" — and that follow-up is where senior and junior answers actually separate.

Want to see the 80%-to-0% drift number measured live, not just quoted? The hands-on notebook is here — zero API key required, about 12 minutes.

Top comments (0)