DEV Community

Maaz Kazi
Maaz Kazi

Posted on

When the prompt is the product, it does not belong in the code

CASEY started from a complaint about resumes. A candidate gets filtered out on the strength of a document that is a static snapshot of them, written once, describing a person who has since changed. Judging someone by it is like running a predictive model on a frozen spreadsheet.

So the product thesis was: replace the document with a conversation, and replace the fixed profile with a plan that keeps changing as the person does. You talk about your background, you pick a direction, and you follow a roadmap built around who you actually are.

The hypothesis worth testing first

The first build existed to answer one question, and it was not "can we ship this". It was: can a spoken conversation carry a psychometric framework?

Career assessment instruments are questionnaires. They work, and almost nobody finishes them, because answering forty Likert items about yourself is miserable and the output feels like a horoscope with citations. The bet was that the same constructs — what someone is motivated by, what they are reaching for, where they might actually fit — could be drawn out in a conversation that does not feel like an assessment at all.

So the MVP was not a product. It was an instrument for finding out whether that held. Everything else was downstream of the answer.

The structure: phases, not a chatbot

A freeform assistant would have been easier and wrong. The journey is explicitly staged.

  1. Discovery — situation, interests, what is actually blocking you.
  2. Orientation — narrow to a direction.
  3. Roadmap — a three-step plan: explore, validate, confirm.
  4. Accountability — recurring check-ins against that plan.

Stages matter because the thing being built is a judgement about a person, accumulated over time. A model that has to infer where it is in that process from conversation history will drift, and it will drift differently for every user.

The backend owns progression. The model supports extraction, not control.

This is the rule I would keep in any product like this.

The model runs the conversation. It does not decide where the user is in the journey, and it does not decide when a stage is complete. Completion is a deterministic backend function with an explicit condition — all three roadmap steps done, for instance. The model may offer to move on. Only the function decides.

It is tempting to let the model handle this, because it can. It will read the transcript and make a sensible call most of the time. But "most of the time" is doing a lot of work in a system where the state it is judging determines what the user is shown next, what gets written to their record, and whether they are charged. A stage boundary is business logic wearing a conversational costume.

The corollary is that the model is excellent at the other half: pulling structured signal out of unstructured talk. Extraction, summarisation, insight — yes. Flow control — no.

Two models, two jobs

The coaching conversation and the memory layer are deliberately not the same model.

The conversation runs on a voice-native stack, because latency and prosody are the product when someone is talking about their career and feeling exposed. Post-session summarisation and key-insight extraction run separately, on a cheaper text model, writing into a summaries table that becomes cross-session memory.

That separation is a cost decision and an architectural one. The expensive real-time path does only what has to happen in real time. Everything reflective happens after the call ends, where latency does not matter and a smaller model is sufficient.

The title of this post

Session prompts and phase configuration live in database tables, not in application code.

This was the single best structural decision in the build, and it is the one that sounds least interesting. Coaching behaviour — how CASEY opens a Discovery session, what it is trying to establish before moving on, the tone it holds when someone says something difficult — is tunable without a deploy. An admin surface exposes it directly.

The reason is not developer convenience. It is that the prompt is the product. In a conventional app, copy is presentation and logic is behaviour. In this one, the prompt is the behaviour: it is the thing you iterate on daily, the thing a non-engineer has the best judgement about, and the thing you most want to change in response to a real session that went badly yesterday.

Putting that behind a pull request and a deploy cycle means the person with the best instinct about it needs an engineer to act, and iteration slows to the speed of releases. Putting it in a table means you fix a coaching failure the afternoon you find it.

The test for whether this applies to your product: if the thing you edit most often is a string in a prompt file, it is not code, it is content, and you are versioning it in the wrong place.


Written by Maaz Kazi, who built CASEY. More at https://maazkazi.com

Top comments (0)