DEV Community

Micheal Heypico
Micheal Heypico

Posted on

Prompt versioning: treat prompts like code (with a 3-model test harness)

Prompt versioning: treat prompts like code (with a 3-model test harness)

Prompts are code. They have regressions, environments, and review requirements - they just don't have a compiler to warn you.

What we do at HeyPico

Every prompt lives in git with a version tag. Every change runs against a fixed eval set before shipping, tested on three models minimum: one frontier (Claude or GPT class), one cheap (DeepSeek or GLM class), one wild card. Why three? Because providers fail differently, and a prompt tuned to one model's quirks silently breaks when you switch.

The failures are humbling:

  • A prompt that worked on Claude lost tone constraints on a smaller model.
  • A JSON-output prompt returned wrapped markdown on one provider, clean JSON on another.
  • An instruction that everyone follows ("answer in <=2 sentences") gets ignored by exactly the cheap models you'd use for bulk work.

The setup, concretely

  1. prompts/ directory, one file per prompt, versioned.
  2. evals/ directory, one JSONL per prompt with input + expected properties.
  3. A runner that hits each model through one API endpoint and diffs behavior.
  4. Merge blocked until eval diff is clean.

The multi-model routing part is one API key with 32 models - the eval runner is ~200 lines of Python. Ping us if you want to replicate it.

Top comments (0)