Prompt versioning: treat prompts like code (with a 3-model test harness)
Prompts are code. They have regressions, environments, and review requirements - they just don't have a compiler to warn you.
What we do at HeyPico
Every prompt lives in git with a version tag. Every change runs against a fixed eval set before shipping, tested on three models minimum: one frontier (Claude or GPT class), one cheap (DeepSeek or GLM class), one wild card. Why three? Because providers fail differently, and a prompt tuned to one model's quirks silently breaks when you switch.
The failures are humbling:
- A prompt that worked on Claude lost tone constraints on a smaller model.
- A JSON-output prompt returned wrapped markdown on one provider, clean JSON on another.
- An instruction that everyone follows ("answer in <=2 sentences") gets ignored by exactly the cheap models you'd use for bulk work.
The setup, concretely
-
prompts/directory, one file per prompt, versioned. -
evals/directory, one JSONL per prompt with input + expected properties. - A runner that hits each model through one API endpoint and diffs behavior.
- Merge blocked until eval diff is clean.
The multi-model routing part is one API key with 32 models - the eval runner is ~200 lines of Python. Ping us if you want to replicate it.
Top comments (0)