I kept paying for the same LLM answer twice. An agent retries, a user double-clicks, two workers ask the identical question a second apart, and each one is a fresh model call on the bill. The other thing that kept biting me: when the provider has a bad ten minutes, requests don't fail, they pile up until timeouts cascade through the whole app.
So I put a small Rust server between my code and the model. It's called tokio-prompt-orchestrator, it's MIT, and you can drop it in front of an existing OpenAI or Anthropic client without changing a line of that client.
Drop it in front of what you already have
Start it, point your client's base URL at it:
export OPENAI_BASE_URL=http://127.0.0.1:8080/v1
Your existing code keeps working, and now the same question asked twice is answered from the dedup window instead of a second model call:
curl -s -i localhost:8080/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Summarize this ticket"}]}'
x-orchestrator-dedup: cached
It also gives you a circuit breaker (503 in OpenAI's error format while the provider is down, instead of a pile-up), timeouts, a spend cap (429 when you hit it) and a dead-letter queue. The Anthropic POST /v1/messages endpoint gets the same treatment, so the official Anthropic SDK works unchanged too.
You can try all of it with no key: orchestrator --provider echo answers with your own prompt.
The interesting part: "same question, different words"
Exact dedup only catches identical prompts. The obvious next step is semantic dedup: embed the prompt, and if it's close enough to a recent one, reuse that answer.
Embeddings alone are not safe for this. Measured with BGE-small:
| Pair | Similarity | Answer reused |
|---|---|---|
| "What is the capital of France?" / "Which city is the capital of France?" | 0.959 | yes |
| "How many ounces are in a pound?" / "How many oz in one lb?" | 0.931 | yes |
| "Convert 10 miles to kilometers" / "Convert 10 kilometers to miles" | 0.992 | no |
| "Is 17 a prime number?" / "Is 21 a prime number?" | 0.845 | no |
"Convert 10 miles to kilometers" and "Convert 10 kilometers to miles" score 0.99, higher than any real paraphrase I tried. A pure-similarity cache would happily hand one user the answer to the opposite question.
So a semantic match also has to keep the same numbers, and its shared words in the same order. That costs a few extra model calls (a reworded "how do I reverse a list in Python" that changes word order gets a fresh call), but it errs toward a second call, never toward someone else's answer. If you're building an LLM cache yourself, I'd check your hits on real traffic before lowering any threshold.
orchestrator --semantic-dedup runs the embedder locally (fastembed, BGE-small, a 128 MB one-time download, no API key).
Also: answer from your own docs
orchestrator --docs ./my-docs indexes a folder of Markdown and text with tantivy (BM25) and puts the best passages in front of each question. If search fails or takes over 2 seconds, the prompt goes through without context: retrieval never drops a request.
Try it
- Linux: a one-line install, no dependencies
- Windows: a single
.exe - macOS or from source:
cargo install --locked tokio-prompt-orchestrator --features web-api,tantivy - Or as a Rust library
Repo: https://github.com/Mattbusel/tokio-prompt-orchestrator
Site (replays a real run step by step): https://tokio-prompt-orchestrator.vercel.app/
I'd especially like to hear from anyone running semantic caching in production: what threshold do you use, and what false hits have you seen?
Top comments (0)