1. The Siren Song of Universal Protocol Normalization
Every backend engineer who builds an LLM proxy experiences the same architectural epiphany:
"Why maintain three separate routing engines when I can parse incoming payloads into a single universal struct, mutate the model name, and serialize them into whatever vendor schema the target API expects?"
[Claude Code] ───(Anthropic JSON)───> [Unified IR] ───(OpenAI JSON)───> [DeepSeek/OpenAI]
This pattern—formalized in popular libraries like LiteLLM and adopted by countless multi-agent proxies—looks elegant on paper. It promises an $O(N+M)$ adapter architecture where any client can converse with any upstream model.
However, when you run heavy, autonomous, long-context engineering workloads through an IR gateway for 48 hours straight, this abstraction springs lethal leaks.
In VMR (bigfatsea/vmr), our architecture strictly forbids cross-protocol schema translation. We maintain physical protocol isolation across three entrypoints:
-
openai-completions(/v1/chat/completions) -
anthropic-messages(/v1/messages) -
openai-responses(/v1/responses)
Here is the engineering reality that forced our hand.
2. Where Schema Translation Fails in Practice
A. The Structural Collapse of Tool Call Schemas
Coding agents do not simply send strings back and forth. They emit deeply nested JSON schemas describing terminal commands, AST search filters, and file mutation diffs.
Consider the protocol divergence:
-
Anthropic Messages: Tool calls live inside content blocks (
type: "tool_use"), and input arguments are strongly typed JSON objects (input: { "path": "...", "content": "..." }). -
OpenAI Chat Completions: Tool calls live in a dedicated
tool_callsarray, and function arguments are serialized as an escaped string literal (function.arguments: "{\"path\":\"...\"}"). - OpenAI Responses: The structure shifts again, using different top-level container keys and event names.
When an intermediary gateway ingests an Anthropic request and translates it to an OpenAI-compatible payload, it must execute:
$$\text{JSON Object} \longrightarrow \text{Internal Struct} \longrightarrow \text{Serialized JSON String} \longrightarrow \text{Outer Payload}$$
A single missing escape, an unexpected null coercion, or a subtle float-to-int rounding error silently mutates the tool schema. The upstream model receives a corrupted schema definition and responds with malformed arguments. The agent runtime fails to parse the response and enters a catastrophic retry loop, burning tens of thousands of tokens on a hallucinated tool failure.
B. The Erasure of Thinking Signatures
With the rise of reasoning models (Claude 3.7+ Thinking, DeepSeek-R1, OpenAI o-series), models do not merely output tokens—they emit cryptographic verification blocks and structured thinking chains.
Anthropic's protocol utilizes cryptographic thinking signatures within the message payload to ensure continuity across multi-turn interactions. If an IR gateway naively strips these blocks, flattens them into plaintext Markdown, or fails to mirror vendor-specific headers during forward passes, the upstream provider detects a session integrity violation and returns a fatal HTTP 400.
An intermediary proxy has no business guessing how an upstream model's internal cognitive traces should be parsed or reshuffled.
C. Streaming SSE Desynchronization
Streaming Server-Sent Events (SSE) cannot be treated as stateless byte chunks. Different vendors emit streaming deltas at fundamentally different boundaries:
- Anthropic emits strict lifecycle events:
content_block_start,content_block_delta, andcontent_block_stop. - OpenAI emits incremental
deltachunks with immediate token arrays.
Attempting to run real-time bidirectional stream transcoding inside proxy memory requires maintaining an elaborate state machine for every concurrent connection. When upstream networks experience transient jitter, synthetic event frames get dropped or delivered out of sequence. To the developer sitting at the terminal, the agent simply hangs indefinitely with a blinking cursor.
3. The Byte-Faithful Alternative
VMR's design principle is straightforward: What enters the network card is what hits the upstream wire.
Client (Claude Code / OpenClaw)
│
├──> [/v1/messages] ─────────> Anthropic Native Pipe ─────────> Upstream Anthropic Endpoint
│ (Zero Serialization / Exact Bytes)
│
└──> [/v1/chat/completions] ──> OpenAI Native Pipe ────────────> Upstream OpenAI Endpoint
(Zero Serialization / Exact Bytes)
Zero-Copy Forwarding
When VMR routes a request from Claude Code to an Anthropic-compatible provider, it does not unmarshal the body into an internal representation. The raw HTTP byte slice is piped directly into the upstream transport.
The single modification VMR performs on the request path is a surgical, non-destructive replacement of the top-level model key:
- We do not use regular expressions (which risk corrupting user prompt text containing
"model": "..."). - We do not perform a full JSON decode/encode cycle (which destroys original key ordering and formatting).
- We use a specialized, allocation-free byte scanner that locates the exact offset of the top-level
"model"token and splices in the target upstream model identifier.
The Immediate Benefits:
- Immunity to Protocol Drift: When an LLM vendor releases a new capability—such as prompt caching flags, PDF document blocks, or custom audio modalities—VMR supports it automatically on day zero. We have no adapter code to update.
- Microsecond Latency Floor: By eliminating full JSON unmarshaling and garbage collector churn, VMR processes routing decisions with sub-millisecond overhead ($p95 < 10\text{ms}$ total proxy latency).
-
Forensic Integrity (
audit.jsonl): Because incoming and outgoing byte streams are preserved intact, our offline forensic recorder captures the literal wire bytes. When an agent malfunctions,vmr analyze -comparecan reconstruct the exact byte-level divergence between runs.
4. Architectural Boundaries: What VMR Does vs. Doesn't Do
To preserve this contract, VMR enforces clear boundaries (formalized in our architecture guard tests):
| Concern | VMR Architecture Choice | Rationale |
|---|---|---|
| Cross-Protocol Translation | Strictly Prohibited | Never convert Anthropic $\leftrightarrow$ OpenAI. Route Anthropic clients to Anthropic upstreams; OpenAI clients to OpenAI upstreams. |
| Model Virtualization | Surgical In-Place Rewrite | Single virtual model name (smart-coder) maps to ordered physical endpoints without altering nested payload structures. |
| Failover Mechanics | Passive Health Machine | Classify upstream HTTP error classes (429, 503, invalid keys) without using synthetic probe queries that burn user tokens. |
| Prompt Cache Protection | Session-Sticky Affinity | Anchor client sessions to specific upstream physical endpoints to maintain $50\%\sim70\%$ prompt cache discount rates. |
| Security Inspection | Agent Guard (mode: block) |
Linear-time Aho-Corasick credential scanning and Unicode zero-width rune scrubbing at the boundary, without mutating clean payloads. |
5. Conclusion: Simplicity as a Production Invariant
In modern distributed systems, protocol conversion layers are often technical debt disguised as convenience.
For quick hobbyist experiments or basic single-turn chatbots, IR-based gateways remain useful shortcuts. But when building the operational substrate for autonomous, multi-turn Coding Agents executing thousands of shell commands and filesystem edits, precision trumps universality.
By treating upstream protocol boundaries with respect and keeping the wire payload byte-faithful, you eliminate an entire class of phantom failures—leaving your agent free to do what it was actually built to do: write software.
Top comments (3)
The top-level model splice is the boundary I would exercise most heavily. A fixture set with nested model keys, escaped key names, duplicate top-level keys and malformed JSON would make the scanner’s acceptance policy explicit; rejecting ambiguous inputs is easier to reason about than allowing the proxy and upstream parser to select different values.
For streaming, byte preservation also needs a completion contract around it. Once any response bytes reach the client, failover to another endpoint cannot safely look like continuation of the same turn. Recording whether headers, events or partial tool arguments were forwarded would make that decision auditable without introducing cross-protocol translation.
Pass-through keeps the tool schema intact, which matters because a translated wrapper often drops the exact enum or required field the model was told to fill. The failure mode is not the bytes on the wire. It is the agent guessing a parameter the schema no longer requires. How do you catch a call that looks valid after translation?****
Pass-through keeps the tool schema intact, which matters because a translated wrapper often drops the exact enum or required field the model was told to fill. The failure mode is not the bytes on the wire. It is the agent guessing a parameter the schema no longer requires. How do you catch a call that looks valid after translation?
iin1006h23