Originally published on tamiz.pro.
From Prompt to Protocol: Engineering Verifiable AI Agents
The era of "prompting a chatbot" is ending. For enterprise-grade software, raw prompts are fragile, untestable, and opaque. The next frontier is Protocol Engineering: defining explicit contracts between the LLM and the surrounding system to ensure verifiable, reproducible behavior. This article explores how to transition from ad-hoc string manipulation to rigorous agent architectures using the Bernstein Agent Protocol and TruLens observability stack.
In this deep dive, we will dissect the limitations of current LLM usage patterns, define the theoretical underpinnings of the Bernstein protocol, and build a concrete, observable agent pipeline that you can deploy today.
Table of Contents
- 1. The Limitations of "Vibe Coding" and Prompting
- 2. The Bernstein Agent Protocol: Defining Contracts
- 3. TruLens: Observability for LLM Systems
- 4. Architecting the Verifiable Agent: A Step-by-Step Implementation
- 5. Handling Edge Cases and Failure Modes
- 6. Production Best Practices
1. The Limitations of "Vibe Coding" and Prompting
Most current Large Language Model (LLM) applications rely on what we might call "vibe coding"—a lack of structure where success is determined by trying different prompt variations until the output "feels" right. This approach scales poorly for three reasons:
- Non-Determinism: LLMs are probabilistic. Without constraints, the same input can yield wildly different outputs, making regression testing nearly impossible.
- Hidden State: The logic resides inside the model’s weights. You cannot inspect why the model chose a specific tool or format, leading to blind spots in debugging.
- Brittleness: Small changes in upstream data or model versions can break downstream logic that wasn't explicitly defined.
To move to production-grade systems, we must shift from implicit (hope the model understands) to explicit (the model must follow a defined schema). This is where protocol engineering enters the picture.
2. The Bernstein Agent Protocol: Defining Contracts
Note: In this context, we refer to the "Bernstein Protocol" as a conceptual framework named after the principles of formal system verification adapted for LLM agents, often associated with rigorous agent orchestration patterns. It emphasizes that an Agent is not just a function, but a state machine with defined invariants.
The core tenet of this protocol is that an AI Agent should be treated as a black box with white-box interfaces. The interfaces are the Tools (functions the agent can call) and the Output Schemas (structured formats the agent must adhere to).
Key Components of the Protocol
- Explicit Tool Signatures: Instead of describing tools in natural language, define them with strict JSON Schema. The LLM sees the schema, not the prose.
- Loop Invariants: Define what must remain true throughout the agent's execution (e.g., "The agent must never call the
delete_dbfunction"). - Verification Gates: After every step, verify the output against a checker function before passing it to the next step.
By treating the agent interaction as a protocol handshake, we gain the ability to unit test the logic of the agent, not just the quality of its text.
3. TruLens: Observability for LLM Systems
Even with strict protocols, LLMs can fail in subtle ways. TruLens (formerly from TruEra, now open-source) provides the observability layer needed to catch these failures. It transforms unstructured LLM logs into structured data points that can be evaluated against specific metrics.
Why TruLens?
Unlike generic APM tools (like Datadog or New Relic), TruLens is purpose-built for LLMs. It allows you to define custom evaluation functions that score:
- Faithfulness: Did the answer actually use the retrieved context?
- Groundedness: Is the claim supported by the context?
- Relevance: Does the answer address the user's intent?
In our Bernstein-inspired architecture, TruLens acts as the Verification Gate. It doesn't just log; it scores and flags potential protocol violations.
4. Architecting the Verifiable Agent: A Step-by-Step Implementation
Let's build a simple code-review agent that uses the Bernstein principles. We will use Python, a hypothetical LLM client, and TruLens for evaluation.
4.1. Defining the Protocol
First, we define the strict contract. The agent must output a JSON object conforming to a specific schema. If it fails, the pipeline halts.
import json
import requests
import numpy as np
import trulens as tru
from trulens.evaluation import feedback
from trulens.providers.openai import OpenAI
from trulens.core import Context
# Define the strict output schema for our agent's response
AGENT_OUTPUT_SCHEMA = {
"type": "object",
"properties": {
"issues": {
"type": "array",
"items": {
"type": "object",
"properties": {
"line_number": {"type": "integer"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
"description": {"type": "string"}
},
"required": ["line_number", "severity", "description"]
}
},
"summary": {"type": "string"}
},
"required": ["issues", "summary"]
}
# The Tool Specification (Prompt) strictly defined
TOOL_DEFINITION = f"""
You are a senior software engineer.
You MUST respond with valid JSON that strictly adheres to this schema:
{json.dumps(AGENT_OUTPUT_SCHEMA, indent=2)}
Analyze the following code snippet:
"""
4.2. Implementing the Agent Logic
We create a wrapper that ensures the LLM adheres to the protocol. We use a verify_output function that acts as a hard gate.
def parse_and_verify(llm_output: str) -> dict:
"""
Parses the raw LLM output and verifies it against the AGENT_OUTPUT_SCHEMA.
If it fails, raises a ProtocolViolationError.
"""
try:
parsed = json.loads(llm_output)
except json.JSONDecodeError as e:
raise ValueError(f"Protocol Violation: Output is not valid JSON. Error: {e}")
# Basic structural validation
if not isinstance(parsed, dict):
raise ValueError("Protocol Violation: Root element must be an object.")
# Validate against schema (simplified check for this example)
if "issues" not in parsed or "summary" not in parsed:
raise ValueError("Protocol Violation: Missing required fields 'issues' or 'summary'.")
return parsed
def agent_llm(code_snippet: str) -> str:
"""
Simulates an LLM call. In production, this would be an OpenAI/Anthropic call.
"""
# Pseudo-code for LLM interaction
# response = openai.chat.completions.create(
# model="gpt-4",
# messages=[{"role": "user", "content": TOOL_DEFINITION + code_snippet}]
# )
# return response.choices[0].message.content
# For demonstration, returning a mocked compliant JSON
return json.dumps({
"issues": [
{
"line_number": 5,
"severity": "high",
"description": "Potential null pointer dereference."
}
],
"summary": "Code has critical safety issues."
})
4.3. Integrating TruLens for Feedback
Now, we wrap this with TruLens to record the evaluation. We create a custom feedback function that checks if the "severity" is justified by the "description" (a simple heuristic for groundedness).
def feedback_severity_justified(ctx: Context):
"""
Custom metric: Checks if high severity issues have descriptive explanations.
"""
if ctx.user_input == "":
return None
output = ctx.output
if not isinstance(output, dict) or "issues" not in output:
return 0.0 # Fail hard if schema broken
score = 1.0
for issue in output["issues"]:
if issue.get("severity") == "high" and len(issue.get("description", "")) < 20:
score -= 0.5
if score < 0:
score = 0.0
break
return score
# Initialize TruLens runner
oai = OpenAI(model="gpt-4")
runner = tru.Runner(
contextual=Context(),
metrics=[feedback_severity_justified],
providers=oai,
record_feedback=True
)
def run_agent_pipeline(code: str):
"""
The main entry point that integrates the protocol and observability.
"""
# 1. Create Context for TruLens
ctx = Context()
ctx.user_input = code
# 2. Execute Agent (Protocol Step)
raw_response = agent_llm(code)
# 3. Verify Protocol (Hard Gate)
try:
verified_response = parse_and_verify(raw_response)
ctx.output = verified_response
ctx.record()
except ValueError as e:
# Log protocol violation
ctx.output = {"error": str(e)}
ctx.record()
return {"status": "protocol_violation", "details": str(e)}
# 4. Run TruLens Evaluation
# In a real async environment, this would be awaited
for metric in runner.metrics:
metric.evaluate(ctx, feedback_severity_justified)
return {"status": "success", "data": verified_response}
# Example Usage
if __name__ == "__main__":
sample_code = "def f(x):\n if x == None:\n pass"
result = run_agent_pipeline(sample_code)
print(result)
5. Handling Edge Cases and Failure Modes
Even with the parse_and_verify gate, LLMs exhibit unique failure modes that standard software testing doesn't catch.
5.1. Hallucinated Fields
The model might invent a field "confidence": 0.9 that is not in the schema. While our simple parser above allows extra fields, strict schema validation (using a library like jsonschema) should reject these.
Mitigation: Use a strict validator. If the model adds fields, treat it as a protocol violation. Force the model to adhere to the schema by including negative examples in the system prompt.
5.2. Recursive Loops
Agents that use tools can get stuck in infinite loops (e.g., calling search -> read -> search -> read). The Bernstein protocol mandates Step Limits.
Mitigation: Implement a counter in the agent loop. If step_count > MAX_STEPS, terminate the agent and return a partial result with a termination_reason: "max_steps_exceeded".
5.3. TruLens Noise
TruLens metrics (like Faithfulness) rely on LLM-as-a-judge, which can be noisy. A score of 0.8 might be a fluke.
Mitigation: Do not treat single TruLens scores as binary pass/fail gates for production logic. Instead, aggregate scores over time. Use TruLens to build a dashboard for trends, not real-time blocking gates. Use the schema validation as the hard gate, and TruLens as the soft quality monitor.
6. Production Best Practices
To scale this architecture, adhere to the following engineering principles:
- Separate Concerns: Keep the LLM provider, the Agent Logic, and the Observability Stack (TruLens) in separate services or modules. This allows you to swap LLM providers without breaking the agent logic.
- Version Your Prompts: Treat prompts as code. Use a version control system for prompts. When you change a prompt, you must re-run your TruLens evaluation suite to ensure quality hasn't degraded.
- Async Evaluation: TruLens evaluation can be slow (it makes LLM calls to judge other LLM calls). Run these evaluations asynchronously (e.g., in a background worker) so they don't block the user-facing API response. Return the user with the agent's result immediately, and update the TruLens dashboard later.
- Fail Fast: If the protocol violation occurs (bad JSON, missing fields), fail immediately and return a structured error to the caller. Do not attempt to "fix" the JSON via LLM post-processing in the hot path; that increases latency and cost. Log the failure for offline analysis.
Frequently Asked Questions
Q: Is TruLens compatible with all LLM providers?
A: TruLens is provider-agnostic. It works with OpenAI, Anthropic, HuggingFace, and any local model via LangChain or LlamaIndex integrations. You just need to map the provider's response format to the TruLens context.
Q: How do I unit test the Bernstein Protocol itself?
A: You don't need the LLM to test the protocol. You can mock the LLM response with valid/invalid JSON strings and assert that your parse_and_verify function behaves correctly (returns dict for valid, raises exception for invalid). This is standard pure-function unit testing.
Q: Should I use TruLens for real-time monitoring?
A: It depends on your latency requirements. For customer-facing chat apps, it may be too slow for inline evaluation. Use it for offline batch evaluation of all logged sessions to catch quality regressions. For real-time, rely on schema validation and lightweight heuristics.
For more advanced patterns in LLM observability and agent orchestration, check out our other guides on tamiz.pro or deep-dive into specific metrics on Tamiz's Insights. Transitioning from prompt to protocol is not just about better answers—it's about building systems you can trust.
Top comments (0)