DEV Community

karmendra pandey
karmendra pandey

Posted on

How to Audit an ML Design Before It Ships

How to Audit an ML Design Before It Ships

Most ML failures are not model failures. They are design failures — and they are almost always visible before launch to anyone who knows where to look.

Training-serving skew nobody checked. An evaluation set that doesn't match production. A feedback loop that will quietly poison the next retraining run. Cost and latency budgets nobody wrote down. I've watched these kill systems that had perfectly good models. After 19+ years of shipping software and reviewing AI research, I now run every ML design through the same five-gate audit before it ships. Here it is.

Gate 1: Data — provenance, leakage, drift

Ask three questions. Where did every training example come from, and can you prove it? Could any information from the future — or from the test set — have leaked into training? And what happens when the real world drifts away from your training distribution?

The classic killer here is eval leakage: the "99% accurate" classifier that fails on its first real day because the eval set was accidentally drawn from training data. It happens more than anyone admits. Demand a written data provenance statement and a leakage check as merge requirements, not nice-to-haves.

Gate 2: Evaluation — does the metric match the mission?

A model can ace its metric and still fail its job. Offline accuracy means nothing if the production decision has different costs for different errors — a fraud model optimized for accuracy will happily approve everything when fraud is rare.

For every model, write down: what decision does this model actually drive, what does a mistake cost in each direction, and does our eval set look like production traffic? If the answers are vague, the design isn't ready.

Gate 3: Operations — latency, cost, fallback

Nobody writes down the latency budget until the first user complaint. Nobody prices the prediction until the first cloud bill. For each model, record: p95 latency budget, cost per 1,000 predictions at expected volume, and what happens when the model is down or too slow. "The page breaks" is not a fallback strategy — cache the last good prediction, degrade to a simpler model, or fail open explicitly and loudly.

For agentic systems this gate matters 10x more: every unnecessary tool call burns tokens and inflates every downstream context window. Cost discipline starts with call discipline.

Gate 4: Security — adversarial inputs and access control

Who can feed inputs to this model, and what can they make it do? Cover the basics: input validation, rate limiting, access controls on the model endpoint and its training data. For LLM-backed systems, add prompt-injection review — treat every tool the agent can call as a privilege, and scope it like one. I model agents as IAM principals: unique identity, least-privilege roles, session-scoped credentials.

Gate 5: Accountability — who owns the decision?

When the model is wrong — and it will be — who is accountable, and what is the remediation path? Every production ML system needs a named owner, a rollback plan, and a written policy for the failure modes you know exist. "The model decided" is not an accountability structure.

The one-page audit

Turn these gates into a single page your team fills out at design review:

  1. Data: provenance documented? leakage check passed? drift monitors planned?
  2. Evaluation: metric matches the mission? eval set production-representative?
  3. Operations: latency budget, cost per prediction, fallback behavior defined?
  4. Security: inputs validated? access scoped? agent tools least-privilege?
  5. Accountability: named owner? rollback plan? known failure modes written down?

Any blank answer blocks the ship. It takes an hour, and it catches the failures that postmortems always describe as "obvious in hindsight."


Karmendra Pandey is a Practice Architect in AI & ML at TEKsystems. He builds production agentic AI systems on AWS and peer-reviews AI research on PREreview.

Top comments (0)