DEV Community

Cover image for Running an LLM as a Pipeline Step Without 2am Surprises
Vaishnav Prabhu
Vaishnav Prabhu

Posted on

Running an LLM as a Pipeline Step Without 2am Surprises

Teams are dropping LLM calls into data pipelines everywhere now — to classify tickets,
extract fields from messy text, summarize records, tag content. It feels like adding
one more transform. It isn't. A SQL transform is a pure function: same input, same
output, every time. An LLM step is a slow, rate-limited call to an external service
that can return something different each run, or something that isn't even valid.

Treat it like the unreliable network dependency it is, and the 2am surprises mostly
go away. Here's the short playbook.

1. Force structure, then validate it

Never let free-form text flow downstream. Ask the model for structured output
(JSON that matches a schema) and validate every response against that schema before
you trust it. If it doesn't parse or a required field is missing, that's a failure —
retry it, route it to a dead-letter table, or fall back. The moment you depend on a
model "usually" returning clean output is the moment a malformed response corrupts a
downstream table.

2. Make re-runs cheap and safe

LLM calls cost money and time, so a naive re-run can be expensive and non-reproducible.
Two habits fix this:

  • Cache by input hash. If you've already processed this exact input with this exact prompt and model version, reuse the result. Re-runs become near-free.
  • Pin your versions. Record the model name and prompt version with each output. "The numbers changed" is impossible to debug if the model silently updated underneath you.

3. Budget for failure up front

Wrap the call the way you'd wrap any flaky API: a timeout, retries with backoff, and a
cost ceiling so a runaway loop can't quietly burn your budget. Decide in advance
what happens when the model is down or slow — skip, use last-known-good, or fail the
run loudly. Don't let a provider outage become a silent data gap.

4. Treat evals as a data-quality gate

This is the big one. Before an AI step's output reaches production, gate it like any
other quality check:

  • % valid — what share of responses passed schema validation?
  • Accuracy on a sample — hold a small labeled set and measure agreement each run.
  • Drift — did the distribution of outputs shift sharply versus a recent baseline?

If those fall below threshold, the run goes red and stops — exactly like a row-count or
uniqueness gate. An LLM step with no eval is a green checkmark hiding unknown quality.

5. Observe the things that only AI steps have

Log tokens, latency, and cost per run alongside a quality metric, and alert on drift.
These are the signals that tell you a prompt change or a model update quietly degraded
things — long before a human notices weird results in a report.

Takeaways

  • An LLM step is a non-deterministic external call, not a pure transform — engineer it like a flaky dependency.
  • Demand structured output and validate it; cache by input; pin model and prompt versions.
  • Gate outputs with evals the same way you gate data quality, and observe tokens, cost, latency, and drift.

Related in The Grain*: modeling your warehouse so these models have correct, usable
data to work with in the first place.*

Top comments (0)