DEV Community

Ujjwal Dubey
Ujjwal Dubey

Posted on

Measure Cost per Handled Enquiry: A Small Logging Pattern for LLM Workflows

Pre-build cost estimates are guesses. After launch you can measure. The most useful number I know for a small-business workflow is cost per handled enquiry in AI automation: everything it cost to take one customer message from arrival to a resolved state.

Disclosure: I run NxFlowAI, an automation agency. The pattern below is generic.

What to count

For each enquiry, record:

  • model tokens in and out (per call)
  • messaging events (per message sent, by category, using your provider's billing categories)
  • automation-platform runs or tasks
  • human minutes (time spent approving or rewriting a draft)

A minimal event log

import json, time, uuid

LOG = "cost_events.jsonl"

def log_event(enquiry_id, kind, qty, meta=None):
    with open(LOG, "a") as f:
        f.write(json.dumps({"ts": time.time(), "enquiry_id": enquiry_id,
                            "kind": kind, "qty": qty, "meta": meta or {}}) + "\n")

# inside your workflow
eid = str(uuid.uuid4())
log_event(eid, "tokens_in", 812, {"model": "your-model"})
log_event(eid, "tokens_out", 164, {"model": "your-model"})
log_event(eid, "wa_message", 1, {"category": "service"})
log_event(eid, "platform_run", 3)
log_event(eid, "human_minutes", 2, {"step": "approve_quote_reply"})
Enter fullscreen mode Exit fullscreen mode

Roll it up with your own rates

from collections import defaultdict

RATES = {  # fill in from your own invoices; units must match qty
    "tokens_in": None, "tokens_out": None,
    "wa_message": None, "platform_run": None, "human_minutes": None,
}

def cost_per_enquiry(path=LOG):
    per = defaultdict(float)
    for line in open(path):
        e = json.loads(line)
        rate = RATES.get(e["kind"])
        if rate is not None:
            per[e["enquiry_id"]] += e["qty"] * rate
    return per
Enter fullscreen mode Exit fullscreen mode

What the numbers tell you

  1. Human minutes dominate more often than tokens. If approvals take most of the cost, look at which drafts are rejected and why.
  2. Long prompts show up fast. If token cost per enquiry is creeping up, check what context you are stuffing into every call.
  3. Retries hide in platform runs. A spike in runs per enquiry usually means a failing step.

Keep it boring

Append-only JSONL, one roll-up script, a weekly look. That is enough for most small businesses. Ship a dashboard later if anyone actually opens it.

We set this up as part of monitoring after the 72-hour audit and build. If you are still deciding whether to go custom at all, our custom AI vs off-the-shelf chatbot comparison covers the trade-offs.

Top comments (0)