DEV Community

Cover image for The LLM Interprets, the Backend Owns the Truth: Week 1 of Adding AI to a WhatsApp ERP
Kirera paul murithi
Kirera paul murithi

Posted on

The LLM Interprets, the Backend Owns the Truth: Week 1 of Adding AI to a WhatsApp ERP

I Spent Week 1 of Adding AI to My ERP Deciding What the AI Is Not Allowed to Do

SokoFlow AI Build Log, Week 1 of 6: tool boundaries, a timezone detour, and a bug with two correct answers

Two correct answers. One question: "Compare this month with the previous period."

My code picked one of them, silently, and never told me it had a choice.

That bug is why Week 1 of adding AI to SokoFlow had almost nothing to do with AI.

TL;DR

  • Week 1 wasn't about wiring an LLM into SokoFlow. It was about deciding what the LLM gets to own, and what it doesn't.
  • The rule I locked in: the LLM interprets language, the backend owns the truth. The model picks a tool and proposes arguments. It never executes anything, never supplies shop_id, and never decides what "today" means.
  • "Previous period" has two valid meanings. My code was choosing one without telling me. Some bugs aren't implementation failures, they're failures to define what the thing means.
  • I treat LLM output like an untrusted client request: validate the shape, then the business rules, then allow one bounded retry for cheap model slips.
  • Provider: I benchmarked three models across Groq and OpenRouter and chose Groq's openai/gpt-oss-20b: 10/10 queries routed correctly, 1309 ms average latency, about $0.00005 per call.

Where SokoFlow was

Welcome to the new series. SokoFlow is a headless WhatsApp ERP and conversational commerce engine for Kenyan SMEs. There's no app to download and no UI to learn. For a shopkeeper, the conversation is the interface. The 16-week MVP is shipped, and I wrote about it along the way. You can read the series here.

The MVP is deterministic and rule-based on purpose, and that isn't changing. But it only understands specific commands that trigger specific flows. For an ERP aimed at small shops, that's enough. I've also always wanted to explore AI systems, and extending something I already understand felt like the best way to do it.

So Phase 3 adds exactly one AI feature: natural-language shop analytics. A shopkeeper asks "which products are running low?" or "how does this week compare to last week?", and an LLM, using function calling against SokoFlow's existing internal services, figures out which query to run.

The guardrails I set up front:

  • No vector DB, no RAG, no fine-tuning.
  • No new data store, no dashboard, no multi-shop analytics.
  • No writes through the LLM. "Record a sale" stays FSM-only.

Week 1 was the Tool Surface & Provider Decision: finalize 5–8 tool schemas, benchmark LLM providers on function-calling reliability, and decide fallback policy and guardrails. The goal was not to connect an LLM to SokoFlow. It was to understand where an LLM should fit without handing it responsibilities that belong to the backend.

The problem: two messages that look alike but aren't

Imagine a shopkeeper sends one of these.

Message B: sell 2 bread

The existing system handles this well:

Intent:   RECORD_SALE
Product:  bread
Quantity: 2
Enter fullscreen mode Exit fullscreen mode

It's constrained. The intent resolver recognizes the pattern and the FSM walks the user through whatever's missing.

Message A: How much did I sell today?

There's no fixed conversational flow to enter. Something has to infer at least two things: what operation does the user want? (a sales summary) and what time period? (today).

Here's the key move. I'm not asking the LLM to answer the question. I'm asking it to interpret the question and pick an operation that can.

User's natural language
        ↓
"What does this person want?"
        ↓
"Sales summary for today"
        ↓
Which SokoFlow operation can provide that?
        ↓
get_sales_summary(...)
        ↓
Actual database/service logic
Enter fullscreen mode Exit fullscreen mode

This is tool calling: instead of replying with prose, the model replies with a structured request, usually JSON, naming a tool and its arguments. It doesn't run anything. Your code does.

When the model outputs get_sales_summary(period="today"), that isn't an answer. It's a request to SokoFlow to perform an operation. The LLM sits at the interpretation boundary between messy human language and SokoFlow's structured world, and it never becomes the source of truth.

One difference from the textbook loop: in the generic version, the tool result goes back to the model so it can write the final reply. I'm deliberately not doing that. SokoFlow's ResponseFormatter turns the structured result into a reply using templates and i18n. My reasoning: a shopkeeper's revenue figure shouldn't pass through a text generator that might rephrase a number. Predictable beats eloquent.

The architecture barely changes:

WhatsApp → Webhook → FSM / Intent Resolver → ANALYTICS_QUERY
   → LLM → tool selection + structured arguments
   → validation / executor → existing SokoFlow services
   → PostgreSQL / Redis → structured result
   → Response Formatter → WhatsApp reply → Interaction Logger
Enter fullscreen mode Exit fullscreen mode

The intent resolver flags free text as ANALYTICS_QUERY only when it doesn't match an existing structured intent. And every LLM decision gets logged before it executes, so failures can be diagnosed afterwards.

Hallucination, and why the boundary matters

An LLM isn't a database of answers. It learns statistical relationships, so it's good at recognizing that "How much did I sell…" lives near concepts like sales, amounts and time periods. That's why it can handle many phrasings of one question.

It's also why it can invent things. Suppose a shopkeeper asks "How much did I sell today?" and the model replies with get_monthly_profit. That tool doesn't exist (this is a made-up example), but the model would have invented a capability.

The model doesn't know how to operate SokoFlow. It only knows the operations we show it: names, descriptions, required arguments. So every rule below comes from one stance: the model's output is a suggestion, never an instruction.

The seven tools, and the decisions hiding inside them

I kept the surface small and read-only. Seven tools, each with a strict typed schema and bounded output.

Tool The decision that mattered
get_sales_summary shop_id is injected by the backend, never supplied by the LLM. Zero sales is a valid result, not an error.
get_top_products Optional ranking dimension (units, revenue, or both when the question is broad). A limit defaults to 10 so WhatsApp doesn't get 500 rows.
get_stock_level Users say "sugar", not product_id=42. A product-resolution step must turn names into trusted IDs.
get_low_stock_items Defaults to each product's own low_stock_threshold. Explicit thresholds like "fewer than 3 units" don't exist yet in this phase.
get_product_price Same name-to-ID problem as stock. Current price comes from Product.price, not the historical Sale.unit_price snapshot.
get_slow_moving_items Products whose total units sold over the period are at or below max_sales_count, which defaults to 2 unless overridden.
get_sales_trend The LLM interprets "this week vs last week"; the backend computes the exact boundaries. The metric (revenue, units, transaction count, or all) depends on the question.

Three of these decisions are worth explaining.

The LLM never supplies shop_id. For a shopkeeper in Nairobi, the worst-case bug is seeing someone else's sales numbers. Scope comes from the authenticated session, injected server-side. The model can't ask for another shop because the argument doesn't exist in its schema.

"Sugar" must become a trusted ID before any tool runs. I already had fuzzy matching with pg_trgm and candidate handling from the MVP, so I reused it. And the system doesn't guess: if several products match, or the best match scores low, it asks the shopkeeper to clarify. A missing product is a domain error, and I created a custom error for it.

The LLM interprets the meaning of "today", but it doesn't decide what date it is. That belongs to the application, using the shop's timezone.

Explicit over ambiguous: the schema that argued with itself

Consider "How much did I sell this week?" The LLM can tell the user means a weekly period. But should it calculate timestamps? I first sketched a schema like this:

relative_period: day | week | month | year
exact_start_date: string | optional
exact_end_date: string | optional
Enter fullscreen mode Exit fullscreen mode

Now the schema allows contradictory input:

relative_period = "week"
start_date = "2026-09-01"
end_date = "2026-09-30"
Enter fullscreen mode Exit fullscreen mode

Who wins? Nobody can say, because those fields aren't different parameters. They're two ways of expressing one concept: the requested reporting period. So the schema should force the model to answer one question first: am I expressing this period relatively, or as an explicit range?

period_type
├── relative
│   └── relative_period: day | week | month | year
│
└── date_range
    ├── start_date
    └── end_date
Enter fullscreen mode Exit fullscreen mode

That gives each layer a clean job:

  • LLM: chooses the strategy and supplies the semantic values.
  • Schema validation: ensures the chosen strategy has the required fields.
  • Backend: derives actual boundaries using the shop timezone, calendar rules, ZoneInfo and half-open intervals.
  • Database: receives deterministic boundaries and just queries.

period_type is the discriminator, and the application enforces it with explicit checks:

if self.period_type == PeriodType.RELATIVE:
    if self.relative_period is None:
        raise ValueError("relative_period is required when period_type is 'relative'.")
    if self.date_range is not None:
        raise ValueError("date_range must be null when period_type is 'relative'.")
elif self.period_type == PeriodType.DATE_RANGE:
    if self.date_range is None:
        raise ValueError("date_range is required when period_type is 'date_range'.")
    if self.relative_period is not None:
        raise ValueError("relative_period must be null when period_type is 'date_range'.")
Enter fullscreen mode Exit fullscreen mode

A small detour: what does "today" even mean?

Time keeps coming back in this project, and in backend work in general. "Show me today's sales" sounds trivial. But a backend doesn't live inside the user's wall clock. It deals with instants on a global timeline. Nairobi is UTC+3, so a sale rung up at 1 a.m. on the 5th is still the 4th in UTC. A shop owner closing late and a server clock can disagree about which day that sale belongs to.

So the shop's timezone defines "today". The backend converts the reference instant into that timezone, derives the calendar date, and turns relative periods into calendar boundaries, not durations. "This month" isn't "the last 30 days". It's:

[start of this month, start of next month)
Enter fullscreen mode Exit fullscreen mode

That's a half-open interval: the start is included, the end isn't. It beats inventing an endpoint like 23:59:59.999999, because the next period's start is exactly this one's end, with no gaps and no overlap.

I used ZoneInfo, astimezone(), datetime.combine() and timedelta, but the APIs weren't the lesson. The lesson was the model: an instant, a timezone, a local calendar and boundaries are four different concepts. In SokoFlow, the week starts on Monday, and "today" is always evaluated in the shop's timezone.

The bug with two right answers

Back to the hook. A user asks: "Compare this month with the previous period." What does "previous period" mean?

I caught this while manually reviewing one of the reports. The previous month's range started on August 2 instead of August 1.

Here's how get_sales_trend worked when I found it. If the caller gives an explicit comparison period, the handler resolves it directly:

if input_data.comparison_period:
    prev_start, prev_end = resolve_period_boundaries(
        input_data.comparison_period
    )
Enter fullscreen mode Exit fullscreen mode

If not, it derived one automatically:

else:
    prev_start, prev_end = resolve_previous_period_boundaries(
        input_data.period
    )
Enter fullscreen mode Exit fullscreen mode

And that helper did it by subtracting the current period's length:

duration = curr_end - curr_start
prev_end = curr_start
prev_start = prev_end - duration
Enter fullscreen mode Exit fullscreen mode

For September, that gives:

Current:   Sep 1 ──────── Oct 1   (30 days)
Previous:  Aug 2 ──────── Sep 1   (30 days)   ← equal duration
Enter fullscreen mode Exit fullscreen mode

But a shopkeeper asking about "last month" means all of August:

Current:   Sep 1 ──────── Oct 1   (30 days)
Previous:  Aug 1 ──────── Sep 1   (31 days)   ← previous calendar period
Enter fullscreen mode Exit fullscreen mode

Both are mathematically valid. Only the domain tells you which is correct. SokoFlow already thinks in calendar units (day, week, month, year), and months have different lengths (30, 31, 28 or 29 days), so "same number of days" isn't a reliable meaning of "equivalent period."

What I chose: by default, "previous period" means the previous equivalent calendar period. The duration-based version isn't wrong, it's a different question: "the last 30 days vs the 30 days before that." That should be asked for explicitly, not inherited by accident:

Previous calendar period → "previous month", "previous week"
Previous duration window → "the 30 days immediately before this 30-day window"
Enter fullscreen mode Exit fullscreen mode

An explicitly provided comparison_period still resolves through resolve_period_boundaries(). I went with the calendar option and recorded the decision in an ADR (architecture decision record), so the meaning is written down instead of living in one helper function.

The lesson: some backend bugs aren't failures of implementation. They're failures to define what the implementation is supposed to mean. And it matters more with an LLM in front, because the model will happily translate a vague phrase into a vague argument.

Never trust the client, even when the client is an LLM

When the model produces a tool call, is that a request or a response? Both, depending on where you stand. In HTTP terms I sent the user's message and the model responded. In application terms it's a request that my backend executes. And how do we treat requests? We don't blindly trust them. That was my turning point.

The tempting shortcut is passing the model's arguments straight to the handler:

await tool.handler(shop_id, raw_args, db)
Enter fullscreen mode Exit fullscreen mode

That's dangerous. The model can hallucinate arguments, add parameters that don't exist, or omit required ones, and any of those can crash application code. So the arguments go through a gate, in two stages:

LLM → raw_args → VALIDATION (Pydantic schema)
                    ├── ❌ invalid → stop
                    └── ✅ valid → handler
Enter fullscreen mode Exit fullscreen mode
  • Stage 1, schema validation: does this data have the right shape and types? A missing required date, or a date that isn't a date, fails here.
  • Stage 2, domain validation: is it valid but nonsensical? A start date after the end date has the right types and still makes no business sense.

The JSON Schema the model sees is generated from the same Pydantic models that validate its output, so there's one definition instead of two that can drift apart. Unexpected extra arguments aren't silently stripped. They cause a validation error and the call is rejected.

The failure cases I designed for

  • Missing tool: the model emits a name like get_sales_summery. Caught at registry lookup. (An illustrative example, not one I've observed yet.)
  • Malformed or missing arguments: caught by stage 1.
  • Domain validation failure: caught by stage 2.
  • Product ambiguity: "How much is milk?" matches several products. We don't guess, we ask.
  • Infrastructure failure: the database or Redis is down. The shopkeeper did nothing wrong.

Who is at fault?

Not all errors are the same, and the awkward part is that a missing tool isn't the user's fault. Returning a generic "invalid request" would be bogus. The user's contract with SokoFlow is "I ask a question in natural language." The LLM's contract with my tool layer is "produce a valid tool call from the tools exposed." An invented tool name is a model failure, not a user failure.

That left two options:

  1. Retry once with the LLM, with a constrained message: "Invalid tool name. Select one of the available tools."
  2. Fail the AI operation with a generic fallback like "I couldn't process that request right now," with no raw model error, no stack trace, and no internals exposed.

I chose option 1. "How much did I make this week?" is a perfectly valid question, and failing it over a one-character generation slip punishes the shopkeeper for the model's typo. I worried this would balloon into a retry system. It doesn't need to:

attempt = 1
LLM → tool call
      ↓
ToolRegistry
      ↓
ToolNotFound
      ↓
if attempt == 1:
    retry once
else:
    fail + log
Enter fullscreen mode Exit fullscreen mode

The principle I locked in: recover from cheap, model-local errors when recovery is deterministic and bounded. Don't build elaborate machinery for failures that belong to the application or infrastructure.

The registry: deliberately boring

The ToolRegistry is the single source of truth. It owns the schemas the model sees, the routing map from tool name to the code that implements it, and the validation definitions.

registry.register_tool(
    name="get_sales_summary",
    description=(
        "Calculates total revenue and transaction count for a shop over a requested "
        "reporting period."
    ),
    input_model=GetSalesSummaryInput,
    output_model=GetSalesSummaryOutput,
    handler=handle_get_sales_summary,
)
Enter fullscreen mode Exit fullscreen mode

It doesn't run queries. A string like "get_sales_summary" maps to one handler, which calls the service method that does the work (here, SalesService.fetch_sales()). Anything not in the map doesn't exist as far as the executor is concerned.

The provider decision

I wanted to stay on free or cheap models, like with the MVP. My early runs were noisy (OpenRouter returned 404s), so I tightened the benchmark script: it runs the same 10 sample shopkeeper queries against each model and scores function-calling reliability, latency and cost.

Provider / model Requests Generated calls 2xx Routed (tool + schema) Avg latency Avg cost / call Statuses
Groq (openai/gpt-oss-20b) 10 10 10 100% (10/10) 1309 ms $0.00005 200: 10
OpenRouter Free (openrouter/free) 10 8 10 20% (2/10) 4590 ms $0.0000 200: 10
OpenRouter (meta-llama/llama-3.3-70b-instruct) 10 10 10 100% (10/10) 1905 ms $0.00042 200: 10

A query only counted as correctly routed if the model picked the exact expected tool and its arguments passed my strict local Pydantic validation (schema and domain checks).

Decision: Groq with openai/gpt-oss-20b. It tied with Llama 3.3 70B at 10/10, so latency and cost broke the tie: 1309 ms vs 1905 ms, and about 8x cheaper per call. That clears the Week 1 bar of 8/10 correctly routed.

One honest caveat: 10 queries is a small sample, and the two top models tied. The real test comes in Week 5, with a 25–40 query eval set.

What I'd do differently

I found the "previous period" bug by eyeballing a report, not because something told me the definition was missing. Next time, the meaning of a business phrase gets written down before the first resolver is written, not after.

What's next

Week 1 was less about writing code and more about deciding what the AI layer should be responsible for. The architecture has a shape, the boundaries are clearer, and I have something concrete to build against.

Week 1 KPIs:

  • Finalized and froze the schemas for all 7 read-only tools (sales summary, top products, stock level, low stock, price, sales trend, slow moving).
  • Built a benchmark script with 10 real shopkeeper-style queries.
  • Benchmarked 3 models across 2 providers on function-calling accuracy, latency and cost per call. The chosen model routed 10/10 queries correctly (target: 8/10).

Next week the question changes from "what should this system look like?" to "can I actually make it work?" The plan is the core function-calling loop: LLMClient, ToolExecutor, and three tools (sales summary, stock level, top products) wired end-to-end from a manual harness, with p95 latency measured.

The LLM interprets language. The backend owns the truth. See you next week.

Top comments (0)