DEV Community

Cover image for The Gateway, the Registry and the Executor: Keeping Groq at the Edge of My Backend
Kirera paul murithi
Kirera paul murithi

Posted on

The Gateway, the Registry and the Executor: Keeping Groq at the Edge of My Backend

SokoFlow AI Build Log, Week 2 of 6: the LLMClient, the ToolExecutor, and one component that almost became a god object

The Week 2 goal said "wire three tools end-to-end."

By the end of the week, all three tools executed against PostgreSQL and came back as structured results. I still wouldn't call it end-to-end, and the reason is the most useful thing I learned this week.

TL;DR

  • Week 2 built the provider boundary (LLMClient) and the execution boundary (ToolExecutor), then ran three tools through them: get_sales_summary, get_stock_level and get_top_products.
  • LLMClient is a translator and gateway. It keeps Groq's concepts at the edge so the rest of SokoFlow speaks SokoFlow's language. It does not know what a tool means.
  • I got two things wrong first: I let LLMClient become a "god" object, and I let ToolExecutor own domain logic. A simple "Milk" query showed me why that's wrong.
  • I verified the execution layer, not the full LLM round trip. I'm not going to claim a loop works when I only tested half of it.

Where we are

SokoFlow is a headless WhatsApp ERP for Kenyan SMEs. There's no app and no UI to learn; for a shopkeeper, the conversation is the interface. Phase 3 adds one AI feature: natural-language shop analytics. A shopkeeper asks "which products are running low?" or "how does this week compare to last week?", and an LLM, using function calling against SokoFlow's existing services, figures out which query to run.

Last week I designed seven read-only tools and benchmarked three models across two providers on the same 10 shopkeeper queries. Groq's openai/gpt-oss-20b won. But right now those tools sit in isolation, and on their own they aren't worth much. This week's goal was to connect at least three of them.

That takes three components: an LLMClient, a ToolRegistry and a ToolExecutor. Each has its own task and its own boundary. So why do we need all three?

The problem: provider coupling

The first thing I had to understand was the provider boundary. Last week I picked one provider for this phase, and my first instinct was the obvious one: call the Groq SDK directly wherever SokoFlow needs the model.

SokoFlow component
        ↓
    Groq SDK
        ↓
       Groq
Enter fullscreen mode Exit fullscreen mode

I quickly saw the flaws in that design. The biggest was provider coupling: Groq's concepts would leak into the whole application. The second was cross-cutting duplication: logging, latency tracking, retries, error mapping and metrics would get scattered everywhere.

And switching providers isn't a hypothetical for me. Llama 3.3 70B on OpenRouter also scored 10/10 in last week's benchmark, so a swap is a realistic future decision.

LLMClient: one gateway to the provider

There had to be a better way, and that's where I learned the idea of an LLM client. What if there were exactly one gateway to the provider? The rest of the application talks to that client and doesn't care who's behind it.

SokoFlow
    │
    ▼
LLMClient
    │
    ▼
Provider SDK / API
    │
    ▼
Groq
Enter fullscreen mode Exit fullscreen mode

The deeper purpose isn't just making providers replaceable. It's keeping provider-specific concerns at the edge of the system, so the rest of SokoFlow can reason in SokoFlow's own language. The LLMClient plays two roles:

  • Translator: SokoFlow input becomes a provider request, and the provider response becomes SokoFlow output.
  • Gateway: it's the one controlled point through which SokoFlow talks to the external LLM.

If Groq exposes response.choices[0].message.tool_calls..., nothing outside LLMClient should ever need to know that structure exists.

The "god object" I almost built

I was treating LLMClient as the silver bullet. It was the right answer, but I kept giving it more to do. At first I figured it would also handle tool semantics, tool existence checks, and the retry policy for semantically bad model responses. It was turning into a "god" object that knows far too much.

So I defined three contracts for the provider boundary.

Input contract. What SokoFlow sends to the LLM: the conversation messages, the available tool definitions, and the model configuration. The rule: SokoFlow expresses the request in a SokoFlow-owned representation, never a provider SDK request object. The client translates it into the provider's format.

Output contract. What SokoFlow actually needs back. The first time I saw a raw response, there was a lot in it (abbreviated here):

{
    "choices": [...],
    "usage": {...},
    "system_fingerprint": "...",
    "id": "...",
    "created": 123456,
    ...
}
Enter fullscreen mode Exit fullscreen mode

Should the client hand back the entire Groq response object, or a small SokoFlow-owned result? Returning the whole object would recreate the exact coupling I was avoiding, with Groq details spread through the app. So the client normalizes the response into a small result with only what the application needs.

Error contract. Groq's SDK exposes its own exceptions, like groq.APIConnectionError and groq.APITimeoutError. If application code catches those, error handling is coupled to Groq too. Instead, provider-specific failures are translated at the LLMClient boundary into provider-agnostic application errors, wherever SokoFlow needs to react to them.

The responsibility split: who handles a bad tool call?

Should LLMClient validate tool names and arguments? To answer that, I ran a thought experiment. Suppose the provider returns:

tool = "delete_everything"
arguments = {"random invalid": true}
Enter fullscreen mode Exit fullscreen mode

Who handles it? My answer: not the LLMClient. As far as the client is concerned, it got a response from the provider. No infrastructure or provider error happened, so it did its job as gateway and translator. It shouldn't know what get_sales_summary or delete_everything means, which periods are valid, or what the business rules are. So the design decision was to normalize the request and let downstream components deal with it.

(One more piece I'll keep mentioning: the Orchestrator. It will sit above these components, take the model's response, run the tool, and decide what happens next. It wasn't finished this week. More on that below.)

Why the ToolRegistry exists

Continuing the same thought experiment: the model asked for delete_everything. Who detects that no such tool exists?

From Week 1, the ToolRegistry is the single source of truth for every tool SokoFlow exposes and the backend handler behind each one. So if any component should detect a missing tool, it's the registry. Its job is to answer one question:

"What capabilities does SokoFlow expose to the model, and how does a proposed capability map into the backend?"

But that left a gap. We now know the tool exists. What about running it?

The ToolExecutor: the execution boundary

That's the third component. LLMClient talks to the provider, the registry confirms a supported tool was requested, and the ToolExecutor owns the execution boundary:

  • Are the arguments structurally valid?
  • Convert them into the typed input the handler expects.
  • Call the handler.
  • Normalize the outcome.

What I got wrong first: I started by designing the executor to handle all domain logic. Then I considered this:

get_stock_level(product_name="Milk")
Enter fullscreen mode Exit fullscreen mode

Suppose "Milk" matches three products: Milk 500ml, Milk 1L and Milk 2L. The model supplied a perfectly valid string according to the schema. But the domain operation knows the product is ambiguous and needs clarification. A ToolExecutor containing product-resolution rules would be doing the domain's job.

So the split became clear: schema and contract validation belong to the executor. Business and domain validation stay with the existing domain services. This revises what I described last week, when I put domain validation in the same gate as schema validation.

The final failure taxonomy

With that boundary, there are three distinct failure scenarios:

  1. Unknown tool, like tool_name = "delete_inventory". The registry lookup fails and the executor rejects it.
  2. Structurally invalid arguments, like get_top_products(limit=10, dimension="profit"). Schema validation fails and the executor rejects it.
  3. Structurally valid, domain-invalid request, like get_stock_level(product_name="Milk"). The schema says it's fine, but the backend can't uniquely resolve the product. That's not a schema failure. The executor validated the request and reached the handler, and the domain operation answers with AmbiguousProduct. The Orchestrator then decides whether to ask the shopkeeper to clarify.

The result envelope

One more decision. Once the executor finds the tool, validates the arguments and calls the handler, what does it return?

  1. The raw handler result, a Pydantic model such as a revenue summary (for example, {"total_revenue": 4250, "transaction_count": 17}).
  2. An executor-owned result that wraps the handler's output:
{
    "status": "success",
    "tool": "get_sales_summary",
    "data": {
        ...
    }
}
Enter fullscreen mode Exit fullscreen mode

I chose the second. It separates execution metadata (did this run succeed?) from domain data (what the handler returned), so the Orchestrator doesn't need to understand every domain result just to know whether execution worked.

The key limit: the executor owns the envelope, but it must not reshape the domain data inside it. Otherwise it grows a universal "mega-schema" and gets coupled to every tool:

ToolExecutionResult
├── revenue
├── products
├── stock
├── transaction_count
├── ...
Enter fullscreen mode Exit fullscreen mode

Instead, the envelope looks like this:

ToolExecutionResult
├── execution metadata
└── data → domain result
Enter fullscreen mode Exit fullscreen mode

The first three tools

With the boundaries set, I wired three analytics tools through the deterministic backend path: get_sales_summary, get_stock_level and get_top_products. Each one executes through the existing ToolRegistry and ToolExecutor, reuses SokoFlow's services, reaches PostgreSQL, and returns a structured result. I covered the tools themselves in the Week 1 blog.

I didn't wire in the full WhatsApp conversation (the next section explains why). Instead I leaned on the principle that guided the MVP: tests. Each tool got multiple unit tests, and there's at least one integration test touching all the related components.

Tool What the contract covers Verified by
get_sales_summary An explicit reporting period, backend-owned date calculations, shop-timezone-aware relative periods, half-open boundaries, valid zero-sales results, and total_revenue and transaction_count in the output Unit tests (mocks and patches) for validation and execution; integration test reaching PostgreSQL through the existing repository; result wrapped in ToolExecutionResult
get_stock_level Backend-owned product resolution; existing stock and product services; ambiguous matches handled by the domain Unit tests; integration test reaching PostgreSQL; result wrapped in ToolExecutionResult
get_top_products Units and revenue rankings, an optional ranking dimension, both rankings when none is given, and a limit applied to both. Structural validation rejects invalid dimensions like profit and invalid limits like zero before the handler runs Unit tests; integration test reaching PostgreSQL; result wrapped in ToolExecutionResult

The part where I changed the plan

The Week 2 objective was to wire these three tools end-to-end outside the WhatsApp flow, using a standalone script or REPL harness that also measured p95 latency.

Partway through, I hit a distinction I hadn't thought about: testing tool execution is not the same as testing the complete LLM-to-tool interaction.

The three tools had registered handlers connected to the service and repository layers, with unit tests and PostgreSQL-backed integration tests. But exercising the full LLM-driven path is harder. A genuine flow needs the model's tool call to pass through the Orchestrator, which coordinates model responses, tool execution and what happens next. At this stage, the Orchestrator and its conversation-level integration weren't complete.

I considered making a real model call with temporary workarounds. But that would have meant writing a half-baked orchestration layer ahead of schedule, and it would have made it harder to tell a tool-execution bug from an orchestration bug.

So I made a deliberate scope decision. I kept Week 2 on the provider boundary, the registry, the executor and the three tool handlers. The standalone harness feeds normalized, model-shaped tool calls straight into the ToolExecutor, so I can test successful execution and failure cases without building the WhatsApp path early.

There's an important qualification. That harness verifies the backend execution path. It does not prove that a real LLM can pick a tool, have the Orchestrator run it, and get the result back in a full round trip with latency measured and recorded. That work moves to Week 3, where analytics connects to the conversation flow and gets verified through the WhatsApp simulator.

The lesson: an end-to-end test is only meaningful when its boundaries are clear. I didn't want to claim the whole function-calling loop worked when I'd only verified its execution layer. I'd rather move the unfinished integration to the week where it belongs than force it early and trust a result I can't interpret.

What I'd do differently

I wrote the goal "end-to-end" before checking what an end-to-end run actually depends on. It depends on the Orchestrator, and that didn't exist yet. Next time I would list the dependencies of an end-to-end goal before I commit to the word.

Where Week 2 landed

Here's how responsibility is divided now:

Component Owns
LLMClient Provider interaction and the normalized tool call
ToolRegistry Exposed tools, schemas and handler mappings
ToolExecutor Lookup, structural validation, safe invocation, and the execution result envelope
Domain Service Business rules and the actual operation
Orchestrator What happens next, including retries and recovery

And against the Week 2 target of three tools wired end-to-end from a harness, with p95 latency recorded:

  • Done: the provider boundary, the executor and three tools running through the execution layer, tested against PostgreSQL.
  • Carried to Week 3: a real LLM tool call going through the Orchestrator and executor in a full round trip, with p95 latency measured.

What's next

Week 3 is the fun part: the WhatsApp and conversation integration. That means adding a new ANALYTICS_QUERY intent to the existing resolver, so complex free-text messages are routed to the LLM. It also means handling session context, so a follow-up like "and yesterday?" can reuse what came before.

Stay locked in as SokoFlow slowly gains intelligence.

Check out the source code on Github.
Read more from the 4-Month MVP Blogs.

Top comments (0)