DEV Community

Cover image for Not another chatbot: we built an open-source agent harness for customer service
blackmouse572
blackmouse572

Posted on

Not another chatbot: we built an open-source agent harness for customer service

Most customer service bots are a model, a prompt and a webhook. Ecbot is an agent harness for customer service: the runtime that sits between a model and your customers' chat apps, and decides what the model sees, what it's allowed to do, and when it has to stop and call a person.

Coding agents made the word "harness" familiar. The model writes the code, but the harness gives it tools, manages its context, runs the loop and keeps it from doing damage. Customer service needs the same thing, with different tools and different dangers. Ecbot's harness gives the model a shop's documents, the shop's own APIs and a set of written playbooks. It is omnichannel, so it keeps one memory of each customer across every channel, and it hands the chat to a human when the customer gets angry or the model is unsure.

There is no vendor lock-in. Ecbot routes through OpenRouter, so you can pick any model per agent, including DeepSeek and MiniMax, which keep the bill low and stay smart. The harness is the part you'd otherwise spend months writing.

It's open source under AGPL-3.0, and it's the same engine we run as a hosted product, with every feature included. Six channels work today: Facebook Messenger, Zalo OA, WhatsApp Business, Telegram, a website widget and a REST API channel.

 Messenger · Zalo · WhatsApp · Telegram · Widget · REST
                         │  webhooks
 ┌───────────────────────▼────────────────────────────┐
 │ Harness                                            │
 │  durable inbox → dedupe → debounce → lease         │
 │  customer identity (one person, many apps)         │
 │  context: history + knowledge (RAG) + skills       │
 │  tools: your REST APIs, MCP servers, Composio      │
 │  guardrails in / out → handoff to operator inbox   │
 └───────────────────────┬────────────────────────────┘
                         │  stream
                   any LLM (OpenRouter)
Enter fullscreen mode Exit fullscreen mode

Below are the five problems that made us build a harness instead of a chatbot, and the real code for each, so you can take the ideas even if you never run Ecbot.

The stack, briefly

  • apps/api: NestJS 11, Postgres (MikroORM), Redis and BullMQ. Owns channels, conversations, knowledge, tools and the operator API.
  • apps/ai: Python 3.12, FastAPI and LangChain. Runs generation, RAG and guardrails. The API streams from it over HTTP.
  • apps/app: the operator inbox and dashboard (React 19, Vite, TanStack Query).
  • apps/edge: an optional Cloudflare Worker that receives platform webhooks. You can skip it and point webhooks at the API.

Models go through OpenRouter, so you bring one key and pick the model per agent. Embeddings live in Postgres with pgvector, next to the rest of the data.

Problem 1: the platform retries, and you restart

Messenger and Zalo expect a fast 200 on their webhook. If you're slow, they retry. If you 200 first and process in memory, a deploy halfway through a reply loses the message for good.

Ecbot writes every incoming message to a durable BullMQ job (Redis with AOF on) before it acknowledges the webhook. A worker drains the queue. If the process dies mid-reply, the job is still there and runs again.

Retries mean duplicates, so there is a single dedupe point at the top of the processing pipeline. It's a Redis SET NX on ${platform}:${externalMessageId} with a 24-hour TTL. We put it there instead of relying on the BullMQ job id because messages can arrive by two paths (straight to the API, or forwarded by the edge Worker), and a single check covers both.

For outages longer than a platform's retry window, adapters can implement an optional reconcile(account, lookback) method that backfills missed messages on a schedule.

Problem 2: people type in bursts

Chat users don't write paragraphs. They send four short messages in a row. Ecbot collapses messages that arrive inside a 3-second window into one debounce burst and replies once, to all of them, while the platform shows a typing indicator.

That solves most bursts. It doesn't cover the customer who sends a fourth message while the model is already streaming a reply to the first three.

Problem 3: the reply is already stale when it's ready

For that case each conversation has a generation lease: a counter in Redis. Every inbound message increments it. A generation records the value when it starts and checks it again while streaming and right before it sends. If the counter moved, a newer message exists, so the generation aborts and nothing goes out. The next turn answers with full context.

// GenerationLeaseService (simplified)
async bump(conversationId: string): Promise<number> {
  const epoch = await redis.incr(`gen-lease:${conversationId}`);
  await redis.expire(`gen-lease:${conversationId}`, 60 * 60);
  return epoch;
}

async isCurrent(conversationId: string, epoch: number): Promise<boolean> {
  return (await this.current(conversationId)) === epoch;
}
Enter fullscreen mode Exit fullscreen mode
// ReplyGenerationService (simplified)
const ac = new AbortController();
const stream = await ai.streamChat(turnContext, ac.signal);

await streaming.deliver({
  stream,
  abort: ac,
  isCurrent: () => lease.isCurrent(conversationId, myEpoch),
});
Enter fullscreen mode Exit fullscreen mode

Aborting the AbortController closes the HTTP stream between the API and the Python service, so the model stops generating and you stop paying for tokens nobody will read. The agent keeps no state between turns. Each turn rebuilds its context from the database, which is what makes throwing a generation away safe.

Problem 4: one person, three apps

A customer asks about a dress on Facebook, orders on WhatsApp, then asks on Telegram where it is. To most bots, that's three strangers.

Ecbot separates the Customer (the human, with name, phone, notes and tags) from the ContactPoint (one platform identity, unique on (workspace, platform, externalSenderId)). One customer has many contact points. Replies route through the contact point the message came from, while the profile, tags and history belong to the customer.

We don't merge customers automatically. When two profiles share a phone number or email, Ecbot creates a merge suggestion and an operator confirms it. The agent never sees pending suggestions, only confirmed customers. A wrong auto-merge would leak one person's order history to another, and we didn't think a heuristic should make that call.

Problem 5: knowing when to stop

An agent that sells also has to know when to hand the chat to a person. Ecbot's escalation fires when:

  • the customer asks for a human (handoff keyword),
  • the agent's confidence drops below a threshold you set,
  • an input guardrail flags the message as prompt injection or abuse,
  • the agent applies a customer tag you marked triggersHandoff, such as "Angry".

Escalation turns the bot off for that conversation and notifies operators, who reply from the same inbox. When they resolve the conversation, the bot comes back on with a clean slate.

Tools and skills: what the harness lets the model do

The agent acts through tools, which come in two kinds: HTTP tools that wrap your own REST endpoints (JSON schema in, description for the model), and MCP tools, either from the Composio marketplace or any MCP server URL you give it. Every call is stored as a ToolInvocation with args, result, latency and errors, so an operator can see what the agent did in a conversation. Platform credentials are encrypted at rest (AES-256-GCM) and secrets are kept out of the model's context entirely.

Skills are markdown playbooks: "when a customer asks for a refund: apologise, ask for the order code, check the order…". The agent only sees each skill's name and description. When a conversation needs one, it calls load_skill(slug) and pulls the full text into context. That keeps the system prompt short even when a shop has twenty playbooks.

Adding a channel

Each channel is a self-contained adapter in apps/api/src/modules/platform/adapters/. The base class is small:

export abstract class PlatformAdapter {
  abstract readonly type: ENUM_ACCOUNT_TYPE;
  abstract readonly capabilities: AdapterCapabilities;

  abstract verifyChallenge(req: Request): Response | null;
  abstract verifySignature(rawBody: string, headers: Headers | Record<string, string>): boolean;
  abstract parse(rawBody: string): PlatformWebhookEvent[];
  abstract fetchSenderProfile(account: AccountEntity, senderId: string): Promise<PlatformUserProfile>;
  reconcile?(account: AccountEntity, lookback: Date): Promise<PlatformWebhookEvent[]>;
  // ...plus doSend() for outbound
}
Enter fullscreen mode Exit fullscreen mode

Dedupe, debounce, the lease, guardrails and handoff sit above the adapter, so a new channel gets all of them for free. Instagram (#22), TikTok Shop (#23), Shopee (#24) and LINE (#80) are open and specced, labelled good first issue.

Running it

You need Node 20+, pnpm 9, Docker, and uv with Python 3.12+.

git clone https://github.com/blackmouse572/ecbot.git
cd ecbot
pnpm install
pnpm setup:local                 # writes the .env files; asks for OPENROUTER_API_KEY
docker compose up -d             # Postgres + pgvector, Redis, JWKS, Bull Board
pnpm --filter api migration:up
pnpm dev:app                     # api + operator UI
pnpm --filter ai dev             # AI service, second shell
Enter fullscreen mode Exit fullscreen mode

The UI is on localhost:5173 and Swagger on localhost:8080/docs.

The honest catch: Facebook and Zalo need your own platform app with messaging permissions approved, and that review can take a while. Telegram, the website widget and the REST channel need no approval, so start with a Telegram bot. You can talk to your agent within a few minutes of the stack coming up.

If you're building something similar, I'd like to hear how you handle stale replies. A per-conversation counter was the simplest thing that worked for us, and I'm sure it isn't the only way.

The repo is at github.com/blackmouse572/ecbot. Clone it, connect a Telegram bot, and tell us in the comments (or on Discord) where it broke.

Top comments (0)