DEV Community

Cover image for Our chatbot not replying in groups turned out to be a Slack plumbing bug, not an AI bug
Lars Winstand
Lars Winstand

Posted on Originally published at standardcompute.com

Our chatbot not replying in groups turned out to be a Slack plumbing bug, not an AI bug

If your bot works in DMs but goes weirdly silent in Slack channels, start by blaming your event handling.

Not GPT-5.
Not Claude.
Not your prompt.

We lost a bunch of time learning this the dumb way.

Our bot was great in 1:1 chats. In Slack DMs it could hold context, call tools, summarize results, and generally look like a competent AI teammate.

Then we dropped the exact same bot into a busy channel.

It became a ghost.

Sometimes it replied.
Sometimes it replied twice.
Sometimes the backend finished successfully and Slack showed nothing.
Sometimes it answered in the wrong thread.

That kind of failure is extra annoying because it makes you debug the wrong layer first.

We blamed the model.
We swapped GPT-5 for Claude.
We trimmed prompts.
We argued about whether Llama or Qwen would be more reliable in channels.

None of that mattered.

The real problem was that we treated Slack group chats like DMs, and Slack absolutely does not work that way.

The first bug: DMs and mentions are different event types

This is the first thing I’d check in any Slack bot.

A direct message to your app comes in as a message event with channel_type="im".
A mention in a channel comes in as app_mention.

Those are different event streams with different behavior.

DM example:

{
  "event": {
    "type": "message",
    "channel_type": "im",
    "channel": "D024BE91L",
    "text": "Hello hello can you hear me?"
  }
}
Enter fullscreen mode Exit fullscreen mode

Channel mention example:

{
  "event": {
    "type": "app_mention",
    "channel": "C123ABC456",
    "text": "<@U0LAN0Z89> is it everything a river should be?",
    "ts": "1515449522.000016"
  }
}
Enter fullscreen mode Exit fullscreen mode

That one distinction explains a lot of “works in DMs, fails in channels” bugs.

In DMs, the shape is simple: user talks to bot.

In channels, everything gets more fragile:

  • mention parsing
  • channel membership
  • permissions
  • thread routing
  • rate limits
  • duplicate events

If your code has one generic handleIncomingMessage() path for everything, there’s a decent chance that’s your bug.

Why it only broke in busy channels

Quiet channels hide bad architecture.

If your bot mostly lives in DMs or low-traffic channels, a synchronous handler can look fine for weeks.

Then one incident channel gets busy.
Five people mention the bot.
One tool call takes 8 seconds.
One GitHub request stalls.
One retry comes in.
And suddenly your “AI reliability issue” is obviously just broken event lifecycle handling.

Slack expects a fast HTTP 200 acknowledgment.

The practical rule is simple:

you have about 3 seconds to ack

If you don’t ack quickly, Slack may retry the event, and now you’re in duplicate-processing land.

This is the dangerous flow:

  1. Receive app_mention
  2. Call GPT-5 or Claude
  3. Query your vector DB
  4. Call GitHub
  5. Call Jira
  6. Summarize everything
  7. Finally reply to Slack

That feels natural when you first build it.

It’s also how you build a bot that disappears under load.

The boring fix that actually worked

We moved all slow work out of the request path.

That means:

  • ack immediately
  • enqueue work
  • process in a background worker
  • post the final answer later

Here’s the shape in Slack Bolt for Python:

from slack_bolt import App

app = App(process_before_response=True)

def ack_fast(ack):
    ack("Accepted")

def run_long_process(respond, body, logger):
    user_text = body["event"]["text"]

    # slow work goes here
    # call model
    # call tools
    # build final answer

    respond("Completed")
Enter fullscreen mode Exit fullscreen mode

That’s the idea, but in production I’d usually push this into a queue instead of doing everything inside the listener.

Something more like this:

from fastapi import FastAPI, Request, BackgroundTasks
import os
import json

app = FastAPI()

@app.post("/slack/events")
async def slack_events(req: Request, background_tasks: BackgroundTasks):
    payload = await req.json()

    # Slack URL verification
    if payload.get("type") == "url_verification":
        return {"challenge": payload["challenge"]}

    event = payload.get("event", {})
    event_id = payload.get("event_id")

    # 1. dedupe by event_id
    # 2. enqueue background work
    background_tasks.add_task(process_event, event_id, event)

    # ack immediately
    return {"ok": True}


def process_event(event_id: str, event: dict):
    # idempotency check here
    # model call here
    # tool calls here
    # post reply here
    pass
Enter fullscreen mode Exit fullscreen mode

The important part is not the framework.

The important part is the split:

  • webhook path does validation + dedupe + enqueue
  • worker does the expensive stuff

Reliable Slack bots are event systems first and AI apps second.

The second bug: Slack can drop messages even when your backend succeeded

This one was nastier.

We assumed that if our worker finished and called chat.postMessage, the message would show up.

That is not safe in busy channels.

Slack rate-limits chat.postMessage at roughly 1 message per second per channel.

Per channel.

That means if your bot posts:

  • “thinking...”
  • “checking GitHub...”
  • “checking Datadog...”
  • “found 3 issues...”
  • “here’s the answer”

and three people mention it in the same thread at once, you’ve basically built a rate-limit machine.

The ugly part is that your logs can still look fine while users see silence.

What we changed

We stopped treating outbound messages like free writes.

The new rules were:

  1. Ack immediately
  2. Queue all work
  3. Dedupe retries
  4. Serialize outbound posts per channel
  5. Throttle to about 1 msg/sec/channel
  6. Prefer one solid answer over five cute updates

That improved reliability more than any prompt tweak.

Not because the model got smarter.
Because the plumbing stopped fighting the app.

Thread state is where a lot of bots quietly fall apart

In DMs, state is easy.

In channels, state is thread-shaped.

If you reply without the right thread_ts, your answer lands in the wrong place or becomes useless noise in the channel.

A lot of teams do this:

  • receive a new message
  • call conversations.replies
  • rebuild the whole thread
  • send that to the model
  • repeat forever

That works at first.
It’s also expensive, slow, and increasingly brittle.

A better pattern is to keep your own compact thread state.

Store:

  • channel
  • root thread_ts
  • normalized user turns
  • tool outputs
  • model outputs
  • a compact rolling summary

Use Slack history as recovery, not as your primary memory system.

Here’s the tradeoff:

Approach What actually happens
Rebuild thread from Slack every turn Easy to start, but slower, noisier, and more fragile under load
Maintain your own thread state Slightly more engineering, much more reliable for long-running agents

If you’re running workflows in n8n, Make, Zapier, OpenClaw, or your own worker stack, this matters even more.

Those tools make it easy to ship a bot quickly.
They also make it easy to hide state bugs until traffic spikes.

A practical event pipeline that doesn’t embarrass you

If I were rebuilding this from scratch, I’d do it like this.

1) Split ingress by conversation type

Handle DMs and mentions separately.

def route_event(event: dict):
    event_type = event.get("type")
    channel_type = event.get("channel_type")

    if event_type == "message" and channel_type == "im":
        return "dm"

    if event_type == "app_mention":
        return "channel_mention"

    return "ignore"
Enter fullscreen mode Exit fullscreen mode

Different paths should have different logic for:

  • permissions
  • mention parsing
  • reply routing
  • thread handling

2) Ack first, think later

Anything slow goes to background work.

That includes:

  • GPT-5 calls
  • Claude calls
  • embeddings/search
  • GitHub lookups
  • Jira lookups
  • Notion reads
  • Google Drive scans
  • internal tool calls

Your webhook should be boring.
Boring is good.

3) Treat retries as normal

Slack retries are not edge cases.

Store event_id and make processing idempotent.

Pseudo-code:

def process_event(event_id: str, event: dict):
    if already_processed(event_id):
        return

    mark_processing(event_id)

    try:
        handle_event(event)
        mark_processed(event_id)
    except Exception:
        mark_failed(event_id)
        raise
Enter fullscreen mode Exit fullscreen mode

If you skip this, duplicate replies are just a matter of time.

4) Rate-limit outbound messages per channel

One queue per channel is a sane default.

Pseudo-code:

from collections import defaultdict
from queue import Queue
import time

channel_queues = defaultdict(Queue)

def post_with_throttle(channel_id: str, message: dict):
    q = channel_queues[channel_id]
    q.put(message)

    while not q.empty():
        next_msg = q.get()
        slack_client.chat_postMessage(**next_msg)
        time.sleep(1.0)
Enter fullscreen mode Exit fullscreen mode

In real code you’d use a proper worker, lock, or async queue, but the design point stands:

throttle per channel, not globally

5) Prefer edits over spam

If you want progress updates, update one message instead of posting five new ones.

That usually looks cleaner for users and is friendlier to rate limits.

6) Own your thread state

Slack is a transport layer.
It should not be your source of truth.

Minimal thread record example:

{
  "channel": "C123ABC456",
  "thread_ts": "1712345678.123456",
  "participants": ["U111", "U222"],
  "messages": [
    {"role": "user", "text": "check prod errors"},
    {"role": "assistant", "text": "Looking into GitHub and Datadog"}
  ],
  "summary": "Investigating production error spike after deploy",
  "last_updated": "2026-10-10T12:00:00Z"
}
Enter fullscreen mode Exit fullscreen mode

That gives you a stable context source for long-running agents.

Slack and Discord have the same class of failure

The labels differ, but the architecture lesson is the same.

Slack has:

  • app_mention
  • threads
  • chat.postMessage
  • channel-specific behavior

Discord has:

  • interactions
  • deferred responses
  • followups
  • edit windows

Different APIs, same rule:

respond fast, queue slow work, dedupe retries, and control outbound writes

If you do long-running agent work inline, both platforms will punish you.

Quick local debugging checklist

When a bot works in DMs but not in channels, this is the checklist I’d run in order.

Verify the incoming event type

Log the event shape.

import json

def debug_event(payload):
    print(json.dumps(payload, indent=2))
Enter fullscreen mode Exit fullscreen mode

Check whether you’re actually receiving app_mention and not assuming all traffic is message.

Verify channel membership and scopes

Make sure the app is actually in the channel and has the scopes it needs.

Verify fast ack timing

Log how long your webhook takes before returning 200.

import time

start = time.time()
# validate + enqueue
elapsed = time.time() - start
print(f"ack path took {elapsed:.3f}s")
Enter fullscreen mode Exit fullscreen mode

If that number is drifting upward, you’re moving work back into the request path.

Verify dedupe

Log event_id and retry headers.

If duplicate replies exist, you probably don’t have idempotency under control.

Verify thread routing

Make sure replies use the correct thread_ts.

Verify outbound rate limiting

If the bot is chatty in one channel, assume rate limits are part of the problem until proven otherwise.

Where this connects to AI infra

This bug looked like a model problem because model calls were the most visible slow step.

That’s common in agent systems.

People blame GPT-5, Claude, Grok, or tool-calling reliability when the real issue is orchestration:

  • bad queue design
  • bad retries
  • bad event routing
  • bad rate limiting
  • bad state management

That’s also why predictable AI infrastructure matters.

When you’re building bots, automations, or long-running agents, you want the freedom to offload the slow work without obsessing over every token or every retry path.

That’s the appeal of Standard Compute: you can keep your app OpenAI-compatible, swap in a flat-rate endpoint, and let agents run in the background without per-token cost anxiety creeping into every architecture decision.

Especially if you’re wiring together Slack, GitHub, Jira, Notion, n8n, Make, Zapier, or custom workers, the expensive part is usually not one single prompt. It’s the whole loop.

The annoying truth

Our bot never needed a better personality.
It needed better manners.

The real fix was not prompt engineering.
It was:

  • separate handling for message.im and app_mention
  • ack within 3 seconds
  • background workers for slow tasks
  • dedupe on retries
  • thread-aware replies
  • per-channel outbound throttling
  • state outside Slack

If your chatbot stops replying in groups while behaving perfectly in DMs, start with the boring plumbing.

That’s the stuff that actually fixes production bots.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to