I realized our setup was fake-stable the first time a background agent stayed “running” while doing absolutely nothing.
No crash.
No stack trace.
Docker said the container was up. The process existed. CPU was low. Memory looked fine.
Meanwhile, webhook jobs were stacking up like a sink full of dishes.
That was the week I stopped trusting "container is running" as a useful definition of healthy.
If you’re moving from a laptop-and-good-vibes setup to a real headless AI server for always-on agents, this failure mode shows up fast.
Usually the stack looks something like this:
- n8n or a custom webhook service
- Stripe / GitHub / Discord / Slack webhooks
- cron jobs
- long-running GPT-5 or Claude tasks
- maybe OpenClaw or another gateway in front
- a few Docker containers and a prayer
Everything looks fine until one worker hangs without technically dying.
The fix was not a new framework.
It was three boring things:
- separate ingress from execution
- make jobs replay-safe
- put workers behind one OpenAI-compatible endpoint
That’s the whole post.
The architecture mistake that caused most of the pain
The biggest mistake was letting webhook intake, cron scheduling, and LLM-heavy execution all happen in the same process.
That works on a laptop.
It gets weird in production.
A slow Claude call delays a Stripe webhook.
A giant PDF parse blocks the event loop.
A CPU spike from OCR makes your Discord bot miss heartbeats.
One stuck worker makes the whole service look haunted.
The clean split is the same pattern n8n uses in queue mode:
- one main instance handles timers and webhooks
- jobs get pushed into Redis
- separate workers pull jobs and execute them
- shared state goes into a real database
That separation matters more than most teams think.
Ingress should stay lightweight.
Execution should be allowed to be slow, retryable, and disposable.
A minimal version of the pattern
[Stripe/GitHub/cron/webhooks]
|
v
[main ingress service]
|
v
[Redis queue]
|
-------------------
| | |
| --- | --- |
v v v
[worker] [worker] [worker]
| | |
| --- | --- |
---------+--------
|
v
[OpenAI-compatible endpoint]
|
v
[GPT-5.4 / Claude Opus 4.6 / Grok 4.20]
If you’re using n8n, this is basically why queue mode exists.
Also: once you do this, local-only state starts breaking things.
That means:
- SQLite stops being a serious option for shared execution state
- local filesystem storage becomes a problem for binary data
- PostgreSQL and S3-compatible storage start looking very reasonable
This is the boring infrastructure tax for not losing jobs.
A queue does not save you from bad job design
A queue gives you another chance.
That’s it.
Whether that second chance helps or hurts depends on idempotency.
If a worker does this:
- charge a card
- send an email
- write to Salesforce
- crash before ack
...then the queue may redeliver the job.
If the task is not replay-safe, you just built a duplicate side-effect machine.
This is not theoretical. It’s how these systems behave.
With Celery, for example, tasks can be redelivered if the worker dies before acknowledgment. Features like acks_late help, but they do not magically make your side effects safe.
The real fix is application-level:
- store idempotency keys
- record processing state in a database
- make external side effects safe to retry
Stripe will absolutely teach you this lesson
Stripe’s webhook model is basically a giant sign that says: expect duplicates, expect retries, design accordingly.
If your endpoint is down, Stripe retries undelivered events for up to 3 days.
You can also recover manually by listing undelivered events.
Example:
curl -G https://api.stripe.com/v1/events \
-u "sk_test_...:" \
-d ending_before=evt_001 \
-d "types[]=payment_intent.succeeded" \
-d "types[]=payment_intent.payment_failed" \
-d delivery_success=false
The important part is not the curl command.
The important part is that your system has to know whether it already processed evt_001.
A simple pattern is a table like this:
CREATE TABLE processed_events (
event_id TEXT PRIMARY KEY,
event_type TEXT NOT NULL,
processed_at TIMESTAMP NOT NULL DEFAULT NOW()
);
Then your worker logic becomes:
async function handleStripeEvent(event) {
const alreadyProcessed = await db.processed_events.findUnique({
where: { event_id: event.id }
});
if (alreadyProcessed) {
return { ok: true, duplicate: true };
}
await doSideEffects(event);
await db.processed_events.create({
data: {
event_id: event.id,
event_type: event.type,
}
});
return { ok: true };
}
That pattern is not glamorous.
It is also the difference between “retries are safe” and “why did we send 4 receipts?”
Docker restart policies are useful, but they did not solve the real problem
Yes, use restart policies.
They help with crashes.
docker run -d --name worker --restart unless-stopped my-worker-image
But restart policies only help when the process exits.
My actual problem was worse:
- the process was alive
- the container was alive
- the worker was not making progress
That’s a different class of failure.
Docker does not know your Node.js event loop is wedged.
Docker does not know your Python worker is stuck in a bad call.
Docker does not know your queue heartbeat stopped.
That’s why “container is up” is not enough.
BullMQ was the first tool that made this obvious
BullMQ has a much more useful idea of health than Docker does.
It marks jobs as stalled when an active worker stops renewing its lock.
By default, that check is around 30 seconds.
That is exactly the kind of signal I needed.
Because a worker that is still running but hasn’t renewed a lock in 30 seconds is not healthy in any meaningful sense.
A basic BullMQ worker looks like this:
import { Queue, Worker } from 'bullmq';
import IORedis from 'ioredis';
const connection = new IORedis(process.env.REDIS_URL);
const queue = new Queue('jobs', { connection });
const worker = new Worker(
'jobs',
async job => {
console.log(`Processing job ${job.id}`);
// your LLM / automation work here
return await runTask(job.data);
},
{ connection }
);
worker.on('completed', job => {
console.log(`Job ${job.id} completed`);
});
worker.on('failed', (job, err) => {
console.error(`Job ${job?.id} failed`, err);
});
And a retry config with jitter should be your default, not an afterthought:
await queue.add('test-retry', { foo: 'bar' }, {
attempts: 8,
backoff: {
type: 'fixed',
delay: 1000,
jitter: 0.5,
},
});
That jitter: 0.5 matters a lot.
Without jitter, a worker fleet tends to fail together and retry together.
That’s how a small outage turns into a retry storm.
CPU-heavy jobs can make a healthy-looking worker useless
This one bit me hard.
If you run CPU-heavy work in the same Node.js process that needs to keep queue heartbeats alive, you can create stalls without crashing anything.
Examples:
- OCR
- PDF parsing
- image conversion
- local embedding generation
- giant JSON transforms
The worker process is technically alive.
But queue progress says otherwise.
That’s why sandboxed processors or separate worker processes are worth it.
If a task can block the event loop, isolate it.
What should supervise what?
The best answer I found was: use multiple layers, but give each layer one job.
| Option | What it’s actually good at |
|---|---|
| systemd watchdog + service restart | Detects hung processes via heartbeat instead of only exit codes. Great for VM or bare-metal workers. |
| Docker restart policy | Restarts exited containers. Good default. Weak for hung-but-still-running workers. |
| BullMQ or Celery | Decouples ingress from execution, supports retries and redelivery, and tolerates worker death if jobs are replay-safe. |
My ranking for real-world usefulness:
- queue semantics
- stall detection / watchdogs
- container restart policies
Container restarts save you from crashes.
Queues save you from reality.
The other thing that cleaned this up: one OpenAI-compatible endpoint
This was a much bigger deal than I expected.
Before, different workers were calling different providers directly:
- one service used OpenAI
- another used Anthropic
- another used Groq
- somebody had an OpenRouter integration
- retry logic was duplicated in multiple places
That gets ugly fast.
Especially when every layer retries differently.
OpenAI-style clients already need backoff for 429s and transient failures. If your queue retries jobs and your model client also retries aggressively, you can easily create a retry storm from both sides.
Putting all workers behind one OpenAI-compatible endpoint fixed a lot:
- one auth pattern
- one SDK shape
- one timeout policy
- one retry policy
- one place to route requests between models
- one place to observe failures and rate limits
That architecture is a great fit for AI agents and automations.
It also makes provider swaps much less painful.
Why this matters for teams running always-on agents
If you have:
- n8n workers
- Make or Zapier handoffs
- custom Node/Python agents
- webhook-driven automations
- internal tools firing LLM jobs all day
...then the model layer should not be reimplemented in every service.
A single OpenAI-compatible endpoint gives you a clean boundary.
This is exactly why products like Standard Compute are interesting for this setup.
You point your workers at one endpoint, keep your existing OpenAI-compatible SDKs, and avoid wiring provider-specific behavior into every queue consumer.
The other practical benefit is cost predictability.
When agents run 24/7, per-token billing gets annoying fast. A flat monthly model is a lot easier to reason about when you’re scaling background jobs, retries, and automations across a worker fleet.
For this kind of architecture, that matters more than people admit.
The catch with a single endpoint
A single endpoint also centralizes failure.
That is real.
If your gateway is down or misconfigured, every worker feels it.
So if you do this, you still need:
- health checks
- request timeouts
- circuit breakers where appropriate
- good logs and metrics
- fallback behavior for critical paths
Still, I would take one well-observed gateway over a zoo of half-maintained provider clients every time.
Especially for teams that want interchangeable workers.
When this is overkill
If your setup is tiny, don’t cargo-cult a distributed system.
If you have:
- one internal script
- no external webhooks
- no concurrency
- no long-running jobs
...then a single process with systemd, strict timeouts, and a dead-letter path may be enough.
You do not need Redis, BullMQ, PostgreSQL, S3, and six workers to rename files once an hour.
But once you add:
- Stripe or GitHub webhooks
- concurrent jobs
- long-running GPT-5 / Claude tasks
- multiple team members expecting reliability
...the “simple” setup stops being simple.
It becomes fragile.
And fragile always looks cheap right before it gets expensive.
The baseline I trust now
If I were rebuilding a headless AI server setup today, this is the baseline:
1) Separate ingress from execution
- webhooks and cron hit a lightweight main service
- execution happens in workers pulling from Redis via BullMQ or Celery
2) Make every job replay-safe
- store idempotency keys
- record processing state in a database
- assume duplicates will happen
3) Use supervision that can detect hangs
- Docker restart policies for crashes
- BullMQ stalled-job detection for queue workers
- systemd watchdog where OS-level supervision makes sense
4) Put model access behind one OpenAI-compatible endpoint
- one client shape
- one retry policy
- one place for routing, auth, quotas, and observability
5) Move shared state off local disk
- PostgreSQL instead of SQLite for multi-worker setups
- S3-compatible storage instead of local files when workers share binaries
6) Tune retries once
- queue retries and model-client retries should cooperate
- add jitter so workers do not stampede together
A practical checklist
If you want something actionable, start here:
# 1. run Redis
docker run -d \
--name redis \
--restart unless-stopped \
-p 6379:6379 \
redis
# 2. split your app into:
# - ingress service
# - worker service
# - shared database
# 3. add health checks that measure progress, not just process existence
# 4. make every external side effect idempotent
# 5. point all workers at one OpenAI-compatible endpoint
If you do only those five things, your setup will already be less fragile than a surprising number of “production” agent stacks.
The weirdly comforting part
The most reliable headless AI server setup I’ve used is boring on purpose.
Not the fanciest agent framework.
Not the prettiest architecture diagram.
Not the stack with the most logos in it.
Just:
- queue in the middle
- watchdog on the worker
- one OpenAI-compatible endpoint in front
That setup assumes workers will freeze.
It assumes webhooks will retry.
It assumes jobs will replay.
It assumes somebody will eventually ship CPU-heavy nonsense into a process that was supposed to stay responsive.
That’s why it works.
And if you’re running always-on AI agents, “nothing broke at 3 a.m.” is a much better definition of success than “the container was still up.”
Top comments (0)