I thought my n8n LLM bug was OpenAI and it turned out to be 5 boring workflow failures
I had one of those bugs that makes you immediately blame the model provider.
The workflow was still spinning.
The user had no response.
There was no obvious error.
And the AI step was the last interesting thing in the execution.
So I did the obvious bad move: I blamed OpenAI, retried the run, and piled more retries onto a workflow that was already jammed.
The thing that broke the illusion was simple:
The provider logs showed the model had already responded.
n8n was the thing still hanging.
That was the moment I had to admit the bug was in my orchestration layer, not in the model.
If you run AI agents in n8n, Make, Zapier, OpenClaw, or your own queue-based workflow stack, this matters a lot:
A lot of “the LLM stopped responding” bugs are not model outages. They’re worker backlog, Redis payload issues, missing timeouts, BullMQ stalls, or agent loops.
I lost hours to this once. You probably don’t need to.
The fast sanity check that changed how I debug this
Before touching prompts, retries, or provider dashboards, compare two timestamps:
- When the provider says the request completed
- When your workflow says the execution completed
If the provider finished at 12:01:04 and your client is still hanging at 12:01:45, the model is not your first suspect anymore.
That’s an orchestration problem until proven otherwise.
Why n8n makes this look like a model outage
When you’re tired and staring at a stuck execution, it feels like the model call is frozen.
But in queue mode, the path is longer than people remember:
- the client stays connected to the main or webhook instance
- the execution may run on a worker
- the result has to come back through Redis
- the workflow still has to serialize, pass data around, and finish cleanly
Any slowdown in that path can look exactly like “OpenAI stopped responding.”
That’s why this bug is so annoying. The model can be done, but the user still sees a spinner.
The 5 workflow bugs I now check before blaming the model
These are the five issues I found, in the order I’d check them now.
1) n8n was waiting on a worker, not on the LLM
What it looked like
The execution sat long enough that it felt like a dead model call.
Why it fooled me
From the outside, these look identical:
- model is slow
- worker is backed up
If all you have is a hanging HTTP request, your brain picks the AI step every time.
What I checked
I compared provider-side completion time with workflow completion time, then checked queue delay and worker throughput.
What proved it
The provider had already returned the response.
The worker path had not finished processing it.
What I’d do now
Check worker pressure first.
# Example things to inspect in your deployment
# Docker
docker ps
docker logs n8n-worker --tail 200
# Kubernetes
kubectl get pods
kubectl logs deployment/n8n-worker --tail=200
If you’re using queue mode, look for:
- too few workers
- workers stuck on other jobs
- queue backlog spikes
- long execution times unrelated to provider latency
If upstream finished and your workflow didn’t, the queue is guilty until proven innocent.
2) Redis was carrying a huge response through the queue
What it looked like
Small runs worked.
Large runs felt flaky or frozen.
Why it fooled me
Big outputs make people suspect token generation or streaming first.
That’s reasonable, but wrong surprisingly often.
What I checked
I compared small-output runs vs large-output runs and looked at where the delay appeared.
What proved it
The slowdown tracked payload size, not model latency.
That means the bottleneck was response handoff, serialization, queue transport, or downstream processing.
Practical fix
Stop shipping giant blobs through your workflow if you don’t need to.
Bad pattern:
{
"messages": [...],
"full_context": "very large text...",
"tool_results": [...],
"raw_model_output": "huge output..."
}
Better pattern:
{
"response_id": "resp_123",
"summary": "short result for next step",
"storage_key": "s3://bucket/run-123/output.json"
}
If your agent produces large intermediate artifacts:
- store them in object storage
- pass references, not full payloads
- trim chat history aggressively
- avoid unnecessary JSON expansion between nodes
A giant JSON blob moving through Redis is one of the most boring ways to fake an LLM outage.
3) I had no real timeout, so the workflow could wait forever
What it looked like
Requests didn’t fail cleanly.
They just kept hanging.
Why it fooled me
No error means people assume “the model is still thinking.”
That’s a terrible default assumption.
What I checked
I audited timeout settings in three places:
- the HTTP client
- the n8n workflow or AI node
- any proxy or webhook layer in front
What proved it
There was no meaningful upper bound forcing the run to fail fast and surface the actual bottleneck.
Practical fix
Set explicit timeouts everywhere.
Example Node.js request with a timeout:
const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 45000);
try {
const res = await fetch("https://api.openai-compatible-endpoint.com/v1/chat/completions", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": `Bearer ${process.env.API_KEY}`,
},
body: JSON.stringify({
model: "gpt-5.4",
messages: [
{ role: "user", content: "Summarize this ticket" }
]
}),
signal: controller.signal,
});
const data = await res.json();
console.log(data);
} finally {
clearTimeout(timeout);
}
And in workflow systems, I’d strongly recommend:
- per-step timeouts
- execution-level timeouts
- retry limits
- dead-letter handling for repeated failures
A timeout is not just for protection. It’s a debugging tool.
4) BullMQ stalled because the worker blocked the event loop
What it looked like
Jobs were retried or marked stalled even though the code looked fine at first glance.
Why it fooled me
The agent appeared to stop mid-task, and the LLM call was the easiest thing to blame.
What I checked
I looked for CPU-heavy work inside the same Node.js worker process handling queue jobs.
Typical offenders:
- giant JSON parsing
- expensive regex work
- PDF extraction
- big Markdown transformations
- synchronous loops over large payloads
What proved it
The worker blocked the event loop long enough for BullMQ to treat the job as stalled.
If you block the event loop for 30+ seconds, that’s not an OpenAI problem.
That’s your worker design.
Minimal example of the kind of thing that hurts
// Bad: CPU-heavy work in the queue worker process
worker.process(async (job) => {
const huge = JSON.parse(job.data.bigPayload);
// pretend this is expensive
for (let i = 0; i < 500000000; i++) {
// blocking the event loop
}
return huge;
});
Better options
- move CPU-heavy work to a separate service
- chunk large transforms
- use worker threads for CPU-bound work
- keep queue workers focused on orchestration, not heavy processing
Useful checks:
# Watch CPU and memory
htop
# If containerized
docker stats
If your queue worker is doing too much non-I/O work, your “LLM outage” may just be a blocked Node process.
5) LangGraph was looping until recursion limits kicked in
What it looked like
The agent kept “working” but never gave a useful final answer.
Why it fooled me
From the user’s perspective, a loop and a hang feel the same.
No answer is no answer.
What I checked
I traced graph transitions and counted repeated states.
What proved it
The graph was revisiting the same path until recursion limits stopped it.
What I’d do now
Instrument state transitions aggressively.
Pseudo-example:
visited = []
def log_state(state_name, state):
visited.append(state_name)
print({
"state": state_name,
"visited_count": len(visited),
"recent_path": visited[-5:]
})
And add hard guards:
- max iterations
- repeated-state detection
- explicit terminal states
- tool-call limits
If your agent can re-enter the same state forever, it’s not autonomous. It’s a polite infinite loop.
My current debugging order for “LLM stopped responding”
This is the checklist I wish I had used first.
1. Check provider status page and request logs
2. Confirm whether the provider actually completed the request
3. Compare provider completion timestamp with n8n/BullMQ/LangGraph timing
4. Check queue backlog and worker health
5. Inspect payload size and response handoff path
6. Verify timeouts at client, workflow, and infrastructure layers
7. Look for event loop blocking or CPU-heavy worker tasks
8. Trace agent loops and recursion behavior
9. Only then blame the model
That order is less dramatic than posting “OpenAI is down again,” but it gets to the truth faster.
A practical example: compare provider logs to workflow timing
If you’re using an OpenAI-compatible endpoint, your app code probably already gives you enough hooks to do this.
Example:
const startedAt = Date.now();
const response = await client.chat.completions.create({
model: "gpt-5.4",
messages: [{ role: "user", content: "Classify this support ticket" }],
});
console.log({
providerReturnedAtMs: Date.now() - startedAt,
responseId: response.id,
});
Then compare that with:
- n8n execution duration
- worker completion time
- webhook response time
- downstream node timing
If provider time is 4 seconds and user-visible completion is 45 seconds, the model is not your biggest problem.
Why this gets expensive fast in agent workflows
This is the part people underestimate.
False outage diagnosis is expensive.
Every blind retry can:
- clog your queue more
- duplicate downstream work
- increase latency for unrelated jobs
- burn more model usage
- make the original bug harder to see
That last part matters a lot if you’re paying per token and debugging high-volume automations. Retries during incident response can turn into their own bill.
That’s one reason I like using OpenAI-compatible endpoints that are easy to swap in and out. If you can test the same workflow against another compatible provider without rewriting your app, it becomes much easier to separate provider issues from orchestration issues.
And if your automation stack runs constantly, flat-rate compute is honestly a better fit than per-token pricing. During debugging, retries and test runs stop feeling like financial damage.
That’s the appeal of something like Standard Compute:
- OpenAI-compatible API
- works with existing SDKs and HTTP clients
- useful for n8n, Make, Zapier, OpenClaw, and custom agent workflows
- flat monthly pricing instead of per-token billing
For teams running agents all day, that’s not just a pricing detail. It changes how safely you can test, retry, and debug.
The actual lesson
The lesson is not that OpenAI, Anthropic, or Grok never fail.
They do.
The lesson is that workflow bugs impersonate model outages extremely well.
So now my default assumption is:
If the provider already answered, the bug is probably mine.
That assumption has saved me more time than any prompt tweak ever has.
If you’re debugging hanging LLM workflows in n8n or another automation stack, start with the response path, not the model.
That’s usually where the real problem is hiding.
Top comments (0)