Most broken n8n workflows aren't broken by n8n. I read 398 public reports from n8n's community forum and Reddit where someone's automation stopped working, and in 317 of them the thread found the cause. Only about 1 in 10 was a bug in n8n. The rest were a setting, another app's rules, or the workflow itself.
That sounds like good news, but it's also why these problems drag on. n8n can't warn you about a setting it doesn't know is wrong, and many of these failures make no noise at all: a schedule that never fires, a webhook pointing at an address nobody can reach, a Google sign-in that expires every seven days.
The templates most people start from don't help much here. I ran the 200 most-viewed workflows on n8n's template library through my n8n workflow checker: of their 368 HTTP Request steps, 333 have no retry and no error handling, so one failed call ends the whole run.
Below are the causes in the order I'd check them, each with the fix from n8n's own documentation, and what I found in those templates at the end.
What 398 broken n8n workflows had in common
The reports run from January 2025 to September 2026: 340 threads from n8n's community forum and 58 from Reddit's r/n8n. I sorted each one by the cause the thread gave, from the poster, an accepted answer, an n8n staff member or another reply, or by the main symptom when nobody found one.

Self-hosting and webhooks lead. Of the reports with a stated cause, about 1 in 10 was a bug in n8n.
Two things stood out. Self-hosting problems were the biggest group, and most of them weren't about n8n's code at all: they were about memory, Docker and the proxy in front of it. And the group that hurts most, workflows that work when you click the button but never run on their own, was also the least solved: 16 of those 34 threads never found a cause.
It works when you click it, but never runs on its own
Start here, because the most common reason changed recently. Since n8n 2.0 in December 2025, the old Active toggle is gone. Your edits save automatically as a draft, and nothing goes live until you click Publish. Live runs keep using the last published version, so a workflow can show "Published, has changes": it's running, but not the version you fixed. A publish can also partly fail, with some triggers switched on and others not.
Then check what starts it, and when:
- A real trigger. The Manual Trigger only runs when you click. A workflow needs a Schedule, Webhook or app trigger to run by itself, and since 2.0 n8n won't publish one without it.
- No pinned test data. Data you pinned while testing is ignored in live runs, so a workflow that only worked because of pinned data behaves differently once published.
- The right timezone. The Schedule Trigger uses the workflow's timezone, or the instance's if none is set, and self-hosted n8n defaults to New York time. Set it under Workflow settings, or with GENERIC_TIMEZONE on your server. Changing the schedule after publishing does nothing until you publish again.
- Runs missed while n8n was down are gone. By default, schedules live in memory, so a run whose time passes during a restart or crash is skipped, not caught up. A durable scheduler that can catch up arrived in n8n 2.36, but it's off unless you turn it on.
- Email triggers that check for new mail. Several Gmail threads came down to the "only unread emails" filter, or to checking for new mail every minute. Polling less often, every five minutes, fixed one of them.
The webhook worked in testing, then stopped
Every Webhook node has two addresses. The test URL, with /webhook-test/ in it, only listens for 120 seconds after you click Listen for test event. The production URL, with /webhook/, only works while the workflow is published, and its runs appear in the Executions tab rather than on the canvas. A webhook that worked in testing and then went quiet is very often still pointing at the test address.
On a self-hosted server, the bigger problem is the address itself. In 22 webhook reports, the cause was n8n not knowing its own public address, so it handed out links like http://localhost:5678 that nothing outside your server can reach. n8n builds webhook addresses from N8N_HOST, which defaults to localhost. Behind a reverse proxy, n8n's proxy guide says to set N8N_WEBHOOK_URL to your public https address and N8N_PROXY_HOPS to 1. Most guides still say WEBHOOK_URL: that name is deprecated from version 2.35, though it still works with a warning.
Setting the variable isn't always enough. In several threads it was written in an .env file that never reached the program: a Docker Compose file with no env_file line, or n8n started as a system service without its environment file. Telegram and WhatsApp were involved in about 30 of the webhook reports: both only accept a webhook address they can reach over https, and both allow one webhook per bot or app, so a second workflow quietly takes over from the first. And on n8n Cloud, a webhook that takes longer than 100 seconds to answer fails with a 524 error, so long jobs should answer right away and do the work afterwards.
Self-hosting: memory, Docker and "Connection lost"
Self-hosting was the biggest group, with 70 reports. Running n8n yourself is cheap, but you become the person who looks after the server, and these are the problems that came up again and again:
- "Connection lost" in the editor. It was the most common error message in the whole set, and in about 10 threads the cause was a proxy like NGINX not passing WebSocket connections through. If it appears in the middle of a run instead, the instance may have crashed.
- Running out of memory. n8n puts no cap on how much data a step loads, so one big spreadsheet or file can crash it, often showing up as "Connection lost" or a 503 rather than a clear memory error. n8n suggests working in smaller batches, such as 200 rows per run instead of 10,000. Testing by hand uses more memory than live runs, because n8n keeps a copy for the editor.
- Data gone after an update. The official Docker setup keeps everything in a volume at /home/node/.n8n. In one thread the volume pointed at /root/.n8n instead, so every update started from a blank setup screen.
- "Credentials could not be decrypted." n8n encrypts your saved logins with a key it creates on first launch and stores in that same folder. Lose the folder, or start n8n with a different key, and every saved credential becomes unreadable. Even if your data lives in Postgres, keep that volume and back up the key.
- Runs stuck in "queued". In queue mode, workers do the actual work. In the most-viewed thread of this group, the extra containers had been started without their worker command, so jobs waited forever. Workers also need the same encryption key as the main instance.
Google sign-in that keeps expiring
Credentials were 45 reports, and most involved Google. Two causes came up again and again. The first, in 15 threads, was a redirect address mismatch, which Google reports as redirect_uri_mismatch: it only accepts the exact address you registered, so copy the OAuth Redirect URL from n8n into Google Cloud Console character for character, including https and the port. On a self-hosted server that address often said localhost, for the same reason webhooks do.
The second is a Google app left in Testing mode. n8n's Google docs note that apps in Testing with External users lose their sign-in after seven days, so a workflow runs for a week and then stops. Publishing the app in Google Cloud Console ends that. On n8n Cloud, many Google nodes offer a managed "Sign in with Google" that avoids building your own Google app at all.
It broke after an n8n update
Forty-three reports were workflows that broke after updating, and about 23 of them involve the 2.x versions that started in December 2025. n8n called 2.0 a hardening release, and it changed defaults on purpose:
- Code nodes can no longer read environment variables.
- The Execute Command node is switched off, and file nodes can only reach one folder.
- OAuth callback addresses now need authentication, and the old Python Code node was removed.
- Activate became Publish, so saving no longer changes what's live.
The rest were plain regressions, and the fix that worked most often was moving to a newer Stable release (about 13 threads), with three people rolling back instead. Two n8n Cloud users were on a Beta version without knowing it, which broke every Code node they had. n8n recommends the Stable track for anything important, pinning an exact version, updating at least monthly so the jumps stay small, and a full backup first. Since 2.0 there's also a Migration Report under Settings that lists what will break before you upgrade.
If you self-host, plan for the next one now. n8n says version 3.0 is due in October 2026: it supports only Docker for self-hosting, cuts the Code node time limit from five minutes to one, and removes the old Cron, Interval and Function nodes along with the first version of the AI Agent node.
The AI Agent doesn't use its tools
Fifty reports were about the AI Agent node, and the biggest group was an agent that answered without using its tools, or called them with empty values. n8n's agent docs and the accepted answers point to the same fixes:
- Describe each tool. The description is how the agent decides when to use it, so say plainly what it does and when to call it.
- Let the model fill the inputs. When a tool always got empty input, the fix an n8n staff member gave was $fromAI(), which lets the model supply that value.
- Use a model that calls tools well. Small local models sometimes write the tool call as plain text instead of making it. A larger model fixed it in one thread, and n8n 2.40 added a Force Tool Call on First Iteration option for exactly this.
- Look at what it actually did. When a tool fails, the error goes back to the agent as text and can disappear inside its friendly answer. Turn on Return Intermediate Steps to see every tool call.
- Store memory somewhere real. Simple Memory doesn't last between sessions and doesn't work in queue mode. For a chat that should remember people, use a database-backed memory, and give it a session ID, which a Telegram trigger doesn't provide by itself.
An agent that answers confidently without checking its tools is the same problem I wrote about in why AI chatbots give wrong answers: the fix is usually in what the model is given, not in the model.
Outside APIs, rate limits and n8n Cloud limits
A 429 error means the other service is telling you to slow down, and n8n shows it as "The service is receiving too many requests from you." The HTTP Request docs give two fixes: turn on Retry On Fail in the node's settings, or send requests in batches, such as one per second. Not every 429 is about speed, though. In one thread it was an OpenAI account out of credits, and in another a provider cut the limit when the account balance dropped.
On n8n Cloud, the plan limits are the other surprise. When you hit your monthly execution limit, n8n's help center says new runs fail at once with "Execution limit reached". They aren't queued and they aren't re-run later, even though your triggers keep firing. The limit resets on the first of each calendar month, not your billing date. Memory is small too: the Starter plan gets 320 MiB, and n8n itself uses about 180 of that, so a few large files or several AI steps at once can crash the workspace. About 8 Cloud users in the reports hit exactly that.
What I found in the 200 most-viewed n8n templates
Most people don't start from a blank canvas. They import a template, so I took the 200 most-viewed workflows on n8n's public template library, with 11.7 million views between them, and ran each one through my free n8n workflow checker, which reads the same workflow file you import into n8n.

Templates show the happy path. Retries and alerts are the part you have to add yourself.
To be fair to the authors, templates are starting points, and an error alert usually lives in a separate workflow, so it wouldn't be part of a template anyway. But that's the point: when you import one, nothing in it tells you what will happen when a step fails. Of 368 HTTP Request steps, only 12 retry, and 333 have neither a retry nor any error handling. Thirty-seven templates only run when you click, and 13 still carry pinned test data that live runs ignore.
Three templates had an expression pointing at a step that had been renamed. In one, the step was now called "Appointment Scheduling Agent1" while the expression still asked for "Appointment Scheduling Agent", so that step fails as soon as it runs. The good news: none of the 200 contained a real API key. Keys live in credentials, which never travel with a workflow, so after importing you connect your own.
What keeps an n8n workflow running
Almost every cause in this post is quiet. The workflow doesn't crash in front of you. It stops firing, or fails at 3am, and the run sits in the Executions list where nobody looks.
I learned that on one of my own products. It ran a daily job that published broken content about nine days a month, and nobody was told, because it never threw an error. What fixed it wasn't a smarter job: it was checking the result before publishing, and sending an alert when it failed.
In n8n, that means three things. Create one error workflow that starts with an Error Trigger and emails or messages you, and select it in each workflow's settings; one error workflow can serve all of them. Turn on Retry On Fail for every step that calls another service. And where a run can fail without an error, like an empty result or a half-written record, add a check that stops the run with a Stop And Error node, so it reaches your alert too. Since version 2.5, n8n emails the owner once on the first live failure if no error workflow is set, but only once, so don't rely on it.
If you'd like a quick look first, paste your workflow into the n8n workflow checker. It runs in your browser, so the workflow never leaves your computer. There are more guides on automations in Automate and connect your tools.
And if you'd rather hand it over, I take these on as client projects: send me the workflow and where it stops.
Common questions
Why is my n8n workflow not running automatically?
Since n8n 2.0, edits are saved as drafts and only go live when you publish, so check for "Published, has changes". The workflow also needs a real trigger, not the Manual Trigger. If it runs at the wrong hour, set the timezone, and remember that scheduled runs missed while n8n was down are skipped by default.
Why does my n8n webhook work in test but not in production?
The test URL (/webhook-test/) only listens for 120 seconds after you click Listen for test event. The production URL (/webhook/) only works while the workflow is published, and its runs show up in the Executions tab, not on the canvas. Make sure the other app calls the production URL.
Why does my n8n webhook URL show localhost?
Self-hosted n8n builds webhook addresses from N8N_HOST, which defaults to localhost. Behind a reverse proxy, set N8N_WEBHOOK_URL (WEBHOOK_URL on versions before 2.35) to your public https address and N8N_PROXY_HOPS to 1, then restart. In Docker, check that the variable actually reaches the container.
What does "Connection lost" mean in n8n?
The editor lost its live connection to the server. Behind a reverse proxy like NGINX it usually means WebSocket upgrade headers are not being forwarded. If it appears in the middle of a run, the instance may have crashed, often from running out of memory.
Why does my n8n AI Agent not call its tools?
Give each tool a clear description of when to use it, and let the model fill tool parameters with $fromAI(). Small models sometimes write the tool call as plain text instead of making it; try a stronger model or the Force Tool Call on First Iteration option added in n8n 2.40. Turn on Return Intermediate Steps to see what it actually called.
Why do my n8n Google credentials keep expiring?
If your Google Cloud app is in Testing mode with External users, Google expires the sign-in after seven days. Publish the app in Google Cloud Console. A redirect_uri_mismatch error is different: copy the OAuth Redirect URL from n8n into Google exactly, including https and the port.
What happens when n8n Cloud hits its execution limit?
According to n8n's help center, new executions fail immediately with "Execution limit reached", they are not queued, and they do not re-run later. Triggers keep firing, so you lose those runs. The limit resets at the start of each calendar month, not on your billing date.
Why did n8n break after I updated it?
Version 2.0 (December 2025) changed defaults on purpose: Code nodes can no longer read environment variables, Execute Command is disabled, and file nodes can only reach one folder. Check the Migration Report under Settings before upgrading, back up first, and pin a Stable version rather than updating to the newest Beta.
This post first appeared on tusharbhendarkar.com. I fix and build automations in n8n, Zapier and Make, and I build the free n8n workflow checker. If something here is broken for you, email me at tusharbhendarkar44@gmail.com.
Top comments (0)