I’ve watched this happen a bunch of times.
The demo is clean:
- webhook hits n8n
- GPT-5 generates a summary
- Claude Opus cleans it up
- result gets posted to Discord or written into Airtable
- everybody says “nice, ship it”
Then someone adds one very normal requirement:
- wait for approval
- retry if HubSpot is down
- pause until the customer replies
- resume tomorrow morning if Stripe finally sends the webhook
That’s the moment your “AI workflow” stops being a sequence of API calls and starts becoming distributed systems work.
And a lot of AI automation demos completely hide that fact.
If your workflow can outlive one HTTP request, you need to design for background execution on purpose.
The first lie your automation tells you
Your first successful automation teaches the wrong lesson.
It teaches you that AI pipelines are just request/response code with a few extra steps:
- receive webhook
- call model
- save output
- return 200
That works until the workflow needs to wait.
A human approval.
A rate-limited API.
A follow-up webhook.
A retry after failure.
A scheduled resume.
Now your code has to survive outside the lifespan of the original request.
That’s where the wheels come off.
Why serverless gets awkward fast
A lot of people discover this through timeouts.
AWS Lambda has a hard 15-minute max execution timeout.
Cloudflare Workers are even stranger for orchestration-heavy jobs because you also care about CPU time, not just wall-clock time.
So if your agent needs to:
- wait 2 hours for approval
- retry a flaky CRM call
- pause until another system emits an event
- fan out work across multiple downstream tools
...your nice little webhook handler is no longer the right abstraction.
At that point you usually bolt on:
- a queue
- a scheduler
- a worker pool
- maybe a workflow engine
That’s how a neat Friday demo becomes a very real Monday incident.
What actually changes when the work moves into the background?
More than most teams expect.
Take n8n.
In the demo, it feels like one app running one flow.
In queue mode, it becomes a split architecture:
- the main n8n instance handles timers and webhooks
- execution IDs get pushed into Redis
- worker processes pick up jobs and run them
- results and state land in the database
That is not “just automation tooling” anymore.
That is a background job system.
n8n queue mode means you now own this stuff
- Redis as broker
- worker processes that can stall or crash
- durable DB state
- shared encryption keys across nodes
- queue backlog behavior
- worker concurrency tuning
A webhook-triggered AI flow that looked simple in the demo is now a distributed system wearing a low-code hat.
That sounds dramatic, but it’s also why production bugs show up in places the workflow builder never warned you about.
The logic was fine.
The runtime wasn’t.
Concurrency is weirder than people think
Most teams hear “concurrency” and assume it means “how many runs exist at once.”
It often does not.
n8n regular mode
In self-hosted n8n regular mode, production executions are unlimited by default unless you set a limit.
export N8N_CONCURRENCY_PRODUCTION_LIMIT=20
n8n queue mode
In queue mode, workers can be tuned directly:
n8n worker --concurrency=5
That’s already more subtle than people expect.
Now compare that to Inngest.
Inngest changes the mental model
In Inngest, concurrency limits active steps, not total runs.
That matters a lot for AI workflows.
If a function is:
- sleeping
- waiting for approval
- paused for an external event
- throttled between steps
...it doesn’t necessarily consume active concurrency the whole time.
That’s a much better fit for always-on agents than the classic model where every waiting job burns a worker slot.
Here’s a tiny example:
import { inngest } from "./client";
export default inngest.createFunction(
{
id: "generate-ai-summary",
concurrency: 10,
},
{ event: "ai/summary.requested" },
async ({ event, step }) => {
const data = await step.run("get-data", async () => {
return getDataFromExternalSource(event.data.id);
});
const summary = await step.run("generate-summary", async () => {
return callLLM(data);
});
await step.run("save-summary", async () => {
return db.summaries.insertOne({ id: event.data.id, summary });
});
}
);
The syntax is not the interesting part.
The checkpointing is.
If save-summary fails, you don’t want to rerun the whole thing from the top and burn another expensive model call if you don’t need to.
For AI automations, that difference matters a lot.
The real issue is not waiting. It’s state.
This is the part that usually changes how people think about orchestration.
Long-running automation is not mainly a timeout problem.
It’s a state problem.
Temporal gets this better than almost anyone.
The core idea is simple:
Workflows should survive crashes, deploys, pauses, and long waits by replaying event history, not by hoping some process memory stays alive forever.
If your agent pipeline depends on in-memory state surviving overnight, you built a trap.
If your workflow needs to continue after:
- a worker restart
- a deployment
- a node crash
- a multi-hour wait
...plain request/response code is the wrong abstraction.
Why Temporal feels different
Temporal is opinionated in a useful way.
Workflow execution timeout is unlimited by default.
Workflows can run for a very long time.
Retries and replay are built into the model instead of being bolted on later.
A minimal TypeScript sketch looks like this:
import { proxyActivities } from "@temporalio/workflow";
import type * as activities from "./activities";
const { generateSummary, postToDiscord } = proxyActivities<typeof activities>({
startToCloseTimeout: "5 minutes",
retry: {
initialInterval: "1 second",
backoffCoefficient: 2,
maximumInterval: "10 minutes",
maximumAttempts: 5,
},
});
export async function aiApprovalWorkflow(input: { docId: string }) {
const summary = await generateSummary(input.docId);
// imagine waiting for approval signal here
// then continue after approval arrives
await postToDiscord(summary);
}
That isn’t overengineering if your workflow behaves more like a business process than a script.
It’s just being honest about the shape of the problem.
Put flaky stuff in the right place
This is where a lot of AI teams get burned.
They put LLM calls directly inside workflow logic and then wonder why retries become messy and replay becomes dangerous.
Temporal’s model is the right one here:
Put failure-prone or non-deterministic work into Activities.
That includes things like:
- calling GPT-5
- calling Claude Opus
- calling Grok
- hitting search APIs
- posting to Discord
- waiting on HubSpot or Stripe webhooks
Why?
Because workflow code needs to stay deterministic for replay.
Activities are where retries belong.
If GPT-5 times out, you want the model call retried.
You do not want your orchestration logic inventing a new branch of reality because replay took a different path.
For AI systems, the flaky edges are the product:
- model calls
- tool calls
- webhooks
- approvals
- rate-limited APIs
That’s exactly where demos break in production.
Three good options for always-on AI work
There is no universal winner, but there are very clear tradeoffs.
| Option | What it’s really good at |
|---|---|
| n8n queue mode | Best when you already live in n8n and need webhook/timer ingestion separated from execution; Redis brokers execution IDs, workers run jobs, and you need to manage operational details like shared encryption keys and durable DB state |
| Temporal | Best when workflows are long-running, failure-prone, and need to survive crashes cleanly; event-history replay and automatic Activity retries are the point |
| Inngest | Best for event-driven automations with waits, pauses, resumptions, and step-level checkpointing without as much queue plumbing |
My opinionated version:
- If your workflow is short, idempotent, and easy to rerun, don’t reach for Temporal just to feel sophisticated.
- If your workflow waits on humans or external systems, plain request handlers are a bad joke.
- If your team already runs n8n and the pain is mostly operational, queue mode is the honest next step.
- If your agent acts more like a business process than a script, Temporal or Inngest usually fit better.
Do you actually need durable execution?
Not always.
Some jobs are fine with a queue plus workers.
If a task:
- finishes in 30 seconds
- can be retried end-to-end
- is idempotent
- doesn’t need fine-grained resume behavior
...a full workflow engine may be unnecessary.
But the more common mistake is the opposite one.
Teams keep pretending their automation is “just a webhook” long after it has become a living process with:
- pauses
- retries
- fan-out
- human review
- external dependencies
- resumptions across deploys
At that point, avoiding durable execution is not simplicity.
It’s denial.
The hidden ops tax nobody mentions in the demo
Background execution solves timeout pain.
It also creates a new pile of work.
Now you need to care about:
- queue backlogs
- retry storms
- shared secrets and encryption keys
- observability across sleeping vs stuck vs dead runs
- deterministic replay behavior
- worker scaling and drain behavior
- idempotency for downstream side effects
This is the part that makes an always-on AI system feel less like “calling an LLM” and more like operating a small factory.
That framing is healthier.
Because once the workflow is always on, the hard part is rarely the prompt.
It’s the lifecycle.
A practical design checklist
If your AI automation can pause, retry, wait, or outlive a web request, ask these questions before you ship:
- Where does workflow state live?
- What happens if the worker dies mid-step?
- Can one failed model call resume from a checkpoint?
- Does waiting consume concurrency?
- Can the workflow survive a deploy?
- Are downstream side effects idempotent?
- Can I tell the difference between sleeping, retrying, and stuck?
If you can’t answer those clearly, the demo is ahead of the architecture.
One more thing: your LLM bill gets weird when retries start
There’s another production problem people underestimate.
Once workflows become long-running and retry-heavy, token-based pricing gets annoying fast.
A failed AI step that gets replayed or retried a few times is not just an engineering problem. It becomes a billing problem too.
That’s especially painful for teams running agents in:
- n8n
- Make
- Zapier
- OpenClaw
- custom worker fleets
If your automations run all day and you’re routing across models like GPT-5, Claude Opus, and Grok, per-token pricing turns every retry policy into a finance discussion.
That’s one reason Standard Compute is interesting for this kind of workload.
It gives you an OpenAI-compatible API with flat monthly pricing, so background agents can run continuously without somebody watching token spend like a hawk. If your orchestration layer is already complicated, removing billing unpredictability helps a lot.
The rule I wish more teams started with
If your AI workflow can outlive a request, design it like a background job from day one.
Not after the first timeout.
Not after the first failed approval flow.
Not after the first customer says, “it worked yesterday.”
Start with the boring questions:
- where state lives
- how retries work
- what gets checkpointed
- what consumes concurrency
- what survives crashes and deploys
The happy-path demo is still useful.
It proves the idea.
But always-on AI systems do not fail on the happy path.
They fail in the gaps between steps:
- while waiting
- while sleeping
- while retrying
- while resuming
- while recovering
That hidden layer is most of the architecture.
Once you see that, you stop building AI automations like glorified request handlers.
And your production incidents get a lot less surprising.
Top comments (1)
"It's not a timeout problem, it's a state problem." That reframes it well. Most teams just reach for a bigger timeout when the real fix is to stop keeping workflow state in process memory at all. Putting LLM calls in Temporal Activities specifically so replay stays deterministic is the detail people skip, then they wonder why retries corrupt their workflow state.