How to Make a Scheduled AI Agent Resume After a Failed Run Instead of Starting Over
Your scheduled agent runs fine for weeks. Then one night it crashes halfway through: a rate limit, a timeout, one bad record. The retry fires, and the agent starts from the very beginning. It re-fetches everything it already processed, re-sends emails that already went out, re-burns thousands of tokens re-reading context it already understood. The failure did not just cost you one run. It cost you the whole run, twice.
This is one of the least talked about failure modes in AI automation. Everyone worries about the agent making mistakes. Fewer people think about what happens when the run itself dies mid-flight. A retry that restarts from scratch is not a recovery plan. It is a second chance to make the same expensive mess.
Why "retry" is not "resume"
Most automation platforms give you retries, not resumes. Take n8n: hitting retry on a failed execution re-runs the workflow with the original input. In queue mode, if a worker crashes, the job is reclaimed and restarted from the beginning. n8n does not checkpoint mid-workflow state another worker can pick up. The progress-saving setting exists for UI visibility, not for resuming the exact state of a crashed run.
The same is true for the agent itself. A scheduled AI agent wakes up with no memory of the previous run. Even a successful run leaves nothing behind except its output. So when a run fails at step 14 of 20, the retry is a new agent, in a new run, with zero knowledge that steps 1 through 13 already happened.
Retry answers the question "should we try again?" Resume answers "where were we?" Confusing the two is where the damage happens.
What a blind restart actually costs
A restart-from-zero failure is expensive in three ways, and only one of them is obvious.
First, duplicate side effects. If the agent had already sent 60 of 120 follow-up emails before it crashed, the retry sends all 120. Contacts get duplicates. CRM records get created twice. Making every write idempotent fixes the worst of this, and you should do that regardless, but idempotency only prevents the damage. It does not recover the lost progress.
Second, wasted compute. The retry re-fetches the data, re-reads the files, re-reasons over the same decisions. If the run cost 80,000 tokens before the crash, a full restart doubles it. For agents that run on expensive models against large datasets, one nightly crash can quietly burn more budget than a week of successful runs.
Third, and most underrated: lost reasoning. The crashed run had figured things out. It had read the tricky records, decided which ones were duplicates, learned that one vendor's date format needed special handling. All of that was in its context window, and the context window died with the process. The retry does not just redo the work. It redoes the thinking.
The checkpoint pattern (and its limit)
The standard fix is checkpointing. Instead of one long run, process in small batches and save a progress marker after each one: the last processed page, the last record ID, a cursor. The next run reads the cursor and continues from there. Combined with idempotent writes, a crashed run picks up at the last safe point instead of at zero.
This pattern works, and it is worth implementing. But it solves only the machine half of the problem. A cursor tells the next run where to continue. It does not tell the agent what the previous run knew.
Checkpoints are machine state. "Last page: 20" is a fact, not a memory. It does not say why the agent skipped three records on page 19, or that the API started rate-limiting at 2:40 AM, or that the agent switched to a fallback source after the primary timed out twice. The resumed run inherits the position but not the understanding. It will happily re-learn the rate limit the hard way, on the same night, in the same run.
You need two layers of resume: the cursor that says where, and the memory that says what happened.
Giving the resumed run its past back
A resume-ready agent saves more than progress markers. At each checkpoint it also records what it learned: decisions made, anomalies spotted, workarounds used, failures hit. Not a freeform log dump, but saved context the next run actually loads before it acts.
The mechanism matters less than the habit. What matters is that the agent's first step on every run, especially a resumed one, is to load what previous runs knew: what was already done, what was in flight when things broke, and why things were done that way. A resumed run that starts by reading the last run's memory behaves like a person picking up a dropped task. It checks where it left off, remembers what was going wrong, and continues.
This works across tools, too. A memory layer that lives outside any single platform means the n8n workflow that crashed, the retry that fires next, and the session you open in the morning to investigate can all read the same run history. The agent's memory should not die with the platform process that crashed.
The resume-ready checklist
- Process in small batches and save a progress cursor after each successful batch.
- Make every write idempotent, so a resumed run can never double-apply an action.
- At each checkpoint, save decisions and anomalies alongside the cursor: what the agent learned, what failed, what workaround it used.
- On every run start, load the previous run's memory before acting, and treat a non-clean shutdown as the normal case, not the exception.
- Expire stale cursors. A checkpoint from three weeks ago is a lie waiting to happen; the data it points to has changed.
Where a memory layer fits
Checkpoint cursors live naturally in a database. The run memory, the reasoning, the decisions, the "API started rate-limiting at 2:40 AM" part, needs somewhere to live too, and it needs to be readable by whatever runs next, in whatever tool it runs in.
Vilix AI is built for exactly this: one shared memory layer your agents read and write over MCP. It is cloud-hosted, so there is no database to provision or keep alive through a crash. The same memory follows your agents across tools: the n8n workflow, the scheduled script, and the chat session you use to debug the failure all read the same run history. It stores full conversation history, not just extracted facts, so the resumed run gets the reasoning back, not just a summary. Free plan forever, a 7-day Pro trial with no credit card, and you can export everything or delete it all anytime in a portable format.
The goal is simple. When your agent crashes at 3 AM, the next run should not wake up blind. It should wake up, remember where it stopped and what went wrong, and keep going. Retry is giving up on the run and hoping the next one goes better. Resume is finishing what you started. Build your agents for the second one.
Top comments (0)