DEV Community

openamer
openamer

Posted on

One heartbeat, not 84 cron jobs: the scheduler at the core of a self-hosted agent

One heartbeat, not 84 cron jobs: the scheduler at the core of a self-hosted agent

Why an autonomous desktop agent should own a clock instead of outsourcing it to the operating system's job runner — and what that buys you when things fail.

If you run an agent that is supposed to keep working while you sleep, you eventually face a scheduling problem. The obvious answer is the system's own job runner: cron on Linux, Task Scheduler on Windows. Start with five jobs. Then twelve. Then you are at eighty-four, each with its own interval, its own log, its own way of failing silently — and no single place where you can answer the only question that matters at 3am: what is actually running, and what is stuck?

OpenAmer is an open-source desktop AI agent (Apache-2.0) that runs locally on Windows. This post is the deep-dive I get asked for most: the ASI heartbeat — a single in-process loop that replaced our pile of scheduled jobs.

The problem with one job per concern

Job-per-concern looks clean in a diagram and rots in practice, for three reasons:

  1. No shared state. Each job wakes up blind. It can't tell whether the thing it depends on already ran, is mid-run, or failed — so it either does redundant work or assumes success.
  2. Failure is invisible. A failing cron job is a line in a log nobody reads. Eight failing jobs still look like a healthy schedule.
  3. No back-off. A job that fails every tick keeps failing every tick. There is nowhere to say "this subsystem is unhealthy, slow it down and isolate it."

The deeper issue is architectural: the schedule lives outside the agent. The agent's own logic cannot see, reason about, or adjust its own cadence. That is backwards for a system whose whole job is to keep itself running.

What the heartbeat actually is

The heartbeat is a single loop that owns cadence for the whole system. Each subsystem declares an interval; the loop checks whether it is due and, if so, runs it — in the same process, via a direct function call.

# tools/asi/heartbeat.py  (conceptually)
class Heartbeat:
    SUBSYSTEMS = {
        "learning": 300,   # every 5 min
        "system":   300,   # self-heal, resource monitor
        "senses":   1800,  # circadian, trend scout
        "swarm":    1800,
        "infra":    1800,  # env check, browser, plugin, mesh
        "meta":     3600,  # reflection, goal, research
        "security": 14400, # bugbot, CVE scan, code review
        "a2a":      14400, # brain export, peer comms
        "darwin":   900,   # evolution, autopatch, publish, probe
        "outreach": 10800,
    }

    def tick(self, system=None, force=False):
        for name, interval in self.SUBSYSTEMS.items():
            if system and name != system:
                continue
            if force or self.is_due(name, interval):
                self.run(name)          # direct call, not a subprocess
                self.mark_ran(name)     # persisted timestamp
Enter fullscreen mode Exit fullscreen mode

Last-run timestamps live in a small JSON state file (memory/asi_heartbeat.json), so the cadence survives across process restarts. The loop is driven by one scheduler entry — asi-heartbeat-tick, every five minutes — and every subsystem rides on top of it.

Ten subsystems, one clock

The subsystems are not arbitrary; they are the organs the agent needs to stay alive and improve:

Subsystem Cadence What it covers
learning 5m internet learner, active learn, knowledge transfer
system 5m self-healer, resource monitor, traffic cop
darwin 15m evolution, autopatch, publish, probe
senses 30m circadian rhythm, watchtower, trend scout
swarm 30m swarm intelligence, autonomous loop
infra 30m env check, browser, plugin, mesh, cache
meta 60m reflection, self-rewriter, goal, research
security 240m bugbot, CVE scan, pen test, code review
a2a 240m brain export, peer communication
outreach 180m social, GitHub, growth report, funding

Cadence lives in one dict. Changing "how often does the agent learn from the internet" is a one-line edit, not a hunt through a job runner's UI.

The part that matters most: in-process, not subprocess

The heartbeat is only half the story. The other half is what a job runs.

Previously each capability was an external script invoked as a process:

# BEFORE: a process per capability
import subprocess
subprocess.run([sys.executable, "scripts/training/self_model.py"])
Enter fullscreen mode Exit fullscreen mode

Now each subsystem is imported and called directly:

# AFTER: a function call
from tools.asi import self_model
state = self_model.gather_state()
Enter fullscreen mode Exit fullscreen mode

Three things fall out of this:

  • No boot tax. A subprocess pays interpreter start, import cost and (de)serialization on every invocation. A function call pays none of it — which is what makes a five-minute cadence affordable on a CPU-only laptop instead of an expensive luxury.
  • No shell-escaping surface. The subprocess boundary is a place bugs live: quoting, argument injection, encoding. A call boundary is a place bugs don't.
  • Shared state for free. Because everything is in-process, a subsystem can read the heartbeat's own state, and the loop can see a subsystem's result — the "no shared state" problem from the cron world simply disappears.

The same five capabilities are also exposed as native agent tools — asi_status, asi_think, asi_learn, asi_remember, asi_trigger — and as a CLI:

openamer asi status      # full system + heartbeat health
openamer asi heartbeat   # tick the loop (optionally one subsystem)
openamer asi trigger <capability>
Enter fullscreen mode Exit fullscreen mode

Honest trade-offs

This is not free lunch, and the trade is deliberate:

  • A single loop is a single loop. If the heartbeat process dies, everything downstream stalls until it comes back. That's why the loop is the smallest possible component — a few hundred lines with no heavy imports at module level — and why an outside watchdog only has to keep one thing alive instead of eighty-four.
  • In-process means shared failure domain. A crash in a subsystem can take the loop with it. We accept this because the alternative (process isolation) is what made the fleet expensive and unobservable in the first place; each subsystem is written to fail soft and return, not raise.
  • Eventually, not exactly. Five-minute granularity is fine for learning, healing and outreach; it is the wrong tool for sub-second control loops. The heartbeat schedules background work, not request/response.

Why it matters

The point is not that a heartbeat is clever. It is that an autonomous system should own its own clock, its own state and its own health — in one place it can measure. Once scheduling is a function call inside the agent, "is the agent healthy?" stops being an archaeology exercise across a job runner and becomes a single status query.

Code and the heartbeat module live here: https://github.com/openamer/openamer — the scheduler is under tools/asi/heartbeat.py.

If you've run a long-lived agent in production: what finally made you move scheduling into the system, or what made you keep it outside? I'm especially curious about the failure-isolation patterns people land on.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to