DEV Community

jackymenCZ (jackymenCZ)
jackymenCZ (jackymenCZ)

Posted on

Your Agent Says "I Did Nothing." Does It Know Why? — Observation Mode in Sentinel

Part of the Sentinel series on building autonomous agents that are cheap, honest, and auditable. ~20 min read.

Part 1 — For everyone

The one-sentence problem

Imagine a security guard who, at the end of every shift, writes exactly one word in the logbook: "Nothing."

Nothing happened? Or nothing was checked? Was the building quiet, or was the guard asleep? Did he decide the back door was fine, or did he just never walk past it? A week of "Nothing." tells you absolutely nothing — and that's the trap most AI agents fall into.

Today's autonomous agents are built to act. We celebrate when they write code, fix a bug, deploy something. But a good agent spends most of its time deciding not to act — and that decision is invisible. It gets logged as "SKIP," the agent's version of "Nothing."

Why "Nothing" is four different things

When Sentinel (our autonomous code-maintenance agent) looks at a file and decides to leave it alone, that single "SKIP" is hiding four completely different situations: 35-cite-0

  1. "Nothing needed." The file is finished, trusted, proven. Leaving it alone is correct and there's nothing to learn. (In the code: NO_VALUE.)
  2. "I couldn't actually judge this." The agent didn't recognize the file, had no pattern to compare against, and skipped by default. That's not confidence — that's a blind spot. (INSUFFICIENT_DATA.)
  3. "Nothing needed right now, but I'm watching." The agent is making a bet: this will stay stable. (OBSERVE.)
  4. "I wanted to think, but couldn't afford it." The daily budget ran out, so it deferred — a retry-later, not a decision. (Also OBSERVE, with a shorter re-check.)

These four look identical in a normal log. They are not the same thing. The first is wisdom. The second is ignorance disguised as wisdom. And that disguise is dangerous — because an agent that confidently reports "all clear" when it really means "I have no idea" is worse than one that admits it's confused.

"Eyes open, hands in pockets"

Observation Mode is Sentinel's fix. It doesn't make the agent act more. It makes the agent honest about its silence. The internal motto in the code is literally "eyes open, hands in pockets." 35-cite-1

Here's the clever part. When Sentinel decides "nothing needed now, but I'm watching," it doesn't just note that — it writes down a prediction: "I bet this file stays stable until next Tuesday." Then, on a later visit, it comes back and grades its own bet: was I right? 35-cite-2

That turns a dead log entry into something that can be scored. Over time, the agent builds a track record: "When I said a file was fine, I was right 94% of the time." That's the difference between an agent that claims to be careful and one that can prove it.

The guard who writes real reports

Back to our security guard. Observation Mode is the difference between:

  • "Nothing." (useless)
  • "03:00 — checked the back door, locked, no activity. Expect it to stay that way; I'll re-check at 06:00." (useful — and checkable)

And crucially: if someone does break in at 04:00, the next guard sees the prediction ("expected quiet until 06:00") sitting right next to the reality ("alarm at 04:00") — an instant, honest record that the bet was wrong. Nobody has to reconstruct what happened. The miss is self-documenting.

One more thing that matters to non-engineers: this never makes the agent do anything. It only labels and records. A SKIP is still a SKIP. We're not giving the agent more power — we're giving it better memory of its own judgment. (The nervous engineer's favorite kind of feature: one that literally cannot break production because it never touches the decision.)


Part 2 — For developers

Where it sits

Observation Mode is libs/core/observation-mode.js — local, cheap, pure, runs inside Executor.processFile. It never calls the network, never stores source code, and never changes a decision. The executor still returns SKIP. 35-cite-3

The module exports exactly five functions plus some constants: classify, observe, resolveIfDue, stats, view. 35-cite-4

Two of them mutate memory (observe, resolveIfDue); the other three are read-only. That split is the whole design: writes happen only on the decision paths, reads happen only at the API boundary.

The pipeline position

It hooks the executor at two points. At the top of every processFile, it resolves any watch that's come due. On each SKIP return path, it labels the skip:

flowchart TD
    A["processFile(file)"] --> B["resolveIfDue() — grade any open bet"]
    B --> C["... decision economy gates ..."]
    C --> D{"SKIP?"}
    D -->|"novelty-reuse"| E["observe()"]
    D -->|"lifecycle skip"| E
    D -->|"local LLM-skip"| E
    D -->|"intent-gate skip"| E
    D -->|"budget-deferred"| E
    D -->|"EVOLVE / REJECT"| F["(not observed — file changes or gate acts)"]
    E --> G["SKIP returned unchanged"]

resolveIfDue is called once at the top: 35-cite-5 and observe is wired into each SKIP return — novelty-reuse 35-cite-6 , lifecycle skip 35-cite-7 , local LLM-skip 35-cite-8 , intent-gate skip 35-cite-9 , and budget-deferred 35-cite-10 .

⚠️ A trap we actually hit: the executor's SKIP paths are ordered, and an earlier one can return before your hook runs. The Novelty Gate reuses a prior SKIP for an unchanged file before the lifecycle block — so a naively placed hook looks dead in e2e while its unit tests pass. The fix is to call observe() on every return path, and to assert on the log markers (♻️ NOVELTY REUSE, 🏁 LIFECYCLE SKIP) in e2e tests, not just the final action. (Identical SKIP can come from very different code paths.) 35-cite-11

classify — a pure function of (reason, envelope, lifecycle)

The entire taxonomy is one pure function. No I/O, no clock dependency beyond what's passed in, trivially testable: 35-cite-12

The decision order matters (first match wins):

Order Condition Category Watch?
1 reason contains budget OBSERVE yes, 6h retry
2 certified/locked (frozen) NO_VALUE no
3 confident known skip, still under_observation OBSERVE yes
3b confident known skip, trusted module NO_VALUE no
4 unknown role / matchedPatterns === 0 INSUFFICIENT_DATA no
5 anything else OBSERVE yes

The budget check is deliberately first — a deferral is a retry-later, not a judgment, so it must not be misread as a confident "nothing needed." 35-cite-13

The prediction: a falsifiable bet

Only OBSERVE produces a prediction. It's a small, honest structure: 35-cite-14

const prediction = {
    stable: true,
    risk,
    estimatedProblemProbability: problemProbability(envelope)
};
Enter fullscreen mode Exit fullscreen mode

The problem probability is clamp((1 - confidence) * riskFactor, 0.01, 0.9) — higher when the agent is unsure, higher when the file is dangerous: 35-cite-15 with risk factors low: 0.5, medium: 1, high: 1.5, critical: 2. 35-cite-16

The re-check horizon scales with how proven the module is — you don't re-interrogate a 7-day-trusted module as often as a brand-new one: under_observation: 24h, stable: 3d, certified/locked: 7d, and 6h for budget deferrals. 35-cite-17

The checkpoint anti-pattern we avoided

Here's a subtle bug that would quietly kill the whole feature. If you recompute reevaluateAt every time a busy scanner re-observes an open watch, the deadline keeps sliding into the future and the bet is never graded. (It's the "I'll start the diet tomorrow" bug, in code form.)

So the checkpoint is set once and held while the watch stays open: 35-cite-18

reevaluateAt: prev?.reevaluateAt ?? (now + (classification.reevaluateInMs ?? REEVAL_MS.default)),
rollbackBaseline: prev?.rollbackBaseline ?? rollbackStreakOf(lifecycle)
Enter fullscreen mode Exit fullscreen mode

The firstTs and rollbackBaseline are preserved the same way, and observations is bumped — so you keep an accurate count of how many times you looked without resetting the clock. There's a dedicated test for exactly this (see below).

resolveIfDue — grading the bet

A watch resolves when either the window elapsed or an incident happened (rollback streak rose above the baseline, or the module got locked). The outcome is correct = (predictedStable === observedStable): 35-cite-19

const incident = currentStreak > num(obs.rollbackBaseline) || status === "locked";
const due = now >= num(obs.reevaluateAt);
if (!incident && !due) return null;

const observedStable = !incident;
const predictedStable = obs.prediction?.stable !== false;
const correct = predictedStable === observedStable;
Enter fullscreen mode Exit fullscreen mode

The incident path is what makes it early-warning rather than post-mortem: a rollback mid-window closes the watch immediately as WRONG, before the scheduled re-check. The resolved record feeds stats() (resolvedCorrect / resolvedWrong → prediction accuracy) and becomes the raw log for the next layers: reflection, outcome learning, calibration v2. 35-cite-2

Memory shape & bounds

State lives at memory.observation_log = { open, resolved, totals }, persisted via the normal snapshot (schema field added in memory-manager defaults + both normalize paths). 35-cite-20 It's bounded like every other Sentinel memory layer: open to OPEN_LIMIT = 500 (oldest lastTs evicted), resolved to RESOLVED_LIMIT = 200. 35-cite-21 35-cite-22 (If you read the earlier memory articles: yes, this obeys the same "never grow unbounded" discipline as the rest of the brain.)

It's exposed read-only at GET /observations (view) and summarized in GET /analytics and GET /dashboard (stats) — both pure, never mutating. 35-cite-23

The honest limitation

The only incident signal wired today is per-file rollback/lock inside memory. External signals — branch/MR observations, human confirmation, production runtime errors — are future extensions, not shipped. So a prediction graded "CORRECT" means "nothing broke that Sentinel's own memory saw," not "nothing broke anywhere." That's a real boundary, and it's stated in the module and the context doc rather than hidden. 35-cite-24

Feature flag

ON by default; kill it with config.OBSERVATION_MODE.ENABLED === false or OBSERVATION_MODE_ENABLED="false". When off, nothing is written. 35-cite-25


The tests

Yes — there are tests. tests/unit/observation-mode.test.js has 18 unit tests using the native node:test runner (node --test), organized in four blocks. 35-cite-26

classify — the taxonomy (6 tests): frozen → NO_VALUE with no prediction; confident skip is NO_VALUE on a trusted module but OBSERVE on a young one; budget-deferred → OBSERVE with a ≤6h window; unknown role / zero matched patterns → INSUFFICIENT_DATA; default known skip → OBSERVE with a bounded probability; and problem-probability monotonicity (rises with risk). 35-cite-27

That last one is a nice property test — it asserts highRisk.estimatedProblemProbability > lowRisk... at equal confidence, pinning the direction of the formula, not a magic number: 35-cite-28

observe — watch lifecycle (5 tests): opening a watch sets the prediction + 24h window + observations: 1 and tallies the category; NO_VALUE/INSUFFICIENT_DATA tally but open no watch; switching an OBSERVE file to NO_VALUE closes the prior watch; re-observing preserves firstTs/reevaluateAt while bumping the count (the checkpoint anti-pattern guard); and bounded eviction beyond OPEN_LIMIT. 35-cite-29 35-cite-30

The checkpoint-stability test is the one that matters most for correctness:

ObservationMode.observe(memory, 'a.js', { ..., now: 1000 });
ObservationMode.observe(memory, 'a.js', { ..., now: 5000 });
const rec = memory.observation_log.open['a.js'];
assert.equal(rec.firstTs, 1000);
assert.equal(rec.reevaluateAt, 1000 + 24 * 60 * 60 * 1000); // did NOT slide to 5000+24h
assert.equal(rec.observations, 2);
Enter fullscreen mode Exit fullscreen mode


35-cite-31

resolveIfDue — grading (4 tests): not due + no incident → null, keeps watching; window elapsed, no incident → correct: true, reason window-elapsed, watch closed, resolvedCorrect incremented; a rollback mid-window → correct: false, reason incident-detected, before the window; and no open watch → null. 35-cite-32

The incident test proves the early-warning path works independent of the clock:

const r = ObservationMode.resolveIfDue(memory, 'a.js', {
    lifecycle: { status: 'under_observation', rollbackStreak: 1 }, now: 2000 // long before 24h
});
assert.equal(r.correct, false);
assert.equal(r.reason, 'incident-detected');
Enter fullscreen mode Exit fullscreen mode


35-cite-33

stats / view — read-only aggregation (3 tests): end-to-end category counts + prediction accuracy after a correct resolve; empty memory returns a safe zero/null shape (accuracy: null, not NaN); and view lists distilled open watches + recent resolved with a boolean dueForReeval. 35-cite-34

Every test drives the real functions with injected now, so nothing waits on a real clock — a 24-hour window is tested in microseconds by passing now: 1000 + 24*60*60*1000 + 1. That's the pattern the testing-skill doc recommends for time-windowed resolution: mutate/supply the checkpoint, don't sleep. 35-cite-35


Takeaway for agent builders

If you're building an autonomous agent, the lesson transfers directly: "did nothing" is not an outcome — it's four outcomes wearing the same coat. Separate them. For the "watching" case, record a falsifiable prediction with a fixed re-check deadline, and grade it on revisit. You get a self-documenting track record of your agent's judgment for free, you catch the "confident nonsense" state (INSUFFICIENT_DATA masquerading as NO_VALUE), and — if you keep it pure and read-only like this — it cannot possibly break the thing it's observing.

The guard who writes "Nothing." gets fired eventually. The one who writes down what he expected and whether he was right gets promoted. Build the second kind.

Top comments (0)