DEV Community

Cover image for Sentinel: The Command Plane
Philip Shaw
Philip Shaw

Posted on Originally published at glitchedpixel.io

Sentinel: The Command Plane

Everything in this series so far has been about reading from devices. This is the part where Sentinel starts writing to them, and a mistake here can leave a heater running or unlock a door twice.

A writable point has two values. reported is what the device says and desired is what the system wants. Nothing writes to a device directly: a rule or a person sets a claim, arbitration resolves the claims on a point into desired, and a reconciler drives reported toward it. A device that reboots into the wrong state is corrected on its next report, and nobody has to send the command again.

Every claim carries an owner and one of six bands, and the band decides who wins. A person's manual claim has to expire, because otherwise one touch of a switch disables the automation on that point until someone notices in November that the heating never came back on. Frost protection can be suspended for half an hour by someone who names it and gives a reason. An interlock can't be suspended at all.

A write on the bus is fire-and-forget, so each command is followed through a state machine to an outcome. For a relay whose state can't be read, the outcome is unverified and the point is marked assumed, which a rule can choose not to trust. A core that restarts in the middle of a write marks the command timeout and works out what to do next from desired and reported. It never sends the same write a second time. And no command is created for a device that can't be reached: the claim stands and waits, so an interlock asserted during cold start is issued the moment its driver is ready, with no deadline it could have missed in the meantime.

Rules never send commands. They set claims and release them, and that is what lets the rules engine, the next part, be tested without a device in the loop.

What this document owns: desired and reported as separate values on a writable point, the durable claim set and the projection resolved from it, the convergence states and the bounded retry, the command state machine, per-point serialisation and queue coalescing, the four confirmation modes, deliverability and re-drive, the six bands and how arbitration resolves them, manual expiry and withdrawal, suspension of a named owner, how a rule's band is declared and checked at registry load, fail-safe versus fail-clear on input loss, the limit a confirmation mode puts on what an interlock can protect, the intent API and the direct-command escape hatch, validation before the wire, and idempotency of an intent. It does not own core_epoch and the fencing behind it, which belong to Core Runtime; how a principal is established, which belongs to Security Model; the command subject, the driver's acknowledgement, the instance states that deliverability reads and the idempotent flag on a driver assignment, which belong to Driver Contract; the timer that claim expiries share with freshness deadlines, which belongs to Watchdog & Derived Points; or the shape of a rule and when it sets or releases a claim, which belong to Rules Engine.

The command plane is where the system stops being an observer and becomes able to be wrong in ways that matter. Two things carry the weight: the state machine that turns a fire-and-forget write into a verified outcome, and arbitration, which decides whose intent wins.

Desired and reported are separate points

A writable point has two values, not one. reported is what the device says. desired is what the system wants. They are different fields on the same point, with independent quality.

Nothing writes to a device directly. Something sets a claim, arbitration resolves the claim set into desired, and a reconciler drives reported toward it. This is the difference between "send an off command" and "this light should be off", and it is what makes the system self-healing: a device that reboots into the wrong state gets corrected on its next report without anyone re-issuing anything.

The claim set is durable; desired is a projection of it. Arbitration has to fall through to the next-highest claim when one is withdrawn, and to restore a suspended claim when its suspension expires. Both require every active claim to be stored, not just the winner.

CREATE TYPE band AS ENUM (        -- declaration order is precedence order
  'default','automation','schedule','manual','safety','interlock');

CREATE TYPE convergence AS ENUM (
  'no_intent','undeliverable','converging','converged','diverged');

CREATE TABLE claim (
  point_id    text        NOT NULL,
  owner       text        NOT NULL,   -- 'rule:frost-protect', 'user:alice', 'system:default'
  band        band        NOT NULL,
  value_num   double precision,
  value_bool  boolean,
  value_text  text,
  set_at      timestamptz NOT NULL,
  expires_at  timestamptz,            -- required for band 'manual'
  reason      text,
  PRIMARY KEY (point_id, owner)
);

CREATE TABLE point_desired (
  point_id     text PRIMARY KEY,
  value_num    double precision,
  value_bool   boolean,
  value_text   text,
  owner        text,                  -- null when no claim is active
  band         band,
  resolved_at  timestamptz NOT NULL,
  convergence  convergence NOT NULL,
  last_command uuid
);
Enter fullscreen mode Exit fullscreen mode

Postgres compares enum values in declaration order, so the band is the comparison. There is no separate priority number to keep in step with it.

One claim per owner per point. A rule asserting a different value updates its own row rather than accumulating claims.

Convergence is evaluated on every reported transition. desired set with no matching reported after the confirmation deadline is a divergence, and divergence is a first-class alertable condition, surfaced as a self point per writable point so that ordinary rules and dashboards can act on it.

Divergence gets a bounded retry: reconcile, wait, reconcile once more, then stop and mark diverged. Unbounded reconciliation against a device that physically cannot reach the desired state means a relay clicking every thirty seconds until someone notices.

What the reconciler does when the device is not reachable at all is a different question, dealt with under deliverability below.

Command state machine

                  ┌──────────→ rejected
                  │
created → queued → sent → acked → verifying → verified
                    │       │         │
                    │       │         └────→ unverified
                    ├───────┴─────────────→ failed
                    └─────────────────────→ timeout
Enter fullscreen mode Exit fullscreen mode
  • created — intent accepted, arbitration passed, command_id assigned
  • queued — waiting on the per-point serialiser
  • sent — published to cmd.{host}.{instance}.{point_id}
  • acked — driver reports transport accepted
  • rejected — driver refused before attempting (unknown point, bad value, instance not ready, stale epoch)
  • verifying — awaiting confirmation per the point's mode
  • verified — reported now matches desired
  • unverified — optimistic mode, or verification window elapsed with no contradiction; point quality becomes assumed
  • failed — driver reported an error
  • timeout — no driver response within the deadline

Every transition is a control_event, carrying the principal that caused it and the core epoch that issued it. Terminal states carry the elapsed time, which gives per-device command latency as a self point without any extra instrumentation.

Every command carries core_epoch. It is the fencing token from Core Runtime, and a driver rejects any command whose epoch is below the highest it has seen, with reason stale_epoch. That rejection is treated differently from every other rejection: an unknown point or a bad value is a configuration error, but a stale epoch means two cores existed and one of them was still issuing commands. It escalates rather than retrying.

Serialise per point. One in-flight command per point, others queued. Concurrent writes to the same relay are never what anyone wanted, and a queue depth of more than one or two is itself a signal worth exposing.

Coalesce the queue. If two commands for the same point are queued and neither has been sent, the older is dropped. A rule flapping should produce one write, not a backlog that plays out after the condition has passed. Coalescing never crosses a sent boundary — the ordering guarantee has to hold for anything already on the wire.

Confirmation modes

inline — the driver's device_ack carries the new value. Compare against desired, go to verified or failed. Latency is one round trip.

async_event — wait for a subsequent observation on the point. Correlate by value match within the verification window, because most protocols do not echo command IDs. Ambiguity is real here: if something else changed the point to the same value in that window, it will be attributed wrongly. That is accepted, and the window kept tight, typically a couple of seconds.

poll_verify(delay) — the core schedules a read after the delay. This is the Modbus path, and it is the core issuing a read, not the driver remembering to. Deliberately a re-read rather than waiting for the next scheduled poll, because a 30 s poll interval is too long to hold a command open.

optimistic — no verification exists. The command goes to unverified, reported becomes the desired value with quality assumed, and it stays assumed until something real contradicts or confirms it. This is honest about the write-only relay whose state cannot be read, and the assumed quality means a rule can choose not to trust it.

Verification timers are durable, in control_event, rehydrated at restart. A core that restarts mid-verification resumes the wait rather than losing the command.

On restart, in-flight commands do not resume being sent. Anything in sent or verifying at shutdown is marked timeout on recovery, and the reconciler re-derives from desired versus reported. Replaying a write whose outcome is unknown is how a door gets unlocked twice.

Deliverability and re-drive

Reconciliation only means anything when the device is reachable. Two conditions decide that, and the driver plane already publishes both.

A point is deliverable when its driver instance is ready or degraded, and the point's own quality is not unavailable. An instance still connecting and a device the driver has declared off the network are the same thing from here: a command would go into the dark.

Undeliverable is a state of the intent, not the outcome of a command. No command is created while the target is undeliverable. The claim stands, desired holds its value, convergence sits at undeliverable, and nothing is queued against a wall-clock deadline. A deadline that expires for reasons unconnected to the device says nothing about the device.

This is why convergence is an enum rather than the converged boolean that would otherwise be the obvious choice. A boolean has to mean three different things — no intent, not yet delivered, and genuinely failing — and that is the same null-means-several-things trap the semantic model exists to escape.

The retry budget counts only attempts made while deliverable. The bounded retry above exists to stop a relay clicking at a device that cannot reach the desired state. A device that is not there has not refused anything. Burning both attempts against an unreachable instance produces a diverged that means nothing, and consumes the budget the real divergence would have needed later.

Re-drive on the transition to deliverable. When an instance reaches ready, or an observation clears a point's unavailable, every non-converged claim on the affected points is re-driven and its retry budget resets. It is the same reconciliation a reported transition triggers, on a different edge.

Spread the re-drive. A driver host with two hundred points reaching ready would otherwise issue two hundred commands in the same instant, at a device and a transport that have just finished coming up. Modbus in particular does not tolerate that. Concurrency is bounded per instance and the issue jittered. The per-point serialiser and the per-point rate limit do not help here, because the burst is across points rather than within one.

Cold start is the ordinary case of this, not a special one. At core start every instance is connecting, so every claim is undeliverable and nothing is sent. Claims made during warmup — in practice the fail-safe interlocks, since every other rule is gated — are recorded immediately and issued the moment their instance becomes deliverable. The protective claim is never racing a timer it can lose; it is waiting on a condition it can only win.

Undeliverable is not silent for protective claims. An ordinary claim can wait indefinitely and nobody needs to know. An interlock or safety claim sitting undeliverable is unenforced protection, which is exactly as serious as an interlock that has diverged. Convergence is exposed as a self point per writable point; for interlock_protected points, undeliverable past a short bound escalates rather than waiting quietly.

Arbitration

The part that is cheap now and expensive later.

Every claim carries an owner and a band. The band is the whole of the precedence model.

band typical owner notes
interlock protective rules that may not be overridden requires interlock_class
safety frost protection, over-temperature cutouts suspendable, never outrankable
manual a person expires_at required
schedule time-based rules
automation ordinary rules
default system:default optional resting state from the descriptor

Resolution. The active claim set for a point is every claim whose expires_at is unset or in the future and whose owner is not currently suspended. The winner is the highest band; ties within a band go to the most recent set_at. The resolver is a pure function from the active set to a winner, and it runs on every claim insert, update, withdrawal, expiry and suspension change.

default is an ordinary claim, and it is optional. A point descriptor may declare default_value; when it does, the registry materialises a non-expiring claim at band default owned by system:default. When it does not, no such claim exists and there is no special case in the resolver.

An empty active set means no intent. point_desired.owner is null, convergence is no_intent, and the reconciler does nothing. Releasing the last claim on a point leaves the device where it is rather than driving it somewhere. Asserting a resting state is something a point opts into by declaring default_value.

That is right for a light and wrong for anything protective, so the descriptor carries require_intent. A point with require_intent: true must declare default_value, checked at registry load, which makes an empty active set unreachable and desired never null. Any point marked interlock_protected gets require_intent implicitly, and the load fails if it has no default.

Manual claims expire. Band manual requires expires_at, defaulting to an hour or two and configurable per point. On expiry the claim is deleted, arbitration re-resolves, and the point falls through to whatever automation wants. Without expiry, every manual touch permanently disables automation for that point, and the discovery comes in November when the heating never came back on.

Claim expiries run on the same timer mechanism as freshness deadlines, and like them they are derived rather than stored: recomputed from expires_at when the claim set loads, so a restart cannot lose one.

Withdrawal, not just setting. An owner can release its claim, and the point falls through to the next-highest active claim, or to nothing. Rules that have stopped applying must release rather than continue asserting, otherwise arbitration resolves against stale claims forever.

Every resolution is logged to control_event with the winner, the full losing set, and the principal responsible. "Why is this light on" is then a query against the claim set, not an exercise in reasoning about rule ordering.

Two load-time checks:

  • Two rules claiming the same point at the same band is a lint warning. It is almost always unintentional, and last-writer-wins between them is rarely what anyone meant.
  • Two interlock rules claiming the same point with different values is an error, not a warning. Resolving that at runtime is worse than failing the build.

Safety bands

interlock and safety are separate bands, and the difference between them is not precedence. manual outranks neither. The difference is that a safety claim can be suspended and an interlock claim cannot.

Suspension is a distinct mechanism from precedence. It is not that manual wins the comparison; it is that the suspended claim is temporarily withdrawn from arbitration on explicit request.

Suspension is scoped to the owner, not to the point. A frost-protection rule holding three pumps is suspended once, for all three. Per-point suspension would let someone servicing the system suspend one pump, believe protection was off, and be wrong about the other two.

CREATE TABLE claim_suspension (
  owner        text PRIMARY KEY,
  suspended_by text        NOT NULL,
  suspended_at timestamptz NOT NULL,
  expires_at   timestamptz NOT NULL,   -- never open-ended
  reason       text        NOT NULL
);
Enter fullscreen mode Exit fullscreen mode
POST /claims/rule:frost-protect/suspend
  { for: 30m, reason: "servicing the pump" }
Enter fullscreen mode Exit fullscreen mode

The request is rejected if the named owner holds any claim at band interlock. Otherwise the suspension is recorded, every point that owner claims is re-resolved, and the response returns those points with their new winners — so the caller sees the actual blast radius of the decision. The suspension and its expiry are both control_event entries, and on expiry the claims return automatically.

expires_at is not nullable and there is no unsuspend-forever. A suspension that outlives the attention of the person who set it is the failure this mechanism exists to prevent.

Modelling it as suspension of a named owner rather than as a precedence comparison forces the person to name what they are overriding. "Turn this on" is ambiguous. "Suspend frost protection for thirty minutes" is a decision someone can be held to, and it is legible in the log afterwards.

Which band a claim gets

The band belongs to the rule, declared in the registry, and validated at load rather than trusted at runtime:

- id: pump-dry-run-cutout
  band: interlock
  interlock_class: thermal
  ...
Enter fullscreen mode Exit fullscreen mode

band: interlock requires interlock_class (thermal | electrical), and the registry rejects it otherwise. That forces the classification to be explicit and reviewable in a diff rather than emerging from whoever wrote the rule.

Two load-time checks:

  • No point may have interlock claims from more than one rule unless they assert the same value. Two interlocks disagreeing is a design error, and resolving it by precedence at runtime is worse than failing the build.
  • Any point that a rule targets with band: interlock gets interlock_protected: true on its descriptor. That flag is what the API and UI key off, and it means protection is a property of the point that is visible without walking the rule set.

Fail-safe versus fail-clear

The harder question, and the one the band split does not answer on its own: what does an interlock do when it cannot evaluate?

An over-temperature cutout depends on a temperature point. If that point goes stale, bad, or unavailable, the rule has no basis for a decision. Two possible behaviours, and they are opposite:

  • fail_safe — assert the protective value. Cut the heater. Correct for anything where the failure mode is thermal runaway or overcurrent.
  • fail_clear — release the claim. Correct where asserting the protective value is itself harmful — a circulation pump forced off during a heat cycle can cause the boiler overheat the interlock was meant to prevent.

There is no sensible default, so it is required:

- id: heatsink-overtemp
  band: interlock
  interlock_class: thermal
  on_input_loss: fail_safe
  fail_value: false
Enter fullscreen mode Exit fullscreen mode

fail_safe requires fail_value. The registry rejects an interlock rule without on_input_loss — this is the one place where inference is refused outright, because getting it wrong is silent until the day it is not.

For fail_safe interlocks, the asserted claim persists until inputs return to live and the rule evaluates cleanly, not merely until data reappears. A single good reading after a fault does not clear a cutout.

Confirmation mode constrains what an interlock can protect

An interlock on a point with confirmation_mode: optimistic cannot guarantee anything — it asserts a value with no way to verify it, and the point sits at quality assumed. That is acceptable for a light and not for a thermal cutout.

This is a load-time rejection: interlock_class: thermal or electrical requires the target point to have a verifiable confirmation mode (inline, async_event, or poll_verify). Protecting a write-only relay requires a second point that reads the actual state — a current sensor on the circuit, an auxiliary contact — rather than an interlock that cannot see what it is doing.

Interlock non-convergence is treated differently from ordinary non-convergence in both directions it can fail. diverged gets a shorter deadline, no bounded-retry give-up, and escalation rather than a settled state. undeliverable escalates too, on the bound described under deliverability. An interlock that has failed to take effect is the loudest thing this system can say, and it does not matter whether it failed because the device refused or because the device was not there.

The intent API

POST   /points/{id}/claim
  { value, expires_in?, reason?, idempotency_key? }

DELETE /points/{id}/claim        withdraws the caller's own claim
Enter fullscreen mode Exit fullscreen mode

Owner and band are derived from the authenticated principal, never supplied in the body. A rule's band comes from its registry declaration; a person's is manual. An API that takes owner and a priority number as parameters lets any caller assert whatever authority it likes, which is not an authorisation model. The mechanism for establishing the principal is in Security Model.

Every mutating call names a principal, and there is no default. A request that arrives without one is rejected rather than attributed to a service identity, and the principal is written to every control_event the call produces. That requirement is what makes the claim log an audit trail rather than an activity feed: nothing can change state without saying who asked.

The response is the resolved arbitration outcome and the resulting convergence state, plus a command_id if this claim won and a write was issued. A claim that won on an undeliverable point returns undeliverable and no command ID, which is a normal answer and not an error. A losing claim is still recorded, and the response returns the current winner. That is not an error either, and the caller may well want to know it was overruled.

Withdrawal removes only the caller's own claim. Removing someone else's is suspension, which is the mechanism above and carries its rules.

A direct-command escape hatch exists for commissioning and debugging, gated behind its own capability and loudly logged, bypassing arbitration. Two limits on it, both non-negotiable: it does not override an asserted interlock claim, because an escape hatch that can defeat a thermal cutout is the absence of one; and enablement is bounded and expiring exactly like a suspension, with both edges in control_event, rather than a mode someone leaves on. It is needed when a device is misbehaving late at night, and building it deliberately is better than the alternative of someone reaching for a raw publish.

Validation before the wire

Rejection at the core, before anything is queued:

  • point is direction: out or inout
  • value type and enum membership check
  • range check against the descriptor
  • rate limit per point, from the descriptor — a physical relay has a duty cycle, and this is the last defence against a rule loop wearing out hardware

Deliverability is checked here too, but it is not a rejection. An undeliverable target means the claim is accepted and the command deferred until re-drive, never queued against a deadline waiting for an instance to come up.

Command values are supplied canonical and converted to native on the way out, using the same table as ingest, inverted. Same registry version, same affine handling. An asymmetric conversion path is a bug waiting to happen, so it is one function used in both directions.

Idempotency

idempotency_key on the intent, deduplicated over a short window. A UI double-click or a bus redelivery must not produce two commands.

The idempotent flag on the driver assignment is separate and narrower: it tells the driver whether a single transport-level retry is safe. A coil write is idempotent. A pulse or a momentary is not, and is marked so — the driver never retries it, and a lost command becomes a timeout that the reconciler handles rather than a double-fire.

What this buys the rules engine

Rules never send commands. They set claims with a value and a band, and release them when they stop applying. Everything about delivery, retry, verification, convergence and conflict happens below them.

That makes rules testable in isolation: a rule's output is a claim, and a claim is data that can be asserted against without a device in the loop.

Exit criteria

Arbitration is a pure function, so it is a fixture table. Claim set in, winner out, with no device and no bus. The cases that need covering are the ones the model exists for rather than the obvious ones:

  • withdrawal falls through to the next-highest active claim, and to nothing when there is none
  • an expired manual claim falls through without anyone withdrawing it
  • a suspension removes an owner's claims across every point it holds, and they return automatically on expiry
  • suspending an owner that holds an interlock claim is refused
  • an empty active set resolves to no_intent and the reconciler does nothing
  • a point with default_value falls through to the band-default claim instead

Interlock cold start delivers rather than times out. Assert a fail-safe interlock claim before any driver instance is ready; assert that no command is created, that convergence is undeliverable, and that the command is issued the moment the instance reaches ready. A timeout here is a failure, not a pass.

The retry budget is not consumed while undeliverable. Take a point unavailable, leave a claim standing, bring it back, and assert the bounded retry still has its full allowance. The bug this catches is a diverged that means nothing, which then hides the real one later.

Fencing. A command carrying a core_epoch below the driver's highest is rejected with stale_epoch, and the core treats that outcome as an escalation rather than a retry.

Restart does not replay writes. Commands in sent or verifying at shutdown become timeout on recovery, and the reconciler re-derives from desired versus reported. Assert that no command is re-published — this is the door-unlocked-twice case.

Every mutating path names a principal. A request without one is rejected, and every control_event written during the fixture run has principal and core_epoch populated. The NOT NULL constraints make this hard to get wrong, which is the point of putting them in the schema.


A break until November. This is the last new article here for October. The specification and the Dev Diary both pause after this part and pick up again at the start of November.

The next part - The Rules Engine. An engine kept small because almost every trigger a control system needs is already a derived point or a quality transition. What is left is a predicate over transitions, with a claim as its output.

The latest Dev Diary - Auditing the Watchdog. Build step four read against the previous part: a sentence with two halves and a test asserting the one that worked, and a stale point whose checkpoint never moved.

Start of the series: An Introduction. The map, the two decisions every later document is downstream of, and why a specification at this scale is being published in public while the system it describes gets built.

Top comments (0)