I have sat next to the mirror-phone wall at 18:45 while Claude Code was this close to doing the wrong thing.
Not a model failure. A posture failure.
The host had started with incident-pause frozen in env / argv / a skill file. Someone on the floor said, out loud, the phone already knew the posture — the host did not. Claude Code was still holding the old process. The only “safe” move anyone trusted was:
- Kill Claude Code (or its MCP server)
- Edit a file
- Restart the host
- Lose live device context and the current tool call
I have watched that restart more times than I want to admit. It feels responsible. It is a ceremony. device leaf gone stale while the agent is mid-turn does not wait for ceremonies.
The Aha: SRE degradation pause as live posture on the mirror-phone wall is not a binary you reboot. It is a function that should read live posture from Kiponos.io on every call. The host stays up. The leaf moves.
The problem: SRE degradation pause as live posture on the mirror-phone wall lived in the process, not in the turn
Claude Code is good at calling tools. It is not born with a shared, instant, restart-free control plane.
So teams hide SRE degradation pause as live posture on the mirror-phone wall in the only places agent frameworks actually ship:
| Where the gate hid | What you restart | What you lose |
|---|---|---|
| MCP server env / argv | The MCP process | Open tool sessions |
| Skill file on disk | The agent turn, sometimes the host | Context the model already paid for |
| Host-local JSON | Whatever still has the file open | Agreement between two agents |
Hard-coded if on incident-pause
|
A release | The incident clock |
The mirror-phone wall already knew. Claude Code did not, because it had started earlier.
That is the missing piece: the framework gave you tools. It did not give you a live hub.
What teams believe
| Belief | Production |
|---|---|
| We'll catch it next turn | The mirror-phone wall already knew this turn |
| Restart Claude Code — it is cheap | Cheap until 18:45 ate live device context and the current tool call |
| The skill file is the source of truth | Skills instruct. They do not fan out |
| Put the SDK in the SPA | Connect tokens do not belong in a browser |
The Aha: local get, live write, host stays up
Kiponos holds a nested tree. Java and Python SDKs keep the latest values in memory, patched over WebSocket deltas. The hot path inside a Claude Code tool is a local get — no HTTP RTT per mirror phone lookup.
Hub leaf for this essay:
examples/sre-agent-pause-mesh/incident-pause = paused
Runnable proof: examples/java/sre-agent-pause-mesh
Public SDKs: Java, Python, plus React/Angular server peers (createFromEnv). Never put Connect tokens in the SPA.
Config tree (mirror phone + peers)
examples/
sre-agent-pause-mesh/
incident-pause: paused # SRE degradation pause as live posture on the mirror-phone wall
apps/
mirror-phone/
live:
incident-pause: paused
Integration — Java hot path
Kiponos kip = Kiponos.createForCurrentTeam();
Folder gate = kip.getRootFolder()
.folderOrCreate("examples")
.folderOrCreate("sre-agent-pause-mesh");
if (!gate.hasKey("incident-pause")) {
gate.set("incident-pause", "paused");
}
String posture = gate.get("incident-pause");
// Claude Code tool: refuse the dangerous call when posture moved
Same leaf from a Python tool (Claude Code just calls it):
from kiponos import Kiponos
k = Kiponos.connect(quiet=True) # env: KIPONOS_ID, KIPONOS_ACCESS, KIPONOS
try:
posture = k.get("examples/sre-agent-pause-mesh/incident-pause", "paused")
if str(posture) == "paused":
raise PermissionError("SRE degradation pause as live posture on the mirror-phone wall gated live — host not restarted")
finally:
k.disconnect()
The Claude Code process does not recycle. The next tool call already sees the dashboard edit.
Real scenarios
| Event | Without Kiponos | With Kiponos |
|---|---|---|
| Device leaf gone stale while the agent is mid-turn | Restart Claude Code; lose live device context and the current tool call | Set incident-pause live; next Claude Code tool call already obeys |
| Peer host still on old incident-pause | Paste the value into the other chat | One hub leaf; both processes get() locally |
| mirror-phone wall shows the new posture | Claude Code started earlier so it writes anyway | Dashboard and tool share the same memory tree |
| Incident over, resume | Another Claude Code restart | Set incident-pause back; session continues |
| Admin dashboard showing paused while the agent still writes | Two ceremonies, two lost threads | Same tree, two products, no paste |
Performance (this path, not a generic table)
- Claude Code tool
get()is an in-process map lookup after bootstrap. - One WebSocket per process lifetime — not per mirror phone line.
- A dashboard edit is a delta of
incident-pause, not a config-file reload. - You do not pay model tokens to “please restart Claude Code.”
- A second host converges without a third paste onto admin dashboard showing paused while the agent still writes.
Compare to alternatives
| Approach | Honest fit | Why it still restarts |
|---|---|---|
| Env file + Claude Code reboot | Simple at 09:00 | The freeze is at 18:45 |
| Skill markdown as policy | Good instructions | Not a live bus |
| Redis poll inside the tool | Shared, but RTT on the hot path | You invented a hub with worse UX |
| Feature-flag SaaS | Product experiments | Rarely session-safe for Claude Code |
@RefreshScope / actuator |
JVM apps | Does not restart Claude Code |
When not to use Kiponos
| Situation | Why |
|---|---|
| Tool schema itself changed (new argument) | That is a code/Claude Code restart |
| Secret rotation of Connect tokens | Credentials are not live knobs |
| One-off local script, no peers | A hub is overkill |
| Browser-only “SDK in the SPA” | Forbidden — tokens leak or defaults lie |
Rehearsal beats slides
In staging: set a painful incident-pause, prove Claude Code recovers without a host kill, prove clamps reject nonsense, prove last-known-good when the hub is firewalled. That drill ends half the architecture arguments about SRE degradation pause as live posture on the mirror-phone wall.
Why Claude Code is the wrong restart target
Claude Code is good at calling tools. It is not a control plane. Killing it to flip incident-pause teaches the on-call that judgment requires a process ID. The mirror-phone wall already disagrees.
What the mirror phone operator actually said
At 18:45 someone said, out loud: the phone already knew the posture — the host did not. That sentence is the whole product. If it cannot land in the running Claude Code process in seconds, you do not have posture. You have a wiki.
Getting started (15 minutes)
- TeamPro on kiponos.io → Connect →
KIPONOS_ID/KIPONOS_ACCESS/ profile['my-app']['v1.0.0']['dev']['base']. - Clone github.com/kiponos-io/kiponos-io.
cd examples/java/sre-agent-pause-mesh && cp kiponos.local.env.example kiponos.local.env-
./gradlew test run— printsexamples/sre-agent-pause-mesh/incident-pause=... - In the dashboard, change
incident-pause. Keep the process up. No rebuild. - Point your Claude Code tool at the same leaf. Do not ship a new server binary to flip SRE degradation pause as live posture on the mirror-phone wall.
Further reading
The moral
If flipping SRE degradation pause as live posture on the mirror-phone wall requires restarting Claude Code, you do not have a gate. You have a hope with a process ID.
Agent frameworks already know how to call tools. Kiponos is the live hub they do not ship — so the mirror-phone wall can change its mind without killing the session.
How to try: examples/java/sre-agent-pause-mesh and ./gradlew test.
Top comments (0)