DEV Community

Cover image for MCP host Finished the Turn Blind — Token budget mid-session on the Senses Path
Moshe Avdiel
Moshe Avdiel

Posted on Originally published at github.com

MCP host Finished the Turn Blind — Token budget mid-session on the Senses Path

I have sat next to the senses wall at 21:03 while MCP host was this close to doing the wrong thing.

Not a model failure. A posture failure.

The host had started with max-tokens frozen in env / argv / a skill file. Someone on the floor said, out loud, priority is a leaf, not a restart. MCP host was still holding the old process. The only “safe” move anyone trusted was:

  1. Kill MCP host (or its MCP server)
  2. Edit a file
  3. Restart the host
  4. Lose the turn budget and the open investigation

I have watched that restart more times than I want to admit. It feels responsible. It is a ceremony. a mid-turn interrupt the host treats as noise does not wait for ceremonies.

The Aha: Token budget mid-session is not a binary you reboot. It is a function that should read live posture from Kiponos.io on every call. The host stays up. The leaf moves.

The problem: Token budget mid-session lived in the process, not in the turn

MCP host is good at calling tools. It is not born with a shared, instant, restart-free control plane.

So teams hide Token budget mid-session in the only places agent frameworks actually ship:

Where the gate hid What you restart What you lose
MCP server env / argv The MCP process Open tool sessions
Skill file on disk The agent turn, sometimes the host Context the model already paid for
Host-local JSON Whatever still has the file open Agreement between two agents
Hard-coded if on max-tokens A release The incident clock

The senses wall already knew. MCP host did not, because it had started earlier.

That is the missing piece: the framework gave you tools. It did not give you a live hub.

What teams believe

Belief Production
We'll catch it next turn The senses wall already knew this turn
Restart MCP host — it is cheap Cheap until 21:03 ate the turn budget and the open investigation
The skill file is the source of truth Skills instruct. They do not fan out
Put the SDK in the SPA Connect tokens do not belong in a browser

The Aha: local get, live write, host stays up

Kiponos holds a nested tree. Java and Python SDKs keep the latest values in memory, patched over WebSocket deltas. The hot path inside a MCP host tool is a local get — no HTTP RTT per senses lookup.

Hub leaf for this essay:

examples/agentic-dev-1030-pm-budget-turn/max-tokens = 8000
Enter fullscreen mode Exit fullscreen mode

Runnable proof: examples/java/agentic-dev-1030-pm-budget-turn

Public SDKs: Java, Python, plus React/Angular server peers (createFromEnv). Never put Connect tokens in the SPA.

Config tree (senses + peers)

examples/
  agentic-dev-1030-pm-budget-turn/
    max-tokens: 8000          # Token budget mid-session
apps/
  senses/
    live:
      max-tokens: 8000
Enter fullscreen mode Exit fullscreen mode

Integration — Java hot path

Kiponos kip = Kiponos.createForCurrentTeam();
Folder gate = kip.getRootFolder()
        .folderOrCreate("examples")
        .folderOrCreate("agentic-dev-1030-pm-budget-turn");
if (!gate.hasKey("max-tokens")) {
    gate.set("max-tokens", "8000");
}
String posture = gate.get("max-tokens");
// MCP host tool: refuse the dangerous call when posture moved
Enter fullscreen mode Exit fullscreen mode

Same leaf from a Python tool (MCP host just calls it):

from kiponos import Kiponos

k = Kiponos.connect(quiet=True)  # env: KIPONOS_ID, KIPONOS_ACCESS, KIPONOS
try:
    posture = k.get("examples/agentic-dev-1030-pm-budget-turn/max-tokens", "8000")
    if str(posture) == "8000":
        raise PermissionError("Token budget mid-session gated live — host not restarted")
finally:
    k.disconnect()
Enter fullscreen mode Exit fullscreen mode

The MCP host process does not recycle. The next tool call already sees the dashboard edit.

Real scenarios

Event Without Kiponos With Kiponos
A mid-turn interrupt the host treats as noise Restart MCP host; lose the turn budget and the open investigation Set max-tokens live; next MCP host tool call already obeys
Peer host still on old max-tokens Paste the value into the other chat One hub leaf; both processes get() locally
senses wall shows the new posture MCP host started earlier so it writes anyway Dashboard and tool share the same memory tree
Incident over, resume Another MCP host restart Set max-tokens back; session continues
Mirror phone already showing the new device leaf Two ceremonies, two lost threads Same tree, two products, no paste

Performance (this path, not a generic table)

  • MCP host tool get() is an in-process map lookup after bootstrap.
  • One WebSocket per process lifetime — not per senses line.
  • A dashboard edit is a delta of max-tokens, not a config-file reload.
  • You do not pay model tokens to “please restart MCP host.”
  • A second host converges without a third paste onto mirror phone already showing the new device leaf.

Compare to alternatives

Approach Honest fit Why it still restarts
Env file + MCP host reboot Simple at 09:00 The freeze is at 21:03
Skill markdown as policy Good instructions Not a live bus
Redis poll inside the tool Shared, but RTT on the hot path You invented a hub with worse UX
Feature-flag SaaS Product experiments Rarely session-safe for MCP host
@RefreshScope / actuator JVM apps Does not restart MCP host

When not to use Kiponos

Situation Why
Tool schema itself changed (new argument) That is a code/MCP host restart
Secret rotation of Connect tokens Credentials are not live knobs
One-off local script, no peers A hub is overkill
Browser-only “SDK in the SPA” Forbidden — tokens leak or defaults lie

What the senses operator actually said

At 21:03 someone said, out loud: priority is a leaf, not a restart. That sentence is the whole product. If it cannot land in the running MCP host process in seconds, you do not have posture. You have a wiki.

Pair max-tokens with a sister dial

max-tokens rarely moves alone on the senses wall. Pair it with a timeout, a mute, or a pause so you do not fix Token budget mid-session by inventing a second incident.

Rehearsal beats slides

In staging: set a painful max-tokens, prove MCP host recovers without a host kill, prove clamps reject nonsense, prove last-known-good when the hub is firewalled. That drill ends half the architecture arguments about Token budget mid-session.

Getting started (15 minutes)

  1. TeamPro on kiponos.io → Connect → KIPONOS_ID / KIPONOS_ACCESS / profile ['my-app']['v1.0.0']['dev']['base'].
  2. Clone github.com/kiponos-io/kiponos-io.
  3. cd examples/java/agentic-dev-1030-pm-budget-turn && cp kiponos.local.env.example kiponos.local.env
  4. ./gradlew test run — prints examples/agentic-dev-1030-pm-budget-turn/max-tokens=...
  5. In the dashboard, change max-tokens. Keep the process up. No rebuild.
  6. Point your MCP host tool at the same leaf. Do not ship a new server binary to flip Token budget mid-session.

Further reading

The moral

If flipping Token budget mid-session requires restarting MCP host, you do not have a gate. You have a hope with a process ID.

Agent frameworks already know how to call tools. Kiponos is the live hub they do not ship — so the senses wall can change its mind without killing the session.

How to try: examples/java/agentic-dev-1030-pm-budget-turn and ./gradlew test.

Top comments (0)