DEV Community

Cover image for Inference Batch Size Should Not Wait for a Redeploy — Fintech Live Posture
Moshe Avdiel
Moshe Avdiel

Posted on • Originally published at github.com

Inference Batch Size Should Not Wait for a Redeploy — Fintech Live Posture

The Aha: batchSize is not a property file trophy. It is incident posture — and posture that waits for a jar is already late.

A flash sale needs the weight raised now. Last Tuesday's properties file is not a strategy.

Domain: payments, fraud, ledger. This essay maps hub key batchSize to inference batch size so the lesson stays concrete for fintech operators.

The problem: ceremony between judgment and effect

You already know the right number. Everyone in the war room knows the right number. What you do not have is a path from mouth → running process that is shorter than a release train.

When inference batch size is frozen in YAML, every incident becomes a process argument. When it lives in a hub with clamps, the argument ends and the work begins.

Belief Production
Defaults are fine Defaults become root causes under load
The mesh owns that Mesh freezes share to another YAML dialect
Flags cover this Second system, second delay, second outage mode
We'll hotfix Hotfix is still CI + roll + nerves

The Aha: local read, live write

Kiponos.io holds the tree. The Python SDK keeps the latest value in memory, patched over WebSocket deltas. Hot path: local get — no per-request hub RTT.

examples/
  ops-fintech-inference-batch/
    batchSize: 16   # inference batch size
    hardMax: compiled-in-app
    failClosed: true
Enter fullscreen mode Exit fullscreen mode
policy = kiponos.path("examples", "ops-fintech-inference-batch")
batchSize = min(int(policy.get("batchSize", 16)), HARD_MAX)
worker_pool.resize(batchSize)  # or admission semaphore
Enter fullscreen mode Exit fullscreen mode

Ops sets batchSize in the dashboard (or automation writes the same path). The next evaluation uses the new value. Same jar. Same tests for structure.

What stays in the jar vs the hub

Jar (versioned) Hub (live)
Code paths & clamps Operational numbers
Hard maxima / allowlists Current posture
Schema & types Human judgment under pressure
Fail-closed defaults Temporary incident overrides

Architecture

Dashboard / automation ──write──► Kiponos hub tree
                                      │ WebSocket delta
                                      ▼
                               SDK in-process cache
                                      │ local get
                                      ▼
                               Hot path decision (inference batch size)
Enter fullscreen mode Exit fullscreen mode

No sidecar tax on every request. No second product for "just this one dial."

Clone and learn the pattern

git clone https://github.com/kiponos-io/kiponos-io.git
# See examples/java/* for runnable Super Pattern / Aha modules
# Profile: ['app']['release']['env']['config'] — same shape as production
Enter fullscreen mode Exit fullscreen mode

Getting started: GETTING-STARTED.md · Product: kiponos.io

Scenarios

Moment Frozen YAML Live hub
Incident PR + pipeline Seconds
Peak event Over-provision Dial down/up
Experiment Long-lived branch Same jar
Rollback Redeploy previous Revert hub value
Region skew Copy three files Per-folder values

When not to live-edit

  • Protocol or schema changes that need coordinated rollouts
  • Values that compliance requires code-reviewed only
  • Anything you cannot clamp or allowlist safely
  • Secrets (use a secret manager — never the ops posture tree)

Live knobs are for posture, not for inventing untested systems under fire.

Operational checklist

  1. Name the hub path so humans find it under pressure (examples/ops-fintech-inference-batch/batchSize).
  2. Default safely when the hub is unreachable (fail closed on money paths).
  3. Allowlist writers (dashboard roles + automation identities).
  4. Log the decision, not every get.
  5. Rehearse the flip in staging with a sibling example module.
  6. Document the one-line kill path (revert key).
  7. Record from→to + reason code in the incident timeline.

Why this is not "just another flag"

Feature flags are often product gates. This essay is about ops posture on a hot path: inference batch size for fintech — numbers humans already change verbally in war rooms.

Kiponos makes that verbal decision executable without a second control plane tax on every request.

Rehearsal beats slides

In staging: set a painful value, prove recovery without restart, prove clamps reject nonsense, prove LKG when hub is firewalled. That drill ends half the architecture arguments.

What never goes live without review

Protocol/schema changes, crypto material, and legal freezes stay in code review. Posture numbers war rooms already shout belong in the hub with clamps.

Guardrails that keep compliance calm

Live does not mean unbounded. Compile a hard max. Allowlist writers. Audit actor, old, new, ticket. Fail closed on money paths when the hub is dark.

A note on testing

Unit-test structure with fixed strings (no network). Integration-test the hub path against the public sandbox when you can.

Good tests:

  • Defaults when keys are missing
  • Clamps reject out-of-range values
  • Fail-closed behavior for money paths

Bad tests:

  • Hitting production hubs from CI
  • Asserting wall-clock times for WebSocket delivery

Closing note

Architecture diagrams do not absorb incidents. Steerable posture does — with audit, clamps, and a revert path written before you need it.

One-line runbook

Who may move this key under P1, what is the clamp, what is the revert? Write that sentence before you need it. Posture without a revert path is just another outage mode.

Moral

Posture beats ceremony — every time.

Ship judgment. Leave the jar alone.


Series: Kiponos live ops posture · Pattern library: kiponos-io/docs · SDK examples: examples/java

Top comments (0)