DEV Community

Odd_Background_328
Odd_Background_328

Posted on

Reject Unreserved Jobs Before Queue Age Breaks Slack

You get paged at 02:14 because completions stop landing.
Idle CPU on the free server looks healthy.
The projected wait is already thirty six seconds.

Your interactive slack was only thirty seconds long.
That idle capacity did not buy you time.
Replay that page locally before you trust the chart.

Read the contradiction before you touch the model

Do not start this page by swapping models.
Start with the clock that already failed you.
Queue age ate the slack before any token returned.

Utilization stayed low because the work sat unscheduled.
That split is the signal you act on.
Act on age versus slack before you touch retries.

Name the cost you actually pay

Free model access feels cheap until retries stack.
A late free server still burns your token budget.
You pay twice when the client retries blind.

You pay in wall time even when tokens are free.
Free capacity is the wrong bet without a bound.
A grant is not a seat with your name on it.

Time, tokens, and queue age move on different clocks.
A cheap token can still miss a hard deadline.
A fast CPU can still hide a stuck reservation.

Put a gate in front of the free slot

Put one admission gate in front of the worker.
Keep a local reservation ledger beside the queue.
Do not enqueue a job without a hold.

Release the hold on success, cancel, or expiry.
The free server must stay behind that gate.
One slow slot should not accept unbounded company.

Local topology

  • Run a single-process gate bound to port 8080.
  • Keep an in-memory ledger for every token hold.
  • Add a fake free-slot worker with injected delay.
  • Retry from the client only with the same key.

This is a lab drill, not a production trace.
Expected lines below are labeled, not measured live.
Nothing here is a vendor latency claim.

Declared workload

Send forty jobs, each asking for eight hundred tokens.
Keep the deadline slack fixed at thirty seconds.
Inject a free-slot service delay of twelve seconds.

Cap the client retry at one replay only.
Set the reservation TTL to twenty seconds flat.
One free slot serves one held job at a time.

Defend this threshold in the review, not a hunch.
Reject the job when projected wait exceeds remaining slack.
Do not reject work on CPU utilization alone.

Queue age tracks waiting while utilization tracks running.
Slack is the user promise you still have left.
The rationale stays simple and strictly operational here.

A busy CPU can still finish inside slack.
An idle CPU can still miss if the queue is old.
Age versus slack is the only admission test.

Ask which number you will page on tonight.
Write that number down before the next deploy.
Page on slack breach, not on a pretty CPU chart.

Reserve before the completion call

The gate below is a proposal you can run locally.
It uses only the Python standard library here.
No model name is hardcoded in this gate.

No remote quota is assumed by this process.
Projected wait equals held jobs plus one, times delay.
That formula is declared, not observed in production.

#!/usr/bin/env python3
"""Local admission gate. Proposal, not a live benchmark."""

from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import json
import time
import uuid

SLACK_S = 30.0
SERVICE_DELAY_S = 12.0
HOLD_TTL_S = 20.0
RETRY_CAP = 1
jobs = {}
holds = {}


def now():
    return time.time()


def held_count():
    return sum(1 for job in jobs.values() if job["state"] == "held")


def project_wait():
    # Declared model: one slot, each hold blocks SERVICE_DELAY_S.
    return (held_count() + 1) * SERVICE_DELAY_S


class Gate(BaseHTTPRequestHandler):
    def log_message(self, fmt, *args):
        return

    def _send(self, code, payload):
        raw = json.dumps(payload).encode()
        self.send_response(code)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(raw)))
        self.end_headers()
        self.wfile.write(raw)

    def do_POST(self):
        length = int(self.headers.get("Content-Length", "0"))
        body = json.loads(self.rfile.read(length) or b"{}")
        if self.path == "/admit":
            self._admit(body)
        elif self.path == "/complete":
            self._complete(body)
        elif self.path == "/expire":
            self._expire()
        else:
            self._send(404, {"error": "unknown"})

    def do_GET(self):
        if self.path != "/stats":
            self._send(404, {"error": "unknown"})
            return
        spent = sum(job.get("tokens_spent", 0) for job in jobs.values())
        wasted = sum(job.get("tokens_wasted", 0) for job in jobs.values())
        self._send(200, {
            "held": held_count(),
            "projected_wait_s": project_wait(),
            "slack_s": SLACK_S,
            "tokens_spent": spent,
            "tokens_wasted": wasted,
        })

    def _admit(self, body):
        key = body.get("idempotency_key") or str(uuid.uuid4())
        tokens = int(body.get("tokens", 0))
        attempt = int(body.get("attempt", 0))
        existing = jobs.get(key)
        if existing and existing["state"] in ("held", "done"):
            self._send(409, {
                "action": "reject_duplicate",
                "idempotency_key": key,
                "state": existing["state"],
                "tokens_held": existing["tokens"],
            })
            return
        if attempt > RETRY_CAP:
            self._send(429, {
                "action": "reject_retry_cap",
                "attempt": attempt,
                "retry_cap": RETRY_CAP,
            })
            return
        wait_s = project_wait()
        if wait_s > SLACK_S:
            jobs[key] = {
                "idempotency_key": key,
                "tokens": tokens,
                "attempt": attempt,
                "enqueued_at": now(),
                "state": "rejected",
                "tokens_spent": 0,
                "tokens_wasted": 0,
            }
            self._send(429, {
                "action": "reject_unreserved",
                "projected_wait_s": wait_s,
                "slack_s": SLACK_S,
                "held": held_count(),
                "reason": "projected_wait_exceeds_slack",
            })
            return
        holds[key] = {"tokens": tokens, "expires_at": now() + HOLD_TTL_S}
        jobs[key] = {
            "idempotency_key": key,
            "tokens": tokens,
            "attempt": attempt,
            "enqueued_at": now(),
            "state": "held",
            "tokens_spent": 0,
            "tokens_wasted": 0,
        }
        self._send(202, {
            "action": "reserve",
            "idempotency_key": key,
            "tokens_held": tokens,
            "projected_wait_s": wait_s,
            "hold_ttl_s": HOLD_TTL_S,
        })

    def _complete(self, body):
        key = body["idempotency_key"]
        job = jobs.get(key)
        if not job or job["state"] != "held":
            self._send(404, {"action": "missing_hold", "idempotency_key": key})
            return
        if holds.get(key, {}).get("expires_at", 0) < now():
            job["state"] = "expired"
            job["tokens_wasted"] = job["tokens"]
            holds.pop(key, None)
            self._send(409, {"action": "hold_expired", "idempotency_key": key})
            return
        spent = holds.pop(key)["tokens"]
        job["state"] = "done"
        job["tokens_spent"] = spent
        self._send(200, {
            "action": "commit",
            "idempotency_key": key,
            "tokens_spent": spent,
            "held_remaining": held_count(),
        })

    def _expire(self):
        released = 0
        wasted = 0
        stamp = now()
        for key, hold in list(holds.items()):
            if hold["expires_at"] <= stamp:
                job = jobs[key]
                job["state"] = "expired"
                job["tokens_wasted"] = hold["tokens"]
                wasted += hold["tokens"]
                holds.pop(key, None)
                released += 1
        self._send(200, {
            "action": "expire",
            "released": released,
            "tokens_wasted": wasted,
        })


def main():
    server = ThreadingHTTPServer(("127.0.0.1", 8080), Gate)
    server.serve_forever()


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Run the gate in one local shell first.
Keep a second shell for the drill commands.
Do not point this process at production traffic.

Inject the late slot you already felt

Start the gate, then send three scripted calls.
The first call should reserve inside the slack.
The second call replays the same idempotency key.

The third call should find two holds already waiting.
Each hold stands in for twelve seconds of slot time.
Two holds project thirty six seconds, which misses slack.

python3 gate.py
Enter fullscreen mode Exit fullscreen mode
curl -s -X POST http://127.0.0.1:8080/admit -H 'Content-Type: application/json' -d '{"idempotency_key":"job-17","tokens":800,"attempt":0}'
curl -s -X POST http://127.0.0.1:8080/admit -H 'Content-Type: application/json' -d '{"idempotency_key":"job-18","tokens":800,"attempt":0}'
curl -s -X POST http://127.0.0.1:8080/admit -H 'Content-Type: application/json' -d '{"idempotency_key":"job-17","tokens":800,"attempt":1}'
curl -s -X POST http://127.0.0.1:8080/admit -H 'Content-Type: application/json' -d '{"idempotency_key":"job-19","tokens":800,"attempt":0}'
curl -s http://127.0.0.1:8080/stats
Enter fullscreen mode Exit fullscreen mode

Expected output for the first call if delay stays at twelve seconds.
Treat the numbers as labeled output from this declared delay.

{"action":"reserve","idempotency_key":"job-17","tokens_held":800,"projected_wait_s":12.0,"hold_ttl_s":20.0}
Enter fullscreen mode Exit fullscreen mode

Expected output for the replay of job seventeen is a duplicate reject.
The hold must stay at eight hundred tokens.

{"action":"reject_duplicate","idempotency_key":"job-17","state":"held","tokens_held":800}
Enter fullscreen mode Exit fullscreen mode

Expected output for job nineteen is a slack reject.
The held count in that body should still read two.

{"action":"reject_unreserved","projected_wait_s":36.0,"slack_s":30.0,"held":2,"reason":"projected_wait_exceeds_slack"}
Enter fullscreen mode Exit fullscreen mode

Label that body expected, not a vendor benchmark.
Your clock and JSON spacing can shift tiny fields.
The action code is the field you should trust.

Watch fields that separate time from tokens

Emit these fields on every admission decision you make.
Skip vanity charts until these counters exist in logs.

  • Record action for reserve, reject, commit, or expiry.
  • Record projected_wait_s before you touch the server.
  • Record slack_s as the promise you still owe.
  • Record tokens_held apart from tokens already spent.
  • Record tokens_wasted when a hold expires unused.
  • Record the idempotency_key so retries cannot hide spend.

Alert when reject rate climbs while CPU stays idle.
That pairing means the free slot is late, not saturated.
Page on queue age over slack, not on token price alone.

A small spend check

Save stats JSON, then compute waste with jq.
This check is a drill helper, not a bill.
It shows amplification before finance sees a surprise.

curl -s http://127.0.0.1:8080/stats | tee stats.json
jq '{amplification: ((.tokens_spent + .tokens_wasted) / (if .tokens_spent == 0 then 1 else .tokens_spent end))}' stats.json
Enter fullscreen mode Exit fullscreen mode

Expected shape after one commit and no expiry is amplification near one.
If expiry runs first, wasted tokens lift that ratio.
A ratio above one means retries or dead holds.

Stop the client loop before you raise the cap.
Commit the useful hold before you call expire.
Otherwise the ratio blames a cleanup you chose.

Decide when the free path is the wrong bet

Use free capacity only for work that can wait.
Shed it when any line below is true.

  • The projected wait is greater than remaining slack.
  • A replay would reserve the same tokens again.
  • Hold TTL is shorter than the free-slot delay.
  • You cannot name the neighbor jobs ahead of you.
  • The caller will retry without the same key.

Interactive work should not sit on an unbounded free slot.
Batch work can wait if the ledger still has room.
A free token grant does not create a latency SLO.

Idle CPU is not evidence that your turn will start.
Unknown neighbors can consume the slot you counted on.
That is why a shared free server is a lab bet.

Decision table you can paste into the runbook

Signal Below threshold Action
Projected wait versus slack Wait is under thirty seconds Reserve and call
Same key seen again State is held or done Reject duplicate
Hold age versus TTL Hold is past twenty seconds Expire and waste
CPU versus queue age CPU idle, age over slack Shed, do not scale on CPU

Read the row before you add capacity.
Scaling an idle slot does not shrink queue age.
Shedding the late job does shrink that age.

Where a free assistant fits this drill

Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode is an open-source coding assistant you can run in this drill.
The operator states free model access and a free server option.

The operator note states an allowance of ten million tokens.
Treat that allowance as a grant, not a reserved capacity contract.
Terms can change before you plan a quarter on them.

This draft does not claim a measured quota, duration, or hardware shape.
Use the free server to rehearse this gate.
Do not use it as the only path for a deadline.

If the free slot stalls, your ledger should already have said no.
Offers change, so re-read the terms on the day you drill.
Do not paste a screenshot of idle CPU into the budget review.

Bring the reject reason and the projected wait instead.
Those two fields survive a model swap.
A chart of idle CPU will not.

Fail, clean up, and roll back

When the gate rejects, do not blind-retry on the client.
Return the 429 to the caller with the reason code.
Keep the idempotency key so a later retry stays single-spend.

Expire holds on a timer so a crash cannot leak budget.
Call the expire path before you trust stats after a stall.
A leaked hold looks like load and spends nothing useful.

Follow these rollback steps after every local drill.
Do not skip the closed-port check after you stop.

  1. Stop gate.py with Ctrl-C in the drill shell.
  2. Drop the in-memory ledger by exiting the process.
  3. Set SERVICE_DELAY_S back to the declared twelve seconds.
  4. Confirm port 8080 is closed before the next run.
  5. Do not leave a retry loop pointed at a shared endpoint.

If you later front a real queue, ship the gate beside it.
Feature-flag admission so you can fail open only for batch.
Fail closed for interactive work when slack is already gone.

Rollback is the flag, not a model swap at 02:14.
Practice the flag flip in staging before you need it.
A runbook line you have not drilled is a wish.

Who should not use this gate

Skip this gate if your jobs have no deadline.
Skip it if every call is unique and never retried.
Skip it if you already reserve capacity with a hard quota.

Skip it if you need multi-node fairness this week.
This ledger is local and dies with the process.
A restart forgets holds unless you add durable storage.

Do not use it to hide spend from finance.
Do not use it to bypass provider rate limits.
Do not treat expected JSON as a production SLO.

A single process will not save a region-wide stall.
Clock skew between shells can fake a late hold.
Declare one clock source before you trust expiry.

Close on the action, not the model

You saw idle CPU and a missed deadline together.
The action is reject unreserved work, not add retries.
Queue age versus slack is the threshold to write down.

Token holds stop a replay from spending the grant twice.
Free capacity stays a lab bet until that gate holds.
Confirm current free-server terms, then run this laptop drill.

Top comments (0)