DEV Community

Dakota Lin
Dakota Lin

Posted on

I Kept the Run Index. Warmup Took the Win.

The second run is not a real speedup. It is a warm socket with a long memory. I kept the run-index graph, and warmup took the credit.

You changed one small line in the client. The next call came back and felt fast. Did the model improve, or did the pipe stay open?

I stopped trusting the warm feeling in my gut. I timed every loop pass in strict run order. I labeled cold and warm before I touched the patch.

This note is a method, not a trophy case. The script below is a proposal you can run. I will not invent a chart from a wish.

The graph I kept

What should that kept graph even show you? Not one proud bar you saved after lunch. It should show latency against the run index.

Run zero is the stranger at the door. It pays for DNS, TLS, and a cold import. Run one may still be kindling, not fire.

I drop those two from the comparison on purpose. I decide that rule before I see the bars. If I pick the rule after, I am shopping.

A free server is useful here for one reason. Your laptop fan is a liar with a motor. It speeds up, then it thermal-throttles, then it lies again.

I would point the same harness at MonkeyCode for repeats. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Free model access is what lets that loop repeat.

A free server option keeps the laptop fan out. I am not naming a model, a quota, or a box. Those facts move, and I will not freeze a rumor.

Read the live page before you trust a limit. A free pool can queue you behind strangers. That wait is not your patch, and it is not a product score.

The loop you can run

Here is the harness I actually want to run. It is stdlib only, so you can audit every line. Point it at a URL you are allowed to call.

#!/usr/bin/env python3
"""Label cold versus warm, and keep the run index."""

import argparse
import csv
import json
import os
import time
import urllib.request

def one_call(url: str, payload: bytes, timeout: float) -> dict:
    wall_start = time.perf_counter_ns()
    cpu_start = time.process_time_ns()
    req = urllib.request.Request(
        url,
        data=payload,
        headers={"Content-Type": "application/json"},
        method="POST",
    )
    with urllib.request.urlopen(req, timeout=timeout) as resp:
        body = resp.read()
        status = resp.status
    wall_ns = time.perf_counter_ns() - wall_start
    cpu_ns = time.process_time_ns() - cpu_start
    return {
        "status": status,
        "wall_ms": wall_ns / 1e6,
        "cpu_ms": cpu_ns / 1e6,
        "bytes_out": len(payload),
        "bytes_in": len(body),
    }

def phase_for(index: int) -> str:
    if index == 0:
        return "cold"
    if index == 1:
        return "kindling"
    return "warm"

def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("--url", required=True)
    parser.add_argument("--runs", type=int, default=12)
    parser.add_argument("--out", default="runs.csv")
    parser.add_argument("--timeout", type=float, default=60.0)
    args = parser.parse_args()
    payload = json.dumps({
        "prompt": "Reply with one short sentence about queues.",
        "max_tokens": 32,
    }).encode()
    rows = []
    for i in range(args.runs):
        row = one_call(args.url, payload, args.timeout)
        row["run_index"] = i
        row["phase"] = phase_for(i)
        row["label"] = os.environ.get("RUN_LABEL", "baseline")
        rows.append(row)
    fields = [
        "label", "run_index", "phase", "status",
        "wall_ms", "cpu_ms", "bytes_out", "bytes_in",
    ]
    with open(args.out, "w", newline="") as fh:
        writer = csv.DictWriter(fh, fieldnames=fields)
        writer.writeheader()
        writer.writerows(rows)
    warm = [r["wall_ms"] for r in rows if r["phase"] == "warm"]
    cold = rows[0]["wall_ms"]
    print(f"cold_wall_ms={cold:.1f}")
    if warm:
        mid = sorted(warm)[len(warm) // 2]
        print(f"warm_median_ms={mid:.1f}")
        print(f"cold_over_warm={cold / mid:.2f}")
    print("kept_graph=wall_ms by run_index, split on label")

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

The JSON body is a stand-in, not a vendor schema. Swap fields to match the server you actually call. Treat every printed number as unexecuted until you run it.

Set a label, then run the baseline before the patch. Change one thing, then run the patched label. Same URL, same payload, and the same run count.

export RUN_LABEL=baseline
python3 harness.py --url "$MODEL_URL" --runs 12 --out baseline.csv
export RUN_LABEL=patched
python3 harness.py --url "$MODEL_URL" --runs 12 --out patched.csv
python3 sketch.py baseline.csv patched.csv
Enter fullscreen mode Exit fullscreen mode

The sketch script is the graph I keep. Each row is a run, not a vibe. Cold rows stay on the page so I cannot forget them.

#!/usr/bin/env python3
"""Sketch wall time by run index from your own CSV."""

import csv
import sys

def load(path: str) -> list:
    with open(path, newline="") as fh:
        return list(csv.DictReader(fh))

def main() -> None:
    rows = []
    for path in sys.argv[1:]:
        rows.extend(load(path))
    if not rows:
        raise SystemExit("pass at least one csv")
    peak = max(float(r["wall_ms"]) for r in rows) or 1.0
    for r in rows:
        n = int(round(40 * float(r["wall_ms"]) / peak))
        bar = "#" * max(n, 1)
        label = r["label"][:8]
        idx = int(r["run_index"])
        phase = r["phase"][:8]
        print(f"{label:8} {idx:02d} {phase:8} {bar}")

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

I omit sample bars so nobody copies a fake chart. Run the sketch yourself, then argue with your own rows.

How do I read that sketch without lying? I cover the warm rows with one hand. I ask whether the cold row moved at all.

If only the warm rows shrink, I blame the socket. Connection reuse will do that trick every time. A prompt cache will smile the same way.

If cold and warm both shrink after the patch, I lean in. I still want a second hour, not a second guess. One quiet minute is not a whole season.

Watch the CPU column while you stare at the wall. If CPU stays tiny while wall grows, you waited. You did not win a secret compute prize.

That split is a guardrail, not the whole story. I already know tails can hide in the network. Today I care who gets credit for the drop.

Do not average the cold row into the warm rows. The mean will hide the stranger and bless the regulars. I want the run index, so order stays visible.

A median of warm rows is fine as a note. It is not fine as a public headline. Headlines eat the label, then they eat the truth.

I also keep bytes out and bytes in. A shorter reply is not a faster client. Did you change the work, or did you change the scale?

If bytes_in falls and wall falls with it, stop. You shrank the job, and then you clapped. That is a diet, not a faster kitchen.

Same payload both times, or the graph is theater. I pin the prompt text in the script for that reason. I do not improve the prompt between the labels.

What about retries, those quiet little cheats of time? A retry can turn one miss into two bills of time. I log status, and I do not drop errors from the file.

A failed row stays in the sketch as a gap marker. Deleting it makes the warm median look brave. Would you delete a crash from a race recap?

Timeouts need a pre-set budget before you start. I pass a timeout on purpose, and I write it down. If I raise it mid-test, I started a new experiment.

The free server can sleep between your bursts. A pause for coffee becomes a new cold start. I note the gap, or I pretend the room never cooled.

Shared capacity is a neighbor, not a lab bench. Someone else's job can sit in your wall time. That is why this loop cannot rank vendors.

Who should walk away

I would not use this method for a public bake-off. I would not use it to promise a speedup in a launch note. I would not use it if I need pinned cores and a quiet NIC.

You should skip it if your boss wants a single number. You should skip it if you cannot store the raw CSV. You should skip it if the payload is still changing shape.

Students can still learn the habit from the shape. Run it once against a local stub if you have no server. The stub will not prove a model, but it will prove your labels.

A local stub is a metronome, not an opponent. It teaches you whether run zero looks different. If run zero looks the same, your timer is blind.

Here is a tiny stub you can hate in private. It sleeps longer on the first hit, then less. That sleep is fake, and the lesson is not.

#!/usr/bin/env python3
"""Local metronome for labels. First hit sleeps longer."""

from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import time

HITS = {"n": 0}

class Handler(BaseHTTPRequestHandler):
    def do_POST(self) -> None:
        n = HITS["n"]
        HITS["n"] = n + 1
        delay = 0.4 if n == 0 else 0.05
        time.sleep(delay)
        body = b'{"ok":true}'
        self.send_response(200)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(body)))
        self.end_headers()
        self.wfile.write(body)

    def log_message(self, fmt: str, *args) -> None:
        return

if __name__ == "__main__":
    ThreadingHTTPServer(("127.0.0.1", 8765), Handler).serve_forever()
Enter fullscreen mode Exit fullscreen mode

The stub listens on port 8765 on localhost only. Start it in one shell, then aim the harness there. Call that stub before you touch a remote pool.

If your sketch cannot see the fake cold hit, stop. Your graph is broken, and the model is innocent.

I like that order more than a hot take. I put tooling before talent, every single time. Labels come before pride, or the bar lies.

When the stub looks honest, I move the URL outward. I keep the same CSV columns, so the eye does not relearn. A new column is a new argument, and I am tired.

How many runs is enough for a note like this? Twelve runs make a sketch, not a paper. If the bars jump like popcorn, I add runs.

I do not add adjectives to a noisy sketch. I also refuse to tune after I see a winner. Tuning the timeout to flatter the patch is a costume.

Take the costume off, or do not publish the photo. Store the script hash next to the CSV if you can. A one-line drift will impersonate a real victory.

I want the blame to land on a commit, not on a mood.

git rev-parse --short HEAD | tee run_commit.txt
python3 -c "import hashlib, pathlib; p=pathlib.Path('harness.py'); print(hashlib.sha256(p.read_bytes()).hexdigest())" | tee run_files.txt
Enter fullscreen mode Exit fullscreen mode

That pair is my alibi when the bar looks pretty. Without it, future me will argue with past me. Past me cheats when the bar looks good.

Limitations sit in the room even when the sketch looks clean. Free model access can change shape without a speech. Free server time can queue, cap, or vanish for an hour.

I am not claiming those options are permanent. I am not claiming a region, a chip, or a rate. If the operator page disagrees with this note, believe the page.

Clock resolution can flatter a tiny client change. A sub-millisecond win on a network hop is noise. Do not frame noise and hang it in the hall.

JSON parsing can dwarf a small logic tweak. If you only time the function you edited, you will miss the dump. Time the whole call, then decide where to cut.

I still would not start a rewrite from one sketch. I would start with a sharper question instead. Which run got faster, and which run only got familiar?

Keep the ugly cold bar in the picture. Cropping it is how warmup steals the win. That theft is common, and I am done applauding it.

If you need a place to repeat the loop, start at MonkeyCode. Read the live limits, then keep the run-index graph.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to