DEV Community

Dakota Ma
Dakota Ma

Posted on

Do Not Grade a Closed Socket

A closed socket is an infrastructure event, and it must never become a task score in the eval ledger. When those two events share one column, the next comparison blames the prompt for a failure the model did not author. Separate transport status, schema status, and task score before any baseline diff is allowed to run. That separation is the practical rule, and the sections below show one small harness that enforces it.

Golden cases stay trustworthy only when a missed connection cannot overwrite an accepted task score from an earlier run. A free endpoint suits shadow traffic because an extra pass can be cheap, yet uptime is not a property of the prompt. If the client writes a timeout as zero, a short outage hardens into a false regression that later reviewers will treat as real. The remedy is clerical before it is statistical, because the outcome class has to exist before a numeric grade is legal.

Where a free shadow lane belongs

MonkeyCode fits this workflow only as a place to obtain free model access and a free server for shadow runs that stay off the baseline. Disclosure: This article was prepared as part of MonkeyCode's product outreach, so the relationship is stated before any method depends on the product. This note names neither a model nor a quota, since those commercial terms were not verified for this draft. Readers should confirm the live offer before a scheduled job is pointed at any free endpoint.

Think of the harness as a customs desk that uses three stamps rather than one shared stamp. The parcel may be refused at the gate, opened and found mislabeled, or accepted and then checked against the invoice. A timeout is a gate refusal, a missing answer key is a labeling fault, and a wrong field value is an invoice fault. Mixing those stamps causes the weekly report to describe road conditions instead of the cargo itself.

A silent regression, in this ledger, is a narrower event than a red row on a dashboard. The fixture hash must match, both transport and schema must pass, and the task score must fall for that same model alias. A lower score under a different alias is a model-swap delta, which is useful, but it is not evidence that the prompt drifted. A lower score beside a changed fixture hash is a case edit, and it should be reviewed as a test change rather than a deployment bug.

Classify before you diff

The artifact below is a proposal, and it has not been executed against a live model or a live server. It stores one outcome row per case and refuses to invent a task score when the call fails. The baseline file remains untouched until a separate command explicitly promotes one fully accepted candidate run. Pass your own wrapper into run_case, and make that wrapper raise TimeoutError instead of hiding the outage inside a normal payload.

# Unexecuted proposal. Requires Python 3.10+. Not a live benchmark.
from __future__ import annotations

import hashlib
import json
import sys
import time
from dataclasses import dataclass
from typing import Any, Callable

TRANSPORT_OK = 'transport_ok'
TRANSPORT_TIMEOUT = 'transport_timeout'
TRANSPORT_ERROR = 'transport_error'
SCHEMA_OK = 'schema_ok'
SCHEMA_FAIL = 'schema_fail'

@dataclass(frozen=True)
class GoldenCase:
    case_id: str
    prompt: str
    expect: dict[str, str]
    fixture_sha: str

@dataclass
class Outcome:
    case_id: str
    fixture_sha: str
    model_alias: str
    response_id: str | None
    transport: str
    schema: str
    task_score: float | None
    detail: str
    elapsed_ms: int

def fixture_sha(payload: dict[str, Any]) -> str:
    body = {key: value for key, value in payload.items() if key != 'fixture_sha'}
    raw = json.dumps(body, sort_keys=True, separators=(',', ':')).encode()
    return hashlib.sha256(raw).hexdigest()[:16]

def grade_task(answer: dict[str, Any], expect: dict[str, str]) -> tuple[float, str]:
    keys = sorted(expect)
    if not keys:
        return 0.0, 'empty_expect'
    hits = sum(1 for key in keys if answer.get(key) == expect[key])
    missing = [key for key in keys if key not in answer]
    detail = 'ok' if hits == len(keys) else 'missing=' + ','.join(missing)
    return hits / len(keys), detail

def run_case(
    case: GoldenCase,
    alias: str,
    timeout_s: float,
    client: Callable[[str, str, float], dict[str, Any]],
) -> Outcome:
    started = time.perf_counter()
    try:
        raw = client(alias, case.prompt, timeout_s)
    except TimeoutError:
        return _closed(case, alias, started, TRANSPORT_TIMEOUT, 'timeout')
    except OSError as exc:
        return _closed(case, alias, started, TRANSPORT_ERROR, type(exc).__name__)
    elapsed = int((time.perf_counter() - started) * 1000)
    response_id = raw.get('id') if isinstance(raw, dict) else None
    answer = raw.get('answer') if isinstance(raw, dict) else None
    if not isinstance(answer, dict):
        return Outcome(
            case.case_id, case.fixture_sha, alias, response_id,
            TRANSPORT_OK, SCHEMA_FAIL, None, 'missing_answer', elapsed,
        )
    score, detail = grade_task(answer, case.expect)
    return Outcome(
        case.case_id, case.fixture_sha, alias, response_id,
        TRANSPORT_OK, SCHEMA_OK, score, detail, elapsed,
    )

def _closed(case: GoldenCase, alias: str, started: float, transport: str, detail: str) -> Outcome:
    elapsed = int((time.perf_counter() - started) * 1000)
    return Outcome(
        case.case_id, case.fixture_sha, alias, None,
        transport, SCHEMA_FAIL, None, detail, elapsed,
    )

def classify_delta(base: Outcome, cand: Outcome) -> str:
    if base.fixture_sha != cand.fixture_sha:
        return 'fixture_move'
    if base.model_alias != cand.model_alias:
        return 'alias_swap'
    if cand.transport != TRANSPORT_OK:
        return 'transport_miss'
    if cand.schema != SCHEMA_OK or cand.task_score is None:
        return 'schema_miss'
    if base.task_score is None:
        return 'baseline_unscored'
    if cand.task_score < base.task_score:
        return 'silent_regression'
    if cand.task_score > base.task_score:
        return 'score_rise'
    return 'unchanged'

def summarize(pairs: list[tuple[Outcome, Outcome]]) -> dict[str, int]:
    counts = {
        'transport_miss': 0,
        'schema_miss': 0,
        'silent_regression': 0,
        'other': 0,
    }
    for base, cand in pairs:
        kind = classify_delta(base, cand)
        if kind in counts:
            counts[kind] += 1
        else:
            counts['other'] += 1
    return counts

def pin_path(path: str) -> None:
    with open(path, encoding='utf-8') as handle:
        for line in handle:
            if not line.strip():
                continue
            data = json.loads(line)
            data['fixture_sha'] = fixture_sha(data)
            print(json.dumps(data, sort_keys=True))

def diff_paths(baseline: str, candidate: str) -> dict[str, int]:
    base, cand = _read_outcomes(baseline), _read_outcomes(candidate)
    shared = sorted(set(base) & set(cand))
    return summarize([(base[key], cand[key]) for key in shared])

def _read_outcomes(path: str) -> dict[str, Outcome]:
    rows: dict[str, Outcome] = {}
    with open(path, encoding='utf-8') as handle:
        for line in handle:
            if not line.strip():
                continue
            data = json.loads(line)
            rows[data['case_id']] = Outcome(**data)
    return rows

def main(argv: list[str]) -> int:
    if len(argv) == 3 and argv[1] == 'pin':
        pin_path(argv[2])
        return 0
    if len(argv) == 4 and argv[1] == 'diff':
        print(json.dumps(diff_paths(argv[2], argv[3]), sort_keys=True))
        return 0
    print('usage: python harness.py pin FILE', file=sys.stderr)
    print('       python harness.py diff BASE CAND', file=sys.stderr)
    return 2

if __name__ == '__main__':
    raise SystemExit(main(sys.argv))
Enter fullscreen mode Exit fullscreen mode

The admission rule is the line that protects the baseline from empty sockets and from partial bodies. A row may enter a score comparison only when transport is clean, the schema check passed, and the task score is an actual number. Everything else remains in the ledger for operations review, but the diff command skips it when it computes prompt regressions. The point is which rows are scoreable at all, not the price of a rerun and not a blanket ban on promotion.

Each golden case carries an identifier, a prompt, an expectation object, and a hash of the canonical JSON. The hash pins the prompt and the expected object, so either edit changes the case while the identifier stays friendly. Expectations should stay small and typed, because exact equality on free prose will flap on harmless wording changes. This example expects a flat object of strings, and the grader scores matching keys without credit for extra keys the model invented.

{"case_id":"city-07","prompt":"Return a JSON object with city and country for the capital of France.","expect":{"city":"Paris","country":"France"}}
{"case_id":"city-09","prompt":"Return a JSON object with city and country for the capital of Japan.","expect":{"city":"Tokyo","country":"Japan"}}
Enter fullscreen mode Exit fullscreen mode

Pin the case file to stdout first, and keep that pinned copy beside the code rather than inside the accepted baseline path. A second command diffs the candidate ledger against the promoted baseline and counts only the shared case identifiers. The printer should show transport misses, schema misses, true task drops, and a remainder bucket, so each pile stays named. If transport misses dominate, the route is the suspect, and if task drops dominate on a stable hash, the prompt change is the suspect.

python harness.py pin cases/cities.jsonl > cases/cities.pinned.jsonl
python harness.py diff runs/accepted.jsonl runs/shadow.jsonl
Enter fullscreen mode Exit fullscreen mode

Suppose the accepted baseline holds a score of one for case city-07 under alias primary, with a placeholder fixture hash. The shadow file shows a timeout and a null score on that same hash, so the diff prints a transport miss and opens no regression. A second row shows a clean transport, a clean schema, the same hash, and a score drop under alias primary, which is the silent regression. A third row with a new hash and a lower score is printed as a fixture move, because the test itself changed under the reviewer.

mkdir -p runs
cat > runs/accepted.jsonl <<'EOF'
{"case_id":"city-07","fixture_sha":"placeholder-hash","model_alias":"primary","response_id":"resp-base-07","transport":"transport_ok","schema":"schema_ok","task_score":1.0,"detail":"ok","elapsed_ms":500}
{"case_id":"city-09","fixture_sha":"placeholder-hash","model_alias":"primary","response_id":"resp-base-09","transport":"transport_ok","schema":"schema_ok","task_score":1.0,"detail":"ok","elapsed_ms":480}
{"case_id":"city-11","fixture_sha":"old-hash","model_alias":"primary","response_id":"resp-base-11","transport":"transport_ok","schema":"schema_ok","task_score":1.0,"detail":"ok","elapsed_ms":510}
EOF
cat > runs/shadow.jsonl <<'EOF'
{"case_id":"city-07","fixture_sha":"placeholder-hash","model_alias":"primary","response_id":null,"transport":"transport_timeout","schema":"schema_fail","task_score":null,"detail":"timeout","elapsed_ms":8001}
{"case_id":"city-09","fixture_sha":"placeholder-hash","model_alias":"primary","response_id":"resp-shadow-09","transport":"transport_ok","schema":"schema_ok","task_score":0.0,"detail":"missing=city","elapsed_ms":640}
{"case_id":"city-11","fixture_sha":"new-hash","model_alias":"primary","response_id":"resp-shadow-11","transport":"transport_ok","schema":"schema_ok","task_score":0.5,"detail":"missing=country","elapsed_ms":700}
EOF
python harness.py diff runs/accepted.jsonl runs/shadow.jsonl
Enter fullscreen mode Exit fullscreen mode

On these illustrative rows the classifier should count one transport miss, one silent regression, and one remainder for the fixture move. That remainder is city-11, whose hash changed, so a lower score there is not allowed to wear the regression label. Nothing in that count is a measurement of a model, a server, or a quota, and it only checks that the three piles stay apart.

Alias drift is not prompt drift

Provider aliases move without a release note, and a friendly name can point at a different backend next week. Store the alias you requested and any response identifier the API actually returned, then treat a mismatch between those strings as its own column. Without that pair, a silent model substitution looks exactly like a prompt regression, because the fixture hash and the case identifier did not change. The free server does not remove this duty, since a convenient endpoint can still rotate what sits behind a stable name.

Limits and non-uses

This grader is intentionally weak, and it will not judge explanations, tone, or partial numeric tolerance. A client that swallows timeouts and returns error text will be filed as a schema miss rather than a transport miss. The sample measures neither latency, nor cost, nor human agreement, and no such figures should be inferred from the function shapes. Free model access can change or disappear, so a harness that assumes a permanent shadow lane will fail in a quiet way.

Teams that need a safety evaluation, a legal review, or a calibrated study of rater agreement should not adopt this sketch as their gate. Anyone sending private customer content to a third-party free endpoint should stop before the first request, because convenience does not answer a data-handling question. A release process that treats one green shadow pass as enough evidence will misuse the columns and recreate false confidence. If your client cannot distinguish a timeout from a normal body, fix that boundary before you trust any score the ledger prints.

The useful habit is small and repeatable: classify the socket, classify the body, then grade only the rows that survived both checks. Keep shadow aliases in a separate directory so a cheap or free pass cannot be copied over the accepted file by a tidy script. A free server option can park that shadow process once you have read the current terms and kept the baseline local. The method still works if you remove the product entirely and run the same classifier against any other client.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

Top comments (0)