DEV Community

Riley Wu
Riley Wu

Posted on

Four Scores Before You Leave Local

Local help should remain the default path for daily coding. A free remote server wins only on narrow, non-secret work. Four scores decide the hop before any token leaves the desk.

Developers often treat the network as a free lunch. The lunch is free only when the plate is empty of secrets. Offline days and slow links still charge you in waiting.

This note is a decision method, not a product tour. The scores come from checks you can run on one laptop. No vendor benchmark is assumed, and none is invented here.

The hop is a tax

A completion request looks small until it carries a secret. File paths, tokens, and customer names ride inside ordinary prompts. Once they cross the wire, local undo cannot pull them back.

Latency has the same shape as a tax on each edit loop. A local model pays that tax in compute on your machine. A remote model pays it in round trips plus queue time.

Offline work breaks the remote path without a warning banner. Trains, flights, and flaky cafe links are normal developer weather. A workflow that requires the cloud will stall in that weather.

Cost is the fourth vote, and free is not the same as fit. A free server can be the right tool for a redacted fixture. It is the wrong tool when the fixture still holds a secret.

Four scores, one route

Score each factor from zero to two before you send anything. A zero on any factor means the turn stays on the machine. Only a full pass on all four factors opens the server.

A timer alone cannot see a secret or an offline flag. Cost also stays invisible if you only watch the clock. The card exists so those missing votes get a number.

Latency score starts with a budget you choose for the task. Interactive edits often need a reply inside one second. Batch refactors can wait longer without breaking the editor flow.

Measure the empty hop before you measure the model. A health check with no repository text shows network time alone. If that empty hop already breaks the budget, do not send code.

Secret score is binary in practice, even if the scale has steps. Any live credential, customer row, or private diff scores zero. Redacted fixtures and public docs can score two after a scan.

Offline score asks whether the task must finish without a network. If the answer is yes, the remote option scores zero today. Queue the request on disk and retry when the link returns.

Cost score asks whether a free server changes the decision at all. If local compute already fits the budget, free remote adds little. Use the free path when local hardware cannot hold the context.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. The operator supplied two availability claims for this draft. MonkeyCode offers free model access and a free server option.

Model names, token quotas, hardware, and duration are not stated here. Those details were not verified against a primary source for this draft. Treat free access as a trial lane, not as a permanent contract.

A proposal you can run

The artifact below is a proposal, not a recorded run. It encodes the four scores in one small function. You can paste it into a scratch file and run the tests locally.

Fail closed works like a turnstile with four arms. One open arm does not let the bag through. A secret hit or a missing hop time keeps the arm down.

# Proposal only. Not executed for this draft.
# Score 0 keeps the turn local. Score 2 allows that factor.
# Any zero fails the remote hop closed.

from dataclasses import dataclass

@dataclass
class TaskCard:
    name: str
    latency_ms_budget: int
    empty_hop_ms: int | None
    secret_hits: int
    must_finish_offline: bool
    local_can_hold_context: bool
    free_server_available: bool

def score_latency(card: TaskCard) -> int:
    if card.empty_hop_ms is None:
        return 0
    if card.empty_hop_ms > card.latency_ms_budget:
        return 0
    if card.empty_hop_ms > card.latency_ms_budget // 2:
        return 1
    return 2

def score_secrets(card: TaskCard) -> int:
    if card.secret_hits > 0:
        return 0
    return 2

def score_offline(card: TaskCard) -> int:
    if card.must_finish_offline:
        return 0
    return 2

def score_cost(card: TaskCard) -> int:
    if not card.free_server_available and not card.local_can_hold_context:
        return 0
    if card.local_can_hold_context:
        return 1
    return 2

def decide(card: TaskCard) -> str:
    scores = (
        score_latency(card),
        score_secrets(card),
        score_offline(card),
        score_cost(card),
    )
    if any(value == 0 for value in scores):
        return "local"
    if all(value == 2 for value in scores):
        return "free_server"
    return "hold"

def sample_cards() -> list[TaskCard]:
    return [
        TaskCard("private-diff", 800, 120, 1, False, False, True),
        TaskCard("flight-edit", 800, 90, 0, True, False, True),
        TaskCard("public-fixture", 5000, 140, 0, False, False, True),
    ]

def test_cards() -> None:
    cards = sample_cards()
    assert decide(cards[0]) == "local"
    assert decide(cards[1]) == "local"
    assert decide(cards[2]) == "free_server"

if __name__ == "__main__":
    test_cards()
    for card in sample_cards():
        print(card.name, decide(card))
Enter fullscreen mode Exit fullscreen mode

Fill empty_hop_ms from a health probe that carries no prompt. Use the total time field, converted to milliseconds. A failed probe should leave the hop time empty.

# Unexecuted probe. Health route only. No repo files.
export HEALTH_URL="https://example.invalid/health"
curl -o /dev/null -s -w 'dns:%{time_namelookup} connect:%{time_connect} total:%{time_total}\n' "$HEALTH_URL"
Enter fullscreen mode Exit fullscreen mode

Read dns, connect, and total as three separate fields. Dns time points at resolver trouble, not at the model. Connect time points at the path, and total includes the server ack.

Do not paste a prompt into that probe to make it realistic. Realism here would smuggle the secret you meant to protect. Realism belongs in a local fixture after redaction, not on the wire.

Fill secret_hits from a local scan, not from a remote classifier. Count matches for keys, private blocks, and your own token prefixes. A count above zero must force the decision back to local.

# Local secret scan proposal. Tune patterns to your formats.
rg -n --hidden -g '!.git' -e 'AKIA[0-9A-Z]{16}' -e '-----BEGIN ' .
Enter fullscreen mode Exit fullscreen mode

Redaction is a local edit, done before any score can reach two. Replace names, ids, and urls with stable placeholders. Then scan again, because a placeholder can still hide a live key.

Read the three sample cards

Consider a public fixture that holds no customer data. The laptop cannot fit the context, and the link is up. The empty hop sits inside a five second batch budget.

That card can score two on every factor. The decision name is free_server, not a promise of quality. You still review the patch before it lands.

Now flip one field and watch the arm drop. Add a single secret hit to the same card. The function returns local, even when the server is free and fast.

A hold result is a pause, not a silent send. Latency scored one, or local compute already fits. Keep the turn local until a human raises the budget on purpose.

Write held cards to a local queue instead of a chat box. A JSON line per task is enough for a later replay. Replay only after the four scores pass again.

A queue line can store the task name and the four scores. Leave the prompt body out of that line. A later run should rebuild the card from the repo, not from the log.

Short edit loops should prefer the local path even when the server is free. A two hundred millisecond hop can still feel late inside a tight edit. Longer jobs can spend that same hop because nobody is waiting.

Context size is a local constraint, not a moral one. If the window will not fit, local fails for a mechanical reason. That is the cleanest case for a free server, once secrets are gone.

Record the hop time next to the decision in a plain text log. Do not copy a number from a blog into that log. Your link is the only latency figure that belongs in the card.

A free lane can be slow, full, or simply gone that day. The score does not retry forever in the background. On a timeout, return the card to the local queue.

Add the three asserts to your usual test command. A secret card must stay on the local path. An offline card must stay local, and a clean batch card may leave.

Raise a latency budget in the card, not in a hidden config. The number should sit next to the task name in review. Reviewers can then see why the hop was allowed.

When the free server wins

A free server wins when four facts line up on the same card. The hop fits the budget, and the scan is clean. The task can wait for a link, but local hardware cannot hold the window.

Use the gate when a local assistant is already in place. An optional remote lane is the only cloud piece required. Solo developers and small teams can share one secret list.

Who should skip this

Skip this gate if you cannot name what counts as a secret. A vague policy will mark everything safe and leak the rest. Teams without a scanner should stay local until that gap closes.

Skip the free server if your work must be reproducible next quarter. Unverified quotas can change, and this draft does not lock a number. Air-gapped repos should ignore the remote branch entirely.

Limits

The script does not call a model and does not open a socket. It only maps fields you already measured on the machine. Wrong inputs produce a confident wrong route, so measure first.

These checks do not prove a vendor is safe or fast. They prove your task cleared your own bars on this day. Run them again when the network, the repo, or the budget changes.

If a local loop is already your default, score one redacted fixture. Compare that card with the free server option only. Publish the four numbers beside the patch, not a slogan.

Top comments (0)