DEV Community

Cover image for Your OpenCode Model List Is Rotting: ocprobe
Sunny JayaRaju
Sunny JayaRaju

Posted on

Your OpenCode Model List Is Rotting: ocprobe

Your OpenCode whitelist is a list of promises made by other people's APIs. This is about the small bash tool I wrote to check them — and one bug in it that turned out to be far more interesting than the feature it came from.


TL;DR

ocprobe is a command-line tool that manages OpenCode model catalogs. It diffs the upstream catalog against your whitelist, probes the models that changed, shows you what it found, and only applies changes after you confirm — with a set of guards that are more interesting than the diffing.

The reason I care about it: running OpenCode against several model providers means your whitelist is a list of promises made by other people's infrastructure. Providers retire models, keys go stale, quotas run out, and a model that worked yesterday can 404 today. The failure is silent until something breaks mid-task, and the tool that reports the problem — the model picker — is not the thing that can tell you which entry is now lying.

Currently at v3.1.3, MIT licensed, installable with Homebrew in two commands.

The section I'd skip to if you only read one thing is the case study — it's about a security allowlist that passed every test on the machine it was written on and was quietly broken on every other platform.


The problem: your catalog rots quietly

Here's the manual loop this replaces, which is straight from the README's own description of the problem:

You diff catalogs by hand, probe models one by one, and risk wiping your whitelist with a bad apply.

Three separate failure modes hide in that sentence:

Models disappear. Providers retire models. An entry that was valid when you added it becomes a 404 later, and it stays in your whitelist until you try to use it.

Models fail without being gone. A key expires, a quota is exhausted, a provider has a bad afternoon. The model exists; your access to it doesn't.

Models change billing behaviour. This is the one that surprised me most, and it's why ocprobe has a status you might not expect. A probe can fail with BILLING_ERROR, which means your account can't currently afford that model's default output size — the model is completely fine and will work the moment credits exist. Blacklisting that would permanently hide a working model on a free or credit-limited key. So BILLING_ERROR is reported for review but never blacklisted, and it doesn't even count toward the failure gate.

That distinction only matters because ocprobe is willing to hide models at all. Which is the part that needs to be dangerous on purpose.

What it actually does

brew tap SunnyJayaRaju/ocprobe
brew install ocprobe
Enter fullscreen mode Exit fullscreen mode

The commands that matter day to day:

Command What it does
ocprobe doctor Health check: config, DB, auth, disk
ocprobe audit Full cycle: diff → probe → confirm → apply (the default)
ocprobe check Dry-run only; exits 1 if changes are pending
ocprobe validate Probes every model and manages the blacklist. Dry-run by default
ocprobe probe <model> Test one model right now
ocprobe status Whitelisted models plus recent probe results
ocprobe watch check + alerts on an interval, never auto-applies
ocprobe scheduler install Same thing as a launchd (macOS) or systemd (Linux) service
ocprobe session list / backup / restore / cleanup

A reasonable first run, which is also the sequence I'd suggest:

ocprobe doctor                 # is anything even set up correctly?
ocprobe check                  # what would change, without changing it?
ocprobe audit                  # actually probe, then confirm the apply
ocprobe validate               # what would be blacklisted, without doing it
ocprobe validate --apply       # do it, with a backup and a verification pass
Enter fullscreen mode Exit fullscreen mode

The default probe prompt is Reply with exactly: OK. That's not arbitrary — it's what makes the session-isolation guarantee below checkable.

Blacklisting is gated

ocprobe validate is the part that hides models from your picker, so it has rules about when it's allowed to do that.

Non-terminal failures — TIMEOUT, AUTH_ERROR, ERROR, UNCLEAR — need two consecutive failures before a model is added to the blacklist:

  • First failure → TENTATIVE. Surfaced for review, not blacklisted.
  • Second consecutive failure → CONFIRMED. Blacklisted.
  • A WORKS at any point resets the counter.
  • EOL / NOT_FOUND → CONFIRMED immediately. Those are terminal, and waiting buys nothing.
  • BILLING_ERROR → never blacklisted, and doesn't advance the counter.

There are two more brakes on the same path. A modality skip list means embedding, reranking, image/audio/video and moderation models are never probed at all — you didn't configure them for chat, so failing them is meaningless. And an AUTH_ERROR provider-wide abort: if more than 40% of a provider's models come back AUTH_ERROR, the whole provider is skipped, because that pattern means your key or your quota is the problem, not the models.

Lesson #1: the safe default for a tool that hides things is to require a second, independent piece of evidence. One failure is a bad afternoon. Two in a row is a fact.

The guards that keep it from making things worse

Because a tool that edits your config and deletes your sessions can absolutely cause the outage it was installed to prevent:

  • Session isolation. Only sessions created during this run, carrying the exact probe prompt in their first message, are ever deleted. Real conversations are never touched.
  • Mass-removal guard. If more than 50% of your whitelist would be removed in one run, ocprobe refuses to apply. Override is opt-in via OCPROBE_ALLOW_MASS_REMOVE=1.
  • Backups on every apply. Timestamped, next to the file: opencode.json.ocprobe-backup-YYYYMMDD-HHMMSS.
  • Graveyard cooldown. A model you deliberately removed won't reappear as a "new" model for 24 hours.
  • Restore is validated per statement. Covered below, because the interesting bug is there.

A case study: an allowlist whose correctness depended on your sqlite build

ocprobe session backup writes your OpenCode sessions to a SQL dump, and ocprobe session restore <file.sql> reads one back. A SQL dump is not something you want to execute by hand — so the restore is enforced through SQLite's authorizer callback, which lets you approve or deny each operation as the statement runs.

The policy ended up small and explicit, living in lib/session_restore.py:

ALLOWED_TABLES = frozenset(("session", "message", "part", "todo"))
ALLOWED_FUNCS  = frozenset(("unistr", "replace", "char"))
Enter fullscreen mode Exit fullscreen mode

Only inserts into the four tables cmd_session_backup actually emits. ATTACH, DETACH, DROP, DELETE, UPDATE, ALTER, every CREATE (which shows up as an insert into sqlite_master, so the table allowlist catches it), every PRAGMA, every other function, and all reads — so INSERT ... SELECT can't quietly copy data out of somewhere. The whole file runs as one transaction, so a rejected dump leaves the database byte-identical, and CI asserts that against the installed package rather than against a working tree where every file is present by definition.

That reads like a solid design. Here's the hole in it.

The replace entry in ALLOWED_FUNCS looks redundant. A real dump doesn't call a function called replace. It's there because of this, quoted from the module's own header:

SQLITE_FUNCTION callback for the conflict target of INSERT OR REPLACE ("replace"), which sqlite 3.54 on macOS does not report at all, so an allowlist of just {"unistr"} passed every local test and the macOS CI leg while breaking restore for every Linux user.

Read that again, because it's the whole lesson. Whether INSERT OR REPLACE causes a function callback at all is a property of your SQLite build, not of the program asking. The allowlist was not just listing which functions were permitted; it was implicitly asserting which functions would ever be offered to it. On macOS, replace never arrives, so its absence from the allowlist was untestable. On the Ubuntu runner, it arrives, isn't on the list, and the restore is denied — for a completely legitimate dump, with an error that names the function rather than the mistake.

There was a second hole of the same shape. INSERT INTO <table> DEFAULT VALUES is a real SQLITE_INSERT on a permitted table, so the authorizer approved it — and it wrote a row of NULLs over real data. It was refused only after a first-keyword check was added ahead of SQLite, which strips a BOM, Unicode whitespace and -- / /* */ comments before looking at the word. That change also stopped the code depending on another untested assumption: REINDEX, ANALYZE and VACUUM were already denied on the builds that had been tested, and after the change that's no longer load-bearing.

Two things make this worth writing down:

  1. An allowlist has two halves — what you permit, and what you expect to be asked about. Only the first is visible in the code. The second is a property of the environment, and the environment is not covered by your tests unless you run your tests in it.
  2. A denier that fires on a legitimate input looks exactly like a denier working. The Linux restore wasn't obviously broken; it was a strict-looking tool refusing a valid file. The only way to tell those apart is to have a test that asserts a known-good input is accepted, on every platform you claim to support.

This is why the project now runs its unit and integration suites on macOS and Ubuntu, plus a separate leg that runs the entire suite under a compiled bash 4.3.30 — the minimum the README advertises. macOS ships bash 3.2, and expanding a declared-but-empty array is an error there before 4.4. Five such expansions aborted. That minimum used to be asserted in the README and never tested.

Lesson #2: "passes on my machine" is not a portability story, it's a measurement you only took once.

Install and try it

brew tap SunnyJayaRaju/ocprobe
brew install ocprobe

ocprobe doctor          # config, DB, auth, disk
ocprobe check           # dry-run: what would change
ocprobe validate        # dry-run: what would be blacklisted
Enter fullscreen mode Exit fullscreen mode

Both of the dry-runs exit 1 when they have something to tell you, so they drop straight into CI or a pre-commit hook if that's useful to you. ocprobe watch is the low-maintenance option — it re-runs check and raises alerts, and never applies anything on its own.

Where it lives

MIT licensed. Contributions are welcome — the repo has one rule worth knowing up front, which I think is the right one: any PR that adds a feature, changes a command's behaviour, or changes a flag must update README.md and CHANGELOG.md in the same PR.


What it doesn't do

A few honest limits, so you can decide whether it's for you:

  • It's a management tool, not a proxy. It probes models on a schedule or on demand. It does not sit in the path of your requests and fail over when one breaks.
  • The policy engine is experimental and off by default. ocprobe policy gives you declarative never_add / never_remove rules for audit and check, but a missing policy file is a true no-op. It never touches validate.
  • Restores of older dumps still work, but only because the round-trip test covers the escaping both current and older SQLite builds produce for values containing newlines. That's a compatibility promise with a test behind it, not a general one.
  • Three config keys are accepted but do nothing: catalog.force_refresh, scheduler.enabled and scheduler.run_at_load. They were declared in the schema and shipped in generated configs, and nothing ever read them — so scheduler.enabled: false did not disable the scheduler. They're still accepted so existing configs keep validating, they log a warning if set to a non-default value, and ocprobe probe --force-refresh is the flag that actually works.
  • The validate path is the risky one. It's gated, backed up and verifiable, but it does hide models from your picker by design. Start with --apply off and read the diff.

Top comments (0)