Magpie is a menu bar app that lets you point a coding agent's harness at a different model. Codex's harness on DeepSeek. Claude Code's harness on Kimi. Pick from a dropdown, get a different brain behind the same interface.
https://github.com/yetone/magpie
The dropdown isn't the interesting part. What's interesting is what the dropdown implies: the coding agent has quietly split into two independently swappable halves, and we've spent the last eighteen months talking about it as one thing.
The model is the half everyone argues about. The harness is the half that actually holds your work.
Two parts, one product
A harness is everything that isn't weights. The loop. The tool schemas. The permission model. The context assembly. The prompt scaffolding that tells the model it's allowed to edit files and how to ask for confirmation. The retry logic when a tool call comes back malformed. The compaction strategy when you blow the window.
That's the part you actually interact with. When you say "Claude Code is good at refactors," you're not describing a model. You're describing a harness that happens to be wired to one. Swap the weights underneath and a surprising amount of the behavior survives — the file-editing flow, the diff presentation, the way it asks before running something destructive.
Magpie's premise is that this is a legitimate architectural boundary and not a party trick. I think that's right. I also think it means we've been measuring the wrong thing for a year.
The lock-in was never the weights
Here's the uncomfortable part. If the harness and the model are separable, then your lock-in isn't the model either. Models are the easiest thing in this stack to replace. You change a string in a config file and you're on a different provider by lunch.
The lock-in is the config.
I went and looked at mine. On this machine, the things that define how my agent behaves live in:
-
~/.claude/settings.json— permissions, hooks, environment -
~/.claude.json— project history, MCP server registrations -
.mcp.jsonin a couple of repos, and one in my home directory I'd forgotten about -
CLAUDE.mdandAGENTS.md, at three different levels of the tree - a directory of custom slash commands
- a directory of subagent definitions
-
~/.codex/config.toml, which has its own provider block, its own approval policy, its own MCP list
That's the actual product surface. That's where my judgment lives — which commands I trust, which servers I've wired in, what I've told the thing about my codebase. And none of it is in a repo. None of it gets a pull request. None of it gets a diff.
I grepped my own allowlist while writing this. Thirty-one entries. Nine of them I could not explain. Some of them I added at 1am during an incident and never removed. One of them is broad enough that I'd flag it in review if a colleague proposed it.
The MCP wiring is the worst offender
MCP servers are the clearest example because they're not just config, they're context. Every server you register injects its tool schemas into the prompt whether you use those tools or not. I still have a database server wired in from a project that ended in March. It's been quietly eating a few thousand tokens of my window on every single session for four months, and I only noticed because I was looking for something else.
That's not a bug in any one harness. It's what happens when the layer that matters most is the layer with no review process. You wouldn't let a dependency sit in package.json for four months after you stopped importing it. But ~/.claude.json isn't package.json. It's a dotfile. It's invisible by design.
Swapping makes it worse before it makes it better
Here's my actual worry about Magpie, and it's not a criticism of the project — it's a criticism of where we are.
If I can swap harnesses from a menu bar, I now have two or three config formats to keep coherent instead of one. Each harness has its own permission syntax, its own MCP registration shape, its own instruction file convention. The swap is easy. The re-onboarding isn't. What looks like a model swap is really "now go re-teach the new harness everything the old one knew about you."
Which means the menu bar is a great demo and a mediocre workflow until somebody solves the portability problem underneath it. Maybe that's the next thing yetone builds. Maybe it's the thing the ecosystem needs and nobody's building because it's boring.
What I actually want
I want the agent profile to be a versioned artifact. One directory in the repo — call it agents/ — containing the harness config, the allowlist, the MCP servers, the instruction files, the hooks. Reviewed like code. Diffed like code. Rolled back like code. And a thin shim that materializes it into whatever dotfile the currently-selected harness expects.
I'm not asking for a standard. Standards at this stage of a technology are how you get three competing ones and a migration guide. I'm asking for a convention that one person can adopt on a Tuesday.
I haven't run Magpie in anger yet. It's early, it's a menu bar app, and menu bar apps are where good ideas go to get a nice icon and no test suite. Maybe the config portability problem is out of scope and I'm projecting. But the architectural claim underneath it — that the harness is a separable, swappable component — is correct, and once you accept it, the dotfile situation stops looking like a quirk and starts looking like the main risk.
Tomorrow I'm moving my allowlist into the repo. Nothing reads it yet. That's fine. At least it'll show up in a diff the next time I do something at 1am.
Top comments (2)
This is an interesting take, and had similar thoughts myself. For me although all big LLM pushers are talking about better AI, smarter reasoning and so on, I find small tiny teams from places you would least expect doing incredible work and research on what AI really means and how it should be architected with a future proof mindset. One example that caught my attention and I am exploring is a Greek org called Arpa Logical Systems or can't recall the full name, but they compare the AI models to ancient Greek philosopher takes on what makes a human. Eg. brain = model, body = runtime, aura = harness, craftsmanship = skillware, and so forth. So I think, as AI evolves, we start to focus less on engineering AI, and more on reverse engineering humans. What do you think?
tr.ee/dev-to