DEV Community

Cover image for Crewforth 3.0: giving Claude Code a crew, and gates that don't forget
Barış Yerlikaya
Barış Yerlikaya

Posted on

Crewforth 3.0: giving Claude Code a crew, and gates that don't forget

I've been using Claude Code every day for months, at work and on side projects. It's the best coding assistant I've used. It also has two habits that kept costing me time:

  1. It's one assistant doing every job at once. The same context plans the feature, writes the migration, reviews its own diff and writes the commit message. Nothing stops it from skipping the planning, and nobody else looks at the diff.
  2. Rules are suggestions. I can write "never force-push" and "ask before committing" in CLAUDE.md, and most of the time it listens. Under pressure, in a long session or in auto mode, it sometimes doesn't.

Crewforth is what I built to fix both. Version 3.0 shipped this week. It used to be called Claude Starter Kit.

A crew instead of one assistant

Crewforth adds four kinds of things to Claude Code, all standard Claude Code features:

  • 12 subagents, each owning one domain: crew-planner, crew-backend-expert, crew-database-expert, crew-frontend-expert, crew-devops-expert, crew-security-expert, crew-privacy-agent, crew-performance-expert, crew-test-expert, crew-review-agent, crew-commit-agent and crew-session-manager.
  • 39 skills that hold the method: how to plan a change, how to review a migration, how to write a runbook, and so on.
  • 10 commands you start with /crew-…: /crew-plan, /crew-review, /crew-ship, /crew-handoff, /crew-doctor and a few more.
  • Hooks that enforce the rules that matter (more on these below).

The work follows five stages:

  1. Understand. An unclear request goes to crew-planner first. It scopes the change and writes acceptance criteria before any code exists.
  2. Produce. The owner builds it: server work to crew-backend-expert, schema changes to crew-database-expert. A routing hook names the owner beside your request, so you can see who picked it up.
  3. Audit. Security, privacy and performance audits run in parallel on the diff.
  4. Close. crew-review-agent reviews the exact diff with a plan, a severity and a category for every finding. Only critical and high findings block. Then crew-commit-agent proposes the commit and waits for your approval.
  5. Hand off. Before the context window fills up, a handover note is written. The next session reads it and picks up where you stopped.

It's stack-agnostic. Earlier versions had a .NET-specific path; 3.0 dropped it. The stack is detected from your repository or recorded once in CLAUDE.md, and the same agents work on Go, TypeScript, Python or C#.

Gates, not reminders

This is the part I care about most. A rule in a prompt is a request. A hook is code that runs outside the model, on every tool call, whatever the model decided.

Crewforth turns a small set of rules into hooks:

  • A destructive shell command (git push --force, git reset --hard, a recursive rm -rf, git commit --no-verify) is refused before it runs.
  • A commit waits for your explicit approval, with the review on record.
  • The staged diff is scanned for secrets before it can reach history.
  • Turning the gates off (git config core.hooksPath …) is itself refused.

They hold in every permission mode, auto mode included, because the model never gets a vote.

Measured, including what didn't hold

It's easy to write "makes Claude safer" in a README. I wanted numbers, so every claim runs the same prompt with and without Crewforth, in a throwaway workspace, and grades what is left on disk. The rule for publishing a claim is written down before the run.

The first run of the main experiment looked great: without Crewforth, 6 of 10 sessions left a world-writable file under pressure; with it, 0 of 10. Then I rewrote parts of the prompts for 3.0 and ran it again under the same rule. Without Crewforth it was 4 of 10 this time, below the threshold I had set. So the claim came off the README.

Both runs are on the evals page, next to the experiments where Crewforth made no difference at all. I would rather show you that than a single flattering number.

Try it

New project:

npx crewforth init
Enter fullscreen mode Exit fullscreen mode

Existing repository:

npx crewforth adopt
Enter fullscreen mode Exit fullscreen mode

adopt puts everything on a separate branch, staged and uncommitted. main is never touched, and you review the diff before anything lands.

As a Claude Code plugin:

/plugin marketplace add crewforth/crewforth
/plugin install crewforth@crewforth
Enter fullscreen mode Exit fullscreen mode

Then run /crew-doctor inside Claude Code to check the setup.

Two more things:

  • npx crewforth studio opens a local panel that draws every delegation live: who runs, who waits, who failed.
  • The installer speaks English and Turkish. A Turkish install answers in Turkish, even after a bare command.

One note on context: Claude Code gives the skill list a slice of the context window. On a 1M-token model it fits by default. On a 200k model, /crew-doctor shows you the one setting that makes room and what it costs per turn, so you can decide.

If you used Claude Starter Kit

Nothing to do. The old package forwards to the new one, so your next update moves the project to 3.0 and renames the components. Your own files and your own lines in CLAUDE.md are left alone.

What's next

3.0.1 is already collecting small fixes. After that I want an eval that measures whether Claude picks the right skill from its description alone, because that's the part of the system I can currently only reason about, not measure.

If you try it, I'd like to hear where the gates get in your way. That's the feedback I can't produce myself.

Top comments (2)

Collapse
 
henry786 profile image
Henry •

Having dedicated subagents check the diff before a push is pretty brilliant. At The Printing World, having hard gates like that would stop us from sending wrong dielines or packaging specs to the printers when someone rushes an update.

Collapse
 
byerlikaya profile image
Barış Yerlikaya •

Thanks! One small correction: the review happens before the commit, not the push. crew-review-agent reviews the exact diff, and the commit waits for your approval. The hard stops are hooks, so they hold even when someone is in a hurry.

The same idea should carry over to print specs: an automated check on what actually changed, plus a human sign-off before anything goes out.