Claude Code used to feel like a brilliant teammate with no onboarding. Some days: tests first, small diff, asked before pushing. Other days: no tests, three "cleaned up" files, a PR opened on its own.
The model wasn't the problem. My setup was: a 20 KB CLAUDE.md that contradicted itself, dozens of copied skills fighting each other, and rules nothing enforced.
So I rebuilt ~/.claude like any other system. Here's how, step by step.
Short on time? Start here. Three things fix most of the chaos:
- One rulebook that says it overrides everything else (Step 2).
- Lanes, so Claude states the size of the job before acting (Step 3).
permissions.askforgit pushandgh pr create, so nothing ships without your yes (Step 8).๐ฆ Everything below is in a starter kit you can copy in about 10 minutes: the rulebook template, the 3 hooks, the 4 relay skills and their tests. github.com/dhrupo/claude-code-dev-os
What we're building
~/.claude/
โโโ CLAUDE.md # the one rulebook
โโโ lessons.md # mistakes not to repeat (imported by CLAUDE.md)
โโโ settings.json # hooks, permissions, plugins
โโโ skills/ # one folder per skill
โโโ skills.lock # where each skill came from, pinned
โโโ overlays/ # my edits, re-applied after updates
โโโ hooks/ # enforcement scripts
โโโ tests/ # tests for the setup itself
The starter kit has everything here except skills.lock and overlays/. Those only matter once you start pinning skills from GitHub (Step 6).
Step 1: Make ~/.claude a git repo
Every change becomes one commit, so any bad rule is one git revert away. ~/.claude also holds chat history and credentials, so use an allowlist .gitignore:
*
!.gitignore
!settings.json
!CLAUDE.md
!lessons.md
!skills.lock
!skills/**
!hooks/**
!overlays/**
!tests/**
cd ~/.claude && git init && git add -A && git commit -m "Baseline"
Keep it local. Never add a remote.
Step 2: Write ONE rulebook
Most of my "Claude ignores me" problems were really "Claude got conflicting instructions." The fix is one line at the top of ~/.claude/CLAUDE.md:
## 0. Precedence
This file overrides skills, plugins, slash commands and repo CLAUDE.md files.
Those may ADD project facts (test commands, code style). On conflict, this file wins.
Then the skeleton:
## 1. Defaults โ be concise; never push without asking; state what "done" means before building
## 2. Lanes โ ๐ข / ๐ก / ๐ด
## 3. Relay โ the steps
## 4. Skill map โ one owner per job
## 5. Rules โ test first, smallest change, no regressions, no noise comments
@lessons.md
Keep it small. It loads every session. If you can't say what a line changes, delete it.
Step 3: Lanes
A typo shouldn't get the same process as a payment-flow change. Claude must put the lane on the first line of every reply:
- ๐ข question โ answer only, no code.
- ๐ก small โ one obvious file. 3-line plan in chat. Tests still required.
- ๐ด full โ anything bigger. Full relay + written plan.
Ratchet up, never down: a ๐ก that grows must say so and switch to ๐ด.
๐ด full ยท bug (refund sends the receipt twice) tells me in one glance whether Claude understood the size of the job.
Step 4: The relay
๐ด work goes through 12 steps. Each step has one owner skill, and some steps have a STOP where I decide:
| # | Step | Owner |
|---|---|---|
| 1 | Grill: one question at a time, with a recommended answer โ STOP | grill-me |
| 2 | Plan: one page + list of allowed files โ STOP | openspec-plan |
| 3 | Check the plan: tests named? "won't change" listed? | plan-check |
| 4 | Build: failing test first, smallest fix |
tdd + ponytail
|
| 5 | Check the build: tests, E2E, lint, diff matches plan | build-check |
| 6 | Tidy: delete what isn't needed | ponytail-review |
| 7 | Deep review: bugs, security, performance, regressions | pre-merge-review |
| 8 | Fix findings โ back to 4 | |
| 9 | Step back: read the whole diff cold | step-back |
| 10 | PR: asks before push, asks again before the PR | commit-push-pr |
| 11 | Wait for review โ STOP | |
| 12 | Handle review: verify every finding, don't trust it blindly | receiving-code-review |
๐ก runs a light version (quick grill, then 3 โ 4 โ 5 โ 9). ๐ข skips it all.
Four of the skills here are my own (openspec-plan, plan-check, build-check, step-back), and each is under 60 lines. All four are in the starter kit, where the plan skill is called plan. For example, here's plan-check in full:
---
name: plan-check
description: Use right after the plan is written, before any code. Gates the PLAN.
---
1. **Tests named** โ the exact failing test(s) to write first.
2. **Must-not-change listed** โ what behaviour is at risk and how we protect it.
3. **Smallest change** โ no new dependency, no new public API, fewest files.
4. **No open questions.**
One line per check, pass/fail. Any fail โ fix the plan, not the code.
You don't need my 12 steps. You need named steps, one owner each, and clear STOP points.
Step 5: Skills, one owner per job
Skills from different sources compete. I had three TDD skills and two brainstormers that each claimed to run first. So the rulebook has a routing table: one owner per job, and everyone else steps aside or gets called by the owner. I also turned off 6 design skills that overlapped (skillOverrides in settings.json).
Here's my whole shelf in one pass:
-
Plan:
superpowers:brainstorming,grill-me,grill-with-docs,openspec-plan,plan-check,design-an-interface,improve-codebase-architecture,zoom-out,explain-this,source-driven-development,new-feature -
Build:
tdd(rigorous-coding-workflowdefers to it),bug-fix(builds a fast pass/fail loop first),ponytail,ponytail-review,ponytail-audit,ponytail-debt,context-packaging,handoff -
Review:
build-check,pre-merge-review,security-auditor(read-only),security-repro,infection-gate,phpstan-gate,step-back.pr-review-toolkitand Codex give a second opinion only when asked -
Ship:
commit-push-pr,release-readiness,knowledge-capture,retro,daily-work, plus Addy Osmani'sshipping-and-launch,ci-cd-and-automation,observability-and-instrumentation,performance-optimization,deprecation-and-migration - Understand code: graphify and the code-review-graph MCP. The rule is graph first, grep second
-
My stack: the WordPress agent skills (
wp-plugin-development,wp-rest-api,wp-block-development,wp-playground,wp-phpstan,wp-plugin-directory-guidelines) andplugin-audit
The special ones get a few lines each:
๐จ UI/UX is a pipeline with gates. ui-ux-pro-max picks the direction (palette, type, patterns) โ impeccable builds it well โ break-ui feeds it worst-case data (80-character names, empty lists, emoji) โ impeccable-audit + axe-a11y-audit must pass. Extras: review-animations, animation-vocabulary, mobile-native, and the Chrome DevTools MCP for live debugging.
โ๏ธ Pencil: draw it before you code it. When I say "mock it up", pen-design uses the pen.dev CLI to make a .pen file and a PNG in ~/Documents/designs/<project>/<id>/. It never overwrites (design-v2, design-v3โฆ) and logs each run. For a redesign, a screenshot of the current screen goes in as reference.png. Then the PNG becomes the target: the UI skills build to it, and build-check compares a screenshot of the result side by side. It never runs unless I ask.
๐งช E2E: "unit tests pass" โ "it works". Anything a user can click gets a Playwright test against the real site, red before the fix and green after. playwright-core writes it, playwright-auth-state handles logins, and playwright-cli / playwright-cli-runner are for exploring only. e2e-reviewer scans for tests that can pass while the feature is broken.
๐ Research and scraping, on demand. agent-reach does web and social research. Its queries leave the machine, so nothing sensitive goes in, and social logins only happen when I ask. scrapling scrapes JS-heavy or bot-protected pages locally, and only when plain WebFetch fails.
๐ Consent-gated. Some tools must never run by surprise. strix-local (a live pentest) only targets local or owned hosts and asks Attack <target> now? (y/n) every time. omniroute-session (routing through another model) runs only in a separate session, never machine-wide.
๐ Product layer. For the two products I work on daily (Fluent Forms and Fluent Player), I add trace skills that walk one real flow end to end before a fix, such as trace-submission-flow, trace-payment-flow, trace-ui-api and trace-player-runtime-flow. There are 18 in total. I also add product wrappers around the generic owners: grill-change, ui-ux-skills, fluentform-pre-pr-review, fluentplayer-release-readiness and so on. Copy the pattern: a generic shelf, plus one trace skill for whichever flow of yours breaks most.
Step 6: Pin skills, keep your edits
You copy a skill from GitHub and tweak it. Then updating it wipes your tweak. My fix:
-
skills.lockrecords each skill's repo and pinned commit, plus a status:installed-upstream,installed-ondemand,consent-gated, orskip(with the reason, so future me doesn't reinstall it). -
overlays/<name>.mdholds my edits. For example, myponytailoverlay says "shrink the code, never the tests." -
bin/sync-skills.shfetches each skill at its pinned commit, copies it in, and appends the overlay. Running it twice changes nothing the second time.
{ "name": "tdd", "status": "installed-upstream",
"repo": "https://github.com/mattpocock/skills", "ref": "6fd9479",
"path": "skills/engineering/tdd", "overlay": "tdd" }
Step 7: Plans live outside the repo
Skills don't share memory, and long chats get compacted. So skills hand off through files. Every ๐ด change gets a folder in ~/Documents/openspec/<project>/changes/<id>/ (the starter kit uses the simpler ~/Documents/plans/<project>/<id>/):
proposal.md # what's wrong ยท what changes ยท what WON'T change ยท how we'll know
tasks.md # checkboxes, failing test named first
scope.txt # the files this change may touch
evidence/ # test runs, screenshots, review reports
"What won't change" turns "don't break anything" into a promise you can test. Keeping the folder outside the repo means plans never land in a PR by accident.
Step 8: Hooks: rules ask, hooks enforce
The biggest lesson: a rule in CLAUDE.md is a request. A hook is a guarantee.
{
"permissions": { "ask": ["Bash(git push:*)", "Bash(gh pr create:*)"] },
"hooks": {
"UserPromptSubmit": [{ "hooks": [{ "type": "command", "command": "~/.claude/hooks/reminder.sh" }] }],
"PostToolUse": [{ "matcher": "Edit|Write", "hooks": [{ "type": "command", "command": "~/.claude/hooks/scope-guard.sh" }] }],
"Stop": [{ "hooks": [{ "type": "command", "command": "~/.claude/hooks/stop-check.sh" }] }]
}
}
-
permissions.ask:git pushandgh pr createalways prompt, even in auto mode. This is the simplest guard and the most important. -
Reminder (
UserPromptSubmit): whatever the hook prints is added to Claude's context, so one short paragraph re-anchors the rules on every prompt:
#!/bin/bash
echo "<rules>Rulebook: ~/.claude/CLAUDE.md. First line: lane. Test first. Smallest change. Ask before push.</rules>"
-
Scope guard (
PostToolUse): after each edit, if the file isn't in the plan'sscope.txt, it tells Claude to explain why and ask me. It warns, never blocks, because sometimes the extra file is correct. It just can't happen silently. -
Stop check (
Stop): when Claude finishes, it checksgit status. If source changed but no test did, it returns{"decision": "block", "reason": "..."}and Claude has to keep working. It fires once per new problem, so it can't loop.
My rulebook lists which rules have a hook behind them. Then I know which rules are guaranteed and which are just hoped for.
The full scripts and their tests are in the starter kit.
Step 9: Test the setup
Hooks are code. A broken scope guard looks exactly like one with nothing to report. So ~/.claude/tests/ pipes fake hook input into each script and checks the output:
OUT=$(echo '{"tool_input":{"file_path":"src/Other.js"},"cwd":"'$REPO'"}' | ../hooks/scope-guard.sh)
echo "$OUT" | grep -q "UNPLANNED FILE" && echo ok || echo FAIL
I have 7 test files with 140 assertions, and I run them before every commit to ~/.claude. One more script tracks the context budget: the size of the rulebook plus all skill descriptions. That way bloat shows up as a number, not a feeling.
Step 10: Learn from review
When a reviewer catches a real bug I should have caught, I add one line to lessons.md: the general pattern, plus the check that catches it. The rulebook imports that file, and the review steps walk every lesson against the diff, so each miss turns into a check for later changes.
The same rule tells Claude not to trust review findings blindly. For each finding it checks the real code, then decides: challenge it with evidence, fix the root cause, list it as out of scope, or ask.
What I learned
- One rulebook with precedence fixes most "Claude ignores me" problems.
- One owner per job. Competing skills cause random behaviour.
- Rules ask, hooks enforce. Put a hook behind whatever you can't afford to skip.
- Warn more than you block. Strict hooks get disabled, and then they protect nothing.
- STOP points where you decide (scope, plan, push, PR). Let everything in between run on its own.
-
Treat
~/.claudeas code: git, small commits, tests.
Starter checklist
- [ ]
git initin~/.claudewith an allowlist.gitignore - [ ] One
CLAUDE.mdwith a precedence clause - [ ] Lanes on the first line of every reply
- [ ] Named relay steps, one owner skill each, clear STOP points
- [ ]
permissions.askfor push and PR - [ ] Hooks: reminder, scope guard, stop check, each with a test
- [ ]
lessons.mdfor every bug review caught
Start with the first three plus ask-before-push. That alone fixes most of the chaos. Or grab the whole thing from the starter kit and run bash tests/run.sh.
What does your setup look like? Tell me in the comments. ๐
Top comments (0)