DEV Community

Cover image for How I set up Claude Code: one rulebook, a 12-step relay to use it daily
Dhrupo Nil
Dhrupo Nil

Posted on AI-assisted

How I set up Claude Code: one rulebook, a 12-step relay to use it daily

Claude Code used to feel like a brilliant teammate with no onboarding. Some days: tests first, small diff, asked before pushing. Other days: no tests, three "cleaned up" files, a PR opened on its own.

The model wasn't the problem. My setup was: a 20 KB CLAUDE.md that contradicted itself, dozens of copied skills fighting each other, and rules nothing enforced.

So I rebuilt ~/.claude like any other system. Here's how, step by step.

Short on time? Start here. Three things fix most of the chaos:

  1. One rulebook that says it overrides everything else (Step 2).
  2. Lanes, so Claude states the size of the job before acting (Step 3).
  3. permissions.ask for git push and gh pr create, so nothing ships without your yes (Step 8).

๐Ÿ“ฆ Everything below is in a starter kit you can copy in about 10 minutes: the rulebook template, the 3 hooks, the 4 relay skills and their tests. github.com/dhrupo/claude-code-dev-os

What we're building

~/.claude/
โ”œโ”€โ”€ CLAUDE.md          # the one rulebook
โ”œโ”€โ”€ lessons.md         # mistakes not to repeat (imported by CLAUDE.md)
โ”œโ”€โ”€ settings.json      # hooks, permissions, plugins
โ”œโ”€โ”€ skills/            # one folder per skill
โ”œโ”€โ”€ skills.lock        # where each skill came from, pinned
โ”œโ”€โ”€ overlays/          # my edits, re-applied after updates
โ”œโ”€โ”€ hooks/             # enforcement scripts
โ””โ”€โ”€ tests/             # tests for the setup itself
Enter fullscreen mode Exit fullscreen mode

The starter kit has everything here except skills.lock and overlays/. Those only matter once you start pinning skills from GitHub (Step 6).


Step 1: Make ~/.claude a git repo

Every change becomes one commit, so any bad rule is one git revert away. ~/.claude also holds chat history and credentials, so use an allowlist .gitignore:

*
!.gitignore
!settings.json
!CLAUDE.md
!lessons.md
!skills.lock
!skills/**
!hooks/**
!overlays/**
!tests/**
Enter fullscreen mode Exit fullscreen mode
cd ~/.claude && git init && git add -A && git commit -m "Baseline"
Enter fullscreen mode Exit fullscreen mode

Keep it local. Never add a remote.

Step 2: Write ONE rulebook

Most of my "Claude ignores me" problems were really "Claude got conflicting instructions." The fix is one line at the top of ~/.claude/CLAUDE.md:

## 0. Precedence
This file overrides skills, plugins, slash commands and repo CLAUDE.md files.
Those may ADD project facts (test commands, code style). On conflict, this file wins.
Enter fullscreen mode Exit fullscreen mode

Then the skeleton:

## 1. Defaults   โ€” be concise; never push without asking; state what "done" means before building
## 2. Lanes      โ€” ๐ŸŸข / ๐ŸŸก / ๐Ÿ”ด
## 3. Relay      โ€” the steps
## 4. Skill map  โ€” one owner per job
## 5. Rules      โ€” test first, smallest change, no regressions, no noise comments
@lessons.md
Enter fullscreen mode Exit fullscreen mode

Keep it small. It loads every session. If you can't say what a line changes, delete it.

Step 3: Lanes

A typo shouldn't get the same process as a payment-flow change. Claude must put the lane on the first line of every reply:

- ๐ŸŸข question โ€” answer only, no code.
- ๐ŸŸก small โ€” one obvious file. 3-line plan in chat. Tests still required.
- ๐Ÿ”ด full โ€” anything bigger. Full relay + written plan.
Ratchet up, never down: a ๐ŸŸก that grows must say so and switch to ๐Ÿ”ด.
Enter fullscreen mode Exit fullscreen mode

๐Ÿ”ด full ยท bug (refund sends the receipt twice) tells me in one glance whether Claude understood the size of the job.

Step 4: The relay

๐Ÿ”ด work goes through 12 steps. Each step has one owner skill, and some steps have a STOP where I decide:

# Step Owner
1 Grill: one question at a time, with a recommended answer โ†’ STOP grill-me
2 Plan: one page + list of allowed files โ†’ STOP openspec-plan
3 Check the plan: tests named? "won't change" listed? plan-check
4 Build: failing test first, smallest fix tdd + ponytail
5 Check the build: tests, E2E, lint, diff matches plan build-check
6 Tidy: delete what isn't needed ponytail-review
7 Deep review: bugs, security, performance, regressions pre-merge-review
8 Fix findings โ†’ back to 4
9 Step back: read the whole diff cold step-back
10 PR: asks before push, asks again before the PR commit-push-pr
11 Wait for review โ†’ STOP
12 Handle review: verify every finding, don't trust it blindly receiving-code-review

๐ŸŸก runs a light version (quick grill, then 3 โ†’ 4 โ†’ 5 โ†’ 9). ๐ŸŸข skips it all.

Four of the skills here are my own (openspec-plan, plan-check, build-check, step-back), and each is under 60 lines. All four are in the starter kit, where the plan skill is called plan. For example, here's plan-check in full:

---
name: plan-check
description: Use right after the plan is written, before any code. Gates the PLAN.
---
1. **Tests named** โ€” the exact failing test(s) to write first.
2. **Must-not-change listed** โ€” what behaviour is at risk and how we protect it.
3. **Smallest change** โ€” no new dependency, no new public API, fewest files.
4. **No open questions.**
One line per check, pass/fail. Any fail โ†’ fix the plan, not the code.
Enter fullscreen mode Exit fullscreen mode

You don't need my 12 steps. You need named steps, one owner each, and clear STOP points.

Step 5: Skills, one owner per job

Skills from different sources compete. I had three TDD skills and two brainstormers that each claimed to run first. So the rulebook has a routing table: one owner per job, and everyone else steps aside or gets called by the owner. I also turned off 6 design skills that overlapped (skillOverrides in settings.json).

Here's my whole shelf in one pass:

  • Plan: superpowers:brainstorming, grill-me, grill-with-docs, openspec-plan, plan-check, design-an-interface, improve-codebase-architecture, zoom-out, explain-this, source-driven-development, new-feature
  • Build: tdd (rigorous-coding-workflow defers to it), bug-fix (builds a fast pass/fail loop first), ponytail, ponytail-review, ponytail-audit, ponytail-debt, context-packaging, handoff
  • Review: build-check, pre-merge-review, security-auditor (read-only), security-repro, infection-gate, phpstan-gate, step-back. pr-review-toolkit and Codex give a second opinion only when asked
  • Ship: commit-push-pr, release-readiness, knowledge-capture, retro, daily-work, plus Addy Osmani's shipping-and-launch, ci-cd-and-automation, observability-and-instrumentation, performance-optimization, deprecation-and-migration
  • Understand code: graphify and the code-review-graph MCP. The rule is graph first, grep second
  • My stack: the WordPress agent skills (wp-plugin-development, wp-rest-api, wp-block-development, wp-playground, wp-phpstan, wp-plugin-directory-guidelines) and plugin-audit

The special ones get a few lines each:

๐ŸŽจ UI/UX is a pipeline with gates. ui-ux-pro-max picks the direction (palette, type, patterns) โ†’ impeccable builds it well โ†’ break-ui feeds it worst-case data (80-character names, empty lists, emoji) โ†’ impeccable-audit + axe-a11y-audit must pass. Extras: review-animations, animation-vocabulary, mobile-native, and the Chrome DevTools MCP for live debugging.

โœ๏ธ Pencil: draw it before you code it. When I say "mock it up", pen-design uses the pen.dev CLI to make a .pen file and a PNG in ~/Documents/designs/<project>/<id>/. It never overwrites (design-v2, design-v3โ€ฆ) and logs each run. For a redesign, a screenshot of the current screen goes in as reference.png. Then the PNG becomes the target: the UI skills build to it, and build-check compares a screenshot of the result side by side. It never runs unless I ask.

๐Ÿงช E2E: "unit tests pass" โ‰  "it works". Anything a user can click gets a Playwright test against the real site, red before the fix and green after. playwright-core writes it, playwright-auth-state handles logins, and playwright-cli / playwright-cli-runner are for exploring only. e2e-reviewer scans for tests that can pass while the feature is broken.

๐Ÿ”Ž Research and scraping, on demand. agent-reach does web and social research. Its queries leave the machine, so nothing sensitive goes in, and social logins only happen when I ask. scrapling scrapes JS-heavy or bot-protected pages locally, and only when plain WebFetch fails.

๐Ÿ” Consent-gated. Some tools must never run by surprise. strix-local (a live pentest) only targets local or owned hosts and asks Attack <target> now? (y/n) every time. omniroute-session (routing through another model) runs only in a separate session, never machine-wide.

๐Ÿ  Product layer. For the two products I work on daily (Fluent Forms and Fluent Player), I add trace skills that walk one real flow end to end before a fix, such as trace-submission-flow, trace-payment-flow, trace-ui-api and trace-player-runtime-flow. There are 18 in total. I also add product wrappers around the generic owners: grill-change, ui-ux-skills, fluentform-pre-pr-review, fluentplayer-release-readiness and so on. Copy the pattern: a generic shelf, plus one trace skill for whichever flow of yours breaks most.

Step 6: Pin skills, keep your edits

You copy a skill from GitHub and tweak it. Then updating it wipes your tweak. My fix:

  • skills.lock records each skill's repo and pinned commit, plus a status: installed-upstream, installed-ondemand, consent-gated, or skip (with the reason, so future me doesn't reinstall it).
  • overlays/<name>.md holds my edits. For example, my ponytail overlay says "shrink the code, never the tests."
  • bin/sync-skills.sh fetches each skill at its pinned commit, copies it in, and appends the overlay. Running it twice changes nothing the second time.
{ "name": "tdd", "status": "installed-upstream",
  "repo": "https://github.com/mattpocock/skills", "ref": "6fd9479",
  "path": "skills/engineering/tdd", "overlay": "tdd" }
Enter fullscreen mode Exit fullscreen mode

Step 7: Plans live outside the repo

Skills don't share memory, and long chats get compacted. So skills hand off through files. Every ๐Ÿ”ด change gets a folder in ~/Documents/openspec/<project>/changes/<id>/ (the starter kit uses the simpler ~/Documents/plans/<project>/<id>/):

proposal.md   # what's wrong ยท what changes ยท what WON'T change ยท how we'll know
tasks.md      # checkboxes, failing test named first
scope.txt     # the files this change may touch
evidence/     # test runs, screenshots, review reports
Enter fullscreen mode Exit fullscreen mode

"What won't change" turns "don't break anything" into a promise you can test. Keeping the folder outside the repo means plans never land in a PR by accident.

Step 8: Hooks: rules ask, hooks enforce

The biggest lesson: a rule in CLAUDE.md is a request. A hook is a guarantee.

{
  "permissions": { "ask": ["Bash(git push:*)", "Bash(gh pr create:*)"] },
  "hooks": {
    "UserPromptSubmit": [{ "hooks": [{ "type": "command", "command": "~/.claude/hooks/reminder.sh" }] }],
    "PostToolUse": [{ "matcher": "Edit|Write", "hooks": [{ "type": "command", "command": "~/.claude/hooks/scope-guard.sh" }] }],
    "Stop": [{ "hooks": [{ "type": "command", "command": "~/.claude/hooks/stop-check.sh" }] }]
  }
}
Enter fullscreen mode Exit fullscreen mode
  • permissions.ask: git push and gh pr create always prompt, even in auto mode. This is the simplest guard and the most important.
  • Reminder (UserPromptSubmit): whatever the hook prints is added to Claude's context, so one short paragraph re-anchors the rules on every prompt:
  #!/bin/bash
  echo "<rules>Rulebook: ~/.claude/CLAUDE.md. First line: lane. Test first. Smallest change. Ask before push.</rules>"
Enter fullscreen mode Exit fullscreen mode
  • Scope guard (PostToolUse): after each edit, if the file isn't in the plan's scope.txt, it tells Claude to explain why and ask me. It warns, never blocks, because sometimes the extra file is correct. It just can't happen silently.
  • Stop check (Stop): when Claude finishes, it checks git status. If source changed but no test did, it returns {"decision": "block", "reason": "..."} and Claude has to keep working. It fires once per new problem, so it can't loop.

My rulebook lists which rules have a hook behind them. Then I know which rules are guaranteed and which are just hoped for.

The full scripts and their tests are in the starter kit.

Step 9: Test the setup

Hooks are code. A broken scope guard looks exactly like one with nothing to report. So ~/.claude/tests/ pipes fake hook input into each script and checks the output:

OUT=$(echo '{"tool_input":{"file_path":"src/Other.js"},"cwd":"'$REPO'"}' | ../hooks/scope-guard.sh)
echo "$OUT" | grep -q "UNPLANNED FILE" && echo ok || echo FAIL
Enter fullscreen mode Exit fullscreen mode

I have 7 test files with 140 assertions, and I run them before every commit to ~/.claude. One more script tracks the context budget: the size of the rulebook plus all skill descriptions. That way bloat shows up as a number, not a feeling.

Step 10: Learn from review

When a reviewer catches a real bug I should have caught, I add one line to lessons.md: the general pattern, plus the check that catches it. The rulebook imports that file, and the review steps walk every lesson against the diff, so each miss turns into a check for later changes.

The same rule tells Claude not to trust review findings blindly. For each finding it checks the real code, then decides: challenge it with evidence, fix the root cause, list it as out of scope, or ask.


What I learned

  1. One rulebook with precedence fixes most "Claude ignores me" problems.
  2. One owner per job. Competing skills cause random behaviour.
  3. Rules ask, hooks enforce. Put a hook behind whatever you can't afford to skip.
  4. Warn more than you block. Strict hooks get disabled, and then they protect nothing.
  5. STOP points where you decide (scope, plan, push, PR). Let everything in between run on its own.
  6. Treat ~/.claude as code: git, small commits, tests.

Starter checklist

  • [ ] git init in ~/.claude with an allowlist .gitignore
  • [ ] One CLAUDE.md with a precedence clause
  • [ ] Lanes on the first line of every reply
  • [ ] Named relay steps, one owner skill each, clear STOP points
  • [ ] permissions.ask for push and PR
  • [ ] Hooks: reminder, scope guard, stop check, each with a test
  • [ ] lessons.md for every bug review caught

Start with the first three plus ask-before-push. That alone fixes most of the chaos. Or grab the whole thing from the starter kit and run bash tests/run.sh.

What does your setup look like? Tell me in the comments. ๐Ÿ‘‡

Top comments (0)