At AirAsia, I own production microservices on the Manage My Booking platform — flight changes, fare summaries, price-slash, post-booking ancillaries. Over a dozen services, each with its own repo, pipeline, and accumulated quirks. The platform processes 130 million requests a day.
My problem was never writing code. It was the context-switching overhead. Bouncing between a fare-calculation refactor, a circuit breaker fix three repos over, and a schema migration for the ancillary service — every switch cost 20–30 minutes of mental reload. Entire days evaporated into friction.
I started using Claude Code not to write code for me, but to keep multiple workstreams moving in parallel. It operates directly in the terminal — reads the codebase, runs commands, writes files, stages changes. An agent, not an autocomplete engine.
Permission Scopes: The Non-Negotiable First Step
You cannot hand an AI agent unrestricted access to a production codebase. That is how you get a force-push to main at 2 AM.
I set up scoped permissions per project in .claude/settings.json:
{
"permissions": {
"allow": [
"Read",
"Edit",
"Write",
"Bash(npm run test*)",
"Bash(npm run lint*)",
"Bash(mvn test*)",
"Bash(git diff*)",
"Bash(git status*)",
"Bash(git log*)",
"Bash(git add -p*)"
],
"deny": [
"Bash(git push*)",
"Bash(git checkout main*)",
"Bash(rm -rf*)",
"Bash(docker*)"
]
}
}
The agent can read, write, edit, run tests and linters, view diffs, and stage changes. It cannot push, switch to main, delete directory trees, or touch Docker. Every push and merge goes through me. Payment-adjacent services get tighter restrictions. Internal tooling repos are looser.
The Three-Phase Model
I restructured feature delivery into three phases:
Inception (human only). I define the scope, draw sequence diagrams, identify affected services, and write interface contracts. This produces a spec document — the agent's input. Thirty to sixty minutes here, and the quality directly determines how useful the agent is downstream.
Construction (agents + review). Each agent gets a spec, a repo, and constraints. I run parallel terminal sessions — one agent per service:
- Terminal 1: Agent on booking-service, implementing the new endpoint
- Terminal 2: Agent on pricing-engine, adding fare calculation logic
- Terminal 3: Agent on ancillary-catalog, updating schema and data layer
- Me: Reviewing diffs, running integration tests, handling cross-service architecture decisions
The agents don't talk to each other. I am the coordination layer. When the pricing agent finishes, I review it and feed interface changes to the booking agent as context.
Operations (human only). Post-merge monitoring, production validation, performance profiling. I watch dashboards, check distributed traces in Zipkin, validate circuit breaker behavior. AI agents have no business in production operations.
Results After Six Months
Feature delivery time dropped ~40%. Cross-service features that took 4–5 days now take 2–3. The reduction comes from parallel Construction — three repos at once instead of sequential.
Unit test coverage increased. Writing tests is mechanical work agents handle well. I used to skip edge-case tests when tired. The agent does not get tired.
Code review quality improved. When I write code, I review it with the same mental model — and miss my own blind spots. Agent-written code gets genuinely fresh-eyed review. The separation between author and reviewer becomes real.
Commit history became cleaner. Selective staging and single-purpose commits made git bisect actually useful.
Not everything parallelizes. Tightly coupled changes where Service B depends on Service A's final interface still bottleneck. Those still see gains, but more like 15–20%.
Pitfalls Worth Knowing
The agent optimizes locally. It writes correct code for the service it sees while introducing interface mismatches with services it cannot see. Cross-service contracts must come from a human holding the full system model.
Spec quality is the bottleneck. "Add error handling to the booking endpoint" gets generic try-catch blocks. "Return HTTP 409 with BOOKING_CONFLICT when the fare class is unavailable, include available fare classes in the response body, emit a Kafka event to pricing-update" gets exactly what you need.
Don't let agents refactor while implementing. Mixed diffs — new functionality tangled with cleanup — are impossible to review. Implement against the existing structure; refactor in a separate PR.
Check in every 15–20 minutes. Not because the agent will break something catastrophic (permission scopes prevent that), but because a five-minute course correction at the 20-minute mark beats throwing away 90 minutes of wrong-direction implementation.
The Bigger Picture
This is not about AI replacing developers. It is about restructuring the workflow so human attention goes where it matters: architecture, system-level reasoning, production reliability. The mechanical act of translating a well-defined spec into working code with tests — that is what agents handle well. At the scale we operate at AirAsia, that mechanical work was eating a disproportionate amount of my time.
The spec-writing phase now takes longer than it used to. That tradeoff is worth it.
Originally published on krishnakky.com. Read the full version with detailed Git integration workflow, war stories, and what I'm experimenting with next.
Top comments (0)