DEV Community

Cover image for 349 Tests, Zero Module Mocks: Building Blast Radius Spec-First with Kiro
Scott Burgholzer for AWS Community Builders

Posted on Originally published at blog.scottburgholzer.tech

349 Tests, Zero Module Mocks: Building Blast Radius Spec-First with Kiro

The first three articles covered what Blast Radius does, how to use it, and how it's designed. This one is the build log: the engineering decisions, the patterns that paid off, and the two runtime gotchas that cost me an afternoon each. But the decision that shaped everything else came before the first line of code: I built the whole thing spec-first, using Kiro. That's why a TypeScript monorepo with 349 tests and no module mocking isn't an accident. It's what falls out when you decide the structure before you write the implementation.

Building Spec-First with Kiro

The interesting thing about how Blast Radius came together isn't the origin story, it's the method. Instead of opening an editor and hacking on the first Lambda, I worked through Kiro's spec workflow: a requirements document first, then a design document, then a task breakdown, and only then code. The architecture (canonical format, Step Functions pipeline, the dependency-injection testing strategy) was decided on paper before any of it was written.

That ordering matters more than it sounds. When the design is settled before implementation, the whole system lands as one coherent build rather than a prototype you keep bolting onto. The git history shows it plainly: the foundation arrives in a single initial commit of the whole system, and what follows is mostly refinement and targeted fixes (architecture docs, frontend UX, the CLI and AI gate, CI and release workflows, plus the deployment gotchas I'll get to) rather than a ground-up rewrite. There's no "v1 that I threw away" in the log, because the big structural decisions were made in the design doc, where changing your mind is cheap, instead of in code, where it isn't.

It's also why this article is mostly about choices that held up rather than dead ends I backed out of. The canonical format, the strict package boundaries, dependency injection over mocking: those were design decisions, not things I stumbled into halfway through and refactored toward. Deciding them up front is exactly why they're consistent across every package.

The Kiro spec workflow genuinely changed how I work. Forcing myself to refine the idea before writing code meant I wasn't refactoring my way out of decisions I'd made too early. It also changed my relationship with testing: before Kiro, I mostly tested things by hand and rarely wrote real unit tests. Working spec-first, with tests planned into the tasks, is how a project like this ended up with 349 tests.

So where did things go wrong? Not in the architecture. In the seams between the code and the runtime. The failures that actually cost me time were deployment and platform surprises: an API Gateway timeout that needed tuning, Bedrock model configuration that wasn't what I first reached for, and two Node.js runtime gotchas that passed every local test and only broke in AWS. Those are the war stories, and they show up later because they're the parts a spec can't protect you from. The lesson that frames the rest of this piece: when you design carefully up front, the bugs that remain aren't in your logic. They're at the boundary where your code meets someone else's platform.

A TypeScript Monorepo Without the Ceremony

Blast Radius is one repository with five workspaces:

Package What it holds
@blast-radius/core Shared models, validation, cache, retry, verdict, auth scoping
@blast-radius/lambdas Every Lambda handler in the analysis pipeline
@blast-radius/frontend The React + Vite + Cytoscape.js SPA
@blast-radius/cli The CI/CD integration tool
@blast-radius/infra The CDK stack that deploys the whole thing

The orchestration is plain npm workspaces: no Nx, no Turborepo, no Lerna. That's a deliberate omission. For a project this size, a heavyweight monorepo tool buys you incremental-build caching and task graphs you don't need yet, at the cost of a config surface you have to learn and maintain. npm workspaces gives the one thing that actually matters here: packages can depend on each other by name ("@blast-radius/core": "*") and resolve locally without publishing. The root package.json runs builds and tests across all workspaces with --workspaces, and that's the whole build system.

The dependency direction is strict and one-way: core depends on nothing internal, and everything else depends on core. The Lambda handlers, the CLI, and the frontend all import their shared types (ResourceChangeManifest, ScoredResource, DependencyGraph) from @blast-radius/core. That single rule keeps the package dependencies from ever forming a cycle and means the domain model has exactly one definition. When the scoring formula's output shape changes, it changes in one file, and TypeScript's strict mode (no implicit any, strict null checks) makes every consumer that needs updating light up red.

There's one deliberate exception to the shared-types rule. The frontend doesn't import from @blast-radius/core; its api/types.ts keeps a hand-maintained copy of the shapes it consumes (ScoredResource, DependencyGraph, and friends). That's a pragmatic call: the SPA is bundled separately and I didn't want to pull the backend package into the browser build just for its type declarations. The tradeoff is honest to state, because it's the one contract the compiler doesn't enforce end-to-end. The mirror can drift, so it's the seam I watch most carefully when the backend response shape changes.

349 Tests, Zero Module Mocks

Here's the claim, and it's literal: running the suite reports 349 passing tests across 29 test files, and a search for vi.mock( across the codebase returns nothing. To be precise about what that means: there's no module mocking, no intercepting imports so a file gets a fake version of its dependencies. I still use vi.fn() to build fake AWS clients, but those fakes are handed to the code explicitly rather than swapped in behind its back. Two decisions make that possible.

Decision one: dependency injection instead of module mocking

Every Lambda that touches AWS takes an optional deps parameter carrying its SDK clients. In production the handler builds its own; in tests you hand it fakes. The resource resolver's signature is representative:

export async function handler(
  event: ResolverInput,
  deps?: ResolverDeps,
): Promise<ResolverOutput> {
  const resolvedDeps = deps && 'configClient' in deps ? deps : createDefaultDeps();
  // ...
}
Enter fullscreen mode Exit fullscreen mode

That deps && 'configClient' in deps check looks paranoid until you know the footgun behind it. The AWS Lambda runtime invokes your handler with (event, context), passing the Context object as the second argument. So deps is always truthy in production; it's just the wrong object. A naive deps ?? createDefaultDeps() would accept the Context as your dependencies and then explode the first time it tried to call .send() on something that isn't an SDK client, and it would only fail in the real runtime, never in tests. The shape check ('configClient' in deps) distinguishes "real injected dependencies" from "the Lambda Context that happens to be sitting in that argument slot." Every AWS touching handler uses the same pattern with its own key: 'dynamoClient' in deps, 'configClient' in deps, and so on.

Because dependencies come in through the front door, tests never need to intercept module imports. They build a fake ConfigServiceClient (or DynamoDBClient, or LambdaClient) with vi.fn() stubs, pass it in, and assert on the output. No vi.mock, no hoisting quirks, no brittle coupling to import paths. The test reads like the production call, just with different clients.

Decision two: property-based testing where invariants matter

For pure logic (scoring, validation, filtering, sorting, caching), example-based tests only cover the cases you thought of. Blast Radius uses fast-check to test properties instead: statements that must hold for all inputs. fast-check generates hundreds of random cases and shrinks any failure to a minimal counterexample.

They're used across all three layers that have pure logic. In core: manifest validation, the LRU cache, the threshold evaluator, hierarchy flattening, access scoping. In lambdas: the risk assessor's scoring and dependency-chain logic, the adapters, the top-K summary selection. In frontend: graph filters, resource-table sorting, JSON export. A representative property, from the resource-table sort:

it('sortResources produces non-increasing order of impactScore when sorted descending', () => {
  fc.assert(
    fc.property(
      fc.array(arbitraryScoredResource(), { minLength: 0, maxLength: 100 }),
      (resources) => {
        const sorted = sortResources(resources /* desc */);
        // assert each element's score >= the next
      },
    ),
  );
});
Enter fullscreen mode Exit fullscreen mode

The value isn't "more tests." It's that the invariant is stated directly ("sorting never produces an out-of-order pair, for any list up to 100 resources"), and fast-check hunts for the input that breaks it. The validation tests lean on this hardest: they assert the exact error path returned for a manifest missing resourceType, resourceId, provider, or modificationType, across randomized manifest sizes and randomized which-resource-is-broken. That's the kind of coverage you can't reasonably hand-write.

The two decisions reinforce each other. Dependency injection means the AWS-facing handlers are testable at all, and property-based testing means the pure logic underneath is tested hard. Neither one reaches for module mocking to get there.

The Adapter Pattern, In Code

The previous article made the architectural case for the canonical format. Here's how it actually works in the source, because the details are where the pattern earns its keep.

Each IaC format has its own adapter Lambda. The Terraform adapter's job is to turn terraform show -json output into canonical ResourceChange records, and the interesting part is the action mapping, the exact place where three tools' vocabularies diverge:

function mapActions(actions: string[]): ModificationType | undefined {
  if (actions.length === 1) {
    switch (actions[0]) {
      case 'create': return 'Add';
      case 'update': return 'Modify';
      case 'delete': return 'Remove';
      case 'no-op':
      case 'read':   return undefined;   // skip: not a real change
    }
  }
  if (actions.length === 2) {
    const sorted = [...actions].sort();
    if (sorted[0] === 'create' && sorted[1] === 'delete') {
      return 'Replace';               // ["delete","create"] OR ["create","delete"]
    }
  }
  return undefined;
}
Enter fullscreen mode Exit fullscreen mode

Two details worth noting. First, Terraform expresses a replacement as a two-element action array, ["delete", "create"], and the adapter sorts before comparing, so it doesn't matter which order Terraform emits. CloudFormation expresses the same concept as Replacement: "True" on a Modify action, and CDK as changeType: "REPLACE". Three encodings, one canonical Replace. Second, no-op and read map to undefined and get skipped entirely, because they aren't changes, so they never enter the graph. The adapter also normalizes providers (registry.terraform.io/hashicorp/aws becomes aws) so downstream code sees a clean provider string.

The adapters don't call each other or know about a central switch statement. Routing lives in the adapter registry, a Lambda backed by a DynamoDB table. The registry maps a formatId to an adapter Lambda ARN:

  1. The pipeline hands the registry { format, payload }.
  2. The registry does a DynamoDB GetItem on formatId.
  3. If found, it InvokeCommands that adapter Lambda and returns the canonical manifest plus metadata (adapter name, conversion duration, warnings).
  4. If not found, it returns an error listing every supported format, pulled live from a Scan of the same table.

That indirection is what makes "add a new IaC tool = one adapter" literally true. A new format is a new adapter Lambda plus one row in the table. Nothing in the pipeline, the registry, or the analysis engine changes. The table is seeded automatically on CDK deploy (via an AwsCustomResource), so a fresh deployment comes up with CDK, CloudFormation, and Terraform already registered.

And the registry uses the same dependency-injection shape as everywhere else (deps && 'dynamoClient' in deps ? deps : createDefaultDeps()), so its routing logic is tested by passing in a fake DynamoDB client, no import interception required.

The Two Runtime Gotchas

I mentioned these two briefly in the previous article's lessons, but they're worth the fuller version here because they're the sharpest example of the theme running through this whole build. Two bugs in this project shared a personality: they only showed up in the real runtime, passed every local test, and made no sense at the point they surfaced. Both are the kind of thing you can't unit-test your way to.

Gotcha one: Node.js 22 Lambda handlers must be async. On the Node.js 22 runtime, a synchronous handler resolves to null. The runtime returns before your synchronous work is reflected. The adapters were the casualties: originally plain synchronous functions, they'd appear to "succeed" and hand the pipeline a null manifest. The failure then surfaced two or three steps downstream, in discovery, where a null manifest made no sense and the stack trace pointed nowhere useful. The fix is one keyword (every adapter handler is now declared async, which you can see in the Terraform adapter's export async function handler), but the diagnosis ate an afternoon. Lesson: on modern Node Lambda runtimes, make every handler async by default, full stop.

Gotcha two: the Lambda Context masquerading as your dependencies. This is the flip side of the dependency-injection pattern above. Because the runtime passes Context as the second argument, the intuitive deps ?? createDefaultDeps() is quietly wrong: deps is truthy in production, so you'd use the Context as your SDK clients and crash. It's invisible in tests (which pass real fakes or nothing) and only detonates in deployment. The 'configClient' in deps shape check is the fix, and it's applied uniformly across handlers. Both gotchas share the same lesson: the boundary between "your code" and "the runtime" is where the surprises live, so test the boundary, not just the interior.

A CLI That's One File and One Command

The CLI has a deliberately small surface built around one main command. blast-radius analyze is the workflow: the thing you run in CI. The rest are small helpers around it: generate produces input files without submitting, status and export fetch an analysis after the fact, and cdk-diff is kept as an alias for analyze --format cdk. No sprawling subcommand tree, no plugin system, just one verb that matters and a few conveniences.

Three implementation choices make it pleasant to use in CI:

Auto-generation by shelling out. For local use, analyze can build its own input. For CDK it runs cdk synth and creates a CloudFormation changeset; for Terraform it runs terraform plan + terraform show -json; for CloudFormation it creates a changeset from a template. These are literal shell-outs to the tools you already have installed (generate.ts and cdk-diff.ts assemble the aws cloudformation create-change-set / describe-change-set / delete-change-set sequence). The changeset is always deleted, never executed. It's a read-only "what would you do?" query.

Polling with stale detection. The analysis is async, so after submitting, the CLI polls status every 3 seconds up to a 90-second ceiling. The clever bit is failure detection: the backend's catch-everything design marks crashed analyses failed, but as a belt-and-suspenders backstop the CLI also watches the status timestamp. If updatedAt hasn't moved for 5 consecutive polls (about 15 seconds), it assumes the analysis died silently and exits as failed rather than spinning until the timeout:

if (updatedAt && updatedAt === lastUpdatedAt) {
  staleCount++;
  if (staleCount >= 5) { finalStatus = 'failed'; break; }
} else {
  staleCount = 0;
  lastUpdatedAt = updatedAt;
}
Enter fullscreen mode Exit fullscreen mode

Single-file distribution, but not from the CLI build. This is a place where the docs could mislead you if you only read the CLI's package.json, which just runs tsc --build into dist/. The single downloadable blast-radius.js is produced by the release workflow, not the package build. On a version tag, .github/workflows/release.yml runs esbuild packages/cli/src/index.ts --bundle --platform=node --target=node20, prepends a #!/usr/bin/env node shebang, smoke-tests it, and attaches it to the GitHub release. That's why CI users can curl one file and run it with nothing but a Node runtime: no npm install, no node_modules. The bundling is a release-time concern, cleanly separated from local development.

Using CDK to Analyze CDK

There's a pleasant irony in the infra package: Blast Radius analyzes CloudFormation changesets, and it's deployed by CDK, which produces CloudFormation. The tool could, in principle, analyze its own deployment.

Two CDK patterns are worth calling out because they solve real problems:

  • Custom resources for seeding. The adapter registry table needs its three default rows (cdk / cloudformation / terraform) to exist the moment the stack is live. Rather than a manual post-deploy script, an AwsCustomResource writes those rows during deployment, so a fresh stack is immediately functional.
  • Conditional constructs for optional features. Bedrock summaries, auth, and results retention are toggleable via CDK props (enableBedrockSummary, enableAuth, resultsRetentionDays). The stack wires the Bedrock env vars and IAM (foundation-model/* and inference-profile/*) only when summaries are enabled, so a deployment with AI disabled doesn't carry permissions it never uses.

The IAM story is a good example of least-privilege falling out of the design: the resource resolver gets exactly the Config and Resource Explorer read actions it needs, the risk summary gets Bedrock invoke on model and inference-profile ARNs, and nothing gets more than its one job requires.

What I'd Do Differently

An honest retrospective, because every build has one. None of these are things I'd change today, but they're the edges I'd sand down if this grew past a v0.1.

  • The 90-second pipeline ceiling is a soft limit, not a real bound. The Step Functions state machine has a 120-second timeout and the CLI polls for 90; a genuinely huge dependency graph could bump into that. A production-hardening pass would make discovery resumable or paginate very large graphs rather than relying on a single bounded traversal.
  • Coverage is honest but coarse. full / partial / unknown tells you that the graph might be incomplete, not where. Surfacing which specific relationships couldn't be resolved would make the "unknown" case actionable instead of just cautionary.
  • The scoring weights are hand-tuned constants. They produce intuitive results, but they're my intuitions. Making them configurable per-team, or learning them from which analyses actually preceded incidents, is the obvious next step, and deliberately out of scope for a v0.1.
  • One repo, one deploy story. npm workspaces was right for now, but if the frontend and backend release cadences diverge, the single-repo/single-deploy model will start to chafe.

Closing

If there's one thread running through all of this, it's that the boring engineering choices were the ones that paid off. The canonical format, the strict core-depends-on-nothing package boundary, dependency injection over mocking, property tests over hand-picked examples: none of them are clever. They're just decisions that made the next decision easier. Testability wasn't something bolted on at the end; it was a constraint that shaped the code from the first commit, which is why "349 tests, zero module mocks" is a description of the design rather than a heroic afterthought. The bugs that got through weren't in the logic I tested. They were at the boundary where my code met someone else's runtime, and that's exactly where I'd tell anyone building something similar to spend their paranoia.

And that closes the series. Four articles, one arc:

If any of this made you want to see the blast radius of your own next change, the project is open source. Deploy the stack, point it at a changeset, and tell me where it's wrong. Bug reports and pull requests are genuinely welcome. Thanks for reading the whole way through.

Top comments (0)