✓ Human-authored analysis; AI used for formatting and proofreading.
We ran a cloud security analyzer against Rhino Security Labs' CloudGoat. A vulnerable AWS environment designed to train red teams. The analyzer found the documented attack path. It also exposed five bugs in itself that no unit test had caught.
This is about what happens when you test a tool against adversary-designed infrastructure instead of hand-crafted fixtures and why the bugs that surface are structurally different from the bugs unit tests find.
Why labs find different bugs
A unit test verifies a property the developer thought to check. The developer writes the input, writes the expected output, and checks that the function maps one to the other. The bugs it catches are the ones the developer anticipated.
A lab creates conditions the developer didn't anticipate. An adversary-designed scenario has pre-existing infrastructure, unexpected resource combinations, and configurations that exist because an attacker would exploit them. The bugs that surface aren't logic errors in individual functions. They're interaction effects between subsystems that only appear when the input has properties the developer never specified in a test.
All five bugs below have the same structure: each subsystem works correctly in isolation, tested and passing. The bug emerges from the interaction between subsystems when the input has a property that unit test did not specify.
Bug 1: Missing evidence treated as confirmed exposure
What happened: A control that checks for S3 prefix-level public access fired on a SAM CLI bucket that had no prefix exposure data. The finding rendered as a high-severity violation with exposure_source: missing_evidence.
Why unit tests missed it: The exposure evaluator was tested with present evidence (both exposed and not-exposed cases). No test provided absent evidence. A resource where the prefix exposure data simply doesn't exist. Because the developer's fixtures always included it.
The interaction: The evaluator treated missing evidence as "fail closed" Exposed: true. In isolation, fail-closed is a defensible default for a security tool. But the downstream rendering pipeline treated Exposed: true as a confirmed violation and produced a finding with a severity score. The result: a confident-looking high-severity finding backed by no actual evidence. The evaluator made a reasonable default. The renderer made a reasonable assumption. Together, they produced a false positive.
The fix: Added an Inconclusive field to the compliance report. Missing evidence now returns Inconclusive: true, Exposed: false. The engine sets VerdictInconclusive instead of VerdictViolation. No finding is produced for inconclusive verdicts. The distinction between "confirmed safe," "confirmed unsafe," and "unable to determine" is now explicit in the type system.
Why it matters beyond this tool: Every tool that evaluates security state has this decision point: what do you do when the evidence is missing? Fail-open misses real problems. Fail-closed produces false positives. The third option of making inconclusiveness an explicit, typed verdict is the one that prevents confident-looking findings from being backed by absent data.
Bug 2: Risk signals ignored asset-type scope
What happened: Controls like AUTOSCALING.INCOMPLETE.001 and GUARDDUTY.INCOMPLETE.001 appeared in risk_signals for security groups and IAM roles. These are asset types those controls don't target. The main findings path was correct (zero findings from wrong types). Only the risk-signals path was broken.
Why unit tests missed it: The risk-signals feature had its own tests, and they passed. The main evaluation engine had its own tests, and they passed. Neither test suite covered the case where both paths process the same control against the same asset. Because they were tested independently.
The interaction: The main evaluation engine in lifecycles.go checks AppliesToAssetType(a.Type) before evaluating a control against an asset. The risk-signals path in upcoming.go checks AppliesToVendor but not AppliesToAssetType. Same control, same asset, two code paths, different scope gates. The main path filtered correctly. The secondary path didn't.
The fix: One line. Added ctl.AppliesToAssetType(a.Type) in upcoming.go, mirroring the identical check in lifecycles.go. Risk signals dropped from 100+ phantom entries to 13 legitimate signals.
Why it matters beyond this tool: When two code paths evaluate the same data for different purposes (findings vs. risk signals), scope gates must be identical in both. This is a composition bug. Each path is correct in isolation, and the discrepancy only surfaces when the input contains asset types that one path filters and the other doesn't. A lab with diverse asset types (security groups + IAM roles + S3 buckets) creates that condition. A unit test with a single asset type doesn't.
Bug 3: 27 controls missing asset-type declarations
What happened: Even after fixing Bug 2, INCOMPLETE controls for SQS, DynamoDB, ECR, EFS, ECS, RDS, and others still appeared in risk signals for every asset type. Because they had no applicable_asset_types declaration at all.
Why unit tests missed it: The controls compiled. They loaded and evaluated. No test checked whether a control's metadata was complete. It only tested whether it functioned. A control that functions correctly but has no scope declaration is invisible to a scope-checking gate (the gate passes everything with no declaration).
The fix: Added applicable_asset_types to all 27 controls with the correct asset type per service. 55 files changed (27 controls + 27 embedded copies + README).
Why it matters beyond this tool: Schema completeness and functional correctness are independent properties. A control can be functionally perfect (its predicate logic is correct) and metadata-incomplete (it doesn't declare what it applies to). Unit tests check function. Labs check completeness. Because labs provide the diverse inputs that make missing metadata visible.
Bug 4: Cache didn't invalidate on engine behavior change
What happened: After fixing Bugs 1 and 2 and bumping the evaluation version from 1.0 to 1.1, the output's policy fingerprint was byte-identical to the pre-fix run. Two different engine semantics produced the same fingerprint.
Why unit tests missed it: The cache had tests. The fingerprint had tests. Neither test verified that changing the engine version invalidates the cache. The cache keyed on (StaveVersion, ControlsDigest, InputHashesHash, ChainsDigest, ConfigDigest, SourcePaths) — six fields. EvalVersion wasn't one of them. During development, StaveVersion="edge" never changes, so engine code changes silently served cached results from the previous engine version.
The interaction: Three subsystems: the cache key builder, the fingerprint hasher, and the eval-version flag. Each worked correctly in its own scope. But the cache key builder didn't include the eval-version, so the fingerprint hasher never ran on new inputs. The cache returned the old fingerprint before the hasher was invoked. A behavior change in the engine (the whole point of bumping the version) was invisible in the output.
The fix: Added EvalVersion to the cache key struct and to the hash computation. Engine behavior changes now invalidate the cache. A new command (stave fingerprint explain) shows the exact preimage bytes fed to SHA256, so cache behavior is inspectable rather than inferred.
Why it matters beyond this tool: Content-addressed caches are only as correct as their key. Every field that affects the output must be in the key. This is the caching equivalent of a hash collision. This is a collision in the key space because a relevant dimension was omitted. The bug is invisible in testing because test suites don't typically change engine versions between runs and check the cache. A lab where you're actively fixing bugs and re-running creates exactly the condition where the cache serves stale results.
Bug 5: Fingerprint preimage not inspectable
What happened: Debugging Bug 4 required adding temporary print statements and manually clearing the cache. No command existed to inspect what the policy fingerprint actually certifies.
Why it mattered: The fingerprint is the system's reproducibility proof: "this output was produced by this exact combination of engine version, control catalog, and input data." If you can't inspect the preimage, you can't debug fingerprint mismatches, cache hits, or version-skew issues. The fingerprint is opaque.
The fix: Built stave fingerprint explain. This prints the exact preimage bytes, plus diagnostic flags (eval_version_present, asset_types_hashed per control). Turns an opaque hash into an inspectable proof.
Why it matters beyond this tool: Any system that uses content-addressed hashing for reproducibility or caching needs an inspection tool for the preimage. The hash is only useful if you can verify what went into it. This is obvious in retrospect and invisible until a cache bug forces you to debug it.
The pattern
All five bugs share a structure:
Individual subsystem: tested, passing, correct in isolation
Interaction between subsystems: untested, failing under lab conditions
Lab conditions that triggered it: diverse asset types, missing evidence,
engine version changes mid-session, pre-existing infrastructure
Unit tests verify functions. Labs verify the system. The functions were correct. The system had bugs at the seams. The places where subsystems hand off to each other, where scope gates need to match, where cache keys need to be complete, where verdicts need a third state beyond true/false.
Adversary-designed labs are effective at surfacing these bugs because the infrastructure is intentionally unusual. A pentester deploying CloudGoat creates resource combinations that exist because they're exploitable. These unusual combinations hit code paths that typical fixtures don't exercise such as the missing-evidence path, the cross-asset-type evaluation path, the engine-version-change path.
What we changed after
Beyond the five fixes:
A dead-code sweep triggered by the session removed 1,181 lines across 38 files (67 dead functions found by deadcode, 66 confirmed dead, 1 false positive). make deadcode-check is now wired into the CI check suite.
The stave fingerprint explain command is permanently available. It is not a debugging aid that gets deleted after the bug is fixed. The next cache issue will be inspectable from the CLI, not from inserted print statements.
The evaluation version (1.0 → 1.1) is now part of the policy fingerprint's preimage, alongside ApplicableAssetTypes. A scope change or engine behavior change automatically changes the fingerprint. The property that was missing before.
The lesson is that unit tests and labs find different classes of bugs. Unit tests find logic errors in functions. Labs find interaction errors between subsystems. Both are necessary. The five bugs above would still be in the codebase if the only testing was unit tests. Because unit tests can't create the conditions that expose them. A lab with diverse, adversary-designed infrastructure can.
The lab scenario is CloudGoat's iam_privesc_by_attachment from Rhino Security Labs. The tool is Stave.
Top comments (0)