<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Tushar Aggarwal</title>
    <description>The latest articles on DEV Community by Tushar Aggarwal (@aggtushar123).</description>
    <link>https://dev.to/aggtushar123</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4120376%2Fb1c85c65-2865-4e47-b218-b8c87208c12c.jpg</url>
      <title>DEV Community: Tushar Aggarwal</title>
      <link>https://dev.to/aggtushar123</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aggtushar123"/>
    <language>en</language>
    <item>
      <title>I Wanted Claude to Write SOC 2 Compliant Code. The Existing Skills Only Did Half the Job.</title>
      <dc:creator>Tushar Aggarwal</dc:creator>
      <pubDate>Mon, 14 Sep 2026 11:33:57 +0000</pubDate>
      <link>https://dev.to/aggtushar123/i-wanted-claude-to-write-soc-2-compliant-code-the-existing-skills-only-did-half-the-job-18e7</link>
      <guid>https://dev.to/aggtushar123/i-wanted-claude-to-write-soc-2-compliant-code-the-existing-skills-only-did-half-the-job-18e7</guid>
      <description>&lt;p&gt;Every SOC 2 project I have seen follows the same shape. Someone buys a compliance platform, someone drafts eighteen policies, and then six months later an auditor asks to see the code path where tenant A is prevented from reading tenant B's invoices. Everyone looks at the one engineer who wrote that middleware two years ago.&lt;/p&gt;

&lt;p&gt;The policies say the right things. The control matrix maps to the right criteria. But the code was written by people who never read either document, because nobody expected them to.&lt;/p&gt;

&lt;p&gt;I wanted to close that gap at the point where it opens: the moment a developer types a new route handler. If Claude is already writing a lot of that code, Claude should know what CC6.1 requires before it writes &lt;code&gt;app.post('/invoices', ...)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;So I went looking for a Claude skill that did this. I found two good ones, and neither was built for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was already out there
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/kurianoff/claude-skills-soc2-policies" rel="noopener noreferrer"&gt;kurianoff/claude-skills-soc2-policies&lt;/a&gt;&lt;/strong&gt; is a set of three skills for managing policy documents. It ships 17 ready-to-use SOC 2 policy templates, a dashboard for tracking review status, a statement-by-statement review carousel with AI-assisted rewrites, and a Word export with title page, version history, and an audit-trail appendix. If your problem is "we need policies and we need to prove someone reviewed them," it is genuinely well built.&lt;/p&gt;

&lt;p&gt;Two things kept it from solving my problem. It runs only in claude.ai Projects, because it depends on interactive widgets, per-project storage, and file download APIs that Claude Code does not have. And it is entirely about documents. The templates do not reference Trust Service Criteria, and nothing in it looks at a codebase.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/alirezarezvani/claude-skills" rel="noopener noreferrer"&gt;alirezarezvani/claude-skills&lt;/a&gt;&lt;/strong&gt; has a &lt;code&gt;soc2-compliance&lt;/code&gt; skill in its regulatory-affairs collection that is the other half. It has a full Trust Service Criteria reference from CC1 through CC9 plus the four optional categories, a control matrix builder, a gap analyzer for Type I and Type II, an evidence tracker with readiness scoring, and solid guides on evidence collection, vendor tiers, and continuous compliance. It runs anywhere, since it is markdown and standalone Python.&lt;/p&gt;

&lt;p&gt;What it does not do is touch code. The scripts generate and score control matrices from JSON you hand them. The word "policy" appears a handful of times, as an evidence type. There is nothing that says "when you write a route, do this."&lt;/p&gt;

&lt;p&gt;Put side by side, the two barely overlap. One is policies, one is controls and evidence. Neither is engineering.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built instead
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/aggtushar123/soc2-skills" rel="noopener noreferrer"&gt;soc2-dev&lt;/a&gt; is a single folder you copy into any repo's &lt;code&gt;.claude/skills/&lt;/code&gt;. It is MIT licensed and reuses content from both projects above with credit: the 17 policy templates converted to markdown, and the Trust Service Criteria reference.&lt;/p&gt;

&lt;p&gt;The piece that makes it different is a &lt;strong&gt;requirement registry&lt;/strong&gt;. Each Trust Service Criterion is translated into concrete engineering rules with IDs, a severity, and the criteria it serves. There are 66 of them across eight domains:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;ID&lt;/th&gt;
&lt;th&gt;Requirement&lt;/th&gt;
&lt;th&gt;TSC&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AUTH-01&lt;/td&gt;
&lt;td&gt;Every non-public endpoint requires authentication. Deny by default.&lt;/td&gt;
&lt;td&gt;CC6.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AUTH-02&lt;/td&gt;
&lt;td&gt;Authorization is enforced server-side per object, never from client-supplied roles or IDs.&lt;/td&gt;
&lt;td&gt;CC6.1, CC6.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API-03&lt;/td&gt;
&lt;td&gt;Database, shell, and template calls use parameterization. No string-built queries.&lt;/td&gt;
&lt;td&gt;CC6.1, CC7.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DATA-08&lt;/td&gt;
&lt;td&gt;Multi-tenant data access is scoped by tenant on every query.&lt;/td&gt;
&lt;td&gt;CC6.1, C1.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LOG-03&lt;/td&gt;
&lt;td&gt;Logs never contain passwords, tokens, or unmasked PII. A redaction layer runs before shipping.&lt;/td&gt;
&lt;td&gt;C1.2, CC6.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SEC-07&lt;/td&gt;
&lt;td&gt;Approved cryptography only. Keys live in a KMS, never alongside the data they protect.&lt;/td&gt;
&lt;td&gt;CC6.6, C1.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CHG-01&lt;/td&gt;
&lt;td&gt;Protected default branch: PR required, non-author approval, no force pushes.&lt;/td&gt;
&lt;td&gt;CC8.1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Everything else in the skill keys off these IDs. The domain references explain how to implement each one with code in Node, Python, and Go. The stack pattern files give runnable middleware, guards, validators, audit loggers, and encryption helpers for Express, FastAPI, Go, and Spring. The scanner reports findings by ID. The PR template asks about them by ID.&lt;/p&gt;

&lt;h3&gt;
  
  
  Write mode
&lt;/h3&gt;

&lt;p&gt;This is the default and the reason the skill exists. When you ask Claude for a new endpoint in a repo that has the skill installed, it classifies the change, loads only the domain checklists that apply, implements the requirements alongside the feature, and annotates each enforcement point so the control is greppable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;router&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/invoices&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;requireAuth&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                                  &lt;span class="c1"&gt;// SOC2:AUTH-01 deny by default&lt;/span&gt;
  &lt;span class="nf"&gt;validate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;CreateInvoiceSchema&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;                &lt;span class="c1"&gt;// SOC2:API-01 strict schema, unknown fields rejected&lt;/span&gt;
  &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// SOC2:AUTH-02 + DATA-08 tenant comes from the verified token, never the body&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;invoice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;invoices&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;auth&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;audit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;event_type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;invoice.created&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;actor_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;auth&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="na"&gt;target_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;tenant_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;auth&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="na"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;success&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;correlation_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;   &lt;span class="c1"&gt;// SOC2:LOG-01&lt;/span&gt;
    &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;201&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It then updates &lt;code&gt;.soc2/CONTROL_MAP.md&lt;/code&gt;, which maps each requirement to the file and symbol that implements it, the evidence an auditor can inspect, and an owner. And it ends with a summary you paste into the PR:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOC 2 controls in this change
- Implemented: AUTH-01, AUTH-02, API-01, API-05, LOG-01, DATA-08
- Reused existing: requireAuth (AUTH-01), auditLog (LOG-01/02)
- Skipped: API-08 (endpoint is not retried by clients)
- Exceptions applied: none
- CONTROL_MAP.md rows updated: 6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rules Claude follows are simple and they are written down in the skill. Deny by default. Reuse an existing primitive before writing one. Never add a control the change does not touch. Never annotate a line that does not actually enforce anything. If a requirement cannot be met, record an exception with a compensating control and an expiry, never silently skip it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Audit mode
&lt;/h3&gt;

&lt;p&gt;A dependency-free Python scanner finds the violations that are findable by pattern: secrets in source, unauthenticated routes, string-built SQL, PII in logs, wildcard CORS with credentials, MD5 for passwords, TLS verification switched off, open security groups, public databases, root containers, and missing CI gates. Every finding carries a requirement ID, a severity, a confidence, and a fix hint.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;| Sev  | Conf   | Location            | Finding                                         |
|------|--------|---------------------|-------------------------------------------------|
| high | medium | src/app.js:7        | Route has no visible authentication: GET /users/:id |
| high | medium | src/app.js:8        | SQL string concatenated with a variable          |
| high | medium | src/app.js:10       | Log statement references req.body.password       |
| critical | high | .env:1          | Connection string with password found in source |
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It runs as a CI gate with &lt;code&gt;--fail-on high&lt;/code&gt;, and as a pre-commit hook with &lt;code&gt;--staged&lt;/code&gt;, which scans only the change set while still reading the whole repo for stack and middleware detection so a staged route is not falsely reported as unauthenticated because its middleware lives in another file.&lt;/p&gt;

&lt;p&gt;The scanner is honest about what it cannot see. Object-level authorization, tenant scoping, field-level encryption, CSRF, and session timeouts are invisible to any regex. The skill's audit mode names those explicitly as manual-review items, and they are where most real audit exceptions come from.&lt;/p&gt;

&lt;h3&gt;
  
  
  Evidence and policy modes
&lt;/h3&gt;

&lt;p&gt;Evidence mode regenerates the control map, a routes manifest listing every endpoint with its auth requirement and data classification, and an evidence index that says where each control's proof lives: a CI job name, a branch protection export, a log query, an IAM policy dump.&lt;/p&gt;

&lt;p&gt;Policy mode drafts organizational policies from the 17 templates and appends a "Technical enforcement" section to each one listing the requirement IDs that implement it. The Access Control Policy points at AUTH-01 through AUTH-10. The Cryptography Policy points at SEC-07 and DATA-02. That is the link the two earlier projects never made: policy text on one side, enforcement points on the other, with IDs in between.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the earlier repos had that I kept, and what was missing
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;kurianoff&lt;/th&gt;
&lt;th&gt;alirezarezvani&lt;/th&gt;
&lt;th&gt;soc2-dev&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Policy templates&lt;/td&gt;
&lt;td&gt;17, HTML, with review workflow&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;the same 17, markdown, mapped to requirement IDs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TSC reference&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;full CC1-CC9, A1, C1, PI1, P1-P8&lt;/td&gt;
&lt;td&gt;reused with credit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Control matrix and gap analysis&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;scripts over JSON you supply&lt;/td&gt;
&lt;td&gt;derived from the code itself via scanner and control map&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code-level requirements&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;66, with per-stack implementations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enforced while writing code&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;yes, that is the default mode&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Traceability from code to criterion&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;SOC2:&amp;lt;ID&amp;gt;&lt;/code&gt; annotations plus control map&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runs in Claude Code&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tests and CI for the tool itself&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;partial&lt;/td&gt;
&lt;td&gt;24 regression tests, three Python versions, self-scan, gitleaks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I want to be fair to both projects. The kurianoff review carousel is a better policy-review experience than anything I built, and if you are in claude.ai Projects it is worth using for that. The alirezarezvani TSC reference is thorough enough that I did not rewrite it. What neither one had was the idea that compliance is something a developer does at the keyboard, not something a compliance team does to the codebase afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned shipping it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Regex scanners lie in both directions.&lt;/strong&gt; The first pass flagged &lt;code&gt;"DELETE " + url&lt;/code&gt; as SQL injection and &lt;code&gt;next_token&lt;/code&gt; as a leaked secret. A teammate's improvements to the scanner added coverage and also added three high-severity false positives that would have blocked ordinary commits. The fix was a regression suite with a false-positive fixture that must produce nothing at medium or above, alongside the true-positive fixture where every planted violation must fire with the right ID. Scanner changes now require both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A gate that can silently pass is worse than no gate.&lt;/strong&gt; Staged mode resolved git paths against the wrong directory, so running it from a subdirectory reported zero findings and exited clean. On macOS, temp directories go through a symlink and the same thing happened. Tests caught it. If you ship a compliance tool without tests, the first auditor who runs it against your own repo will find that out for you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your own tool should pass its own scan.&lt;/strong&gt; The repo runs the scanner on itself in CI. The first run flagged that the repo had no CI, no PR template, and no CODEOWNERS. That was uncomfortable and correct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub push protection will reject your test fixtures.&lt;/strong&gt; A fake Stripe key in a test file blocked the push. The fixture is now assembled at runtime so the committed source never contains a string that matches a live-key pattern. Gitleaks then flagged a fake Dockerfile token in the same file. An allowlist scoped to the fixture directory fixed it without exempting anything else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/aggtushar123/soc2-skills.git /tmp/soc2-skills
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; .claude/skills
&lt;span class="nb"&gt;cp&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; /tmp/soc2-skills/skills/soc2-dev .claude/skills/soc2-dev
python3 .claude/skills/soc2-dev/scripts/soc2_init.py &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;--company&lt;/span&gt; &lt;span class="s2"&gt;"Your Co"&lt;/span&gt;
python3 .claude/skills/soc2-dev/scripts/soc2_scan.py &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt; md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then, in Claude Code, ask for something like "add a POST /invoices endpoint for the current tenant" or "audit this repo for SOC 2" and watch what it loads.&lt;/p&gt;

&lt;p&gt;It is v0.1.0 and I am calling it a public beta. The Claude-facing half is stable. The scanner's patterns will keep being tuned, and the test suite is there so tuning does not regress. Issues and pull requests are welcome, especially stack patterns for frameworks I have not covered.&lt;/p&gt;

&lt;p&gt;SOC 2 is still an attestation over organizational controls and evidence gathered over an observation period. No skill makes you compliant. But the engineering half can be done right from the first commit instead of reconstructed at audit time, and that is the half this is for.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Repo: &lt;a href="https://github.com/aggtushar123/soc2-skills" rel="noopener noreferrer"&gt;github.com/aggtushar123/soc2-skills&lt;/a&gt;. Credits to &lt;a href="https://github.com/kurianoff/claude-skills-soc2-policies" rel="noopener noreferrer"&gt;Constantine Kurianoff&lt;/a&gt; for the policy templates and &lt;a href="https://github.com/alirezarezvani/claude-skills" rel="noopener noreferrer"&gt;Alireza Rezvani&lt;/a&gt; for the Trust Service Criteria reference, both MIT.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>devops</category>
      <category>soc2</category>
    </item>
    <item>
      <title>The bug where every check passed and the data was still wrong</title>
      <dc:creator>Tushar Aggarwal</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:14:30 +0000</pubDate>
      <link>https://dev.to/aggtushar123/the-bug-where-every-check-passed-and-the-data-was-still-wrong-5fo2</link>
      <guid>https://dev.to/aggtushar123/the-bug-where-every-check-passed-and-the-data-was-still-wrong-5fo2</guid>
      <description>&lt;p&gt;I'm building [&lt;a href="https://github.com/aggtushar123/mongopg-migrate" rel="noopener noreferrer"&gt;mongopg-migrate&lt;/a&gt;], a tool that migrates MongoDB collections onto an &lt;em&gt;existing&lt;/em&gt; Postgres schema you already designed. It's alpha &lt;code&gt;pip install mongopg-migrate&lt;/code&gt;, currently v0.2.0. This is a bug from it, the worst one I've hit so far, because every safeguard the tool has fired correctly, and the migration still came out wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;The tool supports nested arrays (&lt;code&gt;explode:&lt;/code&gt;): a Mongo array becomes a child table, and an array inside that array becomes a grandchild table. A field at any level can also be a &lt;code&gt;lookup:&lt;/code&gt;, resolved against another entity's already-migrated rows and rewritten as a foreign key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;explode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;facilities&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;explode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;categoryParts&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;fields&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;categoryId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;lookup&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;zcategories&lt;/span&gt;   &lt;span class="c1"&gt;# one level deeper than the top&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things need to know about that lookup: &lt;code&gt;entity_dependencies()&lt;/code&gt; (so &lt;code&gt;zcategories&lt;/code&gt; loads &lt;em&gt;before&lt;/em&gt; whatever references it) and &lt;code&gt;validate_structure()&lt;/code&gt; (so a typo'd entity name gets caught early). Both only ever walked the first level of &lt;code&gt;explode&lt;/code&gt;. A lookup one level deeper was invisible to both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;lookup at TOP explode level     -&amp;gt; order ['zcategories', 'hospitals']   correct
same lookup ONE LEVEL DEEPER    -&amp;gt; order ['hospitals', 'zcategories']   backwards
validate on typo'd nested lookup -&amp;gt; []   zero issues found
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On its own, that's a real bug. But it used to be a &lt;em&gt;loud&lt;/em&gt; one, a lookup with an empty id_map raised an unconditional &lt;code&gt;LoadError&lt;/code&gt; and crashed the run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The twist
&lt;/h2&gt;

&lt;p&gt;Separately, the tool had just grown an &lt;code&gt;on_missing&lt;/code&gt; policy: what to do when a reference genuinely doesn't resolve (source doc deleted, a normal case). &lt;code&gt;on_missing: null&lt;/code&gt; writes NULL instead of crashing the whole migration. Reasonable on its own.&lt;/p&gt;

&lt;p&gt;Combine both:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Wrong order schedules the referencing entity &lt;strong&gt;before&lt;/strong&gt; &lt;code&gt;zcategories&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Every lookup misses, because &lt;code&gt;zcategories&lt;/code&gt;'s id_map is empty, not because anything is actually dangling.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;on_missing: null&lt;/code&gt;, built for real dangling references, fires on every miss and writes NULL.&lt;/li&gt;
&lt;li&gt;Row counts are unaffected - NULLs don't change counts.&lt;/li&gt;
&lt;li&gt;The post-migration validator re-checks dangling references &lt;em&gt;after&lt;/em&gt; the full run, by which point &lt;code&gt;zcategories&lt;/code&gt; has loaded, so it finds nothing wrong and reports clean.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;End state: a whole foreign-key column silently NULL, and count diff, dangling-reference check, and structural validation all report success. Each check answered the exact question it was built to answer, correctly. The chain connecting them was wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  the fix: two parts, not one
&lt;/h2&gt;

&lt;p&gt;A regression test for this exact repro would fix the one shape found and miss the actual assumption: a policy for "this one reference is dangling" isn't a safe answer to "the entity it points at has never loaded a single row." Those are different failure modes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause&lt;/strong&gt; recurse through every level of nested &lt;code&gt;explode&lt;/code&gt;, not just the first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_explode_lookup_targets&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;explode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ExplodeSpec&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;targets&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;exp&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;explode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;fspec&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fields&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;fspec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lookup&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;targets&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fspec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lookup&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;targets&lt;/span&gt; &lt;span class="o"&gt;|=&lt;/span&gt; &lt;span class="nf"&gt;_explode_lookup_targets&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;explode&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# recurse
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;targets&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Defense in depth&lt;/strong&gt; even with correct ordering, refuse to apply &lt;code&gt;on_missing&lt;/code&gt; blind. Check whether the referenced entity has loaded &lt;em&gt;anything at all&lt;/em&gt; first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;on_missing&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;OnMissing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ERROR&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;idmap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has_any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lookup_conn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lookup_entity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;lookup_schema&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;LoadError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;lookup_entity&lt;/span&gt;&lt;span class="si"&gt;!r}&lt;/span&gt;&lt;span class="s"&gt; has NO id_map rows at all — this looks like a load-order &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bug or a forgotten prerequisite run, not a genuinely dangling reference. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refusing to apply on_missing=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;on_missing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="si"&gt;!r}&lt;/span&gt;&lt;span class="s"&gt; here.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That second check matters independently; it also catches referencing an entity from a &lt;em&gt;separate&lt;/em&gt; migration run that a human simply never ran. No amount of correct ordering fixes that; checking "has this entity loaded anything, ever" does. It's cached per entity, so it's one extra indexed query the first time a lookup to that entity misses, not one per row.&lt;/p&gt;

&lt;h2&gt;
  
  
  the same shape, twice more
&lt;/h2&gt;

&lt;p&gt;I first wrote this up as a one-off. Two bugs since have changed my mind.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two checks aimed at one mistake are one check.&lt;/strong&gt; &lt;code&gt;--pg-schema&lt;/code&gt; exists so you can migrate into a schema other than &lt;code&gt;public&lt;/code&gt;. It reached &lt;code&gt;introspect_postgres()&lt;/code&gt; and nothing else. Every write named its table with no schema qualifier, so rows landed wherever the connecting role's &lt;code&gt;search_path&lt;/code&gt; pointed, normally &lt;code&gt;public&lt;/code&gt;. Nothing errored; the load was perfectly valid against the tables it found. Then &lt;code&gt;validate&lt;/code&gt; counted those &lt;em&gt;same wrong tables&lt;/em&gt; and printed a clean pass:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[OK] hospitals (hospitals): mongo=4002 postgres=4002
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The migration reported success. The validation agreed. Both were pointed at the wrong tables. It had been there since the flag was added, and it could only ever have hurt someone whose tables don't live in &lt;code&gt;public&lt;/code&gt;, which is to say, someone who wasn't me. All three commands now set &lt;code&gt;search_path&lt;/code&gt; explicitly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A precondition that only held at fixture scale.&lt;/strong&gt; To make &lt;code&gt;lookup:&lt;/code&gt; fast over a slow link, the tool prefetches a referenced entity's whole id_map into memory before the first document. That was unbounded, at ~200 bytes a row, ~1.2 GB at 5M rows and ~12 GB at 50M, allocated up front. It never came close in development, because the largest entity there was ~65k rows. There's now a ceiling (&lt;code&gt;--idmap-prefetch-max&lt;/code&gt;, default 2,000,000); above it, lookups go per row behind a bounded cache, slower, but it can't exhaust memory, and the run says which mode it picked.&lt;/p&gt;

&lt;p&gt;Same pattern all three times: logic that's correct &lt;em&gt;given&lt;/em&gt; an assumption nobody wrote down, composed with other correct-in-isolation logic, producing something wrong that looks clean.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lesson
&lt;/h2&gt;

&lt;p&gt;"Every check passed" is a claim about which checks exist, not a claim about correctness. And two checks that share an assumption are, for catching a violation of that assumption, one check.&lt;/p&gt;

&lt;p&gt;Worth asking of any &lt;code&gt;null&lt;/code&gt;/&lt;code&gt;skip&lt;/code&gt;/&lt;code&gt;retry&lt;/code&gt; fallback in a system: what does this do if the precondition I'm assuming turns out to be silently false?&lt;/p&gt;




&lt;p&gt;Repo: &lt;a href="https://github.com/aggtushar123/mongopg-migrate" rel="noopener noreferrer"&gt;https://github.com/aggtushar123/mongopg-migrate&lt;/a&gt; - alpha, building in the open.&lt;br&gt;
&lt;code&gt;pip install mongopg-migrate&lt;/code&gt;.&lt;/p&gt;




</description>
      <category>postgres</category>
      <category>mongodb</category>
      <category>python</category>
      <category>debugging</category>
    </item>
  </channel>
</rss>
