DEV Community

Ishan Naik
Ishan Naik

Posted on Originally published at github.com

AI Coding Agents Are Leaking Your Secrets: Building a Local Pre-Commit DLP for Claude Code

AgentSweep scanning and redacting secrets

You paste a .env file into Claude Code to debug a database connection. You remove the credentials from your next message, clean the repository, and rotate the database password. Your local conversation history still contains the original paste.

Claude Code keeps JSONL transcripts beneath ~/.claude/projects/. Codex keeps session JSONL beneath ~/.codex/sessions/. A process running as your user can read those files without asking either agent for permission.

I built AgentSweep to scan that history, report credential-shaped strings, and redact them without breaking the history format. I also added a pre-commit hook so developers can check this separate store of secrets before committing code.

The headline describes a credential exposure, not proof that your coding assistant has sent a key to an attacker. Developers create the exposure when they paste secrets or let an agent capture secret-bearing tool output. Attackers exploit the copy left on disk.

The credential store outside your repository

A repository scanner checks the files you stage or commit. Your agent's history lives elsewhere:

Developer home
|
+ .claude/
|  + projects/
|     + <encoded-project-path>/
|        + <session>.jsonl
|        + conversations/<session>.jsonl  (layout varies)
|
+ .codex/
|  + sessions/<year>/<month>/<day>/rollout-*.jsonl
|  + history.jsonl
|
+ workspace/
   + .env
   + .git/
Enter fullscreen mode Exit fullscreen mode

Agent versions use different layouts. AgentSweep walks Claude Code's project trees for *.jsonl rather than assuming that a conversations/ directory exists. It also honors CLAUDE_CONFIG_DIR and CODEX_HOME, because a scan of the default home misses a developer's relocated profile.

Developers accumulate more than API keys in these records: PostgreSQL URLs with passwords, AWS access key IDs, GitHub tokens, private-key blocks, and wallet recovery phrases. An assistant response can repeat a credential from the prompt. Tool output can introduce another copy.

Cursor, Windsurf, and Aider keep local history too, but their storage formats differ. Aider uses Markdown history; other integrations use JSON or SQLite. Treating all coding assistants as JSONL producers would miss those stores. AgentSweep uses source adapters for discovery, string extraction, and format-specific redaction.

Rotating a key invalidates that credential at its provider. It does not remove its bytes from a transcript, backup, or filesystem snapshot. Cleaning Git history does not touch these directories either.

An install script runs with your filesystem access

An attacker who compromises a package can run code during installation. An npm postinstall script inherits the installing user's access. Attackers can use Python package build or installation execution paths for the same purpose; PyPI does not use npm's lifecycle-hook names.

Researchers have documented this class of theft. In Aikido's analysis of the compromised @bitwarden/cli@2026.4.0, the malicious preinstall payload targeted .env, cloud credentials, ~/.claude.json, and ~/.claude/mcp.json, among other files. That report establishes theft of AI-tool configuration credentials. It does not establish that this particular payload harvested conversation JSONL.

A transcript harvester needs the same filesystem access. The diagram below models that attack path:

flowchart TD
    D[Developer installs an npm dependency] --> P[Compromised package runs postinstall]
    P --> U[Payload runs as the developer's user]
    U --> H[Read files beneath ~/.claude/projects/]
    H --> J[Parse JSONL prompts and responses]
    J --> K[Extract tokens and database credentials]
    K --> E[Send credentials to attacker endpoint]
    E --> A[Attacker uses credentials that remain valid]

No exploit in JSONL parsing is necessary. The attacker reads an ordinary file. An owner-only permission mode helps against other local users, but it does not stop malware running under your account.

AgentSweep reduces the stored-history exposure. It cannot undo exfiltration, remove a provider's copy of a prompt, or stop malware from reading a live .env file.

Five stages, with an explicit write boundary

I separated discovery and detection from mutation. scan ends with a findings report. fix adds the redaction step, and developers still need to revoke exposed credentials themselves.

flowchart LR
    D[1. Discover: walk source history roots] --> S[2. Scan: Aho-Corasick and regex rules plus BIP-39]
    S --> F[3. Findings: masked report and locations]
    F --> C{Developer selects fix}
    C -->|No| X[Stop without modifying history]
    C -->|Yes, after safety checks| R[4. Redact: backup, temp write, validation, atomic replace]
    R --> K[5. Revocation guidance: provider URLs]

During scanning and redaction, AgentSweep makes zero network requests. It reads local history, performs local matching, and writes local results. Installation downloads packages, and opening a revocation URL uses your browser. The separate update command contacts PyPI. None of those actions belongs to the scan or redaction pipeline.

I chose that boundary because a credential-cleanup tool should not upload the material it inspects.

Discovery: extract strings without guessing the schema

For JSONL, a source adapter parses records and yields string values with their locations. The scanner can associate a finding with a file, physical line, and nested key path. The redactor can then reach that value in the parsed record.

The source layer matters for SQLite and Markdown too. A database adapter needs row-aware updates and an integrity check. A text adapter needs line-preserving replacement. Sharing detection rules does not require sharing a serialization strategy.

Scan: avoid running 207 patterns on every string

AgentSweep documents 207 regex rules, plus mnemonic detection. Applying each pattern to each prompt would repeat a lot of work on ordinary source code and prose.

I use Aho-Corasick pre-filtering to identify rule keywords in a shared pass. The matcher represents those keywords in a trie with failure links. After finding a keyword, the scanner selects the associated rules for regex evaluation.

A keyword hit does not prove that the string contains a credential. The regex still checks the shape and boundaries. Conversely, a rule without a safe keyword gate must remain eligible without a keyword hit. Otherwise the optimization would introduce false negatives.

For wallet phrases, membership in a word list is insufficient. The BIP-39 detector checks candidate lengths of 12, 15, 18, 21, or 24 words and verifies the checksum. That rejects many stretches of English prose that happen to contain mnemonic words. A valid checksum identifies a mnemonic-shaped value; it does not prove that someone funded the associated wallet.

Regex engines: keep compatibility when adding RE2

The default installation uses Python's re. Developers can install the optional native backend:

uv tool install 'agentsweep[fast]'
Enter fullscreen mode Exit fullscreen mode

AgentSweep uses google-re2 for compatible patterns. RE2 avoids backtracking and provides linear-time matching for those expressions, which helps when developers paste large logs into a session.

RE2 does not support every Python regex construct. Lookarounds, Python-specific anchors, and Unicode semantics require care. AgentSweep retains the Python path for unsupported rules and semantic edge cases, including guarded non-ASCII inputs. It also retains that path for short strings and dense matches where native dispatch would not help.

I did not claim linear-time scanning for the entire mixed pipeline. Python fallback rules still use Python's engine.

The checked-in compatibility audit records 144 RE2 rules and 58 stdlib rules in a 202-rule snapshot. That snapshot predates the README's 207-rule count. Treat it as evidence for the routing design, not a current inventory. Rule counts need a version alongside them.

Start with a read-only report

AgentSweep requires Python 3.11 or newer. Install the CLI in an isolated tool environment, then select a source:

uv tool install agentsweep
agentsweep list-sources --detected
agentsweep scan --source claude-code --no-color
agentsweep scan --source codex --no-color
agentsweep scan --all --detected --json -o findings.json
Enter fullscreen mode Exit fullscreen mode

The scan exit codes give scripts a small contract:

Exit code Meaning
0 No findings
1 Findings exist
2 An error occurred

For a concrete report example, suppose a test transcript contains the AWS documentation sample key AKIAIOSFODNN7EXAMPLE. The following block illustrates the finding information, not an execution transcript or a byte-for-byte promise about terminal formatting:

source       claude-code
file         <project>/session.jsonl
line         1
rule         aws-access-key
preview      AKIA...MPLE
next action  Review the finding; revoke a real exposed key
scan status  Findings exist (exit 1)
Enter fullscreen mode Exit fullscreen mode

Review the report without copying full credentials into an issue or another chat. A detector can recognize credential syntax without knowing whether the provider still accepts the value. AgentSweep stays offline, so it does not test key validity.

For a known false positive, use a narrow .agentsweepignore entry. Suppressing an entire provider rule makes future leaks from that provider invisible to the hook.

Redact JSON values, preserve parseable history

A raw substitution over serialized JSON can damage escaping or consume syntax outside the intended value. I redact string values in parsed records, then serialize the records again.

Consider this illustrative one-record transcript. The AWS key comes from a public documentation example, not a working credential.

Before:

{"type":"user","message":{"content":"AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE"}}
Enter fullscreen mode Exit fullscreen mode

After:

{"type":"user","message":{"content":"AWS_ACCESS_KEY_ID=[REDACTED:aws-access-key]"}}
Enter fullscreen mode Exit fullscreen mode

Value-level diff:

- AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE
+ AWS_ACCESS_KEY_ID=[REDACTED:aws-access-key]
Enter fullscreen mode Exit fullscreen mode

The redactor preserves the record's structure. Re-serialization can change whitespace, so I do not promise byte-for-byte identity outside the secret.

For production history, close the agent and allow the recent-write window to expire before fixing:

agentsweep fix --source claude-code --allow-production
Enter fullscreen mode Exit fullscreen mode

In interactive mode, review the findings and type REDACT when prompted. The current alpha requires the production-root opt-in. Keep backups enabled.

Eight core write protections

I use eight core protections around this operation. The project also documents an alpha production gate and an audit log.

Protection Engineering reason
Parse before replacement Replace string values without editing JSON delimiters
Atomic replacement Write a sibling temporary file, flush it with fsync(), then call os.replace()
Format-aware validation Parse non-empty JSONL records and compare line counts before committing the replacement
Backup with no clobber Create .bak by default with exclusive creation and mode 0600
Path containment Refuse targets outside the selected source's history roots
Symlink refusal Refuse symlink targets rather than follow them into another file
Recent-write gate Refuse files modified within the previous 60 seconds
Running-process gate Refuse redaction when the relevant agent appears to be running

The fsync() and replacement sequence prevents a partial rewrite of the target during normal atomic-replace operation. It does not justify a blanket guarantee about storage hardware or power-loss durability on every filesystem.

The format check validates the rewritten content before os.replace() commits it. Describing that as validation after replacing the original would give readers the wrong failure model.

On POSIX, mode 0600 limits backup access to the owner. Windows uses ACLs with different semantics; do not interpret Python's mode argument as an equivalent Windows security boundary.

Developers can bypass the recent-write and process gates with --force. That flag does not bypass containment, symlink rejection, or invalid-content failures. Prefer closing the agent over overriding its concurrency checks.

Make history scanning a pre-commit checkpoint

Add the project's hook to .pre-commit-config.yaml:

repos:
  - repo: https://github.com/Ishannaik/agent-sweep
    rev: v0.1.9
    hooks:
      - id: agentsweep
Enter fullscreen mode Exit fullscreen mode

The project's README uses this released tag in its example. Review release changes before choosing a newer pin.

Install the hook:

pre-commit install
Enter fullscreen mode Exit fullscreen mode

The hook definition invokes:

agentsweep scan --all --detected
Enter fullscreen mode Exit fullscreen mode

It sets pass_filenames: false and always_run: true. Developers therefore get one history scan per commit regardless of which files they staged. With no detected history root, the scan exits clean.

The hook blocks a commit when the detector finds a credential-shaped value. It does not redact anything during the commit. Developers review the report, revoke real exposed credentials, and run a separate fix operation.

This is a local data-loss-prevention checkpoint with limits. A developer can bypass Git hooks. Malware can read a transcript before the next commit. A clean result means the scanner found nothing within the selected sources and rules, not that the workstation contains no secrets.

Keep a staged-file secret scanner alongside it. The two scans inspect different stores.

Revoke the key, then remove the backup

A redaction changes your local history. The provider still accepts a live key until you revoke it.

AgentSweep includes provider-specific guidance. You can inspect a rule without scanning:

agentsweep explain stripe-live
Enter fullscreen mode Exit fullscreen mode

For common findings, developers can review GitHub token settings, OpenAI API keys, or Anthropic API keys. AWS credentials require the appropriate IAM access-key workflow. AgentSweep prints guidance; it does not invoke those providers' APIs.

Backups contain the original plaintext. A same-user attacker who can read the transcript can also read its .bak file. After rotating the exposed credentials and confirming that you no longer need recovery, purge the backups:

agentsweep purge --source claude-code
Enter fullscreen mode Exit fullscreen mode

If you need recovery before that point:

agentsweep undo --source claude-code
Enter fullscreen mode Exit fullscreen mode

Undo restores the secret-bearing original. Purge removes the backup files; it does not promise secure erasure from SSD blocks, snapshots, or cloud backup systems.

My workflow is to avoid pasting secrets, scan existing history, review findings, revoke exposed keys, redact local copies, and remove recovery backups when I no longer need them. The pre-commit hook provides another checkpoint as new history accumulates.

You can inspect the source, detection rules, and safety code in Ishannaik/agent-sweep. If you report a missed credential format, share a synthetic example with the same shape. Do not attach the transcript you are trying to clean.

Written by Ishan Naik.

Top comments (0)