DEV Community

Cover image for How to audit every credential your AI agent process can reach
Conor Bronsdon
Conor Bronsdon

Posted on Originally published at chainofthought.show

How to audit every credential your AI agent process can reach

Check whether your agent already reaches more than you intend. Agents are often wired like classic software, with long-lived keys, broad tool plugins, and a service account that outlives the task. Tyler Akidau, CTO of Redpanda, puts the stall point plainly: enterprises lack the governance layer, not model quality. Before you mint another identity or buy a platform, audit what this process can already touch.

That audit is how you shrink excessive agency in the OWASP sense (LLM06:2025): one bad turn, ambiguous output, or injected instruction should not inherit the power to do real damage. The trigger can be a hallucination, prompt injection, or a compromised tool. You cannot guarantee model behavior, so you limit what behavior can do.

What counts as "reach" for this review?

Reach is anything the agent runtime can read or call, not what you meant to allow in a prompt. Jitender Aswani, of Starburst, frames the data side: a federated query engine can let an agent touch many systems without copying data into one warehouse, but the catalog still must answer where data lives and who may access it. Connecting is not the same as being authorized. Federation can widen blast radius while permissions stay vague.

For one deployment, inventory surfaces together:

  • Environment variables, secret mounts, and CI injectors tied to the agent pod or worker
  • Tool and MCP configs (including plugins shipped for dev and never removed)
  • Shared service accounts the runtime reuses across instances
  • Log and trace pipelines (tracing can write credentials into plaintext logs)

Shadow AI belongs in the same conversation when teams bypass IT: unsanctioned chatbots, extensions, or personally configured agents, and work pasted into contexts you do not govern. IBM defines shadow AI as unsanctioned employee use without formal IT approval. Microsoft and LinkedIn's 2024 Work Trend Index reported 75% of knowledge workers using generative AI at work and 78% of AI users bringing their own tools. Your official agent is one row in a wider table.

How do I list credentials and what each one grants?

Work runtime by runtime, not "our agent product" in the abstract. For each secret or token the process can read, record four fields: name (or vault path), issuer (IdP, cloud IAM, SaaS admin), grants (API scopes, DB roles, repo permissions), and expiry (or "none" if that is true).

Agents sit near real credentials: tool API keys, connection strings, OAuth tokens. Token leakage is when those secrets escape into model context, user-visible tool output, or plaintext traces. OWASP's guidance on system prompt leakage (LLM07:2025) is blunt: do not treat the system prompt as a secret control, and do not put credentials in it. If your audit finds keys in prompts or skill files, treat that as a leakage risk to fix first.

Sketch for turning logs into a review table (adapt to your vault and orchestrator):

# Sketch only: normalize what one agent worker can read
def audit_row(secret_id, issuer, scopes, expires_at, used_by_tools):
    return {
        "secret_id": secret_id,
        "issuer": issuer,
        "grants": list(scopes),
        "expires_at": expires_at,
        "tools": list(used_by_tools),
    }

# Example: walk the env vars and mounted files the worker can read today
rows = [audit_row("OPENAI_API_KEY", "platform", ["inference"], None, ["planner"])]
Enter fullscreen mode Exit fullscreen mode

Flag rows where grants exceed the current task, expiry is missing, or the same secret appears on multiple unrelated agents.

How do I classify "excess" before I fix it?

OWASP groups excessive agency into excess functionality (tools the job does not need), excess permissions (identity can do more than the task), and related failure modes. Map each audit row to a task statement: "summarize tickets", not "administer Jira".

Tyler Akidau's authorization checklist from the agent identity explainer is a good classifier:

  • Narrowly scoped to the task, not everything the agent might ever need
  • Short-lived (billing access at 2 p.m. should not linger an hour later)
  • Deny capable (agent may read production, never write, even if the human can write)
  • Intersection aware (agent plus user rights are the intersection, never the union)

you can get a guest badge and you can kind of go anywhere that a human will take you. But even if the human has access to the secret server room that guests aren't allowed in, you're still not allowed to go because you've got the guest badge.

Tyler Akidau, CTO, Redpanda, on Chain of Thought ep 60

Jeetu Patel, Cisco's President and Chief Product Officer, adds a second axis on the same explainer: access to email is not permission to email your board. You need action control, not access control alone. Read without send, draft without publish, recommend trade without execute.

What fixes should I apply, in order of effort?

Start with cheap removals, then tighten time and scope, then move enforcement out of the agent's view.

  1. Remove excess functionality. Drop tools and plugins the task does not use (document readers that can delete, dev MCP servers left enabled). This is where I would start on excessive agency.

  2. Stop sharing one credential across concurrent instances. Akidau notes agents clone easily; one key plus many parallel tasks mixes blast radii. Split identities per instance or per task where your IdP allows it.

  3. Add expiry and rotation. Replace "never expires" rows first. Pair with narrowly scoped grants for the single job.

  4. Encode denies and intersections in policy, not prompts. Tyler Akidau is explicit that a repo guidance file is guidance: if it must not happen, enforce it in infrastructure the agent cannot edit. Prompts and guard models fail under injection and hallucination pressure.

  5. Escalate by stakes. After the software audit, ask what still moves money or irreversible data. Charles Guillemet of Ledger argues that securing a high-stakes agent with permissions alone does not give the guarantees finance-grade abuse demands; hardware enters when attacker incentives are high. You do not need a hardware program to delete a stale plugin, but you should not pretend a prompt cap replaces policy for treasury actions.

For broader episode context on identity, federation, and action boundaries, the AI security topic page collects the full conversations.

The only way to have very strong security guarantees is to have the dedicated hardware. And when it comes to your money, the incentives for the attacker are very high. So hardware will be part of the equation.

Charles Guillemet, CTO, Ledger, on Chain of Thought ep 65

Use that quote as a severity triage rule, not as an excuse to skip the audit.

Checklist

  • List every secret and token the agent process can read (env, mounts, tools, CI), with issuer, grants, and expiry.
  • Mark rows that exceed the current task, never expire, or feed logging and tracing without redaction.
  • Include unsanctioned or parallel agent use in scope so shadow setups are not invisible.
  • Classify each row against narrow scope, short life, deny rules, and intersection (union is a bug).
  • Apply fixes in order: remove unused tools, split shared credentials, shorten TTL, then enforce denies outside prompts; reserve hardware-grade controls for high-incentive targets.

The longer explainer, with the episode clips, is on Chain of Thought. It draws on this episode.

Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.

Drafted with AI assistance from the episode transcripts.

Top comments (0)