DEV Community

mgbec for AWS Community Builders

Posted on Originally published at Medium on

The Non-Human Chronicles

I originally built a Compliance Agent in AWS Bedrock with an intended audience of typical users. Then I added some agent to agent communication, using A2A protocol with both Entra and Cognito authentication. This allowed agents from both my AWS and Microsoft ecosystems to use it successfully. https://github.com/mgbec/SCF-Agent-with-A2A

A2A, agents, and these agentic non-human users have different needs and priorities than our typical human users, and bring different security challenges to the table as well. Non-human identities in our workflows are nothing new of course. We have been using service accounts, service principals, and managed identities to access resources for years. OWASP gave us a top 10 for them quite a while ago: https://owasp.org/projects/non-human-identities-top-10.

However, as our modern applications transition from simple automation to agentic, these non-human identities have evolved from predictable service accounts into dynamic decision-makers that can execute complex interactions in increasingly diverse environments. Observing and managing these identities is a critical need and is creating many new security and governance challenges.

Key Shifts in Security Priorities

There are a few changing focus areas as we move from human identities into all the permutations of non-human. Some of them include:

Core Security Priorities for Non-Human Identities

1. Elimination of Persistent Credentials (Secretless Architecture)

  • Short-lived tokens: Prioritize dynamic, time-bound credentials over static passwords or long-lived API keys.
  • Workload identity federation: Use framework-native identities (like AWS IAM Roles, Kubernetes Service Accounts, or OIDC) to allow systems to trust each other without exchanging hardcoded secrets.

2. Discovery and Inventory Management

  • Visibility: Non-human agentic identities are created through code, configuration files, CI/CD pipelines, developer toolchains, and AI agent frameworks and we need to be able to scan those environments for evidence of their existence.
  • Contextual mapping: Every non-human identity must be mapped to a specific workload, owner, and business purpose to prevent orphaned accounts.

3. Granular Least Privilege

  • Scope restriction: While a human might need broad access to a database to run various reports, a machine identity should only have access to the exact tables, rows, or API endpoints required to run its specific code.
  • No interactive login: Machine identities should be explicitly blocked from interactive console or UI logins.

4. Automated Monitoring and High-Speed Response

  • Real-time anomalies: Security systems must watch for machine-speed anomalies, such as an API key suddenly making 10,000 requests per second or accessing data outside its baseline behavior.
  • Automated isolation: Because a breach happens in milliseconds, the response must be automated (e.g., automatically disabling a compromised token) rather than waiting for a human analyst to review an alert.

OWASP NHI and A2A

In this particular project, I put quite a bit of focus on the A2A protocol. If I map out the OWASP NHI Top 10 to A2A interactions, these are some of my takeaways:

NHI1 Improper Offboarding : NHIs not deactivated when no longer needed
Retire M2M clients / Entra app registrations for callers that are decommissioned.

NHI2 Secret Leakage : keys/tokens leaked into code, config, or chat
The calling agent’s Cognito client secret must never land in source or logs.

NHI3 Vulnerable Third-Party NHI : compromised extensions/SaaS with granted access
Vet any third-party agent you grant an ‘invoke’ credential to.

NHI4 Insecure Authentication : deprecated/weak auth methods
Use OAuth2/OIDC JWT validation, not static API keys.

NHI5 Overprivileged NHI : more permissions than the function requires
Keep the token scope to invoke only; don’t widen it into a catch-all.

NHI6 Insecure Cloud Deployment Configurations : static creds, poorly validated OIDC in CI/CD
Validate token claims (issuer, audience) strictly; prefer OIDC/role assumption over static keys.

NHI7 Long-Lived Secrets : credentials that expire too far out or never
Short token TTLs (60 min here) plus a client-secret rotation runbook.

NHI8 Environment Isolation : reusing the same NHI across dev/test/prod
Separate credentials per environment; don’t share one client across stages.

NHI9 NHI Reuse : one identity shared across apps/components
Issue one client per calling agent so a compromise is contained and revocable.

NHI10 Human Use of NHI : people using machine identities for manual tasks
Don’t let operators hand-drive the M2M credential; we want attribution.

Changes in this project

As I think about the factors we will need to address the differences between non-human and human identities, I wanted to make a few changes to this project. I broke them down into three categories, but of course, they are not really that discrete.

Cost Predictability

Costs associated with NHI traffic will differ significantly from human generated traffic. NHI traffic can be schedulable and more structured, with more predictability in a well designed situation. Conversely, with poorly designed constraints, an agent can retry aggressively or spawn many subrequests. In this project I have a concurrency cap in place and should keep a2a_worker_max_concurrency low. The SQS + async worker with a DLQ design retries a failed task at most twice then dead-letters it, so a misbehaving caller can’t infinitely re-drive Bedrock calls.

I should also set up budget alerts, though, to make sure there are no cost surprises. I could also do something similar to another project where I had a model router (https://github.com/mgbec/LLM-Router---Deployed-to-AWS-and-SOC2-compliant) in place, to keep the use of more expensive models to situations that require them.

Performance

My latency profile is ok for machines. Async: message/send returns a submitted task in <1s, the worker runs up to 600s, the caller polls tasks/get. Humans don’t want to wait, but for a non-human this is fine.

Concurrency is the throughput ceiling. My current maximum_concurrency is five. This tells AWS to never run more than five copies of the worker Lambda at the same time. Since each worker handles one task and one agent turn, that means at most five agent turns are actually executing at any instant. Any calls beyond that hang out in SQS until the running task finishes, then they take their turn. Again, this backpressure( the system politely making callers wait instead of overloading) is ok for a NHI. Worst case scenario, if things seem too laggy, we can bump up the concurrency.

Other factors I need to think about though:

-Cold starts hit the worker Lambda and AgentCore runtime on the first call after idle. Humans absorb a one-off delay; a latency-sensitive calling agent may retry on it, so I will need to be careful with timeouts.

-Token TTL vs. session length. M2M access tokens live 60 min. A single long task inside the 600s worker is unaffected (the token was already validated at submit), but if a caller is polling for a very long time it should refresh its token.

Security

This is where non-human identities need the most deliberate handling, because a leaked machine credential is worse than a leaked human password : no MFA, no human noticing a weird login, and it’s often embedded in code or CI.

These are already in place:

-Short token TTL (60 min) limits the blast radius of a stolen M2M token.

-Scoped access : the M2M client only gets the invoke scope, nothing else.

-Admin-create-only pool : no self-signup for the human side.

-Guardrail on every model call : prompt-attack + PII filtering applies regardless of who called.

-Caller isolation : runtimeSessionId = sha256(caller:contextId), so two callers can’t collide sessions or read each other’s context.

Things I need to fix:

-Secret management and rotation. The a2a_m2m client has a real secret. It’s stored locally and not committed, but in the real world we would want to manage and rotate this. Or better yet, try workload identity federation. Rather than storing a long-lived client secret, the workload proves its provenance at runtime (environment attestation or a signed identity token) and exchanges that proof for short-lived access. On AWS, the equivalent is having the calling workload assume an IAM role via its execution environment instead of holding a Cognito client secret.

-One identity per caller. Don’t share one M2M client across every calling agent. Issue a distinct client (or Entra app registration) per caller so you can revoke one without breaking the rest and attribute usage/cost per caller. Today there’s a single shared a2a_m2m client.

-Rate limiting per identity. API Gateway HTTP APIs support throttling; add per-route (ideally per-client) limits so one compromised or buggy credential can’t flood you. This doubles as cost protection.

-Audit & anomaly detection. Log client_id on every request (CloudWatch), so I can answer “which caller drove this spend/this spike.” I could consider a CloudWatch alarm on request rate per client.

Governing Agentic NHI- future directions

Implementing Zero Trust for non-human identities in artificial intelligence requires continuous verification and dynamic risk evaluation.. Because AI agents possess reasoning capabilities and autonomy, traditional security models must be adapted to account for their unique behavioral patterns, such as tool execution and contextual memory. As our workflows evolve, we will need to consider the evolution of:

Identity: Every AI interaction must be anchored to a verifiable identity that is continuously evaluated and explicitly authorized.

Context: Additional signals used to inform access decisions include threat intelligence feeds, session risk scores, and the specific sensitivity of the action being attempted.

Verification and Observability: ZeroTrust relies on continuous signals rather than a one time binary decision. These signals include intent assessment, authentication patterns, and abnormal request patterns.

Decision Intelligence: Modern platforms have transitioned toward decision intelligence by correlating various identity signals to evaluate risk in real-time.

One interesting paper that suggests a layered framework for Zero Trust governance is MDPI Informatics : From Service Accounts to Agentic Identities (Zero Trust governance framework) : https://www.mdpi.com/2227-9709/13/9/146

This paper addresses these gaps by developing the Agentic Identity Governance, Authority, Tool-Control and Evidence (AIGATE) Framework, a Zero Trust-aligned governance architecture for agentic AI as a non-human enterprise identity. It generates a set of agent-specific governance requirements from the diverse literature on agent security, machine identity, IAM, Zero Trust, runtime enforcement, and AI governance. It integrates these requirements into eight connected AIGATE layers and operationalizes least agency alongside least privilege so that both access and autonomy are taken into account separately.

It’s an exciting time to be in security. I’m looking forward to all the new developments, both theoretical and practical. Thanks for reading!

Top comments (0)