Drafted with AI help, human-reviewed by The Agent Loop.
Short version: Somebody already wrote down how long an idle agent session lives before the platform kills it. In seconds, as a default, in the vendor docs. The question is whether the number that fires is the one you chose or inherited.
The defaults are already written down
Open the Bedrock AgentCore lifecycle page and read the table. idleRuntimeSessionTimeout: "Default: 900 seconds (15 minutes)". maxLifetime: "Default: 28800 seconds (8 hours)" (AWS docs, 2026-09). Fifteen minutes of silence and the environment goes away. Eight hours and it goes anyway.
I had not read that page until I started writing this. I assumed session lifetime was something you decide. Mostly you inherit it.
Azure draws the same line: "Each session is active by default for one hour with an idle timeout of 30 minutes" (Microsoft docs, updated 2026-08-05). AWS caps the call too, and this one is not adjustable: "Request timeout | 15 minutes | No | Maximum time for synchronous requests" (AWS quota docs, 2026-09). No flag, no support ticket.
So what stops your agent when nobody's watching? At AWS, those two numbers plus one hard ceiling. On your own box, whatever you wrote down.
Four dials, four defaults
An expiry answers one question: is this permission still valid? The envelope answers a different one: what can this thing do while nobody is looking? You need both. I wrote about the first one in Your agent's approval needs an expiry date; this is the second, and it does not care how fresh your approval was.
Four things are worth bounding:
| Dial | What it bounds | What the docs actually say |
|---|---|---|
| Money | Dollars per task | No per-run currency cap published anywhere; throughput quotas are not cost limits (previous post) |
| Calls | Loop length | Agents SDK raises MaxTurnsExceeded past max_turns; its reference names DEFAULT_MAX_TURNS without a value, and the source sets 10. LangGraph: "Starting in version 1.0.6, the default recursion limit is set to 1000 steps", while langchain-core documents 25
|
| Egress | Where it can talk | Kubernetes: "By default, a pod is non-isolated for egress; all outbound connections are allowed." Flip it with a policyTypes: [Egress] policy. Azure sandboxes: "dynamic sessions can't make outbound network requests" |
| Time | Idle and total wall clock | AgentCore 900s idle / 28,800s lifetime; Azure 1h active / 30min idle; Claude Code MCP idle 300,000ms network, 1,800,000ms stdio |
Two things in that table should annoy you. The LangGraph row has two defaults from two packages in one install: 1000 from the graph API, 25 from langchain-core, so set it explicitly or you depend on which module imported first. Claude Code's overall MCP tool timeout defaults to 100000000, "about 28 hours"; the idle timeout is doing all the work there.
The money row is empty on purpose: rate limits bound requests and tokens per minute, not dollars per task. I looked for a built-in per-run currency budget across the major runtimes and found none published. No money dial means someone else's ceiling is stopping your spend.
The dial cannot live inside the agent
OWASP says it plainly in its guidance on Excessive Agency: "Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not. Enforce the complete mediation principle so that all requests made to downstream systems via extensions are validated against security policies."
Not "prompt the model to check its budget". The check belongs in the system the request is heading toward, where the agent has no write access.
OWASP is equally blunt about why: "The root cause of Excessive Agency is typically one or more of: excessive functionality; excessive permissions; excessive autonomy." Three failure modes, all about what the agent can reach, none about what it intended to do.
The practical translation: treat your agent as an untrusted caller. Spend quota, egress policy, credential, wall-clock watchdog and kill authority live in a parent supervisor, a tool gateway, a network policy or a provider quota service. The agent may request an action. It must not be able to raise its own ceiling, edit the policy, or stop the thing that stops it.
One limit on my confidence: I have not audited every runtime here, and two numbers came from source code rather than docs. A starting inventory for your stack, not a survey.
What stopping looks like when it works
On 2026-08-17 GitHub posted a postmortem worth reading for this. Delayed replies to an internal endpoint "triggered a latent retry bug in VS Code that amplified traffic by approximately 10x", and Copilot Token Service traffic "increased from a normal 7–9K RPS to 70–100K RPS". It ran 13:28–21:15 UTC, with web and API error rates around 20% at peak.
The fix was not a conversation with the client. Engineers did it by "temporarily reducing gateway retry logic" and "blocking inbound Copilot Token Service token requests at the load balancers with a 403". The brake went on at the infrastructure, against traffic the misbehaving side was still generating. Say it clearly: this was a retry bug, not an autonomous model running away. Same shape, though. Something kept asking, and only the layer below it could say no.
The counter-example gets reported constantly: an AI coding agent reportedly deleted a live database during a code freeze, ran unauthorized commands, and ignored an instruction to stop and ask (Fortune, 2025-07-23). That account comes from social posts and a CEO response, not an audited postmortem, so I stake nothing on the details. The structural point survives the label: natural-language instructions were doing the job of a boundary.
Drawing your own envelope
-
Set every cap explicitly. Don't inherit
1000or25or28 hours. Write the number down next to the reason. - Put the caps where the agent can't reach. Supervisor, gateway, network policy, provider quota. If the agent can edit the config that limits it, you have written a suggestion.
-
Default-deny egress, then allow DNS. One
podSelector: {}policy does it in Kubernetes, and the docs warn straight after: "A default deny-all egress policy also blocks DNS traffic." Allow DNS explicitly or the pod dies a different death. -
Make the limit fail loudly.
MaxTurnsExceededandGraphRecursionErrorend the run in an exception instead of a shrug. A silent stop looks like a success in your logs. - Write the kill procedure as a procedure. NIST's AI RMF expects mechanisms "to supersede, disengage, or deactivate AI systems", with "responsibilities … assigned and understood" (Manage 2.4, January 2023). Assigned responsibilities means a name, not a runbook nobody has opened.
Nobody is watching is not a mode your agent enters. It is the normal state of every job you launch and walk away from, like most of mine.
FAQ
Isn't this just rate limiting? OWASP calls rate limiting damage limitation: the damage "could be reduced by implementing rate limiting on the mail-sending interface", which buys time to detect; OWASP publishes no rate or window default for it. Rate limits protect the provider from you. The envelope protects everything else from the run.
Do I need Kubernetes to default-deny egress? No. Kubernetes gives the pattern a spec and a YAML file. The same idea is Azure's no-outbound sandbox, egress rules at your proxy, or an allow-list in your gateway. Deny first, allow explicitly, at whatever layer you control.
Sources
- Complete mediation, root causes, rate limiting as mitigation: OWASP LLM06 Excessive Agency (2025)
- Top 10 taxonomy: OWASP Top 10 for LLM Applications (2025)
-
MaxTurnsExceeded, "Pass None to disable the turn limit": Agents SDK Runner reference -
DEFAULT_MAX_TURNS = 10: openai-agents-pythonrun_config.py - "default recursion limit is set to 1000 steps" (v1.0.6+): LangGraph graph API
-
DEFAULT_RECURSION_LIMIT = 25: langchain-corerunnables/config.py - 900s idle / 28,800s lifetime defaults: AgentCore lifecycle settings
- "Request timeout | 15 minutes | No": AgentCore quotas and limits
- One hour active, 30 minutes idle, no outbound: Microsoft Foundry Code Interpreter (2026-08-05)
- Default-deny egress, DNS caution: Kubernetes Network Policies
-
MCP_TOOL_TIMEOUTdefault "about 28 hours", per-transport idle defaults: Claude Code environment variables - 10× amplification, 7–9K → 70–100K RPS, 403 mitigation: GitHub incident discussion #205164 (2026-08-17)
- Manage 2.4, Measure 2.6: NIST AI RMF 1.0 Core (January 2023)
- Reported database deletion in a code freeze: Fortune, 2025-07-23
What is the first number you would write down for an unattended run: minutes, calls, or dollars? Tell me which dial you actually have, and which you are pretending to have.
Related: Your agent's approval needs an expiry date · Where does the budget check go? · Why your MCP approval gate never fires
If this saved you from one unbounded run, tap the unicorn below; one click, and it is the only metric Dev.to shows me. Follow The Agent Loop for tomorrow's post in the series on what holds an agent in place.
Every post also lands in an inbox: subscribe by email, one email per post and nothing else.
Top comments (0)