<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alexander Spoecker</title>
    <description>The latest articles on DEV Community by Alexander Spoecker (@spoecker).</description>
    <link>https://dev.to/spoecker</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F861328%2Fcf1e2c8d-1476-4d2e-bf2c-bf462178675e.JPG</url>
      <title>DEV Community: Alexander Spoecker</title>
      <link>https://dev.to/spoecker</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/spoecker"/>
    <language>en</language>
    <item>
      <title>Before You Build an Agent: Five Amazon Bedrock Fundamentals, Measured</title>
      <dc:creator>Alexander Spoecker</dc:creator>
      <pubDate>Fri, 02 Oct 2026 15:13:24 +0000</pubDate>
      <link>https://dev.to/spoecker/before-you-build-an-agent-five-amazon-bedrock-fundamentals-measured-5goj</link>
      <guid>https://dev.to/spoecker/before-you-build-an-agent-five-amazon-bedrock-fundamentals-measured-5goj</guid>
      <description>&lt;p&gt;&lt;em&gt;This is the long version of my 30-minute talk at AWS Community Day Thailand, Bangkok, 3 October 2026. The code, the checklist and every dated measurement are in &lt;a href="https://github.com/spoecker/aws-th-community-day-2026-bedrock-fundamentals" rel="noopener noreferrer"&gt;the repo&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I work as a Cloud Solution Architect at Iglu in Chiang Mai, Thailand, and I also teach AWS classes as an AWS Authorized Instructor, mostly online. When I started preparing this talk I set myself one goal: no number goes on a slide unless I measured it myself, and no rule goes on a slide unless I can point at the page it comes from.&lt;/p&gt;

&lt;p&gt;A warning before the numbers. Everything here was measured in one AWS account on the date written next to it, and every rule links the AWS or Anthropic page I read it on, with the date I read it. This stuff moves. Often really fast as we know with AI. In the four weeks I worked on the talk, the Thailand Region's model list grew from 21 to 31, two new Claude models showed up, and one documentation page I cite changed its numbers. So if your numbers differ from mine, don't assume one of us is wrong. Re-run the scripts and tell me what you got. I probably won't manage to keep the numbers updated in this post.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;th&gt;Measured&lt;/th&gt;
&lt;th&gt;Where the calls were made&lt;/th&gt;
&lt;th&gt;File in the repo&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Tokens per language and model&lt;/td&gt;
&lt;td&gt;2026-09-25, identical 2026-10-01 (spot check 2026-09-30)&lt;/td&gt;
&lt;td&gt;ap-southeast-1 (Singapore)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;measurements/01-tokens-2026-09-25.json&lt;/code&gt;, &lt;code&gt;01-tokens-spotcheck-claude-5-5-2026-09-30.json&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stateless calls, then memory in DynamoDB&lt;/td&gt;
&lt;td&gt;2026-09-25&lt;/td&gt;
&lt;td&gt;ap-southeast-7 (Thailand)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;02-memory-stateless-…&lt;/code&gt;, &lt;code&gt;02-memory-with-memory-2026-09-25.json&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context growth, caching, windowing&lt;/td&gt;
&lt;td&gt;2026-09-25&lt;/td&gt;
&lt;td&gt;ap-southeast-1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;03-context-2026-09-25.json&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;117 concurrent calls against the TPM quota&lt;/td&gt;
&lt;td&gt;2026-09-25, identical 2026-10-01 (first 2026-09-14)&lt;/td&gt;
&lt;td&gt;ap-southeast-1, from an EC2 runner in Bangkok&lt;/td&gt;
&lt;td&gt;&lt;code&gt;04-quota-haiku-2026-09-25.json&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Where Bedrock ran each request&lt;/td&gt;
&lt;td&gt;2026-09-14, 25, 29, 30 and 2026-10-01&lt;/td&gt;
&lt;td&gt;ap-southeast-7&lt;/td&gt;
&lt;td&gt;&lt;code&gt;05-routing-2026-*.json&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Applied quotas per Region&lt;/td&gt;
&lt;td&gt;2026-09-14 (re-read 2026-09-30, unchanged)&lt;/td&gt;
&lt;td&gt;both&lt;/td&gt;
&lt;td&gt;&lt;code&gt;quotas-by-region-2026-09-14.json&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Unless a different date is given next to a link, I read the documentation on 2026-09-14 and again on 2026-09-30.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three surprises
&lt;/h2&gt;

&lt;p&gt;The talk is built around three surprises that come with an application on Amazon Bedrock. One of them you meet on the first day of development (Specially if you are used to using only consumer AI products before). The other two wait until real users and real traffic arrive. (Or sometimes when expanding to a new country)&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The model forgets.&lt;/strong&gt; This is the early one. In the first test, turn two knows nothing about turn one, and you find out that conversation history isn't a feature you switch on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The bill is wrong.&lt;/strong&gt; The price per token was right. The number of tokens wasn't, because the estimate was made in English and the customers write Thai.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It throttles at a fraction of the expected traffic.&lt;/strong&gt; Everybody reads the tokens-per-minute quota. Almost nobody reads how it gets used up.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these has anything to do with the model being good or bad, and none of them is an agent problem. They're properties of one plain model call.&lt;/p&gt;

&lt;p&gt;A colleague who reviewed my slides asked me what an agent actually is. The answer I ended up with is the reason for the title: an agent is a loop of model calls that decides the next step and calls tools. Whatever is true for a single call is true for every step of that loop, just more often. So I'd rather get it right at one call.&lt;/p&gt;

&lt;p&gt;The example running through the whole post is an airline's customer chat, with one Thai customer and one request. Everything is plain &lt;code&gt;boto3&lt;/code&gt; and the &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/conversation-inference.html" rel="noopener noreferrer"&gt;Converse API&lt;/a&gt;. There's no framework anywhere in the repo.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Tokens are not words
&lt;/h2&gt;

&lt;p&gt;A token is the chunk a model's tokenizer cuts your text into. You pay per token, and the context window is counted in tokens too. For Claude, Anthropic's glossary says a token &lt;em&gt;"approximately represents 3.5 English characters, though the exact number can vary depending on the language used"&lt;/em&gt; (&lt;a href="https://platform.claude.com/docs/en/about-claude/glossary" rel="noopener noreferrer"&gt;Anthropic glossary&lt;/a&gt;). The count also depends on the model: token counting is model-specific (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_CountTokens.html" rel="noopener noreferrer"&gt;CountTokens API reference&lt;/a&gt;). What I couldn't find anywhere, from AWS or from Anthropic, was a number for Thai. So I measured it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The measurement
&lt;/h3&gt;

&lt;p&gt;I took one sentence a customer might type, in four languages. My Thai is very limited, so I had a native speaker read the Thai sentence before I put it on a slide. Each sentence goes to each model once with &lt;code&gt;maxTokens: 1&lt;/code&gt;, and I read &lt;code&gt;usage.inputTokens&lt;/code&gt; from the response. That field is the one you're billed on: CloudWatch's &lt;code&gt;InputTokenCount&lt;/code&gt; for the same calls added up to exactly these values when I checked on 2026-09-04.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;EN (89 characters): &lt;em&gt;Hello, I would like to change my flight to next Tuesday morning. Is there a fee for that?&lt;/em&gt;&lt;br&gt;
TH (95 characters): &lt;em&gt;สวัสดีครับ ผมต้องการเปลี่ยนเที่ยวบินเป็นเช้าวันอังคารหน้า มีค่าธรรมเนียมสำหรับการเปลี่ยนไหมครับ&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foli20a0qdff4gwwf1l70.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foli20a0qdff4gwwf1l70.png" alt="The same sentence in English and Thai as token chunks: 28 versus 90 on Claude Haiku 4.5" width="800" height="357"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Measured 2026-09-25, Singapore, &lt;code&gt;usage.inputTokens&lt;/code&gt;&lt;/strong&gt; (all four rows identical again on 2026-10-01; the Haiku 4.5 and Nova rows also on 4, 14 and 21 September from both Singapore and Bangkok):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model (inference profile)&lt;/th&gt;
&lt;th&gt;"Hi"&lt;/th&gt;
&lt;th&gt;EN (89 chars)&lt;/th&gt;
&lt;th&gt;DE (95)&lt;/th&gt;
&lt;th&gt;ZH (26)&lt;/th&gt;
&lt;th&gt;TH (95)&lt;/th&gt;
&lt;th&gt;TH ÷ EN&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5 (&lt;code&gt;global.&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;td&gt;35&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.2×&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5 (&lt;code&gt;global.&lt;/code&gt;; Opus 5 counted the same on 21 Sep)&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;td&gt;59&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.4×&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Nova Lite (&lt;code&gt;apac.&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;75&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.6×&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Nova 2 Lite (&lt;code&gt;global.&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;47&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;67&lt;/td&gt;
&lt;td&gt;70&lt;/td&gt;
&lt;td&gt;68&lt;/td&gt;
&lt;td&gt;80&lt;/td&gt;
&lt;td&gt;1.2× (1.6× net of "Hi")&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On 2026-09-30 I did a spot check for Claude Sonnet 5.5 and Opus 5.5, which had both joined my account's model list in the last week of September. Both count "Hi" 11, EN 39, TH 91, so 2.3×. Every count is exactly two above Sonnet 5. To me that looks like two more fixed framing tokens and not a new tokenizer, but that's my guess. It isn't documented.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs9cc7gki9ueztx3xw7ps.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs9cc7gki9ueztx3xw7ps.png" alt="Demo 1 as run on 2026-10-01: the raw Converse request and the usage block of its response" width="800" height="387"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftveiz23maej0bu17o2pj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftveiz23maej0bu17o2pj.png" alt="Demo 1, same run: the token table for four models" width="799" height="261"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What I take from the table
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Thai costs 2.4 to 3.6 times as much as English for the same sentence&lt;/strong&gt;, and where you land in that range depends on the tokenizer. On Claude Haiku 4.5 it's close to one token per Thai character (95 characters, 90 tokens). If your budget comes from the English rule of thumb, it's off by a factor of three before the first customer has typed anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The newest Claude tokenizer made English more expensive. It did not make Thai cheaper.&lt;/strong&gt; Anthropic's models overview (read 2026-09-21) says the current tokenizer, introduced with Opus 4.7, fits about 555k words in 1M tokens where the previous one fit about 750k (&lt;a href="https://platform.claude.com/docs/en/about-claude/models/overview" rel="noopener noreferrer"&gt;Anthropic models overview&lt;/a&gt;). In my measurement the English sentence went from 28 tokens on Haiku 4.5 to 37 on Sonnet 5, which is 32% more. Thai went from 90 to 89. So counts you measured on one model generation don't carry over to the next.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nova 2 Lite charges about 46 tokens of overhead on every call.&lt;/strong&gt; The single word "Hi" costs 47 input tokens. I didn't find this documented anywhere, so treat it as my measurement and nothing more. It does mean that fifty one-word calls cost more than one fifty-word call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every model has some per-call overhead.&lt;/strong&gt; "Hi" is 8 tokens on Haiku 4.5 and 1 on Nova Lite. That's worth knowing before you price a workload with lots of short messages.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  How to count, and where counting doesn't work
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;usage&lt;/code&gt; comes back with every response (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_TokenUsage.html" rel="noopener noreferrer"&gt;TokenUsage&lt;/a&gt;): &lt;code&gt;inputTokens&lt;/code&gt;, &lt;code&gt;outputTokens&lt;/code&gt;, &lt;code&gt;totalTokens&lt;/code&gt;, plus &lt;code&gt;cacheReadInputTokens&lt;/code&gt; and &lt;code&gt;cacheWriteInputTokens&lt;/code&gt; when caching is involved. With streaming it arrives at the very end, in the final &lt;code&gt;metadata&lt;/code&gt; event. A client that stops reading early never sees it.&lt;/p&gt;

&lt;p&gt;Bedrock also has a &lt;code&gt;CountTokens&lt;/code&gt; operation, and it's free (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/count-tokens.html" rel="noopener noreferrer"&gt;user guide&lt;/a&gt;: &lt;em&gt;"doesn't incur charges"&lt;/em&gt;). I expected it to be the easy answer. It wasn't, for three reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;From Singapore and Thailand it rejected every inference profile I tried (&lt;em&gt;"The provided model doesn't support counting tokens"&lt;/em&gt;). In &lt;code&gt;us-east-1&lt;/code&gt; it accepted the base model ID &lt;code&gt;anthropic.claude-haiku-4-5-20251001-v1:0&lt;/code&gt;, which you can't invoke on demand there. So you count with one ID and invoke with another.&lt;/li&gt;
&lt;li&gt;Its answer was always &lt;strong&gt;16 tokens above&lt;/strong&gt; what &lt;code&gt;Converse&lt;/code&gt; billed for the same text on Haiku 4.5 (24 vs 8, 44 vs 28, 106 vs 90; measured 2026-09-04, 14 and 25, and again on 2026-10-01). The API reference says the count &lt;em&gt;"will match the token count that would be charged"&lt;/em&gt;. The closest thing to an explanation I found is on Anthropic's side: counts &lt;em&gt;"may include tokens added automatically by Anthropic for system optimizations. You are not billed for system-added tokens"&lt;/em&gt; (&lt;a href="https://platform.claude.com/docs/en/build-with-claude/token-counting" rel="noopener noreferrer"&gt;Anthropic token counting&lt;/a&gt;). That would fit a fixed offset. I haven't been able to confirm it for Bedrock.&lt;/li&gt;
&lt;li&gt;It only works for Claude, and not even for all of Claude. Models that launch as cross-Region-only on &lt;code&gt;bedrock-runtime&lt;/code&gt; &lt;em&gt;"don't support CountTokens on bedrock-runtime"&lt;/em&gt;. The doc points you to Anthropic's &lt;code&gt;count_tokens&lt;/code&gt; route on the &lt;code&gt;bedrock-mantle&lt;/code&gt; endpoint instead, and that endpoint (Haiku 4.5 model card, read 2026-09-14) exists in seven Regions, none of them in Southeast Asia. Amazon Nova doesn't support it at all (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-lite.html" rel="noopener noreferrer"&gt;Nova 2 Lite model card&lt;/a&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So from Bangkok today, the counter that actually works is a &lt;code&gt;Converse&lt;/code&gt; call with &lt;code&gt;maxTokens: 1&lt;/code&gt;. A few dozen of those cost well under a cent, and what you get back is the billed number.&lt;/p&gt;

&lt;p&gt;There's one more thing here that I can't explain. On 21, 25, 29 and 30 September and on 1 October, my account was refused Claude Sonnet 5 and Opus 5 when I called from &lt;code&gt;ap-southeast-7&lt;/code&gt; (&lt;em&gt;"not available for this account"&lt;/em&gt;, &lt;code&gt;AccessDeniedException&lt;/code&gt;). The same &lt;code&gt;global.&lt;/code&gt; profile answered from &lt;code&gt;ap-southeast-1&lt;/code&gt; every single time, and the availability API in Bangkok reported the model agreement as available. I still don't know why. The lesson I took from it: before you promise anyone a model from a Region, call it from that Region.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do this:&lt;/strong&gt; measure &lt;code&gt;usage.inputTokens&lt;/code&gt; on the model you'll ship, with real text in your users' language. Find out the per-call overhead. And build your cost model on the billed number, not on what a counter estimates.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Conversation history is application state
&lt;/h2&gt;

&lt;p&gt;Does the model remember? The Converse user guide answers that in two sentences (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/conversation-inference.html" rel="noopener noreferrer"&gt;conversation-inference&lt;/a&gt;, re-read 2026-09-30):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Amazon Bedrock doesn't store any text, images, or documents that you provide as content. The data is only used to generate the response."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;"You can maintain conversation context by including all the messages in the conversation in subsequent Converse requests."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So the &lt;code&gt;messages&lt;/code&gt; array is the memory. Your code owns it, and you pay for it again on every turn.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frkl1b1qyrp0j9v571ivl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frkl1b1qyrp0j9v571ivl.png" alt="Call 2 knows nothing about call 1: the messages array is what grows" width="799" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The two-call demo
&lt;/h3&gt;

&lt;p&gt;Two &lt;code&gt;Converse&lt;/code&gt; calls from Bangkok on Claude Haiku 4.5. The system prompt is "You are a concise travel assistant for an airline. Answer in one sentence." This is the run of 2026-09-25:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;request:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;message(s),&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;44&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;input&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;tokens&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;user&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Hi,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;am&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Alex.&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;am&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;flying&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;to&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Bangkok&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;AWS&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Community&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Day&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;on&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;October.&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;model:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;How&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;can&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;assist&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;you&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;with&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;your&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;flight&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;to&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Bangkok&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;on&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;October&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="err"&gt;rd,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Alex?&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;request:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;message(s),&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;31&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;input&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;tokens&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;user&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;What&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;should&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;pack&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;trip?&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;model:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Pack&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;comfortable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;clothing&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;appropriate&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;your&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;destination's&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;climate,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;a&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;valid&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="err"&gt;passport/ID,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;medications,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;toiletries,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;electronics&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;with&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;chargers,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;and&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;any&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;documents&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="err"&gt;needed&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;your&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;airline&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;check-in.&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second answer is generic because the second request contained exactly one message. Bangkok, October and the conference were never sent.&lt;/p&gt;

&lt;h3&gt;
  
  
  Memory in three API calls
&lt;/h3&gt;

&lt;p&gt;Now the same two turns with the history kept in a DynamoDB table. The partition key is &lt;code&gt;userId&lt;/code&gt;, the sort key is &lt;code&gt;ts&lt;/code&gt;, and a &lt;code&gt;ttl&lt;/code&gt; attribute lets idle conversations expire after 24 hours. Each turn is three calls: &lt;strong&gt;Query&lt;/strong&gt; this user's turns, &lt;strong&gt;Converse&lt;/strong&gt; with that history plus the new message, &lt;strong&gt;PutItem&lt;/strong&gt; for both new turns.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdd9x4v0bw5xwulmfh1lc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdd9x4v0bw5xwulmfh1lc.png" alt="Query, Converse, PutItem: memory in three calls, with the table schema" width="799" height="232"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;request:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;message(s)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;DynamoDB),&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;44&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;input&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;tokens&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;user&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Hi,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;am&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Alex.&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;am&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;flying&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;to&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Bangkok&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;AWS&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Community&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Day&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;on&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;October.&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;model:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Great!&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I'd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;be&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;happy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;to&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;help&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;with&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;your&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;flight&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;to&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Bangkok&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;on&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;October&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="err"&gt;rd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;request:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;message(s)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;DynamoDB),&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;95&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;input&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;tokens&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;user&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;What&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;should&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;pack&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;trip?&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;model:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;For&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Bangkok&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;in&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;early&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;October,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;pack&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;light&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;breathable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;clothing,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;sunscreen,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;an&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;umbrella&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="err"&gt;or&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;rain&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;jacket&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(monsoon&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;season),&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;comfortable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;walking&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;shoes,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;and&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;any&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;medications&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;you&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;need.&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;user&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;'bob'&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;asks&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;same&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;question&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;message&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;sent)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;model:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Pack&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;based&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;on&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;your&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;destination's&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;weather&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;and&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;trip&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;length&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second request now carries three messages and 95 input tokens instead of 31. That's the price of memory: you pay for it as input, on every turn. In return the answer is about Bangkok in October.&lt;/p&gt;

&lt;p&gt;Then Bob asks the identical question and gets the generic answer again, because the Query was scoped to his &lt;code&gt;userId&lt;/code&gt; and came back empty. That scoping is what keeps one customer's conversation out of another customer's answer. It's your code, so it's your job to get it right.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5z3ydt1w5tu9zo63mk6w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5z3ydt1w5tu9zo63mk6w.png" alt="Demo 2 as run from Bangkok on 2026-10-01: the two turns with history from DynamoDB. The lines under the title recap what the stateless calls answered a moment earlier" width="800" height="368"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Screenshots in this post are frames from the demo recordings of 1 October 2026. The model's wording differs a little from the 25 September transcript above; the behaviour does not.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The API expects you to replay everything
&lt;/h3&gt;

&lt;p&gt;If you want proof that the API is designed around you sending the history back, look at reasoning. When a Claude model reasons, the assistant message carries a &lt;code&gt;reasoningContent&lt;/code&gt; block with a &lt;code&gt;signature&lt;/code&gt;, and the Converse guide says: &lt;em&gt;"The signature field is a hash of all the messages in the conversation and is a safeguard against tampering of the reasoning used by the model. You must include the signature and all previous messages in subsequent Converse requests. If any of the messages are changed, the response throws an error."&lt;/em&gt; A service that kept the state itself wouldn't need the client to carry a hash around.&lt;/p&gt;

&lt;p&gt;This matters if your code rewrites history. From Claude Fable 5.1 on, the API checks that &lt;em&gt;"the system prompt, the tool list, and all messages before the block are unchanged"&lt;/em&gt;. Summarising older turns or injecting per-turn reminders then fails with &lt;code&gt;Invalid signature in thinking block. The block is bound to a different conversation&lt;/code&gt;, and the error &lt;em&gt;"is permanent for that request — an automatic retry loop will not clear it"&lt;/em&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/claude-messages-thinking-block-binding.html" rel="noopener noreferrer"&gt;thinking block binding&lt;/a&gt;, read 2026-09-14). You have two ways out: strip the thinking blocks from the point where you rewrote history, or send the &lt;code&gt;thinking-binding-controls-2026-08-01&lt;/code&gt; beta with &lt;code&gt;mismatch_behavior: "drop_block"&lt;/code&gt; through &lt;code&gt;additionalModelRequestFields&lt;/code&gt;. Haiku 4.5 and earlier models strip older thinking blocks automatically (&lt;a href="https://platform.claude.com/docs/en/build-with-claude/thinking" rel="noopener noreferrer"&gt;Anthropic, extended thinking&lt;/a&gt;). Plain text history without thinking blocks is never checked. The model just answers whatever you sent.&lt;/p&gt;

&lt;h3&gt;
  
  
  If you'd rather not build it yourself
&lt;/h3&gt;

&lt;p&gt;There are managed options. Each one is a separate service, and the model API stays stateless whichever you pick. Look at the last two columns before you plan around one of them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What it stores&lt;/th&gt;
&lt;th&gt;Retention&lt;/th&gt;
&lt;th&gt;Singapore&lt;/th&gt;
&lt;th&gt;Thailand&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Your own store (the demo: DynamoDB)&lt;/td&gt;
&lt;td&gt;whatever you write&lt;/td&gt;
&lt;td&gt;your TTL&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/sessions.html" rel="noopener noreferrer"&gt;Bedrock Session Management APIs&lt;/a&gt; (preview)&lt;/td&gt;
&lt;td&gt;checkpoints for LangGraph/LlamaIndex-style apps; 1,000 steps per session, 50 MB per step&lt;/td&gt;
&lt;td&gt;idle timeout 1 hour; &lt;em&gt;"automatically deleted after 30 days"&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;yes (launch list, Feb 2025)&lt;/td&gt;
&lt;td&gt;not listed; no &lt;code&gt;bedrock-agent-runtime&lt;/code&gt; endpoint row for &lt;code&gt;ap-southeast-7&lt;/code&gt; in the General Reference (read 2026-09-14)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/memory.html" rel="noopener noreferrer"&gt;Amazon Bedrock AgentCore Memory&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;short-term raw events per actor and session; long-term records extracted asynchronously by strategies&lt;/td&gt;
&lt;td&gt;events 7 to 365 days; long-term records have no built-in TTL&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;no&lt;/strong&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/agentcore-regions.html" rel="noopener noreferrer"&gt;AgentCore Regions table&lt;/a&gt;, read 2026-09-30; Runtime is available in Thailand, Memory is not)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bedrock Agents (Classic) &lt;code&gt;sessionId&lt;/code&gt; and memory&lt;/td&gt;
&lt;td&gt;agent session state&lt;/td&gt;
&lt;td&gt;per agent config&lt;/td&gt;
&lt;td&gt;closed to new customers since 30 July 2026 (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/agents-classic-maintenance-mode.html" rel="noopener noreferrer"&gt;maintenance mode page&lt;/a&gt;, read 2026-09-14)&lt;/td&gt;
&lt;td&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;AgentCore's own documentation says it in one line: &lt;em&gt;"AgentCore Memory addresses a fundamental challenge in agentic AI: statelessness."&lt;/em&gt; The service exists because the model call has no memory.&lt;/p&gt;

&lt;h3&gt;
  
  
  What "doesn't store" does not mean
&lt;/h3&gt;

&lt;p&gt;"Doesn't store" in the Converse guide is about conversation state. It isn't a statement about retention policy. Bedrock's data-retention page (read 2026-09-30) describes a &lt;strong&gt;mode&lt;/strong&gt; that you set per Region and per account: &lt;code&gt;none&lt;/code&gt; (zero data retention), &lt;code&gt;default&lt;/code&gt; (the model's own policy; &lt;em&gt;"AWS may retain the data for safety and abuse-prevention purposes"&lt;/em&gt;), &lt;code&gt;aws_review&lt;/code&gt; and a legacy &lt;code&gt;provider_data_share&lt;/code&gt;. The same page says plainly: &lt;em&gt;"There is no data retention change to Claude models released before Claude Fable 5."&lt;/em&gt; Claude Fable 5 and 5.1 require &lt;code&gt;aws_review&lt;/code&gt;, and under that mode prompts and completions are &lt;em&gt;"retained within the AWS boundary for up to 30 days"&lt;/em&gt; and may be reviewed by AWS. &lt;em&gt;"Your content is not shared with the model provider."&lt;/em&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html" rel="noopener noreferrer"&gt;data retention&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/abuse-detection.html" rel="noopener noreferrer"&gt;abuse detection&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Nothing you send is used to train a model, and the providers have no access to it. Each provider's model runs in a Model Deployment Account owned by the Bedrock team, and &lt;em&gt;"Model providers don't have any access to those accounts"&lt;/em&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html" rel="noopener noreferrer"&gt;data protection&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Haiku 4.5, Sonnet 5, Opus 5 and Nova were not on the retention list when I read it on 30 September. If you use Fable, section 5 has one more sentence for you about &lt;em&gt;where&lt;/em&gt; that retained copy sits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do this:&lt;/strong&gt; decide where history lives, who can read it and when it expires, and do that before the first prompt ships. Send exactly the history you mean to send, scoped per user. And check that the managed-state service you're counting on exists in your Region.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. The context window is a finite budget
&lt;/h2&gt;

&lt;p&gt;Everything you send has to fit, and you pay for everything you send. Anthropic's context-window page lists what counts: the system prompt, every message, tool definitions, tool results, and the output including thinking. Cached prefixes still take up space in the window (&lt;a href="https://platform.claude.com/docs/en/build-with-claude/context-windows" rel="noopener noreferrer"&gt;context windows&lt;/a&gt;). When you overflow on Claude 4.5 and later, you get &lt;code&gt;stopReason: model_context_window_exceeded&lt;/code&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_Converse.html" rel="noopener noreferrer"&gt;Converse API reference&lt;/a&gt;).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Context window&lt;/th&gt;
&lt;th&gt;Max output&lt;/th&gt;
&lt;th&gt;Source (read)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;200K&lt;/td&gt;
&lt;td&gt;64K&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-anthropic-claude-haiku-4-5.html" rel="noopener noreferrer"&gt;model card&lt;/a&gt; (2026-09-14)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5, Opus 5&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://platform.claude.com/docs/en/about-claude/models/overview" rel="noopener noreferrer"&gt;Anthropic models overview&lt;/a&gt; (2026-09-21)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Nova 2 Lite&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;64K&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-lite.html" rel="noopener noreferrer"&gt;model card&lt;/a&gt; (2026-09-14)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A word of caution about "1M". For Sonnet 4 and Sonnet 4.5 on Bedrock, 1M is a preview variant with its own, much smaller quota rows. The Sonnet 4 preview also documents a long-context premium that re-rates the &lt;em&gt;whole&lt;/em&gt; request once it goes above 200K input tokens (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-anthropic-claude-messages-request-response.html" rel="noopener noreferrer"&gt;Claude messages parameters&lt;/a&gt;). Check the pricing page for the model you actually use.&lt;/p&gt;

&lt;h3&gt;
  
  
  The curve
&lt;/h3&gt;

&lt;p&gt;Per turn, the input grows linearly. Over a whole conversation the total grows, in AWS's words, &lt;em&gt;"approximately quadratically"&lt;/em&gt; with its length (&lt;a href="https://aws.amazon.com/blogs/machine-learning/beyond-the-price-per-token-choosing-the-right-openai-model-on-amazon-bedrock-for-your-workload/" rel="noopener noreferrer"&gt;AWS ML blog, 2026-09-11&lt;/a&gt;). Here are ten short airline-support turns on Haiku 4.5 with the full history replayed each time, measured 2026-09-25 from Singapore:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F25vyxrq4kwkmx4gqtqeo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F25vyxrq4kwkmx4gqtqeo.png" alt="Input tokens per turn, 50 to 545 over ten turns" width="800" height="378"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;turn:          1    2    3    4    5    6    7    8    9   10
inputTokens:  50   88  137  191  253  320  377  445  501  545
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's 2,907 input tokens for ten questions of about 20 tokens each. At this size it costs next to nothing. I'm showing it for the shape.&lt;/p&gt;

&lt;h3&gt;
  
  
  Four runs, one lesson
&lt;/h3&gt;

&lt;p&gt;Then I ran the same ten turns four ways. Claude Haiku 4.5, Singapore list prices as I read them on the pricing page on 2026-09-14 (input $1.00, five-minute cache write $1.25, cache read $0.10 per 1M tokens). "Input cost" means input tokens plus cache writes plus cache reads at list price. Output isn't in it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Input tokens, turn 1 → 10&lt;/th&gt;
&lt;th&gt;Cache write (turn 1)&lt;/th&gt;
&lt;th&gt;Cache read (each later turn)&lt;/th&gt;
&lt;th&gt;Input cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A0 short system prompt, full history&lt;/td&gt;
&lt;td&gt;50 → 545&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;$0.003&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A 8K-token airline policy in the system prompt, full history&lt;/td&gt;
&lt;td&gt;7,870 → 8,305&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;$0.081&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B same, &lt;code&gt;cachePoint&lt;/code&gt; after the policy&lt;/td&gt;
&lt;td&gt;33 → 599 (the non-cached part)&lt;/td&gt;
&lt;td&gt;7,837&lt;/td&gt;
&lt;td&gt;7,837&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.020&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C same as A, only the last 4 turns sent&lt;/td&gt;
&lt;td&gt;7,870, then flat between 7,897 and 7,931&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;$0.079&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftbp3cwyzuvny38c4ni99.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftbp3cwyzuvny38c4ni99.png" alt="Demo 3 as run on 2026-10-01: all four runs side by side. This run came out at 50 to 567 tokens and $0.081, $0.020 and $0.079; the table above is the 25 September run" width="800" height="324"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1suy1pne9jteoq308ofz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1suy1pne9jteoq308ofz.png" alt="Run B turn by turn: 7,837 tokens written to the cache on turn 1 and read on every later turn, while inputTokens shows only the growing, non-cached history" width="800" height="407"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Look at C first. Windowing barely moved the bill, because the static policy, and not the history, was the big part of every request. Caching the policy (B) cut the input cost to a quarter. In a chat with a short system prompt and a long history it would be the other way round.&lt;/p&gt;

&lt;p&gt;Either way you give something up: the model no longer knows turn one unless you carry it forward on purpose. The Well-Architected Agentic AI Lens asks for &lt;em&gt;"conversation history bounded by summarization, sliding windows, or semantic compression so prompt size doesn't grow linearly with session length"&lt;/em&gt; and lists sending the full history every time as its first anti-pattern (&lt;a href="https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentperf03-bp02.html" rel="noopener noreferrer"&gt;AGENTPERF03-BP02&lt;/a&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  The caching rules that decide whether it works
&lt;/h3&gt;

&lt;p&gt;All of this is from the Bedrock &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html" rel="noopener noreferrer"&gt;prompt caching&lt;/a&gt; page as I re-read it on 2026-09-30, and from the &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching" rel="noopener noreferrer"&gt;Anthropic prompt caching&lt;/a&gt; page.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;There's a minimum prefix size, and it differs per model.&lt;/strong&gt; &lt;em&gt;"Claude Opus 5 requires at least 512 tokens per cache checkpoint, Claude Sonnet 5 requires at least 1,024 tokens per cache checkpoint, and Claude Haiku 4.5 requires at least 4,096 tokens."&lt;/em&gt; The table on that page (30 Sep) lists Sonnet 5.5, Opus 5.5 and Fable 5.1 at 512. The minimum counts everything before the checkpoint, across &lt;code&gt;tools&lt;/code&gt;, &lt;code&gt;system&lt;/code&gt; and &lt;code&gt;messages&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Below the minimum it fails without telling you.&lt;/strong&gt; &lt;em&gt;"your inference still succeeds, but your prefix isn't cached."&lt;/em&gt; There's no error. The only sign is &lt;code&gt;cacheWriteInputTokens&lt;/code&gt; staying at 0. A 3,000-token system prompt on Haiku 4.5 with a &lt;code&gt;cachePoint&lt;/code&gt; does nothing at all, and I'd bet that's behind most "my cache never hits" questions. With Thai at roughly one token per character, 4,096 tokens is about 4,000 Thai characters but about 14,000 English ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Order matters.&lt;/strong&gt; &lt;em&gt;"Cache checkpoints are processed in this order: tools → system → messages … changing content in an earlier section invalidates the cache for later sections."&lt;/em&gt; So static content goes first, then the checkpoint, then the per-user data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TTL.&lt;/strong&gt; Five minutes by default, refreshed on every hit. One hour is available on Claude 4.5 and later models, at a higher write price (&lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/01/amazon-bedrock-one-hour-duration-prompt-caching/" rel="noopener noreferrer"&gt;What's New, Jan 2026&lt;/a&gt;). Anthropic adds a detail that's easy to miss: the lifetime is measured from the &lt;em&gt;start&lt;/em&gt; of the request that wrote or read the entry, and an entry only becomes available after the first response begins. Ten parallel calls with the same prefix are ten cache writes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The cache is per account.&lt;/strong&gt; That's why per-user data goes after the checkpoint. It's also why my demo prints a warning if anyone in the account ran it in the last five minutes: turn one then shows a read instead of a write.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cache reads don't count against your TPM quota. Cache writes do&lt;/strong&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/quotas-token-burndown.html" rel="noopener noreferrer"&gt;token burndown&lt;/a&gt;). A well-cached workload gets more throughput out of the same quota. Section 4 is about that quota.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Caching works with cross-Region inference&lt;/strong&gt;, with one caveat: &lt;em&gt;"At times of high demand, these optimizations may lead to increased cache writes"&lt;/em&gt;. A different destination Region may not have your entry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Different providers, different multiples.&lt;/strong&gt; On the Bedrock pricing page (Singapore, Anthropic tab, read 2026-09-14) Claude's cache multiples are 1.25× for a five-minute write, 2× for a one-hour write and 0.1× for a read. Nova 2 Lite's cache read was priced at 25% of input and its write at $0. So don't quote one universal cache discount.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Caching doesn't pay in three cases: a prefix under the minimum (nothing happens), a prefix that changes on every request (every call is a write at 1.25× with no reads, so you pay more than without caching), and batch inference (not supported).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do this:&lt;/strong&gt; bound the history on purpose, with a window, a summary, or a conscious decision to pay for all of it. Put the long static prefix first with a &lt;code&gt;cachePoint&lt;/code&gt; after it, and check &lt;code&gt;cacheWriteInputTokens&lt;/code&gt; on the first call. Put per-user data after the checkpoint.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Output-token limits consume quota before the model writes a word
&lt;/h2&gt;

&lt;p&gt;Of the five, this is the one I got wrong first. It's also the one the documentation states most precisely. From the &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/quotas-token-burndown.html" rel="noopener noreferrer"&gt;token burndown&lt;/a&gt; page, re-read 2026-09-30:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;At the start of the request&lt;/strong&gt; – The following sum is deducted from your quotas. The request is throttled if you exceed a quota. &lt;code&gt;Total input tokens + max_tokens&lt;/code&gt;&lt;br&gt;
&lt;strong&gt;During processing&lt;/strong&gt; – The quota consumed by the request is periodically adjusted to account for the actual number of output tokens generated.&lt;br&gt;
&lt;strong&gt;At the end of the request&lt;/strong&gt; – The total number of tokens consumed is &lt;code&gt;Input token count + Cache write input tokens + (Output token count × Burndown rate)&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And from the CloudWatch runtime-metrics page, on the &lt;code&gt;EstimatedTPMQuotaUsage&lt;/code&gt; metric: it &lt;em&gt;"does not reflect the reservation-based token consumption that drives throttling decisions. Throttling is based on the upfront reservation of input tokens plus max_tokens"&lt;/em&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/monitoring-runtime-metrics.html" rel="noopener noreferrer"&gt;runtime metrics&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;maxTokens&lt;/code&gt; isn't only a cap on what the model may write. It's also a reservation against your tokens-per-minute quota. Bedrock takes it before the model has written a word and gives it back when the call ends. The doc has its own example: the same request reserves 36,000 tokens up front with &lt;code&gt;max_tokens&lt;/code&gt; 32,000 and 5,250 with &lt;code&gt;max_tokens&lt;/code&gt; 1,250, and both settle at 9,000. Its conclusion contains the word that matters: &lt;em&gt;"fewer concurrent requests could be made because the max_tokens parameter was set too high."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F82nsxfoqcw28l9x0j6ei.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F82nsxfoqcw28l9x0j6ei.png" alt="The formula: in-flight requests × (input + maxTokens) must stay under TPM" width="799" height="244"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two consequences, before I get to the measurement:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If you leave &lt;code&gt;maxTokens&lt;/code&gt; out on Converse, the reservation is the model's maximum.&lt;/strong&gt; The API reference: &lt;em&gt;"The default value is the maximum allowed value for the model that you are using"&lt;/em&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InferenceConfiguration.html" rel="noopener noreferrer"&gt;InferenceConfiguration&lt;/a&gt;). That's 64K on Haiku 4.5 and 128K on Sonnet 5. On InvokeModel with the native Claude body the field is required, so people pick a big number "to be safe".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The output multiplier applies when the call settles, not to the reservation, and it doesn't change your bill.&lt;/strong&gt; As of 2026-09-30 the page lists &lt;strong&gt;15×&lt;/strong&gt; for Claude 4.8, &lt;strong&gt;10×&lt;/strong&gt; for Opus 5.5, Sonnet 5, Opus 5 and Fable 5.1 (and GPT-5.6 Sol, Terra and Luna on &lt;code&gt;bedrock-runtime&lt;/code&gt;), &lt;strong&gt;5×&lt;/strong&gt; for &lt;em&gt;"all other Anthropic models version 4.7 and below"&lt;/em&gt; (so Haiku 4.5), and &lt;strong&gt;1:1&lt;/strong&gt; for every other model. These tiers changed at least twice in 2026. A flat 5× that I had been quoting for a long time was out of date when I re-read the page for this talk, and the global cross-Region inference page still shows an older list. Cite the burndown page. &lt;em&gt;"You're only billed for your actual token usage."&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The false start
&lt;/h3&gt;

&lt;p&gt;My first version of this demo was a loop: 400 sequential calls at &lt;code&gt;maxTokens&lt;/code&gt; 64,000 from Singapore. It never throttled on tokens. Not once. Every 429 I had ever seen in a loop turned out to be the &lt;em&gt;requests-per-minute&lt;/em&gt; quota.&lt;/p&gt;

&lt;p&gt;The explanation is in the three-stage rule above. A call gives its reservation back when it completes, and a loop only ever has one call in flight. So the in-flight total never gets above one request's worth. The reservation is real. You just don't meet it until you have concurrency, and that's what I rebuilt the demo around.&lt;/p&gt;

&lt;h3&gt;
  
  
  The concurrency proof
&lt;/h3&gt;

&lt;p&gt;Claude Haiku 4.5 via its &lt;code&gt;global.&lt;/code&gt; profile from Singapore. The script reads the applied quotas live from Service Quotas: &lt;strong&gt;5,000,000 TPM, 10,000 RPM&lt;/strong&gt;. The prompt asks for a 300-word story, and every answer came out at about 390 output tokens, so a right-sized &lt;code&gt;maxTokens&lt;/code&gt; is 600. At the model maximum the predicted fit is 5,000,000 ÷ (20 + 64,000) ≈ &lt;strong&gt;78&lt;/strong&gt; calls in flight.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Measured 2026-09-25 from an EC2 runner in Bangkok, calling Singapore:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phase&lt;/th&gt;
&lt;th&gt;In flight × &lt;code&gt;maxTokens&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;Reserved in flight&lt;/th&gt;
&lt;th&gt;ok&lt;/th&gt;
&lt;th&gt;&lt;code&gt;ThrottlingException&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;wall clock&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A, model maximum&lt;/td&gt;
&lt;td&gt;117 × 64,000&lt;/td&gt;
&lt;td&gt;7.49M&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;76&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;41&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;15.5 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B, right-sized&lt;/td&gt;
&lt;td&gt;117 × 600&lt;/td&gt;
&lt;td&gt;0.07M&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;117&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;11.8 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C, control&lt;/td&gt;
&lt;td&gt;19 × 64,000&lt;/td&gt;
&lt;td&gt;1.22M&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;6.9 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D, one throttled call retried with full-jitter backoff&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;ok on the first retry&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On stage the script doesn't mention &lt;code&gt;maxTokens&lt;/code&gt; at first. It states the load in plain tokens, about 49,000 against a quota of 5,000,000 per minute, and asks whether all 117 requests get an answer. The obvious answer is yes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0ggr3isev0439c9qgiyh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0ggr3isev0439c9qgiyh.png" alt="Demo 4 as run on 2026-10-01: 117 requests, about 1% of the quota in real tokens, and 41 of them throttled" width="799" height="163"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv600lg7l9u5lsyot4g5i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv600lg7l9u5lsyot4g5i.png" alt="The reason, one Enter later: every request reserved 64,020 tokens, so about 78 fit" width="799" height="211"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fplihpe6kydkq5qfmsl42.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fplihpe6kydkq5qfmsl42.png" alt="The end of the run: the three phases, the documentation's own arithmetic and what it means" width="800" height="330"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Same prompt, same answers, same 117 users. One number changed, and 41 failures became none. The storm is far below RPM (117 requests in a minute against 10,000), so TPM is the only quota in play. And 76 against a prediction of 78 is about as close as a per-minute window lets you get.&lt;/p&gt;

&lt;p&gt;I repeated it a few times, always with 78 predicted: 74 ok / 43 throttled from my laptop on 2026-09-14; 76 / 41 from the runner the same day; 76 / 41 from a second AWS account calling &lt;em&gt;from Bangkok&lt;/em&gt; on 2026-09-14; 76 / 41 again from the runner on 2026-10-01. Five runs, two accounts, two source Regions: four came in at 76 and one at 74, against 78 predicted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The rule: in-flight requests × (input + maxTokens) must stay under TPM.&lt;/strong&gt; I'm stating that as measured on Claude Haiku 4.5 and consistent with the documentation. Read the open question further down before you generalise it.&lt;/p&gt;

&lt;h3&gt;
  
  
  RPM differs per Region in the same account
&lt;/h3&gt;

&lt;p&gt;While I was at it, I read the applied quotas in both Regions. These are Service Quotas &lt;em&gt;applied&lt;/em&gt; values for one account, read 2026-09-14 and re-read 2026-09-30, unchanged:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quota&lt;/th&gt;
&lt;th&gt;ap-southeast-7 (Bangkok)&lt;/th&gt;
&lt;th&gt;ap-southeast-1 (Singapore)&lt;/th&gt;
&lt;th&gt;documented default&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Global cross-Region model inference requests per minute for Anthropic Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;50&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;… tokens per minute for Anthropic Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;5,000,000&lt;/td&gt;
&lt;td&gt;5,000,000&lt;/td&gt;
&lt;td&gt;5,000,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;… requests per minute for Amazon Nova 2 Lite&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;2,000&lt;/td&gt;
&lt;td&gt;2,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;… requests per minute for Amazon Nova Lite (&lt;code&gt;apac.&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;400&lt;/td&gt;
&lt;td&gt;400&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every TPM value matched the General Reference default in both Regions. Every Bangkok RPM was lower. A second account, opted into &lt;code&gt;ap-southeast-7&lt;/code&gt; that same day, showed the documented defaults there, so the 50 belongs to my account and not to the Region. The Bedrock quotas page says applied values can be below the defaults for reasons including &lt;em&gt;"regional factors, payment history, fraudulent usage"&lt;/em&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/quotas.html" rel="noopener noreferrer"&gt;Bedrock quotas&lt;/a&gt;, read 2026-09-14).&lt;/p&gt;

&lt;p&gt;So I asked for the default through Service Quotas on 2026-09-14. The request was declined automatically within minutes. The reply said model access for an account &lt;em&gt;"may depend on factors such as regional availability, account history, and usage patterns"&lt;/em&gt; and &lt;em&gt;"is subject to change automatically as time passes."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I'd be careful about reading a moral into that. The account that had never used Bangkok got the defaults on the day it opted in. The older, heavily used one has 50. What I can say is this: read your own applied values, in every Region you call from, before you size anything. In my account a Haiku 4.5 loop from Bangkok hits RPM after about 50 calls, long before &lt;code&gt;maxTokens&lt;/code&gt; matters.&lt;/p&gt;

&lt;p&gt;One more thing I noticed. The 429 message was word for word the same for the RPM throttles and the TPM throttles (&lt;em&gt;"Too many requests, please wait before trying again."&lt;/em&gt;), so the error alone doesn't tell you which quota fired. Throttled calls also have no &lt;code&gt;inferenceRegion&lt;/code&gt; in CloudTrail, because they never left.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which error means what
&lt;/h3&gt;

&lt;p&gt;From the &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/troubleshooting-api-error-codes.html" rel="noopener noreferrer"&gt;troubleshooting page&lt;/a&gt;, re-read 2026-09-30:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Error&lt;/th&gt;
&lt;th&gt;HTTP&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;th&gt;Your move&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ThrottlingException&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;429&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;"exceeding the account quotas for Amazon Bedrock"&lt;/em&gt;: TPM, TPD or RPM&lt;/td&gt;
&lt;td&gt;retry with exponential backoff and jitter; right-size &lt;code&gt;maxTokens&lt;/code&gt;; read applied quotas; request an increase&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ServiceUnavailable&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;503&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;"not related to your account-level quotas or rate limits (which return 429 ThrottlingException)"&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;retry; another Region or cross-Region inference; Provisioned Throughput. Not a quota increase&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;overloaded_error&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;529&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;"a transient capacity error and is different from a 429 ThrottlingException"&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;honour &lt;code&gt;Retry-After&lt;/code&gt; if present, then back off&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On &lt;code&gt;ConverseStream&lt;/code&gt; these can arrive &lt;em&gt;inside&lt;/em&gt; the stream, after an HTTP 200. A streaming client has to handle errors after the headers too (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_ConverseStream.html" rel="noopener noreferrer"&gt;ConverseStream reference&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The AWS SDKs already retry for you. Standard mode is &lt;em&gt;"exponential backoff with full jitter"&lt;/em&gt;, three attempts in total, a 1,000 ms base for throttling errors and a 20-second cap (&lt;a href="https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html" rel="noopener noreferrer"&gt;SDK retry behavior&lt;/a&gt;). That page describes behaviour that needs &lt;code&gt;AWS_NEW_RETRIES_2026=true&lt;/code&gt; until it becomes the default, so check which behaviour your SDK version actually runs. And retries don't create capacity. If a quota fires under steady load, the fix is concurrency control on your side, a smaller &lt;code&gt;maxTokens&lt;/code&gt;, or more quota.&lt;/p&gt;

&lt;h3&gt;
  
  
  The open question: Amazon Nova 2 Lite
&lt;/h3&gt;

&lt;p&gt;I have to be straight about the limits of this. I saw the reservation rule on Claude Haiku 4.5. On Nova 2 Lite (Singapore, 8,000,000 TPM, 2,000 RPM) I couldn't make it throttle. 250 calls in flight at &lt;code&gt;maxTokens&lt;/code&gt; 64,000, which is twice the TPM if the rule applied, all succeeded (2026-09-14). So did 160 from Bangkok. A follow-up pushed 190 concurrent calls with 44,795 &lt;em&gt;real&lt;/em&gt; input tokens each and &lt;code&gt;maxTokens&lt;/code&gt; 5, about 8.5M input tokens in flight against an 8M quota, and again nothing was throttled.&lt;/p&gt;

&lt;p&gt;I haven't found a page that explains the difference. The burndown page lists Nova 2 Lite's output rate as 1:1 and says nothing model-specific about the reservation. A few things could explain it, and I'm not claiming any of them: a different or capped reservation for Nova, a reconciliation fast enough that the calls never add up to 8M at one instant, a different quota row gating the profile, or some tolerance for short bursts. &lt;code&gt;extra/nova2_input_storm.py&lt;/code&gt; in the repo is the experiment. A run at twice the quota with real input would settle it and costs about $7. If you know the answer, please open an issue.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do this:&lt;/strong&gt; set &lt;code&gt;maxTokens&lt;/code&gt; to what the answer needs. Size TPM as in-flight requests × (input + maxTokens), and remember the output multiplier for settlement. Read RPM for your Region and your account. Retry 429 with backoff and jitter, and treat 503 and 529 as capacity problems, not as your quota.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Cross-Region inference decides where your request runs
&lt;/h2&gt;

&lt;p&gt;This is the local one. There are three ways to name a model on Bedrock, and each makes a different promise (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/cross-region-inference.html" rel="noopener noreferrer"&gt;cross-Region inference&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/global-cross-region-inference.html" rel="noopener noreferrer"&gt;global cross-Region inference&lt;/a&gt;, re-read 2026-09-30):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fopp0s1w78376umd57qlw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fopp0s1w78376umd57qlw.png" alt="Three ways to name a model, three promises" width="798" height="259"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;In-Region model ID&lt;/strong&gt; (&lt;code&gt;anthropic.claude-…&lt;/code&gt;): runs in the Region you call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Geographic inference profile&lt;/strong&gt; (&lt;code&gt;us.&lt;/code&gt;, &lt;code&gt;eu.&lt;/code&gt;, &lt;code&gt;apac.&lt;/code&gt;, &lt;code&gt;au.&lt;/code&gt;, &lt;code&gt;jp.&lt;/code&gt;): runs somewhere in that geography; the destination list for a geographic profile does not change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Global inference profile&lt;/strong&gt; (&lt;code&gt;global.&lt;/code&gt;): may run in any commercial Region where the model is offered; the list can change.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Some things don't move. In the documentation's words: &lt;em&gt;"There's no additional routing cost for using cross-Region inference. The price is calculated based on the Region from which you call an inference profile."&lt;/em&gt; CloudWatch and CloudTrail log in the source Region. Global is listed as &lt;em&gt;"Approximately 10% savings"&lt;/em&gt; against geographic. On the pricing page (read 2026-09-14) that held for Haiku 4.5 in N. Virginia ($1.00 global vs $1.10 geographic). From Singapore or Thailand there's no geographic price for Claude 4.5 and newer at all, because no &lt;code&gt;apac.&lt;/code&gt; profile exists for them.&lt;/p&gt;

&lt;p&gt;AWS picks the destination. re:Post puts it bluntly: &lt;em&gt;"You can't choose a specific Region to process your request"&lt;/em&gt; (&lt;a href="https://repost.aws/knowledge-center/bedrock-cross-region-inference-routing" rel="noopener noreferrer"&gt;re:Post Knowledge Center&lt;/a&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  The Thailand fact
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;ListFoundationModels&lt;/code&gt; in &lt;code&gt;ap-southeast-7&lt;/code&gt;, one account:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Models listed&lt;/th&gt;
&lt;th&gt;Invocable in-Region on demand&lt;/th&gt;
&lt;th&gt;Inference-profile only&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-04&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-14&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-21&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-25&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-29&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;31&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The list grew on almost every reading. The in-Region count never moved from zero. The 31st model was OpenAI's GPT-6.1 Sol, announced on 29 September (&lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/09/openai-gpt-6-1-sol-on-amazon-bedrock/" rel="noopener noreferrer"&gt;What's New&lt;/a&gt;) and in Bangkok's list the next day. On 1 October each of the 31 had a system-defined inference profile in Bangkok: 27 &lt;code&gt;global.&lt;/code&gt; and 4 &lt;code&gt;apac.&lt;/code&gt;. So every Bedrock call made from Thailand is cross-Region, whether you planned for that or not.&lt;/p&gt;

&lt;p&gt;Global profiles reached Thailand as a source Region for Claude Opus 4.6, Sonnet 4.6 and Haiku 4.5 on 24 February 2026 (&lt;a href="https://aws.amazon.com/blogs/machine-learning/global-cross-region-inference-for-latest-anthropic-claude-opus-sonnet-and-haiku-models-on-amazon-bedrock-in-thailand-malaysia-singapore-indonesia-and-taiwan/" rel="noopener noreferrer"&gt;AWS ML blog&lt;/a&gt;). Singapore, for comparison, listed 40 models on 2026-10-01, of which five are invocable in-Region on demand: Claude 3 Haiku, Claude 3.5 Sonnet, two Cohere embedding models and Claude Sonnet 5. On 4 September it was the first four.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where the calls went
&lt;/h3&gt;

&lt;p&gt;CloudTrail records the destination of every cross-Region call in &lt;code&gt;additionalEventData.inferenceRegion&lt;/code&gt;, as a management event in the source Region, typically 5 to 15 minutes late (about 6 on 1 October). These are all the destinations I saw across my runs of 14, 25, 29 and 30 September and 1 October 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Called from&lt;/th&gt;
&lt;th&gt;Profile&lt;/th&gt;
&lt;th&gt;Served in&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bangkok&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;global.&lt;/code&gt; Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;Melbourne (every run)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bangkok&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;global.&lt;/code&gt; Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;Tokyo, Melbourne (25 Sep) · Ohio (29 Sep) · Sydney (30 Sep, 1 Oct)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bangkok&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;global.&lt;/code&gt; Amazon Nova 2 Lite&lt;/td&gt;
&lt;td&gt;Oregon (every run) · N. Virginia (14 Sep) · Tokyo (14 Sep, 1 Oct)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bangkok&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;apac.&lt;/code&gt; Amazon Nova Lite&lt;/td&gt;
&lt;td&gt;Tokyo (14, 25, 30 Sep, 1 Oct) · Sydney (14, 29 Sep, 1 Oct)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Singapore&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;global.&lt;/code&gt; Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;London, Ohio, Ireland, Oregon, Tokyo (14 Sep)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwspcugen4jqldzume43d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwspcugen4jqldzume43d.png" alt="Demo 5 as run on 2026-10-01: one raw CloudTrail event. Called in ap-southeast-7, served in ap-southeast-4" width="800" height="408"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4f06eyqow69muawcwlzk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4f06eyqow69muawcwlzk.png" alt="The same run: every call of the last three hours, by the Region that served it" width="799" height="372"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Thailand never shows up in the right-hand column. That's what zero in-Region models looks like in practice. The &lt;code&gt;apac.&lt;/code&gt; profile kept Nova Lite inside Asia Pacific. The &lt;code&gt;global.&lt;/code&gt; profiles went wherever there was capacity, and for one model that meant three continents in a single afternoon.&lt;/p&gt;

&lt;h3&gt;
  
  
  What "data stays in Region" does and does not mean
&lt;/h3&gt;

&lt;p&gt;The documentation makes three separate statements here, and it's worth keeping them apart:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Stored data stays at the source.&lt;/strong&gt; &lt;em&gt;"By default, the data remains stored only in the source Region."&lt;/em&gt; Logs, knowledge bases and configuration do not move (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/geographic-cross-region-inference.html" rel="noopener noreferrer"&gt;geographic cross-Region inference&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompts and outputs move for processing.&lt;/strong&gt; &lt;em&gt;"your input prompts and output results might move outside of your source Region during cross-Region inference."&lt;/em&gt; On the AWS network, encrypted in transit, never over the public internet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If anything is retained, it is retained at the destination.&lt;/strong&gt; &lt;em&gt;"If cross-region inference is enabled for these models, retained inputs and outputs are stored in destination Regions (i.e., the region where your inference request is processed)"&lt;/em&gt; (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/abuse-detection.html" rel="noopener noreferrer"&gt;abuse detection&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html" rel="noopener noreferrer"&gt;data retention&lt;/a&gt;). By default nothing is retained; for Claude Fable 5 and 5.1 all traffic is retained up to 30 days, so a Bangkok caller using Fable through its global profile has a 30-day copy in whatever Region served the call.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you have a compliance team, show them the table above before they come and ask.&lt;/p&gt;

&lt;h3&gt;
  
  
  The only control, and the single-Region path
&lt;/h3&gt;

&lt;p&gt;You can't pick or exclude destinations inside a profile. The one documented control is an IAM or SCP condition. A global profile is authorised against the Region-less ARN &lt;code&gt;arn:aws:bedrock:::foundation-model/…&lt;/code&gt;, and for that evaluation the service sets &lt;code&gt;aws:RequestedRegion&lt;/code&gt; to &lt;code&gt;unspecified&lt;/code&gt;. That's why Region-name deny policies &lt;em&gt;"don't target the Region-agnostic global foundation model resource evaluation"&lt;/em&gt;. To block global routing you deny &lt;code&gt;bedrock:*&lt;/code&gt; where &lt;code&gt;aws:RequestedRegion&lt;/code&gt; is &lt;code&gt;unspecified&lt;/code&gt; and &lt;code&gt;bedrock:InferenceProfileArn&lt;/code&gt; matches &lt;code&gt;inference-profile/global.*&lt;/code&gt;. To allow it through a Region-deny SCP you add &lt;code&gt;unspecified&lt;/code&gt; to the allow list or exempt on the profile ARN. If you use Control Tower, don't hand-edit its managed SCPs (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/global-cross-region-inference.html" rel="noopener noreferrer"&gt;global cross-Region inference&lt;/a&gt;). For a geographic profile, blocking any one destination Region in an SCP fails the whole request.&lt;/p&gt;

&lt;p&gt;Which models run in-Region differs per Region, and it changes. On 2026-10-01 a &lt;code&gt;Converse&lt;/code&gt; call with the bare ID &lt;code&gt;anthropic.claude-sonnet-5&lt;/code&gt; succeeded in Singapore and its CloudTrail event carried no &lt;code&gt;inferenceRegion&lt;/code&gt; at all, which is what a call that was not re-routed looks like (&lt;a href="https://aws.amazon.com/blogs/machine-learning/getting-started-with-cross-region-inference-in-amazon-bedrock/" rel="noopener noreferrer"&gt;AWS ML blog&lt;/a&gt;, read 14 Sep). Thailand offered no model that way. The documented single-Region path for the newest Claude models is the &lt;code&gt;bedrock-mantle&lt;/code&gt; endpoint with the bare model ID; the Haiku 4.5 model card (read 2026-09-14) lists it in seven Regions, none in Southeast Asia.&lt;/p&gt;

&lt;p&gt;One accounting aside from the same card: Anthropic models are &lt;em&gt;"offered and billed through AWS Marketplace. Charges appear on your AWS bill and in AWS Cost Explorer under the model provider (not under Amazon Bedrock)"&lt;/em&gt;. If you filter Cost Explorer on the Bedrock service, you won't see your Claude spend.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do this:&lt;/strong&gt; choose the profile type on purpose and know which destinations it allows. Put &lt;code&gt;inferenceRegion&lt;/code&gt; on a dashboard before compliance asks for it. And know that from some Regions, Thailand today, the newest models are global-only.&lt;/p&gt;




&lt;h2&gt;
  
  
  The pre-agent checklist
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgob8w1vup7z54bf9zwpg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgob8w1vup7z54bf9zwpg.png" alt="The five checks, as they assembled during the talk" width="800" height="133"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These are the five checks I'd run on one plain model call before putting agents, RAG or anything else on top. Each names the script that measures it. The same list lives in &lt;a href="https://github.com/spoecker/aws-th-community-day-2026-bedrock-fundamentals/blob/main/CHECKLIST.md" rel="noopener noreferrer"&gt;CHECKLIST.md&lt;/a&gt; in the repo.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Tokens&lt;/strong&gt; (&lt;code&gt;01_tokens.py&lt;/code&gt;): &lt;code&gt;usage.inputTokens&lt;/code&gt; measured on my model with my users' language and real text; the per-call overhead known; the metered number in the cost model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State&lt;/strong&gt; (&lt;code&gt;02_memory.py&lt;/code&gt;): history has a home, an owner and an expiry; every request sends exactly the history I intend and never another user's; the managed-state services I rely on exist in my Region.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget&lt;/strong&gt; (&lt;code&gt;03_context_budget.py&lt;/code&gt;): history bounded on purpose; long static prefixes first with a &lt;code&gt;cachePoint&lt;/code&gt; that meets the model's minimum, &lt;code&gt;cacheWriteInputTokens&lt;/code&gt; checked on the first call; per-user data after the checkpoint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quota&lt;/strong&gt; (&lt;code&gt;04_max_tokens_quota.py&lt;/code&gt;, &lt;code&gt;quotas_by_region.py&lt;/code&gt;): &lt;code&gt;maxTokens&lt;/code&gt; set to what the answer needs; TPM sized as in-flight × (input + maxTokens) plus the settlement multiplier; RPM read for my Region and account; 429 retried with backoff and jitter; 503 and 529 handled as capacity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing&lt;/strong&gt; (&lt;code&gt;05_where_did_it_run.py&lt;/code&gt;): the profile type chosen on purpose; &lt;code&gt;additionalEventData.inferenceRegion&lt;/code&gt; on a dashboard and shown to compliance; the destination list of a global profile understood as changeable.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;An agent is a loop of these calls. Fix them at one call, not at a thousand.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/spoecker/aws-th-community-day-2026-bedrock-fundamentals
&lt;span class="nb"&gt;cd &lt;/span&gt;aws-th-community-day-2026-bedrock-fundamentals
python3 &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
python3 demo1.py            &lt;span class="c"&gt;# tokens (Singapore)&lt;/span&gt;
python3 demo2.py            &lt;span class="c"&gt;# stateless, then memory in DynamoDB (Bangkok)&lt;/span&gt;
python3 demo3.py            &lt;span class="c"&gt;# context budget (Singapore)&lt;/span&gt;
python3 demo4.py            &lt;span class="c"&gt;# the concurrency storm (Singapore)&lt;/span&gt;
python3 demo5.py &lt;span class="nt"&gt;--fire&lt;/span&gt;     &lt;span class="c"&gt;# routing calls; 15 minutes later: python3 demo5.py&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You need Bedrock model access for Claude Haiku 4.5, Claude Sonnet 5, Nova Lite and Nova 2 Lite in the Region you call from, plus &lt;code&gt;cloudtrail:LookupEvents&lt;/code&gt;, &lt;code&gt;servicequotas:ListServiceQuotas&lt;/code&gt; and DynamoDB rights on one table. A full run of all five costs well under $1 at list prices. Each script writes a dated JSON file under &lt;code&gt;measurements/&lt;/code&gt;, and &lt;code&gt;python3 replay.py &amp;lt;1-5&amp;gt;&lt;/code&gt; prints the latest one without calling AWS, so you always have a real, dated result to compare against.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;Every page below was fetched and the quoted text read on it; "14 Sep" means read on 2026-09-14, "30 Sep" means re-read on 2026-09-30.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tokens.&lt;/strong&gt; &lt;a href="https://platform.claude.com/docs/en/about-claude/glossary" rel="noopener noreferrer"&gt;Anthropic glossary&lt;/a&gt; (30 Sep) · &lt;a href="https://platform.claude.com/docs/en/about-claude/models/overview" rel="noopener noreferrer"&gt;Anthropic models overview&lt;/a&gt; (21 Sep) · &lt;a href="https://platform.claude.com/docs/en/build-with-claude/token-counting" rel="noopener noreferrer"&gt;Anthropic token counting&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_CountTokens.html" rel="noopener noreferrer"&gt;CountTokens API reference&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/count-tokens.html" rel="noopener noreferrer"&gt;Count tokens user guide&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_TokenUsage.html" rel="noopener noreferrer"&gt;TokenUsage&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-anthropic-claude-haiku-4-5.html" rel="noopener noreferrer"&gt;Claude Haiku 4.5 model card&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-lite.html" rel="noopener noreferrer"&gt;Amazon Nova 2 Lite model card&lt;/a&gt; (14 Sep) · &lt;a href="https://aws.amazon.com/bedrock/pricing/" rel="noopener noreferrer"&gt;Amazon Bedrock pricing&lt;/a&gt; (14 Sep)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;State.&lt;/strong&gt; &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/conversation-inference.html" rel="noopener noreferrer"&gt;Converse user guide&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/claude-messages-thinking-block-binding.html" rel="noopener noreferrer"&gt;Thinking block binding&lt;/a&gt; (14 Sep) · &lt;a href="https://platform.claude.com/docs/en/build-with-claude/thinking" rel="noopener noreferrer"&gt;Anthropic extended thinking&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/sessions.html" rel="noopener noreferrer"&gt;Session management APIs&lt;/a&gt; (30 Sep) · &lt;a href="https://aws.amazon.com/about-aws/whats-new/2025/02/amazon-bedrock-session-management-apis-genai-applications-preview/" rel="noopener noreferrer"&gt;Session management APIs launch&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/memory.html" rel="noopener noreferrer"&gt;AgentCore Memory&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/agentcore-regions.html" rel="noopener noreferrer"&gt;AgentCore Regions&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/agents-classic-maintenance-mode.html" rel="noopener noreferrer"&gt;Bedrock Agents Classic maintenance mode&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html" rel="noopener noreferrer"&gt;Data retention&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/abuse-detection.html" rel="noopener noreferrer"&gt;Abuse detection&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html" rel="noopener noreferrer"&gt;Data protection&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/general/latest/gr/bedrock.html" rel="noopener noreferrer"&gt;AWS General Reference, Bedrock endpoints and quotas&lt;/a&gt; (14 Sep)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Budget.&lt;/strong&gt; &lt;a href="https://platform.claude.com/docs/en/build-with-claude/context-windows" rel="noopener noreferrer"&gt;Anthropic context windows&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_Converse.html" rel="noopener noreferrer"&gt;Converse API reference&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-anthropic-claude-messages-request-response.html" rel="noopener noreferrer"&gt;Claude messages parameters&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html" rel="noopener noreferrer"&gt;Prompt caching&lt;/a&gt; (30 Sep) · &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching" rel="noopener noreferrer"&gt;Anthropic prompt caching&lt;/a&gt; (14 Sep) · &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/01/amazon-bedrock-one-hour-duration-prompt-caching/" rel="noopener noreferrer"&gt;One-hour prompt caching&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentperf03-bp02.html" rel="noopener noreferrer"&gt;Well-Architected Agentic AI Lens, AGENTPERF03-BP02&lt;/a&gt; (14 Sep) · &lt;a href="https://aws.amazon.com/blogs/machine-learning/beyond-the-price-per-token-choosing-the-right-openai-model-on-amazon-bedrock-for-your-workload/" rel="noopener noreferrer"&gt;AWS ML blog, "Beyond the price per token"&lt;/a&gt; (14 Sep)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quota.&lt;/strong&gt; &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/quotas-token-burndown.html" rel="noopener noreferrer"&gt;How tokens are counted (burndown)&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/monitoring-runtime-metrics.html" rel="noopener noreferrer"&gt;Runtime metrics&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InferenceConfiguration.html" rel="noopener noreferrer"&gt;InferenceConfiguration&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/quotas-runtime.html" rel="noopener noreferrer"&gt;Runtime quotas&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/quotas.html" rel="noopener noreferrer"&gt;Bedrock quotas&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/troubleshooting-api-error-codes.html" rel="noopener noreferrer"&gt;Troubleshooting API error codes&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_ConverseStream.html" rel="noopener noreferrer"&gt;ConverseStream&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html" rel="noopener noreferrer"&gt;SDK retry behavior&lt;/a&gt; (30 Sep) · &lt;a href="https://repost.aws/knowledge-center/bedrock-throttling-error" rel="noopener noreferrer"&gt;re:Post, Bedrock throttling&lt;/a&gt; (14 Sep)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing.&lt;/strong&gt; &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/cross-region-inference.html" rel="noopener noreferrer"&gt;Cross-Region inference&lt;/a&gt; (30 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/geographic-cross-region-inference.html" rel="noopener noreferrer"&gt;Geographic cross-Region inference&lt;/a&gt; (14 Sep) · &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/global-cross-region-inference.html" rel="noopener noreferrer"&gt;Global cross-Region inference&lt;/a&gt; (30 Sep) · &lt;a href="https://aws.amazon.com/blogs/machine-learning/global-cross-region-inference-for-latest-anthropic-claude-opus-sonnet-and-haiku-models-on-amazon-bedrock-in-thailand-malaysia-singapore-indonesia-and-taiwan/" rel="noopener noreferrer"&gt;Global CRIS in Thailand, Malaysia, Singapore, Indonesia and Taiwan&lt;/a&gt; (14 Sep) · &lt;a href="https://repost.aws/knowledge-center/bedrock-cross-region-inference-routing" rel="noopener noreferrer"&gt;re:Post, cross-Region inference routing&lt;/a&gt; (14 Sep) · &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/09/openai-gpt-6-1-sol-on-amazon-bedrock/" rel="noopener noreferrer"&gt;What's New, GPT-6.1 Sol on Amazon Bedrock&lt;/a&gt; (1 Oct)&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Alexander Spoecker is a Cloud Solution Architect at Iglu in Thailand and an AWS Authorized Instructor. This post is my own work and not a statement by AWS or by Iglu. The scripts are MIT licensed.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>genai</category>
      <category>python</category>
    </item>
  </channel>
</rss>
