DEV Community

Alexander Spoecker
Alexander Spoecker

Posted on AI-assisted

Before You Build an Agent: Five Amazon Bedrock Fundamentals, Measured

This is the long version of my 30-minute talk at AWS Community Day Thailand, Bangkok, 3 October 2026. The code, the checklist and every dated measurement are in the repo.

I work as a Cloud Solution Architect at Iglu in Chiang Mai, Thailand, and I also teach AWS classes as an AWS Authorized Instructor, mostly online. When I started preparing this talk I set myself one goal: no number goes on a slide unless I measured it myself, and no rule goes on a slide unless I can point at the page it comes from.

A warning before the numbers. Everything here was measured in one AWS account on the date written next to it, and every rule links the AWS or Anthropic page I read it on, with the date I read it. This stuff moves. Often really fast as we know with AI. In the four weeks I worked on the talk, the Thailand Region's model list grew from 21 to 31, two new Claude models showed up, and one documentation page I cite changed its numbers. So if your numbers differ from mine, don't assume one of us is wrong. Re-run the scripts and tell me what you got. I probably won't manage to keep the numbers updated in this post.

What Measured Where the calls were made File in the repo
Tokens per language and model 2026-09-25, identical 2026-10-01 (spot check 2026-09-30) ap-southeast-1 (Singapore) measurements/01-tokens-2026-09-25.json, 01-tokens-spotcheck-claude-5-5-2026-09-30.json
Stateless calls, then memory in DynamoDB 2026-09-25 ap-southeast-7 (Thailand) 02-memory-stateless-…, 02-memory-with-memory-2026-09-25.json
Context growth, caching, windowing 2026-09-25 ap-southeast-1 03-context-2026-09-25.json
117 concurrent calls against the TPM quota 2026-09-25, identical 2026-10-01 (first 2026-09-14) ap-southeast-1, from an EC2 runner in Bangkok 04-quota-haiku-2026-09-25.json
Where Bedrock ran each request 2026-09-14, 25, 29, 30 and 2026-10-01 ap-southeast-7 05-routing-2026-*.json
Applied quotas per Region 2026-09-14 (re-read 2026-09-30, unchanged) both quotas-by-region-2026-09-14.json

Unless a different date is given next to a link, I read the documentation on 2026-09-14 and again on 2026-09-30.

The three surprises

The talk is built around three surprises that come with an application on Amazon Bedrock. One of them you meet on the first day of development (Specially if you are used to using only consumer AI products before). The other two wait until real users and real traffic arrive. (Or sometimes when expanding to a new country)

  1. The model forgets. This is the early one. In the first test, turn two knows nothing about turn one, and you find out that conversation history isn't a feature you switch on.
  2. The bill is wrong. The price per token was right. The number of tokens wasn't, because the estimate was made in English and the customers write Thai.
  3. It throttles at a fraction of the expected traffic. Everybody reads the tokens-per-minute quota. Almost nobody reads how it gets used up.

None of these has anything to do with the model being good or bad, and none of them is an agent problem. They're properties of one plain model call.

A colleague who reviewed my slides asked me what an agent actually is. The answer I ended up with is the reason for the title: an agent is a loop of model calls that decides the next step and calls tools. Whatever is true for a single call is true for every step of that loop, just more often. So I'd rather get it right at one call.

The example running through the whole post is an airline's customer chat, with one Thai customer and one request. Everything is plain boto3 and the Converse API. There's no framework anywhere in the repo.


1. Tokens are not words

A token is the chunk a model's tokenizer cuts your text into. You pay per token, and the context window is counted in tokens too. For Claude, Anthropic's glossary says a token "approximately represents 3.5 English characters, though the exact number can vary depending on the language used" (Anthropic glossary). The count also depends on the model: token counting is model-specific (CountTokens API reference). What I couldn't find anywhere, from AWS or from Anthropic, was a number for Thai. So I measured it.

The measurement

I took one sentence a customer might type, in four languages. My Thai is very limited, so I had a native speaker read the Thai sentence before I put it on a slide. Each sentence goes to each model once with maxTokens: 1, and I read usage.inputTokens from the response. That field is the one you're billed on: CloudWatch's InputTokenCount for the same calls added up to exactly these values when I checked on 2026-09-04.

EN (89 characters): Hello, I would like to change my flight to next Tuesday morning. Is there a fee for that?
TH (95 characters): สวัสดีครับ ผมต้องการเปลี่ยนเที่ยวบินเป็นเช้าวันอังคารหน้า มีค่าธรรมเนียมสำหรับการเปลี่ยนไหมครับ

The same sentence in English and Thai as token chunks: 28 versus 90 on Claude Haiku 4.5

Measured 2026-09-25, Singapore, usage.inputTokens (all four rows identical again on 2026-10-01; the Haiku 4.5 and Nova rows also on 4, 14 and 21 September from both Singapore and Bangkok):

Model (inference profile) "Hi" EN (89 chars) DE (95) ZH (26) TH (95) TH ÷ EN
Claude Haiku 4.5 (global.) 8 28 48 35 90 3.2×
Claude Sonnet 5 (global.; Opus 5 counted the same on 21 Sep) 9 37 59 33 89 2.4×
Amazon Nova Lite (apac.) 1 21 24 28 75 3.6×
Amazon Nova 2 Lite (global.) 47 67 70 68 80 1.2× (1.6× net of "Hi")

On 2026-09-30 I did a spot check for Claude Sonnet 5.5 and Opus 5.5, which had both joined my account's model list in the last week of September. Both count "Hi" 11, EN 39, TH 91, so 2.3×. Every count is exactly two above Sonnet 5. To me that looks like two more fixed framing tokens and not a new tokenizer, but that's my guess. It isn't documented.

Demo 1 as run on 2026-10-01: the raw Converse request and the usage block of its response

Demo 1, same run: the token table for four models

What I take from the table

  • Thai costs 2.4 to 3.6 times as much as English for the same sentence, and where you land in that range depends on the tokenizer. On Claude Haiku 4.5 it's close to one token per Thai character (95 characters, 90 tokens). If your budget comes from the English rule of thumb, it's off by a factor of three before the first customer has typed anything.
  • The newest Claude tokenizer made English more expensive. It did not make Thai cheaper. Anthropic's models overview (read 2026-09-21) says the current tokenizer, introduced with Opus 4.7, fits about 555k words in 1M tokens where the previous one fit about 750k (Anthropic models overview). In my measurement the English sentence went from 28 tokens on Haiku 4.5 to 37 on Sonnet 5, which is 32% more. Thai went from 90 to 89. So counts you measured on one model generation don't carry over to the next.
  • Nova 2 Lite charges about 46 tokens of overhead on every call. The single word "Hi" costs 47 input tokens. I didn't find this documented anywhere, so treat it as my measurement and nothing more. It does mean that fifty one-word calls cost more than one fifty-word call.
  • Every model has some per-call overhead. "Hi" is 8 tokens on Haiku 4.5 and 1 on Nova Lite. That's worth knowing before you price a workload with lots of short messages.

How to count, and where counting doesn't work

usage comes back with every response (TokenUsage): inputTokens, outputTokens, totalTokens, plus cacheReadInputTokens and cacheWriteInputTokens when caching is involved. With streaming it arrives at the very end, in the final metadata event. A client that stops reading early never sees it.

Bedrock also has a CountTokens operation, and it's free (user guide: "doesn't incur charges"). I expected it to be the easy answer. It wasn't, for three reasons:

  1. From Singapore and Thailand it rejected every inference profile I tried ("The provided model doesn't support counting tokens"). In us-east-1 it accepted the base model ID anthropic.claude-haiku-4-5-20251001-v1:0, which you can't invoke on demand there. So you count with one ID and invoke with another.
  2. Its answer was always 16 tokens above what Converse billed for the same text on Haiku 4.5 (24 vs 8, 44 vs 28, 106 vs 90; measured 2026-09-04, 14 and 25, and again on 2026-10-01). The API reference says the count "will match the token count that would be charged". The closest thing to an explanation I found is on Anthropic's side: counts "may include tokens added automatically by Anthropic for system optimizations. You are not billed for system-added tokens" (Anthropic token counting). That would fit a fixed offset. I haven't been able to confirm it for Bedrock.
  3. It only works for Claude, and not even for all of Claude. Models that launch as cross-Region-only on bedrock-runtime "don't support CountTokens on bedrock-runtime". The doc points you to Anthropic's count_tokens route on the bedrock-mantle endpoint instead, and that endpoint (Haiku 4.5 model card, read 2026-09-14) exists in seven Regions, none of them in Southeast Asia. Amazon Nova doesn't support it at all (Nova 2 Lite model card).

So from Bangkok today, the counter that actually works is a Converse call with maxTokens: 1. A few dozen of those cost well under a cent, and what you get back is the billed number.

There's one more thing here that I can't explain. On 21, 25, 29 and 30 September and on 1 October, my account was refused Claude Sonnet 5 and Opus 5 when I called from ap-southeast-7 ("not available for this account", AccessDeniedException). The same global. profile answered from ap-southeast-1 every single time, and the availability API in Bangkok reported the model agreement as available. I still don't know why. The lesson I took from it: before you promise anyone a model from a Region, call it from that Region.

Do this: measure usage.inputTokens on the model you'll ship, with real text in your users' language. Find out the per-call overhead. And build your cost model on the billed number, not on what a counter estimates.


2. Conversation history is application state

Does the model remember? The Converse user guide answers that in two sentences (conversation-inference, re-read 2026-09-30):

"Amazon Bedrock doesn't store any text, images, or documents that you provide as content. The data is only used to generate the response."

"You can maintain conversation context by including all the messages in the conversation in subsequent Converse requests."

So the messages array is the memory. Your code owns it, and you pay for it again on every turn.

Call 2 knows nothing about call 1: the messages array is what grows

The two-call demo

Two Converse calls from Bangkok on Claude Haiku 4.5. The system prompt is "You are a concise travel assistant for an airline. Answer in one sentence." This is the run of 2026-09-25:

request: 1 message(s), 44 input tokens
user : Hi, I am Alex. I am flying to Bangkok for AWS Community Day on 3 October.
model: How can I assist you with your flight to Bangkok on October 3rd, Alex?

request: 1 message(s), 31 input tokens
user : What should I pack for the trip?
model: Pack comfortable clothing appropriate for your destination's climate, a valid
       passport/ID, medications, toiletries, electronics with chargers, and any documents
       needed for your airline check-in.
Enter fullscreen mode Exit fullscreen mode

The second answer is generic because the second request contained exactly one message. Bangkok, October and the conference were never sent.

Memory in three API calls

Now the same two turns with the history kept in a DynamoDB table. The partition key is userId, the sort key is ts, and a ttl attribute lets idle conversations expire after 24 hours. Each turn is three calls: Query this user's turns, Converse with that history plus the new message, PutItem for both new turns.

Query, Converse, PutItem: memory in three calls, with the table schema

request: 1 message(s) (0 from DynamoDB), 44 input tokens
user : Hi, I am Alex. I am flying to Bangkok for AWS Community Day on 3 October.
model: Great! I'd be happy to help with your flight to Bangkok on October 3rd ...

request: 3 message(s) (2 from DynamoDB), 95 input tokens
user : What should I pack for the trip?
model: For Bangkok in early October, pack light breathable clothing, sunscreen, an umbrella
       or rain jacket (monsoon season), comfortable walking shoes, and any medications you need.

user 'bob' asks the same question (1 message sent)
model: Pack based on your destination's weather and trip length ...
Enter fullscreen mode Exit fullscreen mode

The second request now carries three messages and 95 input tokens instead of 31. That's the price of memory: you pay for it as input, on every turn. In return the answer is about Bangkok in October.

Then Bob asks the identical question and gets the generic answer again, because the Query was scoped to his userId and came back empty. That scoping is what keeps one customer's conversation out of another customer's answer. It's your code, so it's your job to get it right.

Demo 2 as run from Bangkok on 2026-10-01: the two turns with history from DynamoDB. The lines under the title recap what the stateless calls answered a moment earlier

Screenshots in this post are frames from the demo recordings of 1 October 2026. The model's wording differs a little from the 25 September transcript above; the behaviour does not.

The API expects you to replay everything

If you want proof that the API is designed around you sending the history back, look at reasoning. When a Claude model reasons, the assistant message carries a reasoningContent block with a signature, and the Converse guide says: "The signature field is a hash of all the messages in the conversation and is a safeguard against tampering of the reasoning used by the model. You must include the signature and all previous messages in subsequent Converse requests. If any of the messages are changed, the response throws an error." A service that kept the state itself wouldn't need the client to carry a hash around.

This matters if your code rewrites history. From Claude Fable 5.1 on, the API checks that "the system prompt, the tool list, and all messages before the block are unchanged". Summarising older turns or injecting per-turn reminders then fails with Invalid signature in thinking block. The block is bound to a different conversation, and the error "is permanent for that request — an automatic retry loop will not clear it" (thinking block binding, read 2026-09-14). You have two ways out: strip the thinking blocks from the point where you rewrote history, or send the thinking-binding-controls-2026-08-01 beta with mismatch_behavior: "drop_block" through additionalModelRequestFields. Haiku 4.5 and earlier models strip older thinking blocks automatically (Anthropic, extended thinking). Plain text history without thinking blocks is never checked. The model just answers whatever you sent.

If you'd rather not build it yourself

There are managed options. Each one is a separate service, and the model API stays stateless whichever you pick. Look at the last two columns before you plan around one of them.

Option What it stores Retention Singapore Thailand
Your own store (the demo: DynamoDB) whatever you write your TTL yes yes
Bedrock Session Management APIs (preview) checkpoints for LangGraph/LlamaIndex-style apps; 1,000 steps per session, 50 MB per step idle timeout 1 hour; "automatically deleted after 30 days" yes (launch list, Feb 2025) not listed; no bedrock-agent-runtime endpoint row for ap-southeast-7 in the General Reference (read 2026-09-14)
Amazon Bedrock AgentCore Memory short-term raw events per actor and session; long-term records extracted asynchronously by strategies events 7 to 365 days; long-term records have no built-in TTL yes no (AgentCore Regions table, read 2026-09-30; Runtime is available in Thailand, Memory is not)
Bedrock Agents (Classic) sessionId and memory agent session state per agent config closed to new customers since 30 July 2026 (maintenance mode page, read 2026-09-14) same

AgentCore's own documentation says it in one line: "AgentCore Memory addresses a fundamental challenge in agentic AI: statelessness." The service exists because the model call has no memory.

What "doesn't store" does not mean

"Doesn't store" in the Converse guide is about conversation state. It isn't a statement about retention policy. Bedrock's data-retention page (read 2026-09-30) describes a mode that you set per Region and per account: none (zero data retention), default (the model's own policy; "AWS may retain the data for safety and abuse-prevention purposes"), aws_review and a legacy provider_data_share. The same page says plainly: "There is no data retention change to Claude models released before Claude Fable 5." Claude Fable 5 and 5.1 require aws_review, and under that mode prompts and completions are "retained within the AWS boundary for up to 30 days" and may be reviewed by AWS. "Your content is not shared with the model provider." (data retention, abuse detection).

Nothing you send is used to train a model, and the providers have no access to it. Each provider's model runs in a Model Deployment Account owned by the Bedrock team, and "Model providers don't have any access to those accounts" (data protection).

Haiku 4.5, Sonnet 5, Opus 5 and Nova were not on the retention list when I read it on 30 September. If you use Fable, section 5 has one more sentence for you about where that retained copy sits.

Do this: decide where history lives, who can read it and when it expires, and do that before the first prompt ships. Send exactly the history you mean to send, scoped per user. And check that the managed-state service you're counting on exists in your Region.


3. The context window is a finite budget

Everything you send has to fit, and you pay for everything you send. Anthropic's context-window page lists what counts: the system prompt, every message, tool definitions, tool results, and the output including thinking. Cached prefixes still take up space in the window (context windows). When you overflow on Claude 4.5 and later, you get stopReason: model_context_window_exceeded (Converse API reference).

Model Context window Max output Source (read)
Claude Haiku 4.5 200K 64K model card (2026-09-14)
Claude Sonnet 5, Opus 5 1M 128K Anthropic models overview (2026-09-21)
Amazon Nova 2 Lite 1M 64K model card (2026-09-14)

A word of caution about "1M". For Sonnet 4 and Sonnet 4.5 on Bedrock, 1M is a preview variant with its own, much smaller quota rows. The Sonnet 4 preview also documents a long-context premium that re-rates the whole request once it goes above 200K input tokens (Claude messages parameters). Check the pricing page for the model you actually use.

The curve

Per turn, the input grows linearly. Over a whole conversation the total grows, in AWS's words, "approximately quadratically" with its length (AWS ML blog, 2026-09-11). Here are ten short airline-support turns on Haiku 4.5 with the full history replayed each time, measured 2026-09-25 from Singapore:

Input tokens per turn, 50 to 545 over ten turns

turn:          1    2    3    4    5    6    7    8    9   10
inputTokens:  50   88  137  191  253  320  377  445  501  545
Enter fullscreen mode Exit fullscreen mode

That's 2,907 input tokens for ten questions of about 20 tokens each. At this size it costs next to nothing. I'm showing it for the shape.

Four runs, one lesson

Then I ran the same ten turns four ways. Claude Haiku 4.5, Singapore list prices as I read them on the pricing page on 2026-09-14 (input $1.00, five-minute cache write $1.25, cache read $0.10 per 1M tokens). "Input cost" means input tokens plus cache writes plus cache reads at list price. Output isn't in it.

Run Input tokens, turn 1 → 10 Cache write (turn 1) Cache read (each later turn) Input cost
A0 short system prompt, full history 50 → 545 0 0 $0.003
A 8K-token airline policy in the system prompt, full history 7,870 → 8,305 0 0 $0.081
B same, cachePoint after the policy 33 → 599 (the non-cached part) 7,837 7,837 $0.020
C same as A, only the last 4 turns sent 7,870, then flat between 7,897 and 7,931 0 0 $0.079

Demo 3 as run on 2026-10-01: all four runs side by side. This run came out at 50 to 567 tokens and $0.081, $0.020 and $0.079; the table above is the 25 September run

Run B turn by turn: 7,837 tokens written to the cache on turn 1 and read on every later turn, while inputTokens shows only the growing, non-cached history

Look at C first. Windowing barely moved the bill, because the static policy, and not the history, was the big part of every request. Caching the policy (B) cut the input cost to a quarter. In a chat with a short system prompt and a long history it would be the other way round.

Either way you give something up: the model no longer knows turn one unless you carry it forward on purpose. The Well-Architected Agentic AI Lens asks for "conversation history bounded by summarization, sliding windows, or semantic compression so prompt size doesn't grow linearly with session length" and lists sending the full history every time as its first anti-pattern (AGENTPERF03-BP02).

The caching rules that decide whether it works

All of this is from the Bedrock prompt caching page as I re-read it on 2026-09-30, and from the Anthropic prompt caching page.

  • There's a minimum prefix size, and it differs per model. "Claude Opus 5 requires at least 512 tokens per cache checkpoint, Claude Sonnet 5 requires at least 1,024 tokens per cache checkpoint, and Claude Haiku 4.5 requires at least 4,096 tokens." The table on that page (30 Sep) lists Sonnet 5.5, Opus 5.5 and Fable 5.1 at 512. The minimum counts everything before the checkpoint, across tools, system and messages.
  • Below the minimum it fails without telling you. "your inference still succeeds, but your prefix isn't cached." There's no error. The only sign is cacheWriteInputTokens staying at 0. A 3,000-token system prompt on Haiku 4.5 with a cachePoint does nothing at all, and I'd bet that's behind most "my cache never hits" questions. With Thai at roughly one token per character, 4,096 tokens is about 4,000 Thai characters but about 14,000 English ones.
  • Order matters. "Cache checkpoints are processed in this order: tools → system → messages … changing content in an earlier section invalidates the cache for later sections." So static content goes first, then the checkpoint, then the per-user data.
  • TTL. Five minutes by default, refreshed on every hit. One hour is available on Claude 4.5 and later models, at a higher write price (What's New, Jan 2026). Anthropic adds a detail that's easy to miss: the lifetime is measured from the start of the request that wrote or read the entry, and an entry only becomes available after the first response begins. Ten parallel calls with the same prefix are ten cache writes.
  • The cache is per account. That's why per-user data goes after the checkpoint. It's also why my demo prints a warning if anyone in the account ran it in the last five minutes: turn one then shows a read instead of a write.
  • Cache reads don't count against your TPM quota. Cache writes do (token burndown). A well-cached workload gets more throughput out of the same quota. Section 4 is about that quota.
  • Caching works with cross-Region inference, with one caveat: "At times of high demand, these optimizations may lead to increased cache writes". A different destination Region may not have your entry.
  • Different providers, different multiples. On the Bedrock pricing page (Singapore, Anthropic tab, read 2026-09-14) Claude's cache multiples are 1.25× for a five-minute write, 2× for a one-hour write and 0.1× for a read. Nova 2 Lite's cache read was priced at 25% of input and its write at $0. So don't quote one universal cache discount.

Caching doesn't pay in three cases: a prefix under the minimum (nothing happens), a prefix that changes on every request (every call is a write at 1.25× with no reads, so you pay more than without caching), and batch inference (not supported).

Do this: bound the history on purpose, with a window, a summary, or a conscious decision to pay for all of it. Put the long static prefix first with a cachePoint after it, and check cacheWriteInputTokens on the first call. Put per-user data after the checkpoint.


4. Output-token limits consume quota before the model writes a word

Of the five, this is the one I got wrong first. It's also the one the documentation states most precisely. From the token burndown page, re-read 2026-09-30:

At the start of the request – The following sum is deducted from your quotas. The request is throttled if you exceed a quota. Total input tokens + max_tokens
During processing – The quota consumed by the request is periodically adjusted to account for the actual number of output tokens generated.
At the end of the request – The total number of tokens consumed is Input token count + Cache write input tokens + (Output token count × Burndown rate).

And from the CloudWatch runtime-metrics page, on the EstimatedTPMQuotaUsage metric: it "does not reflect the reservation-based token consumption that drives throttling decisions. Throttling is based on the upfront reservation of input tokens plus max_tokens" (runtime metrics).

So maxTokens isn't only a cap on what the model may write. It's also a reservation against your tokens-per-minute quota. Bedrock takes it before the model has written a word and gives it back when the call ends. The doc has its own example: the same request reserves 36,000 tokens up front with max_tokens 32,000 and 5,250 with max_tokens 1,250, and both settle at 9,000. Its conclusion contains the word that matters: "fewer concurrent requests could be made because the max_tokens parameter was set too high."

The formula: in-flight requests × (input + maxTokens) must stay under TPM

Two consequences, before I get to the measurement:

  • If you leave maxTokens out on Converse, the reservation is the model's maximum. The API reference: "The default value is the maximum allowed value for the model that you are using" (InferenceConfiguration). That's 64K on Haiku 4.5 and 128K on Sonnet 5. On InvokeModel with the native Claude body the field is required, so people pick a big number "to be safe".
  • The output multiplier applies when the call settles, not to the reservation, and it doesn't change your bill. As of 2026-09-30 the page lists 15× for Claude 4.8, 10× for Opus 5.5, Sonnet 5, Opus 5 and Fable 5.1 (and GPT-5.6 Sol, Terra and Luna on bedrock-runtime), 5× for "all other Anthropic models version 4.7 and below" (so Haiku 4.5), and 1:1 for every other model. These tiers changed at least twice in 2026. A flat 5× that I had been quoting for a long time was out of date when I re-read the page for this talk, and the global cross-Region inference page still shows an older list. Cite the burndown page. "You're only billed for your actual token usage."

The false start

My first version of this demo was a loop: 400 sequential calls at maxTokens 64,000 from Singapore. It never throttled on tokens. Not once. Every 429 I had ever seen in a loop turned out to be the requests-per-minute quota.

The explanation is in the three-stage rule above. A call gives its reservation back when it completes, and a loop only ever has one call in flight. So the in-flight total never gets above one request's worth. The reservation is real. You just don't meet it until you have concurrency, and that's what I rebuilt the demo around.

The concurrency proof

Claude Haiku 4.5 via its global. profile from Singapore. The script reads the applied quotas live from Service Quotas: 5,000,000 TPM, 10,000 RPM. The prompt asks for a 300-word story, and every answer came out at about 390 output tokens, so a right-sized maxTokens is 600. At the model maximum the predicted fit is 5,000,000 ÷ (20 + 64,000) ≈ 78 calls in flight.

Measured 2026-09-25 from an EC2 runner in Bangkok, calling Singapore:

Phase In flight × maxTokens Reserved in flight ok ThrottlingException wall clock
A, model maximum 117 × 64,000 7.49M 76 41 15.5 s
B, right-sized 117 × 600 0.07M 117 0 11.8 s
C, control 19 × 64,000 1.22M 19 0 6.9 s
D, one throttled call retried with full-jitter backoff ok on the first retry

On stage the script doesn't mention maxTokens at first. It states the load in plain tokens, about 49,000 against a quota of 5,000,000 per minute, and asks whether all 117 requests get an answer. The obvious answer is yes.

Demo 4 as run on 2026-10-01: 117 requests, about 1% of the quota in real tokens, and 41 of them throttled

The reason, one Enter later: every request reserved 64,020 tokens, so about 78 fit

The end of the run: the three phases, the documentation's own arithmetic and what it means

Same prompt, same answers, same 117 users. One number changed, and 41 failures became none. The storm is far below RPM (117 requests in a minute against 10,000), so TPM is the only quota in play. And 76 against a prediction of 78 is about as close as a per-minute window lets you get.

I repeated it a few times, always with 78 predicted: 74 ok / 43 throttled from my laptop on 2026-09-14; 76 / 41 from the runner the same day; 76 / 41 from a second AWS account calling from Bangkok on 2026-09-14; 76 / 41 again from the runner on 2026-10-01. Five runs, two accounts, two source Regions: four came in at 76 and one at 74, against 78 predicted.

The rule: in-flight requests × (input + maxTokens) must stay under TPM. I'm stating that as measured on Claude Haiku 4.5 and consistent with the documentation. Read the open question further down before you generalise it.

RPM differs per Region in the same account

While I was at it, I read the applied quotas in both Regions. These are Service Quotas applied values for one account, read 2026-09-14 and re-read 2026-09-30, unchanged:

Quota ap-southeast-7 (Bangkok) ap-southeast-1 (Singapore) documented default
Global cross-Region model inference requests per minute for Anthropic Claude Haiku 4.5 50 10,000 10,000
… tokens per minute for Anthropic Claude Haiku 4.5 5,000,000 5,000,000 5,000,000
… requests per minute for Amazon Nova 2 Lite 200 2,000 2,000
… requests per minute for Amazon Nova Lite (apac.) 40 400 400

Every TPM value matched the General Reference default in both Regions. Every Bangkok RPM was lower. A second account, opted into ap-southeast-7 that same day, showed the documented defaults there, so the 50 belongs to my account and not to the Region. The Bedrock quotas page says applied values can be below the defaults for reasons including "regional factors, payment history, fraudulent usage" (Bedrock quotas, read 2026-09-14).

So I asked for the default through Service Quotas on 2026-09-14. The request was declined automatically within minutes. The reply said model access for an account "may depend on factors such as regional availability, account history, and usage patterns" and "is subject to change automatically as time passes."

I'd be careful about reading a moral into that. The account that had never used Bangkok got the defaults on the day it opted in. The older, heavily used one has 50. What I can say is this: read your own applied values, in every Region you call from, before you size anything. In my account a Haiku 4.5 loop from Bangkok hits RPM after about 50 calls, long before maxTokens matters.

One more thing I noticed. The 429 message was word for word the same for the RPM throttles and the TPM throttles ("Too many requests, please wait before trying again."), so the error alone doesn't tell you which quota fired. Throttled calls also have no inferenceRegion in CloudTrail, because they never left.

Which error means what

From the troubleshooting page, re-read 2026-09-30:

Error HTTP Meaning Your move
ThrottlingException 429 "exceeding the account quotas for Amazon Bedrock": TPM, TPD or RPM retry with exponential backoff and jitter; right-size maxTokens; read applied quotas; request an increase
ServiceUnavailable 503 "not related to your account-level quotas or rate limits (which return 429 ThrottlingException)" retry; another Region or cross-Region inference; Provisioned Throughput. Not a quota increase
overloaded_error 529 "a transient capacity error and is different from a 429 ThrottlingException" honour Retry-After if present, then back off

On ConverseStream these can arrive inside the stream, after an HTTP 200. A streaming client has to handle errors after the headers too (ConverseStream reference).

The AWS SDKs already retry for you. Standard mode is "exponential backoff with full jitter", three attempts in total, a 1,000 ms base for throttling errors and a 20-second cap (SDK retry behavior). That page describes behaviour that needs AWS_NEW_RETRIES_2026=true until it becomes the default, so check which behaviour your SDK version actually runs. And retries don't create capacity. If a quota fires under steady load, the fix is concurrency control on your side, a smaller maxTokens, or more quota.

The open question: Amazon Nova 2 Lite

I have to be straight about the limits of this. I saw the reservation rule on Claude Haiku 4.5. On Nova 2 Lite (Singapore, 8,000,000 TPM, 2,000 RPM) I couldn't make it throttle. 250 calls in flight at maxTokens 64,000, which is twice the TPM if the rule applied, all succeeded (2026-09-14). So did 160 from Bangkok. A follow-up pushed 190 concurrent calls with 44,795 real input tokens each and maxTokens 5, about 8.5M input tokens in flight against an 8M quota, and again nothing was throttled.

I haven't found a page that explains the difference. The burndown page lists Nova 2 Lite's output rate as 1:1 and says nothing model-specific about the reservation. A few things could explain it, and I'm not claiming any of them: a different or capped reservation for Nova, a reconciliation fast enough that the calls never add up to 8M at one instant, a different quota row gating the profile, or some tolerance for short bursts. extra/nova2_input_storm.py in the repo is the experiment. A run at twice the quota with real input would settle it and costs about $7. If you know the answer, please open an issue.

Do this: set maxTokens to what the answer needs. Size TPM as in-flight requests × (input + maxTokens), and remember the output multiplier for settlement. Read RPM for your Region and your account. Retry 429 with backoff and jitter, and treat 503 and 529 as capacity problems, not as your quota.


5. Cross-Region inference decides where your request runs

This is the local one. There are three ways to name a model on Bedrock, and each makes a different promise (cross-Region inference, global cross-Region inference, re-read 2026-09-30):

Three ways to name a model, three promises

  • In-Region model ID (anthropic.claude-…): runs in the Region you call.
  • Geographic inference profile (us., eu., apac., au., jp.): runs somewhere in that geography; the destination list for a geographic profile does not change.
  • Global inference profile (global.): may run in any commercial Region where the model is offered; the list can change.

Some things don't move. In the documentation's words: "There's no additional routing cost for using cross-Region inference. The price is calculated based on the Region from which you call an inference profile." CloudWatch and CloudTrail log in the source Region. Global is listed as "Approximately 10% savings" against geographic. On the pricing page (read 2026-09-14) that held for Haiku 4.5 in N. Virginia ($1.00 global vs $1.10 geographic). From Singapore or Thailand there's no geographic price for Claude 4.5 and newer at all, because no apac. profile exists for them.

AWS picks the destination. re:Post puts it bluntly: "You can't choose a specific Region to process your request" (re:Post Knowledge Center).

The Thailand fact

ListFoundationModels in ap-southeast-7, one account:

Date Models listed Invocable in-Region on demand Inference-profile only
2026-09-04 21 0 21
2026-09-14 22 0 22
2026-09-21 23 0 23
2026-09-25 28 0 28
2026-09-29 30 0 30
2026-09-30 31 0 31
2026-10-01 31 0 31

The list grew on almost every reading. The in-Region count never moved from zero. The 31st model was OpenAI's GPT-6.1 Sol, announced on 29 September (What's New) and in Bangkok's list the next day. On 1 October each of the 31 had a system-defined inference profile in Bangkok: 27 global. and 4 apac.. So every Bedrock call made from Thailand is cross-Region, whether you planned for that or not.

Global profiles reached Thailand as a source Region for Claude Opus 4.6, Sonnet 4.6 and Haiku 4.5 on 24 February 2026 (AWS ML blog). Singapore, for comparison, listed 40 models on 2026-10-01, of which five are invocable in-Region on demand: Claude 3 Haiku, Claude 3.5 Sonnet, two Cohere embedding models and Claude Sonnet 5. On 4 September it was the first four.

Where the calls went

CloudTrail records the destination of every cross-Region call in additionalEventData.inferenceRegion, as a management event in the source Region, typically 5 to 15 minutes late (about 6 on 1 October). These are all the destinations I saw across my runs of 14, 25, 29 and 30 September and 1 October 2026:

Called from Profile Served in
Bangkok global. Claude Haiku 4.5 Melbourne (every run)
Bangkok global. Claude Sonnet 4.6 Tokyo, Melbourne (25 Sep) · Ohio (29 Sep) · Sydney (30 Sep, 1 Oct)
Bangkok global. Amazon Nova 2 Lite Oregon (every run) · N. Virginia (14 Sep) · Tokyo (14 Sep, 1 Oct)
Bangkok apac. Amazon Nova Lite Tokyo (14, 25, 30 Sep, 1 Oct) · Sydney (14, 29 Sep, 1 Oct)
Singapore global. Claude Sonnet 4.6 London, Ohio, Ireland, Oregon, Tokyo (14 Sep)

Demo 5 as run on 2026-10-01: one raw CloudTrail event. Called in ap-southeast-7, served in ap-southeast-4

The same run: every call of the last three hours, by the Region that served it

Thailand never shows up in the right-hand column. That's what zero in-Region models looks like in practice. The apac. profile kept Nova Lite inside Asia Pacific. The global. profiles went wherever there was capacity, and for one model that meant three continents in a single afternoon.

What "data stays in Region" does and does not mean

The documentation makes three separate statements here, and it's worth keeping them apart:

  1. Stored data stays at the source. "By default, the data remains stored only in the source Region." Logs, knowledge bases and configuration do not move (geographic cross-Region inference).
  2. Prompts and outputs move for processing. "your input prompts and output results might move outside of your source Region during cross-Region inference." On the AWS network, encrypted in transit, never over the public internet.
  3. If anything is retained, it is retained at the destination. "If cross-region inference is enabled for these models, retained inputs and outputs are stored in destination Regions (i.e., the region where your inference request is processed)" (abuse detection, data retention). By default nothing is retained; for Claude Fable 5 and 5.1 all traffic is retained up to 30 days, so a Bangkok caller using Fable through its global profile has a 30-day copy in whatever Region served the call.

If you have a compliance team, show them the table above before they come and ask.

The only control, and the single-Region path

You can't pick or exclude destinations inside a profile. The one documented control is an IAM or SCP condition. A global profile is authorised against the Region-less ARN arn:aws:bedrock:::foundation-model/…, and for that evaluation the service sets aws:RequestedRegion to unspecified. That's why Region-name deny policies "don't target the Region-agnostic global foundation model resource evaluation". To block global routing you deny bedrock:* where aws:RequestedRegion is unspecified and bedrock:InferenceProfileArn matches inference-profile/global.*. To allow it through a Region-deny SCP you add unspecified to the allow list or exempt on the profile ARN. If you use Control Tower, don't hand-edit its managed SCPs (global cross-Region inference). For a geographic profile, blocking any one destination Region in an SCP fails the whole request.

Which models run in-Region differs per Region, and it changes. On 2026-10-01 a Converse call with the bare ID anthropic.claude-sonnet-5 succeeded in Singapore and its CloudTrail event carried no inferenceRegion at all, which is what a call that was not re-routed looks like (AWS ML blog, read 14 Sep). Thailand offered no model that way. The documented single-Region path for the newest Claude models is the bedrock-mantle endpoint with the bare model ID; the Haiku 4.5 model card (read 2026-09-14) lists it in seven Regions, none in Southeast Asia.

One accounting aside from the same card: Anthropic models are "offered and billed through AWS Marketplace. Charges appear on your AWS bill and in AWS Cost Explorer under the model provider (not under Amazon Bedrock)". If you filter Cost Explorer on the Bedrock service, you won't see your Claude spend.

Do this: choose the profile type on purpose and know which destinations it allows. Put inferenceRegion on a dashboard before compliance asks for it. And know that from some Regions, Thailand today, the newest models are global-only.


The pre-agent checklist

The five checks, as they assembled during the talk

These are the five checks I'd run on one plain model call before putting agents, RAG or anything else on top. Each names the script that measures it. The same list lives in CHECKLIST.md in the repo.

  1. Tokens (01_tokens.py): usage.inputTokens measured on my model with my users' language and real text; the per-call overhead known; the metered number in the cost model.
  2. State (02_memory.py): history has a home, an owner and an expiry; every request sends exactly the history I intend and never another user's; the managed-state services I rely on exist in my Region.
  3. Budget (03_context_budget.py): history bounded on purpose; long static prefixes first with a cachePoint that meets the model's minimum, cacheWriteInputTokens checked on the first call; per-user data after the checkpoint.
  4. Quota (04_max_tokens_quota.py, quotas_by_region.py): maxTokens set to what the answer needs; TPM sized as in-flight × (input + maxTokens) plus the settlement multiplier; RPM read for my Region and account; 429 retried with backoff and jitter; 503 and 529 handled as capacity.
  5. Routing (05_where_did_it_run.py): the profile type chosen on purpose; additionalEventData.inferenceRegion on a dashboard and shown to compliance; the destination list of a global profile understood as changeable.

An agent is a loop of these calls. Fix them at one call, not at a thousand.

Reproduce it

git clone https://github.com/spoecker/aws-th-community-day-2026-bedrock-fundamentals
cd aws-th-community-day-2026-bedrock-fundamentals
python3 -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt
python3 demo1.py            # tokens (Singapore)
python3 demo2.py            # stateless, then memory in DynamoDB (Bangkok)
python3 demo3.py            # context budget (Singapore)
python3 demo4.py            # the concurrency storm (Singapore)
python3 demo5.py --fire     # routing calls; 15 minutes later: python3 demo5.py
Enter fullscreen mode Exit fullscreen mode

You need Bedrock model access for Claude Haiku 4.5, Claude Sonnet 5, Nova Lite and Nova 2 Lite in the Region you call from, plus cloudtrail:LookupEvents, servicequotas:ListServiceQuotas and DynamoDB rights on one table. A full run of all five costs well under $1 at list prices. Each script writes a dated JSON file under measurements/, and python3 replay.py <1-5> prints the latest one without calling AWS, so you always have a real, dated result to compare against.


Sources

Every page below was fetched and the quoted text read on it; "14 Sep" means read on 2026-09-14, "30 Sep" means re-read on 2026-09-30.

Tokens. Anthropic glossary (30 Sep) · Anthropic models overview (21 Sep) · Anthropic token counting (14 Sep) · CountTokens API reference (14 Sep) · Count tokens user guide (30 Sep) · TokenUsage (14 Sep) · Claude Haiku 4.5 model card (14 Sep) · Amazon Nova 2 Lite model card (14 Sep) · Amazon Bedrock pricing (14 Sep)

State. Converse user guide (30 Sep) · Thinking block binding (14 Sep) · Anthropic extended thinking (14 Sep) · Session management APIs (30 Sep) · Session management APIs launch (14 Sep) · AgentCore Memory (30 Sep) · AgentCore Regions (30 Sep) · Bedrock Agents Classic maintenance mode (14 Sep) · Data retention (30 Sep) · Abuse detection (14 Sep) · Data protection (30 Sep) · AWS General Reference, Bedrock endpoints and quotas (14 Sep)

Budget. Anthropic context windows (14 Sep) · Converse API reference (14 Sep) · Claude messages parameters (14 Sep) · Prompt caching (30 Sep) · Anthropic prompt caching (14 Sep) · One-hour prompt caching (14 Sep) · Well-Architected Agentic AI Lens, AGENTPERF03-BP02 (14 Sep) · AWS ML blog, "Beyond the price per token" (14 Sep)

Quota. How tokens are counted (burndown) (30 Sep) · Runtime metrics (14 Sep) · InferenceConfiguration (14 Sep) · Runtime quotas (30 Sep) · Bedrock quotas (14 Sep) · Troubleshooting API error codes (30 Sep) · ConverseStream (14 Sep) · SDK retry behavior (30 Sep) · re:Post, Bedrock throttling (14 Sep)

Routing. Cross-Region inference (30 Sep) · Geographic cross-Region inference (14 Sep) · Global cross-Region inference (30 Sep) · Global CRIS in Thailand, Malaysia, Singapore, Indonesia and Taiwan (14 Sep) · re:Post, cross-Region inference routing (14 Sep) · What's New, GPT-6.1 Sol on Amazon Bedrock (1 Oct)

Alexander Spoecker is a Cloud Solution Architect at Iglu in Thailand and an AWS Authorized Instructor. This post is my own work and not a statement by AWS or by Iglu. The scripts are MIT licensed.

Top comments (0)