<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: xbill</title>
    <description>The latest articles on DEV Community by xbill (@xbill).</description>
    <link>https://dev.to/xbill</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3490099%2Fc6a975d0-cd94-485d-82b1-14ed5b344fcf.jpg</url>
      <title>DEV Community: xbill</title>
      <link>https://dev.to/xbill</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/xbill"/>
    <language>en</language>
    <item>
      <title>Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Mon, 28 Sep 2026 19:04:50 +0000</pubDate>
      <link>https://dev.to/gde/count-it-or-compute-it-when-a-tool-returns-rows-the-models-that-count-them-right-spend-the-tokens-2hae</link>
      <guid>https://dev.to/gde/count-it-or-compute-it-when-a-tool-returns-rows-the-models-that-count-them-right-spend-the-tokens-2hae</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This article reports a Kaggle benchmark for a step every agent performs and few people test, counting what a tool returns, and ends with a step by step guide to reproducing it.&lt;/p&gt;

&lt;p&gt;An agent's search, database or API tool usually hands back a list of records. When the user asks how many, the model does the counting. That step passes every quick test with a handful of rows, and nothing flags it when it goes wrong: the model sends the right query and quotes a confident number.&lt;/p&gt;

&lt;p&gt;This benchmark asks ten models the same 68 counting questions with two versions of one tool: &lt;code&gt;count_ids&lt;/code&gt; returns the exact count, &lt;code&gt;list_ids&lt;/code&gt; returns the matching ids.&lt;/p&gt;

&lt;p&gt;With the count, every model but one answered every question about 330 ids correctly. With the ids, the models split in two. The ones that spent 2,700 to 8,200 output tokens per question counted 15 to 21 of 21 lists correctly. The ones that answered in under 600 tokens counted 0 to 10.&lt;/p&gt;

&lt;p&gt;The surprise was who landed where: Claude Opus 5, the most expensive model in the lineup, answered in 579 tokens and counted 9 of 21. Gemma 4 26B, a 26B open-weight model, spent 8,153 and counted 20.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/devto-kaggle" rel="noopener noreferrer"&gt;https://github.com/xbill9/devto-kaggle&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/xbillwork/count-it-or-compute-it" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/xbillwork/count-it-or-compute-it&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  What I Benchmarked
&lt;/h4&gt;

&lt;p&gt;The itch is eleven ids:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0, 2, 3, 20, 21, 22, 23, 10, 11, 12, 13
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;How many are 10 or more? The answer is 8, and every model here gets it. Eleven ids say nothing about three hundred, and eleven is about the size of the lists most quick tests use.&lt;/p&gt;

&lt;p&gt;Language models' trouble with counting is familiar from examples like counting the letters in a word. This benchmark measures it where agents meet it: inside tool use, with every miss traced to the query or the count, and every answer priced in tokens.&lt;/p&gt;

&lt;p&gt;The benchmark holds everything fixed except who does the arithmetic:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;What the model gets&lt;/th&gt;
&lt;th&gt;Who counts&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;count-engine&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;count_ids(where)&lt;/code&gt;, which returns the exact count, minimum and maximum&lt;/td&gt;
&lt;td&gt;The tool&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;count-rows-tool&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;list_ids(where)&lt;/code&gt;, which returns the matching ids&lt;/td&gt;
&lt;td&gt;The model, from the returned list&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;count-python-tool&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Every id in the prompt, plus &lt;code&gt;run_python&lt;/code&gt; with &lt;code&gt;ids&lt;/code&gt; already defined&lt;/td&gt;
&lt;td&gt;The model's code, if it writes any&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In the first two tasks the ids never appear in the prompt, so the only difference is whether the tool returns a number or a list. Both tasks check every filter the model sends by the ids it selects, so &lt;code&gt;id &amp;gt; 9&lt;/code&gt; counts as right for "10 or more", and every wrong answer is traced to either a wrong query or a wrong count.&lt;/p&gt;

&lt;p&gt;Each task asks 68 questions: lists of 11, 110 and 330 ids, seven phrasings of the threshold ("10 or more", "no less than", "under", "between 5 and 9 inclusive" and so on), three seeds each, and the original eleven ids five times. Every threshold is an id in the list, so &lt;code&gt;&amp;gt;&lt;/code&gt; and &lt;code&gt;&amp;gt;=&lt;/code&gt; always give different answers.&lt;/p&gt;




&lt;h4&gt;
  
  
  Models Tested
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Vendor&lt;/th&gt;
&lt;th&gt;Models&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;Gemini 2.5 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, Gemma 4 26B A4B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;GPT-5.4 nano, GPT-5.4 mini, gpt-oss-20b&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The lineup takes a small, a mid-sized and a large model from each vendor, plus the open-weight models from Google and OpenAI, so the comparison covers price, size and what anyone can run. Each model runs with the settings Kaggle's model proxy serves it with. Some of those reason before answering and some answer straight away, and that setting tracks the result more closely than size or price.&lt;/p&gt;

&lt;p&gt;GPT-6 Astra is refused function tools by the proxy (&lt;code&gt;Function tools with reasoning_effort are not supported for gpt-6-astra in /v1/chat/completions&lt;/code&gt;). Gemini 3.5 Flash-Lite and both Qwen 3 Next 80B models returned &lt;code&gt;429&lt;/code&gt; or &lt;code&gt;503&lt;/code&gt; on most calls.&lt;/p&gt;




&lt;h4&gt;
  
  
  Findings
&lt;/h4&gt;

&lt;p&gt;At 330 ids, the size where the models separate:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Rows tool correct&lt;/th&gt;
&lt;th&gt;Output tokens per question, rows tool&lt;/th&gt;
&lt;th&gt;Engine correct&lt;/th&gt;
&lt;th&gt;Rows-tool cost vs engine&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;20/21&lt;/td&gt;
&lt;td&gt;8,153&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;18.6x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;20/21&lt;/td&gt;
&lt;td&gt;2,785&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;13.5x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;2,742&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;14.3x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-20b&lt;/td&gt;
&lt;td&gt;15/21&lt;/td&gt;
&lt;td&gt;2,684&lt;/td&gt;
&lt;td&gt;19/21&lt;/td&gt;
&lt;td&gt;5.2x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;0/21&lt;/td&gt;
&lt;td&gt;580&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;2.6x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;9/21&lt;/td&gt;
&lt;td&gt;579&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;2.4x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;10/21&lt;/td&gt;
&lt;td&gt;401&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;1.9x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;5/21&lt;/td&gt;
&lt;td&gt;143&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;1.5x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;5/21&lt;/td&gt;
&lt;td&gt;39&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;1.9x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;0/21&lt;/td&gt;
&lt;td&gt;35&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;1.7x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Claude rows-tool figures come from their 2026-09-25 run, which asked the same 68 questions with the same tool; their 2026-09-28 rows-tool run stopped on the daily quota. Every other figure comes from the 2026-09-28 runs.&lt;/p&gt;

&lt;h4&gt;
  
  
  1. With the Count in the Tool, Every Model Is Right at a Flat Cost
&lt;/h4&gt;

&lt;p&gt;Every model but one answered all 21 engine questions about 330 ids correctly, at 35 to 451 output tokens per question, and each model's tokens stayed about the same from 11 ids to 330. Turning "no less than 244" into &lt;code&gt;id &amp;gt;= 244&lt;/code&gt; and quoting the number back is something every model here does reliably. The exception, gpt-oss-20b, answered 5 of 68 engine questions with a number other than the count it was given, all of them 0, 1 or 2, and answered once without calling the tool.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. With the Rows, Quick Tests Pass and Larger Lists Fail
&lt;/h4&gt;

&lt;p&gt;Eight of the ten models counted all 26 questions about 11 ids correctly. At 330 ids, five of those eight counted 10 or fewer of 21 correctly. None of them returned an error or a hedge: each answer was a single confident number.&lt;/p&gt;

&lt;h4&gt;
  
  
  3. The Models That Count Right Spend the Tokens
&lt;/h4&gt;

&lt;p&gt;The models split into two groups by how many tokens they spend, and model size and price do not predict the split. That was the surprise. The models that count correctly spend more tokens as the list grows: Gemma 4 26B went from 394 output tokens per question at 11 ids to 8,153 at 330. The models that miscount spend about the same at every size: GPT-5.4 nano spent 35 tokens per question at 11, 110 and 330 ids, which leaves no room to count anything. Claude Opus 5, the most expensive model here, spent 579 output tokens per question at 330 ids and counted 9 of 21 correctly; Gemma 4 26B, an open-weight model, spent 8,153 and counted 20.&lt;/p&gt;

&lt;p&gt;The same models landed in the same group in every rows-tool run, three or four runs per model between 2026-09-25 and 2026-09-28. The order inside a group moves by a few questions between runs.&lt;/p&gt;

&lt;p&gt;Counting right by reasoning costs 5 to 19 times what the engine costs for the same answer. An earlier run with 1,100 ids in the prompt shows the same link from the other side: with output capped at 8,192 tokens, Gemini 3.7 Flash counted 1 of 21 lists correctly, against 19 and 15 of 21 without the cap, and still gave a number every time.&lt;/p&gt;

&lt;h4&gt;
  
  
  4. The Query Was Right Every Time
&lt;/h4&gt;

&lt;p&gt;No answer on either task rested on a filter that selected the wrong ids. Every miss on the rows tool was the model counting the correct list wrong, apart from 3 answers given without calling the tool and 1 that could not be read as a number. The failure is in the arithmetic.&lt;/p&gt;

&lt;h4&gt;
  
  
  5. Python Works When the Model Uses It
&lt;/h4&gt;

&lt;p&gt;With every id in the prompt and a Python tool available, 7 of 10 models scored 68 of 68; Gemini 2.5 Flash called the tool on 1 of 42 questions about 110 and 330 ids and scored 37.&lt;/p&gt;

&lt;h4&gt;
  
  
  What It Changed About How I Think About These Models
&lt;/h4&gt;

&lt;p&gt;In these runs, counting a list was work the model did in tokens, and the models that answered straight away had not done it. A tool that returns rows moves that work onto the model and makes its accuracy depend on a setting the caller may never have looked at. The count belongs in the tool: return the count, the minimum and the maximum, and every model in this lineup answers correctly at a fraction of the tokens.&lt;/p&gt;




&lt;h4&gt;
  
  
  Compare and Contrast
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Engine&lt;/th&gt;
&lt;th&gt;Rows tool&lt;/th&gt;
&lt;th&gt;Python tool&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Who counts&lt;/td&gt;
&lt;td&gt;The tool&lt;/td&gt;
&lt;td&gt;The model, from the returned list&lt;/td&gt;
&lt;td&gt;The model's code, if it writes any&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens per question at 330 ids&lt;/td&gt;
&lt;td&gt;35 to 451&lt;/td&gt;
&lt;td&gt;35 to 8,153&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What went wrong&lt;/td&gt;
&lt;td&gt;Quoting a number other than the count&lt;/td&gt;
&lt;td&gt;Miscounting the right list&lt;/td&gt;
&lt;td&gt;Skipping the tool, no &lt;code&gt;print()&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answers built on a wrong filter&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  So, Which One?
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;🟢 &lt;strong&gt;Engine&lt;/strong&gt; — put the count, minimum and maximum in the tool. Every model is right at a flat cost.&lt;/li&gt;
&lt;li&gt;⚠️ &lt;strong&gt;Python tool&lt;/strong&gt; — reliable when the model uses it, prints the result and quotes it.&lt;/li&gt;
&lt;li&gt;❌ &lt;strong&gt;Rows tool&lt;/strong&gt; — correct only with a model that reasons through the list, at 5 to 19 times the cost per question.&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  What I'd Measure Next
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The same model with reasoning on and off.&lt;/strong&gt; Claude Opus 5 with extended thinking against its default. The runs here show that tokens and accuracy move together across models; this would show whether switching reasoning on is enough to fix one model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Other arithmetic.&lt;/strong&gt; Sums, averages and the largest value over returned rows, which the same tool-shape question applies to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harder filters.&lt;/strong&gt; "At least 10 but under 20" and "outside 5 to 9", where a wrong filter with an exact count would show up.&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  How to Reproduce It
&lt;/h4&gt;

&lt;p&gt;These steps rebuild the benchmark from the repository; the tasks, the runs and every table above come from these commands.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;A Kaggle account, with the Kaggle CLI installed (&lt;code&gt;pip install kaggle&lt;/code&gt;, version 2.2.4 here) and logged in with &lt;code&gt;kaggle auth login&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Python 3 for the local checks and the report; the &lt;code&gt;kaggle_benchmarks&lt;/code&gt; library the tasks import is already installed on Kaggle&lt;/li&gt;
&lt;li&gt;The repository cloned: &lt;code&gt;git clone https://github.com/xbill9/devto-kaggle&lt;/code&gt; and &lt;code&gt;cd devto-kaggle&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Step 1 — Check the Questions Locally
&lt;/h4&gt;

&lt;p&gt;Every question is generated in code, and its expected count comes from running a filter over the ids. &lt;code&gt;check.py&lt;/code&gt; builds all 68 questions with no model calls and no Kaggle library:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 tasks/check.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ok: 68 rows per task, by size {11: 26, 110: 21, 330: 21}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 2 — Push Each Task
&lt;/h4&gt;

&lt;p&gt;Each file in &lt;code&gt;tasks/&lt;/code&gt; is one Kaggle task. The first &lt;code&gt;@kbench.task&lt;/code&gt; in the file returns the share of questions answered correctly, because Kaggle names the task and reads its score from the first one it finds. A second task answers one question and runs once per question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@kbench.task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;count-engine&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;count_engine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;runs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;count_engine_row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;evaluate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;evaluation_data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ROWS&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;n_jobs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;on_failure&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;continue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;summarize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ROWS&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;engine&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each question's result keeps the answer, every filter the model sent, and the tokens and cost Kaggle's model proxy reported for it. The push slug must match the task name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kaggle b t push count-engine &lt;span class="nt"&gt;-f&lt;/span&gt; tasks/count_engine.py &lt;span class="nt"&gt;--wait&lt;/span&gt;
kaggle b t status count-engine
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Task:     count-engine
Version:  13
Status:   Completed
Created:  2026-09-28 16:11:45
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 3 — Run the Models
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;kaggle b auth -y&lt;/code&gt; fetches the short-lived key for Kaggle's model proxy, and &lt;code&gt;kaggle b t models&lt;/code&gt; lists the model slugs. A run takes one &lt;code&gt;-m&lt;/code&gt; per model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kaggle b t run count-rows-tool &lt;span class="nt"&gt;-m&lt;/span&gt; gemini-2.5-flash &lt;span class="nt"&gt;-m&lt;/span&gt; gemini-3.8-flash &lt;span class="nt"&gt;-m&lt;/span&gt; gemma-4-26b-a4b-it &lt;span class="nt"&gt;-m&lt;/span&gt; gpt-oss-20b &lt;span class="nt"&gt;-m&lt;/span&gt; gpt-5.4-nano-2026-03-17 &lt;span class="nt"&gt;-m&lt;/span&gt; gpt-5.4-mini-2026-03-17 &lt;span class="nt"&gt;--wait&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;All runs completed:
  gpt-5.4-mini-2026-03-17: COMPLETED
  gpt-5.4-nano-2026-03-17: COMPLETED
  gpt-oss-20b: COMPLETED
  gemma-4-26b-a4b-it: COMPLETED
  gemini-3.8-flash: COMPLETED
  gemini-2.5-flash: COMPLETED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  🔎 Tip: Cap the Output to Stay Inside the Quota
&lt;/h4&gt;

&lt;p&gt;Kaggle's model proxy reserves the worst-case cost of a call, based on the output-token limit, and refuses the call when that exceeds what is left of the day's quota. With no limit set, GPT-6 Astra reserved $6.40 per call and Claude Opus 5 $3.20. Passing a limit keeps the reservation to cents:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;extra_api_params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_completion_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 4 — Download the Runs and Build the Tables
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;report.py&lt;/code&gt; builds every table in this article from the downloaded runs, including the tokens and cost per question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;t &lt;span class="k"&gt;in &lt;/span&gt;count-engine count-rows-tool count-python-tool&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do &lt;/span&gt;kaggle b t download &lt;span class="nv"&gt;$t&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; results&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
&lt;/span&gt;python3 tasks/report.py results
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;| Model | Size | Engine correct | Engine out tokens | Engine $ | Rows correct | Rows out tokens | Rows $ |
|---|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | 330 | 21/21 | 140 | 0.0010 | 21/21 | 2,742 | 0.0146 |
| GPT-5.4 nano | 330 | 21/21 | 35 | 0.0002 | 0/21 | 35 | 0.0003 |
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 5 — Group the Tasks Into a Benchmark
&lt;/h4&gt;

&lt;p&gt;A benchmark is created in the Kaggle web UI. On the benchmark page, &lt;strong&gt;Add Tasks&lt;/strong&gt; adds the tasks and &lt;strong&gt;Add Models&lt;/strong&gt; the models. Under &lt;strong&gt;Settings&lt;/strong&gt;, set the overall score to &lt;em&gt;Average of task scores&lt;/em&gt;: the default, &lt;em&gt;Percentage of tasks passed&lt;/em&gt;, ignores tasks that return a number. Setting the visibility to &lt;em&gt;Public&lt;/em&gt; is permanent.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: The Leaderboard Shows Each Model's Latest Run
&lt;/h4&gt;

&lt;p&gt;Each leaderboard cell shows the model's most recent run of that task, so a run that fails on the quota replaces a complete one before it. Check the quota before a run, and rerun the pair if one fails. &lt;code&gt;pending.py&lt;/code&gt; lists pairs whose latest run is incomplete.&lt;/p&gt;




&lt;h4&gt;
  
  
  My Benchmark
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.kaggle.com/benchmarks/xbillwork/count-it-or-compute-it" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/xbillwork/count-it-or-compute-it&lt;/a&gt; (the benchmark: the three tasks and their leaderboard)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/xbillwork/count-engine" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/xbillwork/count-engine&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/xbillwork/count-rows-tool" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/xbillwork/count-rows-tool&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/xbillwork/count-python-tool" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/xbillwork/count-python-tool&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to measure whether models count correctly what their tools return. The key to the solution was changing only what the tool returns, a count or a list, and recording every filter and every token.&lt;/p&gt;

&lt;p&gt;The results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 With the count in the tool, every model but gpt-oss-20b answered all 330-id questions correctly.&lt;/li&gt;
&lt;li&gt;⚠️ With the rows, the models that spent 2,700 to 8,200 output tokens per question counted 15 to 21 of 21 lists of 330 correctly.&lt;/li&gt;
&lt;li&gt;❌ The models that answered in under 600 tokens counted 0 to 10 of 21, Claude Opus 5 among them, with no error to show it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each model ran each task on Kaggle's model proxy with its default settings and output capped at 8,192 tokens; the engine and rows-tool figures come from runs on 2026-09-28, apart from the Claude rows-tool figures from 2026-09-25, the Python-tool figures from 2026-09-28, and the token-cap result from earlier in-context runs with lists of 1,100 ids.&lt;/p&gt;

&lt;p&gt;The strategy for benchmarking counting across 10 models was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;This benchmark's code: &lt;a href="https://github.com/xbill9/devto-kaggle" rel="noopener noreferrer"&gt;https://github.com/xbill9/devto-kaggle&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Kaggle Benchmarking Challenge: &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;https://dev.to/challenges/kaggle-2026-09-23&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Kaggle Benchmarks: &lt;a href="https://www.kaggle.com/benchmarks" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The tasks are built on Kaggle's &lt;code&gt;kaggle-benchmarks&lt;/code&gt; library: &lt;a href="https://github.com/Kaggle/kaggle-benchmarks" rel="noopener noreferrer"&gt;https://github.com/Kaggle/kaggle-benchmarks&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Kaggle's benchmark-writing skill, used as the reference for the CLI workflow: &lt;a href="https://github.com/Kaggle/kaggle-skills/blob/main/write-kaggle-benchmarks/SKILL.md" rel="noopener noreferrer"&gt;https://github.com/Kaggle/kaggle-skills/blob/main/write-kaggle-benchmarks/SKILL.md&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Google's QAT Gemma 4 26B-A4B on One TPU v6e: 15.6x the KV Cache and 1.9x the Throughput of FP8</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Sat, 26 Sep 2026 22:35:11 +0000</pubDate>
      <link>https://dev.to/gde/googles-qat-gemma-4-26b-a4b-on-one-tpu-v6e-156x-the-kv-cache-and-19x-the-throughput-of-fp8-36kl</link>
      <guid>https://dev.to/gde/googles-qat-gemma-4-26b-a4b-on-one-tpu-v6e-156x-the-kv-cache-and-19x-the-throughput-of-fp8-36kl</guid>
      <description>&lt;p&gt;This article provides a step by step guide to serving Google's quantization-aware-trained (QAT) Gemma 4 26B-A4B on one Google Cloud TPU v6e chip with vLLM, and compares it with the FP8 build that is the only 26B serving on one chip today. Every per-record output, log and script is committed.&lt;/p&gt;

&lt;p&gt;The QAT 26B serves on one v6e chip at 17.43 GiB of HBM, with 53,888 tokens of KV cache and 1,283 output tokens per second. RedHat's FP8 build uses 27.99 GiB, holds 3,456 tokens and serves 668. Against the full-precision bf16 model, served across four chips as a reference, the QAT checkpoint reads a 3,880-record classification suite 0.4 points lower on an NVIDIA GPU and 1.1 points lower on the TPU; FP8 reads 0.4 points lower. The route is a lossless repack of Google's "unquantized" QAT export and one new method in vLLM's TPU backend; the same repacked checkpoint loads unpatched on vLLM 0.30.0 on an NVIDIA L4.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/gemma4-dev/tree/main/jev-tpu-31b" rel="noopener noreferrer"&gt;https://github.com/xbill9/gemma4-dev/tree/main/jev-tpu-31b&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  Why Measure This?
&lt;/h4&gt;

&lt;p&gt;Gemma 4 26B-A4B is a mixture-of-experts model: 25.8B parameters, 128 experts per layer, 8 active per token. At bf16 it needs 48.07 GiB, and one v6e chip has 28.74 GiB of usable HBM.&lt;/p&gt;

&lt;p&gt;Google trained 4-bit versions of every Gemma 4 size with quantization-aware training, and its model card lists three formats. For 26B-A4B two of them apply:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GGUF (Q4_0)&lt;/strong&gt;, for llama.cpp. vLLM has no GGUF loader.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Unquantized" QAT checkpoints&lt;/strong&gt;, bf16 at 48.07 GiB.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compressed Tensors (w4a16)&lt;/strong&gt;, "for native, optimized inference with vLLM", listed for "Gemma 4 E2B, E4B, 12B, and 31B".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So for vLLM on one chip the only 26B choice is a third-party build. RedHat's FP8 checkpoint fits with 0.75 GiB to spare, enough KV cache for 3,456 tokens, which is one and a half 2,048-token requests.&lt;/p&gt;

&lt;p&gt;This article builds the missing W4A16 checkpoint from Google's own QAT weights and serves it. The result is on Hugging Face as &lt;a href="https://huggingface.co/xbill9/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct" rel="noopener noreferrer"&gt;xbill9/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct&lt;/a&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;A Google Cloud project with v6e quota in a region that offers &lt;code&gt;ct6e-standard-1t&lt;/code&gt;, and the &lt;code&gt;gcloud&lt;/code&gt; CLI logged in&lt;/li&gt;
&lt;li&gt;A Cloud Storage bucket for checkpoints and results&lt;/li&gt;
&lt;li&gt;About 70 GB of local disk and Python 3 with &lt;code&gt;numpy&lt;/code&gt;, for the repack&lt;/li&gt;
&lt;li&gt;Two clones: &lt;code&gt;git clone https://github.com/xbill9/gemma4-dev&lt;/code&gt; and &lt;code&gt;git clone -b gemma4-w4a16-moe https://github.com/xbill9/tpu-inference&lt;/code&gt;, the branch of #3660&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Step 1 — Look Inside the "Unquantized" Export
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;google/gemma-4-26B-A4B-it-qat-q4_0-unquantized&lt;/code&gt; stores bf16 tensors, but the values came out of QAT for Q4_0: every group of 32 weights along the input dimension already sits on a 16-level grid, &lt;code&gt;step × level&lt;/code&gt; with level from −8 to 7. Group size 32 is measured: sampled groups of 64 fail the same test.&lt;/p&gt;

&lt;p&gt;That makes a 4-bit checkpoint a change of container. The one trap is the step. The textbook Q4_0 rule, &lt;code&gt;step = max|w| / 8&lt;/code&gt;, assumes the largest weight in a group sits at level ±8. When a group's peak sits at level 5, that rule derives 5/8 of the true step and re-rounds every weight onto a grid that does not contain it: about 5% median error per group, and every shape check still passes.&lt;/p&gt;

&lt;p&gt;So the repack recovers the step instead: for m from 1 to 8 it tries &lt;code&gt;max|w| / m&lt;/code&gt; and keeps the first that reproduces the whole group, then refines the step by least squares over the 32 values.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 2 — Repack to Compressed-Tensors W4A16
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;huggingface-cli download google/gemma-4-26B-A4B-it-qat-q4_0-unquantized &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--local-dir&lt;/span&gt; ~/models/gemma-4-26B-A4B-it-qat-q4_0-unquantized
python3 repack_q4_0.py repack ~/models/gemma-4-26B-A4B-it-qat-q4_0-unquantized &lt;span class="se"&gt;\&lt;/span&gt;
  ~/models/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct &lt;span class="nt"&gt;--workers&lt;/span&gt; 6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model-layer-000.safetensors: 1186 tensors, 0 kept bf16
model-layer-001.safetensors: 1186 tensors, 0 kept bf16
...
total 15.29 GiB in 31 shards; 0 tensors kept bf16 for off-grid groups; 222 ignored modules
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output is published as &lt;code&gt;xbill9/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct&lt;/code&gt;, so Steps 2 and 3 can be skipped by downloading it. It writes compressed-tensors &lt;code&gt;pack-quantized&lt;/code&gt; symmetric int4 with group size 32, the format of Google's other QAT releases. Attention, the dense MLP and all 3,840 experts are quantized; the router, embeddings, norms and vision tower are copied unchanged. The experts ship fused in the source, &lt;code&gt;experts.gate_up_proj&lt;/code&gt; as &lt;code&gt;[128, 1408, 2816]&lt;/code&gt;, and are written one module per expert, &lt;code&gt;experts.{i}.{gate,up,down}_proj&lt;/code&gt;, which is the layout vLLM already reads for int4 mixture-of-experts checkpoints. A tensor with any group off the grid would stay bf16; none did. It ran in 388 seconds on a 16-core machine with 15 GB of RAM, streaming one layer at a time.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 3 — Verify Every Group
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 repack_q4_0.py verify ~/models/gemma-4-26B-A4B-it-qat-q4_0-unquantized &lt;span class="se"&gt;\&lt;/span&gt;
  ~/models/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;verify&lt;/code&gt; rereads both checkpoints from disk and counts, for every group, whether the stored level is the source value's rank on the stored grid:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tensors&lt;/th&gt;
&lt;th&gt;Groups of 32&lt;/th&gt;
&lt;th&gt;Levels off the grid&lt;/th&gt;
&lt;th&gt;Values bit-identical&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Experts, gate and up&lt;/td&gt;
&lt;td&gt;475,791,360&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;92.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Experts, down&lt;/td&gt;
&lt;td&gt;237,895,680&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;92.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dense MLP&lt;/td&gt;
&lt;td&gt;16,727,040&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;92.0–92.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attention&lt;/td&gt;
&lt;td&gt;34,693,120&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;89.7–90.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All 748 copied tensors are byte-identical. The 7–10% of values that differ in their last bits do so through the scale: Q4_0 carries a 16-bit float step, and the scale is stored at bf16 because that is the precision vLLM's TPU W4A16 layers load it at. The worst relative difference is 1.1e-2. The levels, which carry the QAT training, all match.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — Add W4A16 Experts to vLLM's TPU Backend
&lt;/h4&gt;

&lt;p&gt;Gemma 4 runs only on the JAX path of &lt;code&gt;tpu-inference&lt;/code&gt;, vLLM's TPU backend. Two pieces serve this checkpoint there:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;W4A16 for linear layers&lt;/strong&gt;: &lt;a href="https://github.com/vllm-project/tpu-inference/pull/3653" rel="noopener noreferrer"&gt;tpu-inference #3653&lt;/a&gt;, approved, which serves Google's other QAT sizes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;W4A16 for the experts&lt;/strong&gt;: a &lt;code&gt;WNA16FusedMoEMethod&lt;/code&gt;, &lt;a href="https://github.com/vllm-project/tpu-inference/pull/3660" rel="noopener noreferrer"&gt;tpu-inference #3660&lt;/a&gt;, stacked on #3653.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The experts method loads each expert's packed int4 weights and 32-wide scales, fuses gate and up, and hands them to the GMM kernel unchanged. Three properties of the existing mixture-of-experts path decide how:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The default processing re-quantizes expert weights with one scale per output channel. For QAT weights that replaces the trained grid, so the method lays out the weights itself.&lt;/li&gt;
&lt;li&gt;The fused mixture-of-experts kernel needs quantization blocks that are multiples of 256, so group 32 goes to the GMM backend, whose &lt;code&gt;gmm_v2&lt;/code&gt; kernel takes a scale per 32-wide group.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gmm_v2&lt;/code&gt; keeps activations in bf16 when the group is narrower than the matrix unit, 256 columns on v6e, and dequantizes each weight tile in fast on-chip memory before the multiply. A group of 32 stays weight-only 4-bit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Its unit tests include a forward pass through a real expert layer at the 26B's shape, 2816 hidden and 704 intermediate, against a NumPy reference.&lt;/p&gt;

&lt;p&gt;The same two changes serve the checkpoint across four chips at tensor parallelism 4. There, two layers split their input rows across chips mid-group, 176 of the experts' 704 rows and 528 of the dense MLP's 2112 per chip, so each 32-row scale is stored as two 16-row groups; every weight keeps the same scale. On a v6e-4 the model uses 21.75 GiB across the four chips, holds 407,168 KV tokens and serves 2,073 output tokens per second, and it reads the suite exactly as it does on one chip.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 5 — Launch One v6e Chip
&lt;/h4&gt;

&lt;p&gt;The VM applies the patches to the pinned vLLM TPU image, runs the unit tests on the chip, copies the checkpoint from Cloud Storage, serves it, runs the read, and deletes itself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gcloud compute instances create jev-tpu-31b-moe2 &lt;span class="nt"&gt;--zone&lt;/span&gt; europe-west4-a &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--machine-type&lt;/span&gt; ct6e-standard-1t &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--image-family&lt;/span&gt; ubuntu-accel-2204-amd64-tpu-v5e-v5p-v6e &lt;span class="nt"&gt;--image-project&lt;/span&gt; ubuntu-os-accelerator-images &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--boot-disk-size&lt;/span&gt; 200GB &lt;span class="nt"&gt;--scopes&lt;/span&gt; cloud-platform &lt;span class="nt"&gt;--maintenance-policy&lt;/span&gt; TERMINATE &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--provisioning-model&lt;/span&gt; FLEX_START &lt;span class="nt"&gt;--reservation-affinity&lt;/span&gt; none &lt;span class="nt"&gt;--request-valid-for-duration&lt;/span&gt; 2h &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-run-duration&lt;/span&gt; 4h &lt;span class="nt"&gt;--instance-termination-action&lt;/span&gt; DELETE &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--metadata&lt;/span&gt; jev-code&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;code-tarball&amp;gt;,jev-run&lt;span class="o"&gt;=&lt;/span&gt;2026-09-26-moe2,jev-patches&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"kvshare.diff wna16.diff moe.diff"&lt;/span&gt;,jev-gcs-models&lt;span class="o"&gt;=&lt;/span&gt;gs://&amp;lt;bucket&amp;gt;/jev-tpu-31b/models/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--metadata-from-file&lt;/span&gt; startup-script&lt;span class="o"&gt;=&lt;/span&gt;tpu/startup_quant.sh,jev-arms&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;arms-file&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The server runs with the flags of every other run in the repository, so results pair record for record:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;vllm serve /work/models/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 2048 &lt;span class="nt"&gt;--max-num-seqs&lt;/span&gt; 16 &lt;span class="nt"&gt;--max-logprobs&lt;/span&gt; 32 &lt;span class="nt"&gt;--generation-config&lt;/span&gt; vllm &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--limit-mm-per-prompt&lt;/span&gt; &lt;span class="s1"&gt;'{"image":0,"audio":0,"video":0}'&lt;/span&gt; &lt;span class="nt"&gt;--enable-prefix-caching&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 6 — Serve and Read
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[jev-quant 2026-09-26T13:39:47Z] pytest: 36 passed, 17 warnings in 40.64s
[jev-quant 2026-09-26T13:53:13Z] READY /work/models/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct after 796s
[jev-quant 2026-09-26T13:53:31Z] 26b-q4w4 smoke: ok 20 records, labels returned [6, 6, 5, 5, 6, 2, 2, 2, 2, 2, 1, 2, 3, 3, 4, 2, 2, 2, 2, 2]
[jev-quant 2026-09-26T13:53:55Z] 26b-q4w4 four tasks: 24s for 1200 decisions at concurrency 8
[jev-quant 2026-09-26T13:55:36Z] 26b-q4w4 suite: 58s for 3880 records
[jev-quant 2026-09-26T13:55:56Z] 26b-q4w4 load: 1283.3 tok/s, range 1260.1 to 1288.4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The read is the one used throughout this repository: four 300-example classification tasks, the same three-choice tasks with the options reversed, and Bespoke Labs' 3,880-record public suite, each answer taken as the highest-probability label. Throughput is 16 concurrent requests of exactly 256 output tokens, median of three passes. The first boot compiled for 334 seconds; the saved compile cache removes that on later boots.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Fits on One v6e Chip?
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Memory statistics | total_hbm_limit_gb=31.24GiB | total_hbm_limit_cap_gb=28.74GiB | total_hbm_used_gb=17.43GiB | total_hbm_avail_gb=11.32GiB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;26B-A4B build&lt;/th&gt;
&lt;th&gt;HBM used&lt;/th&gt;
&lt;th&gt;Free&lt;/th&gt;
&lt;th&gt;KV cache&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RedHat FP8&lt;/td&gt;
&lt;td&gt;27.99 GiB&lt;/td&gt;
&lt;td&gt;0.75 GiB&lt;/td&gt;
&lt;td&gt;3,456 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QAT W4A16&lt;/td&gt;
&lt;td&gt;17.43 GiB&lt;/td&gt;
&lt;td&gt;11.32 GiB&lt;/td&gt;
&lt;td&gt;53,888 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 17.43 GiB breaks down, by arithmetic from the checkpoint's tensor headers, into 14.10 GiB of experts, 1.38 GiB of embeddings, 1.07 GiB of vision encoder and 0.86 GiB of attention and dense MLP. The vision encoder is resident because the 26B loads through the multimodal model class even with images switched off. vLLM reports the 53,888 tokens as room for 26 concurrent 2,048-token requests.&lt;/p&gt;




&lt;h4&gt;
  
  
  Does the Accuracy Hold?
&lt;/h4&gt;

&lt;p&gt;The reference is the bf16 model, &lt;code&gt;google/gemma-4-26B-A4B-it&lt;/code&gt;, served with the same flags at tensor parallelism 4 on a v6e-4, where its 48 GiB fits: 61.16 GiB of HBM across four chips and 1,982 output tokens per second. Each build is paired with it on the same records:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Build&lt;/th&gt;
&lt;th&gt;Suite&lt;/th&gt;
&lt;th&gt;Difference from bf16 (95% range)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;bf16, four chips&lt;/td&gt;
&lt;td&gt;76.4%&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RedHat FP8, TPU&lt;/td&gt;
&lt;td&gt;76.0%&lt;/td&gt;
&lt;td&gt;−0.4 points (−0.9 to +0.2)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QAT W4A16, NVIDIA L4&lt;/td&gt;
&lt;td&gt;76.0%&lt;/td&gt;
&lt;td&gt;−0.4 points (−1.2 to +0.3)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QAT W4A16, TPU&lt;/td&gt;
&lt;td&gt;75.3%&lt;/td&gt;
&lt;td&gt;−1.1 points (−1.8 to −0.3)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On the four classification tasks every build is within noise of bf16: every difference's 95% range spans zero. On the suite, the QAT checkpoint on the GPU lands as close to bf16 as FP8 does; the same weights on the TPU lose 0.7 points more. The checkpoint is the same file on both, so that gap belongs to the TPU's int4 path. Paired directly, W4A16 on the TPU reads 0.7 points below FP8 (−1.5 to +0.1).&lt;/p&gt;




&lt;h4&gt;
  
  
  The Same Checkpoint on an NVIDIA GPU
&lt;/h4&gt;

&lt;p&gt;The repacked checkpoint loads on stock vLLM with no patches. On one NVIDIA L4 (&lt;code&gt;g2-standard-8&lt;/code&gt;) with &lt;code&gt;vllm/vllm-openai&lt;/code&gt; at vLLM 0.30.0 and the same flags:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[jev-gpu 2026-09-26T14:11:08Z] READY after 255s
[jev-gpu 2026-09-26T14:11:08Z] memory: Model loading took 14.8 GiB memory and 55.495697 seconds GPU KV cache size: 17,990 tokens, Maximum concurrency for 2,048 tokens per request: 8.78x
[jev-gpu 2026-09-26T14:16:23Z] load: 393.7 tok/s, range 389.5 to 396.3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;vLLM picks its existing Marlin int4 kernels for both the experts and the linear layers. Paired with the TPU run, the four tasks agree within one example each, and the suite reads 0.7 points higher on the L4 (76.0% against 75.3%, 95% range +0.3 to +1.1). The KV cache dtype is ruled out as the cause: a TPU run with &lt;code&gt;--kv-cache-dtype bfloat16&lt;/code&gt; set explicitly allocates the same 53,888 tokens and reads the four tasks record for record, with the suite 0.2 points from the default run. The two platforms run different int4 kernels, &lt;code&gt;gmm_v2&lt;/code&gt; on the TPU and Marlin on the GPU.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: 12B Needs One Flag
&lt;/h4&gt;

&lt;p&gt;Gemma 4 12B ships as &lt;code&gt;Gemma4UnifiedForConditionalGeneration&lt;/code&gt;, which the TPU backend does not register, so vLLM falls back to its PyTorch path. Its decoder is the one &lt;code&gt;Gemma4ForCausalLM&lt;/code&gt; already serves:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;vllm serve google/gemma-4-12B-it &lt;span class="nt"&gt;--hf_overrides&lt;/span&gt; &lt;span class="s1"&gt;'{"architectures": ["Gemma4ForCausalLM"]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That loads bf16 12B on the JAX path at 22.18 GiB against 24.56 GiB on the PyTorch path, and the same flag serves the QAT W4A16 12B at 9.46 GiB with #3653. With the override, vLLM treats the model as text-only and needs no &lt;code&gt;--limit-mm-per-prompt&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Compare and Contrast
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;One v6e chip&lt;/th&gt;
&lt;th&gt;RedHat FP8&lt;/th&gt;
&lt;th&gt;QAT W4A16&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Source weights&lt;/td&gt;
&lt;td&gt;bf16 model, quantized by RedHat&lt;/td&gt;
&lt;td&gt;Google's QAT weights&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HBM used&lt;/td&gt;
&lt;td&gt;27.99 GiB&lt;/td&gt;
&lt;td&gt;🥇 17.43 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache&lt;/td&gt;
&lt;td&gt;3,456 tokens&lt;/td&gt;
&lt;td&gt;🥇 53,888 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens/s&lt;/td&gt;
&lt;td&gt;668&lt;/td&gt;
&lt;td&gt;🥇 1,283&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Suite accuracy&lt;/td&gt;
&lt;td&gt;🥇 76.0%&lt;/td&gt;
&lt;td&gt;75.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flex-start cost per million output tokens&lt;/td&gt;
&lt;td&gt;$0.56&lt;/td&gt;
&lt;td&gt;🥇 $0.29&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runs on stock vLLM TPU today&lt;/td&gt;
&lt;td&gt;🥇 yes&lt;/td&gt;
&lt;td&gt;needs #3653 and #3660&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  So, Which One?
&lt;/h4&gt;

&lt;p&gt;For serving on one v6e chip, the QAT W4A16 build: it holds 15.6 times the KV cache, serves 1.92 times the tokens per second, and reads the suite 0.7 points below FP8. On a GPU the same checkpoint reads as close to bf16 as FP8 does. FP8 remains the build that runs on a stock vLLM TPU image today. On NVIDIA GPUs the same QAT checkpoint needs nothing beyond stock vLLM.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Does It Cost?
&lt;/h4&gt;

&lt;p&gt;At europe-west4's v6e rates of $1.35 per chip-hour on flex-start and $2.97 on demand, the measured throughputs work out to, by arithmetic:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Build&lt;/th&gt;
&lt;th&gt;Flex-start&lt;/th&gt;
&lt;th&gt;On demand&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;QAT W4A16, 1,283 tok/s&lt;/td&gt;
&lt;td&gt;$0.29 per million output tokens&lt;/td&gt;
&lt;td&gt;$0.64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FP8, 668 tok/s&lt;/td&gt;
&lt;td&gt;$0.56&lt;/td&gt;
&lt;td&gt;$1.23&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These assume the chip runs at the measured load around the clock, 16 concurrent requests of 256 tokens.&lt;/p&gt;




&lt;h4&gt;
  
  
  Teardown
&lt;/h4&gt;

&lt;p&gt;Each VM deletes itself when its run ends, and &lt;code&gt;--max-run-duration&lt;/code&gt; deletes it regardless. Confirm nothing is left:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gcloud compute instances list &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'value(name,zone.basename(),status)'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The checkpoint stays in Cloud Storage for the next boot; delete it with &lt;code&gt;gcloud storage rm -r gs://&amp;lt;bucket&amp;gt;/jev-tpu-31b/models/&lt;/code&gt; when done.&lt;/p&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to serve Google's QAT Gemma 4 26B-A4B on one TPU v6e chip with vLLM. The key to the solution was a lossless repack of Google's bf16 QAT export into compressed-tensors W4A16, and a W4A16 experts method for vLLM's TPU backend that keeps the checkpoint's 32-wide scales. The results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 The QAT 26B serves on one v6e chip at 17.43 GiB, against 27.99 GiB for RedHat's FP8 build&lt;/li&gt;
&lt;li&gt;🟢 53,888 tokens of KV cache against 3,456, and 1,283 output tokens per second against 668&lt;/li&gt;
&lt;li&gt;🟢 Every one of 765 million weight groups repacks onto its 4-bit grid, and all 748 unquantized tensors copy byte for byte&lt;/li&gt;
&lt;li&gt;⚠️ Against bf16 the suite reads 1.1 points lower on the TPU (−1.8 to −0.3) and 0.4 lower on an NVIDIA L4 (−1.2 to +0.3); FP8 reads 0.4 lower&lt;/li&gt;
&lt;li&gt;🟢 The repacked checkpoint loads unpatched on vLLM 0.30.0 on an NVIDIA L4&lt;/li&gt;
&lt;li&gt;⚠️ On TPU it needs #3653 and #3660, neither merged yet&lt;/li&gt;
&lt;li&gt;🟢 Across four chips (tensor parallelism 4) it serves at 21.75 GiB with 407,168 KV tokens and reads the suite as it does on one chip&lt;/li&gt;
&lt;li&gt;🟢 Gemma 4 12B serves on the TPU backend's JAX path with one &lt;code&gt;--hf_overrides&lt;/code&gt; flag&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: the bf16 reference and the four-chip QAT run used a v6e-4 (&lt;code&gt;ct6e-standard-4t&lt;/code&gt;) at tensor parallelism 4, every other TPU arm on one TPU v6e chip (&lt;code&gt;ct6e-standard-1t&lt;/code&gt;, flex-start) in europe-west4-a, vLLM &lt;code&gt;0.29.1rc1.dev468+g0b7f11a1e&lt;/code&gt; at &lt;code&gt;vllm/vllm-tpu@sha256:19a1a052…&lt;/code&gt; with #3299, #3653 and #3660 applied, one run per build, &lt;code&gt;--max-model-len 2048&lt;/code&gt;, vLLM's default KV cache dtype. The FP8 read comes from a run two days earlier on the same image without the patches, and its throughput from a second VM the same day; the four-task reads repeat record for record across VMs here, the suite within 0.1 points, and throughput moved about 2% between VMs. The GPU run used one NVIDIA L4 in us-central1-a with vLLM 0.30.0. Costs are arithmetic from list prices and the measured throughput. The repacked checkpoint is unofficial and derived from Google's release under Apache 2.0. Parts of the analysis and writing were done with AI assistance (Claude); every figure comes from the committed output files.&lt;/p&gt;

&lt;p&gt;The strategy for serving Google's QAT Gemma 4 26B on one TPU v6e chip was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;The repacked checkpoint: &lt;a href="https://huggingface.co/xbill9/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct" rel="noopener noreferrer"&gt;https://huggingface.co/xbill9/gemma-4-26B-A4B-it-qat-q4_0-w4a16-ct&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Code, repack, per-record results and logs: &lt;a href="https://github.com/xbill9/gemma4-dev/tree/main/jev-tpu-31b" rel="noopener noreferrer"&gt;https://github.com/xbill9/gemma4-dev/tree/main/jev-tpu-31b&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;tpu-inference #3660, W4A16 experts method: &lt;a href="https://github.com/vllm-project/tpu-inference/pull/3660" rel="noopener noreferrer"&gt;https://github.com/vllm-project/tpu-inference/pull/3660&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;tpu-inference #3653, W4A16 linear method: &lt;a href="https://github.com/vllm-project/tpu-inference/pull/3653" rel="noopener noreferrer"&gt;https://github.com/vllm-project/tpu-inference/pull/3653&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Google's QAT source checkpoint: &lt;a href="https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-unquantized" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-unquantized&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;RedHat FP8 build: &lt;a href="https://huggingface.co/RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic" rel="noopener noreferrer"&gt;https://huggingface.co/RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Google's QAT announcement: &lt;a href="https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/" rel="noopener noreferrer"&gt;https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Bespoke Labs public suite: &lt;a href="https://github.com/bespokelabsai/nimble/blob/0e67403/docs/PUBLIC_BENCHMARKS.md" rel="noopener noreferrer"&gt;https://github.com/bespokelabsai/nimble/blob/0e67403/docs/PUBLIC_BENCHMARKS.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;vLLM TPU documentation: &lt;a href="https://docs.vllm.ai/projects/tpu/en/latest/" rel="noopener noreferrer"&gt;https://docs.vllm.ai/projects/tpu/en/latest/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Cloud TPU v6e: &lt;a href="https://cloud.google.com/tpu/docs/v6e" rel="noopener noreferrer"&gt;https://cloud.google.com/tpu/docs/v6e&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>gemma</category>
      <category>googlecloud</category>
      <category>machinelearning</category>
      <category>llm</category>
    </item>
    <item>
      <title>Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:10:27 +0000</pubDate>
      <link>https://dev.to/gde/gemma-4-on-amazon-sagemaker-qat-weights-decode-205x-faster-than-bf16-on-one-l4-2m9g</link>
      <guid>https://dev.to/gde/gemma-4-on-amazon-sagemaker-qat-weights-decode-205x-faster-than-bf16-on-one-l4-2m9g</guid>
      <description>&lt;p&gt;This article gives a short background on Amazon SageMaker real-time endpoints, then measures Gemma 4 E2B's quantization-aware trained (QAT) checkpoint against the full-size bf16 release on the same NVIDIA L4 endpoint. A suite of Python MCP tools is built to simplify management of the vLLM hosted deployment.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/sagemaker-gemma" rel="noopener noreferrer"&gt;https://github.com/xbill9/sagemaker-gemma&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Models&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;google/gemma-4-E2B-it&lt;/code&gt; (bf16) and &lt;code&gt;google/gemma-4-E2B-it-qat-w4a16-ct&lt;/code&gt; (QAT, 4-bit weights)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hardware&lt;/td&gt;
&lt;td&gt;SageMaker &lt;code&gt;ml.g6.xlarge&lt;/code&gt;, 1x NVIDIA L4, 24 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Region&lt;/td&gt;
&lt;td&gt;&lt;code&gt;us-east-2&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Software&lt;/td&gt;
&lt;td&gt;AWS vLLM SageMaker container, vLLM 0.30.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Result&lt;/td&gt;
&lt;td&gt;QAT decodes at &lt;strong&gt;105.1 tok/s&lt;/strong&gt; against &lt;strong&gt;51.3&lt;/strong&gt;, and serves &lt;strong&gt;1077.25 tok/s&lt;/strong&gt; at 16 parallel requests against &lt;strong&gt;619.1&lt;/strong&gt;, with the same score on 40 checked questions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  SageMaker Hosting in Five Minutes
&lt;/h4&gt;

&lt;p&gt;SageMaker real-time inference is three objects, created in order:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Object&lt;/th&gt;
&lt;th&gt;What it holds&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;A container image, its environment variables, and an IAM role&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint config&lt;/td&gt;
&lt;td&gt;Which model runs on which instance types, and how many instances&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint&lt;/td&gt;
&lt;td&gt;The running HTTPS service, billed per instance-hour while it exists&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Requests go through &lt;code&gt;aws sagemaker-runtime invoke-endpoint&lt;/code&gt;, signed with your AWS credentials. SageMaker health-checks the container, routes traffic to it and writes its log to CloudWatch under &lt;code&gt;/aws/sagemaker/Endpoints/&amp;lt;name&amp;gt;&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The container here is the vLLM build AWS publishes for SageMaker. It reads vLLM settings from &lt;code&gt;SM_VLLM_&lt;/code&gt; environment variables, downloads the model from Hugging Face at start-up, and accepts OpenAI-style chat bodies. Switching checkpoints is one variable: &lt;code&gt;SM_VLLM_MODEL&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;SageMaker JumpStart also lists Gemma 4 as ready-made packages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker list-hub-contents &lt;span class="nt"&gt;--hub-name&lt;/span&gt; SageMakerPublicHub &lt;span class="nt"&gt;--hub-content-type&lt;/span&gt; Model &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"HubContentSummaries[].HubContentName"&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="s1"&gt;'\t'&lt;/span&gt; &lt;span class="s1"&gt;'\n'&lt;/span&gt; | &lt;span class="nb"&gt;grep &lt;/span&gt;gemma-4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;huggingface-llm-gemma-4-31b-it-nvfp4
huggingface-vlm-gemma-4-12b-it
huggingface-vlm-gemma-4-26b-a4b-it
huggingface-vlm-gemma-4-31b-it
huggingface-vlm-gemma-4-31b-it-fp8-block
huggingface-vlm-gemma-4-e2b-instruct
huggingface-vlm-gemma-4-e4b-it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;None of the seven is a Google QAT checkpoint, so this comparison uses the vLLM container with a Hugging Face model ID.&lt;/p&gt;

&lt;p&gt;Three account limits shape every deployment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quota&lt;/strong&gt; is per instance type, per region, counted in instances. This account holds 1 for each single-L4 type in &lt;code&gt;us-east-1&lt;/code&gt;, &lt;code&gt;us-east-2&lt;/code&gt; and &lt;code&gt;us-west-2&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity&lt;/strong&gt; is separate. &lt;code&gt;us-east-1&lt;/code&gt; and &lt;code&gt;us-west-2&lt;/code&gt; each left an L4 request in &lt;code&gt;Creating&lt;/code&gt; for about 30 minutes and then returned &lt;code&gt;InsufficientInstanceCapacity&lt;/code&gt;. &lt;code&gt;us-east-2&lt;/code&gt; placed one within minutes every time it was asked.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A fallback list&lt;/strong&gt; (&lt;code&gt;InstancePools&lt;/code&gt; in the endpoint config) holds quota for every type in it, so two endpoints cannot share one region's single-L4 quota.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With one L4 per region to work with, the two checkpoints ran one after the other on the same instance type in the same region.&lt;/p&gt;




&lt;h4&gt;
  
  
  Where Do I Start?
&lt;/h4&gt;

&lt;p&gt;The deployment itself, from quota check to teardown with the aws CLI and the MCP server, is the first article in this pair: &lt;a href="https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d"&gt;https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This article starts from a working endpoint and changes one thing: the checkpoint.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;The repository above, with &lt;code&gt;aws login&lt;/code&gt; done and &lt;code&gt;mcp&lt;/code&gt; 2.x installed&lt;/li&gt;
&lt;li&gt;A SageMaker quota of at least 1 for &lt;code&gt;ml.g6.xlarge&lt;/code&gt; in a region with L4 capacity&lt;/li&gt;
&lt;li&gt;The full-size endpoint from the first article deployed as &lt;code&gt;gemma-4-e2b&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  What QAT Changes
&lt;/h4&gt;

&lt;p&gt;Google trains the QAT checkpoint with 4-bit weights in the loop, then exports it in the &lt;code&gt;compressed-tensors&lt;/code&gt; format vLLM reads natively. The &lt;code&gt;-w4a16-ct&lt;/code&gt; suffix means 4-bit weights, 16-bit activations. Google publishes the same model in four QAT forms; only this one loads in vLLM:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Checkpoint&lt;/th&gt;
&lt;th&gt;vLLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-qat-w4a16-ct&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;🟢 loads, 4-bit weights&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-qat-q4_0-unquantized&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;⚠️ stored at 16-bit, no memory saving&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-qat-q4_0-gguf&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;❌ GGUF, for llama.cpp and Ollama&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-qat-mobile-*&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;❌ on-device formats&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Deploying it is the first article's Steps 4 to 6 with two variables changed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gemma-4-e2b-qat
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MODEL_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;google/gemma-4-E2B-it-qat-w4a16-ct
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;vLLM names the format when it starts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;quantization=compressed-tensors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  What the Engine Allocates
&lt;/h4&gt;

&lt;p&gt;The container log records the weights, the load time and the KV cache vLLM builds from what is left of the L4's memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model loading took 9.75 GiB memory and 82.751806 seconds
GPU KV cache size: 723,484 tokens, Maximum concurrency for 8,192 tokens per request: 88.32x
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model loading took 8.01 GiB memory and 66.190265 seconds
GPU KV cache size: 867,999 tokens, Maximum concurrency for 8,192 tokens per request: 105.96x
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;th&gt;QAT / bf16&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weights (GiB)&lt;/td&gt;
&lt;td&gt;9.75&lt;/td&gt;
&lt;td&gt;8.01&lt;/td&gt;
&lt;td&gt;0.82&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache (tokens)&lt;/td&gt;
&lt;td&gt;723,484&lt;/td&gt;
&lt;td&gt;867,999&lt;/td&gt;
&lt;td&gt;1.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weight load (s)&lt;/td&gt;
&lt;td&gt;82.75&lt;/td&gt;
&lt;td&gt;66.19&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Create to &lt;code&gt;InService&lt;/code&gt; (min)&lt;/td&gt;
&lt;td&gt;9.9&lt;/td&gt;
&lt;td&gt;10.1&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 4-bit weights save 18% of GPU memory. The checkpoint's own header shows why: only the transformer body is 4-bit.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Part of the QAT file&lt;/th&gt;
&lt;th&gt;GB&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;th&gt;Stored as&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Per-layer embedding&lt;/td&gt;
&lt;td&gt;4.698&lt;/td&gt;
&lt;td&gt;56.5%&lt;/td&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vocabulary embedding&lt;/td&gt;
&lt;td&gt;1.611&lt;/td&gt;
&lt;td&gt;19.4%&lt;/td&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transformer body&lt;/td&gt;
&lt;td&gt;1.056&lt;/td&gt;
&lt;td&gt;12.7%&lt;/td&gt;
&lt;td&gt;packed 4-bit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audio tower&lt;/td&gt;
&lt;td&gt;0.614&lt;/td&gt;
&lt;td&gt;7.4%&lt;/td&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision tower&lt;/td&gt;
&lt;td&gt;0.337&lt;/td&gt;
&lt;td&gt;4.1%&lt;/td&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The memory the smaller weights free goes to the KV cache.&lt;/p&gt;




&lt;h4&gt;
  
  
  How the Measurement Works
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;compare.py&lt;/code&gt; runs the same three measurements against each endpoint, at temperature 0, through the same &lt;code&gt;aws sagemaker-runtime invoke-endpoint&lt;/code&gt; call:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Decode.&lt;/strong&gt; Fixed-length replies of 16 and 512 tokens (&lt;code&gt;ignore_eos&lt;/code&gt;), five of each. The decode rate is (512 − 16) / (median time at 512 − median time at 16), which cancels the aws CLI start-up and the network round trip.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parallel requests.&lt;/strong&gt; 1, 4 and 16 requests at once, 256 tokens each, two batches per level. Throughput is total output tokens divided by the batch's wall time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answers.&lt;/strong&gt; 40 fixed questions with exact answers: 15 two-digit multiplications, 15 three-number sums, 10 capitals. Scored by regular expression.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 compare.py measure docs/runs/2026-09-25-qat-vs-bf16 gemma-4-e2b-qat@us-east-2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;wrote docs/runs/2026-09-25-qat-vs-bf16/measure-gemma-4-e2b-qat.json
{
  "decode_tokens_per_second": 105.1,
  "per_call_fixed_cost_seconds": 0.562,
  "load_tokens_per_second": {
    "1": 85.35,
    "4": 328.7,
    "16": 1077.25
  },
  "quality": "37/40"
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;combine&lt;/code&gt; computes every ratio from the two result files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 compare.py combine compare.json measure-gemma-4-e2b.json measure-gemma-4-e2b-qat.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                               gemma-4-e2b     gemma-4-e2b-qat  ratio
weights_gib                           9.75                8.01  0.82
kv_cache_tokens                     723484              867999  1.2
load_seconds                     82.751806           66.190265  
decode_tokens_per_second              51.3               105.1  2.05
load_c1_tokens_per_second             45.6               85.35  1.87
load_c4_tokens_per_second            171.6               328.7  1.92
load_c16_tokens_per_second           619.1             1077.25  1.74
quality_correct                         37                  37  
identical answers: 35/40
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Decode Speed
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Decode (tokens/s)&lt;/td&gt;
&lt;td&gt;51.3&lt;/td&gt;
&lt;td&gt;🥇 105.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;512-token reply, fastest (s)&lt;/td&gt;
&lt;td&gt;10.491&lt;/td&gt;
&lt;td&gt;🥇 5.392&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;512-token reply, slowest (s)&lt;/td&gt;
&lt;td&gt;10.661&lt;/td&gt;
&lt;td&gt;🥇 5.526&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-call client cost (s)&lt;/td&gt;
&lt;td&gt;0.633&lt;/td&gt;
&lt;td&gt;0.562&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;QAT decodes 2.05x faster.&lt;/strong&gt; The slowest QAT reply finished in about half the time of the fastest bf16 one, so the gap is far larger than the spread between repeats.&lt;/p&gt;

&lt;p&gt;Each decode step reads every transformer layer's weights from GPU memory, so decode speed follows how many bytes those layers take. Those are the layers the QAT export stores at 4 bits. The embeddings are looked up one row per token, so their 16-bit size costs memory and little time.&lt;/p&gt;




&lt;h4&gt;
  
  
  Parallel Requests
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests at once&lt;/th&gt;
&lt;th&gt;bf16 tok/s&lt;/th&gt;
&lt;th&gt;QAT tok/s&lt;/th&gt;
&lt;th&gt;QAT / bf16&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;45.6&lt;/td&gt;
&lt;td&gt;🥇 85.35&lt;/td&gt;
&lt;td&gt;1.87&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;171.6&lt;/td&gt;
&lt;td&gt;🥇 328.7&lt;/td&gt;
&lt;td&gt;1.92&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;619.1&lt;/td&gt;
&lt;td&gt;🥇 1077.25&lt;/td&gt;
&lt;td&gt;1.74&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;QAT leads at every level. The ratio narrows at 16, where a batch shares each weight read across more requests and the per-token saving counts for less. The single-request figures sit below the decode rate because each call also pays the aws CLI start-up.&lt;/p&gt;




&lt;h4&gt;
  
  
  Does It Still Answer Correctly?
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Questions&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Multiplication (15)&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a + b − c (15)&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capitals (10)&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total (40)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;37&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;37&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;35 of the 40 answers are identical character for character. The other five are all sums, and each model misses three of them:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Expected&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;876 + 608 − 558&lt;/td&gt;
&lt;td&gt;926&lt;/td&gt;
&lt;td&gt;❌ 1026&lt;/td&gt;
&lt;td&gt;✅ 926&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;257 + 388 − 290&lt;/td&gt;
&lt;td&gt;355&lt;/td&gt;
&lt;td&gt;✅ 355&lt;/td&gt;
&lt;td&gt;❌ 655&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;782 + 571 − 534&lt;/td&gt;
&lt;td&gt;819&lt;/td&gt;
&lt;td&gt;❌ 829&lt;/td&gt;
&lt;td&gt;✅ 819&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;470 + 238 − 983&lt;/td&gt;
&lt;td&gt;−275&lt;/td&gt;
&lt;td&gt;✅ −275&lt;/td&gt;
&lt;td&gt;❌ −285&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;150 + 937 − 139&lt;/td&gt;
&lt;td&gt;948&lt;/td&gt;
&lt;td&gt;❌ 918&lt;/td&gt;
&lt;td&gt;❌ 1010&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two models make the same number of mistakes on different questions. Forty questions are enough to show a large loss and too few to measure a small one.&lt;/p&gt;




&lt;h4&gt;
  
  
  Re-Measured in A-B-A Order
&lt;/h4&gt;

&lt;p&gt;Each endpoint runs on its own physical instance, and the two ran 15 minutes apart. To check that the gap belongs to the checkpoint, the bf16 endpoint was deployed a second time after QAT and measured again:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Start (UTC)&lt;/th&gt;
&lt;th&gt;Decode tok/s&lt;/th&gt;
&lt;th&gt;16 at once tok/s&lt;/th&gt;
&lt;th&gt;Correct&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;td&gt;18:36&lt;/td&gt;
&lt;td&gt;51.3&lt;/td&gt;
&lt;td&gt;619.1&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QAT&lt;/td&gt;
&lt;td&gt;18:50&lt;/td&gt;
&lt;td&gt;105.1&lt;/td&gt;
&lt;td&gt;1077.25&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;bf16 again&lt;/td&gt;
&lt;td&gt;19:10&lt;/td&gt;
&lt;td&gt;51.5&lt;/td&gt;
&lt;td&gt;625.15&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two bf16 runs agree within 1.21% on every speed figure and give identical answers to all 40 questions. The second bf16 instance loaded its weights in 82.09 s against 82.75 s the first time.&lt;/p&gt;




&lt;h4&gt;
  
  
  Compare to Other Deployments
&lt;/h4&gt;

&lt;p&gt;The same bf16 model on the same GPU has been measured on two other platforms in this series, with &lt;code&gt;vllm bench serve&lt;/code&gt; at 128 output tokens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;1 request tok/s&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cloud Run, NVIDIA L4&lt;/td&gt;
&lt;td&gt;49.63&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EC2 &lt;code&gt;g6.2xlarge&lt;/code&gt;, NVIDIA L4&lt;/td&gt;
&lt;td&gt;46.09&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SageMaker &lt;code&gt;ml.g6.xlarge&lt;/code&gt;, NVIDIA L4&lt;/td&gt;
&lt;td&gt;45.6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The three sit within 9% of each other. SageMaker's figure includes the aws CLI start-up in every call; its decode rate with that removed is 51.3.&lt;/p&gt;

&lt;p&gt;On a Tesla T4, QAT decoded 1.79x faster than bf16 for the same model. On the L4 the ratio is 2.05x.&lt;/p&gt;

&lt;p&gt;The methods differ: the other runs used &lt;code&gt;vllm bench serve&lt;/code&gt; with random prompts, other vLLM versions and their own hosts. Read the rows as a shape.&lt;/p&gt;




&lt;h4&gt;
  
  
  And Price/Performance?
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;ml.g6.xlarge&lt;/code&gt; in &lt;code&gt;us-east-2&lt;/code&gt; is $1.1267 an hour on demand:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws pricing get-products &lt;span class="nt"&gt;--region&lt;/span&gt; us-east-1 &lt;span class="nt"&gt;--service-code&lt;/span&gt; AmazonSageMaker &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filters&lt;/span&gt; &lt;span class="nv"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;TERM_MATCH,Field&lt;span class="o"&gt;=&lt;/span&gt;instanceName,Value&lt;span class="o"&gt;=&lt;/span&gt;ml.g6.xlarge &lt;span class="nv"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;TERM_MATCH,Field&lt;span class="o"&gt;=&lt;/span&gt;regionCode,Value&lt;span class="o"&gt;=&lt;/span&gt;us-east-2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ml.g6.xlarge USE2-Host:ml.g6.xlarge 1.1267000000 Hrs | $1.1267 per Hosting ml.g6.xlarge hour in US East (Ohio)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same hourly price buys twice the tokens. Per million output tokens (arithmetic):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests at once&lt;/th&gt;
&lt;th&gt;bf16 $/M&lt;/th&gt;
&lt;th&gt;QAT $/M&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;6.86&lt;/td&gt;
&lt;td&gt;🥇 3.67&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;1.82&lt;/td&gt;
&lt;td&gt;🥇 0.95&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;0.51&lt;/td&gt;
&lt;td&gt;🥇 0.29&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The endpoint bills while it exists, busy or idle, so these figures hold only while it is kept busy.&lt;/p&gt;




&lt;h4&gt;
  
  
  So, Which One?
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Decode speed&lt;/td&gt;
&lt;td&gt;51.3 tok/s&lt;/td&gt;
&lt;td&gt;🥇 105.1 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16 requests at once&lt;/td&gt;
&lt;td&gt;619.1 tok/s&lt;/td&gt;
&lt;td&gt;🥇 1077.25 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU memory for weights&lt;/td&gt;
&lt;td&gt;9.75 GiB&lt;/td&gt;
&lt;td&gt;🥇 8.01 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checked answers&lt;/td&gt;
&lt;td&gt;37 / 40&lt;/td&gt;
&lt;td&gt;37 / 40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per million tokens at 16&lt;/td&gt;
&lt;td&gt;$0.51&lt;/td&gt;
&lt;td&gt;🥇 $0.29&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On an L4 SageMaker endpoint, the QAT checkpoint is the default choice for Gemma 4 E2B: the same instance, one changed environment variable, about twice the tokens per dollar, and no measured change in answers. Keep bf16 as the reference when a task's accuracy needs a larger evaluation than 40 questions.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Stops the Meter
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;delete_endpoint&lt;/code&gt; removes the endpoint, its config and its model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;🗑️ `gemma-4-e2b-qat` in `us-east-2`
- endpoint: deleted
- endpoint-config: deleted
- model: deleted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;Creating&lt;/code&gt; endpoint refuses deletion, so a deploy runs until it reaches &lt;code&gt;InService&lt;/code&gt; or &lt;code&gt;Failed&lt;/code&gt; before it can be removed.&lt;/p&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to measure what Gemma 4 E2B's QAT checkpoint changes on a SageMaker L4 endpoint. The key to the solution was changing only &lt;code&gt;SM_VLLM_MODEL&lt;/code&gt; between two deployments on the same instance type, and measuring decode speed with the per-call client cost removed. The measured results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 QAT decodes at &lt;strong&gt;105.1 tok/s&lt;/strong&gt; against bf16's &lt;strong&gt;51.3&lt;/strong&gt;, 2.05x&lt;/li&gt;
&lt;li&gt;🟢 At 16 requests at once QAT serves &lt;strong&gt;1077.25 tok/s&lt;/strong&gt; against &lt;strong&gt;619.1&lt;/strong&gt;, 1.74x&lt;/li&gt;
&lt;li&gt;🟢 Both score &lt;strong&gt;37 of 40&lt;/strong&gt; on checked questions, with 35 identical answers&lt;/li&gt;
&lt;li&gt;🟢 A second bf16 deployment after QAT reproduced the first within 1.21%&lt;/li&gt;
&lt;li&gt;⚠️ The 4-bit export saves 18% of weight memory, because the embeddings stay at 16-bit&lt;/li&gt;
&lt;li&gt;⚠️ Getting an L4 took three regions: &lt;code&gt;us-east-1&lt;/code&gt; and &lt;code&gt;us-west-2&lt;/code&gt; returned &lt;code&gt;InsufficientInstanceCapacity&lt;/code&gt; after about 30 minutes each&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: one account, SageMaker &lt;code&gt;ml.g6.xlarge&lt;/code&gt; with one NVIDIA L4 in &lt;code&gt;us-east-2&lt;/code&gt;, vLLM 0.30.0 from the AWS container, three deployments on 2026-09-25 each on its own instance, measured in bf16, QAT, bf16 order. Prompts were short; long-prompt behaviour was not measured. Decode used five replies per length, parallel throughput two batches per level, and quality 40 questions at temperature 0. Every request went through the aws CLI from one client machine.&lt;/p&gt;

&lt;p&gt;The strategy for using MCP for SageMaker deployment and benchmarking was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Repository: &lt;a href="https://github.com/xbill9/sagemaker-gemma" rel="noopener noreferrer"&gt;https://github.com/xbill9/sagemaker-gemma&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Part one, deploying Gemma 4 to SageMaker: &lt;a href="https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d"&gt;https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 E2B QAT w4a16: &lt;a href="https://huggingface.co/google/gemma-4-E2B-it-qat-w4a16-ct" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-E2B-it-qat-w4a16-ct&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 E2B: &lt;a href="https://huggingface.co/google/gemma-4-E2B-it" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-E2B-it&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 on a Tesla T4, QAT vs bf16: &lt;a href="https://dev.to/gde/gemma-4-on-a-tesla-t4-qat-weights-decode-179x-faster-than-bf16-2fi4"&gt;https://dev.to/gde/gemma-4-on-a-tesla-t4-qat-weights-decode-179x-faster-than-bf16-2fi4&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;2B Gemma 4 on Cloud Run with an NVIDIA L4: &lt;a href="https://dev.to/gde/2b-gemma-4-deployment-with-cloud-run-nvidia-l4-mcp-sdk-2x-and-claude-code-4ml3"&gt;https://dev.to/gde/2b-gemma-4-deployment-with-cloud-run-nvidia-l4-mcp-sdk-2x-and-claude-code-4ml3&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;SageMaker real-time inference: &lt;a href="https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;AWS Deep Learning Containers: &lt;a href="https://github.com/aws/deep-learning-containers" rel="noopener noreferrer"&gt;https://github.com/aws/deep-learning-containers&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;compressed-tensors: &lt;a href="https://github.com/neuralmagic/compressed-tensors" rel="noopener noreferrer"&gt;https://github.com/neuralmagic/compressed-tensors&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aws</category>
      <category>sagemaker</category>
      <category>gemma</category>
      <category>vllm</category>
    </item>
    <item>
      <title>Gemma 4 on an Amazon SageMaker Endpoint: AWS CLI, NVIDIA L4, and an MCP Server</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:10:24 +0000</pubDate>
      <link>https://dev.to/gde/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-5cdd</link>
      <guid>https://dev.to/gde/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-5cdd</guid>
      <description>&lt;p&gt;This article provides a step by step deployment guide for Gemma 4 E2B to an Amazon SageMaker hosted GPU enabled system. A suite of Python MCP tools is built to simplify management of the vLLM hosted deployment with Claude Code.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/sagemaker-gemma" rel="noopener noreferrer"&gt;https://github.com/xbill9/sagemaker-gemma&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  What is this project trying to Do?
&lt;/h4&gt;

&lt;p&gt;This project serves Gemma 4 E2B from a SageMaker real-time endpoint on one NVIDIA L4 GPU, using the vLLM container AWS publishes for SageMaker. Every AWS call is a plain &lt;code&gt;aws&lt;/code&gt; CLI command, so each step can be run by hand or by the MCP server.&lt;/p&gt;

&lt;p&gt;A SageMaker real-time endpoint is a managed HTTPS inference server. SageMaker places the container on a GPU instance, health-checks it, routes requests to it and writes its logs to CloudWatch. There is no instance to patch, no security group to open and no load balancer to build.&lt;/p&gt;




&lt;h4&gt;
  
  
  Where do I start?
&lt;/h4&gt;

&lt;p&gt;The strategy for starting MCP development for model management is a incremental step by step approach.&lt;/p&gt;

&lt;p&gt;First, the basic development environment is setup with the required system variables and a working Claude Code configuration.&lt;/p&gt;

&lt;p&gt;Then, the Python MCP server is brought up over stdio and validated with Claude Code in the local environment. The deployment follows as eight steps, each shown as the raw &lt;code&gt;aws&lt;/code&gt; command and the MCP tool that runs it.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;An AWS account and the AWS CLI v2, signed in with &lt;code&gt;aws login&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;A SageMaker endpoint quota of at least 1 for a single-L4 instance type (&lt;code&gt;ml.g6.xlarge&lt;/code&gt;, &lt;code&gt;ml.g6.2xlarge&lt;/code&gt; or &lt;code&gt;ml.g6.4xlarge&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Python 3.11 or newer with &lt;code&gt;mcp&lt;/code&gt; 2.x&lt;/li&gt;
&lt;li&gt;Claude Code or Gemini CLI installed and working&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;jq&lt;/code&gt; for reading JSON replies&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Setup the Basic Environment
&lt;/h4&gt;

&lt;p&gt;Clone the repository and install the one requirement into the system Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/xbill9/sagemaker-gemma
&lt;span class="nb"&gt;cd &lt;/span&gt;sagemaker-gemma
python3 &lt;span class="nt"&gt;-m&lt;/span&gt; pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.env&lt;/code&gt; is gitignored. It holds the settings every tool reads:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; .env
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Copy to .env (gitignored) and edit. sm.py reads it, so the MCP server, CLI and Makefile all see it.
AWS_REGION=us-east-2
MODEL_ID=google/gemma-4-E2B-it
INSTANCE_TYPE=ml.g6.xlarge
ENDPOINT_NAME=gemma-4-e2b
ROLE_NAME=sagemaker-gemma-execution-role
MAX_MODEL_LEN=8192
# Leave empty to use the newest SageMaker vLLM image in the region.
IMAGE_URI=
# Fallback instance types, highest priority first (same GPU keeps runs comparable).
INSTANCE_POOLS=ml.g6.xlarge,ml.g6.2xlarge,ml.g6.4xlarge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gemma 4 is Apache-2.0 and ungated on Hugging Face, so no Hugging Face token is needed.&lt;/p&gt;




&lt;h4&gt;
  
  
  Model Management Tool with MCP Stdio Transport
&lt;/h4&gt;

&lt;p&gt;The simplest MCP transport is stdio: the client launches the server as a local process and talks to it over stdin and stdout. In this project Claude Code is the MCP client. The server is one file, &lt;code&gt;server.py&lt;/code&gt;, on the MCP Python SDK 2.x:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;mcp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MCPServer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;RIG_NAME&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;READ_ONLY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ToolAnnotations&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;readOnlyHint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;idempotentHint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;WRITE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ToolAnnotations&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;destructiveHint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;DESTRUCTIVE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ToolAnnotations&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;destructiveHint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every tool carries one of the three annotations, so a client can tell a status check from a deploy from a delete.&lt;/p&gt;

&lt;p&gt;The tools do no AWS work themselves. They call &lt;code&gt;sm.py&lt;/code&gt;, which runs each request as an &lt;code&gt;aws&lt;/code&gt; CLI subprocess:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;aws&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;REGION&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Run `aws &amp;lt;args&amp;gt; --output json` and return the parsed result.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;cmd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;aws&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The CLI renews an &lt;code&gt;aws login&lt;/code&gt; session on its own, so a server that stays up for hours keeps working credentials. &lt;code&gt;sm.py&lt;/code&gt; also drops any &lt;code&gt;AWS_SESSION_TOKEN&lt;/code&gt; the server inherits from its parent process, because a static token expires inside a long-running server and outranks the login session.&lt;/p&gt;




&lt;h4&gt;
  
  
  Running the Python Code
&lt;/h4&gt;

&lt;p&gt;The project can be linted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make lint
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;All checks passed!
16 files already formatted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and tested:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make &lt;span class="nb"&gt;test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;----------------------------------------------------------------------
Ran 16 tests in 0.010s

OK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tests replace the &lt;code&gt;aws&lt;/code&gt; subprocess with a fake, so they run offline with no credentials. One of them compares the registered tool set and annotations against a fixed list: a tool that failed to register, or a delete that lost its destructive flag, fails the suite. ✅&lt;/p&gt;




&lt;h4&gt;
  
  
  Test the Protocol by Hand
&lt;/h4&gt;

&lt;p&gt;A client speaks JSON-RPC over stdio. Hold stdin open with &lt;code&gt;sleep&lt;/code&gt;, or the server sees end-of-input and exits before it answers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%s\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"probe","version":"0"}}}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'{"jsonrpc":"2.0","method":"notifications/initialized"}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}}'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;sleep &lt;/span&gt;3&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | python3 server.py 2&amp;gt;/dev/null
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Summarised:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 sagemaker-gemma
2 ['check_quotas', 'delete_endpoint', 'deploy_endpoint', 'find_vllm_image', 'get_deployment_config', 'get_endpoint_logs', 'get_endpoint_status', 'get_help', 'list_endpoints', 'query_model', 'verify_model_health']
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;🟢 The server answers the handshake and lists 11 tools.&lt;/p&gt;




&lt;h4&gt;
  
  
  Claude Code .mcp.json
&lt;/h4&gt;

&lt;p&gt;Claude Code reads &lt;code&gt;.mcp.json&lt;/code&gt; in the project directory and launches the server with the system &lt;code&gt;python3&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"sagemaker-gemma"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"python3"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"/home/xbill/sagemaker-gemma/server.py"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gemini CLI reads the same entry from &lt;code&gt;.gemini/settings.json&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Validation with Claude Code
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp get sagemaker-gemma
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sagemaker-gemma:
  Scope: Project config (shared via .mcp.json)
  Status: ✔ Connected
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From inside Claude Code, &lt;code&gt;get_help&lt;/code&gt; returns the resolved settings and the order of work:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;### sagemaker-gemma&lt;/span&gt;

Serving &lt;span class="sb"&gt;`google/gemma-4-E2B-it`&lt;/span&gt; with vLLM on a &lt;span class="gs"&gt;**SageMaker real-time endpoint**&lt;/span&gt;, managed
entirely through the aws CLI.

| Setting | Value |
| --- | --- |
| Region | &lt;span class="sb"&gt;`us-east-2`&lt;/span&gt; |
| Endpoint | &lt;span class="sb"&gt;`gemma-4-e2b`&lt;/span&gt; |
| Instance types (priority order) | &lt;span class="sb"&gt;`ml.g6.xlarge`&lt;/span&gt;, &lt;span class="sb"&gt;`ml.g6.2xlarge`&lt;/span&gt;, &lt;span class="sb"&gt;`ml.g6.4xlarge`&lt;/span&gt; |
| Max model length | &lt;span class="sb"&gt;`8192`&lt;/span&gt; |
| Image | &lt;span class="sb"&gt;`newest SageMaker vLLM image (find_vllm_image)`&lt;/span&gt; |
| Execution role | &lt;span class="sb"&gt;`sagemaker-gemma-execution-role`&lt;/span&gt; |

&lt;span class="gs"&gt;**Order of work:**&lt;/span&gt; check_quotas → deploy_endpoint → get_endpoint_status until
&lt;span class="sb"&gt;`InService`&lt;/span&gt; (about 10 minutes once an instance is placed) → verify_model_health
→ query_model → delete_endpoint.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The steps below follow that order. Each one shows the raw CLI command; &lt;code&gt;deploy_endpoint&lt;/code&gt; runs Steps 2 to 6 in one call.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;AWS_REGION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;us-east-2 &lt;span class="nv"&gt;AWS_PAGER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gemma-4-e2b
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MODEL_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;google/gemma-4-E2B-it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 1 — Check the Endpoint Quota
&lt;/h4&gt;

&lt;p&gt;SageMaker endpoint quotas are per instance type, per region, and count instances.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws service-quotas list-service-quotas &lt;span class="nt"&gt;--service-code&lt;/span&gt; sagemaker &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"Quotas[?QuotaName=='ml.g6.xlarge for endpoint usage'].Value"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[
    1.0
]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;check_quotas&lt;/code&gt; tool reads the same quota for each fallback type across the US regions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;| Instance | us-east-1 | us-east-2 | us-west-1 | us-west-2 |
| --- | ---: | ---: | ---: | ---: |
| `ml.g6.xlarge` | 1 | 1 | - | 1 |
| `ml.g6.2xlarge` | 1 | 1 | - | 1 |
| `ml.g6.4xlarge` | 1 | 1 | - | 1 |

Regions with quota for at least one of these types: us-east-1, us-east-2, us-west-2.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;-&lt;/code&gt; means the type is not offered in that region. A &lt;code&gt;0&lt;/code&gt; means a quota increase request first.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: One Quota Call per Region
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;list-service-quotas&lt;/code&gt; pages through every SageMaker quota in the region, and the Service Quotas API is rate limited per account. Twelve calls at once, one per type per region, came back as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;An error occurred (TooManyRequestsException) when calling the ListServiceQuotas operation (reached max retries: 2)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;check_quotas&lt;/code&gt; makes one call per region, one region at a time, and filters the types from that one result. &lt;code&gt;sm.py&lt;/code&gt; also sets &lt;code&gt;AWS_RETRY_MODE=adaptive&lt;/code&gt; so the CLI backs off and retries.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 2 — Find the vLLM Container
&lt;/h4&gt;

&lt;p&gt;The AWS vLLM repository holds pinned release tags such as &lt;code&gt;0.30.0-gpu-py312-cu130-ubuntu24.04-sagemaker-v1.1&lt;/code&gt;, floating aliases such as &lt;code&gt;0.30-gpu-py312&lt;/code&gt;, and &lt;code&gt;-soci&lt;/code&gt; index tags. Older release lines receive patch rebuilds, so the most recent push can carry an older vLLM. Sort the pinned tags by version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;TAG&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws ecr describe-images &lt;span class="nt"&gt;--registry-id&lt;/span&gt; 763104351884 &lt;span class="nt"&gt;--repository-name&lt;/span&gt; vllm &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"imageDetails[].imageTags[]"&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="s1"&gt;'\t'&lt;/span&gt; &lt;span class="s1"&gt;'\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="s1"&gt;'^[0-9]+\.[0-9]+\.[0-9]+-.*-sagemaker-v[0-9]+\.[0-9]+$'&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-V&lt;/span&gt; | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;IMAGE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;763104351884.dkr.ecr.&lt;span class="nv"&gt;$AWS_REGION&lt;/span&gt;.amazonaws.com/vllm:&lt;span class="nv"&gt;$TAG&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nv"&gt;$IMAGE&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;763104351884.dkr.ecr.us-east-2.amazonaws.com/vllm:0.30.0-gpu-py312-cu130-ubuntu24.04-sagemaker-v1.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;find_vllm_image&lt;/code&gt; tool applies the same rule. The image runs vLLM 0.30.0.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 3 — Create the Execution Role
&lt;/h4&gt;

&lt;p&gt;SageMaker assumes this role to pull the image and write logs. It is created once per account.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws iam create-role &lt;span class="nt"&gt;--role-name&lt;/span&gt; sagemaker-gemma-execution-role &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--assume-role-policy-document&lt;/span&gt; &lt;span class="s1"&gt;'{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Principal":{"Service":"sagemaker.amazonaws.com"},"Action":"sts:AssumeRole"}]}'&lt;/span&gt;
aws iam attach-role-policy &lt;span class="nt"&gt;--role-name&lt;/span&gt; sagemaker-gemma-execution-role &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--policy-arn&lt;/span&gt; arn:aws:iam::aws:policy/AmazonSageMakerFullAccess
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ROLE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws iam get-role &lt;span class="nt"&gt;--role-name&lt;/span&gt; sagemaker-gemma-execution-role &lt;span class="nt"&gt;--query&lt;/span&gt; Role.Arn &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a second run &lt;code&gt;create-role&lt;/code&gt; reports &lt;code&gt;EntityAlreadyExists&lt;/code&gt;; &lt;code&gt;get-role&lt;/code&gt; still sets &lt;code&gt;ROLE&lt;/code&gt;. &lt;code&gt;deploy_endpoint&lt;/code&gt; creates the role only when it is missing.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — Create the Model
&lt;/h4&gt;

&lt;p&gt;A SageMaker model pairs an image with its settings. The container turns each &lt;code&gt;SM_VLLM_&lt;/code&gt; variable into the matching vLLM flag: &lt;code&gt;SM_VLLM_MAX_MODEL_LEN&lt;/code&gt; becomes &lt;code&gt;--max-model-len&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker create-model &lt;span class="nt"&gt;--model-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="nt"&gt;--execution-role-arn&lt;/span&gt; &lt;span class="nv"&gt;$ROLE&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--primary-container&lt;/span&gt; &lt;span class="s2"&gt;"{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Image&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="nv"&gt;$IMAGE&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Environment&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:{
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;SM_VLLM_MODEL&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL_ID&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;SM_VLLM_MAX_MODEL_LEN&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;8192&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;SM_VLLM_GPU_MEMORY_UTILIZATION&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;0.9&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;}}"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ModelArn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:sagemaker:us-east-2:&amp;lt;account-id&amp;gt;:model/gemma-4-e2b"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 5 — Create the Endpoint Config With Fallback Instance Types
&lt;/h4&gt;

&lt;p&gt;The endpoint config says where the model runs. &lt;code&gt;InstancePools&lt;/code&gt; lists up to five instance types in priority order, and SageMaker places the first one with a free instance. All three types below carry one L4, so the model sees the same GPU whichever is placed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker create-endpoint-config &lt;span class="nt"&gt;--endpoint-config-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--production-variants&lt;/span&gt; &lt;span class="s2"&gt;"[{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;VariantName&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;AllTraffic&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ModelName&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="nv"&gt;$NAME&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InitialInstanceCount&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:1,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InstancePools&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:[{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InstanceType&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ml.g6.xlarge&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Priority&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:1},
                       {&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InstanceType&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ml.g6.2xlarge&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Priority&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:2},
                       {&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InstanceType&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ml.g6.4xlarge&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Priority&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:3}],
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ContainerStartupHealthCheckTimeoutInSeconds&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:1800,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ModelDataDownloadTimeoutInSeconds&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:1800}]"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"EndpointConfigArn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:sagemaker:us-east-2:&amp;lt;account-id&amp;gt;:endpoint-config/gemma-4-e2b"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two 1800-second timeouts give the container time to download the weights and compile before SageMaker's health check gives up.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: A Fallback List Holds Quota for Every Type in It
&lt;/h4&gt;

&lt;p&gt;With one Gemma endpoint running on &lt;code&gt;ml.g6.xlarge&lt;/code&gt; from the three-type list above, a second endpoint asking for &lt;code&gt;ml.g6.2xlarge&lt;/code&gt; in the same region was refused:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ResourceLimitExceeded: The account-level service limit 'ml.g6.2xlarge for endpoint usage' is 1 Instances, with current utilization of 1 Instances and a request delta of 1 Instances.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With a quota of 1 per type, a second endpoint goes in another region, or uses types outside the first endpoint's list.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 6 — Create the Endpoint
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker create-endpoint &lt;span class="nt"&gt;--endpoint-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="nt"&gt;--endpoint-config-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"EndpointArn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:sagemaker:us-east-2:&amp;lt;account-id&amp;gt;:endpoint/gemma-4-e2b"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Billing starts when an instance is placed. &lt;code&gt;aws sagemaker wait endpoint-in-service --endpoint-name $NAME&lt;/code&gt; blocks until it is ready; &lt;code&gt;get_endpoint_status&lt;/code&gt; reports the same state and the instance type that was placed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✅ `gemma-4-e2b` in `us-east-2`: **InService**
- Instance: `ml.g6.xlarge`
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The container log is in CloudWatch, and &lt;code&gt;get_endpoint_logs&lt;/code&gt; tails it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws logs &lt;span class="nb"&gt;tail&lt;/span&gt; /aws/sagemaker/Endpoints/&lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="nt"&gt;--follow&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start-up on the L4, in minutes after &lt;code&gt;create-endpoint&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phase&lt;/th&gt;
&lt;th&gt;Minutes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weights loaded, 9.75 GiB in 82.75 s&lt;/td&gt;
&lt;td&gt;7.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache sized, 723,484 tokens&lt;/td&gt;
&lt;td&gt;9.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;InService&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;9.9&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  🔎 Tip: A Creating Endpoint Cannot Be Deleted
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;create-endpoint&lt;/code&gt; cannot be undone until the endpoint settles. A &lt;code&gt;delete-endpoint&lt;/code&gt; sent while it is &lt;code&gt;Creating&lt;/code&gt; is refused:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;aws: [ERROR]: An error occurred (ValidationException) when calling the DeleteEndpoint operation: Cannot update in-progress endpoint "arn:aws:sagemaker:us-east-2:&amp;lt;account-id&amp;gt;:endpoint/gemma-4-e2b".
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The endpoint finishes starting, bills from the moment its instance is placed, and can be deleted once it reaches &lt;code&gt;InService&lt;/code&gt; or &lt;code&gt;Failed&lt;/code&gt;. Check the model ID and instance types before Step 6.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: Capacity Is Separate From Quota
&lt;/h4&gt;

&lt;p&gt;A quota of 1 lets you request one instance; the region still has to have one free. In &lt;code&gt;us-east-1&lt;/code&gt;, two requests for L4 instances each stayed in &lt;code&gt;Creating&lt;/code&gt; for about 30 minutes and then failed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Unable to provision requested ML compute capacity due to InsufficientInstanceCapacity error. Please retry using a different ML instance type or after some time.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same request in &lt;code&gt;us-east-2&lt;/code&gt; placed an instance at once. While SageMaker waits for capacity, no container starts and the CloudWatch log group never appears, which is how &lt;code&gt;get_endpoint_logs&lt;/code&gt; tells a capacity wait from a slow model load. No charge accrues during the wait. When it fails, repeat Steps 4 to 6 in another region where Step 1 shows the same quota.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 7 — Cross Check the Deployed Model
&lt;/h4&gt;

&lt;p&gt;The vLLM container accepts an OpenAI chat body on &lt;code&gt;invoke-endpoint&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'{"messages":[{"role":"user","content":"Why is the sky blue?"}],"max_tokens":256}'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; req.json
aws sagemaker-runtime invoke-endpoint &lt;span class="nt"&gt;--endpoint-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--content-type&lt;/span&gt; application/json &lt;span class="nt"&gt;--body&lt;/span&gt; fileb://req.json out.json
jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.choices[0].message.content'&lt;/span&gt; out.json | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;
jq &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'{model,usage}'&lt;/span&gt; out.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{
    "ContentType": "application/json",
    "InvokedProductionVariant": "AllTraffic"
}
The sky is blue due to a phenomenon called **Rayleigh scattering**. This process is caused by how sunlight interacts with the Earth's atmosphere.
{"model":"google/gemma-4-E2B-it","usage":{"prompt_tokens":15,"total_tokens":271,"completion_tokens":256,"prompt_tokens_details":null,"completion_tokens_details":null}}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;verify_model_health&lt;/code&gt; sends one short request and checks for a reply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✅ model=`google/gemma-4-E2B-it` tokens=2 wall=0.749s reply='ok'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 8 — Teardown
&lt;/h4&gt;

&lt;p&gt;The endpoint bills by the hour until it is deleted. Deleting the config and the model as well leaves nothing behind. &lt;code&gt;delete_endpoint&lt;/code&gt; runs all three:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker delete-endpoint &lt;span class="nt"&gt;--endpoint-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt;
aws sagemaker delete-endpoint-config &lt;span class="nt"&gt;--endpoint-config-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt;
aws sagemaker delete-model &lt;span class="nt"&gt;--model-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each command prints nothing and exits 0. &lt;code&gt;list_endpoints&lt;/code&gt; confirms the account is clear:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;### 0 endpoint(s) matching `gemma`
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to deploy Gemma 4 E2B to an Amazon SageMaker real-time endpoint with the AWS CLI and manage it from an MCP server. The key to the solution was the AWS vLLM SageMaker container, which turns the deployment into three &lt;code&gt;create-&lt;/code&gt; calls and makes the endpoint answer OpenAI-style chat requests. The deployment results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 Eight CLI steps take an account from quota check to a serving endpoint and back to nothing&lt;/li&gt;
&lt;li&gt;🟢 The endpoint reached &lt;code&gt;InService&lt;/code&gt; 9.9 minutes after &lt;code&gt;create-endpoint&lt;/code&gt;, with the weights using 9.75 GiB of the L4's 24 GB&lt;/li&gt;
&lt;li&gt;🟢 The MCP server runs every step from Claude Code, marks each tool read-only, write or destructive, and passes its tests offline&lt;/li&gt;
&lt;li&gt;🟢 &lt;code&gt;InstancePools&lt;/code&gt; gives one endpoint config several fallback instance types&lt;/li&gt;
&lt;li&gt;⚠️ L4 capacity varied by region: &lt;code&gt;us-east-1&lt;/code&gt; refused twice, about 30 minutes each, while &lt;code&gt;us-east-2&lt;/code&gt; placed an instance at once&lt;/li&gt;
&lt;li&gt;⚠️ A fallback list holds quota for every type in it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: one account, &lt;code&gt;ml.g6.xlarge&lt;/code&gt; with one NVIDIA L4, vLLM 0.30.0 from the AWS container, &lt;code&gt;google/gemma-4-E2B-it&lt;/code&gt; at full precision, deployed in &lt;code&gt;us-east-2&lt;/code&gt; on 2026-09-25. Start-up times are from a single deployment.&lt;/p&gt;

&lt;p&gt;The strategy for using MCP for SageMaker deployment was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Repository: &lt;a href="https://github.com/xbill9/sagemaker-gemma" rel="noopener noreferrer"&gt;https://github.com/xbill9/sagemaker-gemma&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 E2B on Hugging Face: &lt;a href="https://huggingface.co/google/gemma-4-E2B-it" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-E2B-it&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;AWS Deep Learning Containers: &lt;a href="https://github.com/aws/deep-learning-containers" rel="noopener noreferrer"&gt;https://github.com/aws/deep-learning-containers&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;SageMaker real-time inference: &lt;a href="https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;CreateEndpointConfig API: &lt;a href="https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_CreateEndpointConfig.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_CreateEndpointConfig.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MCP Python SDK: &lt;a href="https://github.com/modelcontextprotocol/python-sdk" rel="noopener noreferrer"&gt;https://github.com/modelcontextprotocol/python-sdk&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;vLLM: &lt;a href="https://docs.vllm.ai/" rel="noopener noreferrer"&gt;https://docs.vllm.ai/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aws</category>
      <category>sagemaker</category>
      <category>gemma</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Fri, 25 Sep 2026 21:55:07 +0000</pubDate>
      <link>https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-qat-weights-decode-205x-faster-than-bf16-on-one-l4-318m</link>
      <guid>https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-qat-weights-decode-205x-faster-than-bf16-on-one-l4-318m</guid>
      <description>&lt;p&gt;This article gives a short background on Amazon SageMaker real-time endpoints, then measures Gemma 4 E2B's quantization-aware trained (QAT) checkpoint against the full-size bf16 release on the same NVIDIA L4 endpoint. A suite of Python MCP tools is built to simplify management of the vLLM hosted deployment.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/sagemaker-gemma" rel="noopener noreferrer"&gt;https://github.com/xbill9/sagemaker-gemma&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Models&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;google/gemma-4-E2B-it&lt;/code&gt; (bf16) and &lt;code&gt;google/gemma-4-E2B-it-qat-w4a16-ct&lt;/code&gt; (QAT, 4-bit weights)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hardware&lt;/td&gt;
&lt;td&gt;SageMaker &lt;code&gt;ml.g6.xlarge&lt;/code&gt;, 1x NVIDIA L4, 24 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Region&lt;/td&gt;
&lt;td&gt;&lt;code&gt;us-east-2&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Software&lt;/td&gt;
&lt;td&gt;AWS vLLM SageMaker container, vLLM 0.30.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Result&lt;/td&gt;
&lt;td&gt;QAT decodes at &lt;strong&gt;105.1 tok/s&lt;/strong&gt; against &lt;strong&gt;51.3&lt;/strong&gt;, and serves &lt;strong&gt;1077.25 tok/s&lt;/strong&gt; at 16 parallel requests against &lt;strong&gt;619.1&lt;/strong&gt;, with the same score on 40 checked questions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  SageMaker Hosting in Five Minutes
&lt;/h4&gt;

&lt;p&gt;SageMaker real-time inference is three objects, created in order:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Object&lt;/th&gt;
&lt;th&gt;What it holds&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;A container image, its environment variables, and an IAM role&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint config&lt;/td&gt;
&lt;td&gt;Which model runs on which instance types, and how many instances&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint&lt;/td&gt;
&lt;td&gt;The running HTTPS service, billed per instance-hour while it exists&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Requests go through &lt;code&gt;aws sagemaker-runtime invoke-endpoint&lt;/code&gt;, signed with your AWS credentials. SageMaker health-checks the container, routes traffic to it and writes its log to CloudWatch under &lt;code&gt;/aws/sagemaker/Endpoints/&amp;lt;name&amp;gt;&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The container here is the vLLM build AWS publishes for SageMaker. It reads vLLM settings from &lt;code&gt;SM_VLLM_&lt;/code&gt; environment variables, downloads the model from Hugging Face at start-up, and accepts OpenAI-style chat bodies. Switching checkpoints is one variable: &lt;code&gt;SM_VLLM_MODEL&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;SageMaker JumpStart also lists Gemma 4 as ready-made packages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker list-hub-contents &lt;span class="nt"&gt;--hub-name&lt;/span&gt; SageMakerPublicHub &lt;span class="nt"&gt;--hub-content-type&lt;/span&gt; Model &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"HubContentSummaries[].HubContentName"&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="s1"&gt;'\t'&lt;/span&gt; &lt;span class="s1"&gt;'\n'&lt;/span&gt; | &lt;span class="nb"&gt;grep &lt;/span&gt;gemma-4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;huggingface-llm-gemma-4-31b-it-nvfp4
huggingface-vlm-gemma-4-12b-it
huggingface-vlm-gemma-4-26b-a4b-it
huggingface-vlm-gemma-4-31b-it
huggingface-vlm-gemma-4-31b-it-fp8-block
huggingface-vlm-gemma-4-e2b-instruct
huggingface-vlm-gemma-4-e4b-it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;None of the seven is a Google QAT checkpoint, so this comparison uses the vLLM container with a Hugging Face model ID.&lt;/p&gt;

&lt;p&gt;Three account limits shape every deployment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quota&lt;/strong&gt; is per instance type, per region, counted in instances. This account holds 1 for each single-L4 type in &lt;code&gt;us-east-1&lt;/code&gt;, &lt;code&gt;us-east-2&lt;/code&gt; and &lt;code&gt;us-west-2&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity&lt;/strong&gt; is separate. &lt;code&gt;us-east-1&lt;/code&gt; and &lt;code&gt;us-west-2&lt;/code&gt; each left an L4 request in &lt;code&gt;Creating&lt;/code&gt; for about 30 minutes and then returned &lt;code&gt;InsufficientInstanceCapacity&lt;/code&gt;. &lt;code&gt;us-east-2&lt;/code&gt; placed one within minutes every time it was asked.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A fallback list&lt;/strong&gt; (&lt;code&gt;InstancePools&lt;/code&gt; in the endpoint config) holds quota for every type in it, so two endpoints cannot share one region's single-L4 quota.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With one L4 per region to work with, the two checkpoints ran one after the other on the same instance type in the same region.&lt;/p&gt;




&lt;h4&gt;
  
  
  Where Do I Start?
&lt;/h4&gt;

&lt;p&gt;The deployment itself, from quota check to teardown with the aws CLI and the MCP server, is the first article in this pair: &lt;a href="https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d"&gt;https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This article starts from a working endpoint and changes one thing: the checkpoint.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;The repository above, with &lt;code&gt;aws login&lt;/code&gt; done and &lt;code&gt;mcp&lt;/code&gt; 2.x installed&lt;/li&gt;
&lt;li&gt;A SageMaker quota of at least 1 for &lt;code&gt;ml.g6.xlarge&lt;/code&gt; in a region with L4 capacity&lt;/li&gt;
&lt;li&gt;The full-size endpoint from the first article deployed as &lt;code&gt;gemma-4-e2b&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  What QAT Changes
&lt;/h4&gt;

&lt;p&gt;Google trains the QAT checkpoint with 4-bit weights in the loop, then exports it in the &lt;code&gt;compressed-tensors&lt;/code&gt; format vLLM reads natively. The &lt;code&gt;-w4a16-ct&lt;/code&gt; suffix means 4-bit weights, 16-bit activations. Google publishes the same model in four QAT forms; only this one loads in vLLM:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Checkpoint&lt;/th&gt;
&lt;th&gt;vLLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-qat-w4a16-ct&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;🟢 loads, 4-bit weights&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-qat-q4_0-unquantized&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;⚠️ stored at 16-bit, no memory saving&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-qat-q4_0-gguf&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;❌ GGUF, for llama.cpp and Ollama&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-qat-mobile-*&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;❌ on-device formats&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Deploying it is the first article's Steps 4 to 6 with two variables changed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gemma-4-e2b-qat
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MODEL_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;google/gemma-4-E2B-it-qat-w4a16-ct
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;vLLM names the format when it starts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;quantization=compressed-tensors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  What the Engine Allocates
&lt;/h4&gt;

&lt;p&gt;The container log records the weights, the load time and the KV cache vLLM builds from what is left of the L4's memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model loading took 9.75 GiB memory and 82.751806 seconds
GPU KV cache size: 723,484 tokens, Maximum concurrency for 8,192 tokens per request: 88.32x
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model loading took 8.01 GiB memory and 66.190265 seconds
GPU KV cache size: 867,999 tokens, Maximum concurrency for 8,192 tokens per request: 105.96x
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;th&gt;QAT / bf16&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weights (GiB)&lt;/td&gt;
&lt;td&gt;9.75&lt;/td&gt;
&lt;td&gt;8.01&lt;/td&gt;
&lt;td&gt;0.82&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache (tokens)&lt;/td&gt;
&lt;td&gt;723,484&lt;/td&gt;
&lt;td&gt;867,999&lt;/td&gt;
&lt;td&gt;1.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weight load (s)&lt;/td&gt;
&lt;td&gt;82.75&lt;/td&gt;
&lt;td&gt;66.19&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Create to &lt;code&gt;InService&lt;/code&gt; (min)&lt;/td&gt;
&lt;td&gt;9.9&lt;/td&gt;
&lt;td&gt;10.1&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 4-bit weights save 18% of GPU memory. The checkpoint's own header shows why: only the transformer body is 4-bit.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Part of the QAT file&lt;/th&gt;
&lt;th&gt;GB&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;th&gt;Stored as&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Per-layer embedding&lt;/td&gt;
&lt;td&gt;4.698&lt;/td&gt;
&lt;td&gt;56.5%&lt;/td&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vocabulary embedding&lt;/td&gt;
&lt;td&gt;1.611&lt;/td&gt;
&lt;td&gt;19.4%&lt;/td&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transformer body&lt;/td&gt;
&lt;td&gt;1.056&lt;/td&gt;
&lt;td&gt;12.7%&lt;/td&gt;
&lt;td&gt;packed 4-bit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audio tower&lt;/td&gt;
&lt;td&gt;0.614&lt;/td&gt;
&lt;td&gt;7.4%&lt;/td&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision tower&lt;/td&gt;
&lt;td&gt;0.337&lt;/td&gt;
&lt;td&gt;4.1%&lt;/td&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The memory the smaller weights free goes to the KV cache.&lt;/p&gt;




&lt;h4&gt;
  
  
  How the Measurement Works
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;compare.py&lt;/code&gt; runs the same three measurements against each endpoint, at temperature 0, through the same &lt;code&gt;aws sagemaker-runtime invoke-endpoint&lt;/code&gt; call:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Decode.&lt;/strong&gt; Fixed-length replies of 16 and 512 tokens (&lt;code&gt;ignore_eos&lt;/code&gt;), five of each. The decode rate is (512 − 16) / (median time at 512 − median time at 16), which cancels the aws CLI start-up and the network round trip.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parallel requests.&lt;/strong&gt; 1, 4 and 16 requests at once, 256 tokens each, two batches per level. Throughput is total output tokens divided by the batch's wall time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answers.&lt;/strong&gt; 40 fixed questions with exact answers: 15 two-digit multiplications, 15 three-number sums, 10 capitals. Scored by regular expression.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 compare.py measure docs/runs/2026-09-25-qat-vs-bf16 gemma-4-e2b-qat@us-east-2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;wrote docs/runs/2026-09-25-qat-vs-bf16/measure-gemma-4-e2b-qat.json
{
  "decode_tokens_per_second": 105.1,
  "per_call_fixed_cost_seconds": 0.562,
  "load_tokens_per_second": {
    "1": 85.35,
    "4": 328.7,
    "16": 1077.25
  },
  "quality": "37/40"
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;combine&lt;/code&gt; computes every ratio from the two result files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 compare.py combine compare.json measure-gemma-4-e2b.json measure-gemma-4-e2b-qat.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                               gemma-4-e2b     gemma-4-e2b-qat  ratio
weights_gib                           9.75                8.01  0.82
kv_cache_tokens                     723484              867999  1.2
load_seconds                     82.751806           66.190265  
decode_tokens_per_second              51.3               105.1  2.05
load_c1_tokens_per_second             45.6               85.35  1.87
load_c4_tokens_per_second            171.6               328.7  1.92
load_c16_tokens_per_second           619.1             1077.25  1.74
quality_correct                         37                  37  
identical answers: 35/40
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Decode Speed
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Decode (tokens/s)&lt;/td&gt;
&lt;td&gt;51.3&lt;/td&gt;
&lt;td&gt;🥇 105.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;512-token reply, fastest (s)&lt;/td&gt;
&lt;td&gt;10.491&lt;/td&gt;
&lt;td&gt;🥇 5.392&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;512-token reply, slowest (s)&lt;/td&gt;
&lt;td&gt;10.661&lt;/td&gt;
&lt;td&gt;🥇 5.526&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-call client cost (s)&lt;/td&gt;
&lt;td&gt;0.633&lt;/td&gt;
&lt;td&gt;0.562&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;QAT decodes 2.05x faster.&lt;/strong&gt; The slowest QAT reply finished in about half the time of the fastest bf16 one, so the gap is far larger than the spread between repeats.&lt;/p&gt;

&lt;p&gt;Each decode step reads every transformer layer's weights from GPU memory, so decode speed follows how many bytes those layers take. Those are the layers the QAT export stores at 4 bits. The embeddings are looked up one row per token, so their 16-bit size costs memory and little time.&lt;/p&gt;




&lt;h4&gt;
  
  
  Parallel Requests
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests at once&lt;/th&gt;
&lt;th&gt;bf16 tok/s&lt;/th&gt;
&lt;th&gt;QAT tok/s&lt;/th&gt;
&lt;th&gt;QAT / bf16&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;45.6&lt;/td&gt;
&lt;td&gt;🥇 85.35&lt;/td&gt;
&lt;td&gt;1.87&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;171.6&lt;/td&gt;
&lt;td&gt;🥇 328.7&lt;/td&gt;
&lt;td&gt;1.92&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;619.1&lt;/td&gt;
&lt;td&gt;🥇 1077.25&lt;/td&gt;
&lt;td&gt;1.74&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;QAT leads at every level. The ratio narrows at 16, where a batch shares each weight read across more requests and the per-token saving counts for less. The single-request figures sit below the decode rate because each call also pays the aws CLI start-up.&lt;/p&gt;




&lt;h4&gt;
  
  
  Does It Still Answer Correctly?
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Questions&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Multiplication (15)&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a + b − c (15)&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capitals (10)&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total (40)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;37&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;37&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;35 of the 40 answers are identical character for character. The other five are all sums, and each model misses three of them:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Expected&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;876 + 608 − 558&lt;/td&gt;
&lt;td&gt;926&lt;/td&gt;
&lt;td&gt;❌ 1026&lt;/td&gt;
&lt;td&gt;✅ 926&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;257 + 388 − 290&lt;/td&gt;
&lt;td&gt;355&lt;/td&gt;
&lt;td&gt;✅ 355&lt;/td&gt;
&lt;td&gt;❌ 655&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;782 + 571 − 534&lt;/td&gt;
&lt;td&gt;819&lt;/td&gt;
&lt;td&gt;❌ 829&lt;/td&gt;
&lt;td&gt;✅ 819&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;470 + 238 − 983&lt;/td&gt;
&lt;td&gt;−275&lt;/td&gt;
&lt;td&gt;✅ −275&lt;/td&gt;
&lt;td&gt;❌ −285&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;150 + 937 − 139&lt;/td&gt;
&lt;td&gt;948&lt;/td&gt;
&lt;td&gt;❌ 918&lt;/td&gt;
&lt;td&gt;❌ 1010&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two models make the same number of mistakes on different questions. Forty questions are enough to show a large loss and too few to measure a small one.&lt;/p&gt;




&lt;h4&gt;
  
  
  Re-Measured in A-B-A Order
&lt;/h4&gt;

&lt;p&gt;Each endpoint runs on its own physical instance, and the two ran 15 minutes apart. To check that the gap belongs to the checkpoint, the bf16 endpoint was deployed a second time after QAT and measured again:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Start (UTC)&lt;/th&gt;
&lt;th&gt;Decode tok/s&lt;/th&gt;
&lt;th&gt;16 at once tok/s&lt;/th&gt;
&lt;th&gt;Correct&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;td&gt;18:36&lt;/td&gt;
&lt;td&gt;51.3&lt;/td&gt;
&lt;td&gt;619.1&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QAT&lt;/td&gt;
&lt;td&gt;18:50&lt;/td&gt;
&lt;td&gt;105.1&lt;/td&gt;
&lt;td&gt;1077.25&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;bf16 again&lt;/td&gt;
&lt;td&gt;19:10&lt;/td&gt;
&lt;td&gt;51.5&lt;/td&gt;
&lt;td&gt;625.15&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two bf16 runs agree within 1.21% on every speed figure and give identical answers to all 40 questions. The second bf16 instance loaded its weights in 82.09 s against 82.75 s the first time.&lt;/p&gt;




&lt;h4&gt;
  
  
  Compare to Other Deployments
&lt;/h4&gt;

&lt;p&gt;The same bf16 model on the same GPU has been measured on two other platforms in this series, with &lt;code&gt;vllm bench serve&lt;/code&gt; at 128 output tokens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;1 request tok/s&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cloud Run, NVIDIA L4&lt;/td&gt;
&lt;td&gt;49.63&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EC2 &lt;code&gt;g6.2xlarge&lt;/code&gt;, NVIDIA L4&lt;/td&gt;
&lt;td&gt;46.09&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SageMaker &lt;code&gt;ml.g6.xlarge&lt;/code&gt;, NVIDIA L4&lt;/td&gt;
&lt;td&gt;45.6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The three sit within 9% of each other. SageMaker's figure includes the aws CLI start-up in every call; its decode rate with that removed is 51.3.&lt;/p&gt;

&lt;p&gt;On a Tesla T4, QAT decoded 1.79x faster than bf16 for the same model. On the L4 the ratio is 2.05x.&lt;/p&gt;

&lt;p&gt;The methods differ: the other runs used &lt;code&gt;vllm bench serve&lt;/code&gt; with random prompts, other vLLM versions and their own hosts. Read the rows as a shape.&lt;/p&gt;




&lt;h4&gt;
  
  
  And Price/Performance?
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;ml.g6.xlarge&lt;/code&gt; in &lt;code&gt;us-east-2&lt;/code&gt; is $1.1267 an hour on demand:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws pricing get-products &lt;span class="nt"&gt;--region&lt;/span&gt; us-east-1 &lt;span class="nt"&gt;--service-code&lt;/span&gt; AmazonSageMaker &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filters&lt;/span&gt; &lt;span class="nv"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;TERM_MATCH,Field&lt;span class="o"&gt;=&lt;/span&gt;instanceName,Value&lt;span class="o"&gt;=&lt;/span&gt;ml.g6.xlarge &lt;span class="nv"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;TERM_MATCH,Field&lt;span class="o"&gt;=&lt;/span&gt;regionCode,Value&lt;span class="o"&gt;=&lt;/span&gt;us-east-2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ml.g6.xlarge USE2-Host:ml.g6.xlarge 1.1267000000 Hrs | $1.1267 per Hosting ml.g6.xlarge hour in US East (Ohio)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same hourly price buys twice the tokens. Per million output tokens (arithmetic):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests at once&lt;/th&gt;
&lt;th&gt;bf16 $/M&lt;/th&gt;
&lt;th&gt;QAT $/M&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;6.86&lt;/td&gt;
&lt;td&gt;🥇 3.67&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;1.82&lt;/td&gt;
&lt;td&gt;🥇 0.95&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;0.51&lt;/td&gt;
&lt;td&gt;🥇 0.29&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The endpoint bills while it exists, busy or idle, so these figures hold only while it is kept busy.&lt;/p&gt;




&lt;h4&gt;
  
  
  So, Which One?
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;bf16&lt;/th&gt;
&lt;th&gt;QAT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Decode speed&lt;/td&gt;
&lt;td&gt;51.3 tok/s&lt;/td&gt;
&lt;td&gt;🥇 105.1 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16 requests at once&lt;/td&gt;
&lt;td&gt;619.1 tok/s&lt;/td&gt;
&lt;td&gt;🥇 1077.25 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU memory for weights&lt;/td&gt;
&lt;td&gt;9.75 GiB&lt;/td&gt;
&lt;td&gt;🥇 8.01 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checked answers&lt;/td&gt;
&lt;td&gt;37 / 40&lt;/td&gt;
&lt;td&gt;37 / 40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per million tokens at 16&lt;/td&gt;
&lt;td&gt;$0.51&lt;/td&gt;
&lt;td&gt;🥇 $0.29&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On an L4 SageMaker endpoint, the QAT checkpoint is the default choice for Gemma 4 E2B: the same instance, one changed environment variable, about twice the tokens per dollar, and no measured change in answers. Keep bf16 as the reference when a task's accuracy needs a larger evaluation than 40 questions.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Stops the Meter
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;delete_endpoint&lt;/code&gt; removes the endpoint, its config and its model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;🗑️ `gemma-4-e2b-qat` in `us-east-2`
- endpoint: deleted
- endpoint-config: deleted
- model: deleted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;Creating&lt;/code&gt; endpoint refuses deletion, so a deploy runs until it reaches &lt;code&gt;InService&lt;/code&gt; or &lt;code&gt;Failed&lt;/code&gt; before it can be removed.&lt;/p&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to measure what Gemma 4 E2B's QAT checkpoint changes on a SageMaker L4 endpoint. The key to the solution was changing only &lt;code&gt;SM_VLLM_MODEL&lt;/code&gt; between two deployments on the same instance type, and measuring decode speed with the per-call client cost removed. The measured results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 QAT decodes at &lt;strong&gt;105.1 tok/s&lt;/strong&gt; against bf16's &lt;strong&gt;51.3&lt;/strong&gt;, 2.05x&lt;/li&gt;
&lt;li&gt;🟢 At 16 requests at once QAT serves &lt;strong&gt;1077.25 tok/s&lt;/strong&gt; against &lt;strong&gt;619.1&lt;/strong&gt;, 1.74x&lt;/li&gt;
&lt;li&gt;🟢 Both score &lt;strong&gt;37 of 40&lt;/strong&gt; on checked questions, with 35 identical answers&lt;/li&gt;
&lt;li&gt;🟢 A second bf16 deployment after QAT reproduced the first within 1.21%&lt;/li&gt;
&lt;li&gt;⚠️ The 4-bit export saves 18% of weight memory, because the embeddings stay at 16-bit&lt;/li&gt;
&lt;li&gt;⚠️ Getting an L4 took three regions: &lt;code&gt;us-east-1&lt;/code&gt; and &lt;code&gt;us-west-2&lt;/code&gt; returned &lt;code&gt;InsufficientInstanceCapacity&lt;/code&gt; after about 30 minutes each&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: one account, SageMaker &lt;code&gt;ml.g6.xlarge&lt;/code&gt; with one NVIDIA L4 in &lt;code&gt;us-east-2&lt;/code&gt;, vLLM 0.30.0 from the AWS container, three deployments on 2026-09-25 each on its own instance, measured in bf16, QAT, bf16 order. Prompts were short; long-prompt behaviour was not measured. Decode used five replies per length, parallel throughput two batches per level, and quality 40 questions at temperature 0. Every request went through the aws CLI from one client machine.&lt;/p&gt;

&lt;p&gt;The strategy for using MCP for SageMaker deployment and benchmarking was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Repository: &lt;a href="https://github.com/xbill9/sagemaker-gemma" rel="noopener noreferrer"&gt;https://github.com/xbill9/sagemaker-gemma&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Part one, deploying Gemma 4 to SageMaker: &lt;a href="https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d"&gt;https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 E2B QAT w4a16: &lt;a href="https://huggingface.co/google/gemma-4-E2B-it-qat-w4a16-ct" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-E2B-it-qat-w4a16-ct&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 E2B: &lt;a href="https://huggingface.co/google/gemma-4-E2B-it" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-E2B-it&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 on a Tesla T4, QAT vs bf16: &lt;a href="https://dev.to/gde/gemma-4-on-a-tesla-t4-qat-weights-decode-179x-faster-than-bf16-2fi4"&gt;https://dev.to/gde/gemma-4-on-a-tesla-t4-qat-weights-decode-179x-faster-than-bf16-2fi4&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;2B Gemma 4 on Cloud Run with an NVIDIA L4: &lt;a href="https://dev.to/gde/2b-gemma-4-deployment-with-cloud-run-nvidia-l4-mcp-sdk-2x-and-claude-code-4ml3"&gt;https://dev.to/gde/2b-gemma-4-deployment-with-cloud-run-nvidia-l4-mcp-sdk-2x-and-claude-code-4ml3&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;SageMaker real-time inference: &lt;a href="https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;AWS Deep Learning Containers: &lt;a href="https://github.com/aws/deep-learning-containers" rel="noopener noreferrer"&gt;https://github.com/aws/deep-learning-containers&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;compressed-tensors: &lt;a href="https://github.com/neuralmagic/compressed-tensors" rel="noopener noreferrer"&gt;https://github.com/neuralmagic/compressed-tensors&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aws</category>
      <category>sagemaker</category>
      <category>gemma</category>
      <category>vllm</category>
    </item>
    <item>
      <title>Gemma 4 on an Amazon SageMaker Endpoint: AWS CLI, NVIDIA L4, and an MCP Server</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Fri, 25 Sep 2026 21:50:14 +0000</pubDate>
      <link>https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d</link>
      <guid>https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d</guid>
      <description>&lt;p&gt;This article provides a step by step deployment guide for Gemma 4 E2B to an Amazon SageMaker hosted GPU enabled system. A suite of Python MCP tools is built to simplify management of the vLLM hosted deployment with Claude Code.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/sagemaker-gemma" rel="noopener noreferrer"&gt;https://github.com/xbill9/sagemaker-gemma&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  What is this project trying to Do?
&lt;/h4&gt;

&lt;p&gt;This project serves Gemma 4 E2B from a SageMaker real-time endpoint on one NVIDIA L4 GPU, using the vLLM container AWS publishes for SageMaker. Every AWS call is a plain &lt;code&gt;aws&lt;/code&gt; CLI command, so each step can be run by hand or by the MCP server.&lt;/p&gt;

&lt;p&gt;A SageMaker real-time endpoint is a managed HTTPS inference server. SageMaker places the container on a GPU instance, health-checks it, routes requests to it and writes its logs to CloudWatch. There is no instance to patch, no security group to open and no load balancer to build.&lt;/p&gt;




&lt;h4&gt;
  
  
  Where do I start?
&lt;/h4&gt;

&lt;p&gt;The strategy for starting MCP development for model management is a incremental step by step approach.&lt;/p&gt;

&lt;p&gt;First, the basic development environment is setup with the required system variables and a working Claude Code configuration.&lt;/p&gt;

&lt;p&gt;Then, the Python MCP server is brought up over stdio and validated with Claude Code in the local environment. The deployment follows as eight steps, each shown as the raw &lt;code&gt;aws&lt;/code&gt; command and the MCP tool that runs it.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;An AWS account and the AWS CLI v2, signed in with &lt;code&gt;aws login&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;A SageMaker endpoint quota of at least 1 for a single-L4 instance type (&lt;code&gt;ml.g6.xlarge&lt;/code&gt;, &lt;code&gt;ml.g6.2xlarge&lt;/code&gt; or &lt;code&gt;ml.g6.4xlarge&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Python 3.11 or newer with &lt;code&gt;mcp&lt;/code&gt; 2.x&lt;/li&gt;
&lt;li&gt;Claude Code or Gemini CLI installed and working&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;jq&lt;/code&gt; for reading JSON replies&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Setup the Basic Environment
&lt;/h4&gt;

&lt;p&gt;Clone the repository and install the one requirement into the system Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/xbill9/sagemaker-gemma
&lt;span class="nb"&gt;cd &lt;/span&gt;sagemaker-gemma
python3 &lt;span class="nt"&gt;-m&lt;/span&gt; pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.env&lt;/code&gt; is gitignored. It holds the settings every tool reads:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; .env
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Copy to .env (gitignored) and edit. sm.py reads it, so the MCP server, CLI and Makefile all see it.
AWS_REGION=us-east-2
MODEL_ID=google/gemma-4-E2B-it
INSTANCE_TYPE=ml.g6.xlarge
ENDPOINT_NAME=gemma-4-e2b
ROLE_NAME=sagemaker-gemma-execution-role
MAX_MODEL_LEN=8192
# Leave empty to use the newest SageMaker vLLM image in the region.
IMAGE_URI=
# Fallback instance types, highest priority first (same GPU keeps runs comparable).
INSTANCE_POOLS=ml.g6.xlarge,ml.g6.2xlarge,ml.g6.4xlarge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gemma 4 is Apache-2.0 and ungated on Hugging Face, so no Hugging Face token is needed.&lt;/p&gt;




&lt;h4&gt;
  
  
  Model Management Tool with MCP Stdio Transport
&lt;/h4&gt;

&lt;p&gt;The simplest MCP transport is stdio: the client launches the server as a local process and talks to it over stdin and stdout. In this project Claude Code is the MCP client. The server is one file, &lt;code&gt;server.py&lt;/code&gt;, on the MCP Python SDK 2.x:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;mcp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MCPServer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;RIG_NAME&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;READ_ONLY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ToolAnnotations&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;readOnlyHint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;idempotentHint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;WRITE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ToolAnnotations&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;destructiveHint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;DESTRUCTIVE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ToolAnnotations&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;destructiveHint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every tool carries one of the three annotations, so a client can tell a status check from a deploy from a delete.&lt;/p&gt;

&lt;p&gt;The tools do no AWS work themselves. They call &lt;code&gt;sm.py&lt;/code&gt;, which runs each request as an &lt;code&gt;aws&lt;/code&gt; CLI subprocess:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;aws&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;REGION&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Run `aws &amp;lt;args&amp;gt; --output json` and return the parsed result.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;cmd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;aws&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The CLI renews an &lt;code&gt;aws login&lt;/code&gt; session on its own, so a server that stays up for hours keeps working credentials. &lt;code&gt;sm.py&lt;/code&gt; also drops any &lt;code&gt;AWS_SESSION_TOKEN&lt;/code&gt; the server inherits from its parent process, because a static token expires inside a long-running server and outranks the login session.&lt;/p&gt;




&lt;h4&gt;
  
  
  Running the Python Code
&lt;/h4&gt;

&lt;p&gt;The project can be linted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make lint
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;All checks passed!
16 files already formatted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and tested:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make &lt;span class="nb"&gt;test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;----------------------------------------------------------------------
Ran 16 tests in 0.010s

OK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tests replace the &lt;code&gt;aws&lt;/code&gt; subprocess with a fake, so they run offline with no credentials. One of them compares the registered tool set and annotations against a fixed list: a tool that failed to register, or a delete that lost its destructive flag, fails the suite. ✅&lt;/p&gt;




&lt;h4&gt;
  
  
  Test the Protocol by Hand
&lt;/h4&gt;

&lt;p&gt;A client speaks JSON-RPC over stdio. Hold stdin open with &lt;code&gt;sleep&lt;/code&gt;, or the server sees end-of-input and exits before it answers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%s\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"probe","version":"0"}}}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'{"jsonrpc":"2.0","method":"notifications/initialized"}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}}'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;sleep &lt;/span&gt;3&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | python3 server.py 2&amp;gt;/dev/null
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Summarised:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 sagemaker-gemma
2 ['check_quotas', 'delete_endpoint', 'deploy_endpoint', 'find_vllm_image', 'get_deployment_config', 'get_endpoint_logs', 'get_endpoint_status', 'get_help', 'list_endpoints', 'query_model', 'verify_model_health']
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;🟢 The server answers the handshake and lists 11 tools.&lt;/p&gt;




&lt;h4&gt;
  
  
  Claude Code .mcp.json
&lt;/h4&gt;

&lt;p&gt;Claude Code reads &lt;code&gt;.mcp.json&lt;/code&gt; in the project directory and launches the server with the system &lt;code&gt;python3&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"sagemaker-gemma"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"python3"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"/home/xbill/sagemaker-gemma/server.py"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gemini CLI reads the same entry from &lt;code&gt;.gemini/settings.json&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Validation with Claude Code
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp get sagemaker-gemma
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sagemaker-gemma:
  Scope: Project config (shared via .mcp.json)
  Status: ✔ Connected
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From inside Claude Code, &lt;code&gt;get_help&lt;/code&gt; returns the resolved settings and the order of work:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;### sagemaker-gemma&lt;/span&gt;

Serving &lt;span class="sb"&gt;`google/gemma-4-E2B-it`&lt;/span&gt; with vLLM on a &lt;span class="gs"&gt;**SageMaker real-time endpoint**&lt;/span&gt;, managed
entirely through the aws CLI.

| Setting | Value |
| --- | --- |
| Region | &lt;span class="sb"&gt;`us-east-2`&lt;/span&gt; |
| Endpoint | &lt;span class="sb"&gt;`gemma-4-e2b`&lt;/span&gt; |
| Instance types (priority order) | &lt;span class="sb"&gt;`ml.g6.xlarge`&lt;/span&gt;, &lt;span class="sb"&gt;`ml.g6.2xlarge`&lt;/span&gt;, &lt;span class="sb"&gt;`ml.g6.4xlarge`&lt;/span&gt; |
| Max model length | &lt;span class="sb"&gt;`8192`&lt;/span&gt; |
| Image | &lt;span class="sb"&gt;`newest SageMaker vLLM image (find_vllm_image)`&lt;/span&gt; |
| Execution role | &lt;span class="sb"&gt;`sagemaker-gemma-execution-role`&lt;/span&gt; |

&lt;span class="gs"&gt;**Order of work:**&lt;/span&gt; check_quotas → deploy_endpoint → get_endpoint_status until
&lt;span class="sb"&gt;`InService`&lt;/span&gt; (about 10 minutes once an instance is placed) → verify_model_health
→ query_model → delete_endpoint.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The steps below follow that order. Each one shows the raw CLI command; &lt;code&gt;deploy_endpoint&lt;/code&gt; runs Steps 2 to 6 in one call.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;AWS_REGION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;us-east-2 &lt;span class="nv"&gt;AWS_PAGER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gemma-4-e2b
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MODEL_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;google/gemma-4-E2B-it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 1 — Check the Endpoint Quota
&lt;/h4&gt;

&lt;p&gt;SageMaker endpoint quotas are per instance type, per region, and count instances.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws service-quotas list-service-quotas &lt;span class="nt"&gt;--service-code&lt;/span&gt; sagemaker &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"Quotas[?QuotaName=='ml.g6.xlarge for endpoint usage'].Value"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[
    1.0
]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;check_quotas&lt;/code&gt; tool reads the same quota for each fallback type across the US regions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;| Instance | us-east-1 | us-east-2 | us-west-1 | us-west-2 |
| --- | ---: | ---: | ---: | ---: |
| `ml.g6.xlarge` | 1 | 1 | - | 1 |
| `ml.g6.2xlarge` | 1 | 1 | - | 1 |
| `ml.g6.4xlarge` | 1 | 1 | - | 1 |

Regions with quota for at least one of these types: us-east-1, us-east-2, us-west-2.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;-&lt;/code&gt; means the type is not offered in that region. A &lt;code&gt;0&lt;/code&gt; means a quota increase request first.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: One Quota Call per Region
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;list-service-quotas&lt;/code&gt; pages through every SageMaker quota in the region, and the Service Quotas API is rate limited per account. Twelve calls at once, one per type per region, came back as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;An error occurred (TooManyRequestsException) when calling the ListServiceQuotas operation (reached max retries: 2)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;check_quotas&lt;/code&gt; makes one call per region, one region at a time, and filters the types from that one result. &lt;code&gt;sm.py&lt;/code&gt; also sets &lt;code&gt;AWS_RETRY_MODE=adaptive&lt;/code&gt; so the CLI backs off and retries.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 2 — Find the vLLM Container
&lt;/h4&gt;

&lt;p&gt;The AWS vLLM repository holds pinned release tags such as &lt;code&gt;0.30.0-gpu-py312-cu130-ubuntu24.04-sagemaker-v1.1&lt;/code&gt;, floating aliases such as &lt;code&gt;0.30-gpu-py312&lt;/code&gt;, and &lt;code&gt;-soci&lt;/code&gt; index tags. Older release lines receive patch rebuilds, so the most recent push can carry an older vLLM. Sort the pinned tags by version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;TAG&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws ecr describe-images &lt;span class="nt"&gt;--registry-id&lt;/span&gt; 763104351884 &lt;span class="nt"&gt;--repository-name&lt;/span&gt; vllm &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"imageDetails[].imageTags[]"&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="s1"&gt;'\t'&lt;/span&gt; &lt;span class="s1"&gt;'\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="s1"&gt;'^[0-9]+\.[0-9]+\.[0-9]+-.*-sagemaker-v[0-9]+\.[0-9]+$'&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-V&lt;/span&gt; | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;IMAGE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;763104351884.dkr.ecr.&lt;span class="nv"&gt;$AWS_REGION&lt;/span&gt;.amazonaws.com/vllm:&lt;span class="nv"&gt;$TAG&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nv"&gt;$IMAGE&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;763104351884.dkr.ecr.us-east-2.amazonaws.com/vllm:0.30.0-gpu-py312-cu130-ubuntu24.04-sagemaker-v1.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;find_vllm_image&lt;/code&gt; tool applies the same rule. The image runs vLLM 0.30.0.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 3 — Create the Execution Role
&lt;/h4&gt;

&lt;p&gt;SageMaker assumes this role to pull the image and write logs. It is created once per account.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws iam create-role &lt;span class="nt"&gt;--role-name&lt;/span&gt; sagemaker-gemma-execution-role &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--assume-role-policy-document&lt;/span&gt; &lt;span class="s1"&gt;'{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Principal":{"Service":"sagemaker.amazonaws.com"},"Action":"sts:AssumeRole"}]}'&lt;/span&gt;
aws iam attach-role-policy &lt;span class="nt"&gt;--role-name&lt;/span&gt; sagemaker-gemma-execution-role &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--policy-arn&lt;/span&gt; arn:aws:iam::aws:policy/AmazonSageMakerFullAccess
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ROLE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws iam get-role &lt;span class="nt"&gt;--role-name&lt;/span&gt; sagemaker-gemma-execution-role &lt;span class="nt"&gt;--query&lt;/span&gt; Role.Arn &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a second run &lt;code&gt;create-role&lt;/code&gt; reports &lt;code&gt;EntityAlreadyExists&lt;/code&gt;; &lt;code&gt;get-role&lt;/code&gt; still sets &lt;code&gt;ROLE&lt;/code&gt;. &lt;code&gt;deploy_endpoint&lt;/code&gt; creates the role only when it is missing.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — Create the Model
&lt;/h4&gt;

&lt;p&gt;A SageMaker model pairs an image with its settings. The container turns each &lt;code&gt;SM_VLLM_&lt;/code&gt; variable into the matching vLLM flag: &lt;code&gt;SM_VLLM_MAX_MODEL_LEN&lt;/code&gt; becomes &lt;code&gt;--max-model-len&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker create-model &lt;span class="nt"&gt;--model-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="nt"&gt;--execution-role-arn&lt;/span&gt; &lt;span class="nv"&gt;$ROLE&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--primary-container&lt;/span&gt; &lt;span class="s2"&gt;"{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Image&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="nv"&gt;$IMAGE&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Environment&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:{
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;SM_VLLM_MODEL&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL_ID&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;SM_VLLM_MAX_MODEL_LEN&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;8192&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;SM_VLLM_GPU_MEMORY_UTILIZATION&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;0.9&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;}}"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ModelArn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:sagemaker:us-east-2:&amp;lt;account-id&amp;gt;:model/gemma-4-e2b"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 5 — Create the Endpoint Config With Fallback Instance Types
&lt;/h4&gt;

&lt;p&gt;The endpoint config says where the model runs. &lt;code&gt;InstancePools&lt;/code&gt; lists up to five instance types in priority order, and SageMaker places the first one with a free instance. All three types below carry one L4, so the model sees the same GPU whichever is placed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker create-endpoint-config &lt;span class="nt"&gt;--endpoint-config-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--production-variants&lt;/span&gt; &lt;span class="s2"&gt;"[{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;VariantName&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;AllTraffic&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ModelName&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="nv"&gt;$NAME&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InitialInstanceCount&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:1,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InstancePools&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:[{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InstanceType&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ml.g6.xlarge&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Priority&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:1},
                       {&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InstanceType&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ml.g6.2xlarge&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Priority&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:2},
                       {&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;InstanceType&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ml.g6.4xlarge&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Priority&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:3}],
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ContainerStartupHealthCheckTimeoutInSeconds&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:1800,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ModelDataDownloadTimeoutInSeconds&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:1800}]"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"EndpointConfigArn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:sagemaker:us-east-2:&amp;lt;account-id&amp;gt;:endpoint-config/gemma-4-e2b"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two 1800-second timeouts give the container time to download the weights and compile before SageMaker's health check gives up.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: A Fallback List Holds Quota for Every Type in It
&lt;/h4&gt;

&lt;p&gt;With one Gemma endpoint running on &lt;code&gt;ml.g6.xlarge&lt;/code&gt; from the three-type list above, a second endpoint asking for &lt;code&gt;ml.g6.2xlarge&lt;/code&gt; in the same region was refused:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ResourceLimitExceeded: The account-level service limit 'ml.g6.2xlarge for endpoint usage' is 1 Instances, with current utilization of 1 Instances and a request delta of 1 Instances.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With a quota of 1 per type, a second endpoint goes in another region, or uses types outside the first endpoint's list.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 6 — Create the Endpoint
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker create-endpoint &lt;span class="nt"&gt;--endpoint-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="nt"&gt;--endpoint-config-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"EndpointArn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:sagemaker:us-east-2:&amp;lt;account-id&amp;gt;:endpoint/gemma-4-e2b"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Billing starts when an instance is placed. &lt;code&gt;aws sagemaker wait endpoint-in-service --endpoint-name $NAME&lt;/code&gt; blocks until it is ready; &lt;code&gt;get_endpoint_status&lt;/code&gt; reports the same state and the instance type that was placed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✅ `gemma-4-e2b` in `us-east-2`: **InService**
- Instance: `ml.g6.xlarge`
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The container log is in CloudWatch, and &lt;code&gt;get_endpoint_logs&lt;/code&gt; tails it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws logs &lt;span class="nb"&gt;tail&lt;/span&gt; /aws/sagemaker/Endpoints/&lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="nt"&gt;--follow&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start-up on the L4, in minutes after &lt;code&gt;create-endpoint&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phase&lt;/th&gt;
&lt;th&gt;Minutes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weights loaded, 9.75 GiB in 82.75 s&lt;/td&gt;
&lt;td&gt;7.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache sized, 723,484 tokens&lt;/td&gt;
&lt;td&gt;9.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;InService&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;9.9&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  🔎 Tip: A Creating Endpoint Cannot Be Deleted
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;create-endpoint&lt;/code&gt; cannot be undone until the endpoint settles. A &lt;code&gt;delete-endpoint&lt;/code&gt; sent while it is &lt;code&gt;Creating&lt;/code&gt; is refused:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;aws: [ERROR]: An error occurred (ValidationException) when calling the DeleteEndpoint operation: Cannot update in-progress endpoint "arn:aws:sagemaker:us-east-2:&amp;lt;account-id&amp;gt;:endpoint/gemma-4-e2b".
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The endpoint finishes starting, bills from the moment its instance is placed, and can be deleted once it reaches &lt;code&gt;InService&lt;/code&gt; or &lt;code&gt;Failed&lt;/code&gt;. Check the model ID and instance types before Step 6.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: Capacity Is Separate From Quota
&lt;/h4&gt;

&lt;p&gt;A quota of 1 lets you request one instance; the region still has to have one free. In &lt;code&gt;us-east-1&lt;/code&gt;, two requests for L4 instances each stayed in &lt;code&gt;Creating&lt;/code&gt; for about 30 minutes and then failed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Unable to provision requested ML compute capacity due to InsufficientInstanceCapacity error. Please retry using a different ML instance type or after some time.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same request in &lt;code&gt;us-east-2&lt;/code&gt; placed an instance at once. While SageMaker waits for capacity, no container starts and the CloudWatch log group never appears, which is how &lt;code&gt;get_endpoint_logs&lt;/code&gt; tells a capacity wait from a slow model load. No charge accrues during the wait. When it fails, repeat Steps 4 to 6 in another region where Step 1 shows the same quota.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 7 — Cross Check the Deployed Model
&lt;/h4&gt;

&lt;p&gt;The vLLM container accepts an OpenAI chat body on &lt;code&gt;invoke-endpoint&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'{"messages":[{"role":"user","content":"Why is the sky blue?"}],"max_tokens":256}'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; req.json
aws sagemaker-runtime invoke-endpoint &lt;span class="nt"&gt;--endpoint-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--content-type&lt;/span&gt; application/json &lt;span class="nt"&gt;--body&lt;/span&gt; fileb://req.json out.json
jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.choices[0].message.content'&lt;/span&gt; out.json | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;
jq &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'{model,usage}'&lt;/span&gt; out.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{
    "ContentType": "application/json",
    "InvokedProductionVariant": "AllTraffic"
}
The sky is blue due to a phenomenon called **Rayleigh scattering**. This process is caused by how sunlight interacts with the Earth's atmosphere.
{"model":"google/gemma-4-E2B-it","usage":{"prompt_tokens":15,"total_tokens":271,"completion_tokens":256,"prompt_tokens_details":null,"completion_tokens_details":null}}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;verify_model_health&lt;/code&gt; sends one short request and checks for a reply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✅ model=`google/gemma-4-E2B-it` tokens=2 wall=0.749s reply='ok'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 8 — Teardown
&lt;/h4&gt;

&lt;p&gt;The endpoint bills by the hour until it is deleted. Deleting the config and the model as well leaves nothing behind. &lt;code&gt;delete_endpoint&lt;/code&gt; runs all three:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sagemaker delete-endpoint &lt;span class="nt"&gt;--endpoint-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt;
aws sagemaker delete-endpoint-config &lt;span class="nt"&gt;--endpoint-config-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt;
aws sagemaker delete-model &lt;span class="nt"&gt;--model-name&lt;/span&gt; &lt;span class="nv"&gt;$NAME&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each command prints nothing and exits 0. &lt;code&gt;list_endpoints&lt;/code&gt; confirms the account is clear:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;### 0 endpoint(s) matching `gemma`
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to deploy Gemma 4 E2B to an Amazon SageMaker real-time endpoint with the AWS CLI and manage it from an MCP server. The key to the solution was the AWS vLLM SageMaker container, which turns the deployment into three &lt;code&gt;create-&lt;/code&gt; calls and makes the endpoint answer OpenAI-style chat requests. The deployment results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 Eight CLI steps take an account from quota check to a serving endpoint and back to nothing&lt;/li&gt;
&lt;li&gt;🟢 The endpoint reached &lt;code&gt;InService&lt;/code&gt; 9.9 minutes after &lt;code&gt;create-endpoint&lt;/code&gt;, with the weights using 9.75 GiB of the L4's 24 GB&lt;/li&gt;
&lt;li&gt;🟢 The MCP server runs every step from Claude Code, marks each tool read-only, write or destructive, and passes its tests offline&lt;/li&gt;
&lt;li&gt;🟢 &lt;code&gt;InstancePools&lt;/code&gt; gives one endpoint config several fallback instance types&lt;/li&gt;
&lt;li&gt;⚠️ L4 capacity varied by region: &lt;code&gt;us-east-1&lt;/code&gt; refused twice, about 30 minutes each, while &lt;code&gt;us-east-2&lt;/code&gt; placed an instance at once&lt;/li&gt;
&lt;li&gt;⚠️ A fallback list holds quota for every type in it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: one account, &lt;code&gt;ml.g6.xlarge&lt;/code&gt; with one NVIDIA L4, vLLM 0.30.0 from the AWS container, &lt;code&gt;google/gemma-4-E2B-it&lt;/code&gt; at full precision, deployed in &lt;code&gt;us-east-2&lt;/code&gt; on 2026-09-25. Start-up times are from a single deployment.&lt;/p&gt;

&lt;p&gt;The strategy for using MCP for SageMaker deployment was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Repository: &lt;a href="https://github.com/xbill9/sagemaker-gemma" rel="noopener noreferrer"&gt;https://github.com/xbill9/sagemaker-gemma&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 E2B on Hugging Face: &lt;a href="https://huggingface.co/google/gemma-4-E2B-it" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-E2B-it&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;AWS Deep Learning Containers: &lt;a href="https://github.com/aws/deep-learning-containers" rel="noopener noreferrer"&gt;https://github.com/aws/deep-learning-containers&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;SageMaker real-time inference: &lt;a href="https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;CreateEndpointConfig API: &lt;a href="https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_CreateEndpointConfig.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_CreateEndpointConfig.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MCP Python SDK: &lt;a href="https://github.com/modelcontextprotocol/python-sdk" rel="noopener noreferrer"&gt;https://github.com/modelcontextprotocol/python-sdk&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;vLLM: &lt;a href="https://docs.vllm.ai/" rel="noopener noreferrer"&gt;https://docs.vllm.ai/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aws</category>
      <category>sagemaker</category>
      <category>gemma</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Two Iceberg Clients, One Protocol: Where the Time Goes</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Fri, 25 Sep 2026 16:40:44 +0000</pubDate>
      <link>https://dev.to/gde/two-iceberg-clients-one-protocol-where-the-time-goes-4g88</link>
      <guid>https://dev.to/gde/two-iceberg-clients-one-protocol-where-the-time-goes-4g88</guid>
      <description>&lt;p&gt;This article provides a step by step comparison of the Rust and Python clients for Apache Iceberg REST catalogs. It times both clients on the same operations against the same tables, then breaks one request down to see where the time goes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/lakehouse-iceberg-2026" rel="noopener noreferrer"&gt;https://github.com/xbill9/lakehouse-iceberg-2026&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  What is this project trying to Do?
&lt;/h4&gt;

&lt;p&gt;Two Apache clients talk to the same catalogs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1&lt;/strong&gt;, the Rust client, built in release mode&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;pyiceberg&lt;/code&gt; 0.12.0&lt;/strong&gt;, the Python client, on Python 3.14.7&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They run against three catalogs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Apache Polaris&lt;/strong&gt; 1.7.0 in Docker on the same machine, so there is no network time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google BigLake&lt;/strong&gt; and &lt;strong&gt;Microsoft OneLake&lt;/strong&gt; over the internet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three things are measured: how long each call takes, how long each client takes to start, and where the time in one call is spent.&lt;/p&gt;

&lt;p&gt;The benchmark measures speed only. It does not check that answers are correct, and a client that lacks an operation you need is the wrong choice however fast it is.&lt;/p&gt;




&lt;h4&gt;
  
  
  Why Not Just Count Features?
&lt;/h4&gt;

&lt;p&gt;The Rust client supports 13 of the 25 catalog endpoints tested, and &lt;code&gt;pyiceberg&lt;/code&gt; supports 21. Anyone can count that from the two repositories, and it only holds for these two versions.&lt;/p&gt;

&lt;p&gt;How long a call takes is in neither repository, so that is what this article measures.&lt;/p&gt;




&lt;h4&gt;
  
  
  Where do I start?
&lt;/h4&gt;

&lt;p&gt;The strategy for comparing the two clients is an incremental step by step approach.&lt;/p&gt;

&lt;p&gt;First, both clients are timed against a local catalog, where only the client's own time is measured. Then one request is broken down to find where the time goes. The same benchmark then runs against two managed catalogs, the two clients' requests are compared, and startup time is measured last.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Rust 1.94 or newer, and a &lt;strong&gt;release&lt;/strong&gt; build — this run used &lt;code&gt;rustc&lt;/code&gt; 1.98.1&lt;/li&gt;
&lt;li&gt;Python 3.10+ with &lt;code&gt;pyiceberg&lt;/code&gt; 0.12.0&lt;/li&gt;
&lt;li&gt;Docker, for the local Polaris catalog&lt;/li&gt;
&lt;li&gt;A quiet machine — this one is a 16-core Linux host&lt;/li&gt;
&lt;li&gt;Optional: a managed Iceberg REST catalog, for the internet runs&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Step 1 — Build Both Clients and Start Polaris
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;git clone https://github.com/xbill9/lakehouse-iceberg-2026
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;lakehouse-iceberg-2026/iceberg-conformance &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; ./polaris-up.sh
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;POLARIS_CLIENT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;root &lt;span class="nv"&gt;POLARIS_CLIENT_SECRET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;s3cr3t
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ../iceberg-rust-client &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; cargo build &lt;span class="nt"&gt;--release&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The benchmark will not save results from a debug build.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 2 — Time Both Clients Locally
&lt;/h4&gt;

&lt;p&gt;To keep the comparison fair:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Each client runs in its own new process.&lt;/strong&gt; The benchmark script is Python, so running &lt;code&gt;pyiceberg&lt;/code&gt; inside it would give Python a head start.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Startup is timed separately.&lt;/strong&gt; Each process connects and warms up before any call is timed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The clients take turns.&lt;/strong&gt; They alternate going first, so a machine that slows down slows both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every timing is saved&lt;/strong&gt;, and the medians are computed from the saved files.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench.py &lt;span class="nt"&gt;--rounds&lt;/span&gt; 4 &lt;span class="nt"&gt;--iters&lt;/span&gt; 15
&lt;span class="go"&gt;apache-polaris     6 ops, 2 clients, 4 rounds x 15 iters = 720 samples
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench_report.py
&lt;span class="go"&gt;wrote evidence/bench-comparison.txt   8 run(s) across 3 catalog(s)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six read operations both clients support, median time in microseconds, lowest to highest across four runs on two days:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;op&lt;/th&gt;
&lt;th&gt;rust&lt;/th&gt;
&lt;th&gt;pyiceberg&lt;/th&gt;
&lt;th&gt;python ÷ rust&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;list_namespaces&lt;/td&gt;
&lt;td&gt;301.3–768.1&lt;/td&gt;
&lt;td&gt;883.2–1479.4&lt;/td&gt;
&lt;td&gt;1.90x–4.29x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;load_namespace&lt;/td&gt;
&lt;td&gt;312.2–699.3&lt;/td&gt;
&lt;td&gt;784.7–1451.2&lt;/td&gt;
&lt;td&gt;2.08x–3.52x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;head_namespace&lt;/td&gt;
&lt;td&gt;328.4–584.8&lt;/td&gt;
&lt;td&gt;1048.6–1362.7&lt;/td&gt;
&lt;td&gt;2.33x–3.19x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;list_tables&lt;/td&gt;
&lt;td&gt;327.0–552.1&lt;/td&gt;
&lt;td&gt;1003.2–1551.3&lt;/td&gt;
&lt;td&gt;2.56x–3.53x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;load_table&lt;/td&gt;
&lt;td&gt;4634.1–5764.4&lt;/td&gt;
&lt;td&gt;6405.6–7344.8&lt;/td&gt;
&lt;td&gt;1.23x–1.38x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;head_table&lt;/td&gt;
&lt;td&gt;4266.1–5298.0&lt;/td&gt;
&lt;td&gt;5167.9–6155.4&lt;/td&gt;
&lt;td&gt;1.14x–1.21x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rust is 1.90x to 4.29x faster on the four small calls, and 1.14x to 1.38x on the two table calls&lt;/strong&gt;, where the server itself does more work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The numbers vary a lot between runs.&lt;/strong&gt; The Rust time for &lt;code&gt;list_namespaces&lt;/code&gt; ranged from 301.3 to 768.1 microseconds, a spread of 154.9%. That is why the table shows ranges.&lt;/p&gt;

&lt;p&gt;With no network involved, these ratios are as large as the difference can get. Step 4 adds the internet.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 3 — Break One Request Down
&lt;/h4&gt;

&lt;p&gt;The same &lt;code&gt;list_namespaces&lt;/code&gt; request is timed in stages on one &lt;code&gt;pyiceberg&lt;/code&gt; connection: the raw HTTP request, then parsing the JSON, then building Python objects, then the full client call. The same raw HTTP request is then timed from Rust.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench_breakdown.py &lt;span class="nt"&gt;--iters&lt;/span&gt; 120
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Across seven runs on the local catalog, over two days:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The raw HTTP request is 90% to 94% of &lt;code&gt;pyiceberg&lt;/code&gt;'s call time.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;The same HTTP request takes &lt;strong&gt;269.8 to 479.7 microseconds&lt;/strong&gt; from Rust's HTTP library, &lt;code&gt;reqwest&lt;/code&gt;, and &lt;strong&gt;864.3 to 1286.4 microseconds&lt;/strong&gt; from Python's, &lt;code&gt;requests&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Everything &lt;code&gt;pyiceberg&lt;/code&gt; does after the HTTP request adds &lt;strong&gt;57.3 to 134.7 microseconds&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So 85% to 90% of the difference between the two clients is the HTTP library. That percentage is calculated from the two ranges above, run by run.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: Small Differences Are Noise Here
&lt;/h4&gt;

&lt;p&gt;The breakdown tool prints, next to each stage, whether the time that stage added is bigger than the normal run-to-run noise. Only JSON parsing was. How the rest of &lt;code&gt;pyiceberg&lt;/code&gt;'s 57.3 to 134.7 microseconds splits up cannot be measured at this sample size.&lt;/p&gt;

&lt;p&gt;One setting is worth knowing about. &lt;code&gt;requests&lt;/code&gt; re-reads proxy settings from the environment on every call unless &lt;code&gt;trust_env=False&lt;/code&gt; is set. Turning it off saved between 32.9 and 89.9 microseconds in each of five runs, but that saving was bigger than the noise in only one run.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — Time Them Over the Internet
&lt;/h4&gt;

&lt;p&gt;The same benchmark against BigLake and OneLake. Both clients use the same login token, created once before the run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench.py &lt;span class="nt"&gt;--only&lt;/span&gt; google-lakehouse &lt;span class="nt"&gt;--rounds&lt;/span&gt; 4 &lt;span class="nt"&gt;--iters&lt;/span&gt; 15 &lt;span class="nt"&gt;--storage&lt;/span&gt; opendal
&lt;span class="go"&gt;google-lakehouse   6 ops, 2 clients, 4 rounds x 15 iters = 720 samples
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench.py &lt;span class="nt"&gt;--only&lt;/span&gt; microsoft-onelake &lt;span class="nt"&gt;--rounds&lt;/span&gt; 4 &lt;span class="nt"&gt;--iters&lt;/span&gt; 15 &lt;span class="nt"&gt;--storage&lt;/span&gt; opendal
&lt;span class="go"&gt;microsoft-onelake  6 ops, 2 clients, 4 rounds x 15 iters = 720 samples
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two runs each, median time in milliseconds:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;op&lt;/th&gt;
&lt;th&gt;BigLake rust&lt;/th&gt;
&lt;th&gt;BigLake py&lt;/th&gt;
&lt;th&gt;ratio&lt;/th&gt;
&lt;th&gt;OneLake rust&lt;/th&gt;
&lt;th&gt;OneLake py&lt;/th&gt;
&lt;th&gt;ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;list_namespaces&lt;/td&gt;
&lt;td&gt;102.4–103.1&lt;/td&gt;
&lt;td&gt;92.2–106.7&lt;/td&gt;
&lt;td&gt;0.90x–1.04x&lt;/td&gt;
&lt;td&gt;36.0–36.4&lt;/td&gt;
&lt;td&gt;35.9–36.0&lt;/td&gt;
&lt;td&gt;0.99x–1.00x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;load_namespace&lt;/td&gt;
&lt;td&gt;199.4–201.6&lt;/td&gt;
&lt;td&gt;199.1–200.9&lt;/td&gt;
&lt;td&gt;0.99x–1.01x&lt;/td&gt;
&lt;td&gt;44.6–45.1&lt;/td&gt;
&lt;td&gt;45.3–46.3&lt;/td&gt;
&lt;td&gt;1.02x–1.03x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;head_namespace&lt;/td&gt;
&lt;td&gt;198.4–200.4&lt;/td&gt;
&lt;td&gt;199.2–200.0&lt;/td&gt;
&lt;td&gt;0.99x–1.01x&lt;/td&gt;
&lt;td&gt;44.8–45.9&lt;/td&gt;
&lt;td&gt;45.3–46.6&lt;/td&gt;
&lt;td&gt;0.99x–1.04x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;list_tables&lt;/td&gt;
&lt;td&gt;202.7–205.5&lt;/td&gt;
&lt;td&gt;203.2–206.1&lt;/td&gt;
&lt;td&gt;1.00x&lt;/td&gt;
&lt;td&gt;45.0–46.3&lt;/td&gt;
&lt;td&gt;45.4–45.6&lt;/td&gt;
&lt;td&gt;0.98x–1.01x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;load_table&lt;/td&gt;
&lt;td&gt;296.4–297.1&lt;/td&gt;
&lt;td&gt;281.9–284.3&lt;/td&gt;
&lt;td&gt;0.95x–0.96x&lt;/td&gt;
&lt;td&gt;36.3–39.0&lt;/td&gt;
&lt;td&gt;160.7–167.7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4.12x–4.62x&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;head_table&lt;/td&gt;
&lt;td&gt;201.7–205.0&lt;/td&gt;
&lt;td&gt;199.9–201.4&lt;/td&gt;
&lt;td&gt;0.98x–0.99x&lt;/td&gt;
&lt;td&gt;53.2–54.4&lt;/td&gt;
&lt;td&gt;53.4–53.8&lt;/td&gt;
&lt;td&gt;0.98x–1.01x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Apart from &lt;code&gt;load_table&lt;/code&gt;, every ratio is between 0.90x and 1.04x. The local difference is under a millisecond per call, and a round trip of 36 to 200 milliseconds hides it.&lt;/p&gt;

&lt;p&gt;On BigLake, &lt;code&gt;pyiceberg&lt;/code&gt; was 4% to 5% faster on &lt;code&gt;load_table&lt;/code&gt; in both runs, for a reason this test did not identify.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: BigLake Allows 75 Catalog Requests a Minute
&lt;/h4&gt;

&lt;p&gt;BigLake limits catalog requests to 75 a minute per project, and answers faster bursts with HTTP 429. The benchmark discards any timing that did not return success, and &lt;code&gt;bench_breakdown.py --pace-ms&lt;/code&gt; waits between requests. The limit is visible with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;gcloud alpha services quota list &lt;span class="nt"&gt;--service&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;biglake.googleapis.com &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="gp"&gt;    --consumer=projects/&amp;lt;project&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="go"&gt;    --format="table(metric,consumerQuotaLimits[0].unit,consumerQuotaLimits[0].quotaBuckets[0].effectiveLimit)"
biglake.googleapis.com/irc_catalog_requests          1/min/{project}       75
biglake.googleapis.com/irc_read_requests             1/min/{project}       600
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 5 — Compare the Two Clients' Requests
&lt;/h4&gt;

&lt;p&gt;OneLake's &lt;code&gt;load_table&lt;/code&gt; took 4.12x to 4.62x longer through &lt;code&gt;pyiceberg&lt;/code&gt;, while every other OneLake call was even. The two clients send different requests.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pyiceberg&lt;/code&gt; sends the header &lt;code&gt;X-Iceberg-Access-Delegation: vended-credentials&lt;/code&gt; by default (&lt;code&gt;catalog/rest/__init__.py:881&lt;/code&gt;). It asks the catalog to include a temporary storage credential with the table. The Rust client sends no such header.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;bench_delegation.py&lt;/code&gt; times the same &lt;code&gt;pyiceberg&lt;/code&gt; request with and without the header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench_delegation.py &lt;span class="nt"&gt;--only&lt;/span&gt; microsoft-onelake &lt;span class="nt"&gt;--iters&lt;/span&gt; 40
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;row                                  p50       p90
GET, header sent                   146.6     235.5
GET, header removed                 38.1      41.3
whole load_table                   148.5     182.4
whole load_table, no header         36.6      42.4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With the header, OneLake adds a storage credential to the response and takes 108.5 ms longer. Without it, &lt;code&gt;pyiceberg&lt;/code&gt;'s &lt;code&gt;load_table&lt;/code&gt; takes 36.6 ms, inside the Rust client's 36.3–39.0 ms. On Polaris and BigLake the header made no measurable difference.&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;pyiceberg&lt;/code&gt; pays 108.5 to 111.9 ms and gets a credential for reading the table's files. The Rust client pays nothing and gets no credential. The companion article on the Rust client covers what happens when it then tries to read the files.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 6 — Measure Startup Time
&lt;/h4&gt;

&lt;p&gt;Startup here means starting a new process, connecting, and answering each of the six operations once. Median of five starts per client per run:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;catalog&lt;/th&gt;
&lt;th&gt;rust&lt;/th&gt;
&lt;th&gt;pyiceberg&lt;/th&gt;
&lt;th&gt;python takes longer by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Polaris, local, 4 runs&lt;/td&gt;
&lt;td&gt;16.4–26.0 ms&lt;/td&gt;
&lt;td&gt;525.8–557.9 ms&lt;/td&gt;
&lt;td&gt;507.2–535.1 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BigLake, 2 runs&lt;/td&gt;
&lt;td&gt;1376.3–1410.3 ms&lt;/td&gt;
&lt;td&gt;1911.8–2016.7 ms&lt;/td&gt;
&lt;td&gt;535.5–606.4 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OneLake, 2 runs&lt;/td&gt;
&lt;td&gt;955.9–996.2 ms&lt;/td&gt;
&lt;td&gt;1723.5–1772.4 ms&lt;/td&gt;
&lt;td&gt;767.6–776.2 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Over the internet both clients wait on the same round trips, so the ratio drops from 21.4x–33.6x locally to 1.4x on BigLake and 1.8x on OneLake. &lt;strong&gt;The extra half second stays.&lt;/strong&gt; It matters for a CLI, a Lambda function or a short-lived agent, and hardly at all for a service that starts once.&lt;/p&gt;

&lt;p&gt;Where the half second goes, median of five, across seven local runs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;start Python, &lt;code&gt;python3 -c pass&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;9.0–11.2 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;import pyiceberg.catalog.rest&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a further 337.9–366.3 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;connect: fetch the catalog config and log in&lt;/td&gt;
&lt;td&gt;6.8–12.4 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most of it is importing the library. Connecting to the catalog is the smallest part.&lt;/p&gt;

&lt;p&gt;On OneLake the import took 338.9 and 341.4 ms, the same as locally, and fetching the config took 447.3 and 493.8 ms. Both clients fetch the config; only &lt;code&gt;pyiceberg&lt;/code&gt; pays for the import.&lt;/p&gt;




&lt;h4&gt;
  
  
  How This Compares to Published Lambda Numbers
&lt;/h4&gt;

&lt;p&gt;&lt;em&gt;Cold Starts Are Dead&lt;/em&gt; measured AWS Lambda start times of &lt;strong&gt;88.3 ms for Python 3.13 on arm64&lt;/strong&gt; (106.2 ms on x86_64) and &lt;strong&gt;14.1 ms for Rust&lt;/strong&gt; (17.0 ms), at 512 MB, with hello-world functions and &lt;strong&gt;no libraries&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The local Python startup here is about six times that: 525.8 to 557.9 ms against 88.3 ms. The difference is mostly the 337.9 to 366.3 ms spent importing &lt;code&gt;pyiceberg&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Compare and Contrast
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;🦀 iceberg-catalog-rest&lt;/th&gt;
&lt;th&gt;🐍 pyiceberg&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;list_namespaces&lt;/code&gt;, local, µs&lt;/td&gt;
&lt;td&gt;🥇 301.3–768.1&lt;/td&gt;
&lt;td&gt;883.2–1479.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small calls over the internet&lt;/td&gt;
&lt;td&gt;tie, 0.90x–1.04x&lt;/td&gt;
&lt;td&gt;tie&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTTP library&lt;/td&gt;
&lt;td&gt;🥇 &lt;code&gt;reqwest&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;requests&lt;/code&gt;, 85%–90% of the difference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extra startup time&lt;/td&gt;
&lt;td&gt;🥇 —&lt;/td&gt;
&lt;td&gt;507.2–776.2 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of which, importing the library&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;337.9–366.3 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;load_table&lt;/code&gt; on OneLake, ms&lt;/td&gt;
&lt;td&gt;🥇 36.3–39.0, no credential&lt;/td&gt;
&lt;td&gt;160.7–167.7, with a storage credential&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requests storage credentials&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;🥇 yes, by default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoints supported&lt;/td&gt;
&lt;td&gt;13 of 25&lt;/td&gt;
&lt;td&gt;🥇 21 of 25&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  So, Which One?
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;For a long-running service, pick by features.&lt;/strong&gt; Over the internet the per-call difference disappears, and startup happens once. Choose the client that supports the operations and login method you need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For a CLI, a Lambda function or a short-lived agent, startup decides it.&lt;/strong&gt; Python adds half a second or more per process on every catalog tested.&lt;/p&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to measure how fast the Rust and Python Iceberg REST clients are. The key to the solution was running both as separate processes against the same catalogs, timing startup separately, and breaking one request down to see where the time goes. The results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 Locally, Rust is 1.90x to 4.29x faster on small calls, and 85% to 90% of that difference is the HTTP library&lt;/li&gt;
&lt;li&gt;🟢 Over the internet the difference disappears: 0.90x to 1.04x on BigLake and OneLake&lt;/li&gt;
&lt;li&gt;⚠️ Python starts 507.2 to 776.2 ms slower on every catalog, mostly from importing the library (337.9 to 366.3 ms locally)&lt;/li&gt;
&lt;li&gt;⚠️ OneLake's 4x on &lt;code&gt;load_table&lt;/code&gt; comes from &lt;code&gt;pyiceberg&lt;/code&gt; requesting a storage credential by default, which the Rust client never does&lt;/li&gt;
&lt;li&gt;❌ &lt;code&gt;pyiceberg&lt;/code&gt; was 4% to 5% faster on BigLake's &lt;code&gt;load_table&lt;/code&gt;, cause unknown&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: &lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1 in a release build with &lt;code&gt;rustc&lt;/code&gt; 1.98.1 and &lt;code&gt;rustls&lt;/code&gt; 0.23.45, and &lt;code&gt;pyiceberg&lt;/code&gt; 0.12.0 on Python 3.14.7 with &lt;code&gt;requests&lt;/code&gt; 2.34.2 over OpenSSL 3.5.7, on one 16-core Linux host. Apache Polaris 1.7.0 in Docker on the same machine with local file storage: four benchmark runs on 2026-09-17 and 2026-09-18, and seven breakdown runs of 80 to 120 requests each. Google BigLake and Microsoft OneLake over the internet on 2026-09-18: two benchmark runs each. Each benchmark run is 720 timed calls across six read operations; startup is the median of five starts per client per run. No writes were timed. Managed catalogs do not report a version, and each internet run is one region at one point in time.&lt;/p&gt;

&lt;p&gt;The strategy for comparing two Iceberg REST clients by speed was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/xbill9/lakehouse-iceberg-2026" rel="noopener noreferrer"&gt;lakehouse-iceberg-2026 | GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust" rel="noopener noreferrer"&gt;apache/iceberg-rust&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-python" rel="noopener noreferrer"&gt;apache/iceberg-python&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/aws/cold-starts-are-dead-5fod"&gt;Cold Starts Are Dead&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/querygraph/catalog-bench" rel="noopener noreferrer"&gt;querygraph/catalog-bench&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/gde/seven-iceberg-rest-catalogs-what-they-declare-and-what-they-serve-40oj"&gt;Seven Iceberg REST Catalogs: What They Declare, and What They Serve&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>rust</category>
      <category>python</category>
      <category>iceberg</category>
      <category>performance</category>
    </item>
    <item>
      <title>Two Iceberg Clients, One Protocol: Where the Time Goes</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Fri, 25 Sep 2026 16:40:41 +0000</pubDate>
      <link>https://dev.to/aws-builders/two-iceberg-clients-one-protocol-where-the-time-goes-4a2c</link>
      <guid>https://dev.to/aws-builders/two-iceberg-clients-one-protocol-where-the-time-goes-4a2c</guid>
      <description>&lt;p&gt;This article provides a step by step comparison of the Rust and Python clients for Apache Iceberg REST catalogs. It times both clients on the same operations against the same tables, then breaks one request down to see where the time goes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/lakehouse-iceberg-2026" rel="noopener noreferrer"&gt;https://github.com/xbill9/lakehouse-iceberg-2026&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  What is this project trying to Do?
&lt;/h4&gt;

&lt;p&gt;Two Apache clients talk to the same catalogs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1&lt;/strong&gt;, the Rust client, built in release mode&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;pyiceberg&lt;/code&gt; 0.12.0&lt;/strong&gt;, the Python client, on Python 3.14.7&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They run against three catalogs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Apache Polaris&lt;/strong&gt; 1.7.0 in Docker on the same machine, so there is no network time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google BigLake&lt;/strong&gt; and &lt;strong&gt;Microsoft OneLake&lt;/strong&gt; over the internet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three things are measured: how long each call takes, how long each client takes to start, and where the time in one call is spent.&lt;/p&gt;

&lt;p&gt;The benchmark measures speed only. It does not check that answers are correct, and a client that lacks an operation you need is the wrong choice however fast it is.&lt;/p&gt;




&lt;h4&gt;
  
  
  Why Not Just Count Features?
&lt;/h4&gt;

&lt;p&gt;The Rust client supports 13 of the 25 catalog endpoints tested, and &lt;code&gt;pyiceberg&lt;/code&gt; supports 21. Anyone can count that from the two repositories, and it only holds for these two versions.&lt;/p&gt;

&lt;p&gt;How long a call takes is in neither repository, so that is what this article measures.&lt;/p&gt;




&lt;h4&gt;
  
  
  Where do I start?
&lt;/h4&gt;

&lt;p&gt;The strategy for comparing the two clients is an incremental step by step approach.&lt;/p&gt;

&lt;p&gt;First, both clients are timed against a local catalog, where only the client's own time is measured. Then one request is broken down to find where the time goes. The same benchmark then runs against two managed catalogs, the two clients' requests are compared, and startup time is measured last.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Rust 1.94 or newer, and a &lt;strong&gt;release&lt;/strong&gt; build — this run used &lt;code&gt;rustc&lt;/code&gt; 1.98.1&lt;/li&gt;
&lt;li&gt;Python 3.10+ with &lt;code&gt;pyiceberg&lt;/code&gt; 0.12.0&lt;/li&gt;
&lt;li&gt;Docker, for the local Polaris catalog&lt;/li&gt;
&lt;li&gt;A quiet machine — this one is a 16-core Linux host&lt;/li&gt;
&lt;li&gt;Optional: a managed Iceberg REST catalog, for the internet runs&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Step 1 — Build Both Clients and Start Polaris
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;git clone https://github.com/xbill9/lakehouse-iceberg-2026
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;lakehouse-iceberg-2026/iceberg-conformance &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; ./polaris-up.sh
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;POLARIS_CLIENT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;root &lt;span class="nv"&gt;POLARIS_CLIENT_SECRET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;s3cr3t
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ../iceberg-rust-client &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; cargo build &lt;span class="nt"&gt;--release&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The benchmark will not save results from a debug build.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 2 — Time Both Clients Locally
&lt;/h4&gt;

&lt;p&gt;To keep the comparison fair:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Each client runs in its own new process.&lt;/strong&gt; The benchmark script is Python, so running &lt;code&gt;pyiceberg&lt;/code&gt; inside it would give Python a head start.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Startup is timed separately.&lt;/strong&gt; Each process connects and warms up before any call is timed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The clients take turns.&lt;/strong&gt; They alternate going first, so a machine that slows down slows both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every timing is saved&lt;/strong&gt;, and the medians are computed from the saved files.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench.py &lt;span class="nt"&gt;--rounds&lt;/span&gt; 4 &lt;span class="nt"&gt;--iters&lt;/span&gt; 15
&lt;span class="go"&gt;apache-polaris     6 ops, 2 clients, 4 rounds x 15 iters = 720 samples
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench_report.py
&lt;span class="go"&gt;wrote evidence/bench-comparison.txt   8 run(s) across 3 catalog(s)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six read operations both clients support, median time in microseconds, lowest to highest across four runs on two days:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;op&lt;/th&gt;
&lt;th&gt;rust&lt;/th&gt;
&lt;th&gt;pyiceberg&lt;/th&gt;
&lt;th&gt;python ÷ rust&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;list_namespaces&lt;/td&gt;
&lt;td&gt;301.3–768.1&lt;/td&gt;
&lt;td&gt;883.2–1479.4&lt;/td&gt;
&lt;td&gt;1.90x–4.29x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;load_namespace&lt;/td&gt;
&lt;td&gt;312.2–699.3&lt;/td&gt;
&lt;td&gt;784.7–1451.2&lt;/td&gt;
&lt;td&gt;2.08x–3.52x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;head_namespace&lt;/td&gt;
&lt;td&gt;328.4–584.8&lt;/td&gt;
&lt;td&gt;1048.6–1362.7&lt;/td&gt;
&lt;td&gt;2.33x–3.19x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;list_tables&lt;/td&gt;
&lt;td&gt;327.0–552.1&lt;/td&gt;
&lt;td&gt;1003.2–1551.3&lt;/td&gt;
&lt;td&gt;2.56x–3.53x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;load_table&lt;/td&gt;
&lt;td&gt;4634.1–5764.4&lt;/td&gt;
&lt;td&gt;6405.6–7344.8&lt;/td&gt;
&lt;td&gt;1.23x–1.38x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;head_table&lt;/td&gt;
&lt;td&gt;4266.1–5298.0&lt;/td&gt;
&lt;td&gt;5167.9–6155.4&lt;/td&gt;
&lt;td&gt;1.14x–1.21x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rust is 1.90x to 4.29x faster on the four small calls, and 1.14x to 1.38x on the two table calls&lt;/strong&gt;, where the server itself does more work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The numbers vary a lot between runs.&lt;/strong&gt; The Rust time for &lt;code&gt;list_namespaces&lt;/code&gt; ranged from 301.3 to 768.1 microseconds, a spread of 154.9%. That is why the table shows ranges.&lt;/p&gt;

&lt;p&gt;With no network involved, these ratios are as large as the difference can get. Step 4 adds the internet.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 3 — Break One Request Down
&lt;/h4&gt;

&lt;p&gt;The same &lt;code&gt;list_namespaces&lt;/code&gt; request is timed in stages on one &lt;code&gt;pyiceberg&lt;/code&gt; connection: the raw HTTP request, then parsing the JSON, then building Python objects, then the full client call. The same raw HTTP request is then timed from Rust.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench_breakdown.py &lt;span class="nt"&gt;--iters&lt;/span&gt; 120
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Across seven runs on the local catalog, over two days:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The raw HTTP request is 90% to 94% of &lt;code&gt;pyiceberg&lt;/code&gt;'s call time.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;The same HTTP request takes &lt;strong&gt;269.8 to 479.7 microseconds&lt;/strong&gt; from Rust's HTTP library, &lt;code&gt;reqwest&lt;/code&gt;, and &lt;strong&gt;864.3 to 1286.4 microseconds&lt;/strong&gt; from Python's, &lt;code&gt;requests&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Everything &lt;code&gt;pyiceberg&lt;/code&gt; does after the HTTP request adds &lt;strong&gt;57.3 to 134.7 microseconds&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So 85% to 90% of the difference between the two clients is the HTTP library. That percentage is calculated from the two ranges above, run by run.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: Small Differences Are Noise Here
&lt;/h4&gt;

&lt;p&gt;The breakdown tool prints, next to each stage, whether the time that stage added is bigger than the normal run-to-run noise. Only JSON parsing was. How the rest of &lt;code&gt;pyiceberg&lt;/code&gt;'s 57.3 to 134.7 microseconds splits up cannot be measured at this sample size.&lt;/p&gt;

&lt;p&gt;One setting is worth knowing about. &lt;code&gt;requests&lt;/code&gt; re-reads proxy settings from the environment on every call unless &lt;code&gt;trust_env=False&lt;/code&gt; is set. Turning it off saved between 32.9 and 89.9 microseconds in each of five runs, but that saving was bigger than the noise in only one run.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — Time Them Over the Internet
&lt;/h4&gt;

&lt;p&gt;The same benchmark against BigLake and OneLake. Both clients use the same login token, created once before the run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench.py &lt;span class="nt"&gt;--only&lt;/span&gt; google-lakehouse &lt;span class="nt"&gt;--rounds&lt;/span&gt; 4 &lt;span class="nt"&gt;--iters&lt;/span&gt; 15 &lt;span class="nt"&gt;--storage&lt;/span&gt; opendal
&lt;span class="go"&gt;google-lakehouse   6 ops, 2 clients, 4 rounds x 15 iters = 720 samples
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench.py &lt;span class="nt"&gt;--only&lt;/span&gt; microsoft-onelake &lt;span class="nt"&gt;--rounds&lt;/span&gt; 4 &lt;span class="nt"&gt;--iters&lt;/span&gt; 15 &lt;span class="nt"&gt;--storage&lt;/span&gt; opendal
&lt;span class="go"&gt;microsoft-onelake  6 ops, 2 clients, 4 rounds x 15 iters = 720 samples
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two runs each, median time in milliseconds:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;op&lt;/th&gt;
&lt;th&gt;BigLake rust&lt;/th&gt;
&lt;th&gt;BigLake py&lt;/th&gt;
&lt;th&gt;ratio&lt;/th&gt;
&lt;th&gt;OneLake rust&lt;/th&gt;
&lt;th&gt;OneLake py&lt;/th&gt;
&lt;th&gt;ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;list_namespaces&lt;/td&gt;
&lt;td&gt;102.4–103.1&lt;/td&gt;
&lt;td&gt;92.2–106.7&lt;/td&gt;
&lt;td&gt;0.90x–1.04x&lt;/td&gt;
&lt;td&gt;36.0–36.4&lt;/td&gt;
&lt;td&gt;35.9–36.0&lt;/td&gt;
&lt;td&gt;0.99x–1.00x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;load_namespace&lt;/td&gt;
&lt;td&gt;199.4–201.6&lt;/td&gt;
&lt;td&gt;199.1–200.9&lt;/td&gt;
&lt;td&gt;0.99x–1.01x&lt;/td&gt;
&lt;td&gt;44.6–45.1&lt;/td&gt;
&lt;td&gt;45.3–46.3&lt;/td&gt;
&lt;td&gt;1.02x–1.03x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;head_namespace&lt;/td&gt;
&lt;td&gt;198.4–200.4&lt;/td&gt;
&lt;td&gt;199.2–200.0&lt;/td&gt;
&lt;td&gt;0.99x–1.01x&lt;/td&gt;
&lt;td&gt;44.8–45.9&lt;/td&gt;
&lt;td&gt;45.3–46.6&lt;/td&gt;
&lt;td&gt;0.99x–1.04x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;list_tables&lt;/td&gt;
&lt;td&gt;202.7–205.5&lt;/td&gt;
&lt;td&gt;203.2–206.1&lt;/td&gt;
&lt;td&gt;1.00x&lt;/td&gt;
&lt;td&gt;45.0–46.3&lt;/td&gt;
&lt;td&gt;45.4–45.6&lt;/td&gt;
&lt;td&gt;0.98x–1.01x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;load_table&lt;/td&gt;
&lt;td&gt;296.4–297.1&lt;/td&gt;
&lt;td&gt;281.9–284.3&lt;/td&gt;
&lt;td&gt;0.95x–0.96x&lt;/td&gt;
&lt;td&gt;36.3–39.0&lt;/td&gt;
&lt;td&gt;160.7–167.7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4.12x–4.62x&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;head_table&lt;/td&gt;
&lt;td&gt;201.7–205.0&lt;/td&gt;
&lt;td&gt;199.9–201.4&lt;/td&gt;
&lt;td&gt;0.98x–0.99x&lt;/td&gt;
&lt;td&gt;53.2–54.4&lt;/td&gt;
&lt;td&gt;53.4–53.8&lt;/td&gt;
&lt;td&gt;0.98x–1.01x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Apart from &lt;code&gt;load_table&lt;/code&gt;, every ratio is between 0.90x and 1.04x. The local difference is under a millisecond per call, and a round trip of 36 to 200 milliseconds hides it.&lt;/p&gt;

&lt;p&gt;On BigLake, &lt;code&gt;pyiceberg&lt;/code&gt; was 4% to 5% faster on &lt;code&gt;load_table&lt;/code&gt; in both runs, for a reason this test did not identify.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: BigLake Allows 75 Catalog Requests a Minute
&lt;/h4&gt;

&lt;p&gt;BigLake limits catalog requests to 75 a minute per project, and answers faster bursts with HTTP 429. The benchmark discards any timing that did not return success, and &lt;code&gt;bench_breakdown.py --pace-ms&lt;/code&gt; waits between requests. The limit is visible with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;gcloud alpha services quota list &lt;span class="nt"&gt;--service&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;biglake.googleapis.com &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="gp"&gt;    --consumer=projects/&amp;lt;project&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="go"&gt;    --format="table(metric,consumerQuotaLimits[0].unit,consumerQuotaLimits[0].quotaBuckets[0].effectiveLimit)"
biglake.googleapis.com/irc_catalog_requests          1/min/{project}       75
biglake.googleapis.com/irc_read_requests             1/min/{project}       600
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 5 — Compare the Two Clients' Requests
&lt;/h4&gt;

&lt;p&gt;OneLake's &lt;code&gt;load_table&lt;/code&gt; took 4.12x to 4.62x longer through &lt;code&gt;pyiceberg&lt;/code&gt;, while every other OneLake call was even. The two clients send different requests.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pyiceberg&lt;/code&gt; sends the header &lt;code&gt;X-Iceberg-Access-Delegation: vended-credentials&lt;/code&gt; by default (&lt;code&gt;catalog/rest/__init__.py:881&lt;/code&gt;). It asks the catalog to include a temporary storage credential with the table. The Rust client sends no such header.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;bench_delegation.py&lt;/code&gt; times the same &lt;code&gt;pyiceberg&lt;/code&gt; request with and without the header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 bench_delegation.py &lt;span class="nt"&gt;--only&lt;/span&gt; microsoft-onelake &lt;span class="nt"&gt;--iters&lt;/span&gt; 40
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;row                                  p50       p90
GET, header sent                   146.6     235.5
GET, header removed                 38.1      41.3
whole load_table                   148.5     182.4
whole load_table, no header         36.6      42.4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With the header, OneLake adds a storage credential to the response and takes 108.5 ms longer. Without it, &lt;code&gt;pyiceberg&lt;/code&gt;'s &lt;code&gt;load_table&lt;/code&gt; takes 36.6 ms, inside the Rust client's 36.3–39.0 ms. On Polaris and BigLake the header made no measurable difference.&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;pyiceberg&lt;/code&gt; pays 108.5 to 111.9 ms and gets a credential for reading the table's files. The Rust client pays nothing and gets no credential. The companion article on the Rust client covers what happens when it then tries to read the files.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 6 — Measure Startup Time
&lt;/h4&gt;

&lt;p&gt;Startup here means starting a new process, connecting, and answering each of the six operations once. Median of five starts per client per run:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;catalog&lt;/th&gt;
&lt;th&gt;rust&lt;/th&gt;
&lt;th&gt;pyiceberg&lt;/th&gt;
&lt;th&gt;python takes longer by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Polaris, local, 4 runs&lt;/td&gt;
&lt;td&gt;16.4–26.0 ms&lt;/td&gt;
&lt;td&gt;525.8–557.9 ms&lt;/td&gt;
&lt;td&gt;507.2–535.1 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BigLake, 2 runs&lt;/td&gt;
&lt;td&gt;1376.3–1410.3 ms&lt;/td&gt;
&lt;td&gt;1911.8–2016.7 ms&lt;/td&gt;
&lt;td&gt;535.5–606.4 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OneLake, 2 runs&lt;/td&gt;
&lt;td&gt;955.9–996.2 ms&lt;/td&gt;
&lt;td&gt;1723.5–1772.4 ms&lt;/td&gt;
&lt;td&gt;767.6–776.2 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Over the internet both clients wait on the same round trips, so the ratio drops from 21.4x–33.6x locally to 1.4x on BigLake and 1.8x on OneLake. &lt;strong&gt;The extra half second stays.&lt;/strong&gt; It matters for a CLI, a Lambda function or a short-lived agent, and hardly at all for a service that starts once.&lt;/p&gt;

&lt;p&gt;Where the half second goes, median of five, across seven local runs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;start Python, &lt;code&gt;python3 -c pass&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;9.0–11.2 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;import pyiceberg.catalog.rest&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a further 337.9–366.3 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;connect: fetch the catalog config and log in&lt;/td&gt;
&lt;td&gt;6.8–12.4 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most of it is importing the library. Connecting to the catalog is the smallest part.&lt;/p&gt;

&lt;p&gt;On OneLake the import took 338.9 and 341.4 ms, the same as locally, and fetching the config took 447.3 and 493.8 ms. Both clients fetch the config; only &lt;code&gt;pyiceberg&lt;/code&gt; pays for the import.&lt;/p&gt;




&lt;h4&gt;
  
  
  How This Compares to Published Lambda Numbers
&lt;/h4&gt;

&lt;p&gt;&lt;em&gt;Cold Starts Are Dead&lt;/em&gt; measured AWS Lambda start times of &lt;strong&gt;88.3 ms for Python 3.13 on arm64&lt;/strong&gt; (106.2 ms on x86_64) and &lt;strong&gt;14.1 ms for Rust&lt;/strong&gt; (17.0 ms), at 512 MB, with hello-world functions and &lt;strong&gt;no libraries&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The local Python startup here is about six times that: 525.8 to 557.9 ms against 88.3 ms. The difference is mostly the 337.9 to 366.3 ms spent importing &lt;code&gt;pyiceberg&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Compare and Contrast
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;🦀 iceberg-catalog-rest&lt;/th&gt;
&lt;th&gt;🐍 pyiceberg&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;list_namespaces&lt;/code&gt;, local, µs&lt;/td&gt;
&lt;td&gt;🥇 301.3–768.1&lt;/td&gt;
&lt;td&gt;883.2–1479.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small calls over the internet&lt;/td&gt;
&lt;td&gt;tie, 0.90x–1.04x&lt;/td&gt;
&lt;td&gt;tie&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTTP library&lt;/td&gt;
&lt;td&gt;🥇 &lt;code&gt;reqwest&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;requests&lt;/code&gt;, 85%–90% of the difference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extra startup time&lt;/td&gt;
&lt;td&gt;🥇 —&lt;/td&gt;
&lt;td&gt;507.2–776.2 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of which, importing the library&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;337.9–366.3 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;load_table&lt;/code&gt; on OneLake, ms&lt;/td&gt;
&lt;td&gt;🥇 36.3–39.0, no credential&lt;/td&gt;
&lt;td&gt;160.7–167.7, with a storage credential&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requests storage credentials&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;🥇 yes, by default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoints supported&lt;/td&gt;
&lt;td&gt;13 of 25&lt;/td&gt;
&lt;td&gt;🥇 21 of 25&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  So, Which One?
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;For a long-running service, pick by features.&lt;/strong&gt; Over the internet the per-call difference disappears, and startup happens once. Choose the client that supports the operations and login method you need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For a CLI, a Lambda function or a short-lived agent, startup decides it.&lt;/strong&gt; Python adds half a second or more per process on every catalog tested.&lt;/p&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to measure how fast the Rust and Python Iceberg REST clients are. The key to the solution was running both as separate processes against the same catalogs, timing startup separately, and breaking one request down to see where the time goes. The results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 Locally, Rust is 1.90x to 4.29x faster on small calls, and 85% to 90% of that difference is the HTTP library&lt;/li&gt;
&lt;li&gt;🟢 Over the internet the difference disappears: 0.90x to 1.04x on BigLake and OneLake&lt;/li&gt;
&lt;li&gt;⚠️ Python starts 507.2 to 776.2 ms slower on every catalog, mostly from importing the library (337.9 to 366.3 ms locally)&lt;/li&gt;
&lt;li&gt;⚠️ OneLake's 4x on &lt;code&gt;load_table&lt;/code&gt; comes from &lt;code&gt;pyiceberg&lt;/code&gt; requesting a storage credential by default, which the Rust client never does&lt;/li&gt;
&lt;li&gt;❌ &lt;code&gt;pyiceberg&lt;/code&gt; was 4% to 5% faster on BigLake's &lt;code&gt;load_table&lt;/code&gt;, cause unknown&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: &lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1 in a release build with &lt;code&gt;rustc&lt;/code&gt; 1.98.1 and &lt;code&gt;rustls&lt;/code&gt; 0.23.45, and &lt;code&gt;pyiceberg&lt;/code&gt; 0.12.0 on Python 3.14.7 with &lt;code&gt;requests&lt;/code&gt; 2.34.2 over OpenSSL 3.5.7, on one 16-core Linux host. Apache Polaris 1.7.0 in Docker on the same machine with local file storage: four benchmark runs on 2026-09-17 and 2026-09-18, and seven breakdown runs of 80 to 120 requests each. Google BigLake and Microsoft OneLake over the internet on 2026-09-18: two benchmark runs each. Each benchmark run is 720 timed calls across six read operations; startup is the median of five starts per client per run. No writes were timed. Managed catalogs do not report a version, and each internet run is one region at one point in time.&lt;/p&gt;

&lt;p&gt;The strategy for comparing two Iceberg REST clients by speed was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/xbill9/lakehouse-iceberg-2026" rel="noopener noreferrer"&gt;lakehouse-iceberg-2026 | GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust" rel="noopener noreferrer"&gt;apache/iceberg-rust&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-python" rel="noopener noreferrer"&gt;apache/iceberg-python&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/aws/cold-starts-are-dead-5fod"&gt;Cold Starts Are Dead&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/querygraph/catalog-bench" rel="noopener noreferrer"&gt;querygraph/catalog-bench&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/gde/seven-iceberg-rest-catalogs-what-they-declare-and-what-they-serve-40oj"&gt;Seven Iceberg REST Catalogs: What They Declare, and What They Serve&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>rust</category>
      <category>python</category>
      <category>iceberg</category>
      <category>performance</category>
    </item>
    <item>
      <title>What One Rust Client Can Reach Across Seven Iceberg Catalogs</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Fri, 25 Sep 2026 16:20:56 +0000</pubDate>
      <link>https://dev.to/gde/what-one-rust-client-can-reach-across-seven-iceberg-catalogs-22ll</link>
      <guid>https://dev.to/gde/what-one-rust-client-can-reach-across-seven-iceberg-catalogs-22ll</guid>
      <description>&lt;p&gt;This article provides a step by step guide to pointing the Apache Rust client for Iceberg REST catalogs at seven catalogs and recording what works. A Python script runs a small Rust program against each catalog and saves every result.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/lakehouse-iceberg-2026" rel="noopener noreferrer"&gt;https://github.com/xbill9/lakehouse-iceberg-2026&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  What is this project trying to Do?
&lt;/h4&gt;

&lt;p&gt;An earlier article tested what seven Iceberg REST catalogs support. This one asks from the other side: if you write a lakehouse tool in Rust today, will it work with the catalog you already pay for, and what do you need to set up? The client is &lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1 from the Apache Iceberg project, and the catalogs, tests and test table are the same as in that article:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Apache Polaris&lt;/strong&gt; 1.7.0, running locally as the reference&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google BigLake&lt;/strong&gt; and &lt;strong&gt;Microsoft OneLake&lt;/strong&gt;, over the internet&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AWS Glue&lt;/strong&gt; and &lt;strong&gt;AWS S3 Tables&lt;/strong&gt;, over the internet&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Databricks Unity&lt;/strong&gt; and &lt;strong&gt;Snowflake Horizon&lt;/strong&gt;, not run here&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result is a short one. This client implements 13 of the 25 operations the earlier article tested, it logs in to five of the seven catalogs, and on every catalog it logs in to, all 13 answer. Two lines of &lt;code&gt;Cargo.toml&lt;/code&gt; stand between it and those five. The other two are AWS, they require a signature this client cannot produce, and the Rust answer to that is two more catalog crates that do not implement quite the same operations.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Does Each Result Mean?
&lt;/h4&gt;

&lt;p&gt;The test suite is 33 checks covering 25 of the 35 operations in the Iceberg REST specification. Each check gets one of five results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;th&gt;meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;OK&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client sent the request and got an answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;FAILED&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client sent the request and got an error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IMPLICIT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client sends this request on its own, and there is no way to call it directly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;NOT-EXPRESSIBLE&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client has no method for this operation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;NOT-ISSUED&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client supports it, but it is a write, and that run only reads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  Where do I start?
&lt;/h4&gt;

&lt;p&gt;The strategy for testing the client is an incremental step by step approach.&lt;/p&gt;

&lt;p&gt;First, a local Polaris catalog is started and the client's source is read to list which operations it supports. The test program is then built and run against Polaris, reads first and writes after, then against the managed catalogs one login method at a time, and finally it tries to read a table's files.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Rust 1.94 or newer and &lt;code&gt;cargo&lt;/code&gt; — this run used &lt;code&gt;rustc&lt;/code&gt; 1.98.1&lt;/li&gt;
&lt;li&gt;Docker, for the local Polaris catalog&lt;/li&gt;
&lt;li&gt;Python 3.10+ with &lt;code&gt;pyyaml&lt;/code&gt;, for the test scripts&lt;/li&gt;
&lt;li&gt;Optional: logins for any managed catalog — &lt;code&gt;gcloud&lt;/code&gt; for BigLake, &lt;code&gt;az&lt;/code&gt; for OneLake, the AWS CLI for Glue and S3 Tables&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Step 1 — Start Polaris
&lt;/h4&gt;

&lt;p&gt;Polaris runs locally with permissive settings. Anything that fails here is a problem in the test scripts, so it is fixed before any cloud catalog is tested:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;git clone https://github.com/xbill9/lakehouse-iceberg-2026
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;lakehouse-iceberg-2026/iceberg-conformance
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;./polaris-up.sh
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;POLARIS_CLIENT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;root &lt;span class="nv"&gt;POLARIS_CLIENT_SECRET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;s3cr3t
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 2 — Build the Test Program
&lt;/h4&gt;

&lt;p&gt;The versions are pinned exactly, so every result applies to one release:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="py"&gt;iceberg&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="py"&gt;"&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="py"&gt;"
iceberg-catalog-rest = "&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="py"&gt;"
iceberg-storage-opendal = { version = "&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="s"&gt;", features = ["&lt;/span&gt;&lt;span class="err"&gt;opendal-gcs&lt;/span&gt;&lt;span class="s"&gt;", "&lt;/span&gt;&lt;span class="err"&gt;opendal-azdls&lt;/span&gt;&lt;span class="s"&gt;"] }&lt;/span&gt;&lt;span class="err"&gt;
&lt;/span&gt;&lt;span class="py"&gt;reqwest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="py"&gt;version&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"0.12"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="py"&gt;default-features&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="py"&gt;features&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"rustls-tls"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The last two lines add a TLS backend and the cloud storage backends. &lt;em&gt;What You Add Beyond the Client&lt;/em&gt;, near the end, says what happens without each of them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ../iceberg-rust-client
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;cargo build &lt;span class="nt"&gt;--release&lt;/span&gt;
&lt;span class="go"&gt;    Finished `release` profile [optimized] target(s) in 1m 41s
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 3 — List What the Client Supports
&lt;/h4&gt;

&lt;p&gt;Each of the 33 checks is matched to a method in the client's published source code, with the file and line number, and &lt;code&gt;check_refs.py&lt;/code&gt; confirms all 25 of those lines still sit inside the function they name. Counting each endpoint once, since five checks use the same &lt;code&gt;update_table&lt;/code&gt; endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;13 of 25 distinct endpoints are expressible through this client.
Counted per probe the figure is 19 of 33, which is the same fact
weighted by how many probes paper 1 happened to point at one endpoint.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 12 endpoints it cannot reach split two ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Missing — 11 endpoints.&lt;/strong&gt; All seven view operations, scan planning, metrics reporting, the separate credentials endpoint, and &lt;code&gt;commitTransaction&lt;/code&gt;. The client builds eight URLs in total (&lt;code&gt;catalog.rs:177-215&lt;/code&gt;), and none of them is a view, a scan plan, a metrics report or a transaction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stubbed — 1 endpoint.&lt;/strong&gt; &lt;code&gt;Catalog::update_namespace&lt;/code&gt; exists, but returns the error &lt;code&gt;"Updating namespace not supported yet!"&lt;/code&gt; (&lt;code&gt;catalog.rs:659&lt;/code&gt;) and sends nothing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two endpoints it does reach, it reaches only in part, and both are counted among the 13: &lt;code&gt;load_table&lt;/code&gt; cannot ask for &lt;code&gt;?snapshots=all&lt;/code&gt;, and &lt;code&gt;list_namespaces&lt;/code&gt; handles paging internally and never sends &lt;code&gt;pageSize&lt;/code&gt;. It also reaches &lt;code&gt;registerTable&lt;/code&gt;, which is one of the ten operations the earlier article did not cover, so it sits outside this count of 25.&lt;/p&gt;

&lt;p&gt;Line numbers quoted here were read by hand and are archived, with the lines themselves, in &lt;code&gt;evidence/rust-source-citations.txt&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — Run the Reads Against Polaris
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 run_rust.py &lt;span class="nt"&gt;--storage&lt;/span&gt; opendal
&lt;span class="go"&gt;apache-polaris     1 implicit, 14 not-expressible, 11 not-issued, 7 ok
wrote evidence/rust-run-apache-polaris.json
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All 7 supported read operations work. The one &lt;code&gt;IMPLICIT&lt;/code&gt; result is the config request the client sends when it first connects, and the 11 &lt;code&gt;NOT-ISSUED&lt;/code&gt; are writes this run does not send.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 5 — Run the Writes Against Polaris
&lt;/h4&gt;

&lt;p&gt;The client has a method for each of those 11, and the read-only run sends none of them. A method that compiles can still fail on the wire, so this step sends them. Polaris is local, permissive and disposable, so the writes run there and on no other catalog, and every request goes through a small logging proxy that records what the client sent beside what came back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 run_writes.py
&lt;span class="go"&gt;apache-polaris     11 ok, 1 unsupported, 0 failed
wrote evidence/rust-write-surface.txt   22 request(s)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  probe                          verdict      ms     endpoint
  create_namespace               OK           60     reachable
  update_namespace_props         UNSUPPORTED  0      unsupported
  create_table                   OK           64     reachable
  commit_table                   OK           112    reachable
  commit_remove_properties       OK           119    reachable
  commit_add_schema              OK           116    reachable
  commit_set_current_schema      OK           116    reachable
  commit_upgrade_format_version  OK           99     reachable
  rename_table                   OK           3      reachable
  drop_table_purge               OK           16     reachable
  drop_table                     OK           3      reachable
  drop_namespace                 OK           1      reachable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All eleven work, and the scratch namespace is dropped at the end. The one refusal is the stub from Step 3, and the proxy log shows why it takes 0 ms: no request to the properties endpoint appears anywhere in the run.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: The Proxy Log Shows Three Things
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;A properties commit always sends both update kinds.&lt;/strong&gt; Setting a property and removing one produce the same pair, &lt;code&gt;set-properties&lt;/code&gt; and &lt;code&gt;remove-properties&lt;/code&gt;, because the crate's properties action builds both every time (&lt;code&gt;update_properties.rs:95&lt;/code&gt;). The two differ in their contents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Adding a column is one request carrying two updates&lt;/strong&gt;, with a requirement attached:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  POST    200  /v1/quickstart_catalog/namespaces/irc_probe_rust_1790009108/tables/t1
          updates: add-schema, set-current-schema
          requirements: assert-current-schema-id
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Every commit re-reads the table first&lt;/strong&gt;, so one property change is a &lt;code&gt;GET&lt;/code&gt; and then a &lt;code&gt;POST&lt;/code&gt;. That reload is also why a transaction commits against the table's current state, which is open upstream as &lt;code&gt;apache/iceberg-rust&lt;/code&gt; #3134.&lt;/p&gt;

&lt;p&gt;The log also confirms one thing: a two-level namespace goes out as &lt;code&gt;...%1Fchild&lt;/code&gt;, the unit separator the specification asks for.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 6 — Match the Login to Each Catalog
&lt;/h4&gt;

&lt;p&gt;The client can log in with a token, with an OAuth2 client ID and secret, or with fixed extra headers. It has no AWS request signing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;login&lt;/th&gt;
&lt;th&gt;catalogs&lt;/th&gt;
&lt;th&gt;with this client&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OAuth2&lt;/td&gt;
&lt;td&gt;Polaris&lt;/td&gt;
&lt;td&gt;built in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;token from an environment variable&lt;/td&gt;
&lt;td&gt;Unity&lt;/td&gt;
&lt;td&gt;built in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;key-pair JWT&lt;/td&gt;
&lt;td&gt;Horizon&lt;/td&gt;
&lt;td&gt;built in: a &lt;code&gt;credential&lt;/code&gt; with no colon is sent as &lt;code&gt;client_secret&lt;/code&gt; with no &lt;code&gt;client_id&lt;/code&gt; (&lt;code&gt;catalog.rs:238&lt;/code&gt;), which is what Horizon expects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;gcloud&lt;/code&gt; or &lt;code&gt;az&lt;/code&gt; login&lt;/td&gt;
&lt;td&gt;BigLake, OneLake&lt;/td&gt;
&lt;td&gt;a token created outside the client, which the client cannot renew&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS SigV4 signing&lt;/td&gt;
&lt;td&gt;Glue, S3 Tables&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;not supported&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Glue and S3 Tables require SigV4 on the catalog requests themselves. The specification does not describe that — its security schemes are OAuth2 and bearer tokens, and the signing in it covers storage access — but &lt;code&gt;pyiceberg&lt;/code&gt; signs anyway, with &lt;code&gt;rest.sigv4-enabled&lt;/code&gt;, &lt;code&gt;rest.signing-name&lt;/code&gt; and &lt;code&gt;rest.signing-region&lt;/code&gt;, so the Python REST client reaches all seven catalogs and the Rust one reaches five. The request to add it, &lt;code&gt;apache/iceberg-rust&lt;/code&gt; #1236, has been open since April 2025.&lt;/p&gt;

&lt;p&gt;Fixed headers look like a way round it, and they are worth one measurement: a signature minted for &lt;code&gt;GET /v1/config&lt;/code&gt; and passed as a static header got a 200 on that request and a 403 on the next, on Glue and S3 Tables both, with all 7 checks refused through the client. A signature covers the request it signs.&lt;/p&gt;

&lt;p&gt;The token the client does hold is refreshed by &lt;code&gt;regenerate_token()&lt;/code&gt;, which only knows how to repeat an OAuth2 login, so a long-running program on BigLake or OneLake mints its own.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 7 — Reach AWS Through the Other Two Crates
&lt;/h4&gt;

&lt;p&gt;Rust does reach both AWS catalogs. The same Apache project publishes &lt;code&gt;iceberg-catalog-glue&lt;/code&gt; and &lt;code&gt;iceberg-catalog-s3tables&lt;/code&gt;, both 0.10.1, which call the AWS APIs directly and sign as the AWS SDK does. All three crates implement the same &lt;code&gt;Catalog&lt;/code&gt; trait, so a program written against &lt;code&gt;dyn Catalog&lt;/code&gt; swaps between them by changing a dependency.&lt;/p&gt;

&lt;p&gt;What changes with the dependency is the set of operations that answers. Reading each crate's &lt;code&gt;impl Catalog for&lt;/code&gt; block:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  Catalog trait method     rest       glue       s3tables
  update_namespace         refused    sent       refused
  drop_table               sent       sent       refused
  register_table           sent       sent       refused
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The other 12 trait methods are sent by all three. Each refusal says why:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  rest       update_namespace         'Updating namespace not supported yet!'
  s3tables   drop_table               'drop_table is not supported for S3Tables; use purge_table instead'
  s3tables   register_table           'Registering a table is not supported yet'
  s3tables   update_namespace         'Update namespace is not supported for s3tables catalog'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One of those four is the service speaking: S3 Tables requires a purge, which the earlier article measured from the wire, so that refusal is the catalog's own rule carried faithfully by the crate. The other three are crate-level gaps, and they land at runtime, since all three crates satisfy the same trait and compile the same way.&lt;/p&gt;

&lt;p&gt;So a Rust tool covering all seven catalogs carries three catalog implementations and an operation set that varies by which one is loaded. In Python, one client covers all seven.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 8 — Point It at a Managed Catalog
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 run_rust.py &lt;span class="nt"&gt;--only&lt;/span&gt; google-lakehouse &lt;span class="nt"&gt;--only&lt;/span&gt; microsoft-onelake &lt;span class="nt"&gt;--storage&lt;/span&gt; opendal
&lt;span class="go"&gt;google-lakehouse   1 implicit, 14 not-expressible, 11 not-issued, 7 ok
wrote evidence/rust-run-google-lakehouse.json
microsoft-onelake  1 implicit, 14 not-expressible, 11 not-issued, 7 ok
wrote evidence/rust-run-microsoft-onelake.json
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same result as Polaris: all 7 supported read operations work on both, over the internet, with a token minted by &lt;code&gt;gcloud&lt;/code&gt; and by &lt;code&gt;az&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 9 — Read the Table's Files
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;load_table&lt;/code&gt; sets up file access but does not read anything, so each run also reads the table's metadata file through the client. The output lists the names of the settings the storage library received, never their values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;apache-polaris
  fileio config keys: ...
  storage credential keys among them: none
  read: ok, scheme file, 5614 bytes, 21 ms
google-lakehouse
  fileio config keys: ...
  storage credential keys among them: none
  read: ok, scheme gs, 5497 bytes, 451 ms
microsoft-onelake
  fileio config keys: ...
  storage credential keys among them: none
  read: FAILED after 47066 ms
  error: Unexpected =&amp;gt; Failure in doing io operation, source: Unexpected
  (persistent) at read, context: { timeout: 10 } =&amp;gt; io timeout reached
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Local files and Google Cloud Storage read fine, the second one on the Google login already on the machine. No catalog passed along a storage credential, because the client never asks for one: it sends no &lt;code&gt;X-Iceberg-Access-Delegation&lt;/code&gt; header, so the storage library finds a login on its own.&lt;/p&gt;

&lt;p&gt;Azure is where that runs out. The Azure backend, &lt;code&gt;opendal-service-azdls&lt;/code&gt; 0.57.0, has code to use your &lt;code&gt;az login&lt;/code&gt; and is never given a way to run the &lt;code&gt;az&lt;/code&gt; command (&lt;code&gt;backend.rs:297&lt;/code&gt;), which leaves an account key, a SAS token or a service principal secret — and OneLake has no account key. Asking for a credential does not change it: OneLake hands one out, and with the delegation header passed as a static header the read still timed out, after 47181 ms. The client reads the credential from the response (&lt;code&gt;types.rs:226&lt;/code&gt;) and copies only the response's &lt;code&gt;config&lt;/code&gt; settings into file access (&lt;code&gt;catalog.rs:455&lt;/code&gt;). Both halves are open upstream, as #2931 and #1442.&lt;/p&gt;




&lt;h4&gt;
  
  
  What You Add Beyond the Client
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;needed for&lt;/th&gt;
&lt;th&gt;what happens without it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;a &lt;code&gt;reqwest&lt;/code&gt; TLS feature&lt;/td&gt;
&lt;td&gt;any &lt;code&gt;https://&lt;/code&gt; catalog&lt;/td&gt;
&lt;td&gt;every request fails before any response: 7 failed, 0 ok, no HTTP status. Upstream #2888&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;iceberg-storage-opendal&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;loading any cloud-stored table&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;load_table&lt;/code&gt; refuses: &lt;em&gt;"StorageFactory must be provided for RestCatalog"&lt;/em&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a storage login in your environment&lt;/td&gt;
&lt;td&gt;reading any table file&lt;/td&gt;
&lt;td&gt;the read fails or hangs; catalog-issued credentials are read and unused&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a SigV4 signer&lt;/td&gt;
&lt;td&gt;Glue and S3 Tables over REST&lt;/td&gt;
&lt;td&gt;no login at all, and Step 7's two crates are the way round it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two are one line each. &lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1 declares &lt;code&gt;reqwest&lt;/code&gt; with TLS switched off and offers no feature to switch it on, so a build that never names a TLS backend fails on the first &lt;code&gt;https://&lt;/code&gt; request; Cargo merges features across dependencies, so your own line fixes it, and a project already using &lt;code&gt;reqwest&lt;/code&gt; with TLS will never see it. The core &lt;code&gt;iceberg&lt;/code&gt; 0.10.1 crate ships two storage factories, local files and memory (&lt;code&gt;io/storage/local_fs.rs:330&lt;/code&gt;, &lt;code&gt;io/storage/memory.rs:250&lt;/code&gt;), with its README pointing at the other crate on line 70.&lt;/p&gt;




&lt;h4&gt;
  
  
  Compare and Contrast
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;catalog&lt;/th&gt;
&lt;th&gt;login&lt;/th&gt;
&lt;th&gt;ok&lt;/th&gt;
&lt;th&gt;failed&lt;/th&gt;
&lt;th&gt;table files&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Apache Polaris&lt;/td&gt;
&lt;td&gt;🟢 OAuth2&lt;/td&gt;
&lt;td&gt;7 reads, 11 writes&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;🟢 read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google BigLake&lt;/td&gt;
&lt;td&gt;🟢 &lt;code&gt;gcloud&lt;/code&gt; token&lt;/td&gt;
&lt;td&gt;7 reads&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;🟢 read, on the machine's Google login&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Microsoft OneLake&lt;/td&gt;
&lt;td&gt;🟢 &lt;code&gt;az&lt;/code&gt; token&lt;/td&gt;
&lt;td&gt;7 reads&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;❌ timed out&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS Glue&lt;/td&gt;
&lt;td&gt;❌ SigV4&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS S3 Tables&lt;/td&gt;
&lt;td&gt;❌ SigV4&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Databricks Unity&lt;/td&gt;
&lt;td&gt;token, not run&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Snowflake Horizon&lt;/td&gt;
&lt;td&gt;key-pair JWT, not run&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Behind those numbers every catalog that answered showed the same shape: 14 endpoints the client has no method for, and every endpoint it does implement returning an answer.&lt;/p&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to point the Apache Rust Iceberg REST client at seven catalogs and record what works and what it takes. The key to the solution was matching every test to a line in the client's source code, running the same tests against each catalog, and logging every request the client sent. The results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 Every endpoint the client implements answered on every catalog it could log in to: 7 reads on Polaris, BigLake and OneLake, and 11 writes on Polaris, with 0 failures&lt;/li&gt;
&lt;li&gt;🟢 13 of 25 endpoints are implemented; of the other 12, 11 are missing and 1 is a stub that sends no request&lt;/li&gt;
&lt;li&gt;⚠️ Two lines of &lt;code&gt;Cargo.toml&lt;/code&gt; stand in front of that: a TLS backend (upstream #2888) and &lt;code&gt;iceberg-storage-opendal&lt;/code&gt; for cloud storage&lt;/li&gt;
&lt;li&gt;❌ Glue and S3 Tables require SigV4, which this client cannot send, so a Rust tool covering all seven catalogs loads three catalog crates where Python loads one — and those three disagree on three of the 15 &lt;code&gt;Catalog&lt;/code&gt; methods, at runtime (upstream #1236, open since April 2025)&lt;/li&gt;
&lt;li&gt;❌ OneLake's files could not be read: the Azure backend cannot use an &lt;code&gt;az login&lt;/code&gt;, and a credential from the catalog is read and then ignored (upstream #2931, #1442)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: &lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1 with &lt;code&gt;iceberg&lt;/code&gt; 0.10.1, &lt;code&gt;iceberg-storage-opendal&lt;/code&gt; 0.10.1 and &lt;code&gt;reqwest&lt;/code&gt; 0.12.28 with &lt;code&gt;rustls-tls&lt;/code&gt;, built with &lt;code&gt;rustc&lt;/code&gt; 1.98.1. Source read 2026-09-04, versions captured 2026-09-17, catalog runs 2026-09-18, write run 2026-09-21, one run each from one machine in one region. Polaris 1.7.0 ran in Docker with permissive settings and local file storage, and the writes ran there alone, so those results describe one permissive server. Glue and S3 Tables were tested only with the signature experiment, and Unity and Horizon were not run at all, their rows coming from the source. The tests check that each operation answers; the answers themselves are not checked. Three of the seven catalogs were on trial accounts in the earlier article, and managed catalogs do not report a version.&lt;/p&gt;

&lt;p&gt;The strategy for testing what one Rust Iceberg client can reach was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/xbill9/lakehouse-iceberg-2026" rel="noopener noreferrer"&gt;lakehouse-iceberg-2026 | GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust" rel="noopener noreferrer"&gt;apache/iceberg-rust&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/1236" rel="noopener noreferrer"&gt;iceberg-rust #1236 — REST catalog: support AWS sigV4&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/2888" rel="noopener noreferrer"&gt;iceberg-rust #2888 — Add TLS features to iceberg-catalog-rest&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/2931" rel="noopener noreferrer"&gt;iceberg-rust #2931 — Support refreshing vended storage credentials for REST catalog tables&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/1442" rel="noopener noreferrer"&gt;iceberg-rust #1442 — ADLS: Support vended "adls.sas-token.xxx" prefixed tokens&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/3134" rel="noopener noreferrer"&gt;iceberg-rust #3134 — Transaction commits against a base it never validated&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://crates.io/crates/iceberg-catalog-glue" rel="noopener noreferrer"&gt;iceberg-catalog-glue | crates.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://crates.io/crates/iceberg-catalog-s3tables" rel="noopener noreferrer"&gt;iceberg-catalog-s3tables | crates.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/gde/seven-iceberg-rest-catalogs-what-they-declare-and-what-they-serve-40oj"&gt;Seven Iceberg REST Catalogs: What They Declare, and What They Serve&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>rust</category>
      <category>iceberg</category>
      <category>lakehouse</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>What One Rust Client Can Reach Across Seven Iceberg Catalogs</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Fri, 25 Sep 2026 16:20:54 +0000</pubDate>
      <link>https://dev.to/aws-builders/what-one-rust-client-can-reach-across-seven-iceberg-catalogs-24al</link>
      <guid>https://dev.to/aws-builders/what-one-rust-client-can-reach-across-seven-iceberg-catalogs-24al</guid>
      <description>&lt;p&gt;This article provides a step by step guide to pointing the Apache Rust client for Iceberg REST catalogs at seven catalogs and recording what works. A Python script runs a small Rust program against each catalog and saves every result.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/lakehouse-iceberg-2026" rel="noopener noreferrer"&gt;https://github.com/xbill9/lakehouse-iceberg-2026&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  What is this project trying to Do?
&lt;/h4&gt;

&lt;p&gt;An earlier article tested what seven Iceberg REST catalogs support. This one asks from the other side: if you write a lakehouse tool in Rust today, will it work with the catalog you already pay for, and what do you need to set up? The client is &lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1 from the Apache Iceberg project, and the catalogs, tests and test table are the same as in that article:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Apache Polaris&lt;/strong&gt; 1.7.0, running locally as the reference&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google BigLake&lt;/strong&gt; and &lt;strong&gt;Microsoft OneLake&lt;/strong&gt;, over the internet&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AWS Glue&lt;/strong&gt; and &lt;strong&gt;AWS S3 Tables&lt;/strong&gt;, over the internet&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Databricks Unity&lt;/strong&gt; and &lt;strong&gt;Snowflake Horizon&lt;/strong&gt;, not run here&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result is a short one. This client implements 13 of the 25 operations the earlier article tested, it logs in to five of the seven catalogs, and on every catalog it logs in to, all 13 answer. Two lines of &lt;code&gt;Cargo.toml&lt;/code&gt; stand between it and those five. The other two are AWS, they require a signature this client cannot produce, and the Rust answer to that is two more catalog crates that do not implement quite the same operations.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Does Each Result Mean?
&lt;/h4&gt;

&lt;p&gt;The test suite is 33 checks covering 25 of the 35 operations in the Iceberg REST specification. Each check gets one of five results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;th&gt;meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;OK&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client sent the request and got an answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;FAILED&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client sent the request and got an error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IMPLICIT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client sends this request on its own, and there is no way to call it directly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;NOT-EXPRESSIBLE&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client has no method for this operation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;NOT-ISSUED&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the client supports it, but it is a write, and that run only reads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  Where do I start?
&lt;/h4&gt;

&lt;p&gt;The strategy for testing the client is an incremental step by step approach.&lt;/p&gt;

&lt;p&gt;First, a local Polaris catalog is started and the client's source is read to list which operations it supports. The test program is then built and run against Polaris, reads first and writes after, then against the managed catalogs one login method at a time, and finally it tries to read a table's files.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Rust 1.94 or newer and &lt;code&gt;cargo&lt;/code&gt; — this run used &lt;code&gt;rustc&lt;/code&gt; 1.98.1&lt;/li&gt;
&lt;li&gt;Docker, for the local Polaris catalog&lt;/li&gt;
&lt;li&gt;Python 3.10+ with &lt;code&gt;pyyaml&lt;/code&gt;, for the test scripts&lt;/li&gt;
&lt;li&gt;Optional: logins for any managed catalog — &lt;code&gt;gcloud&lt;/code&gt; for BigLake, &lt;code&gt;az&lt;/code&gt; for OneLake, the AWS CLI for Glue and S3 Tables&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Step 1 — Start Polaris
&lt;/h4&gt;

&lt;p&gt;Polaris runs locally with permissive settings. Anything that fails here is a problem in the test scripts, so it is fixed before any cloud catalog is tested:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;git clone https://github.com/xbill9/lakehouse-iceberg-2026
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;lakehouse-iceberg-2026/iceberg-conformance
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;./polaris-up.sh
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;POLARIS_CLIENT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;root &lt;span class="nv"&gt;POLARIS_CLIENT_SECRET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;s3cr3t
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 2 — Build the Test Program
&lt;/h4&gt;

&lt;p&gt;The versions are pinned exactly, so every result applies to one release:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="py"&gt;iceberg&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="py"&gt;"&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="py"&gt;"
iceberg-catalog-rest = "&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="py"&gt;"
iceberg-storage-opendal = { version = "&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="s"&gt;", features = ["&lt;/span&gt;&lt;span class="err"&gt;opendal-gcs&lt;/span&gt;&lt;span class="s"&gt;", "&lt;/span&gt;&lt;span class="err"&gt;opendal-azdls&lt;/span&gt;&lt;span class="s"&gt;"] }&lt;/span&gt;&lt;span class="err"&gt;
&lt;/span&gt;&lt;span class="py"&gt;reqwest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="py"&gt;version&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"0.12"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="py"&gt;default-features&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="py"&gt;features&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"rustls-tls"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The last two lines add a TLS backend and the cloud storage backends. &lt;em&gt;What You Add Beyond the Client&lt;/em&gt;, near the end, says what happens without each of them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ../iceberg-rust-client
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;cargo build &lt;span class="nt"&gt;--release&lt;/span&gt;
&lt;span class="go"&gt;    Finished `release` profile [optimized] target(s) in 1m 41s
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 3 — List What the Client Supports
&lt;/h4&gt;

&lt;p&gt;Each of the 33 checks is matched to a method in the client's published source code, with the file and line number, and &lt;code&gt;check_refs.py&lt;/code&gt; confirms all 25 of those lines still sit inside the function they name. Counting each endpoint once, since five checks use the same &lt;code&gt;update_table&lt;/code&gt; endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;13 of 25 distinct endpoints are expressible through this client.
Counted per probe the figure is 19 of 33, which is the same fact
weighted by how many probes paper 1 happened to point at one endpoint.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 12 endpoints it cannot reach split two ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Missing — 11 endpoints.&lt;/strong&gt; All seven view operations, scan planning, metrics reporting, the separate credentials endpoint, and &lt;code&gt;commitTransaction&lt;/code&gt;. The client builds eight URLs in total (&lt;code&gt;catalog.rs:177-215&lt;/code&gt;), and none of them is a view, a scan plan, a metrics report or a transaction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stubbed — 1 endpoint.&lt;/strong&gt; &lt;code&gt;Catalog::update_namespace&lt;/code&gt; exists, but returns the error &lt;code&gt;"Updating namespace not supported yet!"&lt;/code&gt; (&lt;code&gt;catalog.rs:659&lt;/code&gt;) and sends nothing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two endpoints it does reach, it reaches only in part, and both are counted among the 13: &lt;code&gt;load_table&lt;/code&gt; cannot ask for &lt;code&gt;?snapshots=all&lt;/code&gt;, and &lt;code&gt;list_namespaces&lt;/code&gt; handles paging internally and never sends &lt;code&gt;pageSize&lt;/code&gt;. It also reaches &lt;code&gt;registerTable&lt;/code&gt;, which is one of the ten operations the earlier article did not cover, so it sits outside this count of 25.&lt;/p&gt;

&lt;p&gt;Line numbers quoted here were read by hand and are archived, with the lines themselves, in &lt;code&gt;evidence/rust-source-citations.txt&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — Run the Reads Against Polaris
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 run_rust.py &lt;span class="nt"&gt;--storage&lt;/span&gt; opendal
&lt;span class="go"&gt;apache-polaris     1 implicit, 14 not-expressible, 11 not-issued, 7 ok
wrote evidence/rust-run-apache-polaris.json
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All 7 supported read operations work. The one &lt;code&gt;IMPLICIT&lt;/code&gt; result is the config request the client sends when it first connects, and the 11 &lt;code&gt;NOT-ISSUED&lt;/code&gt; are writes this run does not send.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 5 — Run the Writes Against Polaris
&lt;/h4&gt;

&lt;p&gt;The client has a method for each of those 11, and the read-only run sends none of them. A method that compiles can still fail on the wire, so this step sends them. Polaris is local, permissive and disposable, so the writes run there and on no other catalog, and every request goes through a small logging proxy that records what the client sent beside what came back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 run_writes.py
&lt;span class="go"&gt;apache-polaris     11 ok, 1 unsupported, 0 failed
wrote evidence/rust-write-surface.txt   22 request(s)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  probe                          verdict      ms     endpoint
  create_namespace               OK           60     reachable
  update_namespace_props         UNSUPPORTED  0      unsupported
  create_table                   OK           64     reachable
  commit_table                   OK           112    reachable
  commit_remove_properties       OK           119    reachable
  commit_add_schema              OK           116    reachable
  commit_set_current_schema      OK           116    reachable
  commit_upgrade_format_version  OK           99     reachable
  rename_table                   OK           3      reachable
  drop_table_purge               OK           16     reachable
  drop_table                     OK           3      reachable
  drop_namespace                 OK           1      reachable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All eleven work, and the scratch namespace is dropped at the end. The one refusal is the stub from Step 3, and the proxy log shows why it takes 0 ms: no request to the properties endpoint appears anywhere in the run.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: The Proxy Log Shows Three Things
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;A properties commit always sends both update kinds.&lt;/strong&gt; Setting a property and removing one produce the same pair, &lt;code&gt;set-properties&lt;/code&gt; and &lt;code&gt;remove-properties&lt;/code&gt;, because the crate's properties action builds both every time (&lt;code&gt;update_properties.rs:95&lt;/code&gt;). The two differ in their contents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Adding a column is one request carrying two updates&lt;/strong&gt;, with a requirement attached:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  POST    200  /v1/quickstart_catalog/namespaces/irc_probe_rust_1790009108/tables/t1
          updates: add-schema, set-current-schema
          requirements: assert-current-schema-id
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Every commit re-reads the table first&lt;/strong&gt;, so one property change is a &lt;code&gt;GET&lt;/code&gt; and then a &lt;code&gt;POST&lt;/code&gt;. That reload is also why a transaction commits against the table's current state, which is open upstream as &lt;code&gt;apache/iceberg-rust&lt;/code&gt; #3134.&lt;/p&gt;

&lt;p&gt;The log also confirms one thing: a two-level namespace goes out as &lt;code&gt;...%1Fchild&lt;/code&gt;, the unit separator the specification asks for.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 6 — Match the Login to Each Catalog
&lt;/h4&gt;

&lt;p&gt;The client can log in with a token, with an OAuth2 client ID and secret, or with fixed extra headers. It has no AWS request signing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;login&lt;/th&gt;
&lt;th&gt;catalogs&lt;/th&gt;
&lt;th&gt;with this client&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OAuth2&lt;/td&gt;
&lt;td&gt;Polaris&lt;/td&gt;
&lt;td&gt;built in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;token from an environment variable&lt;/td&gt;
&lt;td&gt;Unity&lt;/td&gt;
&lt;td&gt;built in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;key-pair JWT&lt;/td&gt;
&lt;td&gt;Horizon&lt;/td&gt;
&lt;td&gt;built in: a &lt;code&gt;credential&lt;/code&gt; with no colon is sent as &lt;code&gt;client_secret&lt;/code&gt; with no &lt;code&gt;client_id&lt;/code&gt; (&lt;code&gt;catalog.rs:238&lt;/code&gt;), which is what Horizon expects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;gcloud&lt;/code&gt; or &lt;code&gt;az&lt;/code&gt; login&lt;/td&gt;
&lt;td&gt;BigLake, OneLake&lt;/td&gt;
&lt;td&gt;a token created outside the client, which the client cannot renew&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS SigV4 signing&lt;/td&gt;
&lt;td&gt;Glue, S3 Tables&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;not supported&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Glue and S3 Tables require SigV4 on the catalog requests themselves. The specification does not describe that — its security schemes are OAuth2 and bearer tokens, and the signing in it covers storage access — but &lt;code&gt;pyiceberg&lt;/code&gt; signs anyway, with &lt;code&gt;rest.sigv4-enabled&lt;/code&gt;, &lt;code&gt;rest.signing-name&lt;/code&gt; and &lt;code&gt;rest.signing-region&lt;/code&gt;, so the Python REST client reaches all seven catalogs and the Rust one reaches five. The request to add it, &lt;code&gt;apache/iceberg-rust&lt;/code&gt; #1236, has been open since April 2025.&lt;/p&gt;

&lt;p&gt;Fixed headers look like a way round it, and they are worth one measurement: a signature minted for &lt;code&gt;GET /v1/config&lt;/code&gt; and passed as a static header got a 200 on that request and a 403 on the next, on Glue and S3 Tables both, with all 7 checks refused through the client. A signature covers the request it signs.&lt;/p&gt;

&lt;p&gt;The token the client does hold is refreshed by &lt;code&gt;regenerate_token()&lt;/code&gt;, which only knows how to repeat an OAuth2 login, so a long-running program on BigLake or OneLake mints its own.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 7 — Reach AWS Through the Other Two Crates
&lt;/h4&gt;

&lt;p&gt;Rust does reach both AWS catalogs. The same Apache project publishes &lt;code&gt;iceberg-catalog-glue&lt;/code&gt; and &lt;code&gt;iceberg-catalog-s3tables&lt;/code&gt;, both 0.10.1, which call the AWS APIs directly and sign as the AWS SDK does. All three crates implement the same &lt;code&gt;Catalog&lt;/code&gt; trait, so a program written against &lt;code&gt;dyn Catalog&lt;/code&gt; swaps between them by changing a dependency.&lt;/p&gt;

&lt;p&gt;What changes with the dependency is the set of operations that answers. Reading each crate's &lt;code&gt;impl Catalog for&lt;/code&gt; block:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  Catalog trait method     rest       glue       s3tables
  update_namespace         refused    sent       refused
  drop_table               sent       sent       refused
  register_table           sent       sent       refused
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The other 12 trait methods are sent by all three. Each refusal says why:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  rest       update_namespace         'Updating namespace not supported yet!'
  s3tables   drop_table               'drop_table is not supported for S3Tables; use purge_table instead'
  s3tables   register_table           'Registering a table is not supported yet'
  s3tables   update_namespace         'Update namespace is not supported for s3tables catalog'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One of those four is the service speaking: S3 Tables requires a purge, which the earlier article measured from the wire, so that refusal is the catalog's own rule carried faithfully by the crate. The other three are crate-level gaps, and they land at runtime, since all three crates satisfy the same trait and compile the same way.&lt;/p&gt;

&lt;p&gt;So a Rust tool covering all seven catalogs carries three catalog implementations and an operation set that varies by which one is loaded. In Python, one client covers all seven.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 8 — Point It at a Managed Catalog
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 run_rust.py &lt;span class="nt"&gt;--only&lt;/span&gt; google-lakehouse &lt;span class="nt"&gt;--only&lt;/span&gt; microsoft-onelake &lt;span class="nt"&gt;--storage&lt;/span&gt; opendal
&lt;span class="go"&gt;google-lakehouse   1 implicit, 14 not-expressible, 11 not-issued, 7 ok
wrote evidence/rust-run-google-lakehouse.json
microsoft-onelake  1 implicit, 14 not-expressible, 11 not-issued, 7 ok
wrote evidence/rust-run-microsoft-onelake.json
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same result as Polaris: all 7 supported read operations work on both, over the internet, with a token minted by &lt;code&gt;gcloud&lt;/code&gt; and by &lt;code&gt;az&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 9 — Read the Table's Files
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;load_table&lt;/code&gt; sets up file access but does not read anything, so each run also reads the table's metadata file through the client. The output lists the names of the settings the storage library received, never their values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;apache-polaris
  fileio config keys: ...
  storage credential keys among them: none
  read: ok, scheme file, 5614 bytes, 21 ms
google-lakehouse
  fileio config keys: ...
  storage credential keys among them: none
  read: ok, scheme gs, 5497 bytes, 451 ms
microsoft-onelake
  fileio config keys: ...
  storage credential keys among them: none
  read: FAILED after 47066 ms
  error: Unexpected =&amp;gt; Failure in doing io operation, source: Unexpected
  (persistent) at read, context: { timeout: 10 } =&amp;gt; io timeout reached
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Local files and Google Cloud Storage read fine, the second one on the Google login already on the machine. No catalog passed along a storage credential, because the client never asks for one: it sends no &lt;code&gt;X-Iceberg-Access-Delegation&lt;/code&gt; header, so the storage library finds a login on its own.&lt;/p&gt;

&lt;p&gt;Azure is where that runs out. The Azure backend, &lt;code&gt;opendal-service-azdls&lt;/code&gt; 0.57.0, has code to use your &lt;code&gt;az login&lt;/code&gt; and is never given a way to run the &lt;code&gt;az&lt;/code&gt; command (&lt;code&gt;backend.rs:297&lt;/code&gt;), which leaves an account key, a SAS token or a service principal secret — and OneLake has no account key. Asking for a credential does not change it: OneLake hands one out, and with the delegation header passed as a static header the read still timed out, after 47181 ms. The client reads the credential from the response (&lt;code&gt;types.rs:226&lt;/code&gt;) and copies only the response's &lt;code&gt;config&lt;/code&gt; settings into file access (&lt;code&gt;catalog.rs:455&lt;/code&gt;). Both halves are open upstream, as #2931 and #1442.&lt;/p&gt;




&lt;h4&gt;
  
  
  What You Add Beyond the Client
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;needed for&lt;/th&gt;
&lt;th&gt;what happens without it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;a &lt;code&gt;reqwest&lt;/code&gt; TLS feature&lt;/td&gt;
&lt;td&gt;any &lt;code&gt;https://&lt;/code&gt; catalog&lt;/td&gt;
&lt;td&gt;every request fails before any response: 7 failed, 0 ok, no HTTP status. Upstream #2888&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;iceberg-storage-opendal&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;loading any cloud-stored table&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;load_table&lt;/code&gt; refuses: &lt;em&gt;"StorageFactory must be provided for RestCatalog"&lt;/em&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a storage login in your environment&lt;/td&gt;
&lt;td&gt;reading any table file&lt;/td&gt;
&lt;td&gt;the read fails or hangs; catalog-issued credentials are read and unused&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a SigV4 signer&lt;/td&gt;
&lt;td&gt;Glue and S3 Tables over REST&lt;/td&gt;
&lt;td&gt;no login at all, and Step 7's two crates are the way round it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two are one line each. &lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1 declares &lt;code&gt;reqwest&lt;/code&gt; with TLS switched off and offers no feature to switch it on, so a build that never names a TLS backend fails on the first &lt;code&gt;https://&lt;/code&gt; request; Cargo merges features across dependencies, so your own line fixes it, and a project already using &lt;code&gt;reqwest&lt;/code&gt; with TLS will never see it. The core &lt;code&gt;iceberg&lt;/code&gt; 0.10.1 crate ships two storage factories, local files and memory (&lt;code&gt;io/storage/local_fs.rs:330&lt;/code&gt;, &lt;code&gt;io/storage/memory.rs:250&lt;/code&gt;), with its README pointing at the other crate on line 70.&lt;/p&gt;




&lt;h4&gt;
  
  
  Compare and Contrast
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;catalog&lt;/th&gt;
&lt;th&gt;login&lt;/th&gt;
&lt;th&gt;ok&lt;/th&gt;
&lt;th&gt;failed&lt;/th&gt;
&lt;th&gt;table files&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Apache Polaris&lt;/td&gt;
&lt;td&gt;🟢 OAuth2&lt;/td&gt;
&lt;td&gt;7 reads, 11 writes&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;🟢 read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google BigLake&lt;/td&gt;
&lt;td&gt;🟢 &lt;code&gt;gcloud&lt;/code&gt; token&lt;/td&gt;
&lt;td&gt;7 reads&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;🟢 read, on the machine's Google login&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Microsoft OneLake&lt;/td&gt;
&lt;td&gt;🟢 &lt;code&gt;az&lt;/code&gt; token&lt;/td&gt;
&lt;td&gt;7 reads&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;❌ timed out&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS Glue&lt;/td&gt;
&lt;td&gt;❌ SigV4&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS S3 Tables&lt;/td&gt;
&lt;td&gt;❌ SigV4&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Databricks Unity&lt;/td&gt;
&lt;td&gt;token, not run&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Snowflake Horizon&lt;/td&gt;
&lt;td&gt;key-pair JWT, not run&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Behind those numbers every catalog that answered showed the same shape: 14 endpoints the client has no method for, and every endpoint it does implement returning an answer.&lt;/p&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to point the Apache Rust Iceberg REST client at seven catalogs and record what works and what it takes. The key to the solution was matching every test to a line in the client's source code, running the same tests against each catalog, and logging every request the client sent. The results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 Every endpoint the client implements answered on every catalog it could log in to: 7 reads on Polaris, BigLake and OneLake, and 11 writes on Polaris, with 0 failures&lt;/li&gt;
&lt;li&gt;🟢 13 of 25 endpoints are implemented; of the other 12, 11 are missing and 1 is a stub that sends no request&lt;/li&gt;
&lt;li&gt;⚠️ Two lines of &lt;code&gt;Cargo.toml&lt;/code&gt; stand in front of that: a TLS backend (upstream #2888) and &lt;code&gt;iceberg-storage-opendal&lt;/code&gt; for cloud storage&lt;/li&gt;
&lt;li&gt;❌ Glue and S3 Tables require SigV4, which this client cannot send, so a Rust tool covering all seven catalogs loads three catalog crates where Python loads one — and those three disagree on three of the 15 &lt;code&gt;Catalog&lt;/code&gt; methods, at runtime (upstream #1236, open since April 2025)&lt;/li&gt;
&lt;li&gt;❌ OneLake's files could not be read: the Azure backend cannot use an &lt;code&gt;az login&lt;/code&gt;, and a credential from the catalog is read and then ignored (upstream #2931, #1442)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: &lt;code&gt;iceberg-catalog-rest&lt;/code&gt; 0.10.1 with &lt;code&gt;iceberg&lt;/code&gt; 0.10.1, &lt;code&gt;iceberg-storage-opendal&lt;/code&gt; 0.10.1 and &lt;code&gt;reqwest&lt;/code&gt; 0.12.28 with &lt;code&gt;rustls-tls&lt;/code&gt;, built with &lt;code&gt;rustc&lt;/code&gt; 1.98.1. Source read 2026-09-04, versions captured 2026-09-17, catalog runs 2026-09-18, write run 2026-09-21, one run each from one machine in one region. Polaris 1.7.0 ran in Docker with permissive settings and local file storage, and the writes ran there alone, so those results describe one permissive server. Glue and S3 Tables were tested only with the signature experiment, and Unity and Horizon were not run at all, their rows coming from the source. The tests check that each operation answers; the answers themselves are not checked. Three of the seven catalogs were on trial accounts in the earlier article, and managed catalogs do not report a version.&lt;/p&gt;

&lt;p&gt;The strategy for testing what one Rust Iceberg client can reach was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/xbill9/lakehouse-iceberg-2026" rel="noopener noreferrer"&gt;lakehouse-iceberg-2026 | GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust" rel="noopener noreferrer"&gt;apache/iceberg-rust&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/1236" rel="noopener noreferrer"&gt;iceberg-rust #1236 — REST catalog: support AWS sigV4&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/2888" rel="noopener noreferrer"&gt;iceberg-rust #2888 — Add TLS features to iceberg-catalog-rest&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/2931" rel="noopener noreferrer"&gt;iceberg-rust #2931 — Support refreshing vended storage credentials for REST catalog tables&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/1442" rel="noopener noreferrer"&gt;iceberg-rust #1442 — ADLS: Support vended "adls.sas-token.xxx" prefixed tokens&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/apache/iceberg-rust/issues/3134" rel="noopener noreferrer"&gt;iceberg-rust #3134 — Transaction commits against a base it never validated&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://crates.io/crates/iceberg-catalog-glue" rel="noopener noreferrer"&gt;iceberg-catalog-glue | crates.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://crates.io/crates/iceberg-catalog-s3tables" rel="noopener noreferrer"&gt;iceberg-catalog-s3tables | crates.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/gde/seven-iceberg-rest-catalogs-what-they-declare-and-what-they-serve-40oj"&gt;Seven Iceberg REST Catalogs: What They Declare, and What They Serve&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>rust</category>
      <category>iceberg</category>
      <category>lakehouse</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Running a Jev-Style Decision Model on One TPU v6e: What Fits, What It Costs, and What Changes From a GPU</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Thu, 24 Sep 2026 16:32:26 +0000</pubDate>
      <link>https://dev.to/gde/running-a-jev-style-decision-model-on-one-tpu-v6e-what-fits-what-it-costs-and-what-changes-from-1j0g</link>
      <guid>https://dev.to/gde/running-a-jev-style-decision-model-on-one-tpu-v6e-what-fits-what-it-costs-and-what-changes-from-1j0g</guid>
      <description>&lt;p&gt;This article provides a step by step guide to running a Jev-style decision model on one Google Cloud TPU v6e chip with Gemma 4 and vLLM, and compares it with the same read on one NVIDIA L4 GPU. The measurement was pre-registered, and every per-item output is committed.&lt;/p&gt;

&lt;p&gt;One v6e chip serves Gemma 4 E2B, E4B and 12B at bf16 and a 26B-A4B fp8 build; no 31B checkpoint loads. The same checkpoints give the same answers on the TPU and the L4, and the 26B decides in 27 to 33 ms against 61 ms on the L4. On demand, the TPU costs more per decision than the L4 or Jev; 12B is the size to pick.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/gemma4-dev/tree/main/jev-tpu" rel="noopener noreferrer"&gt;https://github.com/xbill9/gemma4-dev/tree/main/jev-tpu&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  Why Measure This?
&lt;/h4&gt;

&lt;p&gt;A Jev-style decision model answers a typed question with a probability for each allowed option: end the prompt where the answer starts, read the scores of the allowed label tokens, and apply a softmax over those. TypeSafe's Jev does this as a hosted service, and any open model can be read the same way.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://dev.to/gde/plain-gemma-4-26b-vs-jev-on-one-ec2-l4-21-points-behind-overall-level-on-yesno-45-behind-on-15k6"&gt;companion article&lt;/a&gt; measured this read on one NVIDIA L4 and set it beside Jev's published results. This one asks what changes on a TPU: which Gemma 4 sizes fit one v6e chip, whether the read works the same way, and what the chip buys in speed and cost.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;A Google Cloud project with v6e quota in a region that offers &lt;code&gt;ct6e-standard-1t&lt;/code&gt;, and the &lt;code&gt;gcloud&lt;/code&gt; CLI logged in&lt;/li&gt;
&lt;li&gt;A Hugging Face token stored as the Secret Manager secret &lt;code&gt;hf-token&lt;/code&gt;, readable by the Compute Engine default service account&lt;/li&gt;
&lt;li&gt;A Cloud Storage bucket for results, set as &lt;code&gt;BUCKET&lt;/code&gt; in &lt;code&gt;tpu/run.sh&lt;/code&gt; and &lt;code&gt;tpu/startup.sh&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The repository cloned: &lt;code&gt;git clone https://github.com/xbill9/gemma4-dev&lt;/code&gt; and &lt;code&gt;cd gemma4-dev/jev-tpu&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Step 1 — Pre-Register the Measurement
&lt;/h4&gt;

&lt;p&gt;The models, image, serving flags, data, metrics and the order of the quantized attempts are written down and committed before any model call, in &lt;code&gt;PREREGISTRATION.md&lt;/code&gt;; each departure from it is recorded there with a date before the affected results are scored.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git show &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;--oneline&lt;/span&gt; e0cce69 4937fb8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;e0cce69 jev-tpu: sibling of jev for one TPU v6e chip — copied read, data, suite and scoring code (proxy byte-identical); pre-registration for E2B/E4B/12B bf16 and exploratory quantized 26B-A4B/31B probes
4937fb8 jev-tpu: VM boot, serve and run drivers; pre-registration amended before any model call — suite built locally at 0e67403 (13/13 checksums) and re-checked on the VM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 2 — Pick Checkpoints That Fit One Chip
&lt;/h4&gt;

&lt;p&gt;One v6e chip has 31.24 GiB of memory, of which vLLM uses up to 28.74 GiB:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Memory statistics | total_hbm_limit_gb=31.24GiB | total_hbm_limit_cap_gb=28.74GiB | total_hbm_used_gb=24.56GiB | total_hbm_avail_gb=4.19GiB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;E2B, E4B and 12B fit at bf16. The 26B-A4B and 31B do not, so the pre-registration lists quantized builds to try in order:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Checkpoint&lt;/th&gt;
&lt;th&gt;Format&lt;/th&gt;
&lt;th&gt;Size on disk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;E2B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;google/gemma-4-E2B-it&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E4B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;google/gemma-4-E4B-it&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;google/gemma-4-12B-it&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;26B-A4B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;fp8&lt;/td&gt;
&lt;td&gt;26.67 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;31B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;google/gemma-4-31B-it-qat-w4a16-ct&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;w4a16&lt;/td&gt;
&lt;td&gt;21.67 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;31B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;cyankiwi/gemma-4-31B-it-AWQ-4bit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4-bit&lt;/td&gt;
&lt;td&gt;19.47 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 31B fp8 builds are 30.98 GiB, over the chip's usable memory, and were not tried.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 3 — Launch One v6e Chip
&lt;/h4&gt;

&lt;p&gt;The VM runs the whole measurement from its startup script, fetching the code and the suite from Cloud Storage, and deletes itself when the run ends.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gcloud compute instances create jev-tpu-v6e1 &lt;span class="nt"&gt;--zone&lt;/span&gt; europe-west4-a &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--machine-type&lt;/span&gt; ct6e-standard-1t &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--image-family&lt;/span&gt; ubuntu-accel-2204-amd64-tpu-v5e-v5p-v6e &lt;span class="nt"&gt;--image-project&lt;/span&gt; ubuntu-os-accelerator-images &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--boot-disk-size&lt;/span&gt; 200GB &lt;span class="nt"&gt;--scopes&lt;/span&gt; cloud-platform &lt;span class="nt"&gt;--maintenance-policy&lt;/span&gt; TERMINATE &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--provisioning-model&lt;/span&gt; STANDARD &lt;span class="nt"&gt;--max-run-duration&lt;/span&gt; 6h &lt;span class="nt"&gt;--instance-termination-action&lt;/span&gt; DELETE &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--metadata&lt;/span&gt; jev-code&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;code-tarball&amp;gt;,jev-run&lt;span class="o"&gt;=&lt;/span&gt;2026-09-24-v6e1 &lt;span class="nt"&gt;--metadata-from-file&lt;/span&gt; startup-script&lt;span class="o"&gt;=&lt;/span&gt;tpu/startup.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NAME          ZONE            MACHINE_TYPE      PREEMPTIBLE  INTERNAL_IP    EXTERNAL_IP    STATUS
jev-tpu-v6e1  europe-west4-a  ct6e-standard-1t               10.164.15.207  34.32.155.166  RUNNING
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--scopes cloud-platform&lt;/code&gt; lets the VM read the Hugging Face token from Secret Manager at boot, so the token never enters instance metadata.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — Serve Each Model
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;tpu/serve.sh&lt;/code&gt; starts one model at a time with the same flags for every arm and waits for &lt;code&gt;/v1/models&lt;/code&gt;. Abridged, with the token and volume options left out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; vllm &lt;span class="nt"&gt;--privileged&lt;/span&gt; &lt;span class="nt"&gt;--net&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;host &lt;span class="nt"&gt;--shm-size&lt;/span&gt; 10gb vllm/vllm-tpu:nightly &lt;span class="se"&gt;\&lt;/span&gt;
  vllm serve google/gemma-4-12B-it &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 1 &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 2048 &lt;span class="nt"&gt;--max-num-seqs&lt;/span&gt; 16 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-logprobs&lt;/span&gt; 32 &lt;span class="nt"&gt;--generation-config&lt;/span&gt; vllm &lt;span class="nt"&gt;--limit-mm-per-prompt&lt;/span&gt; &lt;span class="s1"&gt;'{"image":0,"audio":0}'&lt;/span&gt; &lt;span class="nt"&gt;--enable-prefix-caching&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[jev-run 2026-09-24T14:42:48Z] READY google/gemma-4-E2B-it after 346s
[jev-run 2026-09-24T14:51:34Z] READY google/gemma-4-E4B-it after 421s
[jev-run 2026-09-24T15:02:01Z] READY google/gemma-4-12B-it after 512s
[jev-run 2026-09-24T15:14:59Z] READY RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic after 616s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The image was &lt;code&gt;vllm/vllm-tpu@sha256:19a1a0526476f902eb83e1057f3d8938f35b457dd7716a30e9ab4f7bee90d507&lt;/code&gt;, vLLM &lt;code&gt;0.29.1rc1.dev468+g0b7f11a1e&lt;/code&gt;. On v6e it stores the KV cache in fp8 by default for every model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;INFO 09-24 14:53:37 [tpu_platform.py:232] Automatically using fp8_e5m2 for FP8 KV cache on TPU v6e.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So "bf16" below describes the weights. The L4 run kept the default 16-bit KV cache.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 5 — Read the Labels Out of the Top 32
&lt;/h4&gt;

&lt;p&gt;On a GPU, the read asks vLLM for the log-probabilities of the label tokens by id. vLLM's TPU backend returns only the top-k log-probabilities, and a request for specific token ids fails. Four request shapes are sent before each model's reads, and the answers are kept:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;== probe 3: {"prompt":[2,106,1645,108],"max_tokens":1,"temperature":0,"logprobs":32,"return_tokens_as_token_ids":true}
{"id":"cmpl-9ee6dc5f0461280b","object":"text_completion", ...
== probe 4: {"prompt":[2,106,1645,108],"max_tokens":1,"temperature":0,"logprobs":5,"logprob_token_ids":[236776,236799,236780],"return_tokens_as_token_ids":true}
{"error":{"message":"list index out of range","type":"InternalServerError","param":null,"code":500}}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the read asks for the top 32 log-probabilities (&lt;code&gt;JEV_TOPK=32&lt;/code&gt;) and takes the label tokens from them. A label inside the 32 gets its exact log-probability; a label outside gets the proxy's fallback value, the lowest returned log-probability minus 5, and every four-task record keeps how many labels came back.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 6 — Run and Score
&lt;/h4&gt;

&lt;p&gt;Each model gets a five-example smoke read, the four tasks (300 examples each from sst2, AG News, DAIR Emotion and tweet_eval irony), the three choice tasks with their options reversed, a latency pass at one request at a time, and Bespoke Labs' 3,880-record public suite, rebuilt with Nimble's converters and checked against all 13 published checksums on the VM.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[jev-run 2026-09-24T14:36:51Z] suite: 13 of 13 subsets match
[jev-run 2026-09-24T15:02:42Z] 12b four tasks: 23s for 1200 decisions at concurrency 8
[jev-run 2026-09-24T15:04:25Z] 12b suite: 60s for 3880 records
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 sweep_score.py &lt;span class="nt"&gt;--prefix&lt;/span&gt; 2026-09-24-v6e1 &lt;span class="nt"&gt;--arms&lt;/span&gt; e2b e4b 12b 26b-fp8
python3 nimble_suite/suite_sweep.py &lt;span class="nt"&gt;--prefix&lt;/span&gt; 2026-09-24-v6e1 &lt;span class="nt"&gt;--arms&lt;/span&gt; e2b e4b 12b 26b-fp8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both write their tables into &lt;code&gt;results/&lt;/code&gt;, and every figure below comes from those files and &lt;code&gt;article_figures.py&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Fits and Loads on One v6e Chip?
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Checkpoint&lt;/th&gt;
&lt;th&gt;Serves&lt;/th&gt;
&lt;th&gt;Time to serve&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;E2B&lt;/td&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;346 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E4B&lt;/td&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;421 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12B&lt;/td&gt;
&lt;td&gt;bf16&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;512 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;26B-A4B&lt;/td&gt;
&lt;td&gt;fp8, 26.67 GiB&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;616 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;31B&lt;/td&gt;
&lt;td&gt;w4a16, 21.67 GiB&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;fails at 120 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;31B&lt;/td&gt;
&lt;td&gt;4-bit, 19.47 GiB&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;fails at 120 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both 31B builds stop with the same error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NotImplementedError: compressed-tensors scheme for layer 'model.language_model.layers.0.self_attn.q_proj' is not yet supported in the JAX path.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gemma 4 runs only on the JAX path of vLLM's TPU backend, which loads fp8 but no 4-bit format for a dense model. The one 31B format it loads is 30.98 GiB, over one chip's usable memory.&lt;/p&gt;




&lt;h4&gt;
  
  
  Does the TPU Change the Answers?
&lt;/h4&gt;

&lt;p&gt;No. E2B and E4B were also read on the L4, with the same checkpoints and prompts. The predicted label matched on 1,186 and 1,193 of 1,200 examples, and accuracy per task differed by at most 1.0 and 0.4 points. The 26B-A4B fp8 build here and the 4-bit AWQ build on the L4 are different quantizations of the same model and differ by at most 1.3 points per task and 0.7 points on the suite.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Majority&lt;/th&gt;
&lt;th&gt;E2B&lt;/th&gt;
&lt;th&gt;E4B&lt;/th&gt;
&lt;th&gt;12B&lt;/th&gt;
&lt;th&gt;26B-A4B fp8&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;sst2&lt;/td&gt;
&lt;td&gt;51.0%&lt;/td&gt;
&lt;td&gt;88.7%&lt;/td&gt;
&lt;td&gt;94.0%&lt;/td&gt;
&lt;td&gt;95.0%&lt;/td&gt;
&lt;td&gt;94.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AG News&lt;/td&gt;
&lt;td&gt;25.0%&lt;/td&gt;
&lt;td&gt;30.0%&lt;/td&gt;
&lt;td&gt;83.3%&lt;/td&gt;
&lt;td&gt;86.3%&lt;/td&gt;
&lt;td&gt;86.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DAIR Emotion&lt;/td&gt;
&lt;td&gt;35.0%&lt;/td&gt;
&lt;td&gt;53.3%&lt;/td&gt;
&lt;td&gt;54.3%&lt;/td&gt;
&lt;td&gt;59.3%&lt;/td&gt;
&lt;td&gt;58.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tweet_eval irony&lt;/td&gt;
&lt;td&gt;60.3%&lt;/td&gt;
&lt;td&gt;73.7%&lt;/td&gt;
&lt;td&gt;84.0%&lt;/td&gt;
&lt;td&gt;84.7%&lt;/td&gt;
&lt;td&gt;91.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Per-task calibration, option-order changes and 90th-percentile times are in &lt;code&gt;results/2026-09-24-v6e1-SWEEP.md&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  How Do the Sizes Compare With Jev?
&lt;/h4&gt;

&lt;p&gt;The Jev comparison is the &lt;a href="https://dev.to/gde/plain-gemma-4-26b-vs-jev-on-one-ec2-l4-21-points-behind-overall-level-on-yesno-45-behind-on-15k6"&gt;L4 article's&lt;/a&gt; subject, with its caveats in full. On the same 3,880 public records:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Questions&lt;/th&gt;
&lt;th&gt;Jev 1.13.0, published&lt;/th&gt;
&lt;th&gt;26B-A4B fp8&lt;/th&gt;
&lt;th&gt;12B&lt;/th&gt;
&lt;th&gt;E4B&lt;/th&gt;
&lt;th&gt;E2B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;All, 3,880&lt;/td&gt;
&lt;td&gt;77.3%&lt;/td&gt;
&lt;td&gt;76.0%&lt;/td&gt;
&lt;td&gt;76.2%&lt;/td&gt;
&lt;td&gt;72.8%&lt;/td&gt;
&lt;td&gt;68.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Yes/no, 1,399&lt;/td&gt;
&lt;td&gt;84.6%&lt;/td&gt;
&lt;td&gt;84.2%&lt;/td&gt;
&lt;td&gt;85.1%&lt;/td&gt;
&lt;td&gt;81.5%&lt;/td&gt;
&lt;td&gt;76.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multiple choice, 1,848&lt;/td&gt;
&lt;td&gt;82.8%&lt;/td&gt;
&lt;td&gt;79.7%&lt;/td&gt;
&lt;td&gt;78.1%&lt;/td&gt;
&lt;td&gt;74.5%&lt;/td&gt;
&lt;td&gt;70.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Five-level rating, 633&lt;/td&gt;
&lt;td&gt;45.2%&lt;/td&gt;
&lt;td&gt;47.4%&lt;/td&gt;
&lt;td&gt;51.0%&lt;/td&gt;
&lt;td&gt;48.5%&lt;/td&gt;
&lt;td&gt;44.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median ECE after 50 labels (Jev 0.071 as shipped)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;0.070&lt;/td&gt;
&lt;td&gt;0.070&lt;/td&gt;
&lt;td&gt;0.077&lt;/td&gt;
&lt;td&gt;0.092&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Jev leads 12B by 1.1 points overall (95% range −0.8 to +3.0) and 26B-A4B by 1.3 (−0.6 to +3.2), and on multiple choice by 4.7 and 3.2. The L4's 4-bit 26B scored 75.3%, 2.1 points behind with a range from 0.2 to 4.0, so read the 26B as one to two points behind Jev on either platform. Jev's per-record answers are unpublished, so these ranges compare two independent proportions, and records that share a passage widen them. Gemma was read in the PR proxy's prompt format with no tuning; a different prompt could move the multiple-choice gap in either direction.&lt;/p&gt;




&lt;h4&gt;
  
  
  Did the Top-32 Read Lose Anything?
&lt;/h4&gt;

&lt;p&gt;Not on accuracy. At least one label came back on every four-task read, so the predicted label is exact; only the probabilities of labels outside the 32 are approximate.&lt;/p&gt;

&lt;p&gt;On the four tasks, moving every missing label from the fallback value to the top-32 bound, the most probability it could have had, changes raw ECE by less than 0.001 and ECE after 50 labels by at most 0.005. That check was added after the run. The suite records do not keep how many labels came back, so the suite calibration figures have no such check; by a lower-bound count, at least 10.2% of 12B's suite records and 17.8% of 26B-A4B's used the fallback value.&lt;/p&gt;

&lt;p&gt;The share of probability on the allowed labels varies too. On irony, 12B's most likely token was a label on 5 of 300 reads, with a median of 3.1% of its probability on the two labels; 26B-A4B's was a label on 161 of 300 AG News reads, with a median of 55.2%. The read rescales the labels to sum to one, so the answer reads as confident either way, and &lt;code&gt;label_mass&lt;/code&gt; in every record shows it.&lt;/p&gt;




&lt;h4&gt;
  
  
  TPU v6e Against the L4
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;One TPU v6e chip&lt;/th&gt;
&lt;th&gt;One NVIDIA L4 (&lt;code&gt;g6.xlarge&lt;/code&gt;)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Memory for the model&lt;/td&gt;
&lt;td&gt;🥇 28.74 GiB&lt;/td&gt;
&lt;td&gt;23034 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Largest Gemma 4 read here&lt;/td&gt;
&lt;td&gt;12B at bf16, 26B-A4B at fp8&lt;/td&gt;
&lt;td&gt;26B-A4B at 4-bit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy, same checkpoint&lt;/td&gt;
&lt;td&gt;same answers on 1,186 to 1,193 of 1,200&lt;/td&gt;
&lt;td&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;26B time per decision, one at a time&lt;/td&gt;
&lt;td&gt;🥇 27 to 33 ms&lt;/td&gt;
&lt;td&gt;61 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Price per hour, on demand&lt;/td&gt;
&lt;td&gt;$2.97&lt;/td&gt;
&lt;td&gt;🥇 $0.8048&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;26B cost per million decisions, on demand&lt;/td&gt;
&lt;td&gt;$10.37&lt;/td&gt;
&lt;td&gt;🥇 at most $5.43&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Label read&lt;/td&gt;
&lt;td&gt;top 32 log-probabilities&lt;/td&gt;
&lt;td&gt;🥇 exact label token ids&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache&lt;/td&gt;
&lt;td&gt;fp8 by default&lt;/td&gt;
&lt;td&gt;🥇 16-bit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In this table 🥇 marks the better value. The 26B decides 1.8 to 2.2 times faster on the TPU. The chip costs 3.7 times as much per hour on demand, or 1.7 times at the flex-start price of $1.35, a mode that had no capacity during this run. The L4's cost comes from its first run's throughput with the client on a home connection, so it is an upper bound; the TPU's comes from the request times at eight in flight.&lt;/p&gt;




&lt;h4&gt;
  
  
  🔎 Tip: Serve the Same Checkpoint on Both Platforms First
&lt;/h4&gt;

&lt;p&gt;Before comparing hardware, serve one checkpoint on both with the same prompts and compare predicted labels record by record. It checks the whole read, from chat template to label tokens; here the TPU's top-32 read and the L4's exact read agreed on 1,186 and 1,193 of 1,200 answers.&lt;/p&gt;




&lt;h4&gt;
  
  
  So, Which One?
&lt;/h4&gt;

&lt;p&gt;For a Jev-style decision service on a TPU, Gemma 4 12B at bf16 plus one temperature fitted on about 50 labels. It matches the 26B on the suite, 76.2% against 76.0%, though the 26B scores 6.3 points higher on irony; it reaches a median ECE of 0.070 after fitting, decides in 26 to 27 ms, and needs no third-party quantized checkpoint. Check &lt;code&gt;label_mass&lt;/code&gt; on your own task first: on irony, 12B put a median of 3.1% of its probability on the labels.&lt;/p&gt;

&lt;p&gt;Choose the TPU when time per decision matters or the rest of the stack is already on Google Cloud: it halves the 26B's time per decision against the L4. Choose the L4, or Jev itself, when cost per decision matters: on demand, 12B on one v6e chip costs $8.19 per million decisions, against Jev's $5.54 and at most $5.43 for the 26B on an L4.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Does It Cost?
&lt;/h4&gt;

&lt;p&gt;Measured from the request times at eight in flight, one v6e chip answers about 274 decisions a second with E2B, 206 with E4B, 101 with 12B and 80 with 26B-A4B. At the europe-west4 on-demand price of $2.97 an hour that is $3.01, $4.00, $8.19 and $10.37 per million decisions; at the flex-start price of $1.35, $1.37, $1.82, $3.72 and $4.71. TypeSafe prices Jev at $5.54 per million decisions at the L4 run's median prompt of 132 tokens. The chip is charged by the hour whether busy or idle, and eight requests in flight is below the 16 the server allows, so these figures hold only while the chip is kept at least this busy.&lt;/p&gt;

&lt;p&gt;The measured run took 48 minutes of on-demand time, $2.39, and all instance time for the measurement came to at most $3.25.&lt;/p&gt;




&lt;h4&gt;
  
  
  Teardown
&lt;/h4&gt;

&lt;p&gt;The VM deletes itself when &lt;code&gt;tpu/run.sh&lt;/code&gt; finishes, and &lt;code&gt;--max-run-duration 6h&lt;/code&gt; with &lt;code&gt;--instance-termination-action DELETE&lt;/code&gt; is the backstop. Confirm it is gone:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gcloud compute instances describe jev-tpu-v6e1 &lt;span class="nt"&gt;--zone&lt;/span&gt; europe-west4-a
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR: (gcloud.compute.instances.describe) Could not fetch resource:
 - The resource 'projects/aisprint-491218/zones/europe-west4-a/instances/jev-tpu-v6e1' was not found
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to run a Jev-style decision model on one TPU v6e chip with Gemma 4 and find what changes from a GPU. The key to the solution was the same prompts, labels and scoring as the L4 run, a pre-registered design, and a label read taken from the top 32 log-probabilities. The results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 One v6e chip serves Gemma 4 E2B, E4B and 12B at bf16 and a 26B-A4B fp8 build&lt;/li&gt;
&lt;li&gt;❌ No 31B checkpoint loads: the w4a16 and 4-bit builds fail on the JAX path, and the fp8 build is 30.98 GiB&lt;/li&gt;
&lt;li&gt;🟢 The same checkpoints give the same answers on the TPU and the L4: 1,186 and 1,193 of 1,200&lt;/li&gt;
&lt;li&gt;🟢 The 26B decides in 27 to 33 ms on the TPU against 61 ms on the L4&lt;/li&gt;
&lt;li&gt;⚠️ On demand, 26B costs $10.37 per million decisions on the TPU against at most $5.43 on the L4 and $5.54 for Jev&lt;/li&gt;
&lt;li&gt;⚠️ vLLM on TPU returns only the top-k log-probabilities; reading labels from the top 32 changes no answer, and changes four-task calibration by at most 0.005&lt;/li&gt;
&lt;li&gt;🟢 12B matches the 26B: 76.2% and 76.0% on the 3,880-record suite, against Jev's 77.3%&lt;/li&gt;
&lt;li&gt;🟢 After 50 labels, 12B and 26B-A4B reach a median ECE of 0.070, against Jev's 0.071 as shipped&lt;/li&gt;
&lt;li&gt;🟢 The whole measurement cost at most $3.25 of instance time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: one TPU v6e chip (&lt;code&gt;ct6e-standard-1t&lt;/code&gt;, on demand) in europe-west4-a, vLLM &lt;code&gt;0.29.1rc1.dev468+g0b7f11a1e&lt;/code&gt; at the image digest above, E2B, E4B and 12B with bf16 weights and 26B-A4B through RedHatAI's fp8 build, every model with vLLM's default fp8 KV cache on v6e, &lt;code&gt;--max-model-len 2048&lt;/code&gt;, one run per model, 300 examples per task plus the 3,880-record public suite, client on the VM. The top-32 read and on-demand capacity are dated deviations from the pre-registration, and the top-32 bound check, added after the run, covers the four tasks only. L4 figures come from the companion run on a &lt;code&gt;g6.xlarge&lt;/code&gt; and &lt;code&gt;g6.4xlarge&lt;/code&gt; in us-east-1, and the two platforms were timed with different clients. Every source dataset was published before Gemma 4 and may be in its training data. No Jev call was made: the Jev figures are Bespoke Labs' published results on the same records, from one run of Jev 1.13.0 by a company that sells a competing model, counting an invalid Jev response as wrong, with Jev's probabilities rounded to two decimals by its API. Parts of the analysis and writing were done with AI assistance (Claude); every figure comes from the committed output files.&lt;/p&gt;

&lt;p&gt;The strategy for running a Jev-style decision model on TPU with Gemma 4 was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Code, pre-registration and per-item results: &lt;a href="https://github.com/xbill9/gemma4-dev/tree/main/jev-tpu" rel="noopener noreferrer"&gt;https://github.com/xbill9/gemma4-dev/tree/main/jev-tpu&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Companion L4 measurement: &lt;a href="https://dev.to/gde/plain-gemma-4-26b-vs-jev-on-one-ec2-l4-21-points-behind-overall-level-on-yesno-45-behind-on-15k6"&gt;https://dev.to/gde/plain-gemma-4-26b-vs-jev-on-one-ec2-l4-21-points-behind-overall-level-on-yesno-45-behind-on-15k6&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Companion review of the independent evidence on Jev: &lt;a href="https://dev.to/gde/jev-after-eight-days-of-independent-tests-level-with-mid-price-llms-behind-the-frontier-1kln"&gt;https://dev.to/gde/jev-after-eight-days-of-independent-tests-level-with-mid-price-llms-behind-the-frontier-1kln&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;vLLM TPU documentation: &lt;a href="https://docs.vllm.ai/projects/tpu/en/latest/" rel="noopener noreferrer"&gt;https://docs.vllm.ai/projects/tpu/en/latest/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;tpu-inference: &lt;a href="https://github.com/vllm-project/tpu-inference" rel="noopener noreferrer"&gt;https://github.com/vllm-project/tpu-inference&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;26B-A4B fp8 checkpoint: &lt;a href="https://huggingface.co/RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic" rel="noopener noreferrer"&gt;https://huggingface.co/RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gemma 4 12B: &lt;a href="https://huggingface.co/google/gemma-4-12B-it" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-12B-it&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Bespoke Labs, public-suite results for Jev 1.13.0: &lt;a href="https://github.com/bespokelabsai/nimble/blob/0e67403/docs/PUBLIC_BENCHMARKS.md" rel="noopener noreferrer"&gt;https://github.com/bespokelabsai/nimble/blob/0e67403/docs/PUBLIC_BENCHMARKS.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Cloud TPU v6e: &lt;a href="https://cloud.google.com/tpu/docs/v6e" rel="noopener noreferrer"&gt;https://cloud.google.com/tpu/docs/v6e&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Guo et al., On Calibration of Modern Neural Networks: &lt;a href="https://arxiv.org/abs/1706.04599" rel="noopener noreferrer"&gt;https://arxiv.org/abs/1706.04599&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>gemma</category>
      <category>googlecloud</category>
      <category>machinelearning</category>
      <category>llm</category>
    </item>
    <item>
      <title>What Nobody Is Using in Your Azure Subscriptions, and What It Costs</title>
      <dc:creator>xbill</dc:creator>
      <pubDate>Thu, 24 Sep 2026 15:32:30 +0000</pubDate>
      <link>https://dev.to/xbill/what-nobody-is-using-in-your-azure-subscriptions-and-what-it-costs-205h</link>
      <guid>https://dev.to/xbill/what-nobody-is-using-in-your-azure-subscriptions-and-what-it-costs-205h</guid>
      <description>&lt;p&gt;This article provides a step by step guide to building an Azure waste scanner from source, running it across every subscription your &lt;code&gt;az login&lt;/code&gt; can see, pricing each finding from the Azure Retail Prices API, and drafting the cleanup. A suite of Python checks is built to cover Compute, Networking, Storage, SQL, Key Vault, Container Registry, App Service, Container Apps, AI Services, Machine Learning, Log Analytics and AKS.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/xbill9/zombiescan-azure" rel="noopener noreferrer"&gt;https://github.com/xbill9/zombiescan-azure&lt;/a&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  The Third Cloud
&lt;/h4&gt;

&lt;p&gt;This is the third scanner in a series. The first covered an AWS account and the second covered Google Cloud projects:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AWS: &lt;a href="https://dev.to/aws-builders/find-the-aws-resources-nobody-is-using-and-what-they-cost-you-35bj"&gt;https://dev.to/aws-builders/find-the-aws-resources-nobody-is-using-and-what-they-cost-you-35bj&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Google Cloud: &lt;a href="https://dev.to/gde/what-nobody-is-using-in-your-google-cloud-projects-and-what-it-costs-1k0"&gt;https://dev.to/gde/what-nobody-is-using-in-your-google-cloud-projects-and-what-it-costs-1k0&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The shape carries over: a CLI, a bundled price table, a cleanup plan that is printed and never run, and a Claude Code plugin with an MCP server over the same engine. What changes is everything underneath — how Azure authenticates, how it answers a list call, how it prices a disk, and which resources keep billing once nobody uses them.&lt;/p&gt;




&lt;h4&gt;
  
  
  What Gets Left Behind
&lt;/h4&gt;

&lt;p&gt;A managed disk survives the VM it was attached to. A public IP outlives the load balancer that held it. A NAT gateway keeps its hourly fee after the last subnet moved away, and a provisioned model deployment bills every PTU every hour whether a request arrives or not. Each one reports nothing.&lt;/p&gt;

&lt;p&gt;Cost Management shows the total. Azure Advisor surfaces candidates. A figure that drives a decision names one resource, in one subscription and resource group, and its monthly cost.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;zombiescan&lt;/code&gt; produces that figure for 30 classes of resource across two packs.&lt;/p&gt;




&lt;h4&gt;
  
  
  At This Point You Should Have…
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;An Azure account and the Azure CLI, signed in with &lt;code&gt;az login&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;uv&lt;/code&gt; on the path, and Python 3.11 or newer&lt;/li&gt;
&lt;li&gt;Read access to the subscriptions you mean to scan — the built-in Reader role covers every call, and &lt;code&gt;zombiescan-scanner-role.json&lt;/code&gt;, under &lt;code&gt;policy/&lt;/code&gt; in the repository, is a custom role holding exactly the read actions the checks use&lt;/li&gt;
&lt;/ul&gt;




&lt;h4&gt;
  
  
  Step 1 — Build It From Source
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/xbill9/zombiescan-azure
&lt;span class="nb"&gt;cd &lt;/span&gt;zombiescan-azure
uv &lt;span class="nb"&gt;sync&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Resolved 18 packages in 0.48ms
Checked 17 packages in 0.13ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;uv sync&lt;/code&gt; reads the committed &lt;code&gt;uv.lock&lt;/code&gt;, so the resolved set is the one the tests ran against. The dependency list is &lt;code&gt;click&lt;/code&gt; and &lt;code&gt;rich&lt;/code&gt;; every Azure call goes over HTTPS with &lt;code&gt;urllib&lt;/code&gt; from the standard library.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run zombiescan &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;zombiescan, version 0.1.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h4&gt;
  
  
  Step 2 — Install It as a Tool
&lt;/h4&gt;

&lt;p&gt;Working from the clone keeps &lt;code&gt;uv run&lt;/code&gt; in front of every command. Installing it puts &lt;code&gt;zombiescan&lt;/code&gt; on the path instead.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv tool &lt;span class="nb"&gt;install &lt;/span&gt;git+https://github.com/xbill9/zombiescan-azure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both routes run the same engine. The rest of this article uses the installed form.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 3 — Authenticate, and See Which Account You Are
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;zombiescan&lt;/code&gt; shells out to &lt;code&gt;az account get-access-token&lt;/code&gt; and reuses the result. No key file is downloaded, no client secret is stored, and there is no &lt;code&gt;azure-identity&lt;/code&gt; dependency.&lt;/p&gt;

&lt;p&gt;An ARM token is issued for exactly one tenant. Microsoft lets one email address be both a work or school account and a personal Microsoft account, and those are two directories with two sets of subscriptions. So the scanner reads the subscription list from &lt;code&gt;az account list --all&lt;/code&gt;, which spans every identity the CLI has signed into, and holds one token per tenant.&lt;/p&gt;

&lt;p&gt;Every scan opens by naming each tenant, the kind of account signed into it, and each subscription:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Scanning as xbill@glitnir.com
Tenant Default Directory — personal Microsoft account
  domain        xbillglitnircom.onmicrosoft.com
  tenant id     40482c55-d00d-4c6d-8903-643d76a74b9c
  signed in as  xbill@glitnir.com
  subscription  Azure subscription 1 (default)
                3db3ce66-50b6-4d11-91ef-5950cf4039ed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;az account list&lt;/code&gt; reports a personal account and a work account as the same user with the same email. The access token tells them apart: a personal Microsoft account signs in through &lt;code&gt;live.com&lt;/code&gt; and its token carries &lt;code&gt;idp: live.com&lt;/code&gt;, a work or school account in its own directory carries no &lt;code&gt;idp&lt;/code&gt; at all, and a guest from another directory names its home issuer there.&lt;/p&gt;

&lt;p&gt;🔎 Tip: every token is fetched on the main thread before any worker starts. Concurrent &lt;code&gt;az account get-access-token&lt;/code&gt; processes contend on the MSAL token cache in &lt;code&gt;~/.azure&lt;/code&gt;, and a corrupted cache costs an &lt;code&gt;az login&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 4 — List the Checks
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;zombiescan checks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  aks-idle-cluster  AKS clusters running no nodes  (aks)
  deallocated-vm  Stopped VMs still paying for disks  (core)
  disabled-key-vault-key  Disabled Key Vault keys still billed  (core)
  empty-ai-services-account  AI Services accounts with no model deployed  (core)
  empty-container-apps-environment  Container Apps environments with no app  (core)
  empty-container-registry  Container registries holding no images  (core)
  empty-resource-group  Resource groups holding nothing  (core)
  empty-vnet  Virtual networks with nothing running in them  (core)
  idle-app-service-plan  App Service plans hosting no apps  (core)
  idle-container-app  Always-on container apps serving nothing  (core)
  idle-dedicated-host  Dedicated hosts running no VMs  (core)
  idle-load-balancer  Load balancers with no backends  (core)
  idle-ml-compute  ML compute running or held idle  (core)
  idle-nat-gateway  NAT gateways with no subnets  (core)
  idle-provisioned-deployment  Provisioned model deployments serving nothing  (core)
  idle-workload-profile  Dedicated workload profiles running no app  (core)
  orphaned-nic  Network interfaces with no VM  (core)
  orphaned-snapshot  Snapshots of deleted disks  (core)
  paused-sql-database  Paused SQL databases still paying for storage  (core)
  stale-key-vault-secret  Secrets with no new version in 90 days  (core)
  unattached-disk  Unattached managed disks  (core)
  unbounded-log-workspace  Log workspaces with no ingestion cap  (core)
  unmanaged-storage-account  Versioned storage accounts with no lifecycle policy  (core)
  unused-availability-test  Availability tests watching a deleted resource  (core)
  unused-capacity-reservation  Capacity reservations holding unused slots  (core)
  unused-dns-zone  DNS zones with no records  (core)
  unused-image  Managed images nothing boots from  (core)
  unused-nsg  Security groups protecting nothing  (core)
  unused-public-ip  Public IP addresses attached to nothing  (core)
  unused-subnet  Subnets reserving a range against nothing  (core)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Twenty-nine ship in the &lt;code&gt;core&lt;/code&gt; pack and one in &lt;code&gt;aks&lt;/code&gt;. AKS sits in its own pack because it carries its own resource provider, its own rate section and its own price fetcher, which is what a third-party pack has to supply.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 5 — Which Resource Providers the Checks Read
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;zombiescan providers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  Microsoft.App  3 check(s)
  Microsoft.CognitiveServices  2 check(s)
  Microsoft.Compute  6 check(s)
  Microsoft.ContainerRegistry  1 check(s)
  Microsoft.ContainerService  1 check(s)
  Microsoft.Insights  3 check(s)
  Microsoft.KeyVault  2 check(s)
  Microsoft.MachineLearningServices  1 check(s)
  Microsoft.Network  9 check(s)
  Microsoft.OperationalInsights  1 check(s)
  Microsoft.Resources  1 check(s)
  Microsoft.Sql  1 check(s)
  Microsoft.Storage  1 check(s)
  Microsoft.Web  1 check(s)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This list decides more on Azure than on the other two clouds. &lt;strong&gt;ARM answers a list call against an unregistered resource provider with HTTP 200 and an empty page.&lt;/strong&gt; A subscription that has never used App Service returns &lt;code&gt;{"value": []}&lt;/code&gt; from &lt;code&gt;/providers/Microsoft.Web/serverfarms&lt;/code&gt;, with no error. A check that simply ran would find nothing and report a clean subscription.&lt;/p&gt;

&lt;p&gt;So every check declares the providers it reads, the engine reads each subscription's registrations once, and a check whose provider is missing is counted as skipped. The same declaration generates the read-only custom role, and the test suite fails when a check names a provider the role does not cover.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 6 — Scan One Subscription
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;zombiescan scan
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 subscription(s), 30 check(s) — read-only

No waste found across 1 subscription. Nothing to clean up.

4 of 30 subscription/check pair(s) were skipped: their resource provider is not registered, so there
is nothing of that kind here. 'zombiescan providers' lists what each check reads.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With no flag it scans the subscription &lt;code&gt;az&lt;/code&gt; is set to. The four skipped pairs are Key Vault's two checks, App Service and SQL: this subscription has never registered those providers, so it has no resources of those kinds.&lt;/p&gt;

&lt;p&gt;Every call is a list, a get or a Resource Graph query. The scan has no code path that deletes, modifies or releases anything.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 7 — Scan Every Subscription, in Every Tenant
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;zombiescan scan &lt;span class="nt"&gt;--all-subscriptions&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--all-subscriptions&lt;/code&gt; takes every enabled subscription in &lt;code&gt;az account list --all&lt;/code&gt;, across every tenant the CLI has signed into, and runs &lt;code&gt;az account list --refresh&lt;/code&gt; first to pick up subscriptions created since the last login. On this account that is one subscription in one tenant, and the sweep of 30 checks took 8.65 seconds.&lt;/p&gt;

&lt;p&gt;A second identity is one &lt;code&gt;az login&lt;/code&gt; away. After it, the same command reaches both directories, requests for each subscription carry that tenant's token, and the header lists both tenants with their account kinds.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 8 — What a Finding Looks Like
&lt;/h4&gt;

&lt;p&gt;This subscription is clean, so the rows below come from the test suite's recorded ARM responses — the same JSON each check is tested against — priced from the real bundled price table.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run python articles/zombiescan-azure/show-fixture-findings.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;49 findings from 30 checks, recorded responses, prices generated 2026-09-23T17:19:30Z
total $32,345.26/month

Check                              Resource                           Monthly
idle-provisioned-deployment        acct-ptu/ptu-idle              $21,900.00
idle-ml-compute                    ws-research/gpu-warm            $4,467.60
idle-dedicated-host                host-idle                       $3,084.25
idle-workload-profile              env-single/d4-solo                $522.62
idle-app-service-plan              prod-plan                         $459.90
unused-capacity-reservation        cr-partial                        $420.48
empty-container-apps-environment   env-empty-dedicated               $297.81
idle-workload-profile              env-mixed/d4-idle                 $224.81
idle-ml-compute                    ws-research/ci-forever            $213.89
unused-capacity-reservation        cr-idle                           $140.16
aks-idle-cluster                   scaled-to-zero                     $73.00
idle-nat-gateway                   egress-gw-old                      $32.85
deallocated-vm                     batch-runner                       $24.90
idle-load-balancer                 api-lb                             $21.90
unattached-disk                    orphan-data                        $19.71
idle-container-app                 app-idle                           $11.83
orphaned-nic                       web-01-nic-old                      $3.65
unattached-disk                    tiny-scratch                        $0.60
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The top of that list is where Azure's waste concentrates now: a 15-PTU regional model deployment serving nothing is $21,900 a month, a GPU cluster whose minimum keeps two idle nodes up is $4,467.60, and an empty &lt;code&gt;DSv3-Type3&lt;/code&gt; dedicated host is $3,084.25. A forgotten disk is $19.71.&lt;/p&gt;

&lt;p&gt;Every finding carries its subscription, resource group, location and full ARM id. No &lt;code&gt;az&lt;/code&gt; command works without the resource group, and the ARM id is the only identifier unique across a tenant.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 9 — Where a Check Runs
&lt;/h4&gt;

&lt;p&gt;A check runs once per subscription. Two properties of ARM make that a complete sweep.&lt;/p&gt;

&lt;p&gt;A list call at &lt;code&gt;/subscriptions/&amp;lt;id&amp;gt;/providers/Microsoft.Compute/disks&lt;/code&gt; returns every disk in every resource group and every region, in one call. And Azure Resource Graph answers a cross-type join in one KQL query, so the check for stopped VMs reads power state for every VM at once without an instance-view call per machine.&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;@check&lt;/code&gt; has no region scope. Each finding reads its own &lt;code&gt;location&lt;/code&gt; off the resource, and &lt;code&gt;--location&lt;/code&gt; filters findings afterwards.&lt;/p&gt;

&lt;p&gt;🔎 Tip: follow &lt;code&gt;nextLink&lt;/code&gt; even when the first page is empty. Listing AI Services accounts across a subscription returned an empty first page with a &lt;code&gt;nextLink&lt;/code&gt;, and the account on the second page. A reader that stopped at page one would report the subscription clean.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 10 — Pinned API Versions
&lt;/h4&gt;

&lt;p&gt;Azure has no "latest". Every request carries a required &lt;code&gt;api-version&lt;/code&gt;, and &lt;code&gt;azure.API_VERSIONS&lt;/code&gt; pins one per resource type, resolved by longest prefix so a sub-type inherits its parent's version.&lt;/p&gt;

&lt;p&gt;A retired version produces &lt;code&gt;InvalidResourceType&lt;/code&gt;, which is 404-shaped, and a 404 reads as "nothing here". The check would stop reporting and the subscription would look that much cleaner. The live test, gated behind &lt;code&gt;ZOMBIESCAN_LIVE=1&lt;/code&gt;, sends every pin to ARM and fails on the first one ARM rejects.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 11 — Where the Prices Come From
&lt;/h4&gt;

&lt;p&gt;The bundled price table is generated from the Azure Retail Prices API, which is public: no credentials and no subscription.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; zombiescan.pricing.refresh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A scan reads &lt;code&gt;table.json&lt;/code&gt; from the installed package, so it runs offline and adds no latency. Rates are looked up by key and region: &lt;code&gt;ctx.pricing.rate("disk.tier_month", region=..., variant="P10 LRS")&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Verified rates in eastus, from the bundled table:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rate&lt;/th&gt;
&lt;th&gt;eastus&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Standard static public IP&lt;/td&gt;
&lt;td&gt;$0.005/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NAT Gateway&lt;/td&gt;
&lt;td&gt;$0.045/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard Load Balancer rules&lt;/td&gt;
&lt;td&gt;$0.025/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AKS Standard control plane&lt;/td&gt;
&lt;td&gt;$0.10/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dedicated host &lt;code&gt;DSv3-Type3&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;$4.225/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;Standard_D2s_v3&lt;/code&gt;, Linux&lt;/td&gt;
&lt;td&gt;$0.096/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provisioned throughput, regional&lt;/td&gt;
&lt;td&gt;$2.00 per PTU-hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provisioned throughput, global&lt;/td&gt;
&lt;td&gt;$1.00 per PTU-hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Container Apps Dedicated management fee&lt;/td&gt;
&lt;td&gt;$0.10/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 730 hours to the month, arithmetic over those hourly rates gives $3.65 for an idle public IP, $32.85 for a NAT gateway, $18.25 for load balancer rules, $73.00 for an AKS Standard control plane and $3,084.25 for an empty dedicated host.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 12 — Four Ways a Price Goes Wrong
&lt;/h4&gt;

&lt;p&gt;Each of these produces a plausible table with the wrong numbers in it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Filter on &lt;code&gt;priceType eq 'Consumption'&lt;/code&gt;.&lt;/strong&gt; The same meter is published as &lt;code&gt;Reservation&lt;/code&gt; and &lt;code&gt;DevTestConsumption&lt;/code&gt; too, at a fraction of the price.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Take the first tier that charges.&lt;/strong&gt; Azure returns a tiered meter's rows in no guaranteed order, and tier 0 is often a free allowance. &lt;code&gt;unit_price()&lt;/code&gt; sorts by &lt;code&gt;tierMinimumUnits&lt;/code&gt; and reads the first nonzero row.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Match the meter name as well as the SKU.&lt;/strong&gt; &lt;code&gt;P80 LRS Disk&lt;/code&gt;, &lt;code&gt;P80 LRS Disk Mount&lt;/code&gt; and &lt;code&gt;P80 LRS Disk Operations&lt;/code&gt; all share the SKU "P80 LRS"; only the first is the capacity charge. The same care covers Windows rows: most VM products spell it "Windows", and a few GPU series abbreviate it to "Win".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Some meters have no ARM region.&lt;/strong&gt; NAT Gateway and Load Balancer are published against "Global", and Azure DNS against a billing geography spelled "Zone 1". Those live in a global section.&lt;/p&gt;

&lt;p&gt;The refresher refuses to write a table that loses or empties a section the previous one had, because an empty section means a matcher stopped matching.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 13 — Managed Disks Are Priced by Tier
&lt;/h4&gt;

&lt;p&gt;A managed disk bills at the tier its provisioned size falls into, flat per month:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Disk&lt;/th&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;eastus per month&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 GiB Premium SSD&lt;/td&gt;
&lt;td&gt;P1&lt;/td&gt;
&lt;td&gt;$0.60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;65 GiB Premium SSD&lt;/td&gt;
&lt;td&gt;P10&lt;/td&gt;
&lt;td&gt;$19.71&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128 GiB Premium SSD&lt;/td&gt;
&lt;td&gt;P10&lt;/td&gt;
&lt;td&gt;$19.71&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 GiB Standard HDD&lt;/td&gt;
&lt;td&gt;S4&lt;/td&gt;
&lt;td&gt;$1.54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32 GiB Standard HDD&lt;/td&gt;
&lt;td&gt;S4&lt;/td&gt;
&lt;td&gt;$1.54&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A 65 GiB Premium disk and a 128 GiB one cost the same. Standard HDD has no rung below S4, so a 4 GiB Standard disk bills as a 32 GiB one. The price table is keyed by tier and redundancy — "P10 LRS", "P10 ZRS" — because zone redundancy costs about half as much again: $29.57 for a P10 ZRS.&lt;/p&gt;

&lt;p&gt;Premium SSD v2 and Ultra bill per provisioned GiB, and both bill provisioned IOPS and throughput on top, which the finding says.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 14 — Idle Needs a Week of Metrics
&lt;/h4&gt;

&lt;p&gt;Most checks answer from one list call. Two need usage:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;idle-provisioned-deployment&lt;/code&gt; reads &lt;code&gt;ModelRequests&lt;/code&gt; on each account over seven days, split by &lt;code&gt;ModelDeploymentName&lt;/code&gt;, so one call covers every deployment on the account&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;idle-container-app&lt;/code&gt; reads &lt;code&gt;Requests&lt;/code&gt; on each app with &lt;code&gt;minReplicas&lt;/code&gt; of one or more&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both go through &lt;code&gt;helpers.metric_totals&lt;/code&gt;, which asks Azure Monitor for the &lt;code&gt;Total&lt;/code&gt; over &lt;code&gt;P7D&lt;/code&gt; and returns the sum. A deployment that served nothing has no series at all, and a missing series reads as zero. Both checks declare &lt;code&gt;Microsoft.Insights&lt;/code&gt;, the provider behind the metrics endpoint, so the registration check covers it too.&lt;/p&gt;

&lt;p&gt;On this subscription the container app recorded 0 requests in the seven days, and the gpt-5-mini deployment returned no series. Both stay out of the report: the app has &lt;code&gt;minReplicas: 0&lt;/code&gt; and the deployment is pay-per-token, so each costs nothing idle.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 15 — JSON Output
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;zombiescan scan &lt;span class="nt"&gt;--all-subscriptions&lt;/span&gt; &lt;span class="nt"&gt;--json&lt;/span&gt; findings.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"schema_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"scan"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"generated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-23T17:47:01Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"duration_seconds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;8.65&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"principal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"xbill@glitnir.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"subscriptions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"3db3ce66-50b6-4d11-91ef-5950cf4039ed"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"pairs_attempted"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"pairs_unavailable"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"complete"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"pricing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"generated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-23T17:19:30Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"basis"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Azure Retail Prices API, pay-as-you-go USD list prices"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"excludes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"reservations"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"savings plans"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Azure Hybrid Benefit"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"dev/test rates"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"enterprise agreement pricing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"credits"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"totals"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"monthly_cost"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"annual_cost"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"finding_count"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"by_check"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pricing block dates the table and lists what list prices exclude, so a report read six months later says which rates produced it. &lt;code&gt;pairs_unavailable&lt;/code&gt; carries the skipped count, and &lt;code&gt;complete&lt;/code&gt; is false when every pair failed or was skipped, so a CI job reading the file can tell an all-clear from a scan that reached nothing.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 16 — The HTML Report
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;zombiescan scan &lt;span class="nt"&gt;--all-subscriptions&lt;/span&gt; &lt;span class="nt"&gt;--html&lt;/span&gt; report.html
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One self-contained file with the styles inline, which prints to PDF without fetching anything. With no findings its headline reads "No waste found" and names how many subscription/check pairs ran, how many were skipped for an unregistered provider, and how many failed.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 17 — The Cleanup Plan
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;zombiescan scan &lt;span class="nt"&gt;--all-subscriptions&lt;/span&gt; &lt;span class="nt"&gt;--script&lt;/span&gt; cleanup.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="c"&gt;# zombiescan cleanup plan — generated 2026-09-23T17:47:01Z&lt;/span&gt;
&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="c"&gt;# READ EVERY LINE BEFORE RUNNING THIS.&lt;/span&gt;
&lt;span class="c"&gt;# zombiescan generated this file and did not run it. Some of these&lt;/span&gt;
&lt;span class="c"&gt;# deletions can be undone and some cannot: Key Vault and SQL keep a&lt;/span&gt;
&lt;span class="c"&gt;# recovery window, a released public IP address is gone for good.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every generated command is built by one helper, which adds &lt;code&gt;--resource-group&lt;/code&gt; and &lt;code&gt;--subscription&lt;/code&gt; and wraps interpolated names in shell quoting. The test suite checks both at the source.&lt;/p&gt;

&lt;p&gt;🔎 Tip: &lt;code&gt;--yes&lt;/code&gt; exists only on the &lt;code&gt;az&lt;/code&gt; commands that would otherwise prompt, and passing it to one that does not is an error. &lt;code&gt;az disk delete&lt;/code&gt; takes it; &lt;code&gt;az snapshot delete&lt;/code&gt; and &lt;code&gt;az network nic delete&lt;/code&gt; reject it. The helper holds the list of commands that take it, and the live test re-checks that list against the installed CLI.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 18 — Cleaning Up
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;clean&lt;/code&gt; is a separate command, and a dry run is what it does with no flags.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;zombiescan clean &lt;span class="nt"&gt;--all-subscriptions&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Nothing to clean.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--apply&lt;/code&gt; gates one thing: whether a planned step is sent to Azure. It never changes which steps get planned, so the preview is what runs.&lt;/p&gt;

&lt;p&gt;Planning is read-only. A cleaner may read — re-listing a resource group to confirm it is still empty, re-reading a model deployment's request count, re-reading a capacity reservation's allocation — and it yields the mutations as objects for the runner to send. A plan refuses when the resource has changed since the scan: a host with a VM placed on it, an account that gained a deployment, a reservation that filled up.&lt;/p&gt;

&lt;p&gt;Where the API allows a backup, it comes first. The disk cleaner takes an incremental snapshot before the delete, and a failed step stops the rest of that finding.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;IRREVERSIBLE&lt;/code&gt; marks a step with no recovery window at all. A released public IP is gone, and so is a deleted Container Apps environment's static IP, so both carry the mark. A Key Vault key is held by mandatory soft-delete, a SQL database restores from point-in-time backups, and a deleted AI Services account is recoverable for 48 hours, so none of those does.&lt;/p&gt;

&lt;p&gt;A check with no cleaner reports its findings as unsupported with a reason. &lt;code&gt;idle-container-app&lt;/code&gt; is one: lowering &lt;code&gt;minReplicas&lt;/code&gt; trades the idle charge for a cold start on the next request, and that trade belongs to whoever owns the app.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 19 — The Read-Only Role
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az role definition create &lt;span class="nt"&gt;--role-definition&lt;/span&gt; policy/zombiescan-scanner-role.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Forty-four actions: forty-three reads and Container Registry's &lt;code&gt;listUsages&lt;/code&gt;, the one management-plane call that reports how much a registry stores. The file is generated from the &lt;code&gt;providers=&lt;/code&gt; each check declares, and the suite keeps it in step.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;clean&lt;/code&gt; needs more than this, deliberately: the identity that reports waste should be unable to delete what it reports.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 20 — The Claude Code Plugin
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/plugin marketplace add xbill9/zombiescan-azure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The plugin ships two slash commands, a skill and an MCP server over the same engine.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  scan_subscription  Scan Azure subscriptions for unused resources and price them
  estimate_savings  Total, count and break down the findings in a report
  explain_finding  Explain why a check treats a resource as waste
  list_checks  List the installed checks
  plan_cleanup  Show exactly what `zombiescan clean` would do
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every tool is read-only. &lt;code&gt;plan_cleanup&lt;/code&gt; builds the step objects and stops there; the function that sends them to Azure is unreachable from the server, and a test asserts the module names no other &lt;code&gt;clean.*&lt;/code&gt; attribute.&lt;/p&gt;

&lt;p&gt;The tools return computed figures — totals, counts, breakdowns, cheapest and costliest — and echo the filter they applied. A tool that returned rows for the model to add up would move the arithmetic to the place least able to do it.&lt;/p&gt;




&lt;h4&gt;
  
  
  Step 21 — What a Scan Cannot See
&lt;/h4&gt;

&lt;p&gt;A clean scan answers one question: is anything billing that nothing uses? It says nothing about a project that has gone quiet as a whole. This subscription holds one resource group, &lt;code&gt;research-mesh-rg&lt;/code&gt;, with six resources: a container app and its environment, a Foundry account with one project and a gpt-5-mini deployment, a container registry and a Log Analytics workspace.&lt;/p&gt;

&lt;p&gt;Cost Management reports what the group was charged, and the totals below are Azure's own:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Resource&lt;/th&gt;
&lt;th&gt;Last 30 days&lt;/th&gt;
&lt;th&gt;Since Aug 13&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;researchmeshacr&lt;/code&gt;, Container Registry Basic&lt;/td&gt;
&lt;td&gt;$5.08&lt;/td&gt;
&lt;td&gt;$6.87&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;research-mesh-foundry&lt;/code&gt;, gpt-5-mini tokens&lt;/td&gt;
&lt;td&gt;$0.008&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Log Analytics workspace&lt;/td&gt;
&lt;td&gt;$0.00&lt;/td&gt;
&lt;td&gt;$0.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Container app and environment&lt;/td&gt;
&lt;td&gt;no charges&lt;/td&gt;
&lt;td&gt;no charges&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;$5.09 over the last 30 days, and the registry is 99.8% of it. Its Basic tier bills a flat daily fee whether anyone pulls an image or not. Everything else bills on use, and use is close to zero: the app served 0 requests in seven days and the deployment under a cent of tokens in thirty.&lt;/p&gt;

&lt;p&gt;None of the six trips a check, because each is in use by another: the registry holds the image the app runs, the account has a deployment, and the environment has an app. Finding a group like this takes Cost Management data, which is a different source from anything a scan reads.&lt;/p&gt;




&lt;h4&gt;
  
  
  Three Clouds, Same Resource
&lt;/h4&gt;

&lt;p&gt;The same idle resource costs differently on each cloud. Figures are from each article's bundled price table.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Resource&lt;/th&gt;
&lt;th&gt;AWS (us-east-1)&lt;/th&gt;
&lt;th&gt;Google Cloud (us-central1)&lt;/th&gt;
&lt;th&gt;Azure (eastus)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Idle NAT gateway&lt;/td&gt;
&lt;td&gt;🥈 $32.85/month flat&lt;/td&gt;
&lt;td&gt;🥇 the addresses it holds, $3.65 each&lt;/td&gt;
&lt;td&gt;🥈 $32.85/month flat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128 GB disk, attached to nothing&lt;/td&gt;
&lt;td&gt;🥇 $10.24, gp3 per GB&lt;/td&gt;
&lt;td&gt;🥈 $12.80, pd-balanced per GB&lt;/td&gt;
&lt;td&gt;🥉 $19.71, P10 by tier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Managed Kubernetes, no nodes&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;🥈 $73.00 management fee&lt;/td&gt;
&lt;td&gt;🥇 $0.00 on Free, $73.00 on Standard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unit a check fans out over&lt;/td&gt;
&lt;td&gt;region&lt;/td&gt;
&lt;td&gt;project&lt;/td&gt;
&lt;td&gt;subscription&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Price source&lt;/td&gt;
&lt;td&gt;AWS Price List API&lt;/td&gt;
&lt;td&gt;Cloud Billing Catalog API&lt;/td&gt;
&lt;td&gt;Azure Retail Prices API, no credentials&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two checks exist only on Azure. &lt;strong&gt;Orphaned NICs&lt;/strong&gt;, because a network interface is a resource of its own that outlives its VM and blocks deleting the public IP, subnet and virtual network it references. And &lt;strong&gt;empty resource groups&lt;/strong&gt;, because a resource group is a container inside a subscription.&lt;/p&gt;

&lt;p&gt;🔎 Tip: an Azure public IP bills the same rate attached or idle, so nothing in the price signals that it is unused. On Google Cloud an idle static IP costs more than one in use.&lt;/p&gt;




&lt;h4&gt;
  
  
  What the Checks Leave Alone
&lt;/h4&gt;

&lt;p&gt;A false positive here costs an outage, so each check states the condition it declines to act on.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Treatment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Public IP held by an orphaned NIC, a deallocated VM or an idle load balancer&lt;/td&gt;
&lt;td&gt;Priced on the holder's finding, counted once&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disk reserved by a stopped VM&lt;/td&gt;
&lt;td&gt;Reported under &lt;code&gt;deallocated-vm&lt;/code&gt;, counted once&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disk in the &lt;code&gt;ActiveSAS&lt;/code&gt; state&lt;/td&gt;
&lt;td&gt;Something is reading it; left alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Snapshot whose source disk still exists&lt;/td&gt;
&lt;td&gt;Load-bearing; left alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI Services account with a project but no deployment&lt;/td&gt;
&lt;td&gt;A project can use models hosted elsewhere; left alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Container app with no ingress&lt;/td&gt;
&lt;td&gt;A worker processes queues, so a zero request count proves nothing; left alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource provider not registered&lt;/td&gt;
&lt;td&gt;Skipped and counted; a service never used holds no waste&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  Compare and Contrast
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;zombiescan&lt;/th&gt;
&lt;th&gt;Azure Advisor&lt;/th&gt;
&lt;th&gt;Cost Management&lt;/th&gt;
&lt;th&gt;Cross-subscription SaaS&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Per-resource dollar figure&lt;/td&gt;
&lt;td&gt;🥇 yes&lt;/td&gt;
&lt;td&gt;partial&lt;/td&gt;
&lt;td&gt;per resource, after the fact&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Credentials leave the machine&lt;/td&gt;
&lt;td&gt;🥇 never&lt;/td&gt;
&lt;td&gt;n/a, Azure-side&lt;/td&gt;
&lt;td&gt;n/a, Azure-side&lt;/td&gt;
&lt;td&gt;❌ service principal granted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deletes on request&lt;/td&gt;
&lt;td&gt;🥈 opt-in, dry run first&lt;/td&gt;
&lt;td&gt;❌ no&lt;/td&gt;
&lt;td&gt;❌ no&lt;/td&gt;
&lt;td&gt;🥇 yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runs offline after install&lt;/td&gt;
&lt;td&gt;🥇 yes&lt;/td&gt;
&lt;td&gt;❌ no&lt;/td&gt;
&lt;td&gt;❌ no&lt;/td&gt;
&lt;td&gt;❌ no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Breadth&lt;/td&gt;
&lt;td&gt;30 checks&lt;/td&gt;
&lt;td&gt;broader&lt;/td&gt;
&lt;td&gt;everything billed&lt;/td&gt;
&lt;td&gt;broader&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h4&gt;
  
  
  So, Which One?
&lt;/h4&gt;

&lt;p&gt;Cost Management answers what a subscription spent, down to the resource, and it is the only source that sees a quiet project like &lt;code&gt;research-mesh-rg&lt;/code&gt;. Advisor covers more ground, including right-sizing judgements this tool leaves alone.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;zombiescan&lt;/code&gt; fits the case where the answer has to be per-resource, priced, and produced without granting anything access to the subscriptions. A laptop, an existing &lt;code&gt;az login&lt;/code&gt;, and under ten seconds for this subscription.&lt;/p&gt;




&lt;h4&gt;
  
  
  Cost
&lt;/h4&gt;

&lt;p&gt;A scan costs nothing. ARM list and get calls, Resource Graph queries and Azure Monitor metric reads carry no charge, and the price table ships with the package.&lt;/p&gt;

&lt;p&gt;Regenerating the table calls the Azure Retail Prices API, which is public and free.&lt;/p&gt;




&lt;h4&gt;
  
  
  Tests
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run pytest &lt;span class="nt"&gt;-q&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;380 passed, 10 deselected in 0.34s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The suite runs offline against recorded ARM responses, with no credentials. The ten deselected tests reach a real subscription and run under &lt;code&gt;ZOMBIESCAN_LIVE=1&lt;/code&gt;: they verify every pinned &lt;code&gt;api-version&lt;/code&gt; against ARM and every &lt;code&gt;--yes&lt;/code&gt; command against the installed CLI, two facts about the outside world that a comment cannot keep true.&lt;/p&gt;

&lt;p&gt;Recorded responses are keyed by the tail of the request path, and they include resources that are in use, because half of what these checks do is decline to report.&lt;/p&gt;




&lt;h4&gt;
  
  
  Teardown
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv tool uninstall zombiescan-azure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing is left in any subscription. The tool creates no service principal, no role assignment, no storage account and no stored state. Sign the CLI out with &lt;code&gt;az logout&lt;/code&gt;.&lt;/p&gt;




&lt;h4&gt;
  
  
  Summary
&lt;/h4&gt;

&lt;p&gt;The goal of this article was to audit every Azure subscription a set of &lt;code&gt;az&lt;/code&gt; credentials can reach and attach a monthly cost to each unused resource. The key to the solution was holding one token per tenant, checking provider registration before trusting an empty page, and pricing every finding from the Azure Retail Prices API while keeping the scan read-only. The results were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟢 30 checks in two packs across Compute, Networking, Storage, SQL, Key Vault, Container Registry, App Service, Container Apps, AI Services, Machine Learning, Log Analytics and AKS&lt;/li&gt;
&lt;li&gt;🟢 Every tenant the CLI has signed into, from one &lt;code&gt;az login&lt;/code&gt; each, with the account kind — personal or work — named in the header&lt;/li&gt;
&lt;li&gt;🟢 One call per subscription: ARM lists are subscription-wide and Resource Graph joins resource types in one query&lt;/li&gt;
&lt;li&gt;🟢 Unregistered providers skipped and counted, so an empty page never reads as a clean subscription&lt;/li&gt;
&lt;li&gt;🟢 Disks priced by tier and redundancy, and the priciest classes — provisioned model throughput, GPU clusters, dedicated hosts — priced from their own meters&lt;/li&gt;
&lt;li&gt;🟢 JSON against a published schema, a self-contained HTML report, and a shell cleanup plan that is printed and never run&lt;/li&gt;
&lt;li&gt;🟢 &lt;code&gt;clean&lt;/code&gt; previews by default, re-reads each resource before planning, and marks steps with no recovery window&lt;/li&gt;
&lt;li&gt;⚠️ The subscription scanned here is clean: 0 findings, with 4 of 30 pairs skipped for unregistered providers, so the per-resource figures come from the recorded test responses&lt;/li&gt;
&lt;li&gt;⚠️ Figures are pay-as-you-go list prices, excluding reservations, savings plans, Azure Hybrid Benefit and credits&lt;/li&gt;
&lt;li&gt;⚠️ Two checks judge idleness from seven days of Azure Monitor metrics, which is one window among several reasonable ones&lt;/li&gt;
&lt;li&gt;❌ A resource group that has gone quiet as a whole trips no check; seeing it takes Cost Management data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scope: one personal Microsoft account with one tenant and one subscription, 30 checks, a single sweep per figure quoted, run from one laptop. Rates are eastus from a table generated 2026-09-23; the AWS and Google Cloud figures come from those articles' own price tables. The read-only guarantee holds for the packs in this repository, which are reviewed; a pack installed from PyPI runs with the same credentials and carries no such guarantee.&lt;/p&gt;

&lt;p&gt;The strategy for using per-resource pricing for Azure waste detection was validated with an incremental step by step approach.&lt;/p&gt;




&lt;h4&gt;
  
  
  References
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Repository, MIT: &lt;a href="https://github.com/xbill9/zombiescan-azure" rel="noopener noreferrer"&gt;https://github.com/xbill9/zombiescan-azure&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Findings JSON Schema: &lt;a href="https://github.com/xbill9/zombiescan-azure/blob/main/docs/findings.schema.json" rel="noopener noreferrer"&gt;https://github.com/xbill9/zombiescan-azure/blob/main/docs/findings.schema.json&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Pack author guide: &lt;a href="https://github.com/xbill9/zombiescan-azure/blob/main/docs/PACKS.md" rel="noopener noreferrer"&gt;https://github.com/xbill9/zombiescan-azure/blob/main/docs/PACKS.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Part one, AWS: &lt;a href="https://dev.to/aws-builders/find-the-aws-resources-nobody-is-using-and-what-they-cost-you-35bj"&gt;https://dev.to/aws-builders/find-the-aws-resources-nobody-is-using-and-what-they-cost-you-35bj&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Part two, Google Cloud: &lt;a href="https://dev.to/gde/what-nobody-is-using-in-your-google-cloud-projects-and-what-it-costs-1k0"&gt;https://dev.to/gde/what-nobody-is-using-in-your-google-cloud-projects-and-what-it-costs-1k0&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Azure Retail Prices API: &lt;a href="https://learn.microsoft.com/en-us/rest/api/cost-management/retail-prices/azure-retail-prices" rel="noopener noreferrer"&gt;https://learn.microsoft.com/en-us/rest/api/cost-management/retail-prices/azure-retail-prices&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Azure Resource Graph: &lt;a href="https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview" rel="noopener noreferrer"&gt;https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Resource providers and registration: &lt;a href="https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/resource-providers-and-types" rel="noopener noreferrer"&gt;https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/resource-providers-and-types&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Managed disk pricing: &lt;a href="https://azure.microsoft.com/en-us/pricing/details/managed-disks/" rel="noopener noreferrer"&gt;https://azure.microsoft.com/en-us/pricing/details/managed-disks/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Capacity reservation billing: &lt;a href="https://learn.microsoft.com/en-us/azure/virtual-machines/capacity-reservation-overview" rel="noopener noreferrer"&gt;https://learn.microsoft.com/en-us/azure/virtual-machines/capacity-reservation-overview&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Dedicated host pricing: &lt;a href="https://learn.microsoft.com/en-us/azure/virtual-machines/dedicated-hosts" rel="noopener noreferrer"&gt;https://learn.microsoft.com/en-us/azure/virtual-machines/dedicated-hosts&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>azure</category>
      <category>python</category>
      <category>devops</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
