<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: yyyysa4</title>
    <description>The latest articles on DEV Community by yyyysa4 (@yyyysa4).</description>
    <link>https://dev.to/yyyysa4</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4140270%2F5afac99b-a6a6-4b7d-8aea-613e6de7dca4.png</url>
      <title>DEV Community: yyyysa4</title>
      <link>https://dev.to/yyyysa4</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yyyysa4"/>
    <language>en</language>
    <item>
      <title>RTX PRO 6000 vs H100 for LLM Inference: Which Is More Cost-Effective in 2026?</title>
      <dc:creator>yyyysa4</dc:creator>
      <pubDate>Wed, 07 Oct 2026 03:19:48 +0000</pubDate>
      <link>https://dev.to/yyyysa4/rtx-pro-6000-vs-h100-for-llm-inference-which-is-more-cost-effective-in-2026-40ba</link>
      <guid>https://dev.to/yyyysa4/rtx-pro-6000-vs-h100-for-llm-inference-which-is-more-cost-effective-in-2026-40ba</guid>
      <description>

&lt;p&gt;The NVIDIA H100 is the default answer to "what GPU should I serve this model on?" The RTX PRO 6000 Blackwell costs about two-thirds as much per hour, has more memory, and less than half the memory bandwidth. Which of those facts wins depends on one question that most comparisons skip: &lt;strong&gt;does your model fit on a single card?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This post works through that question using independent benchmarks, current market prices, and arithmetic you can check. No new measurements were taken; sources and a reproduction script are at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Single-GPU models: the RTX PRO 6000 is cheaper per token&lt;/strong&gt; — about 30–40% cheaper than an H100 at October 2026 median on-demand prices.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-GPU models: the advantage shrinks and then reverses.&lt;/strong&gt; Roughly 8% cheaper at 4-way tensor parallelism, roughly 8% &lt;em&gt;more&lt;/em&gt; expensive at 8-way. The RTX PRO 6000 has no NVLink.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The H100 has 2.1× the memory bandwidth. The RTX PRO 6000 has 16 GB more memory&lt;/strong&gt;, which turns into 1.8–2.6× more KV-cache room for 30B–70B models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rule of thumb:&lt;/strong&gt; cost-per-token ratio = price ratio × throughput ratio. Today's price ratio is about 0.63, so the RTX PRO 6000 wins whenever an H100 is less than ~1.6× faster on your workload.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The specs that matter for inference
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;RTX PRO 6000 Blackwell Server Edition&lt;/th&gt;
&lt;th&gt;H100 SXM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;96 GB GDDR7&lt;/td&gt;
&lt;td&gt;80 GB HBM3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory bandwidth&lt;/td&gt;
&lt;td&gt;1,597 GB/s&lt;/td&gt;
&lt;td&gt;3,350 GB/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-GPU interconnect&lt;/td&gt;
&lt;td&gt;PCIe Gen 5 only — no NVLink&lt;/td&gt;
&lt;td&gt;NVLink 900 GB/s, plus PCIe Gen 5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lowest native tensor precision&lt;/td&gt;
&lt;td&gt;FP4&lt;/td&gt;
&lt;td&gt;FP8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MIG partitions&lt;/td&gt;
&lt;td&gt;up to 4&lt;/td&gt;
&lt;td&gt;up to 7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max power&lt;/td&gt;
&lt;td&gt;up to 600 W&lt;/td&gt;
&lt;td&gt;up to 700 W&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;On-demand price, median, Oct 2026&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$2.20 / GPU-hr&lt;/strong&gt; (53 providers)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$3.49 / GPU-hr&lt;/strong&gt; (57 providers)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two notes on that table.&lt;/p&gt;

&lt;p&gt;Peak TFLOPS are left out on purpose. NVIDIA's product pages quote them with different sparsity conventions, so putting them side by side invites a wrong conclusion. For LLM serving, where the decode phase is usually limited by how fast weights and KV cache can be read, memory bandwidth is the more predictive spec.&lt;/p&gt;

&lt;p&gt;"RTX PRO 6000" also covers more than one card. The &lt;strong&gt;Workstation Edition&lt;/strong&gt; lists 1,792 GB/s of bandwidth; the &lt;strong&gt;Server Edition&lt;/strong&gt; most clouds rent lists 1,597 GB/s. That difference shows up in the benchmarks below.&lt;/p&gt;

&lt;h2&gt;
  
  
  What independent benchmarks show
&lt;/h2&gt;

&lt;p&gt;The most useful public data comes from two vLLM benchmark runs published by CloudRift. Both compare the RTX PRO 6000 and the H100 directly, on the same models and settings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run A — November 2025.&lt;/strong&gt; Single GPU, GLM-4.5-Air (AWQ 4-bit), 1,000 input / 1,000 output tokens, 256–512 concurrent requests, RTX PRO 6000 &lt;strong&gt;Workstation&lt;/strong&gt; Edition.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;GPU&lt;/th&gt;
&lt;th&gt;Total throughput&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RTX PRO 6000 (Workstation)&lt;/td&gt;
&lt;td&gt;3,140 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;H100 SXM&lt;/td&gt;
&lt;td&gt;2,987 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Run B — January 2026.&lt;/strong&gt; Google Cloud 8-GPU nodes (G4 with RTX PRO 6000, A3 High with H100), 8,000 input / 8,000 output tokens.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;RTX PRO 6000&lt;/th&gt;
&lt;th&gt;H100&lt;/th&gt;
&lt;th&gt;H100 advantage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GLM-4.5-Air AWQ, one model per GPU&lt;/td&gt;
&lt;td&gt;2,290.69 tok/s&lt;/td&gt;
&lt;td&gt;2,556.03 tok/s&lt;/td&gt;
&lt;td&gt;+11.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder-480B-A35B AWQ, 4-GPU tensor parallel&lt;/td&gt;
&lt;td&gt;1,602.96 tok/s&lt;/td&gt;
&lt;td&gt;2,328.63 tok/s&lt;/td&gt;
&lt;td&gt;+45.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-4.6 FP8, 8-GPU tensor parallel&lt;/td&gt;
&lt;td&gt;1,651.67 tok/s&lt;/td&gt;
&lt;td&gt;2,833.77 tok/s&lt;/td&gt;
&lt;td&gt;+71.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both machines are 8-GPU nodes running identical configurations, so the ratios hold however the throughput was aggregated across the node.&lt;/p&gt;

&lt;p&gt;The two runs disagree on the single-GPU case: the RTX PRO 6000 is 5% ahead in Run A, while the H100 is 12% ahead in Run B. They differ in three ways at once — card edition, sequence length, and host — so neither result is wrong. The most plausible reading is that Run A's Workstation card had about 12% more bandwidth than a typical cloud Server Edition, and Run B's 8K-token sequences make each decoded token read much more KV cache, which favours the H100's bandwidth.&lt;/p&gt;

&lt;p&gt;The multi-GPU rows are not ambiguous. As tensor parallelism widens, the H100's NVLink pulls away: +45% at 4-way, +72% at 8-way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost per million tokens
&lt;/h2&gt;

&lt;p&gt;Throughput alone does not answer a cost question. Combining it with price:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost per token (RTX / H100) = (RTX $/hr ÷ H100 $/hr) × (H100 tok/s ÷ RTX tok/s)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At median on-demand prices ($2.20 vs $3.49, price ratio 0.63):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;H100 throughput vs RTX PRO 6000&lt;/th&gt;
&lt;th&gt;RTX PRO 6000 cost per token&lt;/th&gt;
&lt;th&gt;RTX break-even price&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 GPU, 1K/1K (Run A)&lt;/td&gt;
&lt;td&gt;0.95×&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;40% cheaper&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$3.67/hr&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1 GPU, 8K/8K (Run B)&lt;/td&gt;
&lt;td&gt;1.12×&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;30% cheaper&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$3.13/hr&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4-GPU tensor parallel (Run B)&lt;/td&gt;
&lt;td&gt;1.45×&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8% cheaper&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$2.40/hr&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8-GPU tensor parallel (Run B)&lt;/td&gt;
&lt;td&gt;1.72×&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8% more expensive&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$2.03/hr&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff1cdiv8bospjkeoj8f1t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff1cdiv8bospjkeoj8f1t.png" alt=" " width="799" height="319"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The break-even column is the RTX PRO 6000 hourly price at which it matches an H100 at $3.49/hr. Rent one below that price and it is the cheaper option for that workload.&lt;/p&gt;

&lt;p&gt;In absolute terms, Run A works out to about &lt;strong&gt;$0.195 per million tokens on the RTX PRO 6000 and $0.325 on the H100&lt;/strong&gt;, counting input and output tokens together.&lt;/p&gt;

&lt;p&gt;CloudRift's own cost tables differ from these. Each run used a different price basis — Runpod list prices in November 2025, and an estimated cost of ownership in January 2026. The table above re-prices both runs with one consistent, current source, so the scenarios are comparable with each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the extra 16 GB actually buys
&lt;/h2&gt;

&lt;p&gt;Twenty percent more memory sounds marginal. It is not, because every extra gigabyte lands after the weights are loaded, in the space left for KV cache. For models that nearly fill an 80 GB card, that space roughly doubles.&lt;/p&gt;

&lt;p&gt;KV cache per token for a standard attention model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;2 (K and V) × layers × KV heads × head dim × bytes per element
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With an FP8 KV cache and 4 GB reserved for the runtime:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Weights&lt;/th&gt;
&lt;th&gt;KV per token&lt;/th&gt;
&lt;th&gt;Free for KV on H100 80 GB&lt;/th&gt;
&lt;th&gt;Free for KV on RTX PRO 6000&lt;/th&gt;
&lt;th&gt;32K sequences: H100 → RTX&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder-30B-A3B (BF16)&lt;/td&gt;
&lt;td&gt;56.8 GB&lt;/td&gt;
&lt;td&gt;48 KB&lt;/td&gt;
&lt;td&gt;19.2 GB&lt;/td&gt;
&lt;td&gt;35.2 GB&lt;/td&gt;
&lt;td&gt;12 → &lt;strong&gt;23&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-32B (BF16)&lt;/td&gt;
&lt;td&gt;61.1 GB&lt;/td&gt;
&lt;td&gt;128 KB&lt;/td&gt;
&lt;td&gt;14.9 GB&lt;/td&gt;
&lt;td&gt;30.9 GB&lt;/td&gt;
&lt;td&gt;3 → &lt;strong&gt;7&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Llama-3.3-70B (FP8)&lt;/td&gt;
&lt;td&gt;65.8 GB&lt;/td&gt;
&lt;td&gt;160 KB&lt;/td&gt;
&lt;td&gt;10.2 GB&lt;/td&gt;
&lt;td&gt;26.2 GB&lt;/td&gt;
&lt;td&gt;2 → &lt;strong&gt;5&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Llama-3.3-70B (BF16)&lt;/td&gt;
&lt;td&gt;131.5 GB&lt;/td&gt;
&lt;td&gt;160 KB&lt;/td&gt;
&lt;td&gt;does not fit&lt;/td&gt;
&lt;td&gt;does not fit&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flyopvswh9zqoie3ur9os.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flyopvswh9zqoie3ur9os.png" alt=" " width="800" height="306"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A 70B model at FP8 is the clearest case. On an H100 it fits, but leaves room for only two concurrent 32K-token conversations. On an RTX PRO 6000, five. For long-context, RAG-heavy, or agentic workloads, that headroom decides how much traffic a single card can take.&lt;/p&gt;

&lt;p&gt;The last row matters just as much. A 70B model at BF16 needs two cards on either GPU — and once two cards have to talk to each other, the interconnect comes back into play, and with it the H100's advantage.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the H100 is the better buy
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your model needs four or more GPUs.&lt;/strong&gt; Once weights plus KV cache outgrow two cards (roughly 160–190 GB), you are into 4-way tensor parallelism or wider, where NVLink beats PCIe by a margin that outweighs the price gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single-stream decode speed is the product.&lt;/strong&gt; If latency for one user matters more than total throughput, 2.1× the memory bandwidth is hard to argue with.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You are serving very long sequences on a model that barely fits.&lt;/strong&gt; Run B suggests the bandwidth gap grows with sequence length.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need many small isolated tenants.&lt;/strong&gt; MIG gives you up to seven partitions on an H100 versus four on an RTX PRO 6000.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When the RTX PRO 6000 is the better buy
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The model and its KV cache fit on one 96 GB card.&lt;/strong&gt; That covers most 7B–35B models at BF16 and 70B-class models at FP8 or below.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You scale out with replicas rather than tensor parallelism.&lt;/strong&gt; Data-parallel serving never touches the interconnect, so the PCIe limitation does not apply.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concurrency on long contexts is the constraint.&lt;/strong&gt; The KV-cache table above is the whole argument.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You want to serve FP4-quantized models.&lt;/strong&gt; FP4 tensor cores are a Blackwell feature; Hopper stops at FP8.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A short decision checklist
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Add up weights and the KV cache your target concurrency needs, at the precision you plan to serve.&lt;/li&gt;
&lt;li&gt;If the total fits in 96 GB minus a few GB of overhead, the RTX PRO 6000 is very likely cheaper per token. Benchmark it against your real prompt lengths.&lt;/li&gt;
&lt;li&gt;If you need 2 GPUs, test both — the result could go either way.&lt;/li&gt;
&lt;li&gt;If you need 4 or more GPUs, start from the H100 (or H200), unless your RTX PRO 6000 rate is below the break-even price for your workload.&lt;/li&gt;
&lt;li&gt;Re-run the cost formula with the prices you are actually quoted. Market medians move every month.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is the RTX PRO 6000 faster than the H100 for LLM inference?&lt;/strong&gt;&lt;br&gt;
Not usually per card. In independent vLLM tests the H100 was between 5% slower and 12% faster on a single-GPU model, and 45–72% faster once tensor parallelism across 4–8 GPUs was involved. The RTX PRO 6000's case is cost per token, not raw speed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much does it cost to rent an RTX PRO 6000 versus an H100?&lt;/strong&gt;&lt;br&gt;
As of October 7, 2026, the median on-demand price across tracked providers was $2.20 per GPU-hour for the RTX PRO 6000 and $3.49 for the H100, per getdeploying.com.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the RTX PRO 6000's memory bandwidth?&lt;/strong&gt;&lt;br&gt;
1,597 GB/s for the Server Edition and 1,792 GB/s for the Workstation Edition, both with 96 GB of GDDR7. The H100 SXM has 3,350 GB/s of HBM3.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many tokens per second can an RTX PRO 6000 serve?&lt;/strong&gt;&lt;br&gt;
It depends heavily on model, precision and concurrency. One public data point: about 3,140 total tokens/s on GLM-4.5-Air (AWQ 4-bit) at 256–512 concurrent requests, 1K input / 1K output, on a Workstation Edition card.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is the "RTX 6000 Blackwell" the same GPU as the "RTX PRO 6000"?&lt;/strong&gt;&lt;br&gt;
Yes — both names refer to the 96 GB Blackwell card, sold in Workstation, Max-Q and Server editions. It is a different card from the RTX 6000 Ada (48 GB) and the older Quadro RTX 6000 (24 GB).&lt;/p&gt;

&lt;h2&gt;
  
  
  Method and caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Nothing here was measured by the author.&lt;/strong&gt; Throughput comes from CloudRift's published benchmarks, prices from getdeploying.com's public listings, specs from NVIDIA's product pages, and model shapes from each model's &lt;code&gt;config.json&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prices are medians.&lt;/strong&gt; Your actual price determines the answer, which is why the formula and break-even prices are included.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Throughput ratios are workload-specific.&lt;/strong&gt; The benchmarks used GLM and Qwen3-Coder models; a different model, quantization, or serving stack will shift them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run A likely flatters the RTX PRO 6000 slightly&lt;/strong&gt;, because it used the Workstation Edition, with about 12% more bandwidth than the Server Edition most clouds rent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Units:&lt;/strong&gt; in the VRAM table, GB means GiB (2³⁰ bytes), the unit &lt;code&gt;nvidia-smi&lt;/code&gt; reports.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduce the numbers
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;GiB&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;

&lt;span class="c1"&gt;# Prices: getdeploying.com on-demand medians, 2026-10-07
&lt;/span&gt;&lt;span class="n"&gt;RTX_HR&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;H100_HR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;2.20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;3.49&lt;/span&gt;

&lt;span class="c1"&gt;# (rtx tok/s, h100 tok/s) from CloudRift's published runs
&lt;/span&gt;&lt;span class="n"&gt;RUNS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1 GPU, 1K/1K&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;3140.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;2987.00&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1 GPU, 8K/8K&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;2290.69&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;2556.03&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;4-GPU TP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;      &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;1602.96&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;2328.63&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8-GPU TP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;      &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;1651.67&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;2833.77&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rtx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;h100&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;RUNS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;rel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;RTX_HR&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;H100_HR&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;h100&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;rtx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; RTX cost/token = &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;rel&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; of H100 | break-even $&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;H100_HR&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;h100&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;rtx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/hr&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# KV-cache headroom: FP8 KV cache, 4 GiB runtime reserve, 32K-token sequences
&lt;/span&gt;&lt;span class="n"&gt;MODELS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;  &lt;span class="c1"&gt;# name, params (B), bytes/param, layers, kv_heads, head_dim
&lt;/span&gt;    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen3-Coder-30B-A3B BF16&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;30.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;48&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen3-32B BF16&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;           &lt;span class="mf"&gt;32.8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Llama-3.3-70B FP8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;        &lt;span class="mf"&gt;70.6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;wb&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;L&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kvh&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hd&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;MODELS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;weights&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;1e9&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;wb&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;GiB&lt;/span&gt;
    &lt;span class="n"&gt;kv_seq&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;L&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;kvh&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;hd&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;32_768&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;GiB&lt;/span&gt;
    &lt;span class="n"&gt;fits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;cap&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;cap&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;weights&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;kv_seq&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;cap&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;96&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;26&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; H100: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fits&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; seqs | RTX PRO 6000: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fits&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;96&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; seqs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;NVIDIA, &lt;a href="https://www.nvidia.com/en-us/data-center/rtx-pro-6000-blackwell-server-edition/" rel="noopener noreferrer"&gt;RTX PRO 6000 Blackwell Server Edition&lt;/a&gt; and &lt;a href="https://www.nvidia.com/content/dam/en-zz/Solutions/data-center/rtx-pro-6000-blackwell-workstation-edition/workstation-blackwell-rtx-pro-6000-workstation-edition-nvidia-us-3519208-web.pdf" rel="noopener noreferrer"&gt;Workstation Edition datasheet&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NVIDIA, &lt;a href="https://www.nvidia.com/en-us/data-center/h100/" rel="noopener noreferrer"&gt;H100 Tensor Core GPU&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;CloudRift, &lt;a href="https://www.cloudrift.ai/blog/benchmarking-rtx6000-vs-datacenter-gpus" rel="noopener noreferrer"&gt;RTX PRO 6000 vs H100, H200, and L40S: LLM Inference&lt;/a&gt; (Nov 2025)&lt;/li&gt;
&lt;li&gt;CloudRift, &lt;a href="https://www.cloudrift.ai/blog/benchmarking-b200" rel="noopener noreferrer"&gt;Benchmarking LLM Inference on B200, H200, H100, and RTX PRO 6000&lt;/a&gt; (Jan 2026)&lt;/li&gt;
&lt;li&gt;getdeploying.com, &lt;a href="https://getdeploying.com/gpus/nvidia-rtx-pro-6000" rel="noopener noreferrer"&gt;RTX PRO 6000 rental prices&lt;/a&gt; and &lt;a href="https://getdeploying.com/gpus/nvidia-h100" rel="noopener noreferrer"&gt;H100 rental prices&lt;/a&gt; (checked 2026-10-07)&lt;/li&gt;
&lt;li&gt;Model configs: &lt;a href="https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct" rel="noopener noreferrer"&gt;Qwen3-Coder-30B-A3B-Instruct&lt;/a&gt;, &lt;a href="https://huggingface.co/Qwen/Qwen3-32B" rel="noopener noreferrer"&gt;Qwen3-32B&lt;/a&gt;, &lt;a href="https://huggingface.co/unsloth/Llama-3.3-70B-Instruct" rel="noopener noreferrer"&gt;Llama-3.3-70B-Instruct&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




</description>
      <category>ai</category>
      <category>hardware</category>
      <category>llm</category>
      <category>performance</category>
    </item>
    <item>
      <title>Parameter count is a bad way to pick an open model for one GPU</title>
      <dc:creator>yyyysa4</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:34:54 +0000</pubDate>
      <link>https://dev.to/yyyysa4/parameter-count-is-a-bad-way-to-pick-an-open-model-for-one-gpu-21pi</link>
      <guid>https://dev.to/yyyysa4/parameter-count-is-a-bad-way-to-pick-an-open-model-for-one-gpu-21pi</guid>
      <description>&lt;p&gt;Model selection usually starts with a number: 8B, 27B, 30B. It is the first thing on the model card and the first thing in every comparison thread. It is also close to useless for the question most teams are actually asking, which is &lt;em&gt;how much traffic will this serve on the card I have.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Below is that question worked out for four current open models on a single 96 GB GPU. Everything here is arithmetic on published configs — no benchmarks were run. The script is at the end and takes about a second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Units:&lt;/strong&gt; GB throughout means GiB (2³⁰ bytes), which is what &lt;code&gt;nvidia-smi&lt;/code&gt; reports.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four models
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Total params&lt;/th&gt;
&lt;th&gt;Architecture&lt;/th&gt;
&lt;th&gt;Native context&lt;/th&gt;
&lt;th&gt;License&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/Qwen/Qwen3-8B" rel="noopener noreferrer"&gt;Qwen3-8B&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;8.2B&lt;/td&gt;
&lt;td&gt;Dense, 36 layers, GQA 32Q/8KV, head dim 128&lt;/td&gt;
&lt;td&gt;32,768 (131K with YaRN)&lt;/td&gt;
&lt;td&gt;Apache-2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/Qwen/Qwen3.5-9B" rel="noopener noreferrer"&gt;Qwen3.5-9B&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;~9B&lt;/td&gt;
&lt;td&gt;Hybrid: 3 linear-attention layers per 1 full, 32 total, 4 KV heads, head dim 256&lt;/td&gt;
&lt;td&gt;262,144 (1.01M with YaRN)&lt;/td&gt;
&lt;td&gt;Apache-2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/google/gemma-4-26B-A4B-it" rel="noopener noreferrer"&gt;Gemma-4-26B-A4B-it&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;25.2B (3.8B active)&lt;/td&gt;
&lt;td&gt;MoE, 128 experts, top-8 + 1 shared, 30 layers, 5 full-attention + 25 sliding (window 1024), 8 KV heads, head dim 256&lt;/td&gt;
&lt;td&gt;262,144&lt;/td&gt;
&lt;td&gt;Apache-2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B" rel="noopener noreferrer"&gt;Qwen3.8-27B&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;~27B&lt;/td&gt;
&lt;td&gt;Hybrid: same 3:1 pattern, 64 layers, 4 KV heads, head dim 256&lt;/td&gt;
&lt;td&gt;262,144 (1M with YaRN)&lt;/td&gt;
&lt;td&gt;Apache-2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every figure in that table comes from the model card or &lt;code&gt;config.json&lt;/code&gt; of the repo it links to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Weights: where MoE stops helping
&lt;/h2&gt;

&lt;p&gt;Mixture-of-Experts models are usually introduced with their active parameter count, because that is what determines compute per token. Gemma-4-26B-A4B activates 3.8B parameters out of 25.2B.&lt;/p&gt;

&lt;p&gt;Memory does not work that way. Routing is per-token and unpredictable, so every expert has to be resident. At BF16:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Qwen3-8B — &lt;strong&gt;15.3 GB&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Qwen3.5-9B — &lt;strong&gt;16.8 GB&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Gemma-4-26B-A4B — &lt;strong&gt;46.9 GB&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Qwen3.8-27B — &lt;strong&gt;50.3 GB&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The "4B active" model costs nearly three times the VRAM of the 9B. What you buy for that is compute: it runs at roughly 4B-model speed per token while holding 26B-model knowledge. That is a real and valuable trade. It is just not a memory saving, and model cards that lead with the active count invite exactly that misreading.&lt;/p&gt;

&lt;h2&gt;
  
  
  KV cache: the term nobody budgets
&lt;/h2&gt;

&lt;p&gt;Weights are static. KV cache grows with every token in flight, and it is where single-GPU deployments actually fall over. For a standard attention layer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;bytes per token = 2 (K and V) × layers × kv_heads × head_dim × dtype_bytes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The variable that matters most is &lt;strong&gt;layers&lt;/strong&gt; — specifically, how many layers keep a cache that grows with sequence length. Three of these four models do something to reduce that count:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3-8B&lt;/strong&gt; is conventional. All 36 layers cache. &lt;code&gt;2 × 36 × 8 × 128 × 2 = 144 KB per token&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3.5-9B&lt;/strong&gt; interleaves 3 linear-attention layers for every 1 full-attention layer. Linear attention carries a fixed-size recurrent state instead of a growing cache, so only 8 of 32 layers contribute: &lt;code&gt;2 × 8 × 4 × 256 × 2 = 32 KB per token&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3.8-27B&lt;/strong&gt; uses the same 3:1 pattern over 64 layers, so 16 layers cache: &lt;strong&gt;64 KB per token&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemma-4-26B-A4B&lt;/strong&gt; splits differently: 5 full-attention layers grow with the sequence (40 KB per token), while the other 25 use a 1024-token sliding window and stay flat at about 200 MB per sequence no matter how long the context gets.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So an 8B dense model spends &lt;strong&gt;4.5× more KV memory per token&lt;/strong&gt; than a 9B hybrid. Parameter count predicted the opposite.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft2yzu30iacbcrlsno24t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft2yzu30iacbcrlsno24t.png" alt=" " width="800" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What that buys in concurrency
&lt;/h2&gt;

&lt;p&gt;Take 96 GB, subtract weights, reserve 4 GB for activations and framework overhead, and divide by the KV cost of one 32,768-token sequence:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5j1gndcey429lsxnhz5n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5j1gndcey429lsxnhz5n.png" alt=" " width="800" height="306"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;KV per 32K sequence&lt;/th&gt;
&lt;th&gt;Concurrent 32K sequences&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-8B&lt;/td&gt;
&lt;td&gt;4.50 GB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;17&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-9B&lt;/td&gt;
&lt;td&gt;1.00 GB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;75&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma-4-26B-A4B&lt;/td&gt;
&lt;td&gt;1.45 GB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;31&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8-27B&lt;/td&gt;
&lt;td&gt;2.00 GB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The ranking has almost nothing to do with size. Qwen3.5-9B — nominally the second-smallest model here — serves four times the concurrent long-context traffic of Qwen3-8B, and nearly four times that of the 27B. The 26B MoE, despite spending 47 GB on weights, still fits more concurrent sequences than the 8B dense model does.&lt;/p&gt;

&lt;p&gt;If your workload is RAG, document processing, or anything else that puts long prompts in flight, this table is a better predictor of cost per request than any leaderboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why there is no quality column here
&lt;/h2&gt;

&lt;p&gt;There is an obvious missing dimension, and it is missing on purpose.&lt;/p&gt;

&lt;p&gt;The four model cards report largely disjoint benchmark sets. Gemma-4-26B-A4B-it reports MMLU-Pro 82.6%, GPQA Diamond 82.3%, LiveCodeBench v6 77.1%. Qwen3.5-9B reports MMLU-Pro 82.5 and C-Eval 88.2. Qwen3.8-27B reports GPQA Diamond 89.2 and LiveCodeBench v6 90.3. Qwen3-8B's card reports neither.&lt;/p&gt;

&lt;p&gt;Two of those MMLU-Pro numbers are within 0.1 of each other, which is exactly the kind of coincidence that produces a confident and wrong blog post. Vendor-run evaluations differ in prompt format, few-shot count, decoding parameters, and whether reasoning mode was enabled — none of which are consistently disclosed. Putting them in one table implies a comparison that the underlying numbers do not support.&lt;/p&gt;

&lt;p&gt;The honest version: use these figures to shortlist, then run your own eval on your own prompts. Memory math transfers between deployments. Benchmark scores often do not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this analysis does not tell you
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Nothing about throughput.&lt;/strong&gt; Memory determines what fits. Tokens per second depends on memory bandwidth, kernel quality, batching policy, and the serving stack. A model that fits 75 sequences will not necessarily serve them fast.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quantization changes every number.&lt;/strong&gt; FP8 weights roughly halve the weight column; FP8 or INT8 KV cache halves the cache column. The relative ordering mostly survives, the absolute headroom does not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Linear attention is a trade, not a free win.&lt;/strong&gt; The architectures that cut KV cost do so by compressing history into a fixed-size state. That is very good for throughput and is not free on tasks that need precise long-range recall. Test it rather than assuming the memory win is costless.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real allocators are messier.&lt;/strong&gt; vLLM and SGLang preallocate a KV pool via &lt;code&gt;gpu_memory_utilization&lt;/code&gt;; fragmentation, CUDA graphs, and multimodal encoders all take a cut. Treat the concurrency numbers as ceilings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parameter counts are approximate.&lt;/strong&gt; Qwen publishes "9B" and "27B" rather than exact totals, so those weight figures carry a percent or two of slack.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;GiB&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;
&lt;span class="n"&gt;BYTES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;  &lt;span class="c1"&gt;# BF16 KV cache
&lt;/span&gt;
&lt;span class="c1"&gt;# name, params(B), full_attn_layers, sliding_layers, window, kv_heads, head_dim
&lt;/span&gt;&lt;span class="n"&gt;MODELS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen3-8B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;         &lt;span class="mf"&gt;8.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;36&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;    &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen3.5-9B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="mf"&gt;9.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;    &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Gemma-4-26B-A4B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;25.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;25&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen3.8-27B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;     &lt;span class="mf"&gt;27.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;    &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;kv_for_seq&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;full_l&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sl_l&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;win&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kvh&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;seq&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;growing&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;full_l&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;kvh&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;hd&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;BYTES&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;seq&lt;/span&gt;
    &lt;span class="n"&gt;capped&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;sl_l&lt;/span&gt;   &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;kvh&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;hd&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;BYTES&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seq&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;win&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;growing&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;capped&lt;/span&gt;

&lt;span class="n"&gt;VRAM&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;RESERVE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SEQ&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;96&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;32_768&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;full_l&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sl_l&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;win&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kvh&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hd&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;MODELS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;weights&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;1e9&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;BYTES&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;GiB&lt;/span&gt;
    &lt;span class="n"&gt;per_seq&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;kv_for_seq&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;full_l&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sl_l&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;win&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kvh&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SEQ&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;GiB&lt;/span&gt;
    &lt;span class="n"&gt;free&lt;/span&gt;    &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;VRAM&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;weights&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;RESERVE&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;18&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; weights &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;weights&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;5.1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; GB | &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
          &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;per_seq&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;4.2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; GB/seq | &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;free&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;per_seq&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; concurrent &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;SEQ&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;K seqs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Change &lt;code&gt;SEQ&lt;/code&gt; to your real context length and &lt;code&gt;BYTES&lt;/code&gt; to 1 if you serve an FP8 KV cache. The layer counts come straight from each repo's &lt;code&gt;config.json&lt;/code&gt; — if a model ships a new revision, re-read it rather than trusting this table.&lt;/p&gt;




</description>
      <category>ai</category>
      <category>hardware</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Best models to run on RTX Pro 6000 96GB</title>
      <dc:creator>yyyysa4</dc:creator>
      <pubDate>Thu, 24 Sep 2026 03:43:33 +0000</pubDate>
      <link>https://dev.to/yyyysa4/best-models-to-run-on-rtx-pro-6000-96gb-26kh</link>
      <guid>https://dev.to/yyyysa4/best-models-to-run-on-rtx-pro-6000-96gb-26kh</guid>
      <description>&lt;p&gt;The right model depends on the task, not on the GPU. A 96 GB RTX Pro 6000 holds a 20B to 35B open model with room for context and batching, so you get to pick by the job instead of by a memory ceiling. For coding and agents that means Qwen3 Coder 30B; for low-cost chat, gpt-oss 20B; for retrieval, an embedding model with a reranker; for voice, Whisper and Kokoro. All of them run through the same OpenAI-compatible API at &lt;code&gt;https://api.ecohash.com/v1&lt;/code&gt;, so you start on the shared API and move to a dedicated endpoint or a workspace later without touching your code.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Pick by task, not by memory.&lt;/strong&gt; Coding: &lt;code&gt;qwen3-coder-30b-a3b-instruct&lt;/code&gt;. Chat: &lt;code&gt;gpt-oss-20b&lt;/code&gt;. Retrieval: &lt;code&gt;jina-embeddings-v3&lt;/code&gt; + &lt;code&gt;bge-reranker-v2-m3&lt;/code&gt;. Voice: &lt;code&gt;whisper-large-v3-turbo&lt;/code&gt; + &lt;code&gt;kokoro-82m&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Model by task
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Model ID&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Coding, agents, tool use&lt;/td&gt;
&lt;td&gt;Qwen3 Coder 30B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qwen3-coder-30b-a3b-instruct&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General chat, low cost&lt;/td&gt;
&lt;td&gt;gpt-oss 20B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gpt-oss-20b&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General chat, higher capability&lt;/td&gt;
&lt;td&gt;GLM 5.2, DeepSeek V4 Flash&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;GLM-5.2&lt;/code&gt;, &lt;code&gt;DeepSeek-V4-Flash&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Embeddings for RAG&lt;/td&gt;
&lt;td&gt;Jina v3, Qwen3 Embedding&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;jina-embeddings-v3&lt;/code&gt;, &lt;code&gt;qwen3-embedding-8b&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reranking for RAG&lt;/td&gt;
&lt;td&gt;BGE reranker v2 m3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;bge-reranker-v2-m3&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speech to text&lt;/td&gt;
&lt;td&gt;Whisper Large V3 Turbo&lt;/td&gt;
&lt;td&gt;&lt;code&gt;whisper-large-v3-turbo&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text to speech&lt;/td&gt;
&lt;td&gt;Kokoro 82M&lt;/td&gt;
&lt;td&gt;&lt;code&gt;kokoro-82m&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The higher-capability chat models run through the hosted API, not on a single 96 GB card. Here is what the single-card models measure end-to-end, in July 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Measured&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qwen3-coder-30b-a3b-instruct&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;30 ms TTFT, ~110 tok/s, 32k context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General chat&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gpt-oss-20b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;170 ms TTFT, ~125 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General chat&lt;/td&gt;
&lt;td&gt;&lt;code&gt;llama-3.1-8b-instruct&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;205 ms TTFT, ~143 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speech to text&lt;/td&gt;
&lt;td&gt;&lt;code&gt;whisper-large-v3-turbo&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;WER 4.37, ~0.5 s per 10 s of audio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text to speech&lt;/td&gt;
&lt;td&gt;&lt;code&gt;kokoro-82m&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;first audio ~120 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Against other providers of the same open models, we serve them at full BF16 precision and land competitively on price and speed (&lt;code&gt;llama-3.1-8b-instruct&lt;/code&gt;, EcoHash in purple):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3qkji4zd2xbkf173zdn3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3qkji4zd2xbkf173zdn3.png" alt="Llama-3.1-8B price vs speed: EcoHash vs peers" width="800" height="645"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Full data and charts: &lt;a href="https://github.com/ecohash-ai/ecohash-benchmarks" rel="noopener noreferrer"&gt;ecohash-benchmarks&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Facts
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GPU&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;NVIDIA RTX Pro 6000 Blackwell Server Edition, 96 GB GDDR7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Single-card sweet spot&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;20B to 35B open models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;API base URL&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://api.ecohash.com/v1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Access&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Shared API, dedicated endpoint, GPU workspace&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Price&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;RTX Pro 6000 at $1.89 per GPU-hour (&lt;a href="https://ecohash.com/pricing" rel="noopener noreferrer"&gt;pricing&lt;/a&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Who this is for
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Developers choosing an open model for a coding, chat, retrieval, or voice feature.&lt;/li&gt;
&lt;li&gt;Teams deciding what to run on a 96 GB card versus what to call through a hosted API.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What you can decide
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Which model to start with for your task.&lt;/li&gt;
&lt;li&gt;Whether it fits on a single RTX Pro 6000 or is better used through the hosted API.&lt;/li&gt;
&lt;li&gt;When to move from the shared API to a dedicated endpoint or a workspace.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to choose
&lt;/h2&gt;

&lt;p&gt;Four questions get you most of the way:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;What is the task?&lt;/strong&gt; Code, chat, retrieval, or speech. Start here, not with the hardware.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does it fit in 96 GB&lt;/strong&gt; at your context length and batch size? A 20B to 35B model fits with room to spare, models near 70B need quantization, and 235B and up are better through the hosted API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What are your latency and throughput needs?&lt;/strong&gt; Steady, high-volume traffic is the reason to reserve a dedicated endpoint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do you need to self-host at all?&lt;/strong&gt; For most teams the shared API is enough, and the GPU is our problem, not yours.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Coding
&lt;/h2&gt;

&lt;p&gt;Qwen3 Coder 30B (&lt;code&gt;qwen3-coder-30b-a3b-instruct&lt;/code&gt;) is the default for code generation, refactoring, and tool-using agents. It is a Mixture-of-Experts model with about 30B total parameters, it fits comfortably on one 96 GB card, and it serves a 32,768-token context at about 110 tokens/sec single-stream with a 30 ms time to first token. See &lt;a href="https://ecohash.com/models/qwen3-coder-30b-a3b-instruct" rel="noopener noreferrer"&gt;the model page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For lighter or cheaper high-volume coding, gpt-oss 20B (&lt;code&gt;gpt-oss-20b&lt;/code&gt;) is the smaller option. The three deployment paths for Qwen3 Coder are covered in &lt;a href="https://ecohash.com/blog/qwen3-coder-30b-rtx-pro-6000" rel="noopener noreferrer"&gt;Qwen3 Coder 30B on RTX Pro 6000&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  General chat
&lt;/h2&gt;

&lt;p&gt;Start with gpt-oss 20B (&lt;code&gt;gpt-oss-20b&lt;/code&gt;) for low-cost chat and assistants. It fits on one card, has a long context, and reaches about 8,900 output tokens/sec under load.&lt;/p&gt;

&lt;p&gt;For stronger general capability, we also serve larger models such as GLM 5.2 (&lt;code&gt;GLM-5.2&lt;/code&gt;) and DeepSeek V4 Flash (&lt;code&gt;DeepSeek-V4-Flash&lt;/code&gt;), and very large Mixture-of-Experts models like Qwen3-235B-A22B (&lt;code&gt;Qwen3-235B-A22B&lt;/code&gt;). Those are not single-card deployments. Use them through the shared API or a dedicated endpoint and let us handle the hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval (RAG)
&lt;/h2&gt;

&lt;p&gt;A retrieval pipeline needs two models: an embedding model to index and search text, and a reranker to reorder the top results.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Embeddings:&lt;/strong&gt; &lt;code&gt;jina-embeddings-v3&lt;/code&gt;, &lt;code&gt;jina-embeddings-v4&lt;/code&gt;, or the Qwen3 Embedding family (&lt;code&gt;qwen3-embedding-8b&lt;/code&gt;, &lt;code&gt;qwen3-embedding-4b&lt;/code&gt;, &lt;code&gt;qwen3-embedding-0.6b&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reranking:&lt;/strong&gt; &lt;code&gt;bge-reranker-v2-m3&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of them run through the same key, so one service can embed, rerank, and generate the answer without stitching vendors together. See the &lt;a href="https://ecohash.com/use-cases/rag" rel="noopener noreferrer"&gt;RAG use case&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Voice
&lt;/h2&gt;

&lt;p&gt;A voice agent uses three model types: speech to text, a chat model, and text to speech.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Speech to text:&lt;/strong&gt; &lt;code&gt;whisper-large-v3-turbo&lt;/code&gt; or &lt;code&gt;qwen3-asr-1-7b&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text to speech:&lt;/strong&gt; &lt;code&gt;kokoro-82m&lt;/code&gt; or &lt;code&gt;qwen3-tts&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These speech models are small and do not need an RTX Pro 6000. They run through the API and move onto GPU capacity only when throughput grows. The full build is in &lt;a href="https://ecohash.com/blog/kokoro-whisper-voice-agent" rel="noopener noreferrer"&gt;Build a voice agent with Whisper, Kokoro, and an OpenAI-compatible API&lt;/a&gt;; see also &lt;a href="https://ecohash.com/use-cases/voice-speech" rel="noopener noreferrer"&gt;voice and speech&lt;/a&gt;, &lt;a href="https://ecohash.com/models/kokoro-82m" rel="noopener noreferrer"&gt;Kokoro 82M&lt;/a&gt;, and &lt;a href="https://ecohash.com/models/whisper-large-v3-turbo" rel="noopener noreferrer"&gt;Whisper Large V3 Turbo&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calling these models
&lt;/h2&gt;

&lt;p&gt;The API is OpenAI-compatible. A coding call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.ecohash.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eco_...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qwen3-coder-30b-a3b-instruct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refactor this to be iterative:&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;def fib(n):&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;    return n if n &amp;lt; 2 else fib(n-1) + fib(n-2)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An embeddings call for a RAG index:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;emb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;embeddings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jina-embeddings-v3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EcoHash serves open models through an OpenAI-compatible API.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Shared API, dedicated endpoint, or workspace
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Shared API:&lt;/strong&gt; no GPU to manage, pay per token, good for building and variable traffic. See &lt;a href="https://ecohash.com/inference" rel="noopener noreferrer"&gt;inference&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dedicated endpoint:&lt;/strong&gt; reserved capacity for one model, steadier throughput and more predictable latency. See &lt;a href="https://ecohash.com/dedicated-inference" rel="noopener noreferrer"&gt;dedicated inference&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPU workspace:&lt;/strong&gt; an hourly RTX Pro 6000 for experiments, fine-tuning, and batch jobs. See &lt;a href="https://ecohash.com/gpu-compute" rel="noopener noreferrer"&gt;GPU compute&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Model fit depends on quantization, context length, and batch size. The ranges here are starting points to validate, not guarantees.&lt;/li&gt;
&lt;li&gt;The higher-capability chat models are for the hosted API, not single-card serving.&lt;/li&gt;
&lt;li&gt;Real performance depends on your setup. Reproduce any figure with its method and date.&lt;/li&gt;
&lt;li&gt;Model IDs change as the catalog changes. Check &lt;a href="https://ecohash.com/models" rel="noopener noreferrer"&gt;the model catalog&lt;/a&gt; before you build.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is the best model to run on an RTX Pro 6000 96GB?&lt;/strong&gt; It depends on the task. Coding: &lt;code&gt;qwen3-coder-30b-a3b-instruct&lt;/code&gt;. General chat: &lt;code&gt;gpt-oss-20b&lt;/code&gt;. Retrieval: &lt;code&gt;jina-embeddings-v3&lt;/code&gt; with &lt;code&gt;bge-reranker-v2-m3&lt;/code&gt;. Speech: &lt;code&gt;whisper-large-v3-turbo&lt;/code&gt; with &lt;code&gt;kokoro-82m&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the best coding model for the RTX Pro 6000?&lt;/strong&gt; Qwen3 Coder 30B (&lt;code&gt;qwen3-coder-30b-a3b-instruct&lt;/code&gt;). It fits on one 96 GB card and serves a 32,768-token context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is a good low-cost general chat model?&lt;/strong&gt; gpt-oss 20B (&lt;code&gt;gpt-oss-20b&lt;/code&gt;). It fits on one card and has a long context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I run a 235B model on one RTX Pro 6000?&lt;/strong&gt; No. Use a model like &lt;code&gt;Qwen3-235B-A22B&lt;/code&gt; through the hosted API instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What models do I need for RAG?&lt;/strong&gt; An embedding model and a reranker, for example &lt;code&gt;jina-embeddings-v3&lt;/code&gt; and &lt;code&gt;bge-reranker-v2-m3&lt;/code&gt;, plus a chat model to write the answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do speech models need an RTX Pro 6000?&lt;/strong&gt; No. Whisper and Kokoro run through the API, and move onto GPU capacity only when throughput demands it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many models can I use with one API key?&lt;/strong&gt; All of them. Every model shares one key and one base URL, &lt;code&gt;https://api.ecohash.com/v1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much does this cost?&lt;/strong&gt; API usage is billed per token; a GPU workspace runs at $1.89 per GPU-hour. See &lt;a href="https://ecohash.com/pricing" rel="noopener noreferrer"&gt;pricing&lt;/a&gt; for current rates.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next steps
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Browse all models: &lt;a href="https://ecohash.com/models" rel="noopener noreferrer"&gt;ecohash.com/models&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;See the RTX Pro 6000 page: &lt;a href="https://ecohash.com/gpu-compute" rel="noopener noreferrer"&gt;ecohash.com/gpu-compute&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Reserve capacity: &lt;a href="https://ecohash.com/dedicated-inference" rel="noopener noreferrer"&gt;ecohash.com/dedicated-inference&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




</description>
    </item>
  </channel>
</rss>
