<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: StudioTV</title>
    <description>The latest articles on DEV Community by StudioTV (@studiotv).</description>
    <link>https://dev.to/studiotv</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4175405%2F1a18de61-755f-4b1e-ae72-862af12f265f.png</url>
      <title>DEV Community: StudioTV</title>
      <link>https://dev.to/studiotv</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/studiotv"/>
    <language>en</language>
    <item>
      <title>Will it fit on my GPU? I built a calculator that finally answers it</title>
      <dc:creator>StudioTV</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:42:27 +0000</pubDate>
      <link>https://dev.to/studiotv/will-it-fit-on-my-gpu-i-built-a-calculator-that-finally-answers-it-5120</link>
      <guid>https://dev.to/studiotv/will-it-fit-on-my-gpu-i-built-a-calculator-that-finally-answers-it-5120</guid>
      <description>&lt;p&gt;Every time a new open model drops, the same question floods Reddit and Discord: will it run on my card? On 24 GB? On two of them? In 4-bit? At what context length?&lt;/p&gt;

&lt;p&gt;The usual answer is a back-of-the-envelope formula: parameters times bytes per weight, plus a KV cache computed as if every layer looked at the whole context. For a classic Llama-style model, that works. For most models released in the last year, it is wrong, sometimes by an order of magnitude.&lt;/p&gt;

&lt;p&gt;So I built the &lt;a href="https://studiotvai.com/llm-vram-calculator" rel="noopener noreferrer"&gt;StudioTV LLM VRAM Calculator&lt;/a&gt;. It is free, needs no account, and works with any model on Hugging Face. Here is what it does, and why its numbers may differ from what you are used to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The KV cache is where naive estimates break
&lt;/h2&gt;

&lt;p&gt;Model weights are the easy part. The hard part is the KV cache, the memory that grows with every token of context, and modern architectures cache very differently:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sliding-window attention.&lt;/strong&gt; Gemma 4 31B keeps a 1,024-token window on 50 of its 60 layers. Only 10 layers remember the whole context. At 128K tokens, the real cache is &lt;strong&gt;10.8 GiB&lt;/strong&gt;. Treat every layer as global and you get 120 GiB, eleven times too much.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Linear attention.&lt;/strong&gt; Qwen3.6 35B-A3B caches keys and values on only 10 of its 40 layers; the other 30 keep a small fixed-size state. At 128K: &lt;strong&gt;2.5 GiB&lt;/strong&gt;, not 10.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latent attention (MLA).&lt;/strong&gt; DeepSeek R1 stores a compact 576-wide latent per layer instead of full keys and values for its 128 heads. A multi-head formula predicts about 610 GiB at 128K; the real cache is &lt;strong&gt;8.6 GiB&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compressed attention.&lt;/strong&gt; DeepSeek V4 Flash compresses the sequence and keeps a 128-token window: about seven times less than the naive estimate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The calculator reads each model's configuration layer by layer and computes the cache the way vLLM and llama.cpp actually allocate it, including FP8 cache formats and the extra window llama.cpp keeps on sliding-window layers.&lt;/p&gt;

&lt;h2&gt;
  
  
  One question, answered in four steps
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Your setup.&lt;/strong&gt; Pick one of 53 curated models or paste any Hugging Face link. Choose the workload (inference, full fine-tuning, LoRA or QLoRA), the context length, how many requests you serve at once, your GPU and your engine. There are 60 GPUs, from data-center cards to laptop GPUs, Macs and other unified-memory machines, and four engines: vLLM, SGLang, TensorRT-LLM, and llama.cpp / Ollama / LM Studio.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Your answer.&lt;/strong&gt; How many GPUs you need, whether it fits, and what fills each card.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8yezr1bx50v7xfl2smb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8yezr1bx50v7xfl2smb.png" alt="What fills the card, plus a rough speed and cost. Here an API would be cheaper, and the calculator says so." width="800" height="1364"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;What fills the card, plus a rough speed and cost. Here an API would be cheaper, and the calculator says so.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In this example, Gemma 4 31B in Q4_K_M takes 17.1 GiB of weights, 3.7 GiB of KV cache for one 32K-token request and about 3 GiB of runtime overhead: 23.7 GiB out of the 28.8 GiB usable on the RTX 5090. The calculator also estimates the speed (about 54 tokens per second here) and your cost per million tokens next to API prices. When an API is cheaper than renting the hardware, it tells you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Better quality or a cheaper GPU?&lt;/strong&gt; The “Does it fit?” grid crosses every weight format with every context length on your hardware. Click a cell to apply it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8waz9f9ke3k6g3xyp0wo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8waz9f9ke3k6g3xyp0wo.png" alt="Every weight format against every context length on one RTX 5090." width="800" height="380"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Every weight format against every context length on one RTX 5090.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For Gemma 4 31B on one RTX 5090: Q4_K_M and Q5_K_M up to 32K tokens, Q3_K_M up to 128K. GGUF sizes follow llama.cpp's real quantization rules. Checked against the GGUF files actually published on Hugging Face, the median error is under 1% from Q8_0 down to Q3_K_M. Next to the grid, the calculator lists the cheapest GPUs that fit your setup and the other models that fit your card.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Rent and deploy.&lt;/strong&gt; On-demand prices from RunPod, Vast.ai, Verda and Azure, refreshed every hour, with a price history for each GPU, and a ready-to-paste command for vLLM, SGLang, llama.cpp or TensorRT-LLM.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2oolixk1jk89iy88y2fy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2oolixk1jk89iy88y2fy.png" alt="Hourly rental prices with their history, and the command to launch the model." width="800" height="638"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Hourly rental prices with their history, and the command to launch the model.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;At the time of writing, the RTX 5090 started at $0.46 an hour on Vast.ai, and the calculator pointed out an even cheaper option: two RTX 3090s for $0.29 an hour in total. The command for this setup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="nt"&gt;-hf&lt;/span&gt; unsloth/gemma-4-31B-it-GGUF:Q4_K_M &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt; 32768 &lt;span class="nt"&gt;--parallel&lt;/span&gt; 1 &lt;span class="nt"&gt;-ngl&lt;/span&gt; 99 &lt;span class="nt"&gt;-fa&lt;/span&gt; on
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  When it does not fit
&lt;/h2&gt;

&lt;p&gt;Ask for more than your card can hold and you get an offload plan instead of a dead end: how many MoE experts (&lt;code&gt;--n-cpu-moe&lt;/code&gt;) or layers (&lt;code&gt;-ngl&lt;/code&gt;) to keep in system RAM, and the speed to expect from your RAM, from dual-channel DDR4 to quad-channel DDR5.&lt;/p&gt;

&lt;h2&gt;
  
  
  A page for every model, updated every hour
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffnnzfblgrzlsn7j3ovps.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffnnzfblgrzlsn7j3ovps.png" alt="Each model has its own page, here Gemma 4 31B." width="800" height="768"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Each model has its own page, here Gemma 4 31B.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Each model has its own page: VRAM at every precision and context length, how many GPUs of each type it takes, the cheapest way to run it today, its KV cache curve and the real sizes of the published GGUF files.&lt;/p&gt;

&lt;p&gt;New models show up on their own. A watcher checks Hugging Face every hour, reads each new model's &lt;code&gt;config.json&lt;/code&gt; and safetensors headers without downloading the weights, and publishes a page as soon as it can model the architecture. The “New &amp;amp; trending” section follows Hugging Face trends, Ollama's most pulled models and the latest releases.&lt;/p&gt;
&lt;h2&gt;
  
  
  Ask your AI assistant
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6jcntnandh8sdugi9n82.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6jcntnandh8sdugi9n82.png" alt="The calculator is also an MCP server and a JSON API." width="798" height="262"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The calculator is also an MCP server and a JSON API.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The calculator is also an MCP server. Add it to Claude, ChatGPT or any MCP client, then ask “does Qwen3.6 35B fit on my RTX 4090?”: the assistant calls the calculator and answers with the same numbers. Three tools, no API key: &lt;code&gt;estimate_vram&lt;/code&gt;, &lt;code&gt;models_that_fit&lt;/code&gt; and &lt;code&gt;gpu_prices&lt;/code&gt;. In Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http studiotv-vram https://studiotvai.com/api/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is also a static JSON API and a VRAM badge you can paste into a model card.&lt;/p&gt;

&lt;h2&gt;
  
  
  On your phone too
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F98p36ksqouizjxaj2d9b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F98p36ksqouizjxaj2d9b.png" alt="On a phone, the answer stays pinned at the bottom of the screen while you change the settings." width="800" height="1731"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;On a phone, the answer stays pinned at the bottom of the screen while you change the settings.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How it's built
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Front end:&lt;/strong&gt; React 19, TypeScript and Vite. Every page (the calculator, one page per model, one per GPU) is prerendered to static HTML, so it loads fast and search engines see the numbers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model watcher:&lt;/strong&gt; every hour, a script looks for new models on Hugging Face. &lt;code&gt;config.json&lt;/code&gt; gives the architecture, and the header of each safetensors file, read with an HTTP Range request (a few hundred KiB instead of gigabytes), gives the exact size of every tensor. That is how vision towers or extra prediction layers that engines skip get counted right.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Serving:&lt;/strong&gt; the MCP server is a small Node service, and the JSON API is plain static files. Everything runs in Docker on a single VPS.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A layout check in the build:&lt;/strong&gt; the build opens every changed page in headless Chrome at 320, 390, 768 and 1280 px wide, scrolls it top to bottom, and fails if anything overflows the screen or throws a JavaScript error. It only rechecks pages whose code or structure changed, so it stays quick. I added it after a CSS change silently cut off the right side of the page on iPhones.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Limits, and how you can help
&lt;/h2&gt;

&lt;p&gt;These are estimates. Engine versions and settings change real usage, so keep some headroom and confirm with &lt;code&gt;nvidia-smi&lt;/code&gt;. Speed figures are rough (±30 to 50%) and do not model speculative decoding. Some “Rent” links are affiliate links; this is disclosed on the site, and providers are ranked by price only.&lt;/p&gt;

&lt;p&gt;If a number does not match what you see, paste your vLLM or llama.cpp startup log into the “Compare with your real usage” box: the calculator lines up the real allocation with its estimate. That feedback is exactly what makes it better.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it:&lt;/strong&gt; &lt;a href="https://studiotvai.com/llm-vram-calculator" rel="noopener noreferrer"&gt;studiotvai.com/llm-vram-calculator&lt;/a&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>llm</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
