<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ivaylo Ivanov</title>
    <description>The latest articles on DEV Community by Ivaylo Ivanov (@ivaylopivanov).</description>
    <link>https://dev.to/ivaylopivanov</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4139266%2F81489b59-9725-4085-832e-cb98b70552fc.jpeg</url>
      <title>DEV Community: Ivaylo Ivanov</title>
      <link>https://dev.to/ivaylopivanov</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ivaylopivanov"/>
    <language>en</language>
    <item>
      <title>I Built a €2,300 PC to replace my AI Subscriptions</title>
      <dc:creator>Ivaylo Ivanov</dc:creator>
      <pubDate>Wed, 23 Sep 2026 10:40:23 +0000</pubDate>
      <link>https://dev.to/ivaylopivanov/i-built-a-eu2300-pc-to-replace-my-ai-subscriptions-44pg</link>
      <guid>https://dev.to/ivaylopivanov/i-built-a-eu2300-pc-to-replace-my-ai-subscriptions-44pg</guid>
      <description>&lt;p&gt;I wanted to see whether a reasonably priced desktop PC could replace some of the programming related work I currently delegate to AI subscriptions.&lt;br&gt;
At the time I started building it, in the first week of August, the local model requiring relatively little VRAM but still capable enough for most tasks seemed to be Qwen 3.6 35B A3B. Before buying any hardware, I decided to test the model first.&lt;br&gt;
I went through my cursor history and found a relatively small but complex task that I would normally delegate to Cursor or Codex. I ran the task through OpenRouter, where it cost me $4.73 and took 4 minutes and 23 seconds.&lt;br&gt;
That made me think local models might actually be a viable option within a reasonable budget. I decided to build a machine with 3 simple goals:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Low power consumption.&lt;/li&gt;
&lt;li&gt;32GB of VRAM, so I could run relatively small but capable models.&lt;/li&gt;
&lt;li&gt;Not declaring bankruptcy after buying the hardware.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I ended up with dual RTX 5060ti cards, giving me 32GB of aggregate VRAM.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Motherboard&lt;/td&gt;
&lt;td&gt;SAPPHIRE NITRO+ B850A WIFI 7&lt;/td&gt;
&lt;td&gt;160 €&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;AMD Ryzen 5 7500F (3.7GHz) TRAY&lt;/td&gt;
&lt;td&gt;135 €&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RAM&lt;/td&gt;
&lt;td&gt;Kingston Fury Beast 32GB (2×16GB) DDR5-5600 XMP&lt;/td&gt;
&lt;td&gt;419 €&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU 1&lt;/td&gt;
&lt;td&gt;Inno3D GeForce RTX 5060 ti 16GB TWIN X2 OC&lt;/td&gt;
&lt;td&gt;599 €&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU 2&lt;/td&gt;
&lt;td&gt;Gigabyte GeForce RTX 5060 ti EAGLE MAX OC 16G&lt;/td&gt;
&lt;td&gt;630 €&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PSU&lt;/td&gt;
&lt;td&gt;Corsair RM850e 850W 80+ Gold&lt;/td&gt;
&lt;td&gt;132 €&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Case&lt;/td&gt;
&lt;td&gt;DeepCool CG580&lt;/td&gt;
&lt;td&gt;62 €&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SSD&lt;/td&gt;
&lt;td&gt;ADATA XPG SPECTRIX S65G 500GB&lt;/td&gt;
&lt;td&gt;113 €&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cooler&lt;/td&gt;
&lt;td&gt;XIGMATEK LK 240 Digital&lt;/td&gt;
&lt;td&gt;53 €&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Total
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2,303 €&lt;/strong&gt; / &lt;strong&gt;$2637.84&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The build is capable enough and I'm pleased with the results. I wanted to make this build "sustainable", meaning, having it running in my office, next to me, during my working hours, sometimes during the nights or weekends. Because of this I had 3 concerns:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Noise: The case also has 5 fans that were not mentioned above because after the first few days I turned them off and removed the front panel of the case. This way the noise when both GPUs work at 100% is around 32db and temperature for the top never exceeds 67 degrees, the bottom one stays around 62 degrees. With the fans turned on the noise was around 46db even when idle.&lt;/li&gt;
&lt;li&gt;Electricity: Even at 100% utilisation the GPUs never go above 150W each but most of the time I see them between 100W and 144W. This, considering the efficient CPU which stays idle and way under 100W, gives me around 500W average power draw. For me this was important because if the PC was to consume around 1KW per hour, with the current electricity prices + the upfront cost the math was not mathin - one will be better off just paying for Codex / Claude Code / Cursor subscription if comparing prices only. I'm usually paying between 60 to 120 EUR per month for subscriptions and I always run out of tokens. So based on my math, if I can run this build locally and delegate most of the tasks to the local models, I should break even within 24 to 36 months, without considering electricity costs or potential changes to the subscription models offered by the AI labs. Personally for me the cost for electricity is not that big of a problem because I have solar panels and batteries, so for around 8 months of the year the cost will be exactly 0, for the other 4, assuming no electricity from the solar system, the worst case power consumption will be 500W×10h×30days=150kWh/month or in other words ~22 EUR/month based on the current electricity prices (based in Sofia, Bulgaria).&lt;/li&gt;
&lt;li&gt;Heat: These things produce lots of heat. In August and up until now the weather here was hot and if I didn't run the AC then the heat would have been too much for me. The AC is always turned on during the summer anyway so that was not a problem for me but one should absolutely consider it before buying a GPU that consumes around or above 300W.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;By the time I had the PC up and running the new qwen3.8 was released so I focused on it. I've tested many, many Qwen3.8 27B variations, the one that works the best for 24 to 32GB VRAM is &lt;a href="https://huggingface.co/unsloth/Qwen3.8-27B-GGUF?show_file_info=Qwen3.8-27B-UD-Q4_K_M.gguf" rel="noopener noreferrer"&gt;Qwen3.8-27B-UD-Q4-K-M.gguf&lt;/a&gt; - not too slow and the quality is good enough so that I don't have to step in for every task.&lt;/p&gt;

&lt;p&gt;The initial setup is almost straightforward. After installing all the NVIDIA and CUDA tooling, setting up the vLLM or the llama.cpp is easy. I ended up using llama.cpp with the following params:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;./build/bin/llama-server \
-m ~/models/qwen38/Qwen3.8-27B-UD-Q4_K_M.gguf \
--mmproj ~/models/qwen38/mmproj-F16.gguf \
-ngl 999 \
--split-mode tensor \
--tensor-split 1,1 \
--flash-attn on \
--ctx-size 400000 \
--cache-type-k q4_0 \
--cache-type-v q4_0 \
--cache-prompt \
--spec-type draft-mtp \
--cache-ram 16384 \
--spec-draft-n-max 4 \
--parallel 2 \
--host 0.0.0.0 \
--port 8080 \
--metrics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;which translates to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;It can process 2 parallel requests, each with 200k context which is enough for me. This does ~1100t/s prefill, 60+t/s tg on average (both streams combined). Sometimes the parallel processing can reach 100+ t/s tg:&lt;/p&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;33.00.149.228 I slot print_timing: id  1 | task 18530 | n_gen =    123, tg =  35.15 t/s, tg_3s =  35.43 t/s
33.03.189.061 I slot print_timing: id  1 | task 18530 | n_gen =    258, tg =  39.47 t/s, tg_3s =  44.41 t/s
33.03.679.450 I slot print_timing: id  0 | task 18562 | n_gen =    127, tg =  40.96 t/s, tg_3s =  41.29 t/s
33.06.223.392 I slot print_timing: id  1 | task 18530 | n_gen =    410, tg =  42.85 t/s, tg_3s =  50.09 t/s
33.06.710.003 I slot print_timing: id  0 | task 18562 | n_gen =    275, tg =  44.87 t/s, tg_3s =  48.84 t/s
33.09.273.407 I slot print_timing: id  1 | task 18530 | n_gen =    565, tg =  44.78 t/s, tg_3s =  50.82 t/s
33.09.761.511 I slot print_timing: id  0 | task 18562 | n_gen =    427, tg =  46.52 t/s, tg_3s =  49.81 t/s
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;It consumes almost all the available VRAM:&lt;/p&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 595.84                 Driver Version: 595.84         CUDA Version: 13.2     |
+-----------------------------------------+------------------------+----------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA GeForce RTX 5060 ti     Off |   00000000:01:00.0 Off |                  N/A |
| 55%   66C    P1            143W /  180W |   15135MiB /  16311MiB |     92%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+
|   1  NVIDIA GeForce RTX 5060 ti     Off |   00000000:04:00.0 Off |                  N/A |
| 51%   60C    P1            138W /  180W |   15117MiB /  16311MiB |     95%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|    0   N/A  N/A            1705      G   /usr/lib/xorg/Xorg                       10MiB |
|    0   N/A  N/A            2113      G   /usr/bin/gnome-shell                      3MiB |
|    0   N/A  N/A            2522      C   ...ma.cpp/build/bin/llama-server      15094MiB |
|    1   N/A  N/A            1705      G   /usr/lib/xorg/Xorg                        4MiB |
|    1   N/A  N/A            2522      C   ...ma.cpp/build/bin/llama-server      15094MiB |
+-----------------------------------------------------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;It contains the vision projector and based on my tests, it works well.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Split mode "layer" produces worse results for this configuration.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;I tried to overclock the GPUs but there were no meaningful results.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Code generation is fast (70+ t/s), code reviews are slow (40+ t/s).&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Tests and benchmarks
&lt;/h2&gt;

&lt;p&gt;After the initial delight I've decided to purchase a second hand 3090. The price of the used 3090 was around 900 EUR, way cheaper than the dual 5060 ti so I thought if I can achieve the same results with the cheaper one, why not? So I bought a used one for 830 EUR. Swapped the GPUs and decided to run a semi-controlled test. Took the same prompt that I used for my initial test for the Qwen 3.6 35B A3B. The task: existing project that requires changes in the HTTP endpoint, internal methods, SQL queries and unit tests in Golang. Has clear validation steps and the agent always succeeds. A run takes between 6 to 14 mins, sometimes it figures out the "tricky" parts quickly, sometimes it doesn't. Total files changed: 6, diff on average: +188, -88. It makes from 40 to 60 tool calls. If I have to do the task manually I will probably need between 30 to 60 mins.&lt;/p&gt;

&lt;p&gt;I did 5 runs on each setup, each from a cold start (restart llama after each run) with the following methodology:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;same model&lt;/li&gt;
&lt;li&gt;same quantization&lt;/li&gt;
&lt;li&gt;same prompt&lt;/li&gt;
&lt;li&gt;same context&lt;/li&gt;
&lt;li&gt;same llama configuration&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Single 3090 24GB
&lt;/h3&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;./build/bin/llama-server \
  -m ~/models/qwen38/Qwen3.8-27B-UD-Q4_K_M.gguf \
  --mmproj ~/models/qwen38/mmproj-F16.gguf \
  -ngl 999 \
  --flash-attn on \
  --ctx-size 192144 \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --cache-prompt \
  --spec-type draft-mtp \
  --cache-ram 16384 \
  --spec-draft-n-max 4 \
  --parallel 1 \
  --host 0.0.0.0 \
  --port 8080 \
  --metrics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Mean tg&lt;/th&gt;
&lt;th&gt;Median tg&lt;/th&gt;
&lt;th&gt;Mean tg-3s&lt;/th&gt;
&lt;th&gt;Min&lt;/th&gt;
&lt;th&gt;Max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;61.26 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;61.41&lt;/td&gt;
&lt;td&gt;66.08&lt;/td&gt;
&lt;td&gt;36.71&lt;/td&gt;
&lt;td&gt;77.01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;56.05 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;55.60&lt;/td&gt;
&lt;td&gt;59.89&lt;/td&gt;
&lt;td&gt;34.62&lt;/td&gt;
&lt;td&gt;78.89&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;55.65 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;52.45&lt;/td&gt;
&lt;td&gt;57.81&lt;/td&gt;
&lt;td&gt;33.52&lt;/td&gt;
&lt;td&gt;78.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;61.68 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;62.23&lt;/td&gt;
&lt;td&gt;64.82&lt;/td&gt;
&lt;td&gt;40.02&lt;/td&gt;
&lt;td&gt;76.22&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;58.02 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;56.65&lt;/td&gt;
&lt;td&gt;62.25&lt;/td&gt;
&lt;td&gt;33.72&lt;/td&gt;
&lt;td&gt;78.17&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Average Mean tg: 58.53 t/s&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;prompt eval time = 26520.51 ms&lt;/p&gt;

&lt;p&gt;330W to 350W sustained power consumption, lots of noise&lt;/p&gt;

&lt;p&gt;Note: there are models and instructions that in theory will produce better results with 3090 (eg &lt;a href="https://github.com/noonghunna/club-3090/blob/master/docs/SINGLE-CARD.md" rel="noopener noreferrer"&gt;https://github.com/noonghunna/club-3090/blob/master/docs/SINGLE-CARD.md&lt;/a&gt; and &lt;a href="https://github.com/syv-ai/HyperQwen" rel="noopener noreferrer"&gt;https://github.com/syv-ai/HyperQwen&lt;/a&gt;) but all of them fail to finish my task. Most go into infinite loops eventually or take too long to finish because they drift too much.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. dual 5060ti 32GB
&lt;/h3&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;./build/bin/llama-server \
  -m ~/models/qwen38/Qwen3.8-27B-UD-Q4_K_M.gguf \
  --mmproj ~/models/qwen38/mmproj-F16.gguf \
  -ngl 999 \
  --split-mode tensor \
  --tensor-split 1,1 \
  --flash-attn on \
  --ctx-size 192144 \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --cache-prompt \
  --spec-type draft-mtp \
  --cache-ram 16384 \
  --spec-draft-n-max 4 \
  --parallel 1 \
  --host 0.0.0.0 \
  --port 8080 \
  --metrics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Mean tg&lt;/th&gt;
&lt;th&gt;Median tg&lt;/th&gt;
&lt;th&gt;Mean tg-3s&lt;/th&gt;
&lt;th&gt;Min tg&lt;/th&gt;
&lt;th&gt;Max tg&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;61.30 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;62.67&lt;/td&gt;
&lt;td&gt;64.36&lt;/td&gt;
&lt;td&gt;38.00&lt;/td&gt;
&lt;td&gt;84.67&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;60.53 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;54.30&lt;/td&gt;
&lt;td&gt;63.00&lt;/td&gt;
&lt;td&gt;40.24&lt;/td&gt;
&lt;td&gt;83.66&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;58.48 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;54.04&lt;/td&gt;
&lt;td&gt;61.76&lt;/td&gt;
&lt;td&gt;39.08&lt;/td&gt;
&lt;td&gt;83.34&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;42&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;58.40 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;54.16&lt;/td&gt;
&lt;td&gt;60.69&lt;/td&gt;
&lt;td&gt;42.13&lt;/td&gt;
&lt;td&gt;90.11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;55.42 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;52.99&lt;/td&gt;
&lt;td&gt;59.01&lt;/td&gt;
&lt;td&gt;34.41&lt;/td&gt;
&lt;td&gt;80.08&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;59.97 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;57.84&lt;/td&gt;
&lt;td&gt;64.05&lt;/td&gt;
&lt;td&gt;38.76&lt;/td&gt;
&lt;td&gt;78.19&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Average Mean tg: 58.42 t/s&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;prompt eval time = 28052.82 ms&lt;/p&gt;

&lt;p&gt;250W to 280W power consumption, low noise&lt;/p&gt;

&lt;h3&gt;
  
  
  3. 1x3090 + 1x5060ti 38GB
&lt;/h3&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;./build/bin/llama-server \
  -m ~/models/qwen38/Qwen3.8-27B-UD-Q4_K_M.gguf \
  --mmproj ~/models/qwen38/mmproj-F16.gguf \
  -ngl 999 \
  --split-mode tensor \
  --tensor-split 1,1 \
  --flash-attn on \
  --ctx-size 192144 \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --cache-prompt \
  --spec-type draft-mtp \
  --cache-ram 16384 \
  --spec-draft-n-max 4 \
  --parallel 1 \
  --host 0.0.0.0 \
  --port 8080 \
  --metrics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Mean tg&lt;/th&gt;
&lt;th&gt;Median tg&lt;/th&gt;
&lt;th&gt;Mean tg-3s&lt;/th&gt;
&lt;th&gt;Min&lt;/th&gt;
&lt;th&gt;Max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;27&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;56.18 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;51.08&lt;/td&gt;
&lt;td&gt;58.45&lt;/td&gt;
&lt;td&gt;40.52&lt;/td&gt;
&lt;td&gt;82.64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;63.97 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;64.63&lt;/td&gt;
&lt;td&gt;67.51&lt;/td&gt;
&lt;td&gt;32.02&lt;/td&gt;
&lt;td&gt;86.75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;56.46 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;53.50&lt;/td&gt;
&lt;td&gt;61.09&lt;/td&gt;
&lt;td&gt;38.44&lt;/td&gt;
&lt;td&gt;80.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;62.31 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;62.27&lt;/td&gt;
&lt;td&gt;66.00&lt;/td&gt;
&lt;td&gt;38.94&lt;/td&gt;
&lt;td&gt;86.30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;64.15 t/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;63.76&lt;/td&gt;
&lt;td&gt;69.73&lt;/td&gt;
&lt;td&gt;44.47&lt;/td&gt;
&lt;td&gt;83.21&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Average Mean tg: 60.61 t/s&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;prompt eval time = 26710.29 ms&lt;/p&gt;

&lt;p&gt;3090 capped at 280W&lt;br&gt;
390W to 420W power consumption, moderate noise&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary of the results
&lt;/h2&gt;

&lt;p&gt;Based on the numbers above and current GPU prices in Bulgaria, the single 3090 is the cheaper option for roughly the same generation speed. However, the dual 5060 Ti setup gives me 32GB of VRAM, lower power consumption, much less noise, support for two parallel requests, and a 24-month warranty.&lt;/p&gt;




&lt;p&gt;How about 3 cards instead of 2? The motherboard has 3 slots for GPU (PCIe 5.0 x16, PCIe 4.0 x4, and PCIe 4.0 x2). I initially bought it with the idea to run 2 but the 3rd slot was still free. Every article online will tell you that slow PCIe affects the tg when running multiple GPUs. And it does. I eventually purchased a 3rd 5060 ti, plugged it using Linkup riser cable 4.0 x16 on the 3rd slot (PCIe 4.0 x2). Tried many different configurations but regardless of what I tried the prefill and the tg dropped by 30%.&lt;/p&gt;

&lt;p&gt;While trying different configurations and GPUs I also did various quality tests on 3.8-27B-UD-Q4-K-M.gguf and based on my tests with medium reasoning I can delegate tasks that will be always completed, although not always with the best code quality. If I have to rate it I would say it's somewhere between Opus 4.1 and Opus 4.5 but of course this is very subjective and depends on the tasks and the prompt. Funny enough, I also made tests for PR reviews on a block of code for which I knew upfront the bugs. The models I compared were OpenAI Sol high and 3.8-27B-UD-Q4-K-M xhigh. In 3 runs Sol did not find the critical bugs, the local model found them from the first run. In the end of the day it's a slop casino, maybe a few more runs would have been enough for Sol but I hit the limits. That being said the local model was more than 2 times slower and other tests have shown me how Sol demolished my local model for various tasks. What I'm trying to say is, Qwen is absolutely not a SOTA model but it's definitely a capable model and I do not need SOTA models for 99% of my tasks.&lt;br&gt;
Since I bought the components, prices have gone up, so it's harder to justify this exact build today. If budget isn't a concern, the new Mac Studio with the M5 Ultra and 96GB of RAM is a more attractive option for noise, power consumption, and memory, while also allowing you to run larger and more capable models. If that's outside your budget, the dual 5060 Ti setup is still a good option in my opinion.&lt;/p&gt;

&lt;p&gt;If you had told me three years ago that in 2026 I would be able to do most, if not all, of my programming with local models on a machine costing under €3,000, I wouldn't have believed you.&lt;/p&gt;

&lt;p&gt;I built my own harness for this machine, and I built the entire thing using the local model itself. So I know it works for my workflow and I've now cancelled all my AI subscriptions.&lt;/p&gt;

&lt;p&gt;Alibaba announced they will soon release the new Qwen4 family and we will have an even better 27B model. Exciting times.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>localmodels</category>
    </item>
  </channel>
</rss>
