<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: David </title>
    <description>The latest articles on DEV Community by David  (@purpledoubled).</description>
    <link>https://dev.to/purpledoubled</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3802440%2Fbd0118a6-e9df-4efa-965a-8f8f9c2ef510.png</url>
      <title>DEV Community: David </title>
      <link>https://dev.to/purpledoubled</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/purpledoubled"/>
    <language>en</language>
    <item>
      <title>How to Run Qwen-Image 2.1 Locally: One Model That Draws and Edits</title>
      <dc:creator>David </dc:creator>
      <pubDate>Tue, 22 Sep 2026 10:13:53 +0000</pubDate>
      <link>https://dev.to/purpledoubled/how-to-run-qwen-image-21-locally-one-model-that-draws-and-edits-24bg</link>
      <guid>https://dev.to/purpledoubled/how-to-run-qwen-image-21-locally-one-model-that-draws-and-edits-24bg</guid>
      <description>&lt;p&gt;Alibaba Qwen published Qwen-Image 2.1 on 20 September 2026. The interesting part for anyone running models on their own hardware is not the benchmark table, it is the file count: one set of weights now does text to image &lt;em&gt;and&lt;/em&gt; image editing. There is no separate edit model in this generation.&lt;/p&gt;

&lt;p&gt;The less interesting part, which you should read before you spend 16 GB of bandwidth, is the license. More on that at the end, and it is not a footnote.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is actually in the box
&lt;/h2&gt;

&lt;p&gt;The architecture moved, and it moved downwards:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Image model: 7B parameters, single stream DiT, 32 layers&lt;/li&gt;
&lt;li&gt;Text encoder: Qwen3-VL-8B&lt;/li&gt;
&lt;li&gt;VAE: new, 64 channels, RGBA&lt;/li&gt;
&lt;li&gt;Objective: flow matching&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For comparison, the original Qwen-Image from August 2025 (the one usually written as 2508) had a 20B image model and a Qwen2.5-VL encoder. So 2.1 is not a bigger sibling of the old model, it is a different and smaller one. Anything you tuned against the 20B shape has nothing to attach to.&lt;/p&gt;

&lt;p&gt;New in 2.1 against everything before it: a native alpha channel, up to ten reference images in one pass, local editing driven by a drawn marking rather than only a painted mask, a prefix KV cache to speed up repeated edits, and better typography. Native 2K carried over from 2.0 in February 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory
&lt;/h2&gt;

&lt;p&gt;I have not benchmarked this model, so these are quoted numbers, not mine. ai.rs puts the int8 pipeline at roughly 16 GiB in total: about 6.8 GiB for the image model, 8.7 GiB for the text encoder and 0.6 GiB for the VAE. A 24 GB card keeps all three parts resident. On a 16 GB card the usual approach is to push the text encoder into system RAM and let the card carry the rest. The full bf16 weights are a different machine entirely at 33 to 40 GB.&lt;/p&gt;

&lt;p&gt;ai.rs also reports 5.93 seconds on an RTX 5090. Their hardware, their measurement, quoted as such.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three files
&lt;/h2&gt;

&lt;p&gt;The int8 repack from Comfy-Org is three files that go into three different ComfyUI folders:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ComfyUI/models/diffusion_models/qwen_image_2.1_int8_convrot.safetensors
ComfyUI/models/text_encoders/qwen3vl_8b_int8_convrot.safetensors
ComfyUI/models/vae/qwen_image_2.1_vae_bf16.safetensors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sizes, as Hugging Face lists them and as they measure on disk: 7.26 GB (6.76 GiB), 9.35 GB (8.71 GiB), 676 MB (0.63 GiB). Full bf16 versions of the first two exist at 14.2 GB and 17.5 GB, and there is a w4a8 encoder variant at 6.31 GB if you want to trade quality for memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The version gate you will hit first
&lt;/h2&gt;

&lt;p&gt;ComfyUI added support in 0.37.0, merged on 19 September 2026, and it brought a new encode node:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TextEncodeQwenImage21
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The older Qwen-Image edit nodes do not fit this model. If you are assembling a graph by hand and it fails at the encode step, check your ComfyUI version before you check anything else.&lt;/p&gt;

&lt;p&gt;If you would rather not assemble the graph at all, &lt;a href="https://github.com/PurpleDoubleD/locally-uncensored" rel="noopener noreferrer"&gt;Locally Uncensored&lt;/a&gt; 3.0.1 (AGPL-3.0, Windows and Linux) ships it as a Model Manager bundle called "Qwen-Image 2.1 (Generate and Edit)": 16.1 GB, three files, one click, into the folders above. It refuses politely on an older backend with exactly this line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Qwen-Image 2.1 needs ComfyUI 0.37.0 or newer. Update ComfyUI in Settings.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Which is a small thing, but a better small thing than a graph that dies two nodes in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Values that work
&lt;/h2&gt;

&lt;p&gt;Straight from the official template:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Steps: 25&lt;/li&gt;
&lt;li&gt;CFG: 1&lt;/li&gt;
&lt;li&gt;Sampler: euler&lt;/li&gt;
&lt;li&gt;Scheduler: simple&lt;/li&gt;
&lt;li&gt;Resolution: 1024x1024&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;CFG below 1 is not usable with this model. If you are used to dialling CFG down on other models to soften things, that instinct will only produce noise here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generating and editing
&lt;/h2&gt;

&lt;p&gt;Generation is what you expect: prompt in, negative prompt optional, picture out.&lt;/p&gt;

&lt;p&gt;Editing is the part worth adjusting to. You give it &lt;strong&gt;one&lt;/strong&gt; reference image and a prompt, and there is no mask. You describe the change in words instead of painting where it should happen. In Locally Uncensored the mask path is closed for this model on purpose, and the output follows the reference image while the resolution follows the canvas you picked. Whether you find that liberating or annoying probably depends on how much of your workflow currently involves a brush.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is not wired up yet
&lt;/h2&gt;

&lt;p&gt;In Locally Uncensored 3.0.1, several of the model's own abilities are not exposed. Stated flat, with nothing said about plans:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Several reference images at once. The model takes ten, the app takes one.&lt;/li&gt;
&lt;li&gt;RGBA output. The VAE has the channel, the app does not write it out.&lt;/li&gt;
&lt;li&gt;Masks with this model.&lt;/li&gt;
&lt;li&gt;A prompt enhancer and the KV cache node.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And across the board, not app specific: there is no Lightning or Turbo LoRA for 2.1 as of 21 September 2026. The old Qwen-Image speed LoRAs were built for the 20B architecture and do not fit. If your habit is to reach for a 4 step LoRA, there is nothing to reach for yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The license
&lt;/h2&gt;

&lt;p&gt;Qwen-Image 2.1 is published under the Qwen Research License. Research and evaluation are covered. Commercial use is not, and needs a separate license from Alibaba Qwen.&lt;/p&gt;

&lt;p&gt;This is a real break with the rest of the series. Qwen-Image 2512 and Qwen-Image-Edit 2511 are Apache 2.0 and they stay Apache 2.0. So if you are generating assets for paid work, or building something you intend to sell, the newest model in the family is not the one to build on as published, and the older pair is.&lt;/p&gt;

&lt;p&gt;Worth saying plainly because the naming does not warn you. It reads like the next version of an Apache 2.0 model, and the terms underneath are not the same terms.&lt;/p&gt;

&lt;h2&gt;
  
  
  Short version
&lt;/h2&gt;

&lt;p&gt;A 7B model that draws and edits with one set of weights, about 16 GiB of int8 pipeline, ComfyUI 0.37.0 or newer, 25 steps at CFG 1, one reference image and no mask when editing, and a license that stops at research and evaluation. If that last clause does not block what you are doing, it is a cheap thing to try. If it does, the Apache 2.0 entries in the same series are still there.&lt;/p&gt;

&lt;p&gt;The longer walkthrough with screenshots of each step lives on the &lt;a href="https://locallyuncensored.com/blog/how-to-run-qwen-image-2-1-locally.html" rel="noopener noreferrer"&gt;Locally Uncensored blog&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>tutorial</category>
      <category>comfyui</category>
    </item>
    <item>
      <title>DeepSeek V4.1 Flash is out: 552B, MIT, and nothing can run it yet</title>
      <dc:creator>David </dc:creator>
      <pubDate>Thu, 10 Sep 2026 19:23:12 +0000</pubDate>
      <link>https://dev.to/purpledoubled/deepseek-v41-flash-is-out-552b-mit-and-nothing-can-run-it-yet-3pgj</link>
      <guid>https://dev.to/purpledoubled/deepseek-v41-flash-is-out-552b-mit-and-nothing-can-run-it-yet-3pgj</guid>
      <description>&lt;p&gt;DeepSeek published &lt;strong&gt;DeepSeek-V4.1-Flash&lt;/strong&gt; on Hugging Face on &lt;strong&gt;10 September 2026 at 02:17 UTC&lt;/strong&gt;, MIT licensed and ungated. I spent the day reading the files and poking the model through an OpenAI compatible endpoint instead of reading the announcement. Here is what is worth knowing before you start a download measured in hundreds of gigabytes.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;552B backbone&lt;/strong&gt;, 8B active in prefill and 16B in decode, 1,048,576 token context, reads images.&lt;/li&gt;
&lt;li&gt;Original checkpoint: &lt;strong&gt;510.3 GB&lt;/strong&gt; across 48 shards, FP8 with FP4 routed experts.&lt;/li&gt;
&lt;li&gt;Smallest community quant on day one: &lt;strong&gt;168.9 GB&lt;/strong&gt;. There is no four bit build yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No released runtime can execute any of it.&lt;/strong&gt; llama.cpp support is pull request 28696, opened the same day, still a draft, and it is a &lt;em&gt;convert&lt;/em&gt; PR.&lt;/li&gt;
&lt;li&gt;The KV cache is &lt;strong&gt;890 bytes per token&lt;/strong&gt;. A full million token context is about 930 MB. That is the actual release.&lt;/li&gt;
&lt;li&gt;The reasoning effort is a &lt;strong&gt;1 to 100 dial&lt;/strong&gt;, not a switch, and telling it to stop makes the turn longer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The number that fools people
&lt;/h2&gt;

&lt;p&gt;Eight billion active parameters in prefill, sixteen in decode. Both figures are real and both describe compute per token, not memory. The router picks six of 384 routed experts per token and picks differently for the next one, so all 552B have to stay reachable.&lt;/p&gt;

&lt;p&gt;There is a second store on top: DeepSeek calls it &lt;strong&gt;Engram&lt;/strong&gt;, a conditional memory of 196B parameters looked up sparsely per token. It is in the checkpoint and it is not small.&lt;/p&gt;

&lt;p&gt;Active parameters decide how fast it runs. Resident weights decide whether it starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the quants actually weigh
&lt;/h2&gt;

&lt;p&gt;Measured from the Hub on the evening of the release, summed per build:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Build&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Original FP8 with FP4 experts&lt;/td&gt;
&lt;td&gt;510.3 GB&lt;/td&gt;
&lt;td&gt;48 shards, DeepSeek's own&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MLX 2 bit with MTP&lt;/td&gt;
&lt;td&gt;238.8 GB&lt;/td&gt;
&lt;td&gt;Apple silicon&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GGUF Q2_K&lt;/td&gt;
&lt;td&gt;191.8 GB&lt;/td&gt;
&lt;td&gt;7 shards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GGUF mixed IQ2_XXS + Q2_K&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;168.9 GB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;smallest published&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The smallest one averages 2.25 bits per weight across the routed experts, and its author says plainly that it was made without an importance matrix, on a constant unit importance vector, as a memory bounded choice rather than a calibrated one. Two bit without calibration is not a free lunch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The wall
&lt;/h2&gt;

&lt;p&gt;You can download 168.9 GB today. You cannot run it.&lt;/p&gt;

&lt;p&gt;Support lives in &lt;strong&gt;llama.cpp PR 28696&lt;/strong&gt;, &lt;code&gt;convert : add DeepSeek V4.1 (DeepseekV41ForCausalLM)&lt;/code&gt;, opened 10 September 2026 at 10:23 UTC, about eight hours after the weights, and still marked draft. Read the title: it teaches the converter to write the GGUF. Writing the file and executing it are separate pieces of work.&lt;/p&gt;

&lt;p&gt;The person who published the mixed two bit build says the same thing on the model card, and it is the most useful sentence written about this model so far: it is a weights artifact, and no runtime has yet been shown to execute it end to end.&lt;/p&gt;

&lt;p&gt;This is normal for day one. GLM-5.3-Flash sat behind an unmerged PR for the same reason two weeks ago. It resolves in days or weeks. It has not resolved yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  One trap if you convert it yourself
&lt;/h2&gt;

&lt;p&gt;DeepSeek V4.1 stores FP8 block scales in &lt;strong&gt;32 by 32&lt;/strong&gt; blocks. DeepSeek V4 used &lt;strong&gt;128 by 128&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Carry the old value across and every weight in the model is silently rescaled. No error, no warning. You get a file that loads, runs and produces confident nonsense, which is the worst failure mode there is because it looks like a bad quant rather than a bad conversion.&lt;/p&gt;

&lt;p&gt;If you see token soup out of a self converted build, check the block size before you blame anything else.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that is actually clever
&lt;/h2&gt;

&lt;p&gt;Long context has a cost most tables hide. Every token leaves a key and value entry behind that stays in memory for the life of the conversation. On an agent that reads a repo, runs a command, reads the output and repeats for two hundred steps, that cache is the bill.&lt;/p&gt;

&lt;p&gt;Three things bring it down here:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A causal encoder decoder.&lt;/strong&gt; 40 layers split into a 20 layer causal encoder and a 20 layer decoder. The decoder does not derive its global KV from each of its own layers, it projects it from the final encoder hidden states. One projection instead of twenty caches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CSA2.&lt;/strong&gt; Each attention layer gets one of three static modes, Full, Reindex or Reuse, so layers share main KV and indexer keys. A hierarchical sparse indexer in the decoder restricts later indexing layers to a candidate pool built by the first Full layer, which stops indexing cost growing with context length.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FP4 for the cache itself&lt;/strong&gt;, E2M1 with one E4M3 scale per sixteen channels. Four bit weights are common. Four bit cache is not.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Result: &lt;strong&gt;890 bytes per token&lt;/strong&gt;, which DeepSeek puts at about a quarter of V4 Flash and one four hundred and thirty seventh of V1. A million token window costs roughly 930 MB of cache.&lt;/p&gt;

&lt;p&gt;Their published benchmark table has exactly the shape you would predict. On agentic tests it takes the top row repeatedly against models with far more active compute: Terminal-Bench 2.1 at 90.6, DeepSWE v1.1 at 74.2, AutomationBench at 54.8, Codeforces rating 3471. On single shot reasoning it does not, and it lands slightly below its own V4 Pro sibling on GPQA Diamond. Those are vendor numbers, read them as such, but the pattern is consistent enough to be believable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the reasoning dial does when you tell it to stop
&lt;/h2&gt;

&lt;p&gt;The model card describes a &lt;strong&gt;continuously controllable reasoning effort from 1 to 100&lt;/strong&gt;. Every instruct benchmark was run at 100.&lt;/p&gt;

&lt;p&gt;That collides with the OpenAI compatible convention, where &lt;code&gt;reasoning_effort&lt;/code&gt; is a coarse setting with a &lt;code&gt;none&lt;/code&gt; position. I ran the same question three times each way through DeepInfra, 700 max tokens, nothing else changed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;No signal&lt;/th&gt;
&lt;th&gt;&lt;code&gt;reasoning_effort: "none"&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;42 out, 123 chars reasoning&lt;/td&gt;
&lt;td&gt;82 out, 0 chars&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;42 out, 118 chars reasoning&lt;/td&gt;
&lt;td&gt;53 out, 0 chars&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;42 out, 123 chars reasoning&lt;/td&gt;
&lt;td&gt;51 out, 0 chars&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The off signal did not stop the thinking. It stopped the thinking from being &lt;em&gt;labelled&lt;/em&gt;, and the monologue moved into &lt;code&gt;content&lt;/code&gt;, unmarked. In two of three runs no actual answer fit in the budget.&lt;/p&gt;

&lt;p&gt;If you are building on this model: treat it as always reasoning, and do not offer your users a switch that makes their replies longer and dearer.&lt;/p&gt;

&lt;p&gt;One more thing worth knowing. &lt;code&gt;reasoning_effort: "max"&lt;/code&gt; is accepted, but acceptance proves nothing: an invalid value comes back 422 with the whole enum, which is provider level validation rather than a model capability. Measured on a hard prompt, &lt;code&gt;max&lt;/code&gt; produced 308 characters of reasoning against 337 for &lt;code&gt;high&lt;/code&gt;. There is no higher rung to reach for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I checked because tags lie
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Vision.&lt;/strong&gt; The card says multimodal, so I sent three 8x8 PNGs of different colours with one question each. Red, blue, green, correct every time. The older V4 Flash 0731 returns an error on image input, so this is genuinely new in the line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tools.&lt;/strong&gt; A single &lt;code&gt;get_weather&lt;/code&gt; function definition came back as a real &lt;code&gt;tool_calls&lt;/code&gt; entry with &lt;code&gt;{"city": "Berlin"}&lt;/code&gt;. Native, no prompt translation needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to run locally instead this week
&lt;/h2&gt;

&lt;p&gt;If you want a capable model on hardware you own today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Qwen 3.8 27B&lt;/strong&gt;, 10.7 GB to 54.7 GB depending on the quant, Apache 2.0. The one that fits a single card.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ling 3.0 Flash&lt;/strong&gt;, 27 GB at the smallest build, MIT, 5.1B active.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5.3 Flash&lt;/strong&gt;, 93.1 GB at the smallest build, MIT, reads images.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And if you want V4.1 Flash itself before the runtime lands, it is a hosted model for now, the same way Qwen 3.8 was until its open checkpoint hit HuggingFace on 12 August.&lt;/p&gt;

&lt;p&gt;I maintain &lt;a href="https://locallyuncensored.com/" rel="noopener noreferrer"&gt;Locally Uncensored&lt;/a&gt;, a desktop app for running these models on your own machine, and the hosted catalogue at &lt;a href="https://lu-labs.ai/deepseek-v4-1-flash-online" rel="noopener noreferrer"&gt;lu-labs.ai&lt;/a&gt; picked this one up on release day. The measurements above were taken while deciding how to label it in our own model picker, which is why they exist at all.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>How to run Qwen 3.8 without a GPU (and which Qwen 3.8 you can actually download)</title>
      <dc:creator>David </dc:creator>
      <pubDate>Mon, 07 Sep 2026 21:13:29 +0000</pubDate>
      <link>https://dev.to/purpledoubled/how-to-run-qwen-38-without-a-gpu-and-which-qwen-38-you-can-actually-download-240a</link>
      <guid>https://dev.to/purpledoubled/how-to-run-qwen-38-without-a-gpu-and-which-qwen-38-you-can-actually-download-240a</guid>
      <description>&lt;p&gt;Disclosure, since half of this post is about a service we run: &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;LU Labs&lt;/a&gt; is the hosted layer next to our free open source desktop app for Windows and Linux. The other half is about running Qwen 3.8 on your own hardware, which costs us money rather than making it, and I have written that part first because for a lot of readers it is the right answer.&lt;/p&gt;

&lt;p&gt;"How do I run Qwen 3.8 without a GPU" is really two questions wearing one coat, and they have opposite answers depending on which Qwen 3.8 you mean.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, work out which model you are talking about
&lt;/h2&gt;

&lt;p&gt;Alibaba shipped several things under the same version number in August 2026, and the licences are not the same either.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen 3.8 27B&lt;/strong&gt; is a dense 27 billion parameter model, Apache 2.0, and it runs on a single graphics card or a decent laptop. This is the one you can download.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen 3.8 Max&lt;/strong&gt; and &lt;strong&gt;Qwen 3.8 A95B&lt;/strong&gt; are the flagships. A95B is 2.4 trillion parameters in total with 95 billion active per token, and all 2.4 trillion have to be resident while it works. That is a rack, not a desktop. The weights for A95B are public, but under Alibaba's own qwen3.8-max terms rather than Apache, so if you are checking licences for a commercial build, read them rather than assuming.&lt;/p&gt;

&lt;p&gt;For the flagships, the choice is not local against hosted. It is hosted against not using the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you meant the 27B: run it yourself
&lt;/h2&gt;

&lt;p&gt;Genuinely, run it. At Q4_K_M the file is about 17 GB, which loads whole onto a 24 GB card with room left for context, spills into system RAM on a 16 GB card (slower, still works), and is comfortable on a 32 GB Apple Silicon Mac. IQ4_XS at about 15.7 GB is the largest quant that stays fully on a 16 GB card.&lt;/p&gt;

&lt;p&gt;The one thing that trips almost everyone: Qwen 3.8 ships a chat template, and the template is not decoration. It marks where your message ends and the answer begins, and it carries the reasoning switch. Load a bare GGUF without it and you get one of two failure modes that both look like a bad quantization: it rambles and never stops, or it answers in a clipped voice and forgets the conversation between turns. In llama.cpp the fix is the &lt;code&gt;--jinja&lt;/code&gt; flag. LM Studio and Ollama usually bring the template along with the model, so this mostly bites people who pulled a raw file by hand.&lt;/p&gt;

&lt;p&gt;The longer version, with the full quant table and the vision file, is in &lt;a href="https://lu-labs.ai/blog/how-to-run-qwen-3-8-27b-locally" rel="noopener noreferrer"&gt;how to run Qwen 3.8 27B on your own computer&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Everything below is about the flagships.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reaching Max and A95B in a browser or from code
&lt;/h2&gt;

&lt;p&gt;Both are in our chat catalog. Worth stating clearly, because our own older posts said otherwise: the catalog is now the same on every plan and on every credit pack. There is no shortlist, no model reserved for an expensive tier, and a €5 pack bought with no subscription reaches both flagships exactly like a €99 month does. A bigger plan buys credits, not access.&lt;/p&gt;

&lt;p&gt;In the browser: sign in at &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;lu-labs.ai&lt;/a&gt;, open the chat tab, and both appear under their catalog names, Qwen 3.8 Max and Qwen 3.8 A95B. Nothing downloads and there is no driver to install.&lt;/p&gt;

&lt;p&gt;From code: Settings, then Cloud API keys, create one. It starts with &lt;code&gt;lu_&lt;/code&gt; and is shown once, since we keep a hash rather than the key. Base URL &lt;code&gt;https://lu-labs.ai/api/inference/v1&lt;/code&gt;, OpenAI chat completions shape, so Aider, LibreChat, your own script or plain curl all work unchanged.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://lu-labs.ai/api/inference/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$LU_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"Qwen/Qwen3.8-Max",
       "messages":[{"role":"user","content":"Hello"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://lu-labs.ai/api/inference/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LU_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen/Qwen3.8-2.4T-A95B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both ids are exact catalog strings. Both take tool definitions in the standard &lt;code&gt;tools&lt;/code&gt; parameter and call them natively, so agent scaffolding needs no adapter. Neither reads images: attach a screenshot and the request comes back rejected, which is a property of these two ids rather than a setting you missed. Pick a vision model in the same list for that turn.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a month of this costs
&lt;/h2&gt;

&lt;p&gt;Credits convert straight from the catalog price, and one credit is $0.00001.&lt;/p&gt;

&lt;p&gt;On Qwen 3.8 Max, one million output tokens draws 495,100 credits. On A95B it is 600,000, because A95B is the dearer of the two on both input and output.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you buy&lt;/th&gt;
&lt;th&gt;Credits&lt;/th&gt;
&lt;th&gt;Output tokens on Max&lt;/th&gt;
&lt;th&gt;On A95B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;€5 pack, no subscription&lt;/td&gt;
&lt;td&gt;165,000&lt;/td&gt;
&lt;td&gt;about 333,000&lt;/td&gt;
&lt;td&gt;about 275,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosted, €19 a month&lt;/td&gt;
&lt;td&gt;900,000&lt;/td&gt;
&lt;td&gt;about 1.8 million&lt;/td&gt;
&lt;td&gt;about 1.5 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro, €49 a month&lt;/td&gt;
&lt;td&gt;2,350,000&lt;/td&gt;
&lt;td&gt;about 4.7 million&lt;/td&gt;
&lt;td&gt;about 3.9 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max, €99 a month&lt;/td&gt;
&lt;td&gt;5,000,000&lt;/td&gt;
&lt;td&gt;about 10.1 million&lt;/td&gt;
&lt;td&gt;about 8.3 million&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All of those are ceilings nobody reaches. Input is billed from the same pool at 0.165 credits per token on Max and 0.2 on A95B, a long conversation resends its whole history every turn, and images and video draw on the same balance. Treat the table as the shape of a budget, not a forecast.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Think button is the difference between the two
&lt;/h2&gt;

&lt;p&gt;On Max, reasoning is genuinely optional, and switching it off is the single largest lever on what a session costs. Leave it on for work that has to be derived. Turn it off for extraction, reformatting, translation and summarising, where you already know the shape of the answer and are only paying for the model to talk itself into it.&lt;/p&gt;

&lt;p&gt;On A95B there is no such lever. It reasons every turn by design. There is no parameter that stops it, so plan for a reasoning pass in every single answer and choose the model deliberately rather than by habit.&lt;/p&gt;

&lt;p&gt;One more thing about the effort control, since it looks like the same knob and is not: on these two, effort rungs are not what moves the reasoning. On Max it is the Think button. On A95B nothing moves it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Picking between them
&lt;/h2&gt;

&lt;p&gt;Max is cheaper per token on both sides and it lets you switch the reasoning off, which makes it the default of the two for most work. A95B is the open weight 2.4T checkpoint with 95 billion active per token and a 262,144 token context, and it always reasons, which is either the reason you want it or the reason you do not.&lt;/p&gt;

&lt;p&gt;The honest summary: run the 27B at home when the material should not leave your machine, when you want it offline, and when you would rather pay once for a card than per token forever. Reach for a flagship when the problem is genuinely bigger than the hardware in the room. Most people who use both end up doing exactly that.&lt;/p&gt;

&lt;p&gt;The full guide with the plan details is at &lt;a href="https://lu-labs.ai/blog/how-to-run-qwen-3-8-without-a-gpu" rel="noopener noreferrer"&gt;how to run Qwen 3.8 without a GPU&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Uncensored AI chat online: text without a moderation layer, and the honest limits</title>
      <dc:creator>David </dc:creator>
      <pubDate>Mon, 07 Sep 2026 17:43:50 +0000</pubDate>
      <link>https://dev.to/purpledoubled/uncensored-ai-chat-online-text-without-a-moderation-layer-and-the-honest-limits-51ip</link>
      <guid>https://dev.to/purpledoubled/uncensored-ai-chat-online-text-without-a-moderation-layer-and-the-honest-limits-51ip</guid>
      <description>&lt;p&gt;Ours, said first: we make &lt;a href="https://locallyuncensored.com" rel="noopener noreferrer"&gt;Locally Uncensored&lt;/a&gt;, a free open source desktop app for Windows and Linux that runs open models on your own machine, and &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;LU Labs&lt;/a&gt; is the hosted layer next to it. So this is a guide written by an interested party, and the useful thing we can do about that is be precise instead of enthusiastic.&lt;/p&gt;

&lt;p&gt;The phrase "uncensored AI chat" gets used to mean several different things, most of them oversold. This post says exactly what we do and do not do, because the gap between the marketing version and the real one is where people waste money.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the word means here, for text
&lt;/h2&gt;

&lt;p&gt;Hosted text models answer as they were trained. We do not put a moderation layer in front of them and we do not rewrite what comes back. That is the whole claim. It is smaller than the usual pitch and it is the part that actually changes your experience.&lt;/p&gt;

&lt;p&gt;What it means in practice: no refusal injected by us on top of the model, no silent prompt rewriting, no filtered response substituted for the one the model produced. What the model does is the model's business, and models differ. A finetune tuned for open roleplay behaves differently from a general assistant, and a general assistant will still decline things on its own because that behaviour is baked into its weights. Nobody can sell you a model that never refuses. They can only stop adding refusals of their own, which is what we do.&lt;/p&gt;

&lt;p&gt;The models people come for on the text side are Euryale 70B, Lunaris 8B and MythoMax 13B. They sit in the same catalog as Kimi K3, GLM 5.3, Qwen 3.8 Max and DeepSeek V4 Pro 0813, and that catalog is now the same everywhere: every plan carries all of it, and so does a credit pack bought with no subscription at all. Nothing is held back for a higher tier. What a plan buys is credits, not access.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we do not claim, for images and video
&lt;/h2&gt;

&lt;p&gt;This is the section most articles on this topic skip, so here it is plainly.&lt;/p&gt;

&lt;p&gt;Hosted image and video generation follows our Terms of Service, and prompts are checked before anything is rendered. That check is real and it is not the same regime as text. If you are choosing a service on the assumption that hosted image work here is open the way the text side is, that assumption is wrong, and we would rather you know now than after a purchase.&lt;/p&gt;

&lt;p&gt;We do not promise nudity or adult imagery, on any plan, at any price. There is no tier that changes this.&lt;/p&gt;

&lt;p&gt;And one rule that has no exception anywhere in the product, hosted or local, text or image: content involving minors is refused. Always, everywhere, no argument, no context that makes it acceptable. If that is what you came for, leave.&lt;/p&gt;

&lt;h2&gt;
  
  
  The route that has no layer at all
&lt;/h2&gt;

&lt;p&gt;If what you actually want is a model with nothing between you and it, run it yourself. That is the free desktop app, on Windows or Linux, with the weights on your disk and no account. No queue, no meter, no policy document, no network. It is free and stays free, and we built it first for exactly this reason.&lt;/p&gt;

&lt;p&gt;The honest tradeoff is hardware. A 70B class model does not fit on a normal card, and the models people want most for this are large. So the split most users land on is local for anything private or offline, hosted for the sizes the machine refuses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Using the hosted side
&lt;/h2&gt;

&lt;p&gt;Browser: sign in at &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;lu-labs.ai&lt;/a&gt; and open the Studio. Nothing to install, which is also the only route on a Mac, where we do not ship a desktop build.&lt;/p&gt;

&lt;p&gt;From your own client: mint a key under Settings, then Cloud API keys. It starts with &lt;code&gt;lu_&lt;/code&gt; and it is shown once, because we store a hash rather than the key itself. The endpoint takes the OpenAI chat completions shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://lu-labs.ai/api/inference/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer YOUR_KEY"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"zai-org/GLM-5.3-Flash","messages":[{"role":"user","content":"Hello"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That makes SillyTavern the obvious client for this audience, and it works: custom OpenAI-compatible source, base URL &lt;code&gt;https://lu-labs.ai/api/inference/v1&lt;/code&gt;, key in the key field, model id typed in verbatim. Aider, LibreChat and anything else with a base URL field work the same way.&lt;/p&gt;

&lt;p&gt;One implementation note that matters if you use agents. Most of the roleplay finetunes do not ship a tool calling template, which normally means no function calling. We transport the tool definitions through the prompt instead and parse the calls back out, so agent workflows run on Euryale and Lunaris like they do on anything else. Thinking is per model rather than a global switch: Lunaris 8B has no thinking mode at all, Kimi K3 has a toggle, and GLM 5.3 Flash reasons on every turn with an effort setting instead of an on and off.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Credits, one pool, one credit is $0.00001, and every model bills at its own rate rather than a blended average.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hosted&lt;/strong&gt;, €19 a month, 900,000 credits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro&lt;/strong&gt;, €49 a month, 2,350,000 credits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Max&lt;/strong&gt;, €99 a month, 5,000,000 credits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Packs&lt;/strong&gt;, no subscription needed at all. €5 for 165,000 credits, €10 for 350,000, €25 for 900,000. Credits never expire, and you can buy any pack on your first purchase.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To convert that into something real. Euryale 70B runs at 0.085 credits per output token, so a €5 pack is roughly 1.9 million output tokens on it. GLM 5.3 Flash is 0.05, which is about 3.3 million on the same pack. Kimi K3 is 1.425, which is about 115,000. For long chat sessions on the finetunes that is a lot of reading, and the same pack disappears in an afternoon if you spend it on the biggest reasoner in the list. Rates per model are on the &lt;a href="https://lu-labs.ai/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rest of the honest list
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It is not private the way local is.&lt;/strong&gt; A hosted prompt reaches a GPU that is not yours. Accounts and billing sit in the EU, we do not train on user data and we never sell it. That is a commitment, not a guarantee of physics. The guarantee of physics is the desktop app, offline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The hardware is shared.&lt;/strong&gt; Managed H100, A100 and B200 class GPUs, and at busy hours there is a short queue.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The catalog is curated.&lt;/strong&gt; No bring-your-own checkpoint. If your setup depends on one specific community merge, run it locally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is metered.&lt;/strong&gt; Credits, not unlimited. A pool tells you the truth about what a model costs; an unlimited label does not.&lt;/p&gt;

&lt;p&gt;If the text side is what you were looking for, &lt;a href="https://lu-labs.ai/checkout/start?plan=hosted&amp;amp;src=devto-text-models" rel="noopener noreferrer"&gt;start with a Hosted month or a €5 pack&lt;/a&gt;. The longer version of this guide, with the policy wording in full, is at &lt;a href="https://locallyuncensored.com/blog/uncensored-ai-chat-online-guide.html" rel="noopener noreferrer"&gt;the uncensored AI chat guide&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>An LM Studio alternative for Mac that needs nothing installed</title>
      <dc:creator>David </dc:creator>
      <pubDate>Mon, 07 Sep 2026 17:43:29 +0000</pubDate>
      <link>https://dev.to/purpledoubled/an-lm-studio-alternative-for-mac-that-needs-nothing-installed-451m</link>
      <guid>https://dev.to/purpledoubled/an-lm-studio-alternative-for-mac-that-needs-nothing-installed-451m</guid>
      <description>&lt;p&gt;This is a guide to our own product, so read it with that in mind. &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;LU Labs&lt;/a&gt; is the hosted layer we run next to our open source desktop app for Windows and Linux, and on a Mac the hosted browser Studio is the only route we offer, because we do not ship a Mac desktop build.&lt;/p&gt;

&lt;p&gt;LM Studio is good software. If it works on your Mac, use it. This post is for the Macs where it does not: the low memory machines, the work laptops with locked installs, and the situation where the model you want is simply larger than the memory you have. That last one is the common case and the one people are slowest to accept.&lt;/p&gt;

&lt;h2&gt;
  
  
  The memory wall, in plain terms
&lt;/h2&gt;

&lt;p&gt;Apple Silicon has one memory pool shared between the CPU and the GPU, which is the reason a Mac punches above its price on local inference. It is also the reason a Mac hits a hard stop.&lt;/p&gt;

&lt;p&gt;Take Qwen 3.8 27B, the dense open weight model most people try first. At 4-bit the weights alone are about 17 GB that has to sit in memory before the first token, plus the KV cache for your context, plus macOS, plus whatever your browser is holding. On a 32 GB Mac that is comfortable. On 24 GB it works if the machine is otherwise idle. Below that it does not load, and no download manager or quant picker changes that. LM Studio will show you the model, let you pull it, and then the machine will tell you no.&lt;/p&gt;

&lt;p&gt;And that is the small end of what people ask for. Kimi K3 is 2.8 trillion parameters and Qwen 3.8 A95B is 2.4 trillion. Those do not run on any Mac that exists, at any quantization, and pretending otherwise wastes an afternoon and 40 GB of disk.&lt;/p&gt;

&lt;p&gt;The honest options at that point are a bigger Mac, a much smaller model, or a machine that already has the card.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the hosted route looks like on a Mac
&lt;/h2&gt;

&lt;p&gt;Open &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;lu-labs.ai&lt;/a&gt; in Safari or Chrome, sign in, use the Studio. That is the whole install process. No Homebrew, no Xcode command line tools, no Metal build, no &lt;code&gt;.dmg&lt;/code&gt; you have to right-click open because it is not notarized, no multi gigabyte download sitting on a laptop SSD you were already fighting for space on.&lt;/p&gt;

&lt;p&gt;Every plan reaches the same chat catalogue, and so does a credit pack bought without any subscription. There is no shortlist and no model held back for a higher price: Kimi K3, Qwen 3.8 Max, Qwen 3.8 A95B, GLM 5.3 and GLM 5.3 Flash, DeepSeek V4 Pro 0813 and DeepSeek V4 Flash 0731, DeepSeek V3.2, Gemma 4 26B and gpt-oss 120B all sit on the same list. What a plan buys is credits, not access.&lt;/p&gt;

&lt;p&gt;Two things a local Mac setup usually cannot give you at all come with the same account: 10 image models and 5 video models. Flux 2 Dev, Flux Dev, Flux Schnell, Qwen Image, HiDream, HunyuanImage 2.1, Z-Image Turbo, Chroma, Prefect Pony XL and Neta Lumina on the image side, with inpainting on Flux Dev, Qwen Image Edit, background removal, an eraser and upscale to 2k, 4k or 8k. Video is Wan 2.2 720p, Wan 2.2 Fast, LTX-2 with audio, LTX 2.3 and HunyuanVideo 1.5, at 5 or 8 seconds a clip.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping your existing tools
&lt;/h2&gt;

&lt;p&gt;The part that matters if you already have a workflow: this is an OpenAI-compatible endpoint, so the Mac apps you point at LM Studio's local server can be pointed here instead by changing a base URL.&lt;/p&gt;

&lt;p&gt;Mint a key in Settings, then Cloud API keys. It starts with &lt;code&gt;lu_&lt;/code&gt; and is shown once, since we store a hash and not the key.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://lu-labs.ai/api/inference/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer YOUR_KEY"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"moonshotai/Kimi-K3","messages":[{"role":"user","content":"Hello"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or from Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://lu-labs.ai/api/inference/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lu_...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;zai-org/GLM-5.3-Flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain this stack trace.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Model ids are the exact catalog strings, &lt;code&gt;moonshotai/Kimi-K3&lt;/code&gt; and &lt;code&gt;zai-org/GLM-5.3-Flash&lt;/code&gt;, not the friendly names from the picker. SillyTavern, Aider, LibreChat and anything else that takes a base URL work the same way. If your Mac tool has a field marked "OpenAI compatible", it is already compatible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Credits from one pool, one credit is $0.00001, and each model bills at its own rate rather than a blended average.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hosted&lt;/strong&gt;, €19 a month, 900,000 credits, plus LoRA training 2 a month.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro&lt;/strong&gt;, €49 a month, 2,350,000 credits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Max&lt;/strong&gt;, €99 a month, 5,000,000 credits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Packs&lt;/strong&gt;, no subscription. €5 for 165,000 credits, €10 for 350,000, €25 for 900,000. Credits never expire and you can start on any pack.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anchors, since credits are abstract until you convert them. GLM 5.3 Flash bills 0.05 credits per output token, so a €5 pack is roughly 3.3 million output tokens there. Kimi K3 bills 1.425 credits per output token, so the same pack is about 115,000 output tokens. That spread of nearly 30x is the reason the catalogue is open everywhere: the meter decides what a model costs you, not the plan you are on.&lt;/p&gt;

&lt;p&gt;On the image side, Flux Schnell is 300 credits an image and Flux 2 Dev is 1,200, which makes the €5 pack about 550 quick drafts or 137 Flux 2 Dev renders, and the Hosted month about 3,000 or 750. The &lt;a href="https://lu-labs.ai/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt; carries the full rate card.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;LM Studio on your Mac&lt;/th&gt;
&lt;th&gt;Hosted in the browser&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Install&lt;/td&gt;
&lt;td&gt;Yes, plus the model download&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;€19 a month, or a €5 pack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Works offline&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt stays on the machine&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model size ceiling&lt;/td&gt;
&lt;td&gt;Your unified memory&lt;/td&gt;
&lt;td&gt;The catalog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Images and video&lt;/td&gt;
&lt;td&gt;Separate tooling&lt;/td&gt;
&lt;td&gt;Included&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom checkpoints&lt;/td&gt;
&lt;td&gt;Anything you can download&lt;/td&gt;
&lt;td&gt;Curated catalog only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read that table honestly and it tells you which one you want. If your Mac has the memory for the models you actually use, LM Studio is free and private and there is no argument for paying us. If it does not, or if you also want image and video work, or if you cannot install software on the machine at all, the browser route is the one that exists.&lt;/p&gt;

&lt;p&gt;The limits on our side, stated rather than buried: the hardware is shared, so at busy hours there is a short queue. It is managed H100, A100 and B200 class GPUs, which is more card than any Mac. Accounts and billing are in the EU, we do not train on user data and we never sell it, but a hosted prompt still reaches a GPU that is not yours, and no privacy policy makes that untrue. The catalog is curated, so there is no bring-your-own checkpoint. And there is no Mac desktop build, which is why this article exists in the first place rather than a download link.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lu-labs.ai/checkout/start?plan=hosted&amp;amp;src=devto-macstudio" rel="noopener noreferrer"&gt;Start with a Hosted month or a €5 pack&lt;/a&gt;, or read the longer version at &lt;a href="https://lu-labs.ai/blog/how-to-use-lu-labs-cloud-on-a-mac" rel="noopener noreferrer"&gt;how to use LU Labs Cloud on a Mac&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>macos</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Flux 2 Dev without a 24 GB card: rendering online from €5</title>
      <dc:creator>David </dc:creator>
      <pubDate>Mon, 07 Sep 2026 17:43:08 +0000</pubDate>
      <link>https://dev.to/purpledoubled/flux-2-dev-without-a-24-gb-card-rendering-online-from-eu5-mfb</link>
      <guid>https://dev.to/purpledoubled/flux-2-dev-without-a-24-gb-card-rendering-online-from-eu5-mfb</guid>
      <description>&lt;p&gt;Disclosure first, because this is a guide to our own thing: &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;LU Labs&lt;/a&gt; is the hosted layer we built next to our open source desktop app for Windows and Linux, which is free and runs models locally. This post is about the hosted image side, what a render costs in credits, and where the hosted route is the wrong choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hardware problem, briefly
&lt;/h2&gt;

&lt;p&gt;Flux 2 Dev is a large model. Running it locally at a quality you would actually ship means a 24 GB card, and the workarounds below that line are the usual ones: heavy quantization, sequential CPU offload, or a tiling scheme that turns a render you would have watched into a render you go and make coffee during. All of them work. None of them are fun on a laptop.&lt;/p&gt;

&lt;p&gt;The alternative is to keep the prompt and send the compute elsewhere. That is what this is.&lt;/p&gt;

&lt;h2&gt;
  
  
  The route
&lt;/h2&gt;

&lt;p&gt;Sign in at &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;lu-labs.ai&lt;/a&gt; and open the Studio in your browser. Nothing to install, no Python environment, no &lt;code&gt;xformers&lt;/code&gt; version to reconcile. Every hosted account gets the same 10 image models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Flux 2 Dev&lt;/strong&gt;, the one in the title&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flux Dev&lt;/strong&gt; and &lt;strong&gt;Flux Schnell&lt;/strong&gt;, the older pair, Schnell being the fast draft model&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Qwen Image&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HiDream&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HunyuanImage 2.1&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Z-Image Turbo&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Chroma&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prefect Pony XL&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Neta Lumina&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The image catalog does not scale with the plan. It is identical on every hosted plan and identical on a pack, and the chat catalog works the same way, so the only thing price changes is how many credits you get.&lt;/p&gt;

&lt;h2&gt;
  
  
  Working, not just generating
&lt;/h2&gt;

&lt;p&gt;A first render is rarely the last one, so the editing tools matter more than the model count:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Inpainting on Flux Dev.&lt;/strong&gt; Mask the part that is wrong, describe what should be there instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen Image Edit.&lt;/strong&gt; Instruction-driven edits on an existing image rather than a fresh sample.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Background removal&lt;/strong&gt; and an &lt;strong&gt;eraser&lt;/strong&gt; for the cleanup passes that otherwise send you back to a desktop editor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upscale&lt;/strong&gt; to 2k, 4k or 8k, which is the step people forget until they try to print something.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The workflow that gets the most out of a balance is boring and effective: draft wide on Flux Schnell until the composition is right, then re-render the keeper on Flux 2 Dev, then inpaint the one detail that is off, then upscale. Cheap exploration, expensive commitment, in that order.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs, in whole numbers
&lt;/h2&gt;

&lt;p&gt;Credits, one pool, one credit is $0.00001. Images are priced per render at a fixed rate per model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Credits per image&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Flux Schnell&lt;/td&gt;
&lt;td&gt;300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flux Dev&lt;/td&gt;
&lt;td&gt;1,200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flux 2 Dev&lt;/td&gt;
&lt;td&gt;1,200&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That converts cleanly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Flux Schnell renders&lt;/th&gt;
&lt;th&gt;Flux 2 Dev renders&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;€5 pack, 165,000 credits&lt;/td&gt;
&lt;td&gt;about 550&lt;/td&gt;
&lt;td&gt;137&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosted month, 900,000 credits&lt;/td&gt;
&lt;td&gt;about 3,000&lt;/td&gt;
&lt;td&gt;750&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The €5 pack is the entry point in the title and it is a real one. No subscription is required for it, the credits never expire, and you can buy any pack on the first purchase: €5 for 165,000 credits, €10 for 350,000, €25 for 900,000. If you render in bursts, a pack is the honest shape, because there is no month to waste.&lt;/p&gt;

&lt;p&gt;The Hosted plan is €19 a month for 900,000 credits, and it adds two things a pack does not have: LoRA training, 2 a month, and a monthly refill you do not have to think about.&lt;/p&gt;

&lt;p&gt;The full rate card is on the &lt;a href="https://lu-labs.ai/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same balance does video
&lt;/h2&gt;

&lt;p&gt;Worth knowing before you decide how much to buy, because video is where a balance disappears fastest. Every hosted account also gets 5 video models: Wan 2.2 720p, Wan 2.2 Fast, LTX-2 with audio, LTX 2.3 and HunyuanVideo 1.5, with clips at 5 or 8 seconds. An LTX-2 clip at 5 seconds is 8,000 credits, a different order of cost from a still image. Plan for that, or keep video and stills on separate purchases.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the hosted route is wrong
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;You want a specific community checkpoint.&lt;/strong&gt; The catalog is curated and there is no bring-your-own checkpoint. If your work depends on one particular merge from a model site, hosted will not do it and local will.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You want the render to never leave your machine.&lt;/strong&gt; Hosted means the prompt reaches a GPU that is not yours. Accounts and billing sit in the EU, we do not train on user data and we never sell it, but that is a policy rather than physics. The free desktop app on Windows or Linux is the version where the guarantee is physical, and that is exactly why we shipped it first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You render constantly.&lt;/strong&gt; If a 24 GB card is already in the machine and it is busy every day, local is cheaper than any credit pool, and it always will be. Hosted wins on the days your hardware says no, on a laptop, or on a Mac where we have no desktop build and the browser Studio is the only route.&lt;/p&gt;

&lt;p&gt;One more honest limit: the hardware is shared. It is managed H100, A100 and B200 class GPUs, which is more card than most people have, but at busy hours there is a short queue. A render that normally comes back quickly sometimes waits.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;If you have the card, run Flux 2 Dev locally. If you do not, a €5 pack gets you 137 Flux 2 Dev renders or about 550 quick drafts, with no subscription and no expiry date on the credits, and the editing and upscale tools come with it rather than as a separate purchase.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lu-labs.ai/checkout/start?plan=hosted&amp;amp;src=devto-flux2" rel="noopener noreferrer"&gt;Start with a Hosted month or a €5 pack&lt;/a&gt; and the longer guide is at &lt;a href="https://lu-labs.ai/blog/how-to-generate-flux-2-images-online" rel="noopener noreferrer"&gt;how to generate Flux 2 images online&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>An OpenAI-compatible endpoint with 38 open models, including the unrestricted ones</title>
      <dc:creator>David </dc:creator>
      <pubDate>Sun, 06 Sep 2026 09:39:18 +0000</pubDate>
      <link>https://dev.to/purpledoubled/an-openai-compatible-endpoint-with-38-open-models-including-the-unrestricted-ones-ik1</link>
      <guid>https://dev.to/purpledoubled/an-openai-compatible-endpoint-with-38-open-models-including-the-unrestricted-ones-ik1</guid>
      <description>&lt;p&gt;We built LU Labs as a desktop app that runs open models on your own machine. That part is free and stays free. But a good share of the people who downloaded it wrote back with the same sentence: my laptop cannot hold a 400B model, can you just run it.&lt;/p&gt;

&lt;p&gt;So we run it. Same catalog, same tools, someone else's GPUs. This post is the developer version of the pricing page: what the endpoint is, what the numbers mean, and where it stops.&lt;/p&gt;

&lt;h2&gt;
  
  
  The endpoint
&lt;/h2&gt;

&lt;p&gt;It speaks the OpenAI chat completions shape. Point any client at it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://lu-labs.ai/api/inference/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lu_...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;zai-org/GLM-5.3-Flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain the DSA attention change in GLM-5.3.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keys are minted in account settings, up to five of them, format &lt;code&gt;lu_&lt;/code&gt; plus 40 hex characters. We store a SHA-256 hash and nothing else, so the plaintext exists in the create response and never again. Worth saying plainly because it is the part people get wrong when they build this themselves: the key authenticates for inference only. It cannot read your account and it cannot mint more keys. A leaked key can spend your plan's credits, which is bad, and that is the whole blast radius, which is the point.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/models&lt;/code&gt; returns what your plan can actually reach, not a catalog you then get 403s from.&lt;/p&gt;

&lt;h2&gt;
  
  
  The catalog
&lt;/h2&gt;

&lt;p&gt;53 models: 38 for chat and code, 10 for image, 5 for video.&lt;/p&gt;

&lt;p&gt;Every plan gets the same 15 chat models. That shortlist is not the cheap leftovers. It includes Kimi K3, Hermes 3 405B, Qwen3 Coder 480B and GLM 5.3 Flash, plus all five of the unrestricted ones. Pro and Max add the remaining 23 for the full 38, among them GLM 5.3, Qwen 3.8 A95B and DeepSeek V4 Flash 0731.&lt;/p&gt;

&lt;p&gt;Two properties that are easy to promise and annoying to deliver, so here is where we stand on both:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tools work on all of them.&lt;/strong&gt; Including the unrestricted models, which mostly do not ship with a tool-calling template. For those we transport the tool definitions through the prompt instead of the API field, which means agent mode and the coding agent run on Hermes and Euryale the same as on anything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thinking is per model, not a global switch.&lt;/strong&gt; Some models always think, some never do, some take a toggle, and GLM 5.3 Flash takes an effort level. The &lt;code&gt;/models&lt;/code&gt; response carries that per entry, so a client can render the right control instead of guessing.&lt;/p&gt;

&lt;p&gt;Image, video and every Create tool are identical on every hosted plan. The chat catalog is the only thing that scales with price.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a euro buys
&lt;/h2&gt;

&lt;p&gt;Credits, not a flat rate. One monthly pool, spend it on whatever you want.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Hosted&lt;/th&gt;
&lt;th&gt;Hosted Pro&lt;/th&gt;
&lt;th&gt;Hosted Max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Per month&lt;/td&gt;
&lt;td&gt;€19&lt;/td&gt;
&lt;td&gt;€49&lt;/td&gt;
&lt;td&gt;€99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Credits&lt;/td&gt;
&lt;td&gt;900,000&lt;/td&gt;
&lt;td&gt;2,350,000&lt;/td&gt;
&lt;td&gt;5,000,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chat models&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Images, if you spend it all there&lt;/td&gt;
&lt;td&gt;3,000&lt;/td&gt;
&lt;td&gt;7,800&lt;/td&gt;
&lt;td&gt;16,600&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video clips, same&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;156&lt;/td&gt;
&lt;td&gt;333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens, cheapest to priciest model&lt;/td&gt;
&lt;td&gt;0.6M to 225M&lt;/td&gt;
&lt;td&gt;1.6M to 587M&lt;/td&gt;
&lt;td&gt;3.5M to 1250M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LoRA trainings&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Yearly billing is two months free. Top-ups start at €5 for 165,000 credits, they never expire, and they are spent only after the month's plan credits are gone, so a top-up is never wasted by a quiet month.&lt;/p&gt;

&lt;p&gt;The token range is wide because it is a real range. 225M output tokens is Llama 3.1 8B Turbo, which works out to about €0.08 per million output tokens on the €19 plan. 0.6M is Kimi K3, about €30 per million. Same pool, and you choose. The credits follow each model's own rate rather than a blended average, because a blended average is how a bill surprises you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four things it does not do
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It is not GPT or Claude.&lt;/strong&gt; Open weights only, by choice. If your evaluation depends on a closed frontier model, this is the wrong endpoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not unmetered.&lt;/strong&gt; Cache hits bill at the cached rate rather than list, and a run you abort stops billing at the abort. Both of those were bugs first and are now tests. But there is still a meter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not private the way local is.&lt;/strong&gt; Local mode on your own machine is the private one, and it is free forever for exactly that reason. Hosted means your prompt reaches a GPU that is not yours. We never train on it and never sell it, account data sits in the EU and you can delete it, and that is a promise rather than a physical guarantee. If you need the physical guarantee, download the app and run it offline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not a lock-in.&lt;/strong&gt; The catalog is open weights, so every model you use here you can also pull and run yourself. Export your chats and go. That is not generosity, it is the same argument that made us build the local app first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we did it this way
&lt;/h2&gt;

&lt;p&gt;The honest reason for credits instead of tiers-with-limits: we priced the pool off real wholesale cost, and a model that costs 350 times more per token than another cannot sit behind the same "unlimited" label without one of us lying. A pool tells you the truth and lets you spend it where you want, including all of it on video if that is what you came for.&lt;/p&gt;

&lt;p&gt;Local stays free. The cloud is for the days your hardware says no.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lu-labs.ai/pricing" rel="noopener noreferrer"&gt;Pricing and the full model list&lt;/a&gt; · &lt;a href="https://lu-labs.ai/download" rel="noopener noreferrer"&gt;Download the local app&lt;/a&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>opensource</category>
      <category>api</category>
    </item>
    <item>
      <title>Run GLM-5.3 Locally: Real Quant Sizes, the llama.cpp Surprise, and the reasoning_effort Trap</title>
      <dc:creator>David </dc:creator>
      <pubDate>Sat, 05 Sep 2026 16:25:14 +0000</pubDate>
      <link>https://dev.to/purpledoubled/run-glm-53-locally-real-quant-sizes-the-llamacpp-surprise-and-the-reasoningeffort-trap-2nmg</link>
      <guid>https://dev.to/purpledoubled/run-glm-53-locally-real-quant-sizes-the-llamacpp-surprise-and-the-reasoningeffort-trap-2nmg</guid>
      <description>&lt;p&gt;Z.ai published the GLM-5.3-Flash weights on &lt;strong&gt;27 August 2026 at 10:33 UTC&lt;/strong&gt; and the GLM-5.3 flagship on &lt;strong&gt;28 August 2026 at 15:22 UTC&lt;/strong&gt;. The models were on Z.ai's own API first, though I could not find a primary source that dates that launch, so I am not going to put a day on it. I spent the evening reading the actual repositories instead of the announcement. Here is what is worth knowing before you start a download measured in hundreds of gigabytes.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Two models: &lt;strong&gt;GLM-5.3&lt;/strong&gt; at 753.9B params, weights out 28 August, and &lt;strong&gt;GLM-5.3-Flash&lt;/strong&gt; at 320.8B (18B active), weights out 27 August.&lt;/li&gt;
&lt;li&gt;Smallest usable builds: &lt;strong&gt;216.7 GB&lt;/strong&gt; flagship, &lt;strong&gt;93.1 GB&lt;/strong&gt; Flash.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The flagship runs in stock llama.cpp. Flash does not yet.&lt;/strong&gt; Yes, that way round.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;reasoning_effort&lt;/code&gt; defaults to &lt;code&gt;max&lt;/code&gt;, including when you typo it.&lt;/li&gt;
&lt;li&gt;Flash is &lt;strong&gt;MIT&lt;/strong&gt;. The flagship is not.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The architecture surprise
&lt;/h2&gt;

&lt;p&gt;I expected the smaller model to have better support. The opposite is true, and the reason is interesting.&lt;/p&gt;

&lt;p&gt;I diffed the two config files:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;GLM-5.2&lt;/th&gt;
&lt;th&gt;GLM-5.3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;architectures&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;GlmMoeDsaForCausalLM&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;GlmMoeDsaForCausalLM&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;num_hidden_layers&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;78&lt;/td&gt;
&lt;td&gt;78&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;hidden_size&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;6144&lt;/td&gt;
&lt;td&gt;6144&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;n_routed_experts&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;num_experts_per_tok&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;max_position_embeddings&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1048576&lt;/td&gt;
&lt;td&gt;1048576&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Params (GGUF metadata)&lt;/td&gt;
&lt;td&gt;753,864,139,008&lt;/td&gt;
&lt;td&gt;753,864,139,008&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Identical. Every field, and the parameter count down to the last digit. GLM-5.3 is GLM-5.2 with different post-training, exactly as Z.ai described it at the API launch.&lt;/p&gt;

&lt;p&gt;So the flagship's GGUF architecture is &lt;code&gt;glm-dsa&lt;/code&gt;, which llama.cpp has supported since GLM-5.2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s1"&gt;'LLM_ARCH_GLM'&lt;/span&gt; src/llama-arch.cpp
&lt;span class="o"&gt;{&lt;/span&gt; LLM_ARCH_GLM4,     &lt;span class="s2"&gt;"glm4"&lt;/span&gt;     &lt;span class="o"&gt;}&lt;/span&gt;,
&lt;span class="o"&gt;{&lt;/span&gt; LLM_ARCH_GLM4_MOE, &lt;span class="s2"&gt;"glm4moe"&lt;/span&gt;  &lt;span class="o"&gt;}&lt;/span&gt;,
&lt;span class="o"&gt;{&lt;/span&gt; LLM_ARCH_GLM_DSA,  &lt;span class="s2"&gt;"glm-dsa"&lt;/span&gt;  &lt;span class="o"&gt;}&lt;/span&gt;,
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Flash is the genuinely new architecture. It reports itself as &lt;code&gt;glm5next&lt;/code&gt;, which is &lt;strong&gt;still not in the main branch&lt;/strong&gt; on 2 September 2026. &lt;a href="https://github.com/ggml-org/llama.cpp/pull/27754" rel="noopener noreferrer"&gt;PR 27754&lt;/a&gt; is open, and Unsloth's card points at it or at their desktop app.&lt;/p&gt;

&lt;p&gt;If you get an unknown architecture error loading a Flash GGUF, that is this. Check whether the PR has landed before you spend an hour rebuilding, because this will change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Actual sizes
&lt;/h2&gt;

&lt;p&gt;Unsloth packs, read on the evening of 28 August 2026. The Flash pack is dated 26 August 2026 and the flagship pack 28 August, which is worth noting because a GGUF cannot predate the weights it was built from, so 26 August is the lower bound for Flash. MLX builds of the flagship went up on 2 September 2026 as &lt;code&gt;mlx-community/GLM-5.3-4bit&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  GLM-5.3
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UD-IQ1_S     216.7 GB
UD-IQ1_M     228.5 GB
UD-Q2_K_XL   253.9 GB
UD-Q3_K_XL   343.0 GB
UD-Q4_K_XL   467.3 GB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  GLM-5.3-Flash
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UD-IQ1_S      93.1 GB
UD-IQ1_M      97.6 GB
UD-Q2_K_XL   108.7 GB
UD-IQ3_XXS   120.4 GB
UD-Q3_K_XL   147.5 GB
UD-IQ4_XS    156.8 GB
UD-Q4_K_XL   199.7 GB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are file sizes. Your KV cache is extra, and on a model advertising 1,048,576 tokens of context that is not a rounding error.&lt;/p&gt;

&lt;p&gt;Flash is natively multimodal, and the vision half is a separate 1.13 GB file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mmproj-F16.gguf    1.13 GB
mmproj-BF16.gguf   1.16 GB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Skip it and Flash silently cannot see images. There is no warning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it
&lt;/h2&gt;

&lt;p&gt;Flash, per Unsloth's card, once the PR is in your build:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama serve &lt;span class="nt"&gt;-hf&lt;/span&gt; unsloth/GLM-5.3-Flash-GGUF:UD-Q4_K_XL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The flagship works on a stock build today, if you have somewhere to put 216.7 GB.&lt;/p&gt;

&lt;h2&gt;
  
  
  The default that will bite you
&lt;/h2&gt;

&lt;p&gt;GLM-5.3 takes &lt;code&gt;reasoning_effort&lt;/code&gt; with three values: &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;max&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;From the model card: it &lt;strong&gt;defaults to &lt;code&gt;max&lt;/code&gt; if not passed, or if set to any other value.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is no middle default, and an invalid value does not error, it silently gives you the most expensive setting. So a typo in that field buys you a long think on every message. Start at &lt;code&gt;low&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;One caveat, because I measured it rather than read it. Through a hosted OpenAI-compatible provider the behaviour is not the same: on the same arithmetic prompt the flagship spent 4 output tokens at &lt;code&gt;low&lt;/code&gt;, 10 at &lt;code&gt;medium&lt;/code&gt;, 11 at &lt;code&gt;high&lt;/code&gt; and 48 at &lt;code&gt;max&lt;/code&gt;, and sending no field at all landed at 30, not at &lt;code&gt;max&lt;/code&gt;. &lt;code&gt;medium&lt;/code&gt; is honoured there as a real middle rung even though the card does not list it. So the model card describes the model, and whatever sits in front of it may have its own opinion. Measure your own stack before you trust either.&lt;/p&gt;

&lt;h2&gt;
  
  
  Licensing, which went backwards
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5.2&lt;/strong&gt;: MIT&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5.3-Flash&lt;/strong&gt;: MIT&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5.3 flagship&lt;/strong&gt;: Z.ai's own &lt;code&gt;glm-5.3&lt;/code&gt; licence, not MIT&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same architecture and same parameter count as the MIT-licensed 5.2, different terms. If you are shipping something commercial, that is a real decision and not a formality.&lt;/p&gt;

&lt;p&gt;Z.ai was open about the reasoning: they held the weights back for safety evaluation, and named the concern. Their own numbers put ExploitBench at 54.4 percent versus 24.4 for GLM-5.2, and CyberGym at 84.5 versus 77.2. Those are vendor figures from a vendor harness, unreproduced by anyone outside the company at the time of writing, and the delay makes more sense once you see which benchmark moved most.&lt;/p&gt;

&lt;h2&gt;
  
  
  If the hardware is the blocker
&lt;/h2&gt;

&lt;p&gt;Both of these are now in the LU Labs Cloud catalog, which is the product I work on, so read this section with that in mind. GLM-5.3-Flash sits on the Hosted plan, so it is on every plan, and the GLM-5.3 flagship is on Pro and Max. The catalog is served to the clients rather than compiled into them, so they appear in the model picker in the web app and in the desktop app without anyone installing an update.&lt;/p&gt;

&lt;p&gt;Both models think before every answer, so there is now an Effort button next to the Brain button, and it controls how much the model thinks. It is in the web app today; on the desktop it arrives with the next update. Most reasoning models in the catalog offer three settings, Low, Medium and High, and these two add a fourth above them, Max. It starts on High. A higher setting means more output tokens and therefore more credits, so Low is the saving. Tools run natively on both, and Flash takes images.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I did not check
&lt;/h2&gt;

&lt;p&gt;The benchmarks. I have not run them and neither has anyone else independently yet. Everything above about sizes, parameters, architectures and licences comes from the HuggingFace API, the config files and the llama.cpp source, all of which answer in seconds and none of which are a press release.&lt;/p&gt;

&lt;p&gt;If you are picking a model to actually run this week and you do not have 128 GB of memory, none of this is your model. That is a legitimate answer too.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>We Benchmarked Our Agent Against opencode: Same Task, Same Model, 40 Percent Fewer Credits</title>
      <dc:creator>David </dc:creator>
      <pubDate>Sun, 23 Aug 2026 00:11:55 +0000</pubDate>
      <link>https://dev.to/purpledoubled/we-benchmarked-our-agent-against-opencode-same-task-same-model-40-percent-fewer-credits-14df</link>
      <guid>https://dev.to/purpledoubled/we-benchmarked-our-agent-against-opencode-same-task-same-model-40-percent-fewer-credits-14df</guid>
      <description>&lt;p&gt;Every coding agent says it is efficient. Almost none of them publish the bill. So we ran the boring experiment: the same bugfix, the same model, the same API, the same prices, and a byte identical prompt, once through &lt;a href="https://github.com/sst/opencode" rel="noopener noreferrer"&gt;opencode&lt;/a&gt; and once through the coding agent inside Locally Uncensored.&lt;/p&gt;

&lt;p&gt;Headline: opencode averaged &lt;strong&gt;2157 credits&lt;/strong&gt; over three runs. Our 2.6.6 agent finished the identical task for &lt;strong&gt;1298&lt;/strong&gt;. That is about 40 percent less, and even the cheapest opencode run came in 29 percent above our number.&lt;/p&gt;

&lt;p&gt;The interesting part is not the headline. It is &lt;em&gt;why&lt;/em&gt; the gap exists, and it is not the reason most people guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;p&gt;A cost comparison is only worth reading if everything that drives cost is nailed down. What was held constant:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Held constant&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Task&lt;/td&gt;
&lt;td&gt;Fix a failing test in a small npm repo, then commit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repository&lt;/td&gt;
&lt;td&gt;Three files, a one line bug in &lt;code&gt;add.js&lt;/code&gt;, tests red at the start&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt&lt;/td&gt;
&lt;td&gt;Byte identical, sha256 &lt;code&gt;29cec6c3...cf62687&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;&lt;code&gt;deepseek-ai/DeepSeek-V3.2&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint&lt;/td&gt;
&lt;td&gt;The same OpenAI compatible API for both agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prices&lt;/td&gt;
&lt;td&gt;Same account, same tier, same per token rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Counting&lt;/td&gt;
&lt;td&gt;One wire proxy in front of the API, credits read before and after every run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;opencode&lt;/td&gt;
&lt;td&gt;1.18.21 from npm, wired as an OpenAI compatible provider, &lt;code&gt;opencode run --auto&lt;/code&gt;, otherwise defaults&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Success was defined before the runs, not after:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;npm test&lt;/code&gt; passes&lt;/li&gt;
&lt;li&gt;exactly one commit, with the required message&lt;/li&gt;
&lt;li&gt;only &lt;code&gt;add.js&lt;/code&gt; changed&lt;/li&gt;
&lt;li&gt;clean working tree at the end&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All four runs cleared that bar. Nothing failed, so cost is the only variable that moved.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Credits&lt;/th&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Prompt tokens&lt;/th&gt;
&lt;th&gt;Success&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;opencode, run 1&lt;/td&gt;
&lt;td&gt;1679&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;98,789&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;opencode, run 2&lt;/td&gt;
&lt;td&gt;2433&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;146,058&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;opencode, run 3&lt;/td&gt;
&lt;td&gt;2358&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;146,387&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Locally Uncensored 2.6.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1298&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;74,629&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Locally Uncensored 2.6.5&lt;/td&gt;
&lt;td&gt;4395&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;257,270&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read the last row first. Our own shipped agent from one release earlier is the most expensive thing in that table, by a lot. This is not a chart built so that we win by construction. It is a chart that shows what one efficiency pass is worth, and the previous version of our own software is the loser in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the gap exists
&lt;/h2&gt;

&lt;p&gt;The tempting explanation is that one agent is smarter and needs fewer steps. That is not what happened, and the direction is the reverse of what you would expect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;opencode used fewer turns than we did.&lt;/strong&gt; Eight to eleven requests against our sixteen. If you scored this on steps, opencode wins. The bill went the other way because of what every single request carries.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Per request&lt;/th&gt;
&lt;th&gt;opencode&lt;/th&gt;
&lt;th&gt;Locally Uncensored 2.6.6&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prompt tokens per request&lt;/td&gt;
&lt;td&gt;12,349 to 13,308&lt;/td&gt;
&lt;td&gt;4,664&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool catalogue size&lt;/td&gt;
&lt;td&gt;21,188 bytes&lt;/td&gt;
&lt;td&gt;7,703 bytes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Credits per prompt token&lt;/td&gt;
&lt;td&gt;0.01700 / 0.01666 / 0.01611&lt;/td&gt;
&lt;td&gt;0.01739&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the honest one. &lt;strong&gt;The billing rate is the same.&lt;/strong&gt; Credits per prompt token land within a few percent across all four runs, and ours is marginally the highest of the set. Nobody got a secret discount. The entire difference in the invoice is token volume, not token price.&lt;/p&gt;

&lt;p&gt;Two things drive that volume, and both are familiar to anyone who has built an agent loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The fixed block.&lt;/strong&gt; A tool catalogue of 21,188 bytes against 7,703 bytes is roughly three times the standing overhead, and you pay it again on every call in the loop, whether the model touches those tools or not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context decay.&lt;/strong&gt; As the agent works, the transcript grows. Old tool output that stopped mattering ten steps ago keeps getting resent at full length unless something actively trims it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Put together, a fixed block that size pushes every opencode request past 12,000 tokens. Six agent steps at that weight already approach our total consumption for the entire task.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed between our own 2.6.5 and 2.6.6
&lt;/h2&gt;

&lt;p&gt;The 4395 in the table is not a strawman we built for the article, it is what we shipped in the previous release. Between 2.6.5 and 2.6.6 we went after exactly the two items above: how big the fixed block is, and how much of the transcript gets resent. Measured over our own set of tool driven runs, Locally Uncensored 2.6.6 spent &lt;strong&gt;78.6 percent fewer credits than Locally Uncensored 2.6.5&lt;/strong&gt;, and on the longest run in that set 80.4 percent fewer than 2.6.5. Both figures compare our own agent against our own previous release. Neither of them is the gap to opencode, which is the roughly 40 percent above. The two also come from different sets: on the single task in the table above, our own 2.6.5 would have cost 4395 credits against 1298, which is 70 percent. The opencode comparison is simply what fell out when we pointed the same measurement at somebody else's loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits of this benchmark
&lt;/h2&gt;

&lt;p&gt;This is where vendor benchmarks usually go quiet, so here is how far the number actually carries.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One scenario.&lt;/strong&gt; A tiny repo and a one line bug. It says nothing about a large codebase, a multi file refactor, or a session that runs for an hour. We did not run those.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uneven sample.&lt;/strong&gt; opencode ran three times, we ran once. The spread inside opencode alone is 45 percent, from 1679 to 2433. Our own spread is unknown. One run is a data point, not a distribution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Default settings.&lt;/strong&gt; opencode ran as it ships. It is configurable, and a tuned setup with a trimmed tool set would land somewhere else. We did not tune it in either direction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;opencode is free software.&lt;/strong&gt; The tool costs nothing. Everything measured here is the model bill, which you pay to whichever provider you point at. This is token efficiency, not licence fees.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost is not quality.&lt;/strong&gt; Every run in the table produced correct, committed work. On a harder problem the ranking could look different, and cheapest is never automatically best.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What the measurement does support is a narrower claim than the headline: on short, well scoped agent tasks, opencode can hardly land below us, because the fixed per request overhead sets a floor. Even its best run, with only eight requests, still needed 98,789 tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Credit where it is due
&lt;/h2&gt;

&lt;p&gt;opencode finished the job three times out of three, took fewer turns than we did, and produced a clean diff with the right commit message every time. It is a genuinely good agent and it is open source. Nothing here is an argument to stop using it.&lt;/p&gt;

&lt;p&gt;It is an argument to measure your own loop. Agent bills are made of tokens you never see, and two tools that both feel fast can be a factor of 1.66 apart on the invoice. If you build agents, the two numbers worth putting on a dashboard are prompt tokens per request and the byte size of your tool catalogue. They predict the bill better than step count does.&lt;/p&gt;

&lt;p&gt;Full writeup with the methodology and the raw counts: &lt;a href="https://locallyuncensored.com/blog/opencode-alternative-cost-benchmark.html" rel="noopener noreferrer"&gt;opencode Alternative: We Measured the Cost per Task&lt;/a&gt;. The agent lives inside &lt;a href="https://github.com/PurpleDoubleD/locally-uncensored" rel="noopener noreferrer"&gt;Locally Uncensored&lt;/a&gt; (AGPL, free), and the hosted models we benchmarked against sit in &lt;a href="https://lu-labs.ai" rel="noopener noreferrer"&gt;LU Labs Cloud&lt;/a&gt; if you want them in one picker. You can also point it at a model on your own GPU and skip the API bill entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is opencode expensive?&lt;/strong&gt; opencode is free. The bill is the model. On this one line bugfix with DeepSeek V3.2 the three runs cost 1679, 2433 and 2358 credits, a 45 percent spread between cheapest and dearest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does it use so many tokens?&lt;/strong&gt; Not through extra steps, it used fewer than we did. Each request carries more: 12,349 to 13,308 prompt tokens against our 4,664, with a tool catalogue of 21,188 bytes against 7,703 resent on every call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is the billing rate really identical?&lt;/strong&gt; Yes, and that is the point. Credits per prompt token came out at 0.01739 for us and 0.01700, 0.01666, 0.01611 for opencode. The gap is volume, not price.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I reproduce it?&lt;/strong&gt; Three file npm repo with a red test, prompt pinned at sha256 &lt;code&gt;29cec6c3...cf62687&lt;/code&gt;, model &lt;code&gt;deepseek-ai/DeepSeek-V3.2&lt;/code&gt;, opencode 1.18.21 at defaults, requests counted through a wire proxy, credits read before and after each run.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Run Qwen 3.8 27B Locally: Real GGUF Sizes, the KV Cache Trick, and the Template Trap</title>
      <dc:creator>David </dc:creator>
      <pubDate>Fri, 14 Aug 2026 20:48:03 +0000</pubDate>
      <link>https://dev.to/purpledoubled/run-qwen-38-27b-locally-real-gguf-sizes-the-kv-cache-trick-and-the-template-trap-114j</link>
      <guid>https://dev.to/purpledoubled/run-qwen-38-27b-locally-real-gguf-sizes-the-kv-cache-trick-and-the-template-trap-114j</guid>
      <description>&lt;p&gt;Qwen 3.8 arrived as two different releases with two different licences, and only one of them is something you can put on a card you own. The 2.4 trillion parameter A95B opened up on 12 August under Alibaba's own &lt;code&gt;qwen3.8-max&lt;/code&gt; terms. The one that matters for local work is &lt;strong&gt;Qwen 3.8 27B&lt;/strong&gt;, whose safetensors went up on 13 August at 08:23 UTC with an Apache 2.0 LICENSE file following the next morning. Both dates are off the Hugging Face commit log, not a launch post.&lt;/p&gt;

&lt;p&gt;Here is the practical picture: what it needs, why its long context is unusually cheap, and the one setting that makes people think they downloaded a broken quant.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the model decides everything
&lt;/h2&gt;

&lt;p&gt;27B dense parameters across 64 layers, hidden size 5120. The interesting part is in &lt;code&gt;config.json&lt;/code&gt;, where &lt;code&gt;layer_types&lt;/code&gt; reads &lt;strong&gt;48 linear attention layers and 16 full attention layers&lt;/strong&gt;, alternating three to one (&lt;code&gt;full_attention_interval: 4&lt;/code&gt;). Only those 16 layers keep a KV cache.&lt;/p&gt;

&lt;p&gt;The rest of the shape: 24 attention heads with &lt;code&gt;head_dim&lt;/code&gt; 256 and &lt;strong&gt;4 KV heads&lt;/strong&gt;, a 248,320 token vocabulary, and &lt;code&gt;max_position_embeddings&lt;/code&gt; of &lt;strong&gt;262,144&lt;/strong&gt;. It is a native vision language model, so images and video go in without a wrapper, and the &lt;code&gt;ggml-org&lt;/code&gt; pack also ships a multi token prediction head as a separate file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;p&gt;Sizes below are the file sizes Hugging Face reports for &lt;a href="https://huggingface.co/unsloth/Qwen3.8-27B-GGUF" rel="noopener noreferrer"&gt;&lt;code&gt;unsloth/Qwen3.8-27B-GGUF&lt;/code&gt;&lt;/a&gt;, read on 14 August 2026. Packs differ by a few hundred megabytes, so check the repo you actually pull from. &lt;code&gt;lmstudio-community&lt;/code&gt; has Q4_K_M at 16.8 GB and &lt;code&gt;ggml-org&lt;/code&gt; at 19.0 GB for the same nominal quant.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quant&lt;/th&gt;
&lt;th&gt;Size on disk&lt;/th&gt;
&lt;th&gt;Realistic home&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;UD-IQ2_XXS&lt;/td&gt;
&lt;td&gt;9.0 GB&lt;/td&gt;
&lt;td&gt;12 GB cards, visible quality cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UD-Q2_K_XL&lt;/td&gt;
&lt;td&gt;10.7 GB&lt;/td&gt;
&lt;td&gt;12 GB cards, almost no context left&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UD-Q3_K_XL&lt;/td&gt;
&lt;td&gt;13.4 GB&lt;/td&gt;
&lt;td&gt;16 GB cards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q3_K_M&lt;/td&gt;
&lt;td&gt;13.8 GB&lt;/td&gt;
&lt;td&gt;16 GB cards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IQ4_XS&lt;/td&gt;
&lt;td&gt;15.7 GB&lt;/td&gt;
&lt;td&gt;largest quant that stays whole on 16 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Q4_K_M (sweet spot)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;17.1 GB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;24 GB cards&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q5_K_M&lt;/td&gt;
&lt;td&gt;19.8 GB&lt;/td&gt;
&lt;td&gt;24 GB, less context headroom&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q6_K&lt;/td&gt;
&lt;td&gt;22.9 GB&lt;/td&gt;
&lt;td&gt;24 GB barely, or 32 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q8_0&lt;/td&gt;
&lt;td&gt;29.0 GB&lt;/td&gt;
&lt;td&gt;32 GB or a two card split&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BF16 (from &lt;code&gt;ggml-org&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;53.8 GB&lt;/td&gt;
&lt;td&gt;server cards, or CPU and patience&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mmproj&lt;/td&gt;
&lt;td&gt;0.9 GB&lt;/td&gt;
&lt;td&gt;the vision encoder, separate file&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Machine classes, honestly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;24 GB GPU:&lt;/strong&gt; Q4_K_M whole, with real context headroom. This is the card the model was sized for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;16 GB GPU:&lt;/strong&gt; IQ4_XS fits whole; Q4_K_M works with a few layers offloaded and costs you speed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;12 GB GPU:&lt;/strong&gt; only the 2-bit quants, and you will feel it. A 3060 runs it, slowly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apple Silicon:&lt;/strong&gt; 32 GB unified memory is comfortable at Q4_K_M, 24 GB works if nothing else is open.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where the hybrid layout pays off
&lt;/h2&gt;

&lt;p&gt;Every full attention layer stores a KV cache that grows with the context. At fp16 one token costs &lt;code&gt;2 x 4 heads x 256 dim x 2 bytes = 4 KB&lt;/code&gt; per layer. A conventional 64 layer model pays that on all 64 layers, which is 256 KB per token. Qwen 3.8 27B pays it on &lt;strong&gt;16&lt;/strong&gt; layers, so &lt;strong&gt;64 KB per token&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;KV cache&lt;/th&gt;
&lt;th&gt;Q4_K_M total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;8K&lt;/td&gt;
&lt;td&gt;0.5 GB&lt;/td&gt;
&lt;td&gt;17.6 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32K&lt;/td&gt;
&lt;td&gt;2.0 GB&lt;/td&gt;
&lt;td&gt;19.1 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;8.0 GB&lt;/td&gt;
&lt;td&gt;25.1 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;262K&lt;/td&gt;
&lt;td&gt;16.4 GB&lt;/td&gt;
&lt;td&gt;33.5 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is the difference between a long window on the spec sheet and one you actually turn on. A 128K session that would need 32 GB of cache on a dense model needs 8 GB here.&lt;/p&gt;

&lt;p&gt;With a current llama.cpp:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="nt"&gt;-m&lt;/span&gt; Qwen3.8-27B-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--jinja&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-ngl&lt;/span&gt; 99 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt; 32768
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The template trap
&lt;/h2&gt;

&lt;p&gt;That &lt;code&gt;--jinja&lt;/code&gt; flag is not optional, and it is the single biggest source of "this quant is broken" reports.&lt;/p&gt;

&lt;p&gt;Qwen 3.8 ships its own chat template. Load the model without it and the model has no reliable marker for where your turn ends and its answer begins. The two failure modes both look like a bad conversion: either it rambles past the stop token, or it answers in a clipped voice and loses the conversation between turns.&lt;/p&gt;

&lt;p&gt;There is a second, sharper version of this. The official template wraps each assistant turn in a think block even when the reasoning is empty, then opens another one when generation starts. Across several turns those nest and the history gets truncated. Several GGUF packs already ship a corrected &lt;code&gt;chat_template.jinja&lt;/code&gt;. If you converted the weights yourself, swap the template before you blame the quantization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vision needs the second file
&lt;/h2&gt;

&lt;p&gt;The vision encoder is not inside the language GGUF. Download the &lt;code&gt;mmproj&lt;/code&gt; file, about 0.9 GB, and load it alongside:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="nt"&gt;-m&lt;/span&gt; Qwen3.8-27B-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--mmproj&lt;/span&gt; mmproj-F16.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--jinja&lt;/span&gt; &lt;span class="nt"&gt;-ngl&lt;/span&gt; 99
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Skip it and you have a strong text model that will politely tell you it cannot see the image you just pasted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Thinking is a dial, not a switch
&lt;/h2&gt;

&lt;p&gt;The model reasons before answering by default. Lower &lt;code&gt;reasoning_effort&lt;/code&gt; from the default to medium or low for quick answers, or pass &lt;code&gt;chat_template_kwargs&lt;/code&gt; with &lt;code&gt;enable_thinking: false&lt;/code&gt; to skip the reasoning pass entirely. Locally the cost of thinking is not money, it is your own time watching tokens arrive, which is the more annoying currency.&lt;/p&gt;

&lt;h2&gt;
  
  
  The no-terminal route
&lt;/h2&gt;

&lt;p&gt;If you would rather click than type flags, &lt;a href="https://github.com/PurpleDoubleD/locally-uncensored" rel="noopener noreferrer"&gt;Locally Uncensored&lt;/a&gt; (open source, AGPL) wraps this: install, open the Model Manager, paste a Qwen 3.8 27B GGUF repo, pick the quant that fits your card, chat. It carries the llama.cpp engine, handles the offload split and the template, and keeps everything on your machine with no account and no telemetry.&lt;/p&gt;

&lt;h2&gt;
  
  
  And the 2.4T reality check
&lt;/h2&gt;

&lt;p&gt;The big Qwen 3.8, &lt;code&gt;Qwen/Qwen3.8-2.4T-A95B&lt;/code&gt;, has open weights but not an open licence. Hugging Face reports &lt;code&gt;license: other&lt;/code&gt; with &lt;code&gt;license_name: qwen3.8-max&lt;/code&gt;, so read the terms before you build a product on it, and do not repeat the "Apache" line that has been going around. At 2.4 trillion total parameters with 95 billion active it is a data center model regardless.&lt;/p&gt;

&lt;p&gt;The split that works: run the 27B at home for anything private, offline, or repetitive, and reach the A95B through a hosted API when a task genuinely needs that scale. It is on DeepInfra, and both it and the 27B class of models sit in &lt;a href="https://lu-labs.ai/pricing" rel="noopener noreferrer"&gt;LU Labs Cloud&lt;/a&gt; if you want them next to each other in one picker.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is 27B enough?&lt;/strong&gt; For local work it is the interesting size: big enough for real coding and agent loops, small enough to sit on one consumer card. The hybrid attention means long context does not price you out the way it does on a dense model of the same parameter count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ollama or LM Studio?&lt;/strong&gt; Both pick up the shipped template automatically, which spares you the trap above. The bare-GGUF-plus-llama.cpp route is where people get bitten.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Uncensored builds?&lt;/strong&gt; Apache 2.0 makes finetunes and abliterations fully legal, and community variants started appearing within a day of the weights. The stock model carries standard alignment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I check the sizes myself?&lt;/strong&gt; &lt;code&gt;https://huggingface.co/api/models/&amp;lt;repo&amp;gt;/tree/main&lt;/code&gt; returns every file with its byte count. That is where every number in this post came from.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Run Ling 3.0 Flash Locally: 124B of Knowledge on a 96 GB Machine</title>
      <dc:creator>David </dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:56:06 +0000</pubDate>
      <link>https://dev.to/purpledoubled/run-ling-30-flash-locally-124b-of-knowledge-on-a-96-gb-machine-4d2h</link>
      <guid>https://dev.to/purpledoubled/run-ling-30-flash-locally-124b-of-knowledge-on-a-96-gb-machine-4d2h</guid>
      <description>&lt;p&gt;The two big open releases on everyone's feed right now are Kimi K3 (2.8 trillion parameters, the first open 3T-class model) and Ling 3.0 Flash from Ant Group's inclusionAI. Only one of them can live on hardware a person owns, and it is not the one with the bigger headline. Here is the practical picture for running Ling 3.0 Flash locally, with measured file sizes instead of projections, plus the honest math on why K3 stays in the cloud.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the model decides everything
&lt;/h2&gt;

&lt;p&gt;Ling 3.0 Flash is a 124B parameter Mixture of Experts that activates only &lt;strong&gt;5.1B parameters per token&lt;/strong&gt; (8 of 512 routed experts plus one shared). As with every MoE, total parameters set your memory bill and active parameters set your speed. 124B of knowledge at the memory counter, 5.1B of compute at the speed counter: that combination is why this model runs interactively on machines that would crawl under a dense 70B.&lt;/p&gt;

&lt;p&gt;Two more architecture notes that matter in practice. It uses hybrid linear attention (35 Kimi Delta Attention layers alternating with 7 gated MLA layers), so long contexts grow memory gently instead of quadratically. And it is a native hybrid reasoner: it thinks before answering by default, and the thinking pass can be switched off per request, so you decide when to pay for reasoning.&lt;/p&gt;

&lt;p&gt;The license is plain MIT. Weights are on Hugging Face under &lt;code&gt;inclusionAI/Ling-3.0-flash&lt;/code&gt;, including official fp8, fp4 and int4 variants.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;p&gt;Community GGUF conversions are up, and these sizes are measured from the repos, not estimated:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quant&lt;/th&gt;
&lt;th&gt;Size on disk&lt;/th&gt;
&lt;th&gt;Realistic minimum&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;IQ1_S / IQ1_M&lt;/td&gt;
&lt;td&gt;27 to 30 GB&lt;/td&gt;
&lt;td&gt;36 GB RAM, visible quality cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IQ2_M / Q2_K_XL&lt;/td&gt;
&lt;td&gt;42 to 43 GB&lt;/td&gt;
&lt;td&gt;48 to 64 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IQ3_XXS&lt;/td&gt;
&lt;td&gt;51 GB&lt;/td&gt;
&lt;td&gt;64 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Q4_K_M (sweet spot)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;78 GB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96 GB&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q5_K_M&lt;/td&gt;
&lt;td&gt;92 GB&lt;/td&gt;
&lt;td&gt;128 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q6_K&lt;/td&gt;
&lt;td&gt;105 GB&lt;/td&gt;
&lt;td&gt;128 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q8_0&lt;/td&gt;
&lt;td&gt;136 GB&lt;/td&gt;
&lt;td&gt;192 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Realistic machine classes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Mac with 96 or 128 GB unified memory runs Q4_K_M comfortably, and 64 GB machines run the 2-bit and small 3-bit quants whole.&lt;/li&gt;
&lt;li&gt;A 24 GB GPU plus 64 to 96 GB of DDR5 works well with MoE offload: attention layers and the shared expert on the GPU, routed experts in system RAM. Because only 5.1B parameters fire per token, generation speed stays in usable double digits.&lt;/li&gt;
&lt;li&gt;32 GB or less: skip it and run a model that actually fits. The IQ1 files exist, but you will not enjoy them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With a current llama.cpp:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="nt"&gt;-m&lt;/span&gt; Ling-3.0-flash-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ctx-size&lt;/span&gt; 32768 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--n-gpu-layers&lt;/span&gt; 99 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--n-cpu-moe&lt;/span&gt; 40
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start with a modest context window; the model was trained out to 256K, but every token of window is memory you could spend on a better quant instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  The no-terminal route
&lt;/h2&gt;

&lt;p&gt;If you would rather click than type flags, &lt;a href="https://github.com/PurpleDoubleD/locally-uncensored" rel="noopener noreferrer"&gt;Locally Uncensored&lt;/a&gt; (open source, AGPL) wraps the whole thing: install, open the Model Manager, paste a Ling 3.0 Flash GGUF repo, pick the quant that fits your memory, chat. It handles the llama.cpp engine, the offload split, and keeps everything on your machine with no account and no telemetry.&lt;/p&gt;

&lt;h2&gt;
  
  
  And the Kimi K3 reality check
&lt;/h2&gt;

&lt;p&gt;K3 deserves its headlines: first open 3T-class release, 1M token context, native image input, and it shares the Kimi Delta Attention lineage with Ling. But the smallest usable GGUF conversion, a 1-bit quant, is &lt;strong&gt;466 GB on disk&lt;/strong&gt;. The 2-bit is 861 GB. Q4 is 1.5 TB. Expert-pruned community builds squeeze toward the 512 GB server class, which is heroic and still not a gaming PC.&lt;/p&gt;

&lt;p&gt;So the honest split for the price of one search query: run Ling 3.0 Flash at home, and reach K3 through a hosted API when a task genuinely needs the 1M window or vision. DeepInfra serves both (Ling at $0.03 in / $0.07 out per million tokens, which rounds to free; K3 at $2.85 / $14.25), OpenRouter lists a free Ling tier, and both landed in &lt;a href="https://lu-labs.ai/pricing" rel="noopener noreferrer"&gt;LU Labs Cloud&lt;/a&gt; on every plan this week, with reasoning as a toggle in both cases.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is it actually good?&lt;/strong&gt; inclusionAI benchmarks it at parity with their previous 1T-class flagship on SWE-Bench Pro, agentic tool suites and long-context tasks. Vendor numbers, as always, but the architecture math is real, and the price of testing that claim yourself is a weekend download.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How fast is it locally?&lt;/strong&gt; Speed tracks the 5.1B active count, not the 124B total. On Apple Silicon at Q4 and on GPU-plus-RAM hybrids, expect double-digit tokens per second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Uncensored builds?&lt;/strong&gt; None yet. The MIT license makes finetunes and abliterations fully legal, and a 5.1B-active model is cheap to tune, so expect community variants. The stock model carries standard alignment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How to Run Qwen 3.8 Locally: The 2.4T Max Math, the Open Weight 27B, and What Runs Today</title>
      <dc:creator>David </dc:creator>
      <pubDate>Mon, 03 Aug 2026 15:43:05 +0000</pubDate>
      <link>https://dev.to/purpledoubled/how-to-run-qwen-38-locally-the-24t-max-math-the-open-weight-27b-and-what-runs-today-k31</link>
      <guid>https://dev.to/purpledoubled/how-to-run-qwen-38-locally-the-24t-max-math-the-open-weight-27b-and-what-runs-today-k31</guid>
      <description>&lt;p&gt;Alibaba made Qwen 3.8 Max generally available today (August 3): roughly 2.4 trillion parameters as a Mixture of Experts, a 1 million token context window, image and video input, and launch scores like 92.6 on GPQA Diamond and 86.6 on Terminal-Bench 2.1. The part that matters for this post: the open weight checkpoint is promised for the coming week, and it brings a consumer sized &lt;strong&gt;Qwen 3.8 27B&lt;/strong&gt; along with it. Here is the practical picture for anyone who wants Qwen 3.8 on their own hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the model decides everything
&lt;/h2&gt;

&lt;p&gt;MoE models split the hardware question in two. Total parameters (2.4T) set your memory bill: every expert has to live somewhere. Active parameters (~95B per launch coverage) set your speed. That is why the Max serves cheaply in a datacenter, and also why it will never fit in your tower: all 2.4T parameters must sit in memory, because you never know which experts the next token activates.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers for the Max
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;FP8, the native serving precision: ~2.4 TB of weights.&lt;/li&gt;
&lt;li&gt;Q4, the local standard: ~1.2 TB.&lt;/li&gt;
&lt;li&gt;An extreme 2 bit quant: ~600 GB, with real quality loss.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The largest single machine a consumer can buy in 2026 holds 512 GB of unified memory, less than half of the Q4 weights, before you allocate a single byte of KV cache for that million token context. Someone will chain Mac Studios together for a single digit tokens per second demo within weeks of the weights dropping. It will be a great video and a bad daily driver.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 27B is the actual local story
&lt;/h2&gt;

&lt;p&gt;This is the difference between this launch and Kimi K3 or DeepSeek V4 Pro: the same drop includes an open weight 27B. Projected from the Qwen 3.6 27B precedent:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;VRAM&lt;/th&gt;
&lt;th&gt;Quant class&lt;/th&gt;
&lt;th&gt;Precedent size&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;8 GB&lt;/td&gt;
&lt;td&gt;2 bit (UD-IQ2 class)&lt;/td&gt;
&lt;td&gt;8.7 GB, reduced quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12 GB&lt;/td&gt;
&lt;td&gt;Q3_K_M&lt;/td&gt;
&lt;td&gt;13 GB, RTX 3060 sweet spot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16 GB&lt;/td&gt;
&lt;td&gt;Q4_K_M, recommended&lt;/td&gt;
&lt;td&gt;~16 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;24 GB&lt;/td&gt;
&lt;td&gt;Q6_K, near lossless&lt;/td&gt;
&lt;td&gt;~21 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Community GGUFs of open Qwen releases usually appear within days of the weights. If the 27B inherits even part of the Max's agentic gains (FrontierSWE jumped from 40.7 to 73.5 this generation), it becomes the default local model in its class more or less immediately. And weights mean derivatives: every open Qwen generation has received abliterated and heretic builds within weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to run today, not next week
&lt;/h2&gt;

&lt;p&gt;Until the drop, the newest Qwen you can actually run is Qwen 3.6, and it is genuinely strong: the 27B dense runs from 8 GB VRAM up, the 35B MoE (3B active) is the coding pick at 24 GB. One line with a current llama.cpp era stack:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull qwen3.6:27b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you would rather click than type, the free and open source local AI studio &lt;a href="https://locallyuncensored.com/" rel="noopener noreferrer"&gt;Locally Uncensored&lt;/a&gt; has the Qwen family in its one click model catalog, checks your memory before you download, and speaks to 12 local backends. The full setup walkthrough with quant tables lives here: &lt;a href="https://locallyuncensored.com/blog/how-to-run-qwen-3-6-locally.html" rel="noopener noreferrer"&gt;How to Run Qwen 3.6 Locally&lt;/a&gt;, and the complete Qwen 3.8 hardware math is in &lt;a href="https://locallyuncensored.com/blog/can-you-run-qwen-3-8-locally.html" rel="noopener noreferrer"&gt;Can You Run Qwen 3.8 Locally?&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bother running it at home
&lt;/h2&gt;

&lt;p&gt;Privacy: your prompts never leave the machine. Cost: open weights turn a metered bill into a one time hardware decision. Control: the model you benchmark today is the model you run next year, no silent upstream swaps. For the first time, the frontier launch everyone is hyping comes with a version of itself you will actually own, in the same week's drop.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
  </channel>
</rss>
