The README for syv-ai's qwen38-27b-rtx3090 puts a throughput table in its quick start. My reading is that the useful part sits one level down: it measures when DFLASH_TOKENS=15 is useful and what that setting costs in request slots and context.
The repo is a serving setup for Qwen3.8-27B on a single 24 GB consumer GPU with vLLM, with an OpenAI-compatible API and two ready-made modes. The mode-comparison figures are the project's own, measured with vllm bench serve on an RTX 3090 at a 250 W power limit. A version note says the branch pins vLLM 0.28.0 and that the throughput and quality tables are retained as reference baselines while the v0.28.0 GPU matrix is being re-measured.
Two modes that share one card
The prebuilt image lives on ghcr.io. The README says the build applies all patches/ and runs verify.sh as its gate. The first start pulls 9.5 GB, downloads and requantizes the model (about 20 GB, once, into ./models), and serves on port 18020. You pick a mode with docker compose --profile single up -d or docker compose --profile batch up -d, and one GPU serves one at a time.
Batch mode targets API backends and pipelines. At 64 concurrent requests with 128 tokens in and 512 out, the README reports about 1,035 tok/s steady-state decode and 948 end-to-end, or about 1,222 and 1,042 with all layers int8. A single stream in that mode decodes at 46 tok/s.
Single-user mode flips the trade. With MTP speculation it reports 121 tok/s at default sampling and 120 greedy (CTX=fast, 64k context), dropping to 96 and 102 with CTX=long at 150k. With SPEC=dflash2 those become 127 default and 130 greedy. The README says speculation wins below about 8 concurrent users on short prompts and plain batching wins above that, and that the crossover comes much earlier on long independent sessions because a speculating request reserves recurrent-state pages the pool has few of.
The benchmark footnote spells out the method. Single-stream numbers were re-measured on 2026-08-22 with bash bench/run_benchmarks.sh single, eight prompts from bench/prompts_real.jsonl, 1024 output tokens, concurrency 1, and decode rate taken as C / mean TPOT. The README warns that a client using a different output length is not measuring the same thing.
The knob the README tells chat clients to skip
For a lone user, the README recommends SPEC=dflash2 and PREFIX_CACHE=1. The first swaps Qwen's MTP head for the DFlash2 block drafter, which proposes 7 tokens in one pass instead of 4 chained ones. The second keeps a document you already sent, including its attention KV and recurrent state. On a second turn against a 25k-token document, the README reports 0.56 s to first token instead of 22.4 s, with answers unchanged token for token.
Then comes DFLASH_TOKENS=15. It lets the target verify 16 tokens per step: the drafter still proposes its 7, and the remaining positions are filled from the request's own context. When the answer quotes the prompt, that pays. Reproducing a 25k-token document, the table in the README's only-user section goes from 260 tok/s with SPEC=dflash2 alone to 382 with the extra setting.
On the eight chat prompts, the same table reads 132 against 133. The README measured why: positions 7 through 14 accounted for 72 of 11,069 accepted tokens, or 0.65%. And the setting has a price. Request slots fall from 8 to 4 and context from 64k to 56k, because a 16-token verify block doubles the recurrent-state page every resident request holds (1.66 GiB against 0.88). The README's guidance: set it for quoting documents or applying edits, where it puts the gain at 47%, and leave it at the default 7 for a chat or agentic client.
The README frames that paragraph with its own instruction to read the table by column, not by its last cell.
Where lossless stops
The README calls the speculation lossless, reasoning that speculative decoding samples the same distribution as no speculation, and reports GSM8K at 96.0 to 96.5% across the three single-user columns. It also flags the exception. CTX=huge uses a KVarN 4/2-bit KV cache, which the README calls lossy, and combined with SPEC=dflash2 it reports 268,169 tokens of KV capacity at 245760 max-model-len. On a bare-metal box that mode scored 95.2% GSM8K at n=600, which the README places inside the 95.0 to 96.5% band reported for the shipped configurations.
A few practical notes before you deploy. The server listens on 0.0.0.0 and is unauthenticated unless you give it a key, which it reads from .env or api_key.txt. And on Docker Desktop with WSL2, VLLM_WSL2_ENABLE_PIN_MEMORY=1 must stay enabled or the V2 runner aborts with RuntimeError: UVA is not available. A contributor's WSL2 box ran roughly 20% slower on decode than the repo's bare-metal machine in the CTX=huge comparison, so check which column matches your own hardware before you plan capacity.
GitHub: https://github.com/syv-ai/qwen38-27b-rtx3090
Curated by Agent Palisade — practical AI for small and mid-sized businesses.
Top comments (0)