DEV Community

Cover image for Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context
Dmitry Amelchenko
Dmitry Amelchenko

Posted on

Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context

Understanding My llama.cpp Qwen 3.8 Configuration

I've been tuning llama.cpp for local AI development, and the command line can quickly become a collection of cryptic flags.

Here's what my current configuration does, parameter by parameter.
I'm specifically focusing on maxing out the utilization of my system (which is MBP M5 with 128 GB Unified RAM), for multi-agent coding which requires parallel agents execution.

llama serve \
  -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL \
  --spec-type draft-mtp \
  --spec-default \
  --spec-draft-n-max 8 \
  -ngl 99 \
  -c 524288 \
  --override-kv qwen2.context_length=int:524288 \
  --rope-scaling yarn \
  --yarn-orig-ctx 262144 \
  -b 16384 -ub 4096 \
  -t 16 \
  -tb 16 \
  -np 2 \
  -fa on \
  --cache-type-k f16 \
  --cache-type-v f16 \
  --kv-offload \
  --load-mode none \
  --host 127.0.0.1 \
  --port 8080
Enter fullscreen mode Exit fullscreen mode

The easiest way to understand it is to divide the configuration into several areas:

  • Model
  • Speculative decoding
  • Context
  • GPU
  • Batching
  • CPU
  • KV cache
  • Server configuration

1. Model

-hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL
Enter fullscreen mode Exit fullscreen mode

This tells llama.cpp to download and load the model from Hugging Face.

Breaking it down:

  • unsloth/ — Hugging Face repository owner
  • Qwen3.8-27B — approximately 27 billion parameters
  • GGUF — the model format used by llama.cpp
  • UD-Q4_K_XL — the quantization

Q4 means the model weights are approximately 4-bit quantized.

The trade-off is straightforward: lower precision produces a much smaller model and significantly reduces memory requirements, at some cost to numerical precision.


2. Speculative decoding / MTP

--spec-type draft-mtp
Enter fullscreen mode Exit fullscreen mode

This enables speculative decoding using MTP (Multi-Token Prediction).

Instead of having the main model generate:

token → token → token → token
Enter fullscreen mode Exit fullscreen mode

the system uses a draft mechanism to propose multiple future tokens, which the main model then verifies.

Conceptually:

                 Draft model
                      │
                      ▼
              token token token
                      │
                      ▼
                Main model
                  verifies
                      │
                      ▼
             accept several tokens
Enter fullscreen mode Exit fullscreen mode

When several proposed tokens are accepted, generation can become substantially faster.

For this model, speculative decoding is one of the most important performance-related settings.


3. --spec-default

--spec-default
Enter fullscreen mode Exit fullscreen mode

This enables the default speculative-decoding configuration associated with the selected speculation type.

In this case:

draft-mtp
Enter fullscreen mode Exit fullscreen mode

It's generally not something I'd change unless I was experimenting with the underlying speculative-decoding implementation.


4. Maximum speculative tokens

--spec-draft-n-max 8
Enter fullscreen mode Exit fullscreen mode

This controls the maximum number of speculative tokens proposed ahead.

With:

8
Enter fullscreen mode Exit fullscreen mode

the draft mechanism can attempt to predict up to eight tokens ahead.

Conceptually:

Main model:
A

Draft:
A → B → C → D → E → F → G → H

Main model verifies:
A B C D ✓ ✓ ✓ ✗
Enter fullscreen mode Exit fullscreen mode

The more tokens you speculate, the greater the potential speedup—but only if the draft predictions are good enough.

This is one of the parameters worth benchmarking:

2
4
8
Enter fullscreen mode Exit fullscreen mode

Eight is an aggressive but reasonable value to test.


5. GPU layers

-ngl 99
Enter fullscreen mode Exit fullscreen mode

This is short for:

--n-gpu-layers
Enter fullscreen mode Exit fullscreen mode

It specifies how many model layers should be offloaded to the GPU.

99 effectively means:

Put as many layers as possible on the GPU.

It does not mean "use 99 GPU cores."

Think of it as:

CPU
 │
 ├── some model layers
 │
GPU
 │
 └── most/all model layers
Enter fullscreen mode Exit fullscreen mode

If the model fits comfortably on the GPU, -ngl 99 is generally what you want for performance.


6. Context size

-c 524288
Enter fullscreen mode Exit fullscreen mode

This specifies the maximum context window.

The value is:

524,288 tokens = 512K tokens.

That is an enormous context window.

For comparison:

32K   = 32,768
128K  = 131,072
256K  = 262,144
512K  = 524,288
Enter fullscreen mode Exit fullscreen mode

The important trade-off is that larger context requires more memory, particularly because of the KV cache.

For agentic coding workloads, however, having hundreds of thousands of tokens available can be extremely useful.


7. Context-length override

--override-kv qwen2.context_length=int:524288
Enter fullscreen mode Exit fullscreen mode

This is different from -c.

You're overriding a value stored in the model's GGUF metadata:

qwen2.context_length
Enter fullscreen mode Exit fullscreen mode

and setting it to:

524288
Enter fullscreen mode Exit fullscreen mode

In other words, you're telling llama.cpp to treat the model as having a 512K context length.

This does not magically train the model for 512K context.

That's why the configuration also uses YaRN.


8. RoPE scaling

--rope-scaling yarn
Enter fullscreen mode Exit fullscreen mode

This enables YaRN — Yet another RoPE extension.

RoPE stands for Rotary Position Embedding.

RoPE is part of how the transformer represents token positions:

token 1
token 2
token 3
...
token 262144
Enter fullscreen mode Exit fullscreen mode

When extending the context beyond the model's original trained range, positional scaling is required.

YaRN provides a mechanism for extending that range.

In this configuration:

Original context:
262K

Target context:
524K
Enter fullscreen mode Exit fullscreen mode

So the positional range is being extended by roughly 2×.


9. Original YaRN context

--yarn-orig-ctx 262144
Enter fullscreen mode Exit fullscreen mode

This tells YaRN:

The model's original context length is 262,144 tokens.

So the relevant configuration is:

Original:
262,144

Target:
524,288

Extension:
2×
Enter fullscreen mode Exit fullscreen mode

These two parameters work together:

--rope-scaling yarn
--yarn-orig-ctx 262144
Enter fullscreen mode Exit fullscreen mode

10. Logical batch size

-b 16384
Enter fullscreen mode Exit fullscreen mode

This specifies the maximum number of tokens processed in a logical batch.

The value is:

16,384 tokens
Enter fullscreen mode Exit fullscreen mode

This primarily affects prompt processing / prefill.

For example, if you send a large prompt containing thousands of tokens, a larger batch can allow the GPU to process more tokens efficiently.

Larger batches can increase prompt-processing throughput, but they also consume more memory.

Importantly:

context = 524K
batch   = 16K
Enter fullscreen mode Exit fullscreen mode

is perfectly valid.

The batch size does not limit the context window.


11. Physical / micro batch

-ub 4096
Enter fullscreen mode Exit fullscreen mode

This is the physical or micro-batch size.

It controls how many tokens are actually processed at one time.

The configuration therefore has:

Logical batch:
16,384

Physical batch:
4,096
Enter fullscreen mode Exit fullscreen mode

Conceptually:

16,384 tokens

┌──────────────────┐
│ 4,096 tokens     │
├──────────────────┤
│ 4,096 tokens     │
├──────────────────┤
│ 4,096 tokens     │
├──────────────────┤
│ 4,096 tokens     │
└──────────────────┘
Enter fullscreen mode Exit fullscreen mode

This allows a large logical batch without requiring all 16K tokens to be processed simultaneously.

-ub is therefore particularly important for VRAM usage and prompt-processing performance.


12. CPU threads

-t 16
Enter fullscreen mode Exit fullscreen mode

This specifies the number of CPU threads used for computation.

Here:

16 CPU threads
Enter fullscreen mode Exit fullscreen mode

This does not mean 16 GPU cores.

How useful additional CPU threads are depends heavily on your CPU and on how much of the workload remains on the CPU.


13. Batch-processing CPU threads

-tb 16
Enter fullscreen mode Exit fullscreen mode

This specifies the number of CPU threads used specifically for batch processing.

So the configuration is:

Normal computation: 16 threads
Batch computation:  16 threads
Enter fullscreen mode Exit fullscreen mode

Whether 16 is optimal depends on your CPU.

If you're running on a high-core-count CPU, this is worth benchmarking.

More threads don't automatically mean higher performance.


14. Parallel sequences

-np 2
Enter fullscreen mode Exit fullscreen mode

This enables two parallel sequences/requests.

Conceptually:

                 Model
                   │
          ┌────────┴────────┐
          ▼                 ▼
      Context #1         Context #2
      512K max           512K max
Enter fullscreen mode Exit fullscreen mode

This is useful if you're running two concurrent requests or agents.

There is, however, a memory cost.

With:

512K context
×
2 parallel sequences
Enter fullscreen mode Exit fullscreen mode

the potential KV-cache requirement becomes very large.

If you only ever run one request at a time, -np 1 may provide a better memory/performance balance.


15. Flash Attention

-fa on
Enter fullscreen mode Exit fullscreen mode

This enables Flash Attention.

Flash Attention is an optimized implementation of the attention mechanism designed to reduce memory traffic and improve performance.

It becomes particularly important at long context lengths.

For a 512K configuration, I'd keep:

-fa on
Enter fullscreen mode Exit fullscreen mode

16. K cache

--cache-type-k f16
Enter fullscreen mode Exit fullscreen mode

This specifies the datatype used for the Key portion of the KV cache.

You're using:

F16
Enter fullscreen mode Exit fullscreen mode

or 16-bit floating point.


17. V cache

--cache-type-v f16
Enter fullscreen mode Exit fullscreen mode

This specifies the datatype used for the Value portion of the KV cache.

So the current configuration is:

K = F16
V = F16
Enter fullscreen mode Exit fullscreen mode

This provides high precision, but consumes considerably more memory than:

--cache-type-k q8_0
--cache-type-v q8_0
Enter fullscreen mode Exit fullscreen mode

Given the combination of:

512K context
×
2 parallel sequences
×
F16 KV
Enter fullscreen mode Exit fullscreen mode

this is one of the largest memory-consuming choices in the configuration.


18. KV GPU offload

--kv-offload
Enter fullscreen mode Exit fullscreen mode

This tells llama.cpp to keep the KV cache on the GPU when possible.

That generally improves performance because it avoids repeatedly moving KV data between CPU and GPU.

The desired architecture for maximum performance is therefore approximately:

Model weights → GPU
KV cache     → GPU
Attention    → GPU
Enter fullscreen mode Exit fullscreen mode

assuming you have enough VRAM.


19. Load mode

--load-mode none
Enter fullscreen mode Exit fullscreen mode

This controls the model-loading mechanism.

none means that no special loading mode is being selected.

This isn't a setting I'd normally spend much time optimizing unless you're diagnosing model loading, memory mapping, or startup behavior.


20. Host

--host 127.0.0.1
Enter fullscreen mode Exit fullscreen mode

This makes the server listen only on the local machine.

So the server is accessible through:

127.0.0.1
Enter fullscreen mode Exit fullscreen mode

but isn't directly exposed to other machines on the network.

This is a network/security setting, not an inference-performance setting.


21. Port

--port 8080
Enter fullscreen mode Exit fullscreen mode

The server listens on port:

8080
Enter fullscreen mode Exit fullscreen mode

So your local API is effectively:

http://127.0.0.1:8080
Enter fullscreen mode Exit fullscreen mode

This has essentially no impact on model performance.


Putting everything together

Your command is essentially saying:

Run Qwen 3.8 27B using a Q4 quantization, put as much of the model as possible on the GPU, use MTP speculative decoding with up to eight speculative tokens, support a 512K context by extending the model's 256K positional range with YaRN, process prompts using 16K/4K batches, use 16 CPU threads, support two simultaneous sequences, use Flash Attention, keep the F16 KV cache on the GPU, and expose the model as a local HTTP server on port 8080.

The architecture looks roughly like this:

                    llama.cpp server
                           │
               ┌───────────┴───────────┐
               │                       │
          Request #1              Request #2
          512K max                512K max
               │                       │
               └───────────┬───────────┘
                           │
                       KV Cache
                        F16/F16
                           │
                    Flash Attention
                           │
                  Qwen 3.8 27B Q4
                           │
                     GPU (-ngl 99)
                           │
                   MTP / speculative
                      decoding ×8
                           │
                         Output
Enter fullscreen mode Exit fullscreen mode

What matters most for performance

Not all parameters deserve equal attention.

Highest impact

--spec-draft-n-max 8
-c 524288
-np 2
--cache-type-k f16
--cache-type-v f16
-b 16384
-ub 4096
-fa on
Enter fullscreen mode Exit fullscreen mode

Hardware dependent

-t 16
-tb 16
Enter fullscreen mode Exit fullscreen mode

Usually leave alone

-ngl 99
--kv-offload
Enter fullscreen mode Exit fullscreen mode

Mostly configuration rather than performance

--override-kv
--rope-scaling
--yarn-orig-ctx
--load-mode
--host
--port
Enter fullscreen mode Exit fullscreen mode

The three biggest trade-offs in this particular configuration are:

512K context  ↔  memory

F16 KV       ↔  memory / performance

16K / 4K batch ↔ VRAM / prompt throughput
Enter fullscreen mode Exit fullscreen mode

And for generation speed, the most interesting parameter is probably:

MTP-8 ↔ speculative-token acceptance rate
Enter fullscreen mode Exit fullscreen mode

The optimal configuration therefore isn't necessarily the one with the largest numbers. The goal is to find the point where GPU utilization, memory bandwidth, KV-cache size, batch size, and speculative-token acceptance work together rather than competing with each other.

Top comments (0)