DEV Community

finaltype
finaltype

Posted on Originally published at finaltype.github.io

Running a 35B MoE Model on the R9700 - I Measured the Backend, Concurrency, and MTP One at a Time

Vulkan won, four concurrent slots was the knee, and MTP only helped with a single request

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

In my R9700 purchase review(new tab) I promised to write up the results separately once I had them. This post is that promise.

I measured which setup was fastest when running a 35B MoE model (Qwen3.6-35B-A3B, about 3B active parameters) on the R9700. It's currently dedicated to a backup path rather than serving as a main card, so I've verified speed only and haven't yet measured the quality of its outputs.

First, what is being compared with what

The same R9700 produces completely different numbers depending on the model. The R9700 figure of 56 t/s in my earlier speed comparison post(new tab) was measured by someone else on a 27B dense model. Every number in this post is for the 35B MoE, so the two must not be compared directly.

Also, I standardized time per stock on throughput basis: wall-clock time divided by the number of stocks. I'll get to the pitfalls that muddy this below.

Workload and sampling

It's a multi-agent pipeline where analyzing one stock takes 17 to 18 LLM calls. The 6-stock replay I used as a tuning fixture was 105 calls with 664,557 prompt tokens in total (about 6.3k per call on average).

Prompt length follows the production statistics I published in the 125B serving post(new tab), since it's the same pipeline: about 8.3k on average and about 22k at the maximum. The input-to-output token ratio is roughly 6.6:1, so prefill weighs heavily. I capped generation at 4,096 tokens per call.

Sampling is temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5. This profile originally ran with sampling unspecified, and the first sweep on 09-04 did. The 09-24 retune ran with the values above set explicitly.

I suspected "maybe outputs came out shorter because sampling was unspecified" and compared the two cases. Generation length per call was practically identical, 1,458 versus 1,436 tokens. I rejected that hypothesis.

Setup

The model is a GGUF Q4_K_M (about 22GB), with the KV cache at q8_0, flash attention on, and batch settings -b 4096 -ub 2048. Serving is llama.cpp, and I built both the Vulkan and HIP (ROCm) backends to compare.

I didn't upgrade ROCm; I used the Ubuntu package 7.1.x as it was. Since Vulkan was the first-choice candidate, I judged a ROCm upgrade unnecessary.

I put a 210W power cap on the card, for fan noise reasons. It shares the machine with the 5090.

Vulkan won on backend

Under identical conditions (3 concurrent slots, MTP off), Vulkan took 188.5 seconds per stock and HIP took 256.5. HIP was 1.36x slower. Public benchmarks suggest "HIP is strong at prefill, Vulkan at decode," but on my actual workload (feeding long prompts several times), Vulkan led across the board.

That comparison was even made with the Vulkan side on source that was seven weeks older. Retesting on the latest build pointed the same direction, with HIP 1.3 to 1.4x slower at every concurrency level.

Four concurrent slots was the knee

I varied the number of requests processed at once (np) and measured aggregate decode throughput, on the latest Vulkan build.

Concurrent slots Aggregate decode (t/s)
1 108
2 143
3 163
4 185
6 175
8 172

It bends at 4. Six and eight slots gained nothing or even lost throughput, while the latency for one stock to finish grew almost in proportion to slot count.

In a replay experiment on frozen production input, 4 slots took 172.4 seconds per stock across 12 stocks, about 4.79 hours extrapolated to 100 stocks. Eight slots took 177.6 seconds — no gain — and also used twice the context.

MTP only helped with a single request

MTP (speculative decoding) had two faces on this card too. With one slot it did well, at roughly 133 tokens per second. But with two or more concurrent slots it was useless or harmful.

With MTP on at 3 slots, 188.5 seconds per stock became 297.0, 1.58x slower. Speculative-token acceptance stayed around 50% regardless of backend or concurrency.

The reason is a guess. In a MoE with few active parameters, the tokens to verify for each concurrent request seem to call on different experts, so the verification batch gets more expensive in proportion to the number of requests.

Two things I corrected

First, I revised my "3 slots is best" conclusion to 4. In my first sweep, the 4-slot experiment never finished, so I reported the best of the completed values, 3, as the optimum. I had written down a provisional number as if it were a conclusion.

Second, the premise "MTP acceptance is 77 to 95% on HIP" from a consultation was a misreading. The original report said concurrent requests collapse on a 27B dense model, and I read it backwards. That premise led to the idea of "HIP plus MTP plus 8 slots," which turned out to be useless when measured. I corrected the misreading afterward.

Where I got badly confused in measurement

There were two kinds of seconds per stock. One is throughput basis (wall-clock ÷ stock count). The other is slot-latency basis (how long each slot took to finish one stock), which is the former multiplied by the number of concurrent slots.

I nearly concluded "the new build regressed" by putting 188.5 seconds and 504 seconds side by side. Put on the same basis, the new build was actually 11% faster. Prefill went from 2,476 to 2,884 tokens per second, and decode was unchanged.

Runaway outputs shook the timing. Calls that hit the generation cap made one stock take twice as long, causing about 11% run-to-run variation. So I didn't rely on wall-clock alone; I looked at aggregate decode and prefill throughput alongside it.

A port overlap contaminated a sample. One client from the 6-slot experiment attached to the 8-slot server on the same port as a ninth client. The server didn't reject it even though the model alias differed, so I threw that sample out.

Instantaneous speed swung between 0.6 and 114 at 8 slots. It wasn't a bug; it was structural behavior where chunks from long-prompt slots occupy the batch. So I judged only by seconds per stock, not instantaneous values.

I didn't measure quality

The output quality of this combination (35B MoE on the R9700) is zero measurements. All I've adopted is a speed setting.

Starting September 28, this backup path scores the same 100 stocks every night alongside the main path (27B dense on the 5090). It's a pre-registered experiment that stacks up the two paths' forward RankIC side by side and compares them after four to six weeks. Since a single run's grades carry heavy reproducibility noise, I use the average score from several draws per stock.

This model's self-reproducibility I already published in an earlier post(new tab), from running the same stocks three times (QWK average 0.12, grade agreement 56 to 61%). A model being fast and its signal being trustworthy are two different questions.

Compared with the 5090

The same model runs about 6.8x slower on the R9700 than on the 5090 with vLLM. The cards have different characters, so that's no surprise. That's why I made the R9700 a backup path rather than the main card.

What I didn't try or couldn't

  • Quantization above Q5_K_M has no measurements; I only reasoned about it from the literature
  • I didn't run a -b 16384 batch or -ub sweeps beyond the 4-slot config
  • I excluded the ROCm builds of vLLM and SGLang from the start, since too much wasn't verified on this card
  • Retesting MTP on Vulkan at two or more slots was only proposed, not run

Summary

The backup path's setup right now is Vulkan, 4 slots, MTP off. That's about 172 seconds per stock, roughly 4.8 hours for 100 stocks, and the figure is a provisional value from a 12-stock replay experiment.

I've measured speed, but I've never once measured the output quality of this combination. That comes at the next stage, where I'll accumulate it side by side with the main path and compare. My record of tuning a 27B model on a single 5090 is in a separate post(new tab).

The story of serving a 125B model across the 5090 and R9700, one part on each card, is in the earlier postmortem(new tab).

Top comments (0)