How I cut a 100-stock nightly batch from six-plus hours to under three, and why you shouldn't fully trust that number
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
I decided to use a 27B dense model for the nightly 100-stock analysis batch. I ran experiments on how fast it can go on a single RTX 5090, and this post summarizes where I've got to.
The short version: I cut the production baseline of about 226 seconds per stock to about 102 seconds in a replay experiment. But that 102 is a single repetition on 24 stocks, so it isn't a production measurement yet. I'll write out those limits too.
What this batch actually does
For anyone who wants to reproduce the numbers, here's the workload first.
Analyzing one stock takes 17 LLM calls (a multi-agent pipeline). The per-node averages look like this.
| Node | Avg input | Avg output | Avg latency (3 slots) |
|---|---|---|---|
| Market analyst | 13.3k tokens | 3.0k tokens | 68.4 s |
| Sentiment analyst | 2.95k | 0.87k | 19.7 s |
| Portfolio manager | 9.2k | 0.52k | 12.9 s |
| Trader | 1.2k | 0.35k | 7.9 s |
The total per stock is about 188k input tokens and 25.8k output tokens. Breaking down the time, prefill is about 38 seconds and decode about 540 seconds, so decode is 93%. That's why I judged there was little room in prefill optimization.
The node with the longest input is the Conservative Analyst. The largest input observed over nine production nights was 33,487 tokens (Bear 33,451, Bull 32,699), which is why I set the server's maximum length to 34,816. That's already tight, so I can't shrink it further. Anything over it fails the call with a 400 error.
Sampling settings
Sampling is temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5 — the model card's recommended values for non-thinking (reasoning-off) mode, used as-is.
The nine production nights before 09-24 were different. I had overridden only the temperature to 0.8 on top of the checkpoint defaults (1.0/0.95/20), so the effective sampling was 0.8 / 0.95 / 20 / presence 0. In effect I changed three axes at once.
The effect shows up most clearly in the output tail.
| 9 production nights (29,630 calls) | New-sampling blind run (1,700 calls) | |
|---|---|---|
| Output p50 / p90 / p99 | 1,530 / 2,837 / 5,148 tokens | 1,337 / 2,670 / 4,402 tokens |
| Calls over 6,000 tokens | 104 (0.35%) | 2 (0.12%) |
| Max output | 26,518 tokens | 9,291 tokens |
The presence penalty appears to have clipped the runaway tail. But since three axes changed at once, I couldn't separate which one deserves the credit.
The starting point was 21.7 hours on an XTX
I first ran this model on a single 7900 XTX(new tab) two months ago. On llama.cpp with 4-bit quantization (Q4_K_M), 100 stocks took around 30 hours.
Attaching a separate MTP (speculative decoding) head brought that down to about 21.7 hours, thanks to decode speed going from the mid-30s to the high-50s of tokens per second. That was still more than double my target of staying within roughly 10 hours, the window that fits overnight.
I learned one thing then: comparing by wall-clock time alone gives wrong answers. One run was over-measured by more than 3 hours simply because it happened to produce longer outputs. After that I compared arms by prefill/decode throughput instead of wall-clock time.
Moving to the 5090 and changing the stack
Moving to the 5090, I also switched the serving stack from llama.cpp to vLLM. I use a 4-bit floating-point (NVFP4) checkpoint and compress the KV cache to fp8.
The first production setting had 3 concurrent slots. Nine nights of production running on it measured about 226 seconds per stock, roughly 6.3 hours for 100 stocks. It cleared the target, but without much margin.
Before that, my first attempt with 14 concurrent slots had failed on memory. Back then I only looked at the number and thought, "more slots will make it faster."
Raising concurrency made it 1.77x faster
This time I froze the input (a replay experiment) and changed only the settings. The 3-slot baseline ran all 100 stocks at 204.5 seconds per stock (5.68 hours).
The 14-slot run took 115.8 seconds per stock (3.22 hours) on 28 stocks — 1.77x faster than baseline. Aggregate throughput went from 126 to 233 tokens per second.
But the server logs looked odd. The number of requests actually running at once wasn't 14 — it was 6 to 9. The queue always held 5 to 9 waiting requests, and per-call latency went from a median of 32 seconds to 70.
The bottleneck was KV cache capacity, not slot count
The cause was KV cache capacity. Even with 14 slots configured, once the cache fills up, new requests can't get in. Cache utilization hit 100% at peak.
This model isn't purely dense; it's a hybrid. Of its 64 layers, 48 use linear attention and only 16 use full attention, so KV accounting differs from a plain dense model. That's where the intuition "14 slots = 14x concurrency" breaks.
So the concurrency number I configure is meaningless; what matters is the concurrency actually realized. If I hadn't looked at the number of running requests in the log, I would have wrongly recorded this as "14-way, 1.77x."
Adding MTP doubled it
Next was speculative decoding. This time I turned on the MTP head built into the model, in vLLM, with 2 speculative tokens. At the same time I reduced the slots to 8 and raised GPU memory utilization to 0.95.
Running 24 stocks came to 102.0 seconds per stock (2.83 hours). That's 2.0x the baseline, and 12% faster than the 14-slot run without MTP. Aggregate throughput peaked at 437 tokens per second in the logs, with speculative-token acceptance of 65 to 79% depending on the window.
In this run too, the configuration said 8 slots but realized concurrency was 4 to 6, and KV utilization hit 100% at peak. The bottleneck is still the KV cache. Breaking down the time, decode is about 93% of it, so there's little room in prefill-side optimizations.
Why you shouldn't trust this number
The setting I've adopted is this 102 seconds. But the number has clear limits.
- Three axes changed at once. MTP, slot count, and memory utilization changed together, so MTP's own contribution can't be isolated.
- The sample is small. It's 24 stocks with one repetition. The 2.83 hours for 100 stocks is an extrapolation.
- I only verified it in the replay harness. I haven't run it on the actual production path even once. The first production night is next week.
- I couldn't compare output quality. As the "Quality numbers" section below shows, even running the same input twice gives a low rate of overlapping grades, so a single run can't say anything about quality differences. The direction of the speedup is safe, but quality is a separate question.
I also made one measurement mistake. Seconds per stock already reflects concurrency since it's wall-clock divided by stock count, yet I was about to divide by the slot count again. Dividing twice overstates speed by several times. Luckily review caught it.
Quality numbers
A speedup is meaningless if the output got worse, so here's what I know about quality too. The short version: I haven't been able to confirm quality yet.
- Grade distribution: the new-sampling blind run (100 stocks) gave Hold 66, Underweight 19, Overweight 15. In the earlier era Hold was 41 to 61%. Whether that's due to the sampling change or run noise can't be known without repeated runs.
- Reproducibility: with the same input, the same checkpoint, and the same sampling, the 3-slot run and the 14-slot run scored the same 28 stocks. Grades matched on 14 of 28 (50%), and the Spearman correlation of scores was 0.12 (p=0.54). The two runs' grade distributions (Hold 20/5/3 vs 20/4/4) are practically identical.
- Which stocks get a non-Hold grade: breaking down the 15 non-zero pairs, only 1 had both runs non-zero with the same sign, 0 had opposite signs, and 14 had a zero on one side. The distribution is the same, but which stocks receive the grade is nearly random from run to run. The self-reproducibility correlation I'd measured earlier was 0.04 to 0.53, with grade flips of 18.6%.
- Relation to forward returns: forward RankIC for this combination (27B on the 5090) is zero measurements. In the ranges I've measured under other regimes, the correlation (IC) between this pipeline's grades and forward returns was indistinguishable from zero, on an extremely small sample.
So after speeding up the batch, my plan is to draw each stock several times and use the average score. Measuring IC from a single run's grades would let noise swallow the statistical power.
What I tried and dropped
- Prefix caching: left off, because of isolation issues and a small upper bound from the time breakdown
- Shrinking the context: impossible, since the largest production input is 33,487 tokens and 34,816 is already the limit
- 16 or 24 slots: pointless, since the KV cache was already the ceiling
- Self-quantization (INT4, FP8 variants): put on hold because it either cut into KV headroom on the 5090 or hadn't been quality-verified
What's next
On the first production night next week I'll measure the actual 100-stock time and see how much it differs from the replay experiment's 2.83 hours. That difference is the real conclusion of this post.
I plan to spend the leftover time on drawing each stock multiple times and averaging. I'll write a follow-up once the production result is in.
My record of tuning a 35B MoE model on the R9700 the same way is in the next post(new tab). The conclusions, from backend to concurrency to MTP, came out quite different.
Top comments (0)