DEV Community

Cover image for Ornith-1.0-9B vs. Qwen3.5 vs. Gemma4: A Local LLM Battle Royale
bobby bonam
bobby bonam

Posted on Originally published at bbonam.hashnode.dev AI-assisted

Ornith-1.0-9B vs. Qwen3.5 vs. Gemma4: A Local LLM Battle Royale

Every few weeks a new model lands on Hugging Face with a specific claim: post-trained for agentic coding, tuned for tool use, optimized for terminal workflows. The benchmark numbers that come with these releases are real, but they're aggregate scores over huge, curated task sets. They tell you a model is good on average. They don't tell you whether the specific answer it just gave you is one you can trust.

The good part: checking that no longer requires a research lab. With Ollama and a GGUF file, you can run a purpose-built coding model on an ordinary laptop, put it next to the general-purpose model it was built from, and ask directly: what did the specialization actually buy you?

That's what this post does with Ornith-1.0-9B, DeepReinforce's new agentic-coding release, benchmarked locally against its own base model, Qwen3.5-9B, and a third reference point from a different family, Gemma4-12B.

What is Ornith-1.0-9B?

A 9B model post-trained specifically for agentic coding — terminal use, tool-calling, SWE-Bench-style tasks. MIT licensed, ships with <think> reasoning traces and structured <tool_call> output, and built directly on top of Qwen3.5-9B rather than a new architecture. DeepReinforce reports 69.4% on SWE-Bench Verified.

That last part is what makes this worth testing: Ornith isn't a different base model, it's Qwen3.5-9B after specialized post-training. So instead of arguing about two unrelated architectures in the abstract, you can put the fine-tune next to its own untouched parent, same hardware, same prompts, and see exactly what changed.

Setup

CPU-only, Windows, 32GB RAM, no GPU — a normal dev laptop, not a benchmarking rig. Same quant level (Q4_K_M) across all three models so nobody gets a precision advantage.

ollama pull hf.co/ornith-ai/Ornith-1.0-9B-GGUF:Q4_K_M
ollama pull qwen3.5:9b-q4_K_M
# gemma4-custom:latest already installed from a prior setup
Enter fullscreen mode Exit fullscreen mode

Two things worth flagging before the results:

Gemma4's context window. My gemma4-custom Modelfile bakes in a 262,144-token context. On CPU that's expensive to allocate regardless of prompt length, and it's the reason Gemma4 is slower on every single task below — that's a config choice on my end, not a capability signal.

Reasoning traces are a separate field. Ollama 0.35 splits a model's <think> output into its own thinking field, distinct from content, and several of these models reason by default. First pass through my harness, I capped num_predict at 700 without disabling reasoning — every model spent the whole budget thinking and returned an empty final answer. Lesson: check done_reason, don't assume empty content means the model failed. For the actual scored comparison below, I ran with think: false so I was judging final answers, not internal monologue:

$body = @{
    model    = $modelTag
    messages = @(@{ role = "user"; content = $prompt })
    stream   = $false
    think    = $false
    options  = @{ temperature = 0.2; num_predict = 700 }
} | ConvertTo-Json -Depth 5

Invoke-RestMethod -Uri "http://localhost:11434/api/chat" -Method Post `
    -Body $body -ContentType "application/json; charset=utf-8"
Enter fullscreen mode Exit fullscreen mode

One more discipline point: each model was loaded, run through all five tasks, then explicitly unloaded before the next one started. Letting two CPU-bound models sit resident together skews timing — in one test run it added 50+ seconds of reload overhead to whichever model wasn't already warm.

The tasks

Five identical prompts, same order, temperature 0.2, across all three models.

Prompt: fix the bug

def max_subarray_sum(nums, k):
    max_sum = 0
    for i in range(len(nums) - k):
        window_sum = sum(nums[i:i+k])
        max_sum = max(max_sum, window_sum)
    return max_sum
Enter fullscreen mode Exit fullscreen mode

All three correctly diagnosed the off-by-one (range(len(nums) - k) excludes the final window). Clean sweep on diagnosis. Not on the fix.

Gemma4 and Qwen3.5 both rewrote it as a plain re-summed sliding window with max_sum = float('-inf') — correct for any input. Ornith rewrote it as a rolling-sum optimization:

def max_subarray_sum(nums, k):  # Ornith's fix
    max_sum = sum(nums[:k])
    for i in range(k, len(nums)):
        max_sum = max(max_sum, max_sum + nums[i] - nums[i - k])
    return max_sum
Enter fullscreen mode Exit fullscreen mode

Smart idea, subtle bug: it collapses "current window sum" and "best sum so far" into one variable, which only stays correct if the best window is always the most recent one. I hand-traced nums = [10, 1, 1, 1, 1, 10], k=2 — true answer is 11, Ornith's fix returns 20.

It passed the exact example in the prompt and failed on a case it wasn't shown. If you're handing this class of model real code changes: verify the fix, not just the diagnosis.

Prompt: one-line shell command

Find all files under /var/log larger than 50MB, print path + human-readable size, sorted descending.

Gemma4: find /var/log -type f -size +50M -exec du -h {} + | sort -hr — handles filenames with spaces correctly, since du -h {} + never text-splits the argument.

Ornith and Qwen3.5 both reached for a find -printf/-stat pipe into awk, which silently truncates any filename containing a space (awk's default field-splitting has no concept of "one filename = one token"). Same bug, same root cause, in the base model and its agentic fine-tune — this shell-scripting habit wasn't something Ornith's post-training corrected.

Prompt: strict JSON extraction

Extract into JSON: "Maria Gonzalez, 34 years old, lives in Austin." Keys: name, age, city. No commentary, no fences.

All three got the values right, but Ornith's answer was the tightest: {"name":"Maria Gonzalez","age":34,"city":"Austin"} in 16 tokens, fastest of the three. Qwen3.5 used 21 tokens, Gemma4 used 34 (pretty-printed). If you're piping this into another program, that terseness is a real advantage — and it's the one place Ornith's "built for tool-calling pipelines, not chat" design goal showed up clearly.

Prompt: math logic

35 heads, 94 legs, chickens and rabbits, how many of each?

Clean tie — all three got chickens=23, rabbits=12, correctly reasoned.

Prompt: general knowledge

Explain a bloom filter in 3 sentences, one real-world use case.

Another tie on substance. Ornith's answer was the most technically specific (named the actual mechanism — bit arrays + multiple hash functions — where the others stayed conceptual).

What reasoning mode actually costs

Curious what Ornith's <think> trace looks like, I ran the farmer problem again with reasoning back on. It plans, solves, and explicitly re-verifies its own arithmetic before answering:

Verify:
  Heads: 23 + 12 = 35 (Correct)
  Legs: 23×2 + 12×4 = 46 + 48 = 94 (Correct)
Final Review: Does it meet all constraints? Yes.
Enter fullscreen mode Exit fullscreen mode

Genuinely useful self-check habit. The cost: 768 tokens of reasoning vs. ~220 for a direct answer to the same question — roughly a 3.5x token tax. Worth it when correctness matters more than latency, not worth it for a quick lookup.

Scorecard

Task Ornith-1.0-9B Qwen3.5-9B Gemma4-12B
Bug diagnosis correct correct correct
Bug fix fails on general input correct correct
Shell one-liner breaks on spaces in filenames breaks on spaces in filenames correct
Math logic correct correct correct
JSON output correct, most compact (16 tok) correct correct
General knowledge correct, most detailed correct correct
Avg tok/s (CPU) 4.02 4.16 3.20*

*Gemma4's number reflects its 262K-context config, not raw capability.

Takeaway

Ornith-1.0-9B is a real, specialized release, not hype over an unchanged base model — its terseness and structured-output precision are genuine, measurable wins for agentic/tool-calling use. But on this small test, its post-training didn't make it more reliably correct than its own untouched base model: it matched Qwen3.5 and Gemma4 on reasoning and knowledge, and on the one task where "agentic coding" should matter most, it diagnosed the bug correctly and then shipped a fix that was subtly wrong.

Published benchmark numbers are real signals about average performance — not a guarantee about the specific diff in front of you. Read it. Run the edge case it wasn't shown.

Ran entirely locally via Ollama — happy to share exact prompts/outputs if anyone wants to reproduce or push back on this.

Top comments (0)