<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: bobby bonam</title>
    <description>The latest articles on DEV Community by bobby bonam (@bobbybonam).</description>
    <link>https://dev.to/bobbybonam</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4152995%2Fc67fb9ba-6d78-46f1-a273-81f7ae0b77ce.png</url>
      <title>DEV Community: bobby bonam</title>
      <link>https://dev.to/bobbybonam</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bobbybonam"/>
    <language>en</language>
    <item>
      <title>Ornith-1.0-9B vs. Qwen3.5 vs. Gemma4: A Local LLM Battle Royale</title>
      <dc:creator>bobby bonam</dc:creator>
      <pubDate>Wed, 30 Sep 2026 19:18:02 +0000</pubDate>
      <link>https://dev.to/bobbybonam/ornith-10-9b-vs-qwen35-vs-gemma4-a-local-ai-battle-royale-3b9e</link>
      <guid>https://dev.to/bobbybonam/ornith-10-9b-vs-qwen35-vs-gemma4-a-local-ai-battle-royale-3b9e</guid>
      <description>&lt;p&gt;Every few weeks a new model lands on Hugging Face with a specific claim: post-trained for agentic coding, tuned for tool use, optimized for terminal workflows. The benchmark numbers that come with these releases are real, but they're aggregate scores over huge, curated task sets. They tell you a model is good on average. They don't tell you whether the specific answer it just gave &lt;em&gt;you&lt;/em&gt; is one you can trust.&lt;/p&gt;

&lt;p&gt;The good part: checking that no longer requires a research lab. With Ollama and a GGUF file, you can run a purpose-built coding model on an ordinary laptop, put it next to the general-purpose model it was built from, and ask directly: what did the specialization actually buy you?&lt;/p&gt;

&lt;p&gt;That's what this post does with &lt;strong&gt;Ornith-1.0-9B&lt;/strong&gt;, DeepReinforce's new agentic-coding release, benchmarked locally against its own base model, &lt;strong&gt;Qwen3.5-9B&lt;/strong&gt;, and a third reference point from a different family, &lt;strong&gt;Gemma4-12B&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Ornith-1.0-9B?
&lt;/h2&gt;

&lt;p&gt;A 9B model post-trained specifically for agentic coding — terminal use, tool-calling, SWE-Bench-style tasks. MIT licensed, ships with &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; reasoning traces and structured &lt;code&gt;&amp;lt;tool_call&amp;gt;&lt;/code&gt; output, and built directly on top of Qwen3.5-9B rather than a new architecture. DeepReinforce reports 69.4% on SWE-Bench Verified.&lt;/p&gt;

&lt;p&gt;That last part is what makes this worth testing: Ornith isn't a different base model, it's Qwen3.5-9B &lt;em&gt;after&lt;/em&gt; specialized post-training. So instead of arguing about two unrelated architectures in the abstract, you can put the fine-tune next to its own untouched parent, same hardware, same prompts, and see exactly what changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;p&gt;CPU-only, Windows, 32GB RAM, no GPU — a normal dev laptop, not a benchmarking rig. Same quant level (Q4_K_M) across all three models so nobody gets a precision advantage.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull hf.co/ornith-ai/Ornith-1.0-9B-GGUF:Q4_K_M
ollama pull qwen3.5:9b-q4_K_M
&lt;span class="c"&gt;# gemma4-custom:latest already installed from a prior setup&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things worth flagging before the results:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gemma4's context window.&lt;/strong&gt; My gemma4-custom Modelfile bakes in a 262,144-token context. On CPU that's expensive to allocate regardless of prompt length, and it's the reason Gemma4 is slower on every single task below — that's a config choice on my end, not a capability signal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reasoning traces are a separate field.&lt;/strong&gt; Ollama 0.35 splits a model's &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; output into its own &lt;code&gt;thinking&lt;/code&gt; field, distinct from &lt;code&gt;content&lt;/code&gt;, and several of these models reason by default. First pass through my harness, I capped &lt;code&gt;num_predict&lt;/code&gt; at 700 without disabling reasoning — every model spent the whole budget thinking and returned an &lt;em&gt;empty&lt;/em&gt; final answer. Lesson: check &lt;code&gt;done_reason&lt;/code&gt;, don't assume empty &lt;code&gt;content&lt;/code&gt; means the model failed. For the actual scored comparison below, I ran with &lt;code&gt;think: false&lt;/code&gt; so I was judging final answers, not internal monologue:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$body&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;@{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$modelTag&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;@(@{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;role&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$prompt&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="bp"&gt;$false&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nx"&gt;think&lt;/span&gt;&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="bp"&gt;$false&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nx"&gt;options&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;@{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;temperature&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;num_predict&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;700&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;ConvertTo-Json&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Depth&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;5&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;Invoke-RestMethod&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Uri&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://localhost:11434/api/chat"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Method&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;Post&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="nt"&gt;-Body&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$body&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-ContentType&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"application/json; charset=utf-8"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One more discipline point: each model was loaded, run through all five tasks, then explicitly unloaded before the next one started. Letting two CPU-bound models sit resident together skews timing — in one test run it added 50+ seconds of reload overhead to whichever model wasn't already warm.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tasks
&lt;/h2&gt;

&lt;p&gt;Five identical prompts, same order, temperature 0.2, across all three models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt: fix the bug&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;max_subarray_sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nums&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;max_sum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nums&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;window_sum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nums&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;max_sum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_sum&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;window_sum&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;max_sum&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three correctly diagnosed the off-by-one (&lt;code&gt;range(len(nums) - k)&lt;/code&gt; excludes the final window). Clean sweep on diagnosis. Not on the fix.&lt;/p&gt;

&lt;p&gt;Gemma4 and Qwen3.5 both rewrote it as a plain re-summed sliding window with &lt;code&gt;max_sum = float('-inf')&lt;/code&gt; — correct for any input. Ornith rewrote it as a rolling-sum optimization:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;max_subarray_sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nums&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;  &lt;span class="c1"&gt;# Ornith's fix
&lt;/span&gt;    &lt;span class="n"&gt;max_sum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nums&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nums&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
        &lt;span class="n"&gt;max_sum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_sum&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_sum&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;nums&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;nums&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;max_sum&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Smart idea, subtle bug: it collapses "current window sum" and "best sum so far" into one variable, which only stays correct if the best window is always the most recent one. I hand-traced &lt;code&gt;nums = [10, 1, 1, 1, 1, 10], k=2&lt;/code&gt; — true answer is 11, Ornith's fix returns &lt;strong&gt;20&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It passed the exact example in the prompt and failed on a case it wasn't shown. If you're handing this class of model real code changes: verify the fix, not just the diagnosis.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt: one-line shell command&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Find all files under /var/log larger than 50MB, print path + human-readable size, sorted descending.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Gemma4: &lt;code&gt;find /var/log -type f -size +50M -exec du -h {} + | sort -hr&lt;/code&gt; — handles filenames with spaces correctly, since &lt;code&gt;du -h {} +&lt;/code&gt; never text-splits the argument.&lt;/p&gt;

&lt;p&gt;Ornith and Qwen3.5 both reached for a &lt;code&gt;find -printf&lt;/code&gt;/&lt;code&gt;-stat&lt;/code&gt; pipe into &lt;code&gt;awk&lt;/code&gt;, which silently truncates any filename containing a space (awk's default field-splitting has no concept of "one filename = one token"). Same bug, same root cause, in the base model &lt;em&gt;and&lt;/em&gt; its agentic fine-tune — this shell-scripting habit wasn't something Ornith's post-training corrected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt: strict JSON extraction&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Extract into JSON: "Maria Gonzalez, 34 years old, lives in Austin." Keys: name, age, city. No commentary, no fences.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;All three got the values right, but Ornith's answer was the tightest: &lt;code&gt;{"name":"Maria Gonzalez","age":34,"city":"Austin"}&lt;/code&gt; in 16 tokens, fastest of the three. Qwen3.5 used 21 tokens, Gemma4 used 34 (pretty-printed). If you're piping this into another program, that terseness is a real advantage — and it's the one place Ornith's "built for tool-calling pipelines, not chat" design goal showed up clearly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt: math logic&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;35 heads, 94 legs, chickens and rabbits, how many of each?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Clean tie — all three got chickens=23, rabbits=12, correctly reasoned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt: general knowledge&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Explain a bloom filter in 3 sentences, one real-world use case.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Another tie on substance. Ornith's answer was the most technically specific (named the actual mechanism — bit arrays + multiple hash functions — where the others stayed conceptual).&lt;/p&gt;

&lt;h2&gt;
  
  
  What reasoning mode actually costs
&lt;/h2&gt;

&lt;p&gt;Curious what Ornith's &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; trace looks like, I ran the farmer problem again with reasoning back on. It plans, solves, and explicitly re-verifies its own arithmetic before answering:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Verify:
  Heads: 23 + 12 = 35 (Correct)
  Legs: 23×2 + 12×4 = 46 + 48 = 94 (Correct)
Final Review: Does it meet all constraints? Yes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Genuinely useful self-check habit. The cost: 768 tokens of reasoning vs. ~220 for a direct answer to the same question — roughly a 3.5x token tax. Worth it when correctness matters more than latency, not worth it for a quick lookup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scorecard
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Ornith-1.0-9B&lt;/th&gt;
&lt;th&gt;Qwen3.5-9B&lt;/th&gt;
&lt;th&gt;Gemma4-12B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bug diagnosis&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bug fix&lt;/td&gt;
&lt;td&gt;fails on general input&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shell one-liner&lt;/td&gt;
&lt;td&gt;breaks on spaces in filenames&lt;/td&gt;
&lt;td&gt;breaks on spaces in filenames&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math logic&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JSON output&lt;/td&gt;
&lt;td&gt;correct, most compact (16 tok)&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General knowledge&lt;/td&gt;
&lt;td&gt;correct, most detailed&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avg tok/s (CPU)&lt;/td&gt;
&lt;td&gt;4.02&lt;/td&gt;
&lt;td&gt;4.16&lt;/td&gt;
&lt;td&gt;3.20*&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;*Gemma4's number reflects its 262K-context config, not raw capability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;Ornith-1.0-9B is a real, specialized release, not hype over an unchanged base model — its terseness and structured-output precision are genuine, measurable wins for agentic/tool-calling use. But on this small test, its post-training didn't make it more &lt;em&gt;reliably correct&lt;/em&gt; than its own untouched base model: it matched Qwen3.5 and Gemma4 on reasoning and knowledge, and on the one task where "agentic coding" should matter most, it diagnosed the bug correctly and then shipped a fix that was subtly wrong.&lt;/p&gt;

&lt;p&gt;Published benchmark numbers are real signals about average performance — not a guarantee about the specific diff in front of you. Read it. Run the edge case it wasn't shown.&lt;/p&gt;

&lt;p&gt;Ran entirely locally via Ollama — happy to share exact prompts/outputs if anyone wants to reproduce or push back on this.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>llm</category>
      <category>deeplearning</category>
    </item>
  </channel>
</rss>
