<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Matheus Feijão </title>
    <description>The latest articles on DEV Community by Matheus Feijão  (@beanstechbr).</description>
    <link>https://dev.to/beanstechbr</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158792%2F111669d6-78cd-40a3-915a-e1a20945a5b3.jpeg</url>
      <title>DEV Community: Matheus Feijão </title>
      <link>https://dev.to/beanstechbr</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/beanstechbr"/>
    <language>en</language>
    <item>
      <title>Llama Guard is not a guardrail.</title>
      <dc:creator>Matheus Feijão </dc:creator>
      <pubDate>Sat, 03 Oct 2026 15:37:27 +0000</pubDate>
      <link>https://dev.to/beanstechbr/llama-guard-is-not-a-guardrail-4l2c</link>
      <guid>https://dev.to/beanstechbr/llama-guard-is-not-a-guardrail-4l2c</guid>
      <description>&lt;p&gt;'"We put Granite Guardian and Llama Guard side by side for regulated Brazilian Portuguese. One worked out of the box. The other needed a fine-tuning project before it was useful."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://g.cloud/blog/en/granite-guardian-vs-proprietarios/" rel="noopener noreferrer"&gt;https://g.cloud/blog/en/granite-guardian-vs-proprietarios/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every team shipping LLM features in 2026 eventually asks the same question: &lt;em&gt;what sits between the model and the user?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Meta and IBM both answer it with open-weights safety classifiers — &lt;strong&gt;Llama Guard&lt;/strong&gt; and &lt;strong&gt;Granite Guardian&lt;/strong&gt;. They look interchangeable on a feature matrix. They are not. We learned that the hard way while building &lt;a href="https://g.cloud" rel="noopener noreferrer"&gt;g.cloud&lt;/a&gt;, a guardrail for regulated Brazilian industries (legal, healthcare, finance, government), where a hallucinated citation isn't a bug — it's professional misconduct.&lt;/p&gt;

&lt;p&gt;Here's the honest comparison, including where Llama Guard wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  The taxonomy problem
&lt;/h2&gt;

&lt;p&gt;Llama Guard classifies against Meta's safety taxonomy (MLCommons categories: violence, sexual content, hate, self-harm, etc.). It does this well — for &lt;em&gt;content safety&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;But in a regulated domain, the dangerous output is usually &lt;strong&gt;not&lt;/strong&gt; toxic. It's &lt;em&gt;plausible and wrong&lt;/em&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a lawyer's assistant citing &lt;code&gt;REsp 1.234.567&lt;/code&gt; — a case that doesn't exist;&lt;/li&gt;
&lt;li&gt;a health bot describing a dosage the retrieved document never mentioned;&lt;/li&gt;
&lt;li&gt;a banking assistant politely explaining data that is under bank-secrecy law (in Brazil, LC 105/2001).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these trip a hate/violence/sexual classifier. This is where &lt;strong&gt;Granite Guardian&lt;/strong&gt; plays a different game. Beyond harm categories, it ships detectors for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Groundedness&lt;/strong&gt; — is the response actually supported by the retrieved context? (the RAG hallucination rail)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rationale&lt;/strong&gt; — does the model's justification hold up?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relevance and bias&lt;/strong&gt; — did it drift off-scope or encode a stereotype?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the difference between a &lt;em&gt;moderation filter&lt;/em&gt; and a &lt;em&gt;guardrail&lt;/em&gt;. Moderation asks "is this harmful?" A guardrail asks "is this &lt;strong&gt;true, in-scope, and allowed to leave the perimeter?&lt;/strong&gt;"&lt;/p&gt;

&lt;h2&gt;
  
  
  The language gap (this is where fine-tuning becomes mandatory)
&lt;/h2&gt;

&lt;p&gt;Here's the part most English-first evaluations miss.&lt;/p&gt;

&lt;p&gt;Llama Guard's training data is overwhelmingly English. Meta's own documentation recommends &lt;strong&gt;fine-tuning for custom taxonomies&lt;/strong&gt;, and multilingual performance degrades noticeably out of the box. If your users write in Brazilian Portuguese — with our lovely mess of legal Latin, abbreviations ("art. 1º, §2º"), and code-switching — you are not deploying Llama Guard. You are starting a fine-tuning &lt;em&gt;project&lt;/em&gt;: collecting labeled pt-BR attack/adversarial data, running evals, maintaining the checkpoint as the base model versions churn.&lt;/p&gt;

&lt;p&gt;Granite Guardian is meaningfully stronger multilingual out of the box, and its small MoE variant (&lt;strong&gt;3B with ~800M active parameters&lt;/strong&gt;) runs on modest hardware at production concurrency. For a team that needs a working rail &lt;em&gt;this quarter&lt;/em&gt; in a non-English market, "Apache-2.0 weights that classify Portuguese today" beats "excellent English classifier you must adapt first."&lt;/p&gt;

&lt;p&gt;To be fair: if you &lt;em&gt;do&lt;/em&gt; have the data engineering muscle, a fine-tuned Llama Guard for your exact taxonomy is a legitimate, strong choice. The point is that it's a prerequisite, not a config flag.&lt;/p&gt;

&lt;h2&gt;
  
  
  License: the detail procurement will ask about
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Granite Guardian: Apache-2.0.&lt;/strong&gt; Weights, commercial use, derivatives, self-hosting — no strings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Llama Guard: Llama Community License.&lt;/strong&gt; Fine for most, but it's &lt;em&gt;not&lt;/em&gt; OSI open source, it carries a monthly-active-user threshold, and it has acceptable-use clauses. In regulated procurement (banks, hospitals, government), "Apache-2.0" ends the conversation; "Llama license" starts one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When your differentiator is &lt;em&gt;auditability&lt;/em&gt; — an auditor can literally read and re-run your classifier — the license is part of the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code: both are one afternoon to try
&lt;/h2&gt;

&lt;p&gt;Granite Guardian via &lt;code&gt;llama.cpp&lt;/code&gt; (GGUF, CPU-fine for low volume) — the 3B MoE variant (~800M active params) is small, fast, and Apache-2.0:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="nt"&gt;-hf&lt;/span&gt; ggml-org/granite-guardian-3.2-3b-a800m-GGUF &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--jinja&lt;/span&gt; &lt;span class="nt"&gt;--parallel&lt;/span&gt; 64 &lt;span class="nt"&gt;--cache-type-k&lt;/span&gt; q8_0 &lt;span class="nt"&gt;--cache-type-v&lt;/span&gt; q8_0 &lt;span class="nt"&gt;-fa&lt;/span&gt; on
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The groundedness check — does the answer actually follow from the retrieved context? If not, it never reaches the human:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;granite-guardian&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;context:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;retrieved_chunks&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

response:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;llm_answer&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Is the response grounded in the context? Answer yes or no.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;verdict&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;block_and_log&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# the answer never reaches the human
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Llama Guard 4 (multimodal) via the same pattern, classifying prompt+response against Meta's categories — with the caveat that anything domain-specific (your bank-secrecy rule, your medical-council rule) needs fine-tuning data you'll have to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Llama Guard genuinely wins
&lt;/h2&gt;

&lt;p&gt;Intellectual honesty, because that's the point of the exercise:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;LlamaFirewall's ecosystem&lt;/strong&gt; — PromptGuard 2 (injection), AlignmentCheck (agent reasoning), CodeShield (insecure code) is a coherent suite for &lt;em&gt;agent&lt;/em&gt; safety that IBM doesn't fully match yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multimodal&lt;/strong&gt; — Llama Guard 4 handles images; Granite Guardian is text-first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;English content moderation at scale&lt;/strong&gt; — if that's your whole problem, it's excellent and battle-tested.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The actual lesson
&lt;/h2&gt;

&lt;p&gt;"Add a guardrail" is not one decision. It's three:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Content safety&lt;/strong&gt; (hate/violence/self-harm) → either model works; Llama Guard is proven.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Factuality rails&lt;/strong&gt; (groundedness, citation verification, scope) → Granite Guardian ships these natively; Llama Guard doesn't do this at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Domain rules&lt;/strong&gt; (your jurisdiction, your regulator, your taxonomy) → &lt;em&gt;nobody&lt;/em&gt; ships these. You either fine-tune, or you encode them as versioned rules around the model.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We chose: Granite Guardian for the ML rails (Apache-2.0, works in Portuguese, groundedness native) + Brazilian regulation as &lt;strong&gt;compliance-as-code&lt;/strong&gt; — versioned rules in Git that any auditor can read — + a public, immutable receipt for every interception, timestamped into Bitcoin via OpenTimestamps.&lt;/p&gt;

&lt;p&gt;That stack is live at &lt;a href="https://g.cloud" rel="noopener noreferrer"&gt;g.cloud&lt;/a&gt; — and the rule repository is open, because a guardrail nobody can inspect is just another black box you have to trust. In a regulated market, "trust me" is not a product.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What's your rail #3? I'm genuinely curious how teams encode jurisdiction-specific rules — fine-tuning, rules engines, or something else. Comments open.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Related reading on g.cloud:&lt;/strong&gt; &lt;a href="https://g.cloud/blog/en/o-que-e-guardrail-ia/" rel="noopener noreferrer"&gt;What is an AI guardrail&lt;/a&gt; · &lt;a href="https://g.cloud/blog/en/ibm-granite-guardian/" rel="noopener noreferrer"&gt;IBM Granite Guardian: the engine&lt;/a&gt; · &lt;a href="https://g.cloud/blog/en/recibo-do-guardrail/" rel="noopener noreferrer"&gt;The receipt: Bitcoin-timestamped proof&lt;/a&gt;`&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>security</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
