DEV Community

Matheus Feijão
Matheus Feijão

Posted on

Llama Guard is not a guardrail.

'"We put Granite Guardian and Llama Guard side by side for regulated Brazilian Portuguese. One worked out of the box. The other needed a fine-tuning project before it was useful."

https://g.cloud/blog/en/granite-guardian-vs-proprietarios/

Every team shipping LLM features in 2026 eventually asks the same question: what sits between the model and the user?

Meta and IBM both answer it with open-weights safety classifiers — Llama Guard and Granite Guardian. They look interchangeable on a feature matrix. They are not. We learned that the hard way while building g.cloud, a guardrail for regulated Brazilian industries (legal, healthcare, finance, government), where a hallucinated citation isn't a bug — it's professional misconduct.

Here's the honest comparison, including where Llama Guard wins.

The taxonomy problem

Llama Guard classifies against Meta's safety taxonomy (MLCommons categories: violence, sexual content, hate, self-harm, etc.). It does this well — for content safety.

But in a regulated domain, the dangerous output is usually not toxic. It's plausible and wrong:

  • a lawyer's assistant citing REsp 1.234.567 — a case that doesn't exist;
  • a health bot describing a dosage the retrieved document never mentioned;
  • a banking assistant politely explaining data that is under bank-secrecy law (in Brazil, LC 105/2001).

None of these trip a hate/violence/sexual classifier. This is where Granite Guardian plays a different game. Beyond harm categories, it ships detectors for:

  • Groundedness — is the response actually supported by the retrieved context? (the RAG hallucination rail)
  • Rationale — does the model's justification hold up?
  • Relevance and bias — did it drift off-scope or encode a stereotype?

That's the difference between a moderation filter and a guardrail. Moderation asks "is this harmful?" A guardrail asks "is this true, in-scope, and allowed to leave the perimeter?"

The language gap (this is where fine-tuning becomes mandatory)

Here's the part most English-first evaluations miss.

Llama Guard's training data is overwhelmingly English. Meta's own documentation recommends fine-tuning for custom taxonomies, and multilingual performance degrades noticeably out of the box. If your users write in Brazilian Portuguese — with our lovely mess of legal Latin, abbreviations ("art. 1º, §2º"), and code-switching — you are not deploying Llama Guard. You are starting a fine-tuning project: collecting labeled pt-BR attack/adversarial data, running evals, maintaining the checkpoint as the base model versions churn.

Granite Guardian is meaningfully stronger multilingual out of the box, and its small MoE variant (3B with ~800M active parameters) runs on modest hardware at production concurrency. For a team that needs a working rail this quarter in a non-English market, "Apache-2.0 weights that classify Portuguese today" beats "excellent English classifier you must adapt first."

To be fair: if you do have the data engineering muscle, a fine-tuned Llama Guard for your exact taxonomy is a legitimate, strong choice. The point is that it's a prerequisite, not a config flag.

License: the detail procurement will ask about

  • Granite Guardian: Apache-2.0. Weights, commercial use, derivatives, self-hosting — no strings.
  • Llama Guard: Llama Community License. Fine for most, but it's not OSI open source, it carries a monthly-active-user threshold, and it has acceptable-use clauses. In regulated procurement (banks, hospitals, government), "Apache-2.0" ends the conversation; "Llama license" starts one.

When your differentiator is auditability — an auditor can literally read and re-run your classifier — the license is part of the product.

Code: both are one afternoon to try

Granite Guardian via llama.cpp (GGUF, CPU-fine for low volume) — the 3B MoE variant (~800M active params) is small, fast, and Apache-2.0:

llama-server -hf ggml-org/granite-guardian-3.2-3b-a800m-GGUF \
  --jinja --parallel 64 --cache-type-k q8_0 --cache-type-v q8_0 -fa on
Enter fullscreen mode Exit fullscreen mode

The groundedness check — does the answer actually follow from the retrieved context? If not, it never reaches the human:

resp = client.chat.completions.create(
    model="granite-guardian",
    messages=[{"role": "user", "content": f"""context:
{retrieved_chunks}

response:
{llm_answer}

Is the response grounded in the context? Answer yes or no."""}],
    temperature=0,
)
verdict = resp.choices[0].message.content.strip().lower()
if "no" in verdict:
    block_and_log()  # the answer never reaches the human
Enter fullscreen mode Exit fullscreen mode

Llama Guard 4 (multimodal) via the same pattern, classifying prompt+response against Meta's categories — with the caveat that anything domain-specific (your bank-secrecy rule, your medical-council rule) needs fine-tuning data you'll have to build.

Where Llama Guard genuinely wins

Intellectual honesty, because that's the point of the exercise:

  1. LlamaFirewall's ecosystem — PromptGuard 2 (injection), AlignmentCheck (agent reasoning), CodeShield (insecure code) is a coherent suite for agent safety that IBM doesn't fully match yet.
  2. Multimodal — Llama Guard 4 handles images; Granite Guardian is text-first.
  3. English content moderation at scale — if that's your whole problem, it's excellent and battle-tested.

The actual lesson

"Add a guardrail" is not one decision. It's three:

  1. Content safety (hate/violence/self-harm) → either model works; Llama Guard is proven.
  2. Factuality rails (groundedness, citation verification, scope) → Granite Guardian ships these natively; Llama Guard doesn't do this at all.
  3. Domain rules (your jurisdiction, your regulator, your taxonomy) → nobody ships these. You either fine-tune, or you encode them as versioned rules around the model.

We chose: Granite Guardian for the ML rails (Apache-2.0, works in Portuguese, groundedness native) + Brazilian regulation as compliance-as-code — versioned rules in Git that any auditor can read — + a public, immutable receipt for every interception, timestamped into Bitcoin via OpenTimestamps.

That stack is live at g.cloud — and the rule repository is open, because a guardrail nobody can inspect is just another black box you have to trust. In a regulated market, "trust me" is not a product.

What's your rail #3? I'm genuinely curious how teams encode jurisdiction-specific rules — fine-tuning, rules engines, or something else. Comments open.


Related reading on g.cloud: What is an AI guardrail · IBM Granite Guardian: the engine · The receipt: Bitcoin-timestamped proof`

Top comments (0)