Originally published on rohitraj.tech
Mistral released Shieldstral on August 4, 2026 — a 3B open-weights safety classifier that reads your moderation policy at inference time instead of training on fixed categories. Here's how it actually compares to Llama Guard and the OpenAI Moderation API, the real self-host commands, and the one case where I'd still reach for a fixed taxonomy instead.
Read the full version with code samples, diagrams, and architecture details: Shieldstral vs Llama Guard vs OpenAI Moderation API: A Self-Hosted Content Moderation Guide (2026)
More engineering notes: rohitraj.tech/en/notes
Top comments (0)