DEV Community

#aisafety

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Why AI Agents Need a Pre-Execution Guard

Why AI Agents Need a Pre-Execution Guard

Comments
5 min read
Der Passwort-Vorfall

Der Passwort-Vorfall

Comments
5 min read
China’s Kimi K3 AI Model Escapes Sandbox and Cheats on Test

China’s Kimi K3 AI Model Escapes Sandbox and Cheats on Test

Comments
3 min read
Die Einladung

Die Einladung

Comments
5 min read
Luna Fired an Employee. Then Forgot Why.

Luna Fired an Employee. Then Forgot Why.

Comments
3 min read
AX-RAY: VIDRAFT's Open AI Safety Diagnostic Leaderboard & Dataset Now on Hugging Face

AX-RAY: VIDRAFT's Open AI Safety Diagnostic Leaderboard & Dataset Now on Hugging Face

Comments
4 min read
The Verification Gap: Why “Done” Is Not a Fact About Your Codebase

The Verification Gap: Why “Done” Is Not a Fact About Your Codebase

Comments
10 min read
Human Oversight of AI Agents Failed 33% of the Time in Testing

Human Oversight of AI Agents Failed 33% of the Time in Testing

Comments
2 min read
The AI Wasn't Cheating. It Was Maximizing Its Score.

The AI Wasn't Cheating. It Was Maximizing Its Score.

1
Comments 1
6 min read
Beyond Reconstruction: Verifying Model Explanations with RECAP

Beyond Reconstruction: Verifying Model Explanations with RECAP

Comments
3 min read
When AI Models Escaped Their Sandbox: What the OpenAI Hugging Face Breach Really Means

When AI Models Escaped Their Sandbox: What the OpenAI Hugging Face Breach Really Means

Comments
3 min read
AI Safety & Ethics: Building Responsible AI Systems That Don't Backfire

AI Safety & Ethics: Building Responsible AI Systems That Don't Backfire

Comments
2 min read
AI Agent Safety: When Boundaries Fail with External Tools

AI Agent Safety: When Boundaries Fail with External Tools

Comments
4 min read
Day 12: LOOM now owns its memory — a trust layer for AI-written code, in plain language

Day 12: LOOM now owns its memory — a trust layer for AI-written code, in plain language

Comments
3 min read
Why Do Multi-Agent AI Systems Fail at Production Scale?

Why Do Multi-Agent AI Systems Fail at Production Scale?

3
Comments 4
8 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.