<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Emirhan</title>
    <description>The latest articles on DEV Community by Emirhan (@emirhandemir).</description>
    <link>https://dev.to/emirhandemir</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3837860%2Fc970470a-36ee-4993-9e9f-9c36664018c7.JPG</url>
      <title>DEV Community: Emirhan</title>
      <link>https://dev.to/emirhandemir</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/emirhandemir"/>
    <language>en</language>
    <item>
      <title>Agent security has two failure modes: miss the attack, or break the work.</title>
      <dc:creator>Emirhan</dc:creator>
      <pubDate>Mon, 05 Oct 2026 11:52:49 +0000</pubDate>
      <link>https://dev.to/emirhandemir/agent-security-has-two-failure-modes-miss-the-attack-or-break-the-work-1lck</link>
      <guid>https://dev.to/emirhandemir/agent-security-has-two-failure-modes-miss-the-attack-or-break-the-work-1lck</guid>
      <description>&lt;p&gt;A guardrail that blocks everything is safe.&lt;br&gt;
It is also useless.&lt;br&gt;
So we benchmarked both sides of AI agent security:&lt;/p&gt;

&lt;p&gt;Can you stop dangerous tool calls without breaking legitimate work?&lt;/p&gt;

&lt;p&gt;We ran:&lt;br&gt;
1,652 harmful tool calls across 15 attack families.&lt;br&gt;
24,911 benign tool calls from real agent sessions and public repositories.&lt;br&gt;
Against SolonGate, Claude Code permissions, Invariant, llm-guard, and the allowlists / denylists teams usually build themselves.&lt;br&gt;
The result for SolonGate:&lt;br&gt;
73.7% of harmful calls caught&lt;br&gt;
0.78% of legitimate calls falsely blocked&lt;br&gt;
15% of real sessions broken&lt;br&gt;
MCC: 0.790&lt;br&gt;
0.01 ms per call&lt;br&gt;
Nearly 3× the MCC of the next guard in the benchmark.&lt;/p&gt;

&lt;p&gt;But the number we care about most isn't detection alone.&lt;br&gt;
A security layer has to stop the attack without becoming the thing that stops the agent.&lt;br&gt;
We published the methodology, the misses, the false blocks, and the limitations too.&lt;/p&gt;

&lt;p&gt;And very soon, we're open-sourcing the Agent Security Gateway itself.&lt;br&gt;
Read the benchmark:&lt;br&gt;
&lt;/p&gt;
&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
        &lt;div class="c-embed__cover"&gt;
          &lt;a href="https://solongate.com/blog/agent-guardrail-benchmark/" class="c-link align-middle" rel="noopener noreferrer"&gt;
            &lt;img alt="" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsolongate.com%2Fblog-assets%2Fact-tradeoff.png" height="620" class="m-0" width="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="c-embed__body"&gt;
        &lt;h2 class="fs-xl lh-tight"&gt;
          &lt;a href="https://solongate.com/blog/agent-guardrail-benchmark/" rel="noopener noreferrer" class="c-link"&gt;
            Why do you need us for AI agent security? | SolonGate
          &lt;/a&gt;
        &lt;/h2&gt;
          &lt;p class="truncate-at-3"&gt;
            Choosing an agent guardrail, you weigh SolonGate, Claude Code permissions, Invariant, llm-guard, or roll-your-own. We put all of them on 1,652 harmful tool calls and 24,911 real ones and measured both halves. SolonGate catches the most and breaks the least, by nearly 3x the next guard.
          &lt;/p&gt;
        &lt;div class="color-secondary fs-s flex items-center"&gt;
            &lt;img alt="favicon" class="c-embed__favicon m-0 mr-2 radius-0" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsolongate.com%2Ffavicon.ico" width="64" height="64"&gt;
          solongate.com
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;PS!&lt;br&gt;
Agent security is quickly becoming another hype cycle.&lt;br&gt;
Companies are being valued and funded as if basic security primitives around AI agents are some kind of proprietary magic.&lt;br&gt;
A lot of that “magic” will be inside the Agent Security Gateway we’re open-sourcing in the next few days.&lt;br&gt;
AI security is too important to become another financial bubble built on closed boxes and pitch decks.&lt;br&gt;
Stop treating AI safety as a market opportunity first.&lt;br&gt;
Show your code.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>I let an AI agent loose on my codebase. It tried to read my .env file in 30 seconds.</title>
      <dc:creator>Emirhan</dc:creator>
      <pubDate>Mon, 23 Mar 2026 03:22:24 +0000</pubDate>
      <link>https://dev.to/emirhandemir/i-let-an-ai-agent-loose-on-my-codebase-it-tried-to-read-my-env-file-in-30-seconds-2imj</link>
      <guid>https://dev.to/emirhandemir/i-let-an-ai-agent-loose-on-my-codebase-it-tried-to-read-my-env-file-in-30-seconds-2imj</guid>
      <description>&lt;p&gt;Not a horror story. Well. Kind of.&lt;br&gt;
A few months ago, Çınar and I were building a side project. Nothing fancy. Just two guys, a codebase, and way too much coffee.&lt;br&gt;
We started using Claude Code to speed things up. And honestly? It was great. It was writing code faster than we could review it, jumping between files, running commands, doing things we didn't even ask it to do yet.&lt;br&gt;
That last part should have been a red flag.&lt;br&gt;
One evening I left it running while I went to grab food. Came back. Looked at the terminal. It had read the .env file.&lt;br&gt;
Not because it was malicious. Not because someone hacked it. Just because it could. Nobody told it not to. There was no rule. No policy. No wall.&lt;br&gt;
It saw a file. It read the file. That's it.&lt;br&gt;
And I sat there thinking: this thing has access to everything. Every file. Every command. Every API call. And I have absolutely no idea what it's been doing for the last 20 minutes.&lt;br&gt;
No logs. No audit trail. No "hey, are you sure about this?"&lt;br&gt;
Just vibes.&lt;/p&gt;

&lt;p&gt;That was the moment SolonGate started as an actual idea and not just a shower thought.&lt;br&gt;
We wanted something stupidly simple. Something that sat between the AI and everything it could touch, and asked one question before every single action:&lt;br&gt;
"Is this allowed?"&lt;br&gt;
If yes, go ahead. If no, stop. Log everything either way.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1rauyd5yviwlzru6g4my.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1rauyd5yviwlzru6g4my.jpeg" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No configuration PhD required. No 47-step setup guide. One command and you're protected.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;bashnpx @solongate/proxy &lt;span class="nt"&gt;--&lt;/span&gt; your-server
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it.&lt;/p&gt;

&lt;p&gt;Here's what it looks like in practice.&lt;br&gt;
We tested it with Gemini CLI last week. Asked it to read test.txt. Fine, allowed, no problem.&lt;br&gt;
Then asked it to read .env.&lt;br&gt;
Gemini's response: "I'm sorry, I cannot read the .env file. It seems to be blocked by a policy."&lt;br&gt;
And in the dashboard: read_file: .env — DENY — Policy Rule&lt;br&gt;
Logged. Timestamped. Done.&lt;br&gt;
The agent didn't argue. Didn't try again. Didn't find a creative workaround. Just stopped, reported back, and moved on.&lt;br&gt;
That's the whole point. Not to make AI tools useless. To make them safe enough to actually trust.&lt;/p&gt;

&lt;p&gt;We've blocked 100+ real attacks since we launched. Prompt injection attempts, path traversal, SSRF, credential file grabs. Some of them were tests. Some of them were not.&lt;br&gt;
Every single one is in the audit log. Every decision. Every layer that caught it. Every timestamp.&lt;br&gt;
When someone asks "what did your AI agent do last month?" — you have an answer.&lt;/p&gt;

&lt;p&gt;If you're running Claude Code, Gemini CLI, or any AI tool with file system or network access, and you don't have something like this in place — you're one unlucky prompt away from a bad day.&lt;br&gt;
We built SolonGate so that day doesn't happen :)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.tourl"&gt; solongate.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
