<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Broskidev</title>
    <description>The latest articles on DEV Community by Broskidev (@broskigx).</description>
    <link>https://dev.to/broskigx</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4096551%2F46a39800-ab3c-423f-a6d5-a90d463ea955.png</url>
      <title>DEV Community: Broskidev</title>
      <link>https://dev.to/broskigx</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/broskigx"/>
    <language>en</language>
    <item>
      <title>Construí un Agente de pentesting que aprende de sus errores : OIHK PENTESTING</title>
      <dc:creator>Broskidev</dc:creator>
      <pubDate>Thu, 24 Sep 2026 02:51:23 +0000</pubDate>
      <link>https://dev.to/broskigx/mi-agente-de-pentesting-con-ia-ahora-aprende-de-sus-propios-errores-2cma</link>
      <guid>https://dev.to/broskigx/mi-agente-de-pentesting-con-ia-ahora-aprende-de-sus-propios-errores-2cma</guid>
      <description>&lt;p&gt;Hace unas semanas publiqué la primera versión de &lt;strong&gt;OIHK-pentesting&lt;/strong&gt;: un motor multi-agente que corre pentests autorizados con un planificador LLM, donde la frontera de seguridad vive en el &lt;strong&gt;motor&lt;/strong&gt;, no en el "buen comportamiento" del modelo.&lt;/p&gt;

&lt;p&gt;Desde entonces reconstruí la cabina entera. Este post va de lo nuevo — y de por qué la parte más difícil no fue hacer el agente más fuerte, sino hacer que &lt;strong&gt;se frene solo&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠️ Solo testing autorizado. Es una herramienta para gente con permiso escrito sobre los objetivos que apunta. El motor fuerza el scope igual — pero la responsabilidad es tuya.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Si quieres ver a OIHK en accion entra al github esta mas abajo 😈😈😈
&lt;/h2&gt;

&lt;p&gt;Si preferís solo leer: escribís un objetivo, Baron lo detecta, planifica y la flota ejecuta. Así de simple.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Se acabaron los slash commands. Escribís el objetivo y listo.
&lt;/h2&gt;

&lt;p&gt;La primera versión exponía "atajos de ataque" en la TUI. Estaba mal — el operador no tiene que memorizar maquinaria. Ahora:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;uv run oihk start
&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; escanea 10.10.12.5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Baron&lt;/strong&gt; (el copiloto interactivo) detecta el objetivo, planifica el assessment y delega los pasos de ataque él mismo. La TUI es chat; la inteligencia es suya.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. La flota: 7 versiones de sí mismo en paralelo
&lt;/h2&gt;

&lt;p&gt;Baron puede partir un plan en tareas sin solapamiento y lanzar &lt;strong&gt;siete subagentes reales en paralelo&lt;/strong&gt;, cada uno un loop LLM completo ejecutando comandos reales dentro del sandbox Kali gobernado:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Ari]   ▶ nmap -sV 10.10.12.5
[Bruma] ↳ 22/tcp open, 445/tcp open
[Ciro]  ✔ done — 3 candidatos a finding
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Olas en segundo plano&lt;/strong&gt;: lanza sin bloquear su propio turno, sigue planificando y junta reportes después con &lt;code&gt;fleet_status&lt;/code&gt; (hasta 3 olas en vuelo).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guards por unidad&lt;/strong&gt;: corte por llamadas repetidas, breaker por fallos consecutivos, tope de rondas. Que una unidad crashee nunca tumba la ola.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pre-flight de Docker&lt;/strong&gt;: sin daemon, no hay ola.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. AdapterOne: el agente que recuerda sus errores
&lt;/h2&gt;

&lt;p&gt;La parte que más me entusiasma. Cada miss del assessment — un scan bloqueado, recon vacío, un pivot equivocado — se convierte en una &lt;strong&gt;lección ponderada con decay&lt;/strong&gt; en un ledger JSONL local. En cada run futuro:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;las lecciones más pesadas se inyectan en el prompt de Baron &lt;strong&gt;automáticamente&lt;/strong&gt;,&lt;/li&gt;
&lt;li&gt;puede consultarlas a mitad de run (&lt;code&gt;adapterone_review&lt;/code&gt;),&lt;/li&gt;
&lt;li&gt;puede grabar reglas nuevas él mismo (&lt;code&gt;adapterone_confirm&lt;/code&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Mismo modelo, habilidad compuesta. Es RAG, pero el corpus son &lt;strong&gt;sus propios fracasos&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. /power: el pedal del acelerador de tu máquina
&lt;/h2&gt;

&lt;p&gt;Siete contenedores en paralelo fríen una laptop. Así que el operador tiene un presupuesto:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/power 0    # una instancia bajo demanda, spawn bloqueado
/power 30   # hasta 3 contenedores — las olas se recortan a eso
/power 100  # pool completo de 8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Bajar el presupuesto retira contenedores vivos &lt;strong&gt;al instante&lt;/strong&gt;. El camino de ejecución nunca se bloquea — reutilizar la instancia viva y la primera bajo demanda siempre funcionan. Baron está instruido para no reintentar ni insistir cuando un spawn se rechaza: se adapta y sigue. Tu daemon de Docker sobrevive; tus ventiladores también.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Por qué el motor sigue sin confiar en el modelo
&lt;/h2&gt;

&lt;p&gt;Todo lo anterior corre dentro del mismo stack de políticas que antes, porque agentes más fuertes necesitan &lt;strong&gt;rejas más fuertes&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Scope exacto, DNS-pinnado&lt;/strong&gt; — declarar &lt;code&gt;example.com&lt;/code&gt; NO autoriza sus subdominios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Egress fail-closed&lt;/strong&gt; — allowlist de netfilter dentro del namespace de red del propio sandbox; donde eso no se puede garantizar, el startup aborta.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Findings con evidencia obligatoria&lt;/strong&gt; — un finding necesita una ejecución real, propia y exitosa por una tool gobernada &lt;strong&gt;más&lt;/strong&gt; un registro de validación separado. Una historia convincente no prueba nada.&lt;/li&gt;
&lt;li&gt;El agente nunca ve una shell cruda. Nunca.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scorea &lt;strong&gt;24/24 en su propio benchmark de escenarios vulnerables&lt;/strong&gt; con el solver mock, y todo corre offline contra cualquier modelo local OpenAI-compatible.&lt;/p&gt;




&lt;h2&gt;
  
  
  Probalo
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/Broskigx/Oihk-pentesting.git
&lt;span class="nb"&gt;cd &lt;/span&gt;Oihk-pentesting
uv &lt;span class="nb"&gt;sync
&lt;/span&gt;uv run oihk start
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Licencia MIT, +1.300 tests pasando, beta temprana. Rompe fuerte — issues bienvenidas.&lt;/p&gt;

&lt;p&gt;¿Qué querés que un agente de pentest autónomo aprenda primero de sus propios fracasos? Oihk Es tu opcion opensource y gratis&lt;/p&gt;

</description>
      <category>python</category>
      <category>agents</category>
      <category>opensource</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>I built an autonomous multi-agent AI pentester — and why it's not another GPT wrapper</title>
      <dc:creator>Broskidev</dc:creator>
      <pubDate>Thu, 27 Aug 2026 02:52:12 +0000</pubDate>
      <link>https://dev.to/broskigx/i-built-an-autonomous-multi-agent-ai-pentester-and-why-its-not-another-gpt-wrapper-1l09</link>
      <guid>https://dev.to/broskigx/i-built-an-autonomous-multi-agent-ai-pentester-and-why-its-not-another-gpt-wrapper-1l09</guid>
      <description>&lt;p&gt;Most "AI pentester" projects are a single LLM in a while-loop with a shell. You&lt;br&gt;
give it a target, it runs commands until it decides it found something. That's&lt;br&gt;
how you get &lt;strong&gt;confident nonsense&lt;/strong&gt; — a model that writes a beautiful vulnerability&lt;br&gt;
report for a bug that doesn't exist.&lt;/p&gt;

&lt;p&gt;I wanted the opposite: an engine where a finding has to be &lt;em&gt;earned&lt;/em&gt;. So I built&lt;br&gt;
&lt;a href="https://github.com/Broskigx/Oihk-pentesting" rel="noopener noreferrer"&gt;OIHK&lt;/a&gt; — an autonomous, multi-agent&lt;br&gt;
AI penetration-testing engine. It's open source (MIT) and runs locally.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not one model in a loop — a team of agents
&lt;/h2&gt;

&lt;p&gt;OIHK is a &lt;strong&gt;multi-agent engine&lt;/strong&gt;. A root planner delegates to specialist agents —&lt;br&gt;
recon, discovery, validation, reporting — that all share two things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a &lt;strong&gt;versioned scan plan&lt;/strong&gt; (optimistic concurrency, revision history, resume), and&lt;/li&gt;
&lt;li&gt;an &lt;strong&gt;evidence ledger&lt;/strong&gt; (immutable execution records).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents don't coordinate by vibes in a chat log. They claim explicit plan steps,&lt;br&gt;
attach real evidence, and update state through a revisioned store. The root can't&lt;br&gt;
close a run while critical work is still open.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule I care about most: no evidence, no finding
&lt;/h2&gt;

&lt;p&gt;Here's the design decision the whole thing is built around:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An LLM writing a convincing PoC string is &lt;strong&gt;not&lt;/strong&gt; a finding.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A finding requires a &lt;strong&gt;real, successful, governed tool execution&lt;/strong&gt; &lt;em&gt;and&lt;/em&gt; a&lt;br&gt;
&lt;strong&gt;separate validation record&lt;/strong&gt;. Only a validation agent can turn evidence into a&lt;br&gt;
finding. If there's no execution record and no independent validation, it never&lt;br&gt;
becomes a finding — no matter how confident the model sounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Safety enforced in code, not prompts
&lt;/h2&gt;

&lt;p&gt;Offensive tools + autonomous agents is a scary combo if "be careful" is just a&lt;br&gt;
line in a prompt. In OIHK the guardrails are actual code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;PASSIVE mode is a policy layer&lt;/strong&gt;, not an instruction. Active tools are rejected
even if the agent tries to route them through the generic shell.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scope is exact.&lt;/strong&gt; Declaring &lt;code&gt;example.com&lt;/code&gt; doesn't authorize its subdomains or
resolved IPs. Declared hosts are resolved once and DNS-pinned for the whole run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Egress fails closed.&lt;/strong&gt; The declared scope is compiled into a netfilter
allowlist inside the sandbox's own per-run network namespace. On platforms that
can't guarantee it, startup &lt;strong&gt;aborts&lt;/strong&gt; instead of pretending to be isolated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hardened sandbox:&lt;/strong&gt; read-only rootfs, dropped capabilities, no-new-privileges,
non-root, no sudo surface.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Swap the model, keep the engine
&lt;/h2&gt;

&lt;p&gt;OIHK is provider-agnostic. Any OpenAI-compatible endpoint works (LM Studio by&lt;br&gt;
default), with per-role model routing and no hardcoded provider. You can run a&lt;br&gt;
strong reasoning model as the planner and a fast one for the specialists.&lt;/p&gt;

&lt;h2&gt;
  
  
  It's also its own benchmark
&lt;/h2&gt;

&lt;p&gt;This is my favorite part. OIHK doubles as an &lt;strong&gt;evaluation environment&lt;/strong&gt;: it runs&lt;br&gt;
the &lt;em&gt;real&lt;/em&gt; engine against 16 local, deliberately vulnerable scenarios and scores&lt;br&gt;
the model &lt;strong&gt;programmatically&lt;/strong&gt; — never by asking a model to grade itself.&lt;/p&gt;

&lt;p&gt;There's a deterministic offline &lt;code&gt;mock&lt;/code&gt; solver for CI and demos:&lt;/p&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

bash
uv run oihk eval run-all --model mock
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>python</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
