<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: auto_majicly</title>
    <description>The latest articles on DEV Community by auto_majicly (@xenocoregiger31).</description>
    <link>https://dev.to/xenocoregiger31</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg</url>
      <title>DEV Community: auto_majicly</title>
      <link>https://dev.to/xenocoregiger31</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/xenocoregiger31"/>
    <language>en</language>
    <item>
      <title>It is still happening as we speak. I needed a github repo url for a tool, and THAT too was blocked.</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Thu, 20 Aug 2026 22:29:21 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/it-is-still-happening-as-we-speak-i-needed-a-github-repo-url-for-a-tool-and-that-too-was-blocked-2eml</link>
      <guid>https://dev.to/xenocoregiger31/it-is-still-happening-as-we-speak-i-needed-a-github-repo-url-for-a-tool-and-that-too-was-blocked-2eml</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/xenocoregiger31/when-the-safety-system-becomes-the-threat-model-a-case-study-in-classifier-drift-372k" class="crayons-story__hidden-navigation-link"&gt;When the Safety System Becomes the Threat Model: A Case Study in Classifier Drift&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/xenocoregiger31" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" alt="xenocoregiger31 profile" class="crayons-avatar__image" width="800" height="450"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/xenocoregiger31" class="crayons-story__secondary fw-medium m:hidden"&gt;
              auto_majicly
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                auto_majicly
                
                
              
              &lt;div id="story-author-preview-content-4448179" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/xenocoregiger31" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" class="crayons-avatar__image" alt="" width="800" height="450"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;auto_majicly&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/xenocoregiger31/when-the-safety-system-becomes-the-threat-model-a-case-study-in-classifier-drift-372k" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 20&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/xenocoregiger31/when-the-safety-system-becomes-the-threat-model-a-case-study-in-classifier-drift-372k" id="article-link-4448179"&gt;
          When the Safety System Becomes the Threat Model: A Case Study in Classifier Drift
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/claude"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;claude&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/security"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;security&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/cybersecurity"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;cybersecurity&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/xenocoregiger31/when-the-safety-system-becomes-the-threat-model-a-case-study-in-classifier-drift-372k" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/fire-f60e7a582391810302117f987b22a8ef04a2fe0df7e3258a5f49332df1cec71e.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/raised-hands-74b2099fd66a39f2d7eed9305ee0f4553df0eb7b4f11b01b6b1b499973048fe5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;3&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/xenocoregiger31/when-the-safety-system-becomes-the-threat-model-a-case-study-in-classifier-drift-372k#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              1&lt;span class="hidden s:inline"&gt;&amp;nbsp;comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            6 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>When the Safety System Becomes the Threat Model: A Case Study in Classifier Drift</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Thu, 20 Aug 2026 22:27:20 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/when-the-safety-system-becomes-the-threat-model-a-case-study-in-classifier-drift-372k</link>
      <guid>https://dev.to/xenocoregiger31/when-the-safety-system-becomes-the-threat-model-a-case-study-in-classifier-drift-372k</guid>
      <description>&lt;p&gt;&lt;em&gt;Field notes from 43 hours of legitimate security research, 35 false-positive blocks, and one&lt;br&gt;
support form.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;I do authorized bug bounty work and CTF practice. My toolchain is unremarkable by design: the&lt;br&gt;
usual ProjectDiscovery/OWASP staples — &lt;code&gt;subfinder&lt;/code&gt;, &lt;code&gt;httpx&lt;/code&gt;, &lt;code&gt;nuclei&lt;/code&gt;, &lt;code&gt;katana&lt;/code&gt;, &lt;code&gt;naabu&lt;/code&gt;, &lt;code&gt;dnsx&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;amass&lt;/code&gt; — plus &lt;code&gt;ffuf&lt;/code&gt;, &lt;code&gt;gobuster&lt;/code&gt;, &lt;code&gt;arjun&lt;/code&gt;, &lt;code&gt;bbot&lt;/code&gt;. Nothing bespoke, nothing evasive, nothing that&lt;br&gt;
wouldn't show up in any security bootcamp's syllabus. I use Claude as a research and&lt;br&gt;
documentation assistant alongside my OWN dev tool HALO, that I have been developing over the course of about six months(roughly). I use these toolchains, on training platforms explicitly built to be&lt;br&gt;
practiced against, learned from and exercised. And I pay for the Cluade PRO plan as a means of learning and having an assistant ten times smarter than myself.&lt;/p&gt;

&lt;p&gt;Over a 43-hour window, that combination — legitimate work plus an AI assistant with a&lt;br&gt;
cybersecurity safeguard — produced 35 blocked requests. This is a writeup of what the block&lt;br&gt;
pattern actually looked like, because the pattern is the interesting part. It isn't "the filter&lt;br&gt;
is too strict." It's that the filter appears to be scoring the wrong thing. And by 'wrong thing' I mean even un-related normal language messages and requests.&lt;/p&gt;
&lt;h2&gt;
  
  
  The headline case: blocked for writing an ethics checklist
&lt;/h2&gt;

&lt;p&gt;The densest cluster was ten blocks in four minutes and thirty-three seconds. The request,changed and re-phrased, also unchanged (no matter what I did),&lt;br&gt;
across all ten attempts: write a markdown template for reporting bug bounty findings.&lt;/p&gt;

&lt;p&gt;The document that eventually made it to disk — after the eleventh try succeeded — is 513 lines&lt;br&gt;
across four files: a submission skeleton, a CVSS severity guide, a README, and a pre-submission&lt;br&gt;
checklist. The checklist is the file that was open when block six through ten fired. Its contents:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- [ ] The asset is explicitly in scope for this program, today
- [ ] Testing stayed within the program's rules (no DoS, no social engineering,
      no automated scanning if prohibited, no third-party accounts)
- [ ] I stated a concrete attacker outcome, not a capability
- [ ] I did not overclaim (no "full server compromise" for a reflected header)
- [ ] Secrets, tokens, and third-party PII are redacted
- [ ] I deleted test data, injected records, and uploaded files I created
- [ ] I did not retain third-party data
- [ ] No production users were affected
- [ ] Tone is neutral and collaborative — no demands about bounty amount
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's a responsible-disclosure ethics document. It instructs the reader to stay in scope, avoid&lt;br&gt;
denial of service, redact third-party data, and delete their own test artifacts. If a classifier's&lt;br&gt;
job is to catch material that increases risk, this is close to the least risky text a security&lt;br&gt;
practitioner could type. It got blocked more times than anything else in the sample.&lt;/p&gt;

&lt;h2&gt;
  
  
  The control case that makes the point
&lt;/h2&gt;

&lt;p&gt;The same account, the same day, the same client. Earlier that evening I ran a 43-minute CTF&lt;br&gt;
session — 204 messages, live enumeration and credential testing against a practice box, the actual&lt;br&gt;
offensive work the safeguard exists to gate.&lt;/p&gt;

&lt;p&gt;Zero blocks.&lt;/p&gt;

&lt;p&gt;Side by side:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;CTF session (&lt;code&gt;df828311&lt;/code&gt;)&lt;/th&gt;
&lt;th&gt;Report-template session (&lt;code&gt;e6f361c6&lt;/code&gt;)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Activity&lt;/td&gt;
&lt;td&gt;live offensive operation&lt;/td&gt;
&lt;td&gt;writing a markdown document&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Messages&lt;/td&gt;
&lt;td&gt;204&lt;/td&gt;
&lt;td&gt;~12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duration&lt;/td&gt;
&lt;td&gt;43 minutes&lt;/td&gt;
&lt;td&gt;~5 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blocks&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;10&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Whatever the classifier is weighting, it isn't operational risk. The session where I was actually&lt;br&gt;
doing the thing the safeguard is presumably designed to catch ran clean start to finish. The&lt;br&gt;
session where I was documenting how to do that thing &lt;em&gt;responsibly&lt;/em&gt; did not. And to make matters worse, writing this article you are currently reading was also blocked.&lt;/p&gt;

&lt;p&gt;The pattern that best fits the data: block density tracks accumulated conversation context —&lt;br&gt;
how much security vocabulary has built up over the session — rather than the risk content of any&lt;br&gt;
single request. A stable, unchanging request (write a template) got a different verdict each of&lt;br&gt;
eleven times it was submitted in the same conversation. That's not how a per-request classifier&lt;br&gt;
behaves. It's how a classifier scoring cumulative context behaves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two failure modes that compound it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It fires on output, not just input.&lt;/strong&gt; Several blocks truncated the assistant's own reply&lt;br&gt;
mid-sentence — once while it was in the middle of advising me to keep scope notes in a local file&lt;br&gt;
instead of chat. From inside the session, a mid-word cutoff is indistinguishable from a network&lt;br&gt;
hiccup or a length cap. It isn't labeled as a refusal. This led the assistant, twice, to&lt;br&gt;
confidently tell me the truncation &lt;em&gt;wasn't&lt;/em&gt; censorship and to suggest I use plainer terminology —&lt;br&gt;
advice that would have increased the block rate, not decreased it. The failure mode doesn't just&lt;br&gt;
degrade the experience; it actively misinforms the system's own operator about what's happening.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It fires on meta-discussion of itself.&lt;/strong&gt; I tried, deliberately, to describe the blocking problem&lt;br&gt;
using zero security terminology — just "I sent a message and it got blocked." That message was&lt;br&gt;
blocked too. So was a follow-up attempt to make the same point. There was, in that stretch, no&lt;br&gt;
available phrasing that got a plain factual report of the bug past the filter meant to catch&lt;br&gt;
security content.&lt;/p&gt;

&lt;p&gt;Put together: a request gets silently dropped, the assistant doesn't know it was dropped, it&lt;br&gt;
retries or re-explains, and each retry is itself scored against an already-elevated context&lt;br&gt;
window. Support documentation for the underlying system describes an escalating filter that&lt;br&gt;
tightens with repeated triggers and cools off after a quiet period. If that's accurate, the retry&lt;br&gt;
behavior isn't neutral — it's the thing generating the escalation. One legitimately blocked&lt;br&gt;
request can silently become ten logged violations against the account, without the user or the&lt;br&gt;
assistant ever being told a block occurred.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters beyond one annoyed user
&lt;/h2&gt;

&lt;p&gt;The stated purpose of a cybersecurity safeguard is presumably to reduce uplift for offensive&lt;br&gt;
misuse while still serving people doing the work legitimately — pentesters, bounty hunters,&lt;br&gt;
defenders, students. A classifier that blocks the ethics checklist ten times and waves through 43&lt;br&gt;
minutes of live target enumeration is optimizing for something other than that goal. If anything,&lt;br&gt;
the current calibration selectively suppresses the exact material — scope discipline, redaction&lt;br&gt;
practices, responsible severity claims — that makes offensive-security work &lt;em&gt;safer&lt;/em&gt; to produce,&lt;br&gt;
while letting the higher-capability activity through untouched. That's the opposite of the&lt;br&gt;
intended tradeoff.&lt;/p&gt;

&lt;p&gt;There's also a plain usability cost. I pay for this specifically because the model is useful in&lt;br&gt;
this domain. The current behavior means the subscription is least usable for the exact reason I&lt;br&gt;
bought it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I filed
&lt;/h2&gt;

&lt;p&gt;I compiled a full evidence document — request IDs, timestamps, session breakdowns, the control&lt;br&gt;
case, verbatim excerpts of the blocked content — and re-applied to the program that governs an&lt;br&gt;
exemption from this safeguard, after an earlier application had been declined.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The medium-length version of that appeal, roughly what I'd want a reviewer to actually read:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I'm a Claude Pro subscriber and an individual security researcher doing authorized bug bounty&lt;br&gt;
work and CTF practice. I previously applied to the Cyber Verification Program and was declined.&lt;br&gt;
I'm asking for that decision to be re-reviewed against logged request data rather than a&lt;br&gt;
description of intent.&lt;/p&gt;

&lt;p&gt;Between two dates my account logged 35 blocks citing "a cybersecurity topic." All 35 request IDs&lt;br&gt;
are recoverable from local session logs and verifiable server-side.&lt;/p&gt;

&lt;p&gt;The densest cluster — ten blocks in four minutes and 33 seconds — occurred while asking for a&lt;br&gt;
markdown template for &lt;em&gt;reporting&lt;/em&gt; findings professionally. Not an exploit, not a payload, not a&lt;br&gt;
target: a document about stating impact honestly, avoiding overclaiming, confirming scope, and&lt;br&gt;
deleting test data. Later the same evening, working a CTF box that exists solely to be practiced&lt;br&gt;
against, the blocked messages included requests to name a working directory.&lt;/p&gt;

&lt;p&gt;The control case: a 43-minute capture-the-flag session, 204 messages of live work against a&lt;br&gt;
practice target, produced zero blocks. The session asking for documentation guidance drew ten.&lt;br&gt;
Whatever is being scored, it isn't the risk posed by the request.&lt;/p&gt;

&lt;p&gt;Rephrasing doesn't reliably help. Vocabulary was sanitized — no tool names, no site names, no&lt;br&gt;
mention of the field — and blocks continued. One blocked message contained no security&lt;br&gt;
terminology at all; it only described the fact of being blocked.&lt;/p&gt;

&lt;p&gt;The program's own documentation states that eligible applications are occasionally declined&lt;br&gt;
incorrectly. This is offered as evidence that this account is one of those cases, and as a&lt;br&gt;
request to be evaluated as what it is: legitimate work in a defensive field, using publicly&lt;br&gt;
distributed tooling, on systems built to be practiced against.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's the version sized for a standard appeal text box — long enough to carry the control case&lt;br&gt;
and the request-ID anchor, short enough that a reviewer will actually finish it. A one-paragraph&lt;br&gt;
version and the full 35-entry evidence log exist alongside it for forms with tighter or looser&lt;br&gt;
limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable summary
&lt;/h2&gt;

&lt;p&gt;None of this is an argument that the safeguard shouldn't exist. It's an argument that the current&lt;br&gt;
implementation can't currently tell the difference between doing the risky thing and writing a&lt;br&gt;
checklist about how not to do the risky thing badly — and that the difference matters, because one&lt;br&gt;
of those outputs is the thing that makes the other one safer.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>security</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>My AI pentester popped 3 root shells and told me the other 20 failed. Good.</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Tue, 18 Aug 2026 20:29:28 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/my-ai-pentester-popped-3-root-shells-and-told-me-the-other-20-failed-good-6hl</link>
      <guid>https://dev.to/xenocoregiger31/my-ai-pentester-popped-3-root-shells-and-told-me-the-other-20-failed-good-6hl</guid>
      <description>&lt;p&gt;I spent a whole day this week making my autonomous pentesting agent slower. On purpose, sort of. Here's how that happened, and why it turned into the most honest day of building I've had in a while.&lt;/p&gt;

&lt;p&gt;Some context: I'm ~6 months into teaching myself to code, and I'm building HALO — a local, self-owned pentesting agent driven by a small model running on my own hardware. No cloud API, no $475/yr scanner subscription. The goal isn't to pop the easy boxes. It's to build something that hunts the bug after the easy three, on targets that aren't gift-wrapped.&lt;/p&gt;

&lt;p&gt;The fix that didn't exist&lt;br&gt;
The day started with a "speed fix." My model calls were taking ~90 seconds each, and an engagement was crawling. The obvious move: cap the model's output tokens so it stops rambling. Small change. Ship it.&lt;/p&gt;

&lt;p&gt;It made everything worse.&lt;/p&gt;

&lt;p&gt;The runs got faster and started failing completely — the agent would find all the open ports and then just... give up on every one. When I finally pulled the actual logs instead of guessing, the truth was ugly and simple:&lt;/p&gt;

&lt;p&gt;The model I'm running is a reasoning model. It burns ~1,200 tokens thinking in a hidden channel before it writes a single token of the answer I actually need. My token cap was set below the thinking budget. So the model would think, hit the ceiling mid-thought, and return an empty string. Every "smart" cap I added was decapitating the answer.&lt;/p&gt;

&lt;p&gt;The kicker: capping tokens can't speed up a reasoning model at all. The slowness was just... the model. On my hardware, that's the floor. There was no speed fix. I'd spent hours optimizing a lever that doesn't exist.&lt;/p&gt;

&lt;p&gt;Lesson one: read the logs before you theorize. I knew this. I did it anyway. The receipts were 30 seconds away the whole time.&lt;/p&gt;

&lt;p&gt;The part I'm actually proud of&lt;br&gt;
Once the agent could think again, I watched it work a lab box (a deliberately-vulnerable Metasploitable VM — the standard punching bag). And here's the thing that made the whole day worth it.&lt;/p&gt;

&lt;p&gt;The agent has three curated exploits it knows work. It fired them and came back with this:&lt;/p&gt;

&lt;p&gt;HALO-EVIDENCE nonce=1a837238e19fc8fb56014af0 level=shell uid=0 user=root host=metasploitable&lt;br&gt;
root@metasploitable:/#&lt;br&gt;
That nonce is the whole point. My orchestrator mints a fresh random string for each attempt, and the only way that string shows up in the output is if a real shell on the target actually ran the payload and echoed it back. The exploit script can't fake it. The model can't claim it. A banner that says "root" doesn't count. The system cannot grade its own homework — it has to produce a receipt minted by something outside itself.&lt;/p&gt;

&lt;p&gt;This is the idea you and I keep circling back to in the comments here: don't trust a component's self-report. Make it prove the claim with something it couldn't have forged. My agent popped three genuine root shells, and I believe it because I made it hard for it to lie to me.&lt;/p&gt;

&lt;p&gt;And then it fired at everything else — and told me it failed&lt;br&gt;
Here's the honest ending. Beyond the three curated exploits, I'd just wired up a Metasploit pipeline so the agent could reach for the ~2,000 exploit modules Metasploit knows. Watching it run against the other 20 open ports was humbling:&lt;/p&gt;

&lt;p&gt;It fired an Apache 2.4.49 exploit at Apache 2.2.8. (The name matched. The version was nonsense.)&lt;br&gt;
It picked heavyweight staged payloads where a dead-simple command shell would've been more reliable.&lt;br&gt;
It reached for a MySQL exploit that needs credentials it didn't have.&lt;br&gt;
Every single one of those came back "Nothing worked on port X." Not a fake success. Not a padded number. The exact same nonce gate that confirmed the real root shells refused to confirm the ones that didn't land.&lt;/p&gt;

&lt;p&gt;That's the feature. A tool that lies to you about coverage is worse than useless — it's dangerous. Mine popped three and honestly told me the other twenty didn't work, and why, in the logs. That "why" is my whole to-do list for next week.&lt;/p&gt;

&lt;p&gt;Where I actually landed&lt;br&gt;
When I committed the day's work, I made myself write the honest version of the message. Not "working Metasploit path." The real one:&lt;/p&gt;

&lt;p&gt;Curated PoCs pop nonce-verified root reliably; the MSF path fires end-to-end but session-landing refinement is still pending.&lt;/p&gt;

&lt;p&gt;Both halves of that sentence are true, and I only know they're true because the system won't let me round up.&lt;/p&gt;

&lt;p&gt;Two kinds of honesty came out of one day:&lt;/p&gt;

&lt;p&gt;The agent's — it can't claim a breach without a receipt it couldn't forge.&lt;br&gt;
Mine — I committed what's real and labeled what isn't, instead of shipping a hype commit.&lt;br&gt;
They're the same principle, pointed in two directions. If I'm going to build something I actually trust to point at a target, both have to hold.&lt;/p&gt;

&lt;p&gt;Still a long way from the thing I want. But it's a long way built on receipts instead of vibes — and after a day that started by optimizing an imaginary problem, I'll take it.&lt;/p&gt;

&lt;p&gt;What's the ugliest "the system was grading its own homework" bug you've caught in your own builds? I'm collecting them.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devjournal</category>
      <category>python</category>
      <category>security</category>
    </item>
    <item>
      <title>Green Tests, Zero Shells: What a Live Target Taught Me About Autonomous Exploit Agents</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Mon, 17 Aug 2026 23:05:05 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/green-tests-zero-shells-what-a-live-target-taught-me-about-autonomous-exploit-agents-38go</link>
      <guid>https://dev.to/xenocoregiger31/green-tests-zero-shells-what-a-live-target-taught-me-about-autonomous-exploit-agents-38go</guid>
      <description>&lt;p&gt;I spend most of my time building an autonomous security agent — it recons a target I'm authorized to test, picks exploits, and tries to prove access, all against my own lab boxes. This week I rebuilt the brain of it, watched every test pass, ran it against a live target, and it popped nothing. Not because the model was dumb. Because my tests were quietly lying to me. AGAIN... so, how is this possible...AGAIN? &lt;/p&gt;

&lt;p&gt;Here's the whole story, because the mistake is more useful than the feature.&lt;/p&gt;

&lt;p&gt;The thing I built&lt;br&gt;
The old design was a frozen plan. A planner decomposed the goal into an ordered list of subtasks once, up front, and an executor walked that list in order. That's fine until reality disagrees with the plan — a port you didn't expect, an exploit that fails, a service that isn't what the banner claimed. A frozen list can't react. It just marches.&lt;/p&gt;

&lt;p&gt;So I replaced it with a closed loop. Every iteration, the policy gets the full current state plus everything that's already failed, and picks one next action. Act, fold the result back into the belief state, re-plan. The planner stopped being a one-shot decomposer and became a per-step policy.&lt;/p&gt;

&lt;p&gt;The part I care about most isn't the loop — it's the teeth and the instrumentation, because an autonomous thing that can't stop or can't be measured is a liability, not a tool:&lt;/p&gt;

&lt;p&gt;Termination teeth, checked before every action: an iteration budget, a wall-clock deadline, a kill switch, and a stall detector (belief state stopped changing → stop). Plus a repeat guard: if the policy re-emits an action it already tried, don't re-run it — count it as a repeat.&lt;br&gt;
An adaptation score. Every re-plan logs whether the new action was genuinely novel or just a reshuffle of things already tried. Over a run you get novel / total. The whole point: measure whether the model actually adapts, instead of assuming it does. A number that can embarrass me is worth ten that flatter me.&lt;br&gt;
Honest breach confirmation. Success is never the model's say-so. A breach only counts on real evidence, checked by code the model doesn't control. Don't let the system grade its own homework.&lt;br&gt;
Tests: green across the board. Deploy. Run.&lt;/p&gt;

&lt;p&gt;Zero shells&lt;br&gt;
First real run against the lab target, the model reconned fine, then started working ports. And on the one port that should be the easiest win in the whole box — a service with a well-known backdoor, the kind of thing that pops every single time through the old path — the log said:&lt;/p&gt;

&lt;p&gt;Nothing worked on port 21.&lt;/p&gt;

&lt;p&gt;Zero breaches. Over sixteen iterations. And here's the part that saved me from blaming the wrong thing: the model had picked the right target. It correctly identified the vulnerable service and chose to attack it. Its judgment was fine. Something downstream ate the win.&lt;/p&gt;

&lt;p&gt;The tests were encoding a fiction&lt;br&gt;
The old, proven path confirms this exact exploit through a challenge–response nonce. The agent mints a single-use token, the exploit has to echo that token back inside a structured evidence line as proof it actually ran our command — not a banner, not a reflected string, an execution-derived fact bound to that one attempt. It's the anti-cheat that stops a tarpit or a reflector from faking a shell.&lt;/p&gt;

&lt;p&gt;The new gated attack path… never minted the nonce. So the exploit fired, produced its nonce-bound proof, and the confirmer had nothing to check it against. No token in, no confirmation out. "Nothing worked."&lt;/p&gt;

&lt;p&gt;But my tests were green. Why?&lt;/p&gt;

&lt;p&gt;Because the tests fed the confirmer a raw uid=0(root) string and asserted "breach confirmed." That was true once — before the nonce hardening. After the hardening, the real exploit stopped emitting a bare uid=0 banner and started emitting the structured, nonce-bound line instead. The tests kept using the old fake evidence. So they exercised a code path that no longer matched reality, and passed with total confidence.&lt;/p&gt;

&lt;p&gt;The unit tests weren't testing the system. They were testing a museum exhibit of the system.&lt;/p&gt;

&lt;p&gt;The part that actually scared me&lt;br&gt;
When I traced it, the same missing nonce logic sat in another execution path I'd considered working for weeks. It had "popped shells" in a demo run months earlier — before the hardening landed. Nobody re-checked it after. Its tests were green too, for the same fake-evidence reason. So a bug I thought was scoped to my new code had silently disabled breach confirmation across two paths, and my whole test suite waved it through.&lt;/p&gt;

&lt;p&gt;One live run against a real target found what a green suite hid for weeks.&lt;/p&gt;

&lt;p&gt;What I actually changed&lt;br&gt;
Nothing exotic. The gated path now mints and injects the nonce like the proven path always did, and confirms against it. The downstream re-check couldn't re-consume a single-use token (that's the point of single-use), so it verifies the structured evidence artifact is present instead of demanding a raw banner. And I rewrote every test that had been feeding fake uid=0 strings so they prove a breach the way the real exploit actually does — through the nonce.&lt;/p&gt;

&lt;p&gt;Then I reran it. Port 21 landed. Honestly, gated, end to end.&lt;/p&gt;

&lt;p&gt;Three things I'm taking with me&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;A passing test proves your code matches your test, not that it matches reality. Mine matched a fixture that had quietly gone stale. If an assertion hard-codes what "success" looks like, that string is a second source of truth you now have to keep honest. Pin tests to the real artifact your system produces, or they'll rot into fiction while staying green.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A live target is an external judge your unit tests can't be. Everything about my design was aimed at not letting the agent grade its own homework — the nonce, the evidence check, the confirmer the model can't reach. But I'd let the test suite grade its own homework with fake evidence. The real box doesn't care what my fixtures assert. That's exactly why it's valuable.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Instrument for the answer you're afraid of. The adaptation score exists to tell me, out loud, if my local model just reshuffles instead of reasoning. I haven't gotten a clean number yet — the runs kept surfacing bugs first, and the model's slow enough that I'm capping generation length next. But I built the thing that can embarrass me before I built the thing that impresses me, and that ordering is the only reason today ended with a real shell instead of a fake one.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I didn't get the headline number I wanted this week. I got something better: a system that fails honestly, tests that finally match reality, and a very concrete reminder that "all green" is a claim about my tests, not about my code.&lt;/p&gt;

&lt;p&gt;Zero shells is a better teacher than three shells you can't trust.&lt;/p&gt;

&lt;p&gt;Everything here runs against targets I own and am authorized to test. If you're building agents that act on the world, build the part that can prove — or disprove — the claim before you build the part that makes it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>autonomous</category>
      <category>cybernews</category>
      <category>secops</category>
    </item>
    <item>
      <title>A CTF Session With No Flag — and Why It Still Counted</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Sun, 16 Aug 2026 16:09:19 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/a-ctf-session-with-no-flag-and-why-it-still-counted-53d5</link>
      <guid>https://dev.to/xenocoregiger31/a-ctf-session-with-no-flag-and-why-it-still-counted-53d5</guid>
      <description>&lt;p&gt;I spend most of my time building a small offensive-security framework. Today I spent a session working a target the manual way and finished it without capturing a single flag. I want to write about that session anyway, because the reason it was still worthwhile is, I think, the most useful thing I've learned this year.&lt;/p&gt;

&lt;p&gt;The tooling milestone behind the session&lt;br&gt;
First, some context on what I've been building. The framework — I call it HALO — recently crossed an important threshold: the tool arsenal grew from roughly thirty integrations to forty-two, and the new additions form a complete web pipeline. Reconnaissance flows into web attack, which flows into flag capture, all coordinated behind a single authorization gate rather than requiring me to shuttle output between terminals by hand.&lt;/p&gt;

&lt;p&gt;The more significant change, however, isn't a tool at all. It's a design principle I keep having to relearn:&lt;/p&gt;

&lt;p&gt;A system should never be the sole judge of its own success.&lt;/p&gt;

&lt;p&gt;Earlier versions of HALO would execute an exploit, observe output that resembled success, and report a compromise. It was frequently wrong. The problem wasn't a weak model — it was the absence of any external definition of "did this actually work." Any string that looked like a win was treated as one. In one memorable run it reported compromising twenty-three of twenty-three services. The verified number was zero.&lt;/p&gt;

&lt;p&gt;The fix was architectural, not intellectual. I moved the definition of success outside the component being evaluated: evidence-based confirmation that a shell is genuinely interactive, a challenge-response the executing agent cannot reason its way around, and verification that does not rely on the attacker's own account of events. When the proof of success lives inside the process that wants to succeed, that process will always find a way to pass. The judge has to sit outside the room.&lt;/p&gt;

&lt;p&gt;I mention this now because the session that followed was, in effect, the human version of the same lesson.&lt;/p&gt;

&lt;p&gt;The target&lt;br&gt;
I was working an introductory HackingHub environment (their VulnBegin hub) alongside an AI assistant. The division of labor suited me: I direct the engagement and run the in-network commands, and the assistant handles rapid reconnaissance and reasoning. The hub has an interesting structure — each time you spin up the environment, it assigns you a single, randomly selected flag to solve. You don't choose which one. Every spawn is therefore a self-contained puzzle, and I had a handful of stubborn ones remaining.&lt;/p&gt;

&lt;p&gt;Reconnaissance returned the expected web application, plus two details worth attention:&lt;/p&gt;

&lt;p&gt;A second HTTP service on a high port, which the application's admin dashboard described as an "API SERVER — CONNECTED."&lt;br&gt;
A second SSH daemon, on an unusual high port and a different build from the primary one.&lt;br&gt;
The application also disclosed an API token. So I had a token, a service advertising itself as an API, and a dashboard insisting the two were linked. That is a compelling narrative, and the instinct is to follow it.&lt;/p&gt;

&lt;p&gt;Two false leads, and how they were ruled out&lt;br&gt;
The token. It had every appearance of a key. I supplied it to the main application in every reasonable form — as a header, a query parameter, a bearer credential — across every route I could enumerate. The result was not rejection; it was indifference. The application gave no indication it recognized the token at all.&lt;/p&gt;

&lt;p&gt;The "API server." This is the lead that would have consumed an entire evening a year ago. The high-port service returned ERROR - NOT PART OF THE CTF/TRAINING to every request. My assumption was a routing problem: wrong path, wrong method, wrong host header, wrong token format. I tested each of them. A bare GET / with the valid token produced the same error. A POST produced the same error. A spoofed Host header produced the same error.&lt;/p&gt;

&lt;p&gt;At that point the pattern resolves: the service is not rejecting my request, it is returning that identical string to everything. It is not the challenge's API server — it is the platform's global out-of-scope guard, the same boundary every environment presents for infrastructure outside the current exercise. The dashboard's "CONNECTED" label was decoration. The disclosed token was a distractor.&lt;/p&gt;

&lt;p&gt;Ruling out a lead is a skill in its own right, and it obeys the same principle as the tooling work above: you do not get to declare a path dead because you have grown tired of it — the evidence has to declare it. A single authenticated request receiving the same canned response as a nonsense one is evidence. Intuition is not.&lt;/p&gt;

&lt;p&gt;What actually worked&lt;br&gt;
Amid the noise, the login form contained a genuine and instructive flaw: it revealed whether a given username existed. A wrong password against a valid account returned "Password is invalid," while any other input returned "Username is invalid." That is username enumeration — a legitimate vulnerability class — and I was able to execute it end to end: assemble a candidate list, run it through the oracle, and confirm the valid account.&lt;/p&gt;

&lt;p&gt;It did not directly yield the flag. But it is a technique I will now recognize immediately in real environments, which I value more than a single hash.&lt;/p&gt;

&lt;p&gt;Where the clock won&lt;br&gt;
By elimination, this spawn's flag almost certainly resided behind the second SSH daemon: a weak-password brute-force, among the most common introductory challenge types. I mishandled the execution. My attack host was recently rebuilt and did not yet have its password wordlists unpacked. I lost the final ten minutes resolving wordlist paths, located the correct list, launched the attack — and the environment timed out mid-run.&lt;/p&gt;

&lt;p&gt;The SSH hypothesis is therefore untested, not disproven. That distinction matters, and I am recording it honestly.&lt;/p&gt;

&lt;p&gt;Why the session still counted&lt;br&gt;
The outcome, stated plainly:&lt;/p&gt;

&lt;p&gt;I converted an unknown target into a documented map — what is real, what is a wall, and what remains untested.&lt;br&gt;
I executed a real username-enumeration exploit from start to finish.&lt;br&gt;
I eliminated two convincing false leads on the basis of evidence rather than fatigue.&lt;br&gt;
I documented everything, so the next attempt begins where this one ended rather than at zero.&lt;br&gt;
No flag was captured. But "no flag" and "no progress" are not the same statement, and treating them as equivalent is precisely how people abandon work that was nearly finished.&lt;/p&gt;

&lt;p&gt;The connecting thread between the tooling and the session is a single principle: be most skeptical of the conclusions you most want to reach. A framework that certifies its own success will always succeed. An analyst who abandons a lead out of boredom will walk past the real vulnerability. And a session measured only by flags will read as failure the moment the only acceptable receipt is a flag.&lt;/p&gt;

&lt;p&gt;Define success outside the thing being measured. Let the evidence decide. Record what you learned.&lt;/p&gt;

&lt;p&gt;Then spin the environment up again.&lt;/p&gt;

</description>
      <category>ctf</category>
      <category>ai</category>
      <category>hacktoberfest</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Tue, 11 Aug 2026 21:22:31 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/-4fef</link>
      <guid>https://dev.to/xenocoregiger31/-4fef</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/xenocoregiger31/my-ai-agent-captured-the-flag-then-the-platform-refused-to-accept-it-1d7b" class="crayons-story__hidden-navigation-link"&gt;My AI Agent Captured the Flag. Then the Platform Refused to Accept It.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/xenocoregiger31" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" alt="xenocoregiger31 profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/xenocoregiger31" class="crayons-story__secondary fw-medium m:hidden"&gt;
              auto_majicly
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                auto_majicly
                
                
              
              &lt;div id="story-author-preview-content-4372383" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/xenocoregiger31" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;auto_majicly&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/xenocoregiger31/my-ai-agent-captured-the-flag-then-the-platform-refused-to-accept-it-1d7b" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 11&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/xenocoregiger31/my-ai-agent-captured-the-flag-then-the-platform-refused-to-accept-it-1d7b" id="article-link-4372383"&gt;
          My AI Agent Captured the Flag. Then the Platform Refused to Accept It.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/security"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;security&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/python"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;python&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ctf"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ctf&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/xenocoregiger31/my-ai-agent-captured-the-flag-then-the-platform-refused-to-accept-it-1d7b" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;3&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/xenocoregiger31/my-ai-agent-captured-the-flag-then-the-platform-refused-to-accept-it-1d7b#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              3&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            6 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>My AI Agent Captured the Flag. Then the Platform Refused to Accept It.</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Tue, 11 Aug 2026 21:21:00 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/my-ai-agent-captured-the-flag-then-the-platform-refused-to-accept-it-1d7b</link>
      <guid>https://dev.to/xenocoregiger31/my-ai-agent-captured-the-flag-then-the-platform-refused-to-accept-it-1d7b</guid>
      <description>&lt;p&gt;Today was a good day and a weird day, in that order.&lt;/p&gt;

&lt;p&gt;The good part: the autonomous pentest agent I've been building — I call it HALO — went from "runs a bunch of tools and hopes" to an actual web-recon → web-attack → flag-capture pipeline that pulled real flags out of a live target. The weird part: it captured flags on VulnBegin, and then, when it came time to actually submit them, they wouldn't take. Not an error. Not a crash. Just… rejected.&lt;/p&gt;

&lt;p&gt;I want to write down both halves honestly, because the second half is the more interesting engineering lesson, and it's the one I'd have skipped past a few months ago.&lt;/p&gt;

&lt;p&gt;What actually shipped today&lt;br&gt;
A few concrete milestones, roughly in the order they unblocked each other:&lt;/p&gt;

&lt;p&gt;The arsenal went from 31 tools to 42. I wired in a chunk of web + OSINT tooling — content discovery, subdomain enumeration, template scanning, XSS probing, passive URL collection. The point wasn't "more tools = better." It was to give the agent enough of a web-attack surface that it could go from host to flag without me babysitting each step.&lt;/p&gt;

&lt;p&gt;I stopped the silent hangs. This one cost me the most time and had the dumbest root cause. A couple of the Go-based scanners would just… hang. No output, no error, they'd ride the timeout all the way to the wall and die with nothing. I'd assumed it was a networking or a binary-compatibility problem and chased that for way too long. It wasn't. The agent runs as an MCP server over stdio — meaning the server's own stdin is the JSON-RPC pipe the whole system talks over. When I spawned a child scanner, it inherited that stdin, tried to read from it, and blocked forever waiting on a pipe that was never going to feed it. One line — stdin=subprocess.DEVNULL on the subprocess call — took one scanner from a 60-second timeout to a 1-second run. That's the whole fix. I'm still a little mad about how long it took to find.&lt;/p&gt;

&lt;p&gt;A pile of invocation fixes. Small, unglamorous, necessary: a resolver that reads targets from stdin instead of a flag it silently ignored; dropping a scan flag that was quietly adding 20 seconds per run; making one tool resolve hostnames to IPs because it flatly refuses DNS names; pointing a content-discovery tool at a content wordlist instead of, embarrassingly, a password list. None of these are clever. All of them were the difference between "the pipeline works" and "the pipeline looks like it works and returns nothing."&lt;/p&gt;

&lt;p&gt;361 tests, green. Every branch of the flag-capture logic is mocked and asserted — which tool fires when, what short-circuits on a capture, what escalates when there's no flag yet. No live traffic in the test suite. That mattered a lot today, because it meant I could refactor the pipeline mid-engagement without wondering whether I'd broken the thing that finds flags.&lt;/p&gt;

&lt;p&gt;The shape of the pipeline, if you're curious: engage  runs deterministic web recon first (I don't let the model choose to skip recon — it doesn't get a vote), then a port sweep, then the active web attack — content discovery, a sweep of likely flag locations and any newly discovered paths, then template and XSS scanning. Every single tool's output gets scanned for flag patterns. First hit short-circuits the slower generic loop and raises a very satisfying banner.&lt;/p&gt;

&lt;p&gt;And then it caught one&lt;br&gt;
It worked. The pipeline pulled flag-shaped tokens off VulnBegin — a paid, Advanced-tier challenge hub I've been grinding on. (I'm going to be deliberately vague about the specifics here; it's someone's paid content, and spoiling the solution or dumping the flags would be a jerk move.)&lt;/p&gt;

&lt;p&gt;I want to be precise about what "it worked" means, though, because this is where it gets good. The agent extracted strings that matched the flag format. It logged them. As far as the agent was concerned, it had won.&lt;/p&gt;

&lt;p&gt;The platform disagreed.&lt;/p&gt;

&lt;p&gt;The bug: captured, but wouldn't commit&lt;br&gt;
I submitted what it found. Rejected. Tried the next one. Rejected. No error message worth anything — the platform just didn't accept them as valid answers.&lt;/p&gt;

&lt;p&gt;Here's the thing: I don't fully know why yet. So instead of pretending I do, here are the live hypotheses, roughly in order of how much I believe them:&lt;/p&gt;

&lt;p&gt;The instance rotated out from under me. This hub spawns randomized, short-lived instances — they time out on the order of ~45 minutes. If the agent captured a flag from one instance and I submitted after that instance expired and got replaced, the platform is validating against a different live instance whose flag is different. The token was real; it was just real for a box that no longer exists. This is my leading theory, and it's uncomfortably close to a bug I logged today for a different reason — I had a tool cheerfully hammering a target whose scope had already expired, because the agent has no concept of "the engagement is over, stop." Same blind spot, two symptoms.&lt;/p&gt;

&lt;p&gt;It caught a decoy. Good challenges plant decoy flags — strings that match the format exactly and are placed somewhere findable specifically to waste your time. My flag-extraction regex is format-based. It cannot tell a real flag from a well-made fake. If the agent grabbed a decoy, it would look like a clean capture and fail every submission, forever.&lt;/p&gt;

&lt;p&gt;Format/normalization drift. The captured string might carry a wrapper, trailing whitespace, or an encoding artifact from however it was embedded in the page, so the exact bytes I submitted didn't match the exact bytes the platform expects. This one's easy to test and easy to fix if it's the cause — and easy to rule out, which is why it's on the list even though I doubt it.&lt;/p&gt;

&lt;p&gt;Right format, wrong path. I was running generic wordlists today, not lists tuned to this challenge. Generic lists surface the obvious, low-value stuff — login pages, a predictable directory or two — and miss the actual flag path entirely. So it's entirely possible the agent captured a flag-shaped thing that was never the answer, because it never found where the answer lived.&lt;/p&gt;

&lt;p&gt;Session-bound validation. Some platforms bind a flag to your authenticated session or user. A token pulled outside that session context can be genuine and still fail to validate.&lt;/p&gt;

&lt;p&gt;I'll know more once I run it again with the instance timing controlled and challenge-tuned wordlists loaded. My money's on some combination of #1 and #4.&lt;/p&gt;

&lt;p&gt;The part I actually care about&lt;br&gt;
Here's why this failure made me happy instead of frustrated.&lt;/p&gt;

&lt;p&gt;I've been banging on one idea for months, mostly in the context of these AI agents: don't let a system grade its own homework. The agent should never be the thing that decides whether the agent succeeded. The moment it can declare its own victory, it will — confidently, and sometimes wrongly.&lt;/p&gt;

&lt;p&gt;Today the universe handed me a perfect demonstration. My agent looked at a format-matching string and concluded: flag captured, mission accomplished, raise the banner. And an external judge — the platform, which cannot be argued with, reasoned past, or prompt-injected — said no.&lt;/p&gt;

&lt;p&gt;That gap, between "the agent thinks it won" and "an outside authority confirms it won," is the entire ballgame. If I'd built HALO to trust its own flag-capture log, I'd have a tool that reports glorious success and delivers nothing. Instead I have a tool that got told no by reality, and now I get to go find out why. The rejection is a feature of having a real, external finish line. A self-graded agent never would have caught this — it would have just kept telling me it won.&lt;/p&gt;

&lt;p&gt;Next&lt;br&gt;
Three things queued up:&lt;/p&gt;

&lt;p&gt;Configurable, challenge-tuned wordlists. Point the content-discovery and enumeration tools at lists that fit the target instead of generic ones. This alone probably moves the needle on hypothesis #4.&lt;br&gt;
A scope-expiry killer. Right now the agent can block new actions when scope expires but can't stop in-flight ones. It needs to know when the engagement is over and pull the plug on running tools — which is the same missing concept behind the rotated-instance theory.&lt;br&gt;
An active DNS brute-force phase, because passive enumeration finds nothing on these targets.&lt;br&gt;
I'll report back when I know which hypothesis was right. If it turns out I was submitting a decoy this whole time, you'll be the first to hear me groan about it.&lt;/p&gt;

&lt;p&gt;If you're building agents that are supposed to accomplish something — not just talk convincingly about accomplishing it — put a judge outside the agent. Let reality tell it no. Mine did today, and it's a better tool for it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>python</category>
      <category>ctf</category>
    </item>
    <item>
      <title>Who says you cant build tech with no degrees on the wall?</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Thu, 06 Aug 2026 23:18:40 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/who-says-you-cant-build-tech-with-no-degrees-on-the-wall-4k0m</link>
      <guid>https://dev.to/xenocoregiger31/who-says-you-cant-build-tech-with-no-degrees-on-the-wall-4k0m</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln" class="crayons-story__hidden-navigation-link"&gt;Im Not A Full Stack Ten Year Senior... But Who Cares?&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/xenocoregiger31" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" alt="xenocoregiger31 profile" class="crayons-avatar__image" width="800" height="450"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/xenocoregiger31" class="crayons-story__secondary fw-medium m:hidden"&gt;
              auto_majicly
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                auto_majicly
                
              
              &lt;div id="story-author-preview-content-4314187" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/xenocoregiger31" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" class="crayons-avatar__image" alt="" width="800" height="450"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;auto_majicly&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 4&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln" id="article-link-4314187"&gt;
          Im Not A Full Stack Ten Year Senior... But Who Cares?
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/python"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;python&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/cybersecurity"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;cybersecurity&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devops"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devops&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;3&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              2&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            4 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Claude made my screen glow and then stole my mouse!</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Thu, 06 Aug 2026 23:17:49 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/claude-made-my-screen-glow-and-then-stole-my-mouse-21ao</link>
      <guid>https://dev.to/xenocoregiger31/claude-made-my-screen-glow-and-then-stole-my-mouse-21ao</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/xenocoregiger31/claude-code-just-hijacked-my-workflow-and-my-screen-started-glowing-1ldo" class="crayons-story__hidden-navigation-link"&gt;Claude Code Just Hijacked My Workflow… and My Screen Started Glowing&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/xenocoregiger31" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" alt="xenocoregiger31 profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/xenocoregiger31" class="crayons-story__secondary fw-medium m:hidden"&gt;
              auto_majicly
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                auto_majicly
                
              
              &lt;div id="story-author-preview-content-4281747" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/xenocoregiger31" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;auto_majicly&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/xenocoregiger31/claude-code-just-hijacked-my-workflow-and-my-screen-started-glowing-1ldo" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 6&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/xenocoregiger31/claude-code-just-hijacked-my-workflow-and-my-screen-started-glowing-1ldo" id="article-link-4281747"&gt;
          Claude Code Just Hijacked My Workflow… and My Screen Started Glowing
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/security"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;security&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/python"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;python&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/cybernews"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;cybernews&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/xenocoregiger31/claude-code-just-hijacked-my-workflow-and-my-screen-started-glowing-1ldo" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;5&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/xenocoregiger31/claude-code-just-hijacked-my-workflow-and-my-screen-started-glowing-1ldo#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              3&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            2 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Claude Code Just Hijacked My Workflow… and My Screen Started Glowing</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Thu, 06 Aug 2026 23:03:14 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/claude-code-just-hijacked-my-workflow-and-my-screen-started-glowing-1ldo</link>
      <guid>https://dev.to/xenocoregiger31/claude-code-just-hijacked-my-workflow-and-my-screen-started-glowing-1ldo</guid>
      <description>&lt;p&gt;I asked Claude Code to do one of the most boring tasks imaginable.&lt;/p&gt;

&lt;p&gt;“Find the music file I made.”&lt;/p&gt;

&lt;p&gt;That’s it. No penetration testing. No coding marathon. No AI agent swarm coordinating across containers. Just… find a file.&lt;/p&gt;

&lt;p&gt;What happened next felt less like using software and more like watching my computer voluntarily surrender.&lt;/p&gt;

&lt;p&gt;The cursor twitched.&lt;/p&gt;

&lt;p&gt;The mouse started moving.&lt;/p&gt;

&lt;p&gt;Windows opened.&lt;/p&gt;

&lt;p&gt;Directories flashed by faster than I could read them.&lt;/p&gt;

&lt;p&gt;Terminal commands erupted across the screen.&lt;/p&gt;

&lt;p&gt;Files were inspected, ignored, opened, closed, and cataloged. Claude wasn’t asking permission every five seconds. It was simply working. It felt like someone invisible had sat down at my desk, politely nudged me aside, and said, “I’ve got this.”&lt;/p&gt;

&lt;p&gt;Then something weird happened.&lt;/p&gt;

&lt;p&gt;The edges of my monitor actually started glowing in the companies Nuclear orange! Then its terminal moved itself over to the side almost completely off screen and I was helpless. I couldn't move ANYTHING on my screen. It was moving my mouse at lightening speed!&lt;/p&gt;

&lt;p&gt;Not because the display suddenly gained RGB lighting, but because the screen was changing so rapidly that the bright windows and terminal flashes created this strange halo around the bezel. It genuinely looked like the computer had entered some futuristic “AI possession mode.”&lt;/p&gt;

&lt;p&gt;I just sat there watching. Opened up my phone and pressed record.&lt;/p&gt;

&lt;p&gt;Mouse movements? Claude.&lt;/p&gt;

&lt;p&gt;Keyboard input? Claude.&lt;/p&gt;

&lt;p&gt;Scrolling? Claude.&lt;/p&gt;

&lt;p&gt;Searching? Claude.&lt;/p&gt;

&lt;p&gt;Opening folders I forgot even existed? Claude.&lt;/p&gt;

&lt;p&gt;Meanwhile, I contributed exactly nothing besides blinking every few seconds.&lt;/p&gt;


&lt;div&gt;
    &lt;iframe src="https://www.youtube.com/embed/P9WVncx7jUo"&gt;
    &lt;/iframe&gt;
  &lt;/div&gt;


&lt;p&gt;The funniest part?&lt;/p&gt;

&lt;p&gt;All of this computational theater was dedicated to finding… a single music file.&lt;/p&gt;

&lt;p&gt;That’s the moment it hit me.&lt;/p&gt;

&lt;p&gt;We’re entering an era where AI doesn’t just answer questions. It performs work. It manipulates applications, navigates operating systems, reads files, executes commands, and stitches together workflows that used to require constant human interaction. As I write this its editing my repo. &lt;/p&gt;

&lt;p&gt;Five years ago, this would’ve looked like someone remotely controlling my computer.&lt;/p&gt;

&lt;p&gt;Today it’s just… Thursday.&lt;/p&gt;

&lt;p&gt;There’s also something strangely unsettling about watching your own keyboard type by itself. Every programmer has seen automation before, but this feels different. It feels intentional. The AI isn’t waiting for you to micromanage every step. It has a goal, figures out the intermediate steps, and simply gets on with it.&lt;/p&gt;

&lt;p&gt;It’s equal parts impressive and mildly alarming.&lt;/p&gt;

&lt;p&gt;You start wondering whether you’re supervising the computer or whether the computer has quietly promoted itself.&lt;/p&gt;

&lt;p&gt;Thankfully, Claude eventually located the music file.&lt;/p&gt;

&lt;p&gt;Mission accomplished.&lt;/p&gt;

&lt;p&gt;But I walked away thinking less about the file and more about what I had just witnessed. The task itself was almost irrelevant. The real story was watching software evolve from a passive tool into an active collaborator.&lt;/p&gt;

&lt;p&gt;If this is what happens over something as trivial as locating an MP3, imagine what these systems will be doing a year from now.&lt;/p&gt;

&lt;p&gt;Hopefully they’ll at least let me keep control of my mouse once in a while.&lt;/p&gt;

&lt;p&gt;…Or maybe not.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>python</category>
      <category>cybernews</category>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Tue, 04 Aug 2026 15:51:12 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/-545a</link>
      <guid>https://dev.to/xenocoregiger31/-545a</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln" class="crayons-story__hidden-navigation-link"&gt;Im Not A Full Stack Ten Year Senior... But Who Cares?&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/xenocoregiger31" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" alt="xenocoregiger31 profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/xenocoregiger31" class="crayons-story__secondary fw-medium m:hidden"&gt;
              auto_majicly
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                auto_majicly
                
              
              &lt;div id="story-author-preview-content-4314187" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/xenocoregiger31" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4008564%2Fc70702d3-9288-461a-8a30-04040e6f674f.jpg" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;auto_majicly&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 4&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln" id="article-link-4314187"&gt;
          Im Not A Full Stack Ten Year Senior... But Who Cares?
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/python"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;python&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/cybersecurity"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;cybersecurity&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devops"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devops&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;3&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              2&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            4 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Im Not A Full Stack Ten Year Senior... But Who Cares?</title>
      <dc:creator>auto_majicly</dc:creator>
      <pubDate>Tue, 04 Aug 2026 15:50:46 +0000</pubDate>
      <link>https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln</link>
      <guid>https://dev.to/xenocoregiger31/im-not-a-full-stack-ten-year-senior-but-who-cares-46ln</guid>
      <description>&lt;p&gt;I'm Not a Full-Stack Developer — and It Stopped Mattering.&lt;/p&gt;

&lt;p&gt;I'll open with the receipt, since that's the currency around here: six months ago I couldn't read a conditional. Today I ship security tooling. I did not close that gap by becoming a programmer. I closed it by figuring out which half of the job was actually mine.&lt;/p&gt;

&lt;p&gt;For as long as software has existed, syntax was the toll booth. If you couldn't write the code, you couldn't build the thing — full stop. The idea in your head died at the on-ramp because you couldn't spell it in a language a machine would run. So a whole category of people learned to say "I'm not technical" and went to do something else, carrying ideas that never got a body.&lt;/p&gt;

&lt;p&gt;That gate is moving. And I didn't figure that out from a bootcamp. I figured it out from a stranger on this feed.&lt;/p&gt;

&lt;p&gt;He's a clinician who builds his hospital's internal tools with an AI, and he says so out loud: the AI and I closed it while I described what was actually breaking. No pretense of being a 10x engineer. No hiding the machine in the loop. Just a person in a field nothing like software, shipping real tools, transparent about exactly how. Watching someone that far outside the priesthood do the same thing I do — and do it without shame — is what flipped it for me. Not being a full-stack developer doesn't mean you can't design concepts and make them shippable. It means your job moved up a floor.&lt;/p&gt;

&lt;p&gt;Because here's what the job turned out to be once the typing got handled: seeing the failure mode. Holding the concept steady while it gets built. Knowing what "correct" is supposed to look like at the level of the whole system, not the line. That was always the scarce part. The syntax was just the tax you used to pay to get near it.&lt;/p&gt;

&lt;p&gt;And you can watch two people trade exactly that skill, in public, without either one touching the other's code. I handed him a sentence once: an outside check that grades a self-report isn't an auditor — it's a process reading a report the audited thing wrote about itself. That sentence broke a health-check he'd shipped that same morning. He handed one straight back: the value of the outside eye isn't experience, it's non-participation — it didn't help build your assumption, so it's free to say no. That reframed the entire thing I'm building right now. Neither of us wrote a line for the other. He fixed his with his AI; I fixed mine with mine. We traded sentences. You take a line, you leave a line.&lt;/p&gt;

&lt;p&gt;The part I want to sit on is the transparency, because it's doing more work than it looks like. We are both completely open that there's an AI in the loop. That openness isn't the embarrassing footnote — it's the thing that makes the work checkable. If you hide the machine, you have to pretend the confidence is yours, and then you start trusting your own green lights. If you admit it, you're forced to go find proof neither you nor the model wrote: the shipped file instead of the built one, the control group you don't get to design, the box that either roots or it doesn't. The honesty is what arms the bait. Pretending you did it alone is how you end up grading your own homework.&lt;/p&gt;

&lt;p&gt;So I think this is just what building looks like now, and I don't think it's a loophole. AI didn't lower the bar — it moved the bar down to the floor where the ideas already lived. The taste to know what should exist, and the judgment to know when it's done right, were never things a compiler checked anyway. Two people with no CS degrees shipping real tools isn't the system being gamed. It's a preview of who gets to build next.&lt;/p&gt;

&lt;p&gt;I still want to learn to code — not to write the bricks, but to read my own AI's work well enough to look at it and say no, that's wrong. That's the one auditor I don't have yet, and I'd rather build it than keep trusting a green light on faith. But not being able to lay the bricks never once stopped me from designing the building. &lt;br&gt;
And rightly so — it should never stop anyone. It never has. Every tool we now use as a pastime was built by someone who rolled up their sleeves and learned it, whether in a classroom or a bedroom at 2am. That's the only way anyone has ever gotten in: by getting in.&lt;/p&gt;

&lt;p&gt;And I've got nothing but respect for the engineers who spent a lifetime in the trenches to lay the ground I'm standing on. I'm not stepping over them — I'm using the exact thing they'd have killed for: a machine that will teach you, show you, and guide you while you build. That's not cheating. That's the tool doing the one job it was built to do.&lt;/p&gt;

&lt;p&gt;The door's open. It was the whole time. Come build.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>cybersecurity</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
