<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Muhammad Nauman Hafeez</title>
    <description>The latest articles on DEV Community by Muhammad Nauman Hafeez (@naum4n).</description>
    <link>https://dev.to/naum4n</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4126340%2Fb9016376-9587-4c7b-8314-761aa019b2d2.jpg</url>
      <title>DEV Community: Muhammad Nauman Hafeez</title>
      <link>https://dev.to/naum4n</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/naum4n"/>
    <language>en</language>
    <item>
      <title>Why Your Startup Needs a DevOps Engineer (Even at 5 People)</title>
      <dc:creator>Muhammad Nauman Hafeez</dc:creator>
      <pubDate>Mon, 05 Oct 2026 14:33:02 +0000</pubDate>
      <link>https://dev.to/naum4n/why-your-startup-needs-a-devops-engineer-even-at-5-people-381c</link>
      <guid>https://dev.to/naum4n/why-your-startup-needs-a-devops-engineer-even-at-5-people-381c</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fipbmall3oqmd9sbz8757.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fipbmall3oqmd9sbz8757.png" alt=" " width="800" height="420"&gt;&lt;/a&gt;&lt;br&gt;
Most startup founders think DevOps is something you hire for later - when you're bigger, when you have&lt;br&gt;
product-market fit, when you can afford it.&lt;br&gt;
That's backwards.&lt;br&gt;
Here's why DevOps from day one is a competitive advantage, not a luxury.&lt;/p&gt;

&lt;h2&gt;
  
  
  The "We'll Deal With It Later" Trap
&lt;/h2&gt;

&lt;p&gt;We've seen this story dozens of times:&lt;br&gt;
A startup launches with a scrappy infrastructure. Deployments are manual. The "CI/CD pipeline" is&lt;br&gt;
someone running commands on their laptop. AWS credentials are shared in Slack. Monitoring is "we'll&lt;br&gt;
check if it's down when customers complain."&lt;br&gt;
It works... until it doesn't.&lt;br&gt;
Then comes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The first major outage&lt;/li&gt;
&lt;li&gt;The security incident&lt;/li&gt;
&lt;li&gt;The AWS bill that's 3x what it should be&lt;/li&gt;
&lt;li&gt;The week where no features shipped because everyone was fighting fires&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The cost of fixing these problems later is 10x higher than preventing them early.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What a DevOps Engineer Actually Does for Startups
&lt;/h2&gt;

&lt;p&gt;DevOps isn't about fancy tools or complex architectures. For startups, it's about:### 1. Making Deployments Boring&lt;br&gt;
Your developers should ship code with a single click. No SSH-ing into servers. No "it works on my&lt;br&gt;
machine." No 2-hour deployment windows.&lt;br&gt;
&lt;strong&gt;Result:&lt;/strong&gt; Ship 10x more often with 10x fewer headaches.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Preventing the 3 AM Wake-Up Call
&lt;/h3&gt;

&lt;p&gt;Proper monitoring, alerting, and auto-scaling means problems are caught before customers notice.&lt;br&gt;
&lt;strong&gt;Result:&lt;/strong&gt; Sleep better. Keep customers happy.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Keeping AWS Bills Under Control
&lt;/h3&gt;

&lt;p&gt;We've seen startups burning $15k/month when $5k would do the same job. Right-sizing, spot instances,&lt;br&gt;
cleaning up orphaned resources.&lt;br&gt;
&lt;strong&gt;Result:&lt;/strong&gt; 30-50% cost reduction is typical.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Not Being the Next Breach Headline
&lt;/h3&gt;

&lt;p&gt;IAM roles instead of shared credentials. Secrets in a vault, not in code. Network security that actually&lt;br&gt;
makes sense.&lt;br&gt;
&lt;strong&gt;Result:&lt;/strong&gt; Security that doesn't slow you down.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Letting Developers Actually Develop
&lt;/h3&gt;

&lt;p&gt;When developers spend 20% of their time on infrastructure tasks, that's 20% not spent on features.&lt;br&gt;
&lt;strong&gt;Result:&lt;/strong&gt; Happier developers, faster product iteration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers: DevOps ROI for Startups
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Without DevOps&lt;/th&gt;
&lt;th&gt;With DevOps&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Deployment frequency&lt;/td&gt;
&lt;td&gt;Weekly or less&lt;/td&gt;
&lt;td&gt;Multiple times daily&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deployment time&lt;/td&gt;
&lt;td&gt;2-4 hours&lt;/td&gt;
&lt;td&gt;10 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Change failure rate&lt;/td&gt;
&lt;td&gt;15-30%&lt;/td&gt;
&lt;td&gt;Less than 5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recovery time&lt;/td&gt;
&lt;td&gt;Hours to days&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer time on infra&lt;/td&gt;
&lt;td&gt;20-30%&lt;/td&gt;
&lt;td&gt;Less than 5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud cost efficiency&lt;/td&gt;
&lt;td&gt;50-60%&lt;/td&gt;
&lt;td&gt;85-95%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  When to Bring in DevOps
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;You need DevOps NOW if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Deployments are manual or scary&lt;/li&gt;
&lt;li&gt;You don't know what's running in production- X Your AWS bill surprises you every month&lt;/li&gt;
&lt;li&gt;Developers SSH into production to debug&lt;/li&gt;
&lt;li&gt;There's no staging environment&lt;/li&gt;
&lt;li&gt;You've had an outage with no clear root cause&lt;/li&gt;
&lt;li&gt;Credentials are shared or in code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;You can wait if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're pre-launch and iterating on a prototype&lt;/li&gt;
&lt;li&gt;You're a single developer on a side project&lt;/li&gt;
&lt;li&gt;You're using a fully managed platform (Heroku, Vercel)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Full-Time vs. Consultant
&lt;/h2&gt;

&lt;p&gt;Not every startup needs a full-time DevOps engineer:&lt;br&gt;
| Option | Best For | Cost |&lt;br&gt;
|--------|----------|------|&lt;br&gt;
| &lt;strong&gt;Full-time hire&lt;/strong&gt; | 10+ engineers, complex infrastructure | $150-250k/year |&lt;br&gt;
| &lt;strong&gt;Fractional/retainer&lt;/strong&gt; | 5-15 engineers, need ongoing support | $3-8k/month |&lt;br&gt;
| &lt;strong&gt;Project-based&lt;/strong&gt; | Set up infrastructure right, then maintain in-house | $10-50k one-time |&lt;br&gt;
| &lt;strong&gt;Consulting/audit&lt;/strong&gt; | Know what's broken and prioritize fixes | $0-5k |&lt;br&gt;
Most seed-stage startups should start with a &lt;strong&gt;project engagement&lt;/strong&gt; to set up the foundation, then move&lt;br&gt;
to a &lt;strong&gt;light retainer&lt;/strong&gt; for ongoing support.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Good Startup DevOps Looks Like
&lt;/h2&gt;

&lt;p&gt;You don't need Kubernetes on day one. Good early-stage DevOps is simple:&lt;br&gt;
Infrastructure as code (Terraform)&lt;br&gt;
CI/CD pipeline - automated tests and deployments&lt;br&gt;
Staging environment&lt;br&gt;
Monitoring &amp;amp; alerting&lt;br&gt;
Secrets management&lt;br&gt;
Backups - tested and automated&lt;br&gt;
Basic security&lt;br&gt;
Documentation&lt;/p&gt;

&lt;p&gt;That's it. No service mesh, no multi-region active-active. Those come later if you actually need them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Cost of Waiting
&lt;/h2&gt;

&lt;p&gt;We worked with a startup that waited until Series A to address DevOps. &lt;br&gt;
By then:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AWS bill had grown to $28k/month (should have been $12k)&lt;/li&gt;
&lt;li&gt;Deployments took 3 hours and failed 30% of the time&lt;/li&gt;
&lt;li&gt;Two major outages, one during a sales demo&lt;/li&gt;
&lt;li&gt;Three developers spent 40% of their time on infra&lt;/li&gt;
&lt;li&gt;Security audit revealed credentials in GitHub historyCleaning this up took 3 months and $45k.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;If they'd invested $15k in DevOps at seed stage, they'd have saved $100k+.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting Started
&lt;/h2&gt;

&lt;p&gt;If you're a startup founder or CTO:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Get an audit.&lt;/strong&gt; A good DevOps consultant will tell you what's broken and what's urgent. Many offer
this free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix the foundations.&lt;/strong&gt; CI/CD, infrastructure as code, monitoring. This is a 2-4 week project.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Establish ongoing support.&lt;/strong&gt; Fractional engagement, training, or self-service docs.
The goal isn't perfect infrastructure. It's infrastructure that doesn't slow you down, doesn't surprise you
with costs, and doesn't wake you up at 3 AM.
---
&lt;em&gt;What's your experience with DevOps at a startup? Did you hire early or wish you had? Let me know in
the comments!&lt;/em&gt;
---&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Originally published at
&lt;a href="https://orbucx.com/blog/why-your-startup-needs-devops-engineer.html" rel="noopener noreferrer"&gt;orbucx.com&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We offer free cloud audits for startups - &lt;a href="https://orbucx.com/#contact" rel="noopener noreferrer"&gt;book yours here&lt;/a&gt;.&lt;/p&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>devops</category>
      <category>startup</category>
      <category>aws</category>
      <category>cloud</category>
    </item>
    <item>
      <title>A Supply-Chain Attack Hit Our Repos. Here's How We Locked Down GitHub So It Can't Happen Again</title>
      <dc:creator>Muhammad Nauman Hafeez</dc:creator>
      <pubDate>Wed, 23 Sep 2026 08:57:14 +0000</pubDate>
      <link>https://dev.to/naum4n/a-supply-chain-attack-hit-our-repos-heres-how-we-locked-down-github-so-it-cant-happen-again-57gb</link>
      <guid>https://dev.to/naum4n/a-supply-chain-attack-hit-our-repos-heres-how-we-locked-down-github-so-it-cant-happen-again-57gb</guid>
      <description>&lt;p&gt;&lt;em&gt;This is the follow-up to my previous post: &lt;a href="https://dev.to/naum4n/a-poisoned-npm-package-infected-our-production-server-heres-how-we-found-and-removed-it-2peo"&gt;A Poisoned npm Package Infected Our Production Server, Here's How We Found and Removed It&lt;/a&gt;. If you haven't read that one, start there — it covers the incident itself. This post covers what we changed afterwards so it can't happen again.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;A few weeks ago I spent three of the worst days of my (so far short) DevOps career hunting malware on a production server. A poisoned npm package had made it into one of our builds, and on top of that, an entire repository on our GitHub org had been force-replaced in a single commit — whole codebase swapped out, new &lt;code&gt;package.json&lt;/code&gt;, new CI workflow, everything.&lt;/p&gt;

&lt;p&gt;We cleaned the server. We rotated every credential. But when the adrenaline wore off, one question kept bothering me:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why was it even &lt;em&gt;possible&lt;/em&gt; for one push to replace an entire repository?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The answer was embarrassing: because nothing was stopping it. No branch protection. No required reviews. Anyone with write access could force-push to any branch, delete branches, and push straight to production-deploying branches. We'd been running on trust, and trust doesn't scale — especially not when supply-chain attacks are involved.&lt;/p&gt;

&lt;p&gt;This post is about how we fixed that with GitHub rulesets. I'm about 2.5 years into DevOps, so this isn't a "grizzled veteran" take — it's a "we got burned and here's exactly what we did about it" take. Which, honestly, might be more useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  The goal
&lt;/h2&gt;

&lt;p&gt;After the incident, we defined what a "safe" repo looks like for us:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;No force pushes.&lt;/strong&gt; History rewriting is how the attack rewrote our reality. Never again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No branch deletion.&lt;/strong&gt; Deleting a branch shouldn't be a one-click way to destroy evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every change goes through a pull request.&lt;/strong&gt; No direct pushes to protected branches — even from admins.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;At least one approving review.&lt;/strong&gt; A second pair of eyes on every change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nobody bypasses the rules.&lt;/strong&gt; Not admins. Not the org owner. Nobody.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last one turned out to be the most interesting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we chose rulesets over classic branch protection
&lt;/h2&gt;

&lt;p&gt;GitHub has two mechanisms for this: the older &lt;strong&gt;branch protection rules&lt;/strong&gt; and the newer &lt;strong&gt;rulesets&lt;/strong&gt;. We went with rulesets, for a few reasons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Layering:&lt;/strong&gt; multiple rulesets can apply at once, and GitHub enforces the union of all of them. You can have a baseline ruleset plus a stricter one for release branches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Better targeting:&lt;/strong&gt; one ruleset can target all branches (&lt;code&gt;~ALL&lt;/code&gt;), just the default branch (&lt;code&gt;~DEFAULT&lt;/code&gt;), or name patterns like &lt;code&gt;release/*&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explicit bypass list:&lt;/strong&gt; rulesets have a bypass list that is &lt;em&gt;empty by default&lt;/em&gt;. Classic branch protection has the infamous "Do not allow bypassing the above settings" checkbox that everyone forgets to tick — which silently exempts admins.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Visibility:&lt;/strong&gt; anyone with read access to the repo can see which rules apply. When a push gets rejected, the error message names the ruleset. No mystery.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Classic branch protection still works, but if you're setting things up fresh in 2026, rulesets are the better tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ruleset we deployed
&lt;/h2&gt;

&lt;p&gt;Here's the exact configuration we now apply to every repo in the org. The important top-level decisions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Enforcement:&lt;/strong&gt; Active (not "Evaluate" — evaluate mode only logs, it doesn't block)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Target branches:&lt;/strong&gt; All branches (&lt;code&gt;~ALL&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bypass list:&lt;/strong&gt; empty. This is the whole point.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The full JSON (you can import this directly — steps below):&lt;/p&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
json
{
  "name": "org-branch-guard",
  "target": "branch",
  "enforcement": "active",
  "conditions": {
    "ref_name": {
      "include": ["~ALL"],
      "exclude": []
    }
  },
  "bypass_actors": [],
  "rules": [
    { "type": "non_fast_forward" },
    { "type": "deletion" },
    {
      "type": "pull_request",
      "parameters": {
        "required_approving_review_count": 1,
        "dismiss_stale_reviews_on_push": true,
        "require_last_push_approval": true,
        "required_review_thread_resolution": true,
        "require_code_owner_review": false,
        "automatic_copilot_code_review_enabled": false,
        "allowed_merge_methods": ["merge", "squash", "rebase"]
      }
    }
  ]
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>github</category>
      <category>devops</category>
      <category>security</category>
      <category>npm</category>
    </item>
    <item>
      <title>A Poisoned npm Package Infected Our Production Server, Here's How We Found and Removed It</title>
      <dc:creator>Muhammad Nauman Hafeez</dc:creator>
      <pubDate>Wed, 16 Sep 2026 07:34:56 +0000</pubDate>
      <link>https://dev.to/naum4n/a-poisoned-npm-package-infected-our-production-server-heres-how-we-found-and-removed-it-2peo</link>
      <guid>https://dev.to/naum4n/a-poisoned-npm-package-infected-our-production-server-heres-how-we-found-and-removed-it-2peo</guid>
      <description>&lt;p&gt;How a routine deploy infected our CI server, how we hunted it down, and what actually fixed it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The setup&lt;/strong&gt;&lt;br&gt;
Our on-prem Ubuntu server runs ~20 GitHub Actions self-hosted runners and a dozen Next.js apps under pm2. One shared box, many repos — remember that detail, it matters later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 0: "Is this suspicious?"&lt;/strong&gt;&lt;br&gt;
While running a routine security audit, one check jumped out — two strange processes:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;node -e global.i='7-v1729';global.e='NPM';global.r=require;...&lt;/code&gt;&lt;br&gt;
Heavily obfuscated JavaScript, running as our deploy user, with established HTTPS connections to an unknown IP: 181.214.149.148:443.&lt;/p&gt;

&lt;p&gt;Deobfuscated, the loader did one thing: fetch JSON from a remote server and eval() whatever came back. A classic remote-code stager — the attacker could execute anything, anytime, on our production box.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The forensics trail&lt;/strong&gt;&lt;br&gt;
What we ruled out, step by step:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;❌ SSH brute-force — auth logs clean, all logins from our LAN&lt;/li&gt;
&lt;li&gt;❌ Rogue SSH keys, extra root accounts, cron/systemd persistence — clean&lt;/li&gt;
&lt;li&gt;❌ Malicious install scripts — we reinstalled with ignore-scripts=true… and it came back anyway&lt;/li&gt;
&lt;li&gt;❌ App source code &amp;amp; workflow file — clean&lt;/li&gt;
&lt;li&gt;❌ NODE_OPTIONS/env injection, npmrc tampering, pm2 modules — clean&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What we confirmed:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;✅ The loaders spawned during npm install + build of one specific app, delivered through a CI deploy&lt;/li&gt;
&lt;li&gt;✅ The parent chain: pm2 → npm start → next-server → malware. Every app boot respawned it&lt;/li&gt;
&lt;li&gt;✅ global.e='NPM' — the malware itself tags its delivery vector: a poisoned npm package&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The scariest moment: after a full wipe (rm -rf node_modules .next + fresh install), the malware respawned within a minute — spawned mid-install, then orphaned itself to PPID 1 so it survived pm2 stop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The resolution&lt;/strong&gt;&lt;br&gt;
We never got the exact package name — and that's a realistic detail worth sharing. Here's what the evidence showed: a dependency version that was live-poisoned on the npm registry on deploy day, cached locally, and pulled/patched upstream within 48 hours (this is how real supply-chain waves work — npm yanks compromised versions fast). Once we cleaned the npm cache and reinstalled fresh, the exact same sequence came back clean — verified end-to-end with kernel-level auditd logging armed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What actually fixed it&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Contain: iptables DROP on the C2 IP (made persistent!), stop the infected app, kill the orphaned loaders&lt;/li&gt;
&lt;li&gt;Freeze the vector: stop ALL runners — deploy freeze until clean&lt;/li&gt;
&lt;li&gt;Wipe properly: rm -rf node_modules .next plus npm cache clean --force — the poisoned tarballs live in ~/.npm/_cacache and reinfect you from cache&lt;/li&gt;
&lt;li&gt;Rotate everything: GitHub PATs, runner tokens, org secrets, every .env, SSH passwords. The malware had 24h of access — cleanup does not un-steal secrets&lt;/li&gt;
&lt;li&gt;Verify with a real deploy, then bring runners back one at a time&lt;/li&gt;
&lt;li&gt;Schedule a full host rebuild — after a remote-code stager, "no evidence of a second stage" ≠ "no second stage"&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Lessons learned&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ignore-scripts=true is not enough. It blocks install hooks but NOT require-time payloads that run when your app boots.&lt;/li&gt;
&lt;li&gt;Self-hosted runners are a blast radius multiplier. One poisoned dependency in ONE repo = full compromise of every project on the box. Isolate runners.&lt;/li&gt;
&lt;li&gt;npm's cache can reinfect you. Wiping node_modules without npm cache clean is not a clean install.&lt;/li&gt;
&lt;li&gt;Orphaned processes survive pm2 restarts. Check for PPID 1 processes you don't recognize.&lt;/li&gt;
&lt;li&gt;Block C2 IPs at the firewall AND persist the rule — an in-memory iptables rule dies on reboot.&lt;/li&gt;
&lt;li&gt;Rotate credentials immediately. This is the step everyone delays and the one that matters most.&lt;/li&gt;
&lt;li&gt;A ufw status that says inactive on a production server is a finding in itself. 🙃&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This was my first incident of this kind — if you run self-hosted runners on a shared box, I hope you never need this post. But if you do: contain, freeze, wipe (cache too!), rotate, verify, rebuild.&lt;/p&gt;

</description>
      <category>security</category>
      <category>devops</category>
      <category>node</category>
      <category>cloud</category>
    </item>
  </channel>
</rss>
