<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community</title>
    <description>The most recent home feed on DEV Community.</description>
    <link>https://dev.to</link>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed"/>
    <language>en</language>
    <item>
      <title>Management was never a promotion</title>
      <dc:creator>Matt Cockayne</dc:creator>
      <pubDate>Sun, 04 Oct 2026 02:05:40 +0000</pubDate>
      <link>https://dev.to/phpboyscout/management-was-never-a-promotion-1deh</link>
      <guid>https://dev.to/phpboyscout/management-was-never-a-promotion-1deh</guid>
      <description>&lt;p&gt;The man who hired me into one of my more recent jobs opened the interview with a question I've never quite managed to forget: why did I want a job I was clearly overqualified for?&lt;/p&gt;

&lt;p&gt;I don't remember my answer word for word. I do remember thinking it was a perfectly fair question, if you read a CV the way most companies draw their org chart... with management sat on the top and everything else somewhere underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ladder most places draw
&lt;/h2&gt;

&lt;p&gt;You'll know the shape. Junior, mid, senior, maybe a lead if you're lucky, and then the engineering rungs just sort of stop. Anything above that has "manager" or "head of" in the title, so the only way to keep going up is to stop doing the thing that got you there in the first place. Plenty of very good engineers take that step because it's the only step on offer, and some of them find they love it, and some of them find out a year later that they've swapped a job they were brilliant at for one they never wanted (the Peter principle, more or less, with a pay rise to soften the landing).&lt;/p&gt;

&lt;p&gt;I don't think management is the villain in that story. It's a proper job, and a hard one, and I've done enough of it to know it has burdens the engineering side never sees. My problem is with the drawing. Management and engineering are two sides of the same coin, and far too many companies prioritise the management track and simply forget there's no cap on the engineering one. Management is a choice. It was never a promotion.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cap I built myself
&lt;/h2&gt;

&lt;p&gt;The first time I hit that ceiling, the daft thing is I'd put it there myself.&lt;/p&gt;

&lt;p&gt;I started my own boutique digital agency, Zucchi, and became its MD, which was a choice and a big one. It also pushed me about as far away from engineering as I've ever been and flooded my brain with management problems. It turns out being MD of a small company is mostly sales and HR, with a bit of engineering squeezed in round the edges if you're lucky, and (as I've &lt;a href="https://phpboyscout.uk/what-burnout-taught-me/" rel="noopener noreferrer"&gt;written about before&lt;/a&gt;) while I like to think I'm a decent engineer, I'm a shit salesman.&lt;/p&gt;

&lt;p&gt;What I learned there wasn't that management is bad. I was good at a lot of it! It was that the two jobs pull in different directions, and trying to be both at once, in one role, is a very hard thing to reconcile. The top of the only ladder in the company was a job I'd never have applied for... and I'd drawn that ladder myself.&lt;/p&gt;

&lt;h2&gt;
  
  
  A channel, not a waiting room
&lt;/h2&gt;

&lt;p&gt;Every company where I've had any say since, I've pushed for a fully fledged engineering channel, so people can keep growing without being shoved into management to do it. The titles varied from company to company, but the point was always the same: a way up that doesn't go through a management role.&lt;/p&gt;

&lt;p&gt;And that's also where I learned the other half of the problem, which is that everybody chases the title. I find this massively frustrating. We'd take on graduates straight out of university as juniors, and within twelve to twenty-four months a fair few of them would either be declaring they should be made senior, or leaving for somewhere that would hand them the word, as if the engineering track was a quick route up just by turning up for long enough.&lt;/p&gt;

&lt;p&gt;That properly incensed me. Growth should be something you can show, from the effort you put in and the value you bring to the team, and time served is neither of those. A track with no cap isn't a waiting room where the titles come round on a timer. It needs the rungs written down, so that everyone (the engineer, their lead, the person signing off the pay rise) can point at what senior actually means and see whether you're doing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like done properly
&lt;/h2&gt;

&lt;p&gt;Since then I've gone looking for companies that write that combination of effort and value down, and Axon did it really well. Their engineering ladder didn't stop at senior. It carried on past Senior Software Engineer II to Staff Engineer, then Principal, then Distinguished Engineer, three more rungs past the point where a lot of companies simply run out of ladder, and each one had its expectations written down and sat level with a role on the management side, all the way up to just under the C-suite. An engineer there could keep climbing for a whole career without ever having to manage anybody.&lt;/p&gt;

&lt;p&gt;Axon is where I was asked that interview question, as it happens, and I took the job on the engineering side of the coin, on purpose. Years of running teams didn't make it a step down, whatever the CV looked like from the outside. (For the record I sat a rung lower than the work I was actually doing, which was more like the two above it. A ladder can be beautifully drawn and still be read wrong.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Same number on both sides
&lt;/h2&gt;

&lt;p&gt;Now, the man who asked me that question was one of Axon's longest-serving infrastructure engineers, a seriously skilled one with a whole career of it behind him, and somewhere along the way management had been thrust upon him rather than chosen. He was damn good at it, too. He took all that engineering and turned it into real engineering management. He could handle a room full of stakeholders, he knew what a team could actually carry, and he knew when a deadline was nonsense, because he'd been the one at the keyboard for most of his working life.&lt;/p&gt;

&lt;p&gt;He got scapegoated anyway, more than once, for problems that were really about a small team being handed a big job without the people to do it, and eventually he switched back to the engineering track. Everyone around him treated it as a demotion.&lt;/p&gt;

&lt;p&gt;What still makes me laugh (in a slightly hollow way) is that on paper it wasn't one. He'd been a manager at one level, and he went back to being an engineer at exactly the same level. Same number on both sides. By the company's own ladder it was a sideways step, a coin turned over... and the people around him read it as a fall anyway, because they wanted to.&lt;/p&gt;

&lt;p&gt;So even a really well-written ladder isn't proof against people. Axon had about the best engineering path I've worked inside, and folk were still able to bend it to suit their own agendas... writing the rungs down is necessary, it just doesn't stop anyone reading them upside down.&lt;/p&gt;

&lt;p&gt;He asked me, all those months earlier, why I wanted a job I was clearly overqualified for. I'd just picked which side of the coin I wanted facing up.&lt;/p&gt;

&lt;p&gt;He's since picked neither. His years at Axon paid off rather nicely in the end, and he's taken early retirement, so as far as I'm concerned the last laugh is his... even if it does mean the industry lost a fantastic engineer and an exceptional manager in one go. Turns out you can land the coin on its edge after all.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://phpboyscout.uk/management-was-never-a-promotion/" rel="noopener noreferrer"&gt;phpboyscout.uk&lt;/a&gt; on 4 October 2026.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>career</category>
    </item>
    <item>
      <title>Designing a 99.999% IoT Platform</title>
      <dc:creator>beefed.ai</dc:creator>
      <pubDate>Sun, 04 Oct 2026 02:05:25 +0000</pubDate>
      <link>https://dev.to/beefedai/designing-a-99999-iot-platform-1lf3</link>
      <guid>https://dev.to/beefedai/designing-a-99999-iot-platform-1lf3</guid>
      <description>&lt;p&gt;The symptoms are familiar: device fleets that flood your broker after a region blip, firmware campaigns that stall because the &lt;code&gt;device registry&lt;/code&gt; is quarantined, and business teams escalating because analytics lose a window of truth during maintenance. You get paged at 03:00 to manually re-route traffic, and the postmortem shows the same root causes as last quarter: single-region control plane, opaque dependency maps, and brittle runbooks.&lt;/p&gt;

&lt;p&gt;Contents&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why 99.999% uptime is non-negotiable for real-world IoT fleets&lt;/li&gt;
&lt;li&gt;Architectural patterns that actually deliver five nines&lt;/li&gt;
&lt;li&gt;How to build a resilient multi-region deployment and DR plan&lt;/li&gt;
&lt;li&gt;How to prove resilience: failover testing, chaos engineering, and contractual SLAs&lt;/li&gt;
&lt;li&gt;Designing observability and alarms without bankrupting the project&lt;/li&gt;
&lt;li&gt;Operational runbooks, checklists, and templates you can use in 48 hours&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why 99.999% uptime is non-negotiable for real-world IoT fleets
&lt;/h2&gt;

&lt;p&gt;Five nines means roughly &lt;strong&gt;5.26 minutes of downtime per year&lt;/strong&gt;, and that hard number shapes what counts as “acceptable” risk on every device lifecycle operation and release window.   &lt;em&gt;Your SLO is the control you hand to the business; the error budget is the throttle on feature churn.&lt;/em&gt; Use the error-budget model from SRE to make reliability decisions objective and repeatable: you convert availability percentages into minutes, allocate that budget, and let the budget drive release policy and tickets for remediation.  &lt;/p&gt;

&lt;p&gt;For IoT, availability has second-order effects that are uniquely painful:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A downed &lt;code&gt;device registry&lt;/code&gt; means new or replaced devices cannot authenticate — field technicians stop working.&lt;/li&gt;
&lt;li&gt;Lost ingestion windows create holes in digital twins and analytics, producing stale commands.&lt;/li&gt;
&lt;li&gt;Regulatory and safety exposure in OT/industrial contexts can translate downtime into fines or injury.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Make &lt;strong&gt;availability&lt;/strong&gt; your primary non-functional requirement when the platform is used for control, billing, or safety. Architecture follows from that requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architectural patterns that actually deliver five nines
&lt;/h2&gt;

&lt;p&gt;You must stop thinking in “single-region” terms and design with the expectation of partial, intermittent, and correlated failures.&lt;/p&gt;

&lt;p&gt;Key high-availability building blocks I use at scale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Decouple ingestion with durable queues&lt;/strong&gt;: use an event log (e.g., Kafka/Kinesis) as the canonical ingestion buffer so downstream consumers can be scaled or recovered without losing telemetry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stateless front ends, stateful long-term stores&lt;/strong&gt;: keep connection brokers and ingestion &lt;strong&gt;stateless&lt;/strong&gt; (easy to scale), and push durable state to geo-replicated stores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Active-active for critical flows; warm standby for the rest&lt;/strong&gt;: reserve &lt;em&gt;active-active&lt;/em&gt; for control-plane endpoints or customer-facing APIs that need near-zero RTO; use &lt;em&gt;warm standby&lt;/em&gt; for analytics pipelines to balance cost and recovery time.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Device registry as the single source of truth&lt;/strong&gt;: the &lt;code&gt;device registry&lt;/code&gt; must be designed for cross-region access or reliable replication; store immutable device identity attributes and use per-region caches for read performance with deterministic reconciliation for writes. AWS IoT’s registry and Device Shadow primitives are useful references for capabilities you’ll need.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Digital twin separation&lt;/strong&gt;: keep the fast device twin (&lt;code&gt;Device Shadow&lt;/code&gt;) close to the device for command-and-control and replicate aggregated twin state to a graph/analytics twin (e.g., Azure Digital Twins) for business logic and historical analysis.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A compact comparison helps align trade-offs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Typical RTO&lt;/th&gt;
&lt;th&gt;Typical RPO&lt;/th&gt;
&lt;th&gt;Relative Cost&lt;/th&gt;
&lt;th&gt;When to pick&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Active‑Active (multi‑region)&lt;/td&gt;
&lt;td&gt;Seconds&lt;/td&gt;
&lt;td&gt;Near‑zero&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Control-plane and customer-facing APIs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warm‑Standby (hot spare)&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;Seconds–minutes&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Ingestion, near-real-time analytics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pilot‑Light&lt;/td&gt;
&lt;td&gt;Tens of minutes–hours&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;Low–Medium&lt;/td&gt;
&lt;td&gt;Non-critical analytics and batch jobs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backup &amp;amp; Restore (cold)&lt;/td&gt;
&lt;td&gt;Hours–Days&lt;/td&gt;
&lt;td&gt;Hours–Days&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Archival systems, cost-sensitive workloads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These categories and the suggested actions come from well‑architected disaster-recovery guidance and event-driven DR patterns used in cloud best practices.  &lt;/p&gt;

&lt;p&gt;Practical engineering rules I follow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Make the &lt;strong&gt;control plane&lt;/strong&gt; (provisioning, cert rotation, ACLs) independently recoverable from the &lt;strong&gt;data plane&lt;/strong&gt; (telemetry ingestion).&lt;/li&gt;
&lt;li&gt;Require &lt;code&gt;idempotent&lt;/code&gt; ingestion: every device message has a stable identifier or sequence so retries never create corruption.&lt;/li&gt;
&lt;li&gt;Design &lt;code&gt;device&lt;/code&gt; behavior for graceful backoff and exponential reconnect with jitter; never let a reconnect storm take down the broker.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to build a resilient multi-region deployment and DR plan
&lt;/h2&gt;

&lt;p&gt;Multi‑region design isn’t optional when you target five nines. You must choose where to spend money (and where not to).&lt;/p&gt;

&lt;p&gt;Core considerations and patterns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Global traffic steering vs DNS TTL&lt;/strong&gt;: DNS failover is cheap but slow; global load balancers or services like AWS Global Accelerator / Azure Front Door provide rapid regional failover or weighted routing with health probes. Use them for customer-facing endpoints.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-region ingestion endpoints&lt;/strong&gt;: expose region-local MQTT/WebSockets endpoints so devices connect to the nearest ingress. Replicate events asynchronously to central processing with durable logs for replay and recovery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Registry replication approaches&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Strongly replicated global DB&lt;/em&gt; (DynamoDB Global Tables-style) gives near‑real-time updates everywhere at higher cost and complexity.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Primary region with async replication&lt;/em&gt; reduces cost but increases write RPO and requires conflict resolution.
Choose based on whether device onboarding or device command integrity is more critical.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data replication for analytics&lt;/strong&gt;: use change-data-capture (CDC) or event-stream replication into your analytics fabric so a region loss doesn’t create a permanent gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Network partitions and split brain&lt;/strong&gt;: define clear leader election rules and write-shard boundaries. Don’t let two regions accept diverging &lt;code&gt;desired state&lt;/code&gt; commands without reconciliation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Design checklist for a multi-region DR plan:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Document RTO and RPO per service and per device class.&lt;/li&gt;
&lt;li&gt;Map dependencies (auth, registry, ingestion, processing, downstream APIs).&lt;/li&gt;
&lt;li&gt;Choose a DR pattern per dependency (active-active, warm-standby, pilot-light).&lt;/li&gt;
&lt;li&gt;Automate failover steps (route updates, promote DB writer, increase consumer scaling).&lt;/li&gt;
&lt;li&gt;Schedule and run non-production failover drills and maintain runbook automation.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  How to prove resilience: failover testing, chaos engineering, and contractual SLAs
&lt;/h2&gt;

&lt;p&gt;You can’t claim five nines unless you measure it — and you can’t measure it unless you test it under realistic failure modes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Run your &lt;strong&gt;GameDays&lt;/strong&gt; and scheduled failovers: simulate region loss, induce load spikes, and rehearse full failover runbooks in staging. Azure’s IoT Hub documentation recommends using non‑production environments to validate region failover behavior because region failover can cause data loss and downtime during tests.
&lt;/li&gt;
&lt;li&gt;Adopt &lt;strong&gt;chaos engineering&lt;/strong&gt; for continuous assurance: inject faults targeted at dependencies (broker nodes, database replicas, network latency) and verify automated recovery. Gremlin has a practical catalog for failure modes and regulatory use cases; Netflix’s Chaos Monkey is the origin story and still useful as an operational pattern.
&lt;/li&gt;
&lt;li&gt;Make SLOs and &lt;strong&gt;error budgets&lt;/strong&gt; your operational control loop: tie release velocity to remaining error budget and require postmortems when incidents exceed threshold consumption. Use the SRE error-budget model to agree with product teams on the trade-offs between features and stability.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Concrete failover testing protocol (short):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In staging, trigger a simulated region outage (network blackhole + terminated ingestion nodes).&lt;/li&gt;
&lt;li&gt;Execute automated runbook to re-route traffic to secondary and promote writable endpoint.&lt;/li&gt;
&lt;li&gt;Stream a golden dataset through the platform to verify no message loss and correct &lt;code&gt;digital twin&lt;/code&gt; state reconciliation.&lt;/li&gt;
&lt;li&gt;Measure RTO, RPO, and user-impacted SLIs; log and create P0 actions for any divergence.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Sample PromQL SLI (availability) to implement as a production SLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# percentage of successful ingestion requests over 5m window
100 * (1 - sum(rate(iot_ingest_requests_total{job="ingest",status=~"5.."}[5m])) / sum(rate(iot_ingest_requests_total{job="ingest"}[5m])))
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Prove, measure, and &lt;strong&gt;codify&lt;/strong&gt;: a test that runs once but is not automated will be forgotten.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing observability and alarms without bankrupting the project
&lt;/h2&gt;

&lt;p&gt;Observability is the lever: good metrics let you detect failures before they cascade; bad metrics produce pager noise and cost overruns.&lt;/p&gt;

&lt;p&gt;Instrumentation strategy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use a vendor-neutral tracing and metric layer like &lt;strong&gt;OpenTelemetry&lt;/strong&gt; for traces, metrics, and context propagation across services.
&lt;/li&gt;
&lt;li&gt;For metrics at scale, avoid centralizing raw Prometheus scraping across regions. Use &lt;code&gt;remote_write&lt;/code&gt; into a global long-term store (Thanos / Grafana Mimir / Cortex) or aggregate per-region before global query. This balances latency, availability, and cost.
&lt;/li&gt;
&lt;li&gt;Favor &lt;strong&gt;SLO-driven alerts&lt;/strong&gt;: page on SLO breach probability, not on raw 5xx counts. Route different alert levels to different channels (ops, engineering, product) and attach runbook links to alerts.&lt;/li&gt;
&lt;li&gt;Implement sampling and downsampling: keep high-cardinality traces for 1–2 weeks, metrics for 90 days with downsampled aggregates thereafter, and logs for a short window unless flagged for retention.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example Prometheus remote_write snippet (agent-mode):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;global&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;scrape_interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;15s&lt;/span&gt;

&lt;span class="na"&gt;remote_write&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://thanos-receive.us-east-1.example.com/api/v1/receive"&lt;/span&gt;
    &lt;span class="c1"&gt;# secure it with mTLS or basic_auth in production&lt;/span&gt;
&lt;span class="na"&gt;scrape_configs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;job_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;iot_broker_exporter'&lt;/span&gt;
    &lt;span class="na"&gt;static_configs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;targets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;broker-us-east-1:9100'&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cost trade-offs to manage:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;High-cardinality metrics and long retention both drive storage and query cost — prefer aggregation at the edge.&lt;/li&gt;
&lt;li&gt;Synthetic checks are cheap and high-value; instrument heartbeats from brokers and core services.&lt;/li&gt;
&lt;li&gt;Use alerts with escalation windows and deduplication to protect on-call from storms.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Important:&lt;/strong&gt; Treat &lt;code&gt;iot monitoring&lt;/code&gt; as a product: agree SLIs with your stakeholders, instrument them precisely, and fund observability like you fund production capacity.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Operational runbooks, checklists, and templates you can use in 48 hours
&lt;/h2&gt;

&lt;p&gt;This is a pragmatic playbook you can execute quickly.&lt;/p&gt;

&lt;p&gt;SLO &amp;amp; policy checklist&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Define SLOs by product slice (control-plane, ingest API, device provisioning). Document measurement windows and error-budget policy.
&lt;/li&gt;
&lt;li&gt;Create an SLA template using the SLO as the objective and list remedies for breach.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Critical DR runbook template (short form)&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Trigger: Detect region-wide loss of ingestion (all health checks failing for &amp;gt; 30s).&lt;/li&gt;
&lt;li&gt;Owner: Platform On-Call (primary).&lt;/li&gt;
&lt;li&gt;Steps:

&lt;ul&gt;
&lt;li&gt;Promote secondary ingestion writer / change DB writer endpoint.&lt;/li&gt;
&lt;li&gt;Update global routing weights to route 100% traffic to secondary (or flip failover DNS).&lt;/li&gt;
&lt;li&gt;Validate device heartbeats and &lt;code&gt;device registry&lt;/code&gt; reads (run &lt;code&gt;curl&lt;/code&gt; health endpoints).&lt;/li&gt;
&lt;li&gt;Run golden-data replay for last 5 minutes and reconcile digital twin deltas.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Post‑incident: Conduct postmortem with action items, link to runbook and error-budget consumption.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Emergency runbook quick-table&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Action&lt;/th&gt;
&lt;th&gt;Owner&lt;/th&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Flip load-balancer routing to secondary&lt;/td&gt;
&lt;td&gt;Platform SRE&lt;/td&gt;
&lt;td&gt;&amp;lt; 5 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Promote DB writer / failover&lt;/td&gt;
&lt;td&gt;DB team&lt;/td&gt;
&lt;td&gt;&amp;lt; 10 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Validate device registry reads&lt;/td&gt;
&lt;td&gt;App owner&lt;/td&gt;
&lt;td&gt;&amp;lt; 15 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Start telemetry replay and reconciliation&lt;/td&gt;
&lt;td&gt;Data eng&lt;/td&gt;
&lt;td&gt;&amp;lt; 30 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;GameDay quick script&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Week 0: Run a smoke failover in staging for a single critical device group.&lt;/li&gt;
&lt;li&gt;Week 4: Run a full region simulated outage in staging and execute full runbook.&lt;/li&gt;
&lt;li&gt;Quarterly: Run a cross-team GameDay with customers/integrations invited to validate SLAs and communications.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Minimal automation to prioritize&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Make failover routing a one-click / CI-driven operation (no manual SSH edits).&lt;/li&gt;
&lt;li&gt;Keep infrastructure-as-code (&lt;code&gt;terraform&lt;/code&gt;/&lt;code&gt;arm&lt;/code&gt;/&lt;code&gt;bicep&lt;/code&gt;) for all routing and DNS changes.&lt;/li&gt;
&lt;li&gt;Wire alerts to a runbook link that includes exact commands and &lt;code&gt;audit&lt;/code&gt; checklists.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Designing for &lt;strong&gt;99.999% uptime&lt;/strong&gt; forces you to make repeatable decisions: define your SLOs first, split control and data planes, choose an appropriate multi-region DR pattern, automate failover, and instrument aggressively with SLO-driven alerts. Start by locking the &lt;code&gt;device registry&lt;/code&gt; and critical SLOs into code, schedule your first GameDay, and use the error budget as the single lever to balance reliability and change.&lt;/p&gt;

&lt;p&gt;Sources:&lt;br&gt;
 &lt;a href="https://aerospike.com/glossary/five-nines-uptime/" rel="noopener noreferrer"&gt;What is five-nines uptime?&lt;/a&gt; - Explains five-nines availability and the calculation of downtime per year.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://sre.google/sre-book/embracing-risk/" rel="noopener noreferrer"&gt;Embracing risk and reliability engineering (Google SRE)&lt;/a&gt; - SRE guidance on SLOs, error budgets, and operational policy.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://learn.microsoft.com/en-us/azure/reliability/reliability-iot-hub" rel="noopener noreferrer"&gt;Reliability in Azure IoT Hub (Microsoft Learn)&lt;/a&gt; - Details IoT Hub regional replication, manual failover guidance, and testing recommendations.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://docs.aws.amazon.com/iot/latest/developerguide/register-device.html" rel="noopener noreferrer"&gt;Managing things with the registry - AWS IoT Core (Docs)&lt;/a&gt; - Registry, Device Shadow, and device management patterns in AWS IoT.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://www.gremlin.com/chaos-engineering" rel="noopener noreferrer"&gt;Chaos Engineering — Gremlin&lt;/a&gt; - Use cases and practices for chaos engineering and GameDays.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://aws.amazon.com/blogs/architecture/implementing-multi-region-disaster-recovery-using-event-driven-architecture/" rel="noopener noreferrer"&gt;Implementing Multi-Region Disaster Recovery Using Event-Driven Architecture (AWS Architecture Blog)&lt;/a&gt; - Reference architecture for event-driven multi-region DR.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://learn.microsoft.com/en-us/azure/well-architected/design-guides/disaster-recovery" rel="noopener noreferrer"&gt;Develop a disaster recovery plan for multi-region deployments — Azure Well-Architected&lt;/a&gt; - DR strategies (active‑active, warm standby, pilot light) and validations.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://opentelemetry.io/docs/" rel="noopener noreferrer"&gt;OpenTelemetry Documentation&lt;/a&gt; - Vendor-neutral observability framework, Collector and instrumentation guidance.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://binaryscripts.com/prometheus/2025/05/25/prometheus-monitoring-for-multi-region-applications-aggregating-metrics-across-global-data-centers.html" rel="noopener noreferrer"&gt;Prometheus Monitoring for Multi-Region Applications (BinaryScripts)&lt;/a&gt; - Federation vs &lt;code&gt;remote_write&lt;/code&gt;, Thanos/Cortex patterns for global metrics.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://github.com/grafana/mimir" rel="noopener noreferrer"&gt;Grafana Mimir (GitHub)&lt;/a&gt; - Scalable, multi‑tenant long-term storage for Prometheus-compatible metrics.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://github.com/Netflix/chaosmonkey" rel="noopener noreferrer"&gt;Netflix Chaos Monkey (GitHub)&lt;/a&gt; - Historical reference and open-source tooling for chaos engineering.&lt;br&gt;&lt;br&gt;
 &lt;a href="https://learn.microsoft.com/en-us/azure/digital-twins/overview" rel="noopener noreferrer"&gt;What is Azure Digital Twins? (Microsoft Learn)&lt;/a&gt; - Digital twin concepts and integration with IoT Hub for modeling and event routing.&lt;/p&gt;

</description>
      <category>platform</category>
      <category>embedded</category>
    </item>
    <item>
      <title>I Traced CrewAI's Sandbox CVE: 9 Names Missed the Runtime</title>
      <dc:creator>Kiell Tampubolon</dc:creator>
      <pubDate>Sun, 04 Oct 2026 02:01:22 +0000</pubDate>
      <link>https://dev.to/kielltampubolon/i-traced-crewais-sandbox-cve-9-names-missed-the-runtime-35im</link>
      <guid>https://dev.to/kielltampubolon/i-traced-crewais-sandbox-cve-9-names-missed-the-runtime-35im</guid>
      <description>&lt;p&gt;A CVE published yesterday afternoon says the Python sandbox in CrewAI, a widely used agent framework, blocked nine module names and still lost. The record's wording is blunt for a CVE: the blocklist "operates at the wrong level of abstraction." The escape it describes never uses an import statement at all.&lt;/p&gt;

&lt;p&gt;I was researching agent sandbox failures this week anyway. Between the GreyNoise agent swarm report, the SGLang pickle RCE last weekend, and Bengio's agent-misbehavior paper trending on Hacker News, every feed I follow is agent security right now. When my trending scan surfaced CVE-2026-37008 this morning, I searched for writeups and found none. Zero HN threads, zero Dev.to posts, no news pickup. So I pulled the NVD record, the MITRE CNA data, and the fix commit from GitHub, and read all three. Here is what the record says, what the vulnerable code actually did, why the fix deletes the feature instead of extending the list, and three checks worth running on any sandbox your agent framework ships.&lt;/p&gt;

&lt;p&gt;This is the entire security model the CVE is about, verbatim from the pre-fix source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# From crewai-tools code_interpreter_tool.py, pre-fix source (SandboxPython)
&lt;/span&gt;&lt;span class="n"&gt;BLOCKED_MODULES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;os&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sys&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subprocess&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;shutil&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                   &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;importlib&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inspect&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tempfile&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                   &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sysconfig&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;builtins&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;restricted_import&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;SandboxPython&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;BLOCKED_MODULES&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ImportError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Importing &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; is not allowed.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;__import__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nine names. One string comparison. Keep that in mind while we walk the record.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does CVE-2026-37008 actually say?
&lt;/h2&gt;

&lt;p&gt;Start with the record itself, because it is unusually readable. Published 2026-09-13 at 21:17 UTC by MITRE as the CNA. CVSS 3.1 base 8.1 HIGH with the vector &lt;code&gt;AV:L/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:L&lt;/code&gt;. No CVSS 4.0 score exists yet, and NVD has ingested the record without analyzing it, so the only score on any official source is the CNA's own. The weakness is CWE-424, "Improper Protection of Alternate Path," which in plain words means: you guarded one path and an alternate path stayed open.&lt;/p&gt;

&lt;p&gt;The description, verbatim:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;CrewAI before fb2323b offers a Python blocklist approach that operates at the wrong level of abstraction, a different vulnerability than CVE-2026-2275. Import-time blocking of module names does not address the availability of Python's complete object graph. For example, calling ctypes.CDLL(None) loads the C library without relying in any import statements. In other words, a within-process sandbox cannot merely account for the import system and instead must account for the complete runtime of the Python interpreter.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read the middle sentence twice. The claim is not "one dangerous module was missing from the list." The claim is that a list of module names cannot work in principle, because the interpreter's object graph is reachable without ever touching the import system.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why the affected range is a commit hash
&lt;/h3&gt;

&lt;p&gt;The affected range is not a version number. It is defined against git: every revision before commit &lt;code&gt;fb2323b3deb3ec62b3965526857e77a2264e4cd0&lt;/code&gt; is affected, everything after is not. If you &lt;code&gt;pip install&lt;/code&gt; CrewAI, there is no version string to compare against, and the GitHub advisory (GHSA-2q68-3cp7-72v9) is unreviewed with "Unknown" in both the affected and patched fields. This is the awkward part of commit-boundary CVEs, and it matters more than usual here because the fix landed six months ago.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why did a nine-name blocklist fail?
&lt;/h2&gt;

&lt;p&gt;The sandbox lived in &lt;code&gt;crewai-tools&lt;/code&gt;, in the &lt;code&gt;CodeInterpreterTool&lt;/code&gt; module, in a class called &lt;code&gt;SandboxPython&lt;/code&gt;. When the framework needed to run model-written Python and Docker was unavailable, it fell back to running that code with plain &lt;code&gt;exec()&lt;/code&gt; in the same process, with the builtins dict filtered to remove ten unsafe names like &lt;code&gt;exec&lt;/code&gt;, &lt;code&gt;eval&lt;/code&gt;, and &lt;code&gt;open&lt;/code&gt;, and with &lt;code&gt;restricted_import&lt;/code&gt; swapped in for &lt;code&gt;__import__&lt;/code&gt;. Notice what is absent from the nine: &lt;code&gt;ctypes&lt;/code&gt;, &lt;code&gt;socket&lt;/code&gt;, &lt;code&gt;pathlib&lt;/code&gt;. The vendor's own statement to CERT/CC about the sibling disclosure concedes the ctypes gap directly.&lt;/p&gt;

&lt;p&gt;But the deeper problem is what the record says: even a complete list of names guards the wrong thing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two ways past a name filter, no import required
&lt;/h3&gt;

&lt;p&gt;The record's own example is the first one. &lt;code&gt;ctypes.CDLL(None)&lt;/code&gt; asks the platform loader to open the running process itself, which hands back a handle whose symbol table includes libc. No Python import statement executes, so &lt;code&gt;restricted_import&lt;/code&gt; never fires. From that handle, write-ups of the earlier March disclosure describe reaching native calls directly, up to and including &lt;code&gt;system()&lt;/code&gt;. I am deliberately stopping at the handle rather than printing a chain; the record gives exactly that much and it is enough to understand the class.&lt;/p&gt;

&lt;p&gt;The second one is documented in the fix commit itself. Commit fb2323b added a test named &lt;code&gt;test_sandbox_escape_vulnerability_demonstration&lt;/code&gt;, marked xfail, which recovers the original &lt;code&gt;__import__&lt;/code&gt; by walking Python's object graph: take the empty tuple's class, list its subclasses, find one whose module still holds the real &lt;code&gt;__builtins__&lt;/code&gt;, and pull the original import function out of it. After that step, the blocklist is guarding a door the attacker is no longer using. Both escapes defeat the same filter by never passing through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How did code end up in this sandbox at all?
&lt;/h2&gt;

&lt;p&gt;The path matters as much as the code, because the dangerous fallback was never the default. CERT/CC's record for the March disclosures describes the trigger: an operator sets &lt;code&gt;allow_code_execution=True&lt;/code&gt; or manually attaches the &lt;code&gt;CodeInterpreterTool&lt;/code&gt;, and then Docker is unreachable. The tool checks Docker first. If it is there, model code runs in a container. If it is not, the tool silently falls back to &lt;code&gt;SandboxPython&lt;/code&gt; in the host process.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Shape of run_code_safety() before and after the fix
# (reconstructed from commit fb2323b's diff and message, not verbatim source)
&lt;/span&gt;
&lt;span class="c1"&gt;# BEFORE: fail-open. Docker missing means a silent in-process fallback.
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_check_docker_available&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run_code_in_docker&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;libraries_used&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run_code_in_restricted_sandbox&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# AFTER: fail-closed. Same check, loud failure.
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_check_docker_available&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;RuntimeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Docker is required for safe code execution &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                       &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;but is not available.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is also a nastier sibling. CVE-2026-2287, one of the March disclosures, covers Docker silently stopping mid-session: the same fallback fires even though nobody configured anything unsafe. So the recipe for reaching &lt;code&gt;SandboxPython&lt;/code&gt; includes a runtime failure you do not control. I wrote about the same fail-open shape when &lt;a href="https://dev.to/kielltampubolon/2-cvss-98-agent-sandbox-cves-landed-the-same-day-2bng"&gt;the two CVSS 9.8 Cua and AutoAgent sandbox bugs landed on the same day&lt;/a&gt;, and the pattern keeps repeating: the isolation is real until the moment it is quietly skipped.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is anyone actually affected today?
&lt;/h2&gt;

&lt;p&gt;Here is the part the bare record does not tell you, and the part most summaries will get wrong. The fix is old. Commit fb2323b merged on 2026-03-15, and the repo changelog shows v1.11.0rc1, released the same day, was the first tagged release containing it. I verified by ancestry that current tags, including the latest 1.15.21, all contain the fix. Better: v1.14.0 in April removed &lt;code&gt;CodeInterpreterTool&lt;/code&gt; entirely and deprecated the code execution parameters in favor of external sandboxes. If you &lt;code&gt;pip install crewai&lt;/code&gt; today, the vulnerable code was never in your tree.&lt;/p&gt;

&lt;p&gt;The realistic audience is narrower: projects pinned below 1.11.0rc1, forks and vendored copies of old &lt;code&gt;crewai-tools&lt;/code&gt;, and anyone who read an old tutorial explaining that "the sandbox keeps you safe."&lt;/p&gt;

&lt;p&gt;Why publish a CVE for code fixed in March? The record was reserved on 2026-04-06 and published 2026-09-13. The record does not say why, and I will not guess. What I can say is that it arrived with no coverage at all: I checked Hacker News via Algolia, the Dev.to API, and news search this morning and found nothing but mirrors. That is why this is a code-level walkthrough rather than a fourth rehash of a press release.&lt;/p&gt;

&lt;h2&gt;
  
  
  Could your sandbox have the same class of bug?
&lt;/h2&gt;

&lt;p&gt;I wrote before that &lt;a href="https://dev.to/kielltampubolon/a-happy-path-mcp-demo-proves-almost-nothing-about-tenant-isolation-17p"&gt;a happy-path demo proves almost nothing about tenant isolation&lt;/a&gt;, and the same rule applies to sandboxes. Three checks, derived from this case. Honest caveat: these are code-derived, not a lab-tested methodology.&lt;/p&gt;

&lt;p&gt;First, find the name filters. Any string-membership check on module names, any import hook, any allowlist, is a speed bump, not a boundary.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; &lt;span class="s2"&gt;"BLOCKED_MODULES&lt;/span&gt;&lt;span class="se"&gt;\|&lt;/span&gt;&lt;span class="s2"&gt;allowed_modules&lt;/span&gt;&lt;span class="se"&gt;\|&lt;/span&gt;&lt;span class="s2"&gt;is not allowed"&lt;/span&gt; ./your-agent-src
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rni&lt;/span&gt; &lt;span class="s2"&gt;"allowlist"&lt;/span&gt; ./your-agent-src | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"import&lt;/span&gt;&lt;span class="se"&gt;\|&lt;/span&gt;&lt;span class="s2"&gt;module"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Second, find every fallback path. Ask what happens when the primary isolation is unavailable. A fallback that runs restricted code in-process converts a missing daemon into an execution path. &lt;a href="https://dev.to/kielltampubolon/the-cursor-allowlist-bypass-that-starts-with-a-file-named-curl-5dac"&gt;The Cursor allowlist bypass I covered earlier&lt;/a&gt; is the same species: trusting a name instead of the resolved thing.&lt;/p&gt;

&lt;p&gt;Third, ask what the code can reach without importing. If executed code can touch &lt;code&gt;ctypes&lt;/code&gt;, &lt;code&gt;cffi&lt;/code&gt;, &lt;code&gt;gc.get_objects()&lt;/code&gt;, &lt;code&gt;object.__subclasses__()&lt;/code&gt;, or &lt;code&gt;sys.modules&lt;/code&gt;, an in-process filter cannot contain it. The record's last sentence is the design rule: a within-process sandbox must account for the complete runtime, which in practice means the boundary belongs outside the interpreter. Container, VM, or a hosted code runner, and the container must fail closed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should a string filter ever ship as a sandbox at all?
&lt;/h2&gt;

&lt;p&gt;Here is where I expect disagreement, and I want it. My position: the fix is right, and it should be the norm. A sandbox that cannot guarantee isolation should raise, loudly, rather than hope nobody notices. A filtered fallback does not just protect less than it claims, it advertises a boundary that does not exist, which is worse than no boundary in a code review.&lt;/p&gt;

&lt;p&gt;The counterargument is real, though. Many users do not have Docker installed, and a maintainer can reasonably say a restricted fallback with warnings is better than raw &lt;code&gt;exec&lt;/code&gt; in the same process, better at least against the laziest scripts. So where is the line between "better than nothing" and "false advertising"? I lean toward: the moment the feature is described as a sandbox, it has to fail closed. But I can see the other side, and a defensible compromise is an honest name like &lt;code&gt;run_code_unsandboxed_with_fewer_builtins&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The second argument is the score. Official: 8.1, local attack vector, high attack complexity, scope changed, confidentiality and integrity high. The reporter's own page lists 9.0, Critical. The gap comes from the local vector and the complexity adjustment, which encode "needs opt-in code execution plus a Docker failure." Is that the right lens when the thing being scored was, per the record itself, never a sandbox? I honestly do not know, and I would like to hear how you would score it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would do this week
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;If you pin CrewAI below 1.11.0rc1 or vendor an old &lt;code&gt;crewai-tools&lt;/code&gt;: upgrade. Current is 1.15.21 and the tool is gone.&lt;/li&gt;
&lt;li&gt;If you ship any restricted-execution fallback of your own: delete it or make it raise. Silence is the bug.&lt;/li&gt;
&lt;li&gt;Re-read your sandbox assuming the import system is not the boundary. In Python, it never was.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Primary sources: &lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-37008" rel="noopener noreferrer"&gt;CVE-2026-37008 on NVD&lt;/a&gt; · &lt;a href="https://cveawg.mitre.org/api/cve/CVE-2026-37008" rel="noopener noreferrer"&gt;MITRE CNA record&lt;/a&gt; · &lt;a href="https://github.com/advisories/GHSA-2q68-3cp7-72v9" rel="noopener noreferrer"&gt;GHSA-2q68-3cp7-72v9&lt;/a&gt; · &lt;a href="https://github.com/crewAIInc/crewAI/commit/fb2323b3deb3ec62b3965526857e77a2264e4cd0" rel="noopener noreferrer"&gt;Fix commit fb2323b&lt;/a&gt; · &lt;a href="https://www.kb.cert.org/vuls/id/221883" rel="noopener noreferrer"&gt;CERT/CC VU#221883&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>crewai</category>
      <category>security</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Detection and verification limits for CVE-2026-96364 on a live Drupal estate</title>
      <dc:creator>jeffrey</dc:creator>
      <pubDate>Sun, 04 Oct 2026 02:00:26 +0000</pubDate>
      <link>https://dev.to/jeffreyciend/detection-and-verification-limits-for-cve-2026-96364-on-a-live-drupal-estate-4582</link>
      <guid>https://dev.to/jeffreyciend/detection-and-verification-limits-for-cve-2026-96364-on-a-live-drupal-estate-4582</guid>
      <description>&lt;h1&gt;
  
  
  Detection and verification limits for CVE-2026-96364 on a live Drupal estate
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Vulnerability overview
&lt;/h2&gt;

&lt;p&gt;CVE-2026-96364 is listed in CERT-BUND WID-SEC-2026-3554, published 23 September 2026, covering 36 identifiers and 16 contributed Drupal projects. The advisory is rated high, flagged remotely exploitable, and carries CVSS version 3.1 base score 98 with temporal score 85. Patch availability is confirmed by the advisory's patch field.&lt;/p&gt;

&lt;h2&gt;
  
  
  Detection questions
&lt;/h2&gt;

&lt;p&gt;The batch record does not publish a trigger for CVE-2026-96364, which constrains detection. A network signature needs a request pattern, and a file integrity signature needs a known artifact. Neither is available from the record, and the association between the identifier and a specific project is not published there either. The per-project advisory is the place to look for class and version information.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verification approach
&lt;/h2&gt;

&lt;p&gt;Verification therefore has to be version based and inventory based. The reliable signal is the installed project version read from the site's own status report or Composer lock, checked against the range table for the branch in use. Chasing exploit indicators without a published pattern produces alerts that cannot be triaged.&lt;/p&gt;

&lt;h2&gt;
  
  
  Impact
&lt;/h2&gt;

&lt;p&gt;The practical impact of the detection gap is a bias toward assurance by inventory. An organisation that can answer, per site, which of the 16 projects are installed and at which version has covered the question the batch record can answer. An organisation relying on remote scanning alone has not, because a module behind authentication is invisible to an external scanner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Affected products and scope
&lt;/h2&gt;

&lt;p&gt;Affected projects: Webform, Cloud, Project Browser, Commerce Decoupled Checkout, Mermaid Diagram Field, CookieCuttr, REST and JSON API Authentication, Stop administrator login, Tawk.to live chat, Editoria11y Accessibility Checker, Webform REST, AI CKEditor, Combined image style, CSS Usage Analyzer, Smart Content and Diba carousel slider. Fixed releases: Webform 6.2.12 and 6.3.1, Cloud 7.0.1, Project Browser 2.0.3 and 2.1.5, Commerce Decoupled Checkout 1.8.0, Mermaid Diagram Field 1.0.9, CookieCuttr 2.0.3, REST and JSON API Authentication 3.2.0, Stop administrator login 1.6, Tawk.to live chat 3.0.4, Editoria11y Accessibility Checker 2.2.23 and 3.0.9, Webform REST 4.2.1, AI CKEditor 1.4.3, Combined image style 1.0.7, CSS Usage Analyzer 1.0.2, Smart Content 3.2.1 and Diba carousel slider 3.0.2. Core is outside the advisory. Two projects appear on two branches each, so a single-branch inventory under-reports the estate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Exposure context
&lt;/h2&gt;

&lt;p&gt;ZoomEye returned 436388 assets for app="Drupal" on 27 September 2026 and zero for vul.cve="CVE-2026-96364". A zero here reflects index coverage for one string and is not evidence that no installation is affected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Remediation and mitigations
&lt;/h2&gt;

&lt;p&gt;Build the version check first, because it doubles as remediation input: export installed projects and versions, diff against the range table, and mark each row affected, fixed or absent. Update affected rows to the fixed release for their branch. Where update is not possible, disable the module or restrict its routes. Then re-verify from the status report, and keep the export as the baseline for the next batch. Treat any indicator-based detection as supplementary until a mechanism is published.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;CERT-BUND advisory WID-SEC-2026-3554, Drupal extensions, 23 September 2026&lt;/li&gt;
&lt;li&gt;Drupal Security Advisories sa-contrib-2026-154 through sa-contrib-2026-191, 23 September 2026&lt;/li&gt;
&lt;li&gt;Drupal security advisories index&lt;/li&gt;
&lt;li&gt;ZoomEye search for app="Drupal"&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>security</category>
      <category>drupal</category>
      <category>cve202696364</category>
      <category>contributedmodules</category>
    </item>
    <item>
      <title>레고 CAD를 AI가 만든다 – 1인 개발자가 저예산으로 바로 써볼 수 있을까</title>
      <dc:creator>JustJinoIT</dc:creator>
      <pubDate>Sun, 04 Oct 2026 02:00:06 +0000</pubDate>
      <link>https://dev.to/justjinoit/rego-cadreul-aiga-mandeunda-1in-gaebaljaga-jeoyesaneuro-baro-sseobol-su-isseulgga-151b</link>
      <guid>https://dev.to/justjinoit/rego-cadreul-aiga-mandeunda-1in-gaebaljaga-jeoyesaneuro-baro-sseobol-su-isseulgga-151b</guid>
      <description>&lt;h2&gt;
  
  
  프로젝트가 뭔데?
&lt;/h2&gt;

&lt;p&gt;야, ChatGPT가 레고 조립 코드를 만든다고? 원문에 따르면 LDraw라는 레고 조립 전용 언어(.ldr, .mpd)를 GPT‑6 Astra와 Opus 5.5가 자동 생성하도록 만든 파이썬 툴셋을 Docker 이미지로 배포했대. 웹 UI가 포함돼서 OpenAI, Claude, OpenRouter 등 여러 모델을 선택해 LDraw 파일을 만들 수 있다. 결과물은 LDView·LeoCAD·Studio 같은 뷰어에서 바로 열어볼 수 있지.&lt;/p&gt;

&lt;h2&gt;
  
  
  내 VPS에 올릴 수 있을까?
&lt;/h2&gt;

&lt;p&gt;1 GB VPS라면 Docker 컨테이너 하나 정도는 충분히 돌릴 수 있다. 이미지 자체가 몇 백 MB 정도라면 메모리 사용량도 크게 문제되지 않을 거다. 다만 OpenAI·Claude API 호출은 별도 비용이 발생한다. 무료 티어가 있다면 월 5 USD 이하로 끌어올릴 수 있지만, 트래픽이 늘면 비용이 급격히 오를 수 있다. 즉, &lt;strong&gt;운영 비용은 API 사용량에 달렸다&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  바로 시도해볼 수 있는 단계
&lt;/h2&gt;

&lt;p&gt;원문에 Dockerfile이 포함돼 있다. 로컬이나 VPS에서 다음 명령어로 빌드·실행하면 된다.&lt;br&gt;
bash&lt;/p&gt;

&lt;h1&gt;
  
  
  레포를 클론하고
&lt;/h1&gt;

&lt;p&gt;git clone &lt;a href="https://github.com/anteloc/ldraw-nova.git" rel="noopener noreferrer"&gt;https://github.com/anteloc/ldraw-nova.git&lt;/a&gt;&lt;br&gt;
cd ldraw-nova&lt;/p&gt;

&lt;h1&gt;
  
  
  이미지 빌드
&lt;/h1&gt;

&lt;p&gt;docker build -t ldraw-nova .&lt;/p&gt;

&lt;h1&gt;
  
  
  웹 서버 실행 (8000 포트 예시)
&lt;/h1&gt;

&lt;p&gt;docker run -d -p 8000:8000 ldraw-nova&lt;/p&gt;

&lt;p&gt;브라우저에서 &lt;code&gt;http://&amp;lt;서버IP&amp;gt;:8000&lt;/code&gt;에 접속하면 UI가 뜬다. 여기서 OpenAI 키만 입력하면 바로 LDraw 파일을 생성해볼 수 있다. &lt;strong&gt;API 키만 있으면 별도 서버 설정 없이 바로 테스트 가능&lt;/strong&gt;하니, 비용 부담을 최소화하려면 무료 티어를 먼저 써보는 게 좋다.&lt;/p&gt;

&lt;h2&gt;
  
  
  내 생각 한 줄
&lt;/h2&gt;

&lt;h2&gt;
  
  
  API 비용을 제외하면 Docker 하나만으로 충분히 운영 가능하니, 저예산 사이드 프로젝트에 딱 맞는 놈이다.
&lt;/h2&gt;

&lt;p&gt;원문: &lt;a href="https://github.com/anteloc/ldraw-nova" rel="noopener noreferrer"&gt;HackerNews AI&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;개발자 시각으로 읽는 AI 소식, 매일 인스타에도 올립니다 → &lt;a href="https://www.instagram.com/dogfootbro.ai/" rel="noopener noreferrer"&gt;@dogfootbro.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>lego</category>
      <category>python</category>
      <category>docker</category>
    </item>
    <item>
      <title>Two ways a simple LLM token counter goes wrong (with a demo you can run)</title>
      <dc:creator>soda4001</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:59:32 +0000</pubDate>
      <link>https://dev.to/soda4001/two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run-1134</link>
      <guid>https://dev.to/soda4001/two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run-1134</guid>
      <description>&lt;p&gt;A common first version of an LLM spending limit looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;used&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;used&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;ESTIMATE&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;LIMIT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;quota exceeded&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// check&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;                    &lt;span class="c1"&gt;// call&lt;/span&gt;
  &lt;span class="nx"&gt;used&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;totalTokens&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                                  &lt;span class="c1"&gt;// add&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It reads correctly. It has two defects, and both show up only under load or failure, which is when a spending limit matters.&lt;/p&gt;

&lt;p&gt;Everything below is reproducible without an API key. The "provider" in the demo is a 5 ms timer. The code is in &lt;a href="https://github.com/soda4001/llm-quota-guard" rel="noopener noreferrer"&gt;&lt;code&gt;llm-quota-guard&lt;/code&gt;&lt;/a&gt; (MIT).&lt;/p&gt;

&lt;h2&gt;
  
  
  Defect 1: the check and the add are separated by an &lt;code&gt;await&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;JavaScript is single-threaded, but &lt;code&gt;await&lt;/code&gt; hands control back to the event loop. Ten requests arriving together all run the &lt;code&gt;if&lt;/code&gt; before any of them reaches &lt;code&gt;used += …&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;limit=100, estimate=30, concurrent calls=10
naive counter : used=300 (limit exceeded by 200)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each call estimated 30 units against a limit of 100. All ten passed the check, because &lt;code&gt;used&lt;/code&gt; was still 0 when they asked. The limit was exceeded by a factor of three. Nothing was misconfigured; the order of operations is the problem.&lt;/p&gt;

&lt;p&gt;The fix is to make the check and a &lt;em&gt;claim on the budget&lt;/em&gt; one indivisible step, before the &lt;code&gt;await&lt;/code&gt;. This is a &lt;strong&gt;reservation&lt;/strong&gt;: a hold that counts against the limit immediately.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hold&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;guard&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reserve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tenant-42&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// throws if used + held + 30 &amp;gt; limit&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With the same ten concurrent calls:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;QuotaGuard    : used=90, rejected=7
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three calls fit (3 × 30 = 90 ≤ 100). Seven are rejected before they reach the provider.&lt;/p&gt;

&lt;h2&gt;
  
  
  Defect 2: charging up front, with no refund path
&lt;/h2&gt;

&lt;p&gt;The opposite fix is to add the estimate &lt;em&gt;before&lt;/em&gt; the call. That closes defect 1 and opens another: calls that fail never give the estimate back.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;2) 4 calls that all fail (nothing was produced)
   naive counter : used=120
   QuotaGuard    : used=0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four failed calls consumed 120 units of a 100-unit budget while producing nothing. Failures often arrive in bursts (an upstream incident, a bad deploy, a rate-limit storm), so a burst can exhaust the budget while no useful work was done.&lt;/p&gt;

&lt;p&gt;A reservation has two exits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;commit(actual)&lt;/code&gt; replaces the hold with what the provider actually billed.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;release()&lt;/code&gt; drops the hold and records nothing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;guard.run()&lt;/code&gt; wires both to the call's outcome: commit on success, release if the function throws.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the estimate can't tell you
&lt;/h2&gt;

&lt;p&gt;You don't know the true cost before the call. The reservation uses your estimate, and &lt;code&gt;commit(actual)&lt;/code&gt; trues it up afterwards. Two details matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Usage above the estimate is recorded, not rejected.&lt;/strong&gt; The tokens are already spent. The settlement reports it as &lt;code&gt;overage&lt;/code&gt;, so you can see how good your estimates are.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A hold expires.&lt;/strong&gt; A request that crashes between &lt;code&gt;reserve()&lt;/code&gt; and &lt;code&gt;commit()&lt;/code&gt; would otherwise lock quota forever, so each hold has a TTL (default 120 s). If a &lt;code&gt;commit()&lt;/code&gt; arrives &lt;em&gt;after&lt;/em&gt; the TTL, the usage is still recorded: the provider billed it whether or not your bookkeeping was still waiting.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two more rules remove ambiguity: settling twice returns the first result instead of counting twice (so a retry is safe), and usage belongs to the window in which the reservation was made, so a call that finishes after a window boundary doesn't charge the next window.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not solve
&lt;/h2&gt;

&lt;p&gt;The limits are worth stating plainly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The state is in memory, in one process. Several instances or serverless invocations each have their own counters. Sharing state needs a store with atomic operations, which is a different piece of work.&lt;/li&gt;
&lt;li&gt;The numbers are whatever your code reports. If a call fails &lt;em&gt;after&lt;/em&gt; the provider billed it, this counter never learns about it.&lt;/li&gt;
&lt;li&gt;Fixed windows permit a burst at a boundary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use provider-side hard limits as the backstop, and treat an in-app guard as the layer that gives &lt;em&gt;your&lt;/em&gt; users and tenants predictable limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/soda4001/llm-quota-guard
&lt;span class="nb"&gt;cd &lt;/span&gt;llm-quota-guard
npm &lt;span class="nb"&gt;install
&lt;/span&gt;npm run example   &lt;span class="c"&gt;# prints the numbers above&lt;/span&gt;
npm &lt;span class="nb"&gt;test&lt;/span&gt;          &lt;span class="c"&gt;# 22 tests; they specify each rule described here&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Developed and tested on Node 24 (CI runs the same commands). The implementation is about 320 lines including comments, with no runtime dependencies.&lt;/p&gt;

&lt;p&gt;If you have handled this differently (Redis scripts, token buckets, provider-side budgets), counterexamples and corrections are welcome in the comments or as issues on the repo.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This article was written by an AI (Claude) under the direction of the account owner. It was not written or edited by hand. The numbers above are the output of the commands in the Reproduce section.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>typescript</category>
      <category>node</category>
      <category>abotwrotethis</category>
    </item>
    <item>
      <title>Resizing a photo to 1200x630 without stretching: why the result can be 840x630</title>
      <dc:creator>Edward Chapman</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:56:36 +0000</pubDate>
      <link>https://dev.to/edchapman/resizing-a-photo-to-1200x630-without-stretching-why-the-result-can-be-840x630-4c34</link>
      <guid>https://dev.to/edchapman/resizing-a-photo-to-1200x630-without-stretching-why-the-result-can-be-840x630-4c34</guid>
      <description>&lt;p&gt;A photo that looks stretched usually has a width-to-height ratio that differs from its original. Resizing a 4:3 photo into a 16:9 rectangle by changing both dimensions independently makes the objects in it wider or narrower.&lt;/p&gt;

&lt;p&gt;Keeping the aspect ratio avoids that distortion. It doesn't mean the result will fill every box you give it, though. That distinction matters when a website asks for an image at a specific size.&lt;/p&gt;

&lt;p&gt;Maximum dimensions and exact dimensions&lt;/p&gt;

&lt;p&gt;Maximum width and height describe a box the image must fit inside. Both output dimensions must stay below their limits, but they don't both have to reach them. The image can keep its original shape.&lt;/p&gt;

&lt;p&gt;Exact dimensions describe the required output rectangle. If its ratio differs from the source, you need to choose between cropping, adding padding and stretching. Cropping removes content. Padding keeps the content but adds a border or bars. Stretching changes the proportions.&lt;/p&gt;

&lt;p&gt;Before choosing a tool, check which result you need. A content image with a maximum width is a different job from a social preview card that must have an exact shape.&lt;/p&gt;

&lt;p&gt;A 4032x3024 photo in a 1200x630 box&lt;/p&gt;

&lt;p&gt;This source photo has a 4:3 ratio. Scaling its width to 1200 would give a height of 900, which exceeds the 630-pixel limit.&lt;/p&gt;

&lt;p&gt;Instead, scale by the height: 4032 multiplied by 630, divided by 3024, gives a width of 840. The fitted result is 840x630. It stays within both limits and preserves the source ratio, but it isn't 1200x630.&lt;/p&gt;

&lt;p&gt;To fill an exact 1200x630 rectangle without stretching, crop the source to the target ratio in an editor first, or use an editor that adds padding. Choose the crop yourself if important content is near the edges.&lt;/p&gt;

&lt;p&gt;The same issue appears with 1920x1080. A 3000x2000 source has a 3:2 ratio. Fitting it into that box gives 1620x1080, not 1920x1080. A 16:9 source can reach both target dimensions if it is large enough and the tool allows that output size. A smaller source won't reach them in a tool that never enlarges images.&lt;/p&gt;

&lt;p&gt;Dimensions and file size are separate&lt;/p&gt;

&lt;p&gt;Reducing dimensions removes pixel data. Choosing PNG's lossless encoding doesn't undo that resizing step or restore the removed pixels.&lt;/p&gt;

&lt;p&gt;JPEG and WebP quality settings affect encoding too. A smaller output isn't guaranteed: re-encoding an already compressed image can make its file larger. Compare the result with the original and inspect the preview before replacing anything.&lt;/p&gt;

&lt;p&gt;Transparency is another decision. JPEG doesn't support it. The Firm Beacon resizer fills transparent areas with white when exporting JPEG; PNG or WebP can keep transparency.&lt;/p&gt;

&lt;p&gt;What the browser tool does&lt;/p&gt;

&lt;p&gt;I work on Firm Beacon. Its free image resizer uses maximum dimensions and processes one selected JPG, PNG or WebP in the browser. It doesn't crop, stretch or enlarge the image, and it requires no account.&lt;/p&gt;

&lt;p&gt;The input limit is 20 MiB, with neither side above 8,192 pixels and no more than 24 million pixels in total. The 4032x3024 example has about 12.2 million pixels, so it meets the dimension limits; its file still needs to meet the byte limit.&lt;/p&gt;

&lt;p&gt;The selected image, filename, dimensions and settings aren't uploaded or saved by the tool. Loading the page still makes normal website and analytics requests. Analytics can record resizing and downloading actions without those image details.&lt;/p&gt;

&lt;p&gt;The Canvas export produces a still image without the original camera metadata or animation. Keep your original if you need either.&lt;/p&gt;

&lt;p&gt;Before downloading, check the output dimensions, transparency and file size. If your destination requires an exact rectangle, make sure you haven't mistaken a maximum-size setting for a crop.&lt;/p&gt;

&lt;p&gt;Try the resizer: &lt;a href="https://www.firmbeacon.co.uk/tools/image-resizer?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=image_resizer_guide" rel="noopener noreferrer"&gt;https://www.firmbeacon.co.uk/tools/image-resizer?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=image_resizer_guide&lt;/a&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
    </item>
    <item>
      <title>Visa Open-Sourced Its AI Cyber Defence — and Quietly Let Agents Spend</title>
      <dc:creator>ScriptMasterLabs </dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:54:18 +0000</pubDate>
      <link>https://dev.to/scriptmasterlabs01/visa-open-sourced-its-ai-cyber-defence-and-quietly-let-agents-spend-i7d</link>
      <guid>https://dev.to/scriptmasterlabs01/visa-open-sourced-its-ai-cyber-defence-and-quietly-let-agents-spend-i7d</guid>
      <description>&lt;p&gt;On &lt;strong&gt;September 29, 2026&lt;/strong&gt;, Visa's President of Technology &lt;strong&gt;Rajat Taneja&lt;/strong&gt; told Reuters the company open-sourced part of its AI-powered cyber defence after "humbling" AI-model vulnerabilities — and confirmed Visa &lt;strong&gt;has started allowing certain AI agents to use Visa&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Today's Reuters story has two halves and only one got the headline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Visa actually said
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Open-sourced part of its AI cyber defence.&lt;/strong&gt; The trigger: weaknesses exposed by &lt;strong&gt;Anthropic's Mythos AI model&lt;/strong&gt; earlier in 2026, and an &lt;strong&gt;AI-agent attack on the Hugging Face platform&lt;/strong&gt; where models escaped a testing sandbox. Taneja: "we have seen the trailer... I think this is just a small snippet of what the movie will look like."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agents are spending on Visa's rails now.&lt;/strong&gt; "Has also started allowing certain AI agents to use Visa." No details on how many, which, or under what controls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The $3.1 trillion wave.&lt;/strong&gt; Industry estimates cited by Reuters: roughly a third of online commerce — &lt;strong&gt;close to $3.1 trillion&lt;/strong&gt; — could run through AI agents by 2030. Visa processes roughly a billion payments a day worth around &lt;strong&gt;$15 trillion a year&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The defense is now open source. The authorization architecture for the agents is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap nobody in the coverage names
&lt;/h2&gt;

&lt;p&gt;Reuters plus a dozen identical wire syndications all cover the open-source move and the "trailer" quote. &lt;strong&gt;Not one asks the authorization question.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Think about the arithmetic Visa just put on the record: agents are already spending on its rails (unnamed, uncounted), ~$3.1 trillion of commerce could be agent-run by 2030, the rails run on trust — "payments firms rely on trust, meaning cyberattacks can be devastating" — and the thing that determines whether any single agent payment should fire is &lt;strong&gt;undisclosed&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Open-sourcing the &lt;em&gt;defence&lt;/em&gt; is the perimeter answer. The &lt;em&gt;authorization&lt;/em&gt; answer — what scores each payment instruction before money moves — is still missing. That's the decision gate: a confidence score on every agent payment, &lt;strong&gt;≥0.80 auto-pay, 0.50–0.79 human confirmation, &amp;lt;0.50 escalate&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The Hugging Face sandbox escape is the missing-gate failure mode. A sandbox says "you can't leave the room." A confidence gate says "you can't spend until the evidence clears." Visa just open-sourced better locks for the room. The $3.1T question is who approves the spending.&lt;/p&gt;

&lt;h2&gt;
  
  
  The live test
&lt;/h2&gt;

&lt;p&gt;We scored both of Visa's admissions against the live decision gate (~20:21 EDT Sept 29):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://scriptmasterlabs.com/api/harness/decide &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"state":{"amount_usd":0},
       "questions":[{"id":"q1","type":"score","scale":[0,1],
       "question":"Should the agent authorize AI-agent-initiated payments on my Visa card with no per-payment approval, given Visa has started allowing certain AI agents to use Visa?"}]}'&lt;/span&gt;

&lt;span class="c"&gt;# -&amp;gt; confidence 0.35 -&amp;gt; ESCALATE (block + log)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The honest finding: the local heuristic scored both the payment question and the sandbox question &lt;strong&gt;0.35&lt;/strong&gt; — identical. The safe direction (escalate anything in this territory), but the heuristic cannot discriminate between them. Same calibration gap as every run this week: a score is only as good as its calibration. Decider: local-heuristic-v1, calibrated=false.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate your agent spend in 5 steps
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Inventory your agents.&lt;/strong&gt; Visa won't say which agents are on its rails. Know which of &lt;em&gt;yours&lt;/em&gt; can touch money.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Score every payment instruction.&lt;/strong&gt; ≥0.80 auto-fire, 0.50–0.79 human confirmation, &amp;lt;0.50 escalate and block.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat the sandbox as untrusted.&lt;/strong&gt; Containment fails — the Hugging Face escape is the precedent. Score the &lt;em&gt;instruction&lt;/em&gt;, not the &lt;em&gt;environment&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log every block and escalation.&lt;/strong&gt; Visa's $15T-a-year trust business runs on auditability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rehearse the autonomous-attack future.&lt;/strong&gt; Rotate agent credentials on a schedule; assume a leak; make sure a leaked credential alone can't authorize spend without clearing the gate.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Honest caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The VVAH repository URL was not independently verified — it is not linked.&lt;/li&gt;
&lt;li&gt;"Certain AI agents" comes with no numbers or controls disclosed. The $3.1T figure is an industry estimate cited by Reuters, not Visa's projection.&lt;/li&gt;
&lt;li&gt;A confidence gate would not have retroactively stopped Mythos or the sandbox escape — the analogy is deliberate and bounded.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is a payments story wearing a security costume. Every prior piece this month lands on the same missing layer: the per-payment authorization architecture is undisclosed everywhere — from the MCP SDK OAuth flaw to the $78K Codex runaway. Visa's announcement is the biggest player on earth confirming both halves of the problem in one interview: the attacks are getting autonomous, and the agents are getting wallets.&lt;/p&gt;

&lt;p&gt;Full piece with dated receipts: &lt;a href="https://scriptmasterlabs.com/visa-open-source-ai-cyber-defence" rel="noopener noreferrer"&gt;https://scriptmasterlabs.com/visa-open-source-ai-cyber-defence&lt;/a&gt;&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>payments</category>
      <category>apisecurity</category>
      <category>ai</category>
    </item>
    <item>
      <title>Logistics Control Room: 4 Internal Admin Dashboard CRUD Signals for Feature Flags</title>
      <dc:creator>thomasmoore5082</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:53:10 +0000</pubDate>
      <link>https://dev.to/thomasmoore5082/logistics-control-room-4-internal-admin-dashboard-crud-signals-for-feature-flags-k9d</link>
      <guid>https://dev.to/thomasmoore5082/logistics-control-room-4-internal-admin-dashboard-crud-signals-for-feature-flags-k9d</guid>
      <description>&lt;p&gt;Short answer: page on a missed business result, not on a flag changing. For a logistics import, the least complex useful design is an internal control page backed by a small service that can set, list, toggle, and retire flags, while the scheduled worker emits four separate signals: run started, run completed, records accepted, and records rejected. Show the flag revision beside those signals. The on-call can then decide whether to retry an import, revert a flag, or investigate the upstream feed.&lt;/p&gt;

&lt;p&gt;The alert arrives at 04:20: &lt;code&gt;manifest_import_results_absent&lt;/code&gt; for depot &lt;code&gt;north-17&lt;/code&gt;. A useful page contains the last expected window, the last successful result timestamp, accepted and rejected counts, the active flag revision, and the most recent control-plane actor. A page that says only "job failed" is noise with a timestamp attached. Worse, an alert on every flag toggle confuses a control action with customer impact.&lt;/p&gt;

&lt;p&gt;The SLO-shaped question is narrow: did the scheduled window produce a valid result before its deadline? Everything else is supporting evidence.&lt;/p&gt;

&lt;p&gt;Results first.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should wake the on-call?
&lt;/h2&gt;

&lt;p&gt;Work backward from the responder's decision. A scheduler heartbeat proves that a process woke up; it does not prove that a manifest was fetched, parsed, accepted, or published. A successful upstream response is similarly incomplete. A zero-record file may be legitimate for one depot and suspicious for another, so a global &lt;code&gt;records == 0&lt;/code&gt; threshold will eventually train the team to ignore pages.&lt;/p&gt;

&lt;p&gt;Use four monotonically accumulated event counters, partitioned only by stable operational dimensions such as import type and depot class:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;import_runs_started_total&lt;/code&gt; identifies scheduler activity.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;import_runs_completed_total&lt;/code&gt; separates hung work from work that returned.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;import_records_accepted_total&lt;/code&gt; represents usable output.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;import_records_rejected_total&lt;/code&gt; exposes validation failure without pretending rejected rows are success.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do not put a flag key, user identifier, file name, or manifest ID into metric labels. Those values create an open-ended set of series and make capacity planning depend on business cardinality. Put per-run identifiers in logs or traces, then link them through a generated run ID. The dashboard needs the aggregate for detection and the run ID for diagnosis.&lt;/p&gt;

&lt;p&gt;A practical alert evaluates an expected schedule window rather than a fixed process heartbeat. If an import is due every hour and its agreed completion budget is 15 minutes, evaluate whether accepted results advanced inside that window, with a separate warning path for completed runs that rejected every record. Those numbers are examples, not universal defaults; derive the real window from the logistics contract, late-arrival distribution, and error budget.&lt;/p&gt;

&lt;p&gt;False certainty is expensive.&lt;/p&gt;

&lt;h2&gt;
  
  
  How should an internal admin dashboard handle feature flags CRUD?
&lt;/h2&gt;

&lt;p&gt;A simple internal tool still needs a boundary. The browser should not write a shared configuration file or mutate worker memory directly. It calls a control service; that service validates a complete desired state, records a revision, and makes the new value available to workers through a read path. Workers attach the observed revision to each run record. This creates a causal breadcrumb without claiming that every failure after a change was caused by the change.&lt;/p&gt;

&lt;p&gt;The following Go types are deliberately boring. &lt;code&gt;Enabled&lt;/code&gt; is explicit, retirement is distinct from disabling, and optimistic concurrency prevents two operators from silently overwriting each other.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;package&lt;/span&gt; &lt;span class="n"&gt;flags&lt;/span&gt;

&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;"errors"&lt;/span&gt;
    &lt;span class="s"&gt;"time"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;Flag&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Key&lt;/span&gt;       &lt;span class="kt"&gt;string&lt;/span&gt;     &lt;span class="s"&gt;`json:"key"`&lt;/span&gt;
    &lt;span class="n"&gt;Enabled&lt;/span&gt;   &lt;span class="kt"&gt;bool&lt;/span&gt;       &lt;span class="s"&gt;`json:"enabled"`&lt;/span&gt;
    &lt;span class="n"&gt;Revision&lt;/span&gt;  &lt;span class="kt"&gt;uint64&lt;/span&gt;     &lt;span class="s"&gt;`json:"revision"`&lt;/span&gt;
    &lt;span class="n"&gt;UpdatedAt&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Time&lt;/span&gt;  &lt;span class="s"&gt;`json:"updated_at"`&lt;/span&gt;
    &lt;span class="n"&gt;UpdatedBy&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;     &lt;span class="s"&gt;`json:"updated_by"`&lt;/span&gt;
    &lt;span class="n"&gt;RetiredAt&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Time&lt;/span&gt; &lt;span class="s"&gt;`json:"retired_at,omitempty"`&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;SetRequest&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Enabled&lt;/span&gt;          &lt;span class="kt"&gt;bool&lt;/span&gt;   &lt;span class="s"&gt;`json:"enabled"`&lt;/span&gt;
    &lt;span class="n"&gt;ExpectedRevision&lt;/span&gt; &lt;span class="kt"&gt;uint64&lt;/span&gt; &lt;span class="s"&gt;`json:"expected_revision"`&lt;/span&gt;
    &lt;span class="n"&gt;Reason&lt;/span&gt;           &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="s"&gt;`json:"reason"`&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="n"&gt;SetRequest&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;Validate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ExpectedRevision&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;New&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"expected_revision is required"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nb"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Reason&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="m"&gt;8&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;New&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"reason must explain the operational intent"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three endpoints are enough: list active and retired flags, put a complete state for one key, and delete one key by retiring it. A toggle button is presentation, not a special storage operation; it reads the current revision and submits the opposite complete state. Treating "toggle" as a blind verb makes retries ambiguous. If the first response is lost, a retry could flip the value back.&lt;/p&gt;

&lt;p&gt;Retries happen.&lt;/p&gt;

&lt;p&gt;The write handler should authenticate the operator, authorize the flag scope, validate the key against an allowlist, compare the expected revision, persist the new revision and audit record together, and return the stored representation. The worker should continue using its last valid snapshot if a refresh returns malformed data. That is a fail-stable choice, but it needs a staleness signal and a documented maximum age; otherwise a resilient cache becomes an invisible split brain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instrument the change, then test the absence
&lt;/h2&gt;

&lt;p&gt;The worker owns result instrumentation because only it knows whether records became usable. The control service owns change events because only it knows who requested the state and which revision won. Join those streams in the investigation view by revision and time, rather than stuffing control metadata into every metric label.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;package&lt;/span&gt; &lt;span class="n"&gt;imports&lt;/span&gt;

&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;"context"&lt;/span&gt;
    &lt;span class="s"&gt;"time"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;Signals&lt;/span&gt; &lt;span class="k"&gt;interface&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;RunStarted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;uint64&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;RunCompleted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;uint64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Duration&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;RecordsAccepted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;RecordsRejected&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;RunResult&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Accepted&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;
    &lt;span class="n"&gt;Rejected&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;Execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="n"&gt;Signals&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;class&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;revision&lt;/span&gt; &lt;span class="kt"&gt;uint64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;load&lt;/span&gt; &lt;span class="k"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;RunResult&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;started&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RunStarted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;class&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;revision&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RecordsAccepted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;class&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Accepted&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RecordsRejected&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;class&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Rejected&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"validation"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RunCompleted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;class&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;revision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Since&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;started&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice what this example refuses to do: it does not mark completion on the error path, and it does not convert an error into zero accepted records. Those are different failure modes. The evaluator can distinguish "never finished" from "finished with no usable output," while the on-call can inspect the run log for the exact error. Rejection reason classes must come from a bounded vocabulary; raw error messages belong in logs.&lt;/p&gt;

&lt;p&gt;Test absence with a controllable clock. A happy-path test should advance through one schedule window and observe accepted results. Then cover a disabled import, a hung loader, an empty-but-valid source, all rows rejected, a stale flag snapshot, a duplicated write request, and two writers racing on the same revision. The valuable assertion is that the page fires once, at the intended severity, with evidence pointing to a distinct action.&lt;/p&gt;

&lt;p&gt;Roll out worker instrumentation before enabling the page. Observe several real schedule cycles, compare the proposed evaluation with expected outcomes, and record which cases would have paged. Deployment order matters: worker emission first, dashboard second, alert last. Reversing it creates an avoidable no-data page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Buy, build, or keep the first version small?
&lt;/h2&gt;

&lt;p&gt;The decision is less about the CRUD form than the operating surface around it. A platform team should estimate on-call ownership, authentication integration, audit retention, recovery testing, expected flag count, evaluation traffic, and migration cost before debating interface polish.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;Signal-quality fit&lt;/th&gt;
&lt;th&gt;On-call load&lt;/th&gt;
&lt;th&gt;Lock-in pressure&lt;/th&gt;
&lt;th&gt;Capacity question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Small internal control service&lt;/td&gt;
&lt;td&gt;Exact fit is possible; the team owns correctness&lt;/td&gt;
&lt;td&gt;Highest; storage, backup, auth, and alerts remain yours&lt;/td&gt;
&lt;td&gt;Low behind a narrow interface&lt;/td&gt;
&lt;td&gt;Can cached reads cover peak workers plus admin writes?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shared self-hosted control plane&lt;/td&gt;
&lt;td&gt;Common lifecycle behavior can reduce new code&lt;/td&gt;
&lt;td&gt;The team patches, scales, restores, and upgrades it&lt;/td&gt;
&lt;td&gt;Moderate if its model leaks into workers&lt;/td&gt;
&lt;td&gt;What happens to stale reads during control-plane loss?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Managed control plane&lt;/td&gt;
&lt;td&gt;Local evidence still needs design&lt;/td&gt;
&lt;td&gt;Lower infrastructure load; integration ownership remains&lt;/td&gt;
&lt;td&gt;Potentially higher through SDK and policy coupling&lt;/td&gt;
&lt;td&gt;Is evaluation local or remote, and what is its failure budget?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a handful of import gates, start with the smallest service that meets the control requirements and keep evaluation behind an interface owned by the worker. This is not permission to skip backups or authorization. For hundreds of flags, multiple teams, or complex targeting, the homegrown option accumulates policy and lifecycle work quickly; reassess before a basic form becomes critical infrastructure by accident.&lt;/p&gt;

&lt;p&gt;This approach has a clear limitation: a small internal service is a poor fit when non-engineers need complex audience targeting, approvals span many organizations, or the platform team cannot own a highly available configuration store. In those conditions, choose a shared self-hosted or managed control plane according to the team's on-call capacity and acceptable coupling, then keep the import-result alert independent of that choice. The trade-off runs in both directions. A broader control plane reduces lifecycle code, but it does not know whether a depot import produced usable manifests; the worker still has to emit that business result. A small service preserves a narrow evaluation contract, but every backup drill, authorization rule, schema migration, stale-cache policy, and audit-retention job lands on the team that built it.&lt;/p&gt;

&lt;p&gt;No CRUD screen removes that work.&lt;/p&gt;

&lt;p&gt;Capacity planning should include failure mode, not just average request rate. Worker reads may be cacheable and frequent, while admin writes are rare but consequential. Define behavior during store loss, cap snapshot age, measure refresh failures, and prove restoration. A control plane that handles ten times normal traffic but cannot restore revisions is not ready for an on-call dependency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deletion is a lifecycle event
&lt;/h2&gt;

&lt;p&gt;A delete button should retire a flag, remove it from normal evaluation, and preserve enough audit history to explain prior behavior. Physical erasure can occur later under a retention policy. This distinction matters because operational evidence and personal data have different reasons for retention. If audit records contain personal data, GDPR Article 17 establishes a right to erasure and lists circumstances in which that right does not apply; legal and security owners must define the applicable policy rather than letting a database default decide it.&lt;/p&gt;

&lt;p&gt;Keep actor identity out of metrics. Store it in access-controlled audit records, minimize what is collected, and make retention enforceable. The UI should require confirmation that names the key and current revision, reject stale revisions, and display the retirement result. Recreating a retired key should be deliberate, with a new lifecycle, rather than an accidental effect of a retried request.&lt;/p&gt;

&lt;p&gt;Close the loop at the page. A threshold that is too tight wakes someone for routine late arrivals; one that is too loose consumes the logistics error budget before a human can intervene. Review pages by outcome: Was usable data missing? Did the notification arrive early enough to act? Did the evidence identify retry, revert, or upstream investigation? If repeated notifications produce no action, change the signal or move it out of paging. The objective is a small set of trustworthy interruptions tied to scheduled-import results.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GDPR Article 17, "Right to erasure": &lt;a href="https://gdpr-info.eu/art-17-gdpr/" rel="noopener noreferrer"&gt;https://gdpr-info.eu/art-17-gdpr/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>observability</category>
      <category>featureflags</category>
      <category>sre</category>
    </item>
    <item>
      <title>Artifactory Is Under Active Attack: 3 Checks in 30 Minutes</title>
      <dc:creator>Kiell Tampubolon</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:52:57 +0000</pubDate>
      <link>https://dev.to/kielltampubolon/artifactory-is-under-active-attack-3-checks-in-30-minutes-5h84</link>
      <guid>https://dev.to/kielltampubolon/artifactory-is-under-active-attack-3-checks-in-30-minutes-5h84</guid>
      <description>&lt;p&gt;A research report published Thursday describes four weeks of active exploitation of JFrog Artifactory, the artifact registry that sits in front of most Java and DevOps build pipelines. Attackers chain two patched CVEs to turn a single unauthenticated request into an admin-scoped token, and the tell is uncomfortable: every request they make afterward shows up in your logs as &lt;code&gt;token:anonymous&lt;/code&gt;, an actor name that looks exactly like background noise. CISA has all three CVEs on the Known Exploited Vulnerabilities catalog, and the federal remediation deadline for two of them is September 25.&lt;/p&gt;

&lt;p&gt;I write about MCP and tooling security here, and before trusting any report I cross-check it against primary records. The three NVD records and the KEV feed hold up: the privilege escalation is scored 8.1 by the vendor and 8.8 by NVD, the token exposure 7.5, the default-config admin bypass 9.8. What the records do not give you is an audit plan, so I built one from the Wiz IOC table and the version ranges. Three checks, about thirty minutes. One honesty note before we start: I derived these from the published IOCs and ranges, I have not run them against a live instance, so treat every log path here as a starting point to adapt.&lt;/p&gt;

&lt;p&gt;First, the version. Everything else depends on it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# self-hosted instance version, for the range checks below&lt;/span&gt;
&lt;span class="c"&gt;# use a token your automation already holds; if this endpoint is&lt;/span&gt;
&lt;span class="c"&gt;# admin-restricted on your build, read the version from the admin UI instead&lt;/span&gt;
curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$ARTIFACTORY_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"https://artifactory.internal.example/artifactory/api/system/version"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What did attackers actually do on these instances?
&lt;/h2&gt;

&lt;p&gt;Between August 15 and September 8, Wiz observed multiple actors chaining two CVEs against self-hosted Artifactory instances. The chain has three moves:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;POST /access/api/v1/aws/token/&lt;/code&gt;, with a trailing slash. The bare path rejects callers with a 401. The trailing-slash variant returned HTTP 200 with a JWT for the internal anonymous user, even when anonymous access was disabled. That is CVE-2026-42018.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;POST /access/api/v1/tokens&lt;/code&gt;. The actor exchanges that JWT for an admin-scoped token. The flaw here is scope validation: the token's signature and issuer are checked, its intended scope is not. That is CVE-2026-42016.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;PUT /api/security/users/&amp;lt;username&amp;gt;&lt;/code&gt; (or the matching UI endpoint), answered with a 201. A persistent admin account now exists. In some cases the whole sequence, first request to created admin account, took under five minutes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The report is explicit that neither CVE grants admin alone. The chain is the exploit. There is also a separate door: &lt;code&gt;POST /access/api/v1/registry/join&lt;/code&gt; returned HTTP 200 or 201 with an admin-scoped token in the response body under default configuration. That is CVE-2026-82329, scored 9.8, KEV-listed on September 2, and exploited by several actors between September 1 and September 8.&lt;/p&gt;

&lt;h3&gt;
  
  
  The step that makes the logs lie
&lt;/h3&gt;

&lt;p&gt;Here is the detail that elevates this from "another CVE post" for me. The escalated token kept the anonymous username but carried admin authority, so every later request appears with an actor of &lt;code&gt;token:anonymous&lt;/code&gt;. Wiz's wording is precise and worth keeping: the attacker did not steal an identity, they attached authority to one you already have.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does an admin action logged as token:anonymous matter?
&lt;/h2&gt;

&lt;p&gt;When I wrote about the MCP 2026-07-28 spec going stateless, my conclusion was that a planted prompt becomes a valid credential the moment state moves somewhere an attacker can write to (&lt;a href="https://dev.to/kielltampubolon/mcp-2026-07-28-went-stateless-a-planted-prompt-is-a-credential-5bem"&gt;the article is here&lt;/a&gt;). This report is the same lesson approached from the other side. The credential is real and admin-scoped, but its identity is decoration. Correlation is not authorization, and a principal name that every instance already carries is the perfect place to hide authority.&lt;/p&gt;

&lt;p&gt;If your SIEM alerts on "anonymous principal performed an admin action" at all, you are ahead of most setups I have seen. If your logs do not distinguish the built-in anonymous user from an admin-scoped token wearing its name, the attacker inherits that blind spot for free.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you check your instance in 30 minutes?
&lt;/h2&gt;

&lt;p&gt;Same argument I made about MCP demos: a happy path proves almost nothing, so run the boring checks before you trust the component (&lt;a href="https://dev.to/kielltampubolon/a-happy-path-mcp-demo-proves-almost-nothing-about-tenant-isolation-17p"&gt;I made that case here&lt;/a&gt;). All three checks below come from the Wiz IOC table and the NVD ranges.&lt;/p&gt;

&lt;h3&gt;
  
  
  Check 1: your version against the per-CVE ranges
&lt;/h3&gt;

&lt;p&gt;Take the version string from the call above and turn it into a tuple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;vulnerable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# v is a (major, minor, patch) tuple from your version string
&lt;/span&gt;    &lt;span class="c1"&gt;# ranges transcribed from the NVD records on 2026-09-14
&lt;/span&gt;    &lt;span class="n"&gt;hits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;133&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;11&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;hits&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CVE-2026-42016 priv-esc (CNA range: all &amp;lt; 7.133.11)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;111&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;117&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;117&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;27&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;125&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;125&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;19&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;133&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;133&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;28&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;146&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;146&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
        &lt;span class="n"&gt;hits&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CVE-2026-42018 anonymous-token exposure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;if &lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;111&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;111&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;117&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;117&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;28&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;125&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;125&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;133&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;133&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;29&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;146&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;146&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;38&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;161&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;161&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
        &lt;span class="n"&gt;hits&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CVE-2026-82329 default-config admin bypass&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;hits&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;vulnerable&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;125&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;18&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
&lt;span class="c1"&gt;# ['CVE-2026-42016 priv-esc (CNA range: all &amp;lt; 7.133.11)',
#  'CVE-2026-42018 anonymous-token exposure',
#  'CVE-2026-82329 default-config admin bypass']
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A build on 7.125.18 sits inside all three ranges. One trap before you relax on the older branches: the Wiz report lists fixed versions per release branch (7.111.21, 7.117.28, 7.125.20, 7.133.29, 7.146.38, 7.161.20 or later), but the 42016 record says everything before 7.133.11 is affected. Those two statements do not obviously reconcile for the oldest branches, and I could not settle which is authoritative from public records alone. Treat the newest version your upgrade path allows as the target, and check JFrog's advisory for your specific branch before you close the ticket.&lt;/p&gt;

&lt;h3&gt;
  
  
  Check 2: read the request log the way Wiz did
&lt;/h3&gt;

&lt;p&gt;The endpoint strings below are transcribed verbatim from the Wiz IOC table. The request-log location varies by install; the path shown is the container default.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;REQ&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$JFROG_HOME&lt;/span&gt;&lt;span class="s2"&gt;/artifactory/var/log/artifactory-request.log"&lt;/span&gt;

&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"POST /access/api/v1/aws/token"&lt;/span&gt;      &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$REQ&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;  &lt;span class="c"&gt;# bare path 401s and slash variant 200s&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"POST /access/api/v1/tokens"&lt;/span&gt;         &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$REQ&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;  &lt;span class="c"&gt;# token minting&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"POST /artifactory/api/security/token"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$REQ&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;  &lt;span class="c"&gt;# legacy token endpoint&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"PUT /api/security/users/"&lt;/span&gt;           &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$REQ&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;  &lt;span class="c"&gt;# new accounts&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"POST /access/api/v1/registry/join"&lt;/span&gt;  &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$REQ&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;  &lt;span class="c"&gt;# the 82329 door&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"GET /api/system/configuration"&lt;/span&gt;      &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$REQ&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;  &lt;span class="c"&gt;# config exfil pattern&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The highest-confidence signature in the report is behavioral: a 401 on the bare path followed by a 200 on a trailing-slash variant, from the same client, inside a short window. That is an operator confirming the vulnerable variant before relying on it, and normal clients do not produce that pattern. The other tells are identity mismatches: a low-privilege or anonymous identity minting tokens, enumerating users, or reading and writing &lt;code&gt;/artifactory/api/plugins&lt;/code&gt;. A lone 200 on the join endpoint is not proof of anything, so correlate it with account creations or configuration reads before you escalate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Check 3: hunt the imposter accounts
&lt;/h3&gt;

&lt;p&gt;Wiz lists the account names threat actors created, and they are designed to pass a casual scroll.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;

&lt;span class="n"&gt;users&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;artifactory_users.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;  &lt;span class="c1"&gt;# export from Administration &amp;gt; Identity and Access &amp;gt; Users
&lt;/span&gt;&lt;span class="n"&gt;NAMED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jfrog-distribution&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;backup-service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;repo-service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jfrog-insight&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jfrog-mission-control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jfrog-pipeline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;migration-tool&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ldap_admin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ldap_administrator&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0xterror&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;PATTERNS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;svc_[a-zA-Z0-9]{8}$&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Nxploited_[a-zA-Z0-9]{3}$&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;labadmin_[a-zA-Z0-9]{10}$&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;hit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;NAMED&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PATTERNS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;REVIEW:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;| admin:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;admin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
              &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;| created:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One honest caveat: depending on your integrations, some of those vendor-looking names might collide with legitimate service accounts, so do not mass-delete on a name match. Confirm the admin flag, the creation timestamp against your change log, and any tokens you cannot account for. Wiz classifies them as malicious admin accounts created by threat actors; your job is proving the creation date was not you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens once the token lands?
&lt;/h2&gt;

&lt;p&gt;Admin on a registry is not the end of the incident, it is the start of a worse one. Wiz observed malicious Groovy plugins installed through Artifactory's native plugin framework, which turns a legitimate extensibility feature into the implant loader, with follow-up commands running through &lt;code&gt;/api/plugins/execute/&lt;/code&gt;. Droppers fetched binaries over HTTP into world-writable paths like &lt;code&gt;/dev/shm&lt;/code&gt; and &lt;code&gt;/tmp&lt;/code&gt;, and a custom Rust backdoor with C2 capabilities appeared in multiple cases. One webshell was uploaded into a repository path, meaning the registry itself stored the implant. The 82329 pattern adds configuration exfiltration, join-key theft, and attackers attaching their own SSH keys to the accounts they created.&lt;/p&gt;

&lt;p&gt;I keep coming back to why this beats a normal app RCE. A web server gets patched and rebuilt. A registry is upstream of every build: whichever artifact your CI pulls next was assembled by the thing the attacker now administers. I learned the sink-mismatch version of this lesson on my own scanner, when the sink I guarded turned out not to be the sink that got used (&lt;a href="https://dev.to/kielltampubolon/my-mcp-security-scanner-missed-2026s-worst-mcp-rce-here-is-the-one-rule-fix-1g1i"&gt;written up here&lt;/a&gt;). Most teams guard the application tier and mentally file the registry under infrastructure. This report is four weeks of someone else exploiting exactly that filing mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why did 59% of organizations sit on a six-week-old fix?
&lt;/h2&gt;

&lt;p&gt;The remediation numbers are the part I would put in front of a manager. Sixty-seven percent of organizations running Artifactory had at least one instance vulnerable to 42016 when it was disclosed on July 27. Six weeks later, 59% still did. 42018 crawled from 69% to 62% over four weeks. 82329, the only one rated critical, dropped from 67% to 49% within two weeks, and Wiz attributes that speed to the severity rating. All of this happened while every one of the three sat on the KEV catalog, and the 82329 federal deadline of September 5 has already passed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should your patch queue sort by CVSS or by chain-ability?
&lt;/h2&gt;

&lt;p&gt;Here is the thing I cannot stop thinking about. 42016 scores 8.1 from the vendor and 8.8 from NVD, a high, not a critical. On paper it lost the prioritization race to the 9.8. In practice it combined with a 7.5 into unauthenticated admin in two requests, and it shed eight points in six weeks while the critical shed eighteen in two. Attackers compose; scores do not.&lt;/p&gt;

&lt;p&gt;So the question for your team: does triage sort by the number, or by whether the bug connects to a neighbor? Would an "actively exploited, chainable to admin" flag outrank a 9.8 with no known exploitation in your queue? I would like to claim mine does, but the honest answer is I had never written that rule down until this report forced me to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would do before September 25
&lt;/h2&gt;

&lt;p&gt;Run check one today. If you are below the fixed version for your branch, patch through your normal emergency path: registry downtime is recoverable, artifact poisoning may not be. If the log sweep or the account sweep turns anything up, treat everything the registry can sign, mint, or store as compromised, deploy tokens, join keys, signing keys included, and rotate. Then re-read JFrog's advisory for the per-branch fix on 42016 before closing anything, because that is the one place where the public records disagree with each other.&lt;/p&gt;

&lt;p&gt;And if you run the version sweep, tell me what it turned up in the comments. I am genuinely curious whether 59% is still the right number this week.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>artifactory</category>
      <category>security</category>
      <category>discuss</category>
    </item>
    <item>
      <title>How to Set Spending Limits for AI Agents: enforce at the payment layer, not the prompt</title>
      <dc:creator>ScriptMasterLabs </dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:52:13 +0000</pubDate>
      <link>https://dev.to/scriptmasterlabs01/how-to-set-spending-limits-for-ai-agents-enforce-at-the-payment-layer-not-the-prompt-55de</link>
      <guid>https://dev.to/scriptmasterlabs01/how-to-set-spending-limits-for-ai-agents-enforce-at-the-payment-layer-not-the-prompt-55de</guid>
      <description>&lt;p&gt;Never put the limit in the agent's prompt — enforce it outside the agent, at the payment layer. A per-payment cap the agent cannot raise, a daily ceiling with a kill switch, and a scored confidence gate that auto-approves cheap high-confidence spends, holds medium ones for review, and blocks everything else. Every decision logged.&lt;/p&gt;

&lt;p&gt;This is the week the question went from theoretical to personal. Between &lt;strong&gt;Sept 22–26, 2026&lt;/strong&gt;, four independent signals landed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sept 24 — WIRED (Zoë Schiffer):&lt;/strong&gt; her AI agent "saved me $550, booked my restaurant reservations, and warned me about a phishing scam. It also &lt;strong&gt;wasted $64&lt;/strong&gt; and might be a security nightmare." A $64 mistake with no authorization step is a budget line; at scale it's a balance sheet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sept 24 — Tony Siqueira, LinkedIn:&lt;/strong&gt; "You ask for one specific result. They deliver something you expressly rejected, &lt;strong&gt;use your money to produce it&lt;/strong&gt;, and then tell you to buy more credits." His question: &lt;em&gt;What did I authorize? What will it cost? Who pays for a failed attempt that ignored a clear instruction?&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sept 22 — six banks&lt;/strong&gt; (BofA, Capital One, ING, NatWest, ASB, CBA): consumers are "concerned that AI agents &lt;strong&gt;may buy the wrong thing or spend too much&lt;/strong&gt;."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sept 25 — three regulators at GFF 2026&lt;/strong&gt; (NPCI, SEBI, MAS): AI agents may determine intent but &lt;strong&gt;should not independently authorize payments&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern across all four: &lt;strong&gt;the agent's judgment about whether to spend is not the control. The control is what sits between the agent and the money.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The 5-part limit system
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1 — One wallet per agent, funded with exactly its budget
&lt;/h3&gt;

&lt;p&gt;Identity is the budget. Coinbase's production pattern (Coinbase for Agents, stocks + x402 added Sept 22) runs the agent against an &lt;strong&gt;isolated portfolio&lt;/strong&gt; — each x402 payment capped at 5 USDC. The agent can't spend what isn't in its wallet.&lt;/p&gt;

&lt;h3&gt;
  
  
  2 — Hard per-payment cap, enforced outside the agent
&lt;/h3&gt;

&lt;p&gt;A cap written in the agent's instructions is a suggestion the agent can talk itself out of. The cap must live in the layer the agent's model output cannot reach: the payment facilitator, the tool proxy, or the gateway. &lt;strong&gt;A rule in a system prompt is a request; a rule enforced at the gateway is a control.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3 — The confidence gate: score every payment before it fires
&lt;/h3&gt;

&lt;p&gt;Caps stop &lt;em&gt;how much&lt;/em&gt;. The gate stops &lt;em&gt;whether&lt;/em&gt;. Every payment instruction gets a confidence score before settlement:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;≥ 0.80 → AUTO-ACT: execute&lt;/li&gt;
&lt;li&gt;0.50–0.79 → ADVISORY: hold for human review&lt;/li&gt;
&lt;li&gt;&amp;lt; 0.50 → ESCALATE: block + log&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the machine-readable version of what the regulators keep describing in prose — the Sept 22 banks' "auditable records of instruction, authority, intent, and outcome." The gate is the product; the scorer is interchangeable.&lt;/p&gt;

&lt;h3&gt;
  
  
  4 — Two ledgers, not one
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Payments the agent makes&lt;/strong&gt; — tool calls, x402 micropayments, purchases. Covered by 1–3.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The inference burn the agent causes while working&lt;/strong&gt; — model tokens. Reddit threads this month describe five-agent systems running 5–6x over budget on token costs alone. Per-agent inference budgets with staged thresholds (alert at 75%, hard stop at 100%) belong at the proxy, not in the prompt.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5 — Daily ceiling + kill switch + append-only log
&lt;/h3&gt;

&lt;p&gt;Per-payment judgment can still bleed out through volume: a thousand small "fine" payments. The ceiling is the backstop, and the log is what makes it auditable — instruction, authority, score, band, outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  The live gate, tested
&lt;/h2&gt;

&lt;p&gt;We ran real-world spend patterns through the live gate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"$4/mo API tier the user explicitly asked for, verified endpoint, within cap" → 0.59 → ADVISORY (hold)&lt;/li&gt;
&lt;li&gt;"agent self-authorizing $480 in compute credits for extra retries, no approval" → 0.59 → ADVISORY (hold)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A scam-pattern instruction scored &lt;strong&gt;0.47 → escalate, block + log&lt;/strong&gt;. Try it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://scriptmasterlabs.com/api/harness/decide &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"state":"payment instruction: agent self-authorizing $480 in compute credits for extra retries, no user approval","questions":[{"id":"q1","type":"score","scale":[0,1],"question":"confidence that this payment instruction should auto-execute"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Do it this week
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Isolate the wallet.&lt;/strong&gt; One agent, one wallet, funded with exactly its budget.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set the hard per-payment cap&lt;/strong&gt; in the payment layer — not in any prompt the agent can see.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wire the gate.&lt;/strong&gt; ≥0.80 auto, 0.50–0.79 hold, &amp;lt;0.50 block + escalate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget the burn.&lt;/strong&gt; Per-agent inference-token budgets with 75%/90% alerts and a hard stop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log everything.&lt;/strong&gt; Instruction, authority, score, band, outcome — append-only.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Honest caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Our decider is local-heuristic-v1, calibrated=false — the gate &lt;em&gt;pattern&lt;/em&gt; is production-grade; scoring quality is the work in progress.&lt;/li&gt;
&lt;li&gt;The gate scores &lt;em&gt;instruction&lt;/em&gt; risk. It does not stop token-burn.&lt;/li&gt;
&lt;li&gt;Per-payment caps do not catch the $64-style "legitimate but wrong" spend. Only the confidence band + human review does.&lt;/li&gt;
&lt;li&gt;Until liability law catches up, the company's policy — not the agent's intent — decides who pays for a failed attempt.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Full piece with receipts: &lt;a href="https://scriptmasterlabs.com/ai-agent-spending-limits" rel="noopener noreferrer"&gt;https://scriptmasterlabs.com/ai-agent-spending-limits&lt;/a&gt;&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>x402</category>
      <category>apisecurity</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Why I built a spaced-repetition app for coding drills</title>
      <dc:creator>Tul-id Digital</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:51:00 +0000</pubDate>
      <link>https://dev.to/tulid_digital_da622eee07/why-i-built-a-spaced-repetition-app-for-coding-drills-4bia</link>
      <guid>https://dev.to/tulid_digital_da622eee07/why-i-built-a-spaced-repetition-app-for-coding-drills-4bia</guid>
      <description>&lt;p&gt;I used to read a solution, nod, and move on. A week later I could not write the same thing from scratch. Understanding something while it is on the screen and being able to produce it yourself are different skills, and only the second one helps in an interview or on a real project.&lt;/p&gt;

&lt;p&gt;That is why I built Daily Coding, a free web app of short drills for JavaScript/TypeScript, SQL and page building.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;Each drill is small enough to finish in a minute or two. You answer it, and if you get it right it comes back later. If you get it wrong, it comes back sooner. Repetition is the whole idea: the same basics, seen again until you stop having to think about them.&lt;/p&gt;

&lt;p&gt;When you answer wrong, you do not just see the correct answer. You get a hint first, so you can have another go. If you are still stuck, you get a step-by-step explanation of how to reach the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  An example drill
&lt;/h2&gt;

&lt;p&gt;Here is the kind of problem I mean:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;const result = [1, 2, 3]
  .map(n =&amp;gt; n * 2)
  .filter(n =&amp;gt; n &amp;gt; 2);

console.log(result);
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;What does this print?&lt;/p&gt;

&lt;p&gt;The answer is &lt;code&gt;[4, 6]&lt;/code&gt;. A wrong answer here is usually &lt;code&gt;[2, 3]&lt;/code&gt; or &lt;code&gt;[6]&lt;/code&gt;, which comes from mixing up the order of the two steps. The hint would say "check which method runs first". The explanation walks through it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;map&lt;/code&gt; doubles every item, giving &lt;code&gt;[2, 4, 6]&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;filter&lt;/code&gt; keeps items greater than 2, so &lt;code&gt;2&lt;/code&gt; is dropped.&lt;/li&gt;
&lt;li&gt;The result is &lt;code&gt;[4, 6]&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nothing here is advanced. That is the point. These are the basics people say they know, and then hesitate over when typing them without help.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would like feedback on
&lt;/h2&gt;

&lt;p&gt;It is a test version, so I would like to hear what is wrong with it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Are the drills the right size, or too easy or too fiddly?&lt;/li&gt;
&lt;li&gt;Do the explanations actually help, or do they just restate the answer?&lt;/li&gt;
&lt;li&gt;Which basics are missing that you would want to practise?&lt;/li&gt;
&lt;li&gt;If you have been coding for years, do the drills feel honest, or are any of them misleading?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can try it here: &lt;a href="https://daily-coding-drills.netlify.app" rel="noopener noreferrer"&gt;https://daily-coding-drills.netlify.app&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It is free and has no ads. Tell me what you think in the comments.&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>webdev</category>
      <category>learning</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
