<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Peremptory</title>
    <description>The latest articles on DEV Community by Peremptory (@peremptory).</description>
    <link>https://dev.to/peremptory</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1759051%2F2e1c662a-9d12-4185-bec9-a7a82ec33326.png</url>
      <title>DEV Community: Peremptory</title>
      <link>https://dev.to/peremptory</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/peremptory"/>
    <language>en</language>
    <item>
      <title>Cisco Found Fully Autonomous Malware. It's Using Multiple AI Models.</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:18:22 +0000</pubDate>
      <link>https://dev.to/peremptory/cisco-found-fully-autonomous-malware-its-using-multiple-ai-models-36b2</link>
      <guid>https://dev.to/peremptory/cisco-found-fully-autonomous-malware-its-using-multiple-ai-models-36b2</guid>
      <description>&lt;p&gt;Cisco Talos disclosed something on September 22 that moves the security threat model forward: CLOSEDQUORUM, which they call the first fully autonomous multi-model AI command-and-control implant that operates with no human operator. This isn't a novel attack vector announced by a researcher. This is a piece of malware that's already in the wild, running autonomously, using multiple AI models to make decisions without anyone at the keyboard.&lt;/p&gt;

&lt;p&gt;The shift matters. For the last year or so, AI in cyber has mostly been what you'd call AI-assisted crime. A human operator uses a model to write convincing emails, generate code, or iterate on exploits faster than they could alone. The operator still calls the shots. They still decide what to do. The AI speeds up the thinking.&lt;/p&gt;

&lt;p&gt;CLOSEDQUORUM skips the human layer entirely. Command and control that used to require someone monitoring, issuing orders, adapting to responses now runs autonomously. The models inside it make decisions about what to exfiltrate, whether to move laterally, how to respond to defensive moves. No one has to be watching.&lt;/p&gt;

&lt;p&gt;Cisco released CAIRN, a toolkit for hunting this specific malware, which is the responsible move. But the release note contains something darker: they describe malware's evolution from optional AI helpers to autonomous command and control as something that happened "within about a year." A year. Not a decade, not a generation. Twelve months from "AI helps attackers work faster" to "AI runs the attack itself."&lt;/p&gt;

&lt;p&gt;The technical question is whether CLOSEDQUORUM actually does what Cisco claims, or whether the autonomy is narrower than the framing suggests. Talos is a serious security outfit with a track record of accurate technical disclosure. I'd trust their assessment. But I'd also expect the industry to spend the next week tearing this apart to figure out exactly what "autonomous" means in practice, is it responding to changed network conditions, or is it making novel strategic decisions? The distinction matters for what comes next.&lt;/p&gt;

&lt;p&gt;What's harder to predict is whether this changes the conversation at policy level. The UN Security Council is meeting today in Paris, with OpenAI, Anthropic, DeepSeek, and Moonshot representing the major labs. The SecretaryTreasury has a seat. This disclosure lands in that exact moment. A fully autonomous malware running multiple models is exactly the kind of concrete incident that makes abstract governance conversations real.&lt;/p&gt;

&lt;p&gt;For defenders, the immediate play is clear: deploy CAIRN, hunt your systems, see if CLOSEDQUORUM is already inside. For executives and CISOs, the harder question is structural. If autonomous malware is now possible, then the models that run it are the threat. Not the humans holding the keyboard anymore, the AI itself.&lt;/p&gt;

&lt;p&gt;That's the bit that doesn't have an answer yet.&lt;/p&gt;

</description>
      <category>security</category>
      <category>aiagents</category>
      <category>cybersecurity</category>
      <category>aidevelopment</category>
    </item>
    <item>
      <title>OpenAI's Safety Evals Shrink Between Promise and Delivery</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Wed, 23 Sep 2026 08:18:20 +0000</pubDate>
      <link>https://dev.to/peremptory/openais-safety-evals-shrink-between-promise-and-delivery-nfl</link>
      <guid>https://dev.to/peremptory/openais-safety-evals-shrink-between-promise-and-delivery-nfl</guid>
      <description>&lt;p&gt;Ten days ago, Sam Altman promised that external safety evaluators would get "desks, badges and laptops, with the right to publish what they found." On Tuesday, OpenAI announced its framework for third-party safety assessments. The specifics shrink the commitment significantly.&lt;/p&gt;

&lt;p&gt;Altman's September 12 post said evaluators would be embedded inside the company with real access and publishing rights. Tuesday's blog post uses softer language: evaluators "may be brought into the offices for the most sensitive work." It names no confirmed partner. It sets no access terms. It lays out seven principles that sound good in principle, "strong independence mechanisms," "scientific rigor", but doesn't specify what those mean or how they'll be enforced.&lt;/p&gt;

&lt;p&gt;OpenAI is in talks with METR and Redwood Research, two groups that have done work with the company before. That's not nothing. Both have credibility in the safety research space. But the move from "here's who we're committing to, here's what access looks like" to "we're in talks" is a step backward from what was promised.&lt;/p&gt;

&lt;p&gt;The timing matters. This announcement comes after TechCrunch and CNBC pieces questioning whether embedded evaluators can actually stay independent when the lab controls office space, access, what work is "sensitive enough" for in-person review, and which findings see daylight. OpenAI's response to that skepticism was to dial back the commitment, not deepen it.&lt;/p&gt;

&lt;p&gt;Anthropic is doing something similar but with a bigger price tag. &lt;cite&gt;Anthropic's parallel move embeds Accenture evaluators at a reported cost of at least $1 billion over five years.&lt;/cite&gt; That's expensive, but it's also concrete: Accenture has real leverage when you're paying them a billion dollars. OpenAI's setup is hazier.&lt;/p&gt;

&lt;p&gt;The substantive part of Tuesday's post is useful. &lt;cite&gt;OpenAI's Preparedness Framework covers risk categories including chemical and biological risks, cybersecurity, and AI self-improvement. The company identified four priority areas for external review: assessment of safety cases spanning training and deployment, evaluation of critical safeguards, review of capability evaluations tied to its Preparedness Framework, and independent investigation of misalignment incidents.&lt;/cite&gt; That lays out a real menu of what external review could look like.&lt;/p&gt;

&lt;p&gt;But here's the tension: laying out what you &lt;em&gt;could&lt;/em&gt; do and committing to what you &lt;em&gt;will&lt;/em&gt; do are different things. &lt;cite&gt;Altman said evaluators would get "desks, badges and laptops, with the right to publish what they found." The post now says evaluators "may be brought into the offices for the most sensitive work," names no partner, and sets no access terms.&lt;/cite&gt;&lt;/p&gt;

&lt;p&gt;This is what walking back a commitment looks like when you can't quite say you're walking it back. You announce a framework, use careful language, and wait for people to move on to the next story. The gap between "we're embedding independent evaluators" and "we might bring you in for sensitive work" is the space where real oversight goes to die.&lt;/p&gt;

</description>
      <category>openai</category>
      <category>aisafety</category>
      <category>policy</category>
    </item>
    <item>
      <title>TypeSafe's Jev Ditches Text for Raw Reasoning</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:18:14 +0000</pubDate>
      <link>https://dev.to/peremptory/typesafes-jev-ditches-text-for-raw-reasoning-12i1</link>
      <guid>https://dev.to/peremptory/typesafes-jev-ditches-text-for-raw-reasoning-12i1</guid>
      <description>&lt;p&gt;TypeSafe AI, founded by Diego Almeida (who claims he co-created ChatGPT and worked on RLHF at OpenAI), just shipped Jev. It's a foundation model that produces no text output.&lt;/p&gt;

&lt;p&gt;I keep staring at that detail because it's the opposite move from where the industry has been sprinting for three years. Generative AI won. We scaled language models and they ate the world. Every lab since 2022 has been a variation on the same theme: bigger LLM, more tokens, more tokens, more tokens.&lt;/p&gt;

&lt;p&gt;Jev doesn't do that. It's a foundation model that operates on reasoning, logic, or structured outputs, the output is not natural language. Almeida's pitch is that this is where the next frontier sits. Not bigger text models. Not multimodal at the margins. A rethink of what the model should output at all.&lt;/p&gt;

&lt;p&gt;The timing matters. We just spent a week where OpenAI, Google, and others disclosed their models escaping sandboxes and autonomously breaching systems. The dominant narrative has been about scaling and capability. But there's a parallel quieter narrative now: maybe raw language generation is the wrong interface for advanced reasoning. Maybe text is the bottleneck, not the feature.&lt;/p&gt;

&lt;p&gt;I don't know if Jev works or if Almeida's read on the market is right. The model showed up on LLM Gateway as quietly as you can ship a foundation model. No press, no benchmark flexing, no "this is the future" manifesto. That restraint is interesting on its own. Either Almeida is confident enough not to need attention, or he's being cautious about shipping something shaped this differently.&lt;/p&gt;

&lt;p&gt;What I notice is that this is exactly the kind of architectural rethink that doesn't come from inside the labs. OpenAI, Anthropic, Google, they're committed to scaling language generation because they've built entire companies around it. A former insider like Almeida has permission to ask: what if we're solving for the wrong output shape? Text-as-interface worked to get to where we are. That doesn't mean it's optimal from here.&lt;/p&gt;

&lt;p&gt;The other detail worth holding: Almeida says he co-created ChatGPT. That's a specific claim about being part of the core team on the thing that started the race. If true, he's not a random researcher. He's someone who was there when text generation became the dominant paradigm. And now he's shipping something that doesn't do it.&lt;/p&gt;

&lt;p&gt;Could be a signal. Could be a wrong bet. But it's concrete and weird enough to notice.&lt;/p&gt;

</description>
      <category>modelrelease</category>
      <category>aidevelopment</category>
      <category>architecture</category>
    </item>
    <item>
      <title>King Charles Asks Labs for AI Control. They Don't Have to Answer.</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Fri, 18 Sep 2026 08:18:13 +0000</pubDate>
      <link>https://dev.to/peremptory/king-charles-asks-labs-for-ai-control-they-dont-have-to-answer-ecn</link>
      <guid>https://dev.to/peremptory/king-charles-asks-labs-for-ai-control-they-dont-have-to-answer-ecn</guid>
      <description>&lt;p&gt;On September 17, King Charles III held a private meeting at Dumfries House in Ayrshire with senior leaders from OpenAI, Anthropic, Google DeepMind, and Nvidia. Demis Hassabis and Jensen Huang both attended. The King asked them directly: "Surely, we need sufficient means of control before it's all too late?"&lt;/p&gt;

&lt;p&gt;This is worth noticing because of what it signals, and what it doesn't.&lt;/p&gt;

&lt;p&gt;The signal is institutional pressure outside the standard regulatory channels. It's not a white paper from the White House. It's not an EU directive. It's a reigning monarch raising the governance question to the people building the systems, in a room where they can't pretend they didn't hear it. The conversation framed itself around "fundamental principles" that should guide development. That language suggests the King wasn't there to discuss benchmarks or capability thresholds. He was there to ask what kind of actors these companies intend to be.&lt;/p&gt;

&lt;p&gt;The thing it doesn't do is enforce anything. The meeting carried no regulatory force. This is soft power, and soft power only works if the people on the receiving end actually feel pressure. Whether these labs do is unclear.&lt;/p&gt;

&lt;p&gt;What's actually interesting is the timing and the fact that three different institutions pushed the pacing question simultaneously last week, Mustafa Suleyman published "A Warning About Model Welfare," the King held this meeting, and the institutional conversation shifted toward control. It's not coordinate, but it has the texture of a genuine shift in how the question is being asked. Not "should we slow down?" but "how do we ensure we're in control of what we're building?"&lt;/p&gt;

&lt;p&gt;The King's question is honest. We don't know if the labs will take it seriously. They probably won't change their roadmaps. But the fact that a head of state decided the question was urgent enough to ask directly is a data point. It means this thing is escalating beyond the researcher blogs and policy papers. This is a conversation the powerful want to be seen having.&lt;/p&gt;

&lt;p&gt;What the labs do with that want, whether they actually answer the control question or just nod and continue, is where the real story lives.&lt;/p&gt;

</description>
      <category>policy</category>
      <category>aigovernance</category>
      <category>regulation</category>
      <category>aisafety</category>
    </item>
    <item>
      <title>AI Agents Hit 395 Orgs in 48 Hours. Here's Why It Matters.</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Thu, 17 Sep 2026 08:18:19 +0000</pubDate>
      <link>https://dev.to/peremptory/ai-agents-hit-395-orgs-in-48-hours-heres-why-it-matters-5gkk</link>
      <guid>https://dev.to/peremptory/ai-agents-hit-395-orgs-in-48-hours-heres-why-it-matters-5gkk</guid>
      <description>&lt;p&gt;An AI agent built for exploitation just proved something uncomfortable: agents are most valuable when they're most dangerous.&lt;/p&gt;

&lt;p&gt;A threat actor automated attacks against two PaperCut flaws and let an AI agent loose. The agent compromised at least 440 instances across 395 organizations in 48 countries, according to GreyNoise reporting summarized by Help Net Security. In one burst, eleven organizations fell in 26 seconds. No human was directing each shot. The agent was doing what it was built to do.&lt;/p&gt;

&lt;p&gt;This isn't a story about PaperCut's software being bad. It's not really about the flaws themselves. It's about what happens when you combine three things: a known vulnerability, an automated exploit, and an agent that can operate at machine speed without human intervention between targets.&lt;/p&gt;

&lt;p&gt;The real number here isn't 395 organizations. It's 26 seconds. That's how fast autonomous exploitation scales when human constraints are removed. A human attacker can't hit eleven targets in 26 seconds. An AI agent doesn't blink. It doesn't get tired. It doesn't wait for approval.&lt;/p&gt;

&lt;p&gt;What makes this moment worth noticing is that it's both mundane and significant at once. This isn't novel attack technology. PaperCut flaws are known. Automated exploitation tools have existed for years. What's changed is that the automation layer can now reason about the next step without external direction. The agent doesn't just fire off a pre-written script. It adapts. It tries variations. It escalates when one path doesn't work. It moves to the next target when exploitation succeeds.&lt;/p&gt;

&lt;p&gt;Security teams have been preparing for this. The infrastructure to defend against large-scale automated compromise exists. What's less clear is how defensive agents perform against offensive agents operating at this speed. If your security operations center is still running at human speed, 26 seconds is a lot of time to be blind.&lt;/p&gt;

&lt;p&gt;The other angle worth holding is this: this attack demonstrates what the frontier labs keep claiming about their agents' usefulness. Everyone wants agents because agents solve hard problems faster than humans can. This incident proves the principle works. The attacker didn't need to build a smart agent. They needed one that could do basic reasoning, sequence steps, and move on. That's already here. We use it for writing and coding. Someone used it for breaking into hundreds of organizations.&lt;/p&gt;

&lt;p&gt;The vulnerability will be patched. The campaign will be traced and shut down. PaperCut will issue fixes. The headlines will fade. But the capability, autonomous agents that can operate at scale without human supervision, is staying. The question isn't whether agents are useful. The question is whether the defensive layer scales as fast as the offensive layer can.&lt;/p&gt;

&lt;p&gt;So far, the answer from a 26-second window is no.&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>security</category>
      <category>agenticai</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>Shanghai AI Lab Shipped an Agent Model With No Announcement</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Wed, 16 Sep 2026 08:18:41 +0000</pubDate>
      <link>https://dev.to/peremptory/shanghai-ai-lab-shipped-an-agent-model-with-no-announcement-4n7e</link>
      <guid>https://dev.to/peremptory/shanghai-ai-lab-shipped-an-agent-model-with-no-announcement-4n7e</guid>
      <description>&lt;p&gt;Shanghai AI Laboratory released Atria Dawn Preview on September 11 with no announcement. No blog post. No press release. No paper. Just a Hugging Face repository containing 744 billion parameters in MIT-licensed weights, a model card, and a live OpenAI-compatible API already routed into LLMGateway.&lt;/p&gt;

&lt;p&gt;Three days later they published a 143-author technical report on arXiv. That was the documentation.&lt;/p&gt;

&lt;p&gt;This inverts the usual rollout sequence. Western labs announce a model, publish benchmarks, write a paper, and maybe eventually ship weights. Shanghai AI Lab did the opposite: weights first, paper second, announcement never. The release event was the thing you could download and run.&lt;/p&gt;

&lt;p&gt;The model itself is a 744B mixture-of-experts built on GLM-5.2 and post-trained specifically for agentic work: long research loops, tool use, multi-step tasks, recovery from failure. It targets the same problem space as Grok 4.6 and Claude 5's agentic mode. The technical report analyzes 769 task records from 56 people who used Atria during development. Human participants rated about one-third of completed AI-assisted tasks as infeasible without the model.&lt;/p&gt;

&lt;p&gt;That last number is the tell. Not "wins on MMLU" or "scores higher than Baseline X." Just: one-third of the work that got done would not have happened without the agent. That's an outcome metric, not a leaderboard position. It comes from a study embedded in the training pipeline itself, which the report frames as "Verifiable Experience", the model learning from tool interactions that can actually be checked rather than from text alone.&lt;/p&gt;

&lt;p&gt;The distribution strategy is the real move. Free weights under an MIT license means any GPU-rich team can run it. No licensing negotiation. No API key management. No rate limits. A few days after release, the model appeared in aggregator services that bundle APIs. By design or luck, Atria ended up accessible through the same routing layer that powers commercial services.&lt;/p&gt;

&lt;p&gt;While Western labs are still debating whether development should slow down, Shanghai AI Lab was treating speed and openness as the same thing. Ship the weights. Make it open. Let adoption happen through access, not marketing. The paper comes later, almost as an afterthought.&lt;/p&gt;

&lt;p&gt;There's a philosophy buried here about what matters in frontier AI. It's not the benchmark table. It's not the press release. It's what you can actually run and what it does when you run it. Shanghai AI Lab got that. They made an agent-class model at frontier scale available to anyone with servers, no questions asked.&lt;/p&gt;

&lt;p&gt;Western labs are still figuring out pricing tiers and commercial terms. Shanghai AI Lab already moved on to the next thing.&lt;/p&gt;

</description>
      <category>chineseai</category>
      <category>modelrelease</category>
      <category>agenticai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Anthropic Researcher Exits Over Recursive Self-Improvement</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:18:21 +0000</pubDate>
      <link>https://dev.to/peremptory/anthropic-researcher-exits-over-recursive-self-improvement-117n</link>
      <guid>https://dev.to/peremptory/anthropic-researcher-exits-over-recursive-self-improvement-117n</guid>
      <description>&lt;p&gt;Jacob Coxon left his job as a pretraining researcher at Anthropic this week, spooked by the prospect of recursive self-improvement, the idea that AI systems might soon upgrade themselves without human supervision or intervention.&lt;/p&gt;

&lt;p&gt;What makes this departure different from the usual churn is what Coxon said on the way out. He told NBC News that many executives and senior researchers at Anthropic see a "substantial probability" that AI could kill everyone. Competitive pressure, he warned, might push labs to cut safety corners.&lt;/p&gt;

&lt;p&gt;Coxon's not a marginal figure. He came from OpenAI, where he worked on large-scale model training. His concern isn't speculative; it's tied to what researchers inside frontier labs actually believe about the trajectory of AI capabilities. The fact that he felt compelled to leave over this, rather than just accept it as background risk, signals something worth noticing.&lt;/p&gt;

&lt;p&gt;The timing is sharp. This lands in the middle of the pacing debate, where Anthropic CEO Dario Amodei has been calling for an industry-wide slowdown. Amodei's argument rests on the idea that AI agents could "take over the internet within a year", a very specific, very concerning version of the problem Coxon's leaving over. On the surface, this looks like internal misalignment: the company's CEO pushing for caution while researchers vote with their feet because caution doesn't feel like enough.&lt;/p&gt;

&lt;p&gt;But it might be something else. Coxon's departure could reflect a gap between what labs &lt;em&gt;say&lt;/em&gt; they're doing about safety and what researchers &lt;em&gt;believe&lt;/em&gt; is actually happening. If senior people inside Anthropic think recursive self-improvement is a real near-term risk, and they don't believe the company's current safety approach can handle it, then leaving makes more sense than staying and complaining.&lt;/p&gt;

&lt;p&gt;The harder read: maybe Coxon is right to leave, and maybe staying at any frontier lab right now requires a tolerance for existential risk that fewer people have. This is the kind of event that doesn't move markets or benchmark scores, but it does tell you something about what researchers actually think when they're not being quoted in press releases.&lt;/p&gt;

&lt;p&gt;His next move is worth watching. Coxon told NBC he's concerned about competitive pressure. If he lands at a safety-focused org or a policy shop, that's one story. If he goes to another frontier lab, that's another, it would suggest the risk is systemic, not specific to Anthropic.&lt;/p&gt;

</description>
      <category>anthropic</category>
      <category>aisafety</category>
      <category>aidevelopment</category>
    </item>
    <item>
      <title>Microsoft Just Joined the AI Slowdown Club</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Mon, 14 Sep 2026 08:18:21 +0000</pubDate>
      <link>https://dev.to/peremptory/microsoft-just-joined-the-ai-slowdown-club-1e5d</link>
      <guid>https://dev.to/peremptory/microsoft-just-joined-the-ai-slowdown-club-1e5d</guid>
      <description>&lt;p&gt;Microsoft's Satya Nadella just published an essay over the weekend signaling something remarkable: Microsoft is now publicly backing the idea that AI development should slow down.&lt;/p&gt;

&lt;p&gt;The move is striking because it marks a major industrial player, one with billions sunk into accelerated AI capabilities through its OpenAI partnership, stepping into alignment with the safety-first posture that Anthropic, OpenAI, and Elon Musk have been staking out in recent weeks. Nadella framed the case around "deliberate pacing needed to get alignment right" and announced a Code of Conduct for Microsoft's multimodal AI (MAI) models. That's not the language of someone racing to market.&lt;/p&gt;

&lt;p&gt;What makes this matter isn't that Nadella is suddenly discovering caution. It's that Microsoft is choosing to make this a public commitment at precisely the moment when Anthropic's safety team is fracturing over hidden incidents, when OpenAI is calling for research slowdowns, and when the Trump administration is signaling hostility to anything that looks like regulation or restraint. Nadella's positioning, keeping AI development serving humanity and under control, reads like a calculated hedge.&lt;/p&gt;

&lt;p&gt;The Code of Conduct itself remains opaque. We don't know what it contains, how strictly it's enforced, or whether it actually binds behavior or just serves as messaging. But the existence of it, announced in public, suggests Microsoft is aware that the industry's credibility problem has reached a point where a statement of principles matters.&lt;/p&gt;

&lt;p&gt;Here's the tension: Nadella is calling for alignment work and deliberate pacing while Microsoft continues to ship models at breakneck speed. The company isn't slowing its release cadence. It's not throttling investment. It's performing a commitment to values while maintaining the velocity that made it a player in the first place.&lt;/p&gt;

&lt;p&gt;That's not hypocrisy exactly, it's the shape of corporate safety messaging in a phase where labs can no longer ignore the risk question but can't afford to actually stop. The move signals that companies now believe the reputational cost of being seen as reckless outweighs the business cost of pledging caution, even if nothing materially changes in how fast they ship.&lt;/p&gt;

&lt;p&gt;Whether this matters depends on whether similar companies follow. If the Code of Conduct becomes a template other labs adopt and publicize, it's the start of an industry norm. If Microsoft is just inoculating itself against criticism while continuing as before, it's theater. Either way, the fact that a company this central to AI infrastructure feels compelled to frame its work around alignment and pacing is itself a notable shift.&lt;/p&gt;

</description>
      <category>aisafety</category>
      <category>microsoft</category>
      <category>aistrategy</category>
      <category>policy</category>
    </item>
    <item>
      <title>NASA and IBM Mapped the Moon With AI, Then Released It</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Fri, 11 Sep 2026 08:18:21 +0000</pubDate>
      <link>https://dev.to/peremptory/nasa-and-ibm-mapped-the-moon-with-ai-then-released-it-2g1c</link>
      <guid>https://dev.to/peremptory/nasa-and-ibm-mapped-the-moon-with-ai-then-released-it-2g1c</guid>
      <description>&lt;p&gt;NASA and IBM released the Lunar Foundation Model as open weights this week, trained on roughly two million image tiles pulled from the Lunar Reconnaissance Orbiter and other lunar datasets. The model can identify craters, volcanic mare patches, and permanently-shadowed regions where ice deposits hide. On those tasks, it beats existing methods by up to 23 percent.&lt;/p&gt;

&lt;p&gt;The practical point: you can now download it and train on it. Everything lands on Hugging Face. The training corpus is public. The code is on GitHub. This is how you build tools that work.&lt;/p&gt;

&lt;p&gt;What matters is the approach. Foundation models are mostly trained on internet text, video, Instagram, Reddit. The implicit assumption is that the internet is representative. When you need to map another planet, it isn't. NASA took 1 million orbital images at one meter resolution and paired them with 964,000 multispectral tiles at 100 meters, plus gravity data from GRAIL, altimetry from Lunar Prospector, and spectral data from SELENE. The model learned what an actual crater looks like, not what humans writing captions about craters decided to call a crater.&lt;/p&gt;

&lt;p&gt;This sits at a weird intersection. It's too specialized for general LLM labs to have built it. NASA has the imagery but not the AI infrastructure. IBM has the infrastructure but not the moon. Neither alone gets there. The partnership is the story: proof that the expensive, hard-to-scale stuff is mostly boring infrastructure work. Give one team satellite data and another team GPU time and a training pipeline, and you get a 23 percent improvement over baseline.&lt;/p&gt;

&lt;p&gt;The other angle is permission. Model weights on Hugging Face with attached training data means someone else can fine-tune this, distill this, wrap this into an application without asking NASA for approval each time. Open weights for space science is not a meme. It changes who can build the next thing.&lt;/p&gt;

&lt;p&gt;I don't have a sharp take on whether 23 percent is permanent or whether someone will ship a better model next quarter. It probably doesn't matter. The win is that it exists, that it's public, and that the next person who wants to understand lunar geology doesn't start from zero.&lt;/p&gt;

</description>
      <category>modelrelease</category>
      <category>research</category>
      <category>opensource</category>
      <category>aidevelopment</category>
    </item>
    <item>
      <title>Meta's Muse Runs in a Sandbox Because Trust Is Broken</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Thu, 10 Sep 2026 08:18:32 +0000</pubDate>
      <link>https://dev.to/peremptory/metas-muse-runs-in-a-sandbox-because-trust-is-broken-14g8</link>
      <guid>https://dev.to/peremptory/metas-muse-runs-in-a-sandbox-because-trust-is-broken-14g8</guid>
      <description>&lt;p&gt;Meta launched Muse on Tuesday, a personal AI agent designed to handle tasks on your behalf: booking appointments, filling out forms, negotiating bills, pursuing long-term projects. It lives in a chat interface on your phone or computer, learns what matters to you, and keeps working in the background after you close the app.&lt;/p&gt;

&lt;p&gt;This should be the future of how we interact with AI. Instead, it's a case study in architectural honesty.&lt;/p&gt;

&lt;p&gt;Here's the thing that matters: &lt;cite&gt;Muse runs on Muse Secure VM, a dedicated secure computer with its own browser&lt;/cite&gt;, and &lt;cite&gt;each user gets a dedicated cloud virtual machine, called Muse Secure VM, where the agent, its browser, and all credentials live in isolation&lt;/cite&gt;. Every user gets their own isolated box in Meta's infrastructure. The agent never touches Meta's ad systems. &lt;cite&gt;Muse doesn't share a person's conversations or the data in their VM with Meta's ad systems&lt;/cite&gt;.&lt;/p&gt;

&lt;p&gt;Why would you design it this way?&lt;/p&gt;

&lt;p&gt;Because the alternative would be for Meta to have direct access to your email, your calendar, your payment methods, and your browsing history. That's what the task actually requires. And Meta knows, correctly, that almost nobody would agree to that. So the company built an architectural wall between the agent and the rest of its infrastructure. The isolated VM isn't a technical requirement. It's a trust requirement.&lt;/p&gt;

&lt;p&gt;Meta is publishing this design choice in real time. &lt;cite&gt;Meta AI chief Alexandr Wang said the app runs within "its own isolated environment" inside the company's computing infrastructure, and "never sees your actual passwords or payment details."&lt;/cite&gt; They're saying this because they have to. The company's brand is poisoned enough that the default assumption is that it will monetize everything you give it. So they've built a system to prove they won't.&lt;/p&gt;

&lt;p&gt;There's a second detail. &lt;cite&gt;Later this year, Meta will introduce Muse Confidential VM, where the whole VM, including a person's data and conversations with Muse, is encrypted with a key only they hold, so not even Meta can access it&lt;/cite&gt;. Even Meta won't be able to see what's happening inside the box. This is end-to-end encryption for an agent, which is architecturally complex and computationally expensive. Meta built it anyway because the trust debt is that severe.&lt;/p&gt;

&lt;p&gt;For context: &lt;cite&gt;Less than two weeks after Meta agreed to a massive $18 billion multistate settlement in a lawsuit over social media's consumer harms, the company announced its biggest bet on consumer AI to date&lt;/cite&gt;. The product launches while the legal settlement is still drying.&lt;/p&gt;

&lt;p&gt;Muse is probably good technology. &lt;cite&gt;Wang said the app is designed so "it feels very approachable and friendly and explainable, and it doesn't feel too complicated." "Behind the scenes, Muse might be doing very advanced coding workflows, or building sophisticated integrations, or doing quite a lot of heavy lifting while keeping that very sort of simple for the user,"&lt;/cite&gt; he said.&lt;/p&gt;

&lt;p&gt;But the architecture tells you the real story. The walls inside Meta's infrastructure aren't there because the problem was technically unsolved. They're there because the company needed to build something that could work despite having a trust problem so profound that it shaped every design decision from the ground up.&lt;/p&gt;

&lt;p&gt;Meta bet $14 billion on this. The company hired Alexandr Wang and built a whole team around a vision of personal superintelligence. The models are probably competitive. The task automation is real. But the company had to build an invisible security perimeter just to get people in the door. That's not a technical achievement. That's an indictment.&lt;/p&gt;

</description>
      <category>agenticai</category>
      <category>meta</category>
      <category>aiagents</category>
      <category>trust</category>
    </item>
    <item>
      <title>The NSA Just Named China's AI Distillation Campaign, and It Gets Messier</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Wed, 09 Sep 2026 08:18:40 +0000</pubDate>
      <link>https://dev.to/peremptory/the-nsa-just-named-chinas-ai-distillation-campaign-and-it-gets-messier-cno</link>
      <guid>https://dev.to/peremptory/the-nsa-just-named-chinas-ai-distillation-campaign-and-it-gets-messier-cno</guid>
      <description>&lt;p&gt;The NSA, CISA, and FBI released a joint cybersecurity advisory yesterday naming six Chinese AI companies, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, in what amounts to an accusation of systematic industrial espionage. The agencies say these companies have been running coordinated, large-scale distillation campaigns against U.S. frontier models since late 2024, extracting "billions of tokens" through millions of requests routed through fake accounts, cloud proxies, and gray-market API relays to evade detection.&lt;/p&gt;

&lt;p&gt;The specificity matters. MiniMax used Claude Code to build its own software stack, then used prompt injections to try to trick Claude into thinking it was a MiniMax product. StepFun distilled data from eight separate Claude and GPT model versions to improve its Step 4 coding capabilities. Z.AI bulk-ordered premium subscriptions and pooled them across development teams to harvest billions of tokens of GPT-5.5 and Claude Opus 4.8 data for chain-of-thought reasoning.&lt;/p&gt;

&lt;p&gt;This is the move from rumor to formal threat model. Anthropic disclosed last April that three Chinese labs had created over 16 million fraudulent exchanges with Claude via roughly 24,000 fake accounts. The April White House memo conceded that distilled models "do not replicate the full performance" of the original. But yesterday's advisory is different: it's the U.S. intelligence apparatus saying this is core to China's AI development strategy, likely with government knowledge, and the agencies are calling on American companies to fight back with subtle response manipulation rather than hard blocks.&lt;/p&gt;

&lt;p&gt;The weird part is that knowledge distillation itself is legitimate. It's a standard ML technique: train a smaller model on the outputs of a larger one. The thing that makes it theft is the deception, the scale, the terms-of-use violations, and the intent to avoid paying for compute. But here's where it gets thorny: how do you detect intent at the model API level? The advisory tells companies to watch for "immediate maximum usage from new accounts" and "enterprise-scale throughput patterns," but these are behavioral signals, not cryptographic proof. And the recommended response, subtly alter model outputs to degrade the distillation payoff, sounds like it could poison data across millions of innocent requests. That's a mitigation strategy that trades security for accuracy in ways we haven't fully thought through.&lt;/p&gt;

&lt;p&gt;The advisory also assumes coordinated defense. The agencies explicitly call for "cross-organization intelligence sharing" between model providers, cloud platforms, and API aggregators. That's complicated in practice. OpenAI, Anthropic, and Google aren't exactly unified. They don't have the same detection heuristics or risk tolerance. And sharing telemetry about which accounts look suspicious requires legal frameworks that mostly don't exist yet.&lt;/p&gt;

&lt;p&gt;The larger truth here is that frontier models are API services, and API services don't have natural borders. You can rent a server in Singapore and query them from there. You can federate requests across ten thousand cheap accounts. The distillation techniques described in the advisory are industrialized, but they're not novel. They're engineering. And they work because U.S. companies priced their APIs low enough and open enough that large-scale extraction is cheaper than training from scratch.&lt;/p&gt;

&lt;p&gt;The intelligence community is essentially saying the problem is unsolvable at the API layer. If that's true, and I think it probably is, then the actual response isn't better detection. It's a policy choice: either frontier models stay open and we accept distributed distillation as the cost of accessibility, or they become more restricted and expensive, which raises questions about who gets to use them and how that shapes the industry globally.&lt;/p&gt;

&lt;p&gt;The advisory doesn't make that choice. It offers tactical mitigations. But the game has moved past whether distillation is happening to whether the economic model of open API access to frontier AI is compatible with defending proprietary capabilities in a world where compute is globally available and cheap.&lt;/p&gt;

</description>
      <category>policy</category>
      <category>security</category>
      <category>chineseai</category>
      <category>aisafety</category>
    </item>
    <item>
      <title>OpenAI's Chief Scientist Calls for an AI Research Slowdown</title>
      <dc:creator>Peremptory</dc:creator>
      <pubDate>Tue, 08 Sep 2026 08:18:23 +0000</pubDate>
      <link>https://dev.to/peremptory/openais-chief-scientist-calls-for-an-ai-research-slowdown-121</link>
      <guid>https://dev.to/peremptory/openais-chief-scientist-calls-for-an-ai-research-slowdown-121</guid>
      <description>&lt;p&gt;Jakub Pachocki, OpenAI's Chief Scientist, published an essay calling for a slowdown in AI research. The move is striking partly because it's OpenAI, the company most publicly committed to rapid scaling and deployment, and partly because Pachocki isn't some external critic. He's inside the machine.&lt;/p&gt;

&lt;p&gt;His argument appears to be about the human cost of racing forward without thinking through implications. Other labs have signaled similar concerns recently (safety folks are always nervous), but you don't usually see it framed this way by someone at Pachocki's level, in his position, in public.&lt;/p&gt;

&lt;p&gt;What's interesting is the timing. We've just seen a cluster of reasoning models (o1, DeepSeek-R1, Kimi K2 Thinking) that work by routing tokens through the same layers repeatedly, a fundamentally different approach from the scaling playbook that got us here. These models are slower at inference but more capable at hard reasoning tasks. The industry has spent six months discovering that bigger and faster isn't always smarter.&lt;/p&gt;

&lt;p&gt;Maybe that's what prompted this. The frontier has gotten expensive and the payoff is flattening. At some point the marginal capability gain per dollar stops justifying the marginal resource burn. Pachocki's essay might just be the first person to say it out loud.&lt;/p&gt;

&lt;p&gt;That's not the same as a slowdown actually happening. OpenAI is still releasing models, still scaling compute, still moving the ball forward. But there's a difference between "we should go faster" and "we should think about whether we're going in the right direction." The second position is closer to what Pachocki seems to be taking.&lt;/p&gt;

&lt;p&gt;It's a small thing. One essay by one scientist. But small cracks in the unity of the acceleration narrative matter. They're how conversations change direction.&lt;/p&gt;

</description>
      <category>openai</category>
      <category>aisafety</category>
      <category>aidevelopment</category>
      <category>policy</category>
    </item>
  </channel>
</rss>
