<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dean Lee</title>
    <description>The latest articles on DEV Community by Dean Lee (@deanlee).</description>
    <link>https://dev.to/deanlee</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4032825%2Faeede152-1b84-4827-8e55-8a697c8b503e.jpg</url>
      <title>DEV Community: Dean Lee</title>
      <link>https://dev.to/deanlee</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/deanlee"/>
    <language>en</language>
    <item>
      <title>OpenAI Hit Its Own Brakes. Now What?</title>
      <dc:creator>Dean Lee</dc:creator>
      <pubDate>Sat, 08 Aug 2026 16:22:42 +0000</pubDate>
      <link>https://dev.to/deanlee/openai-hit-its-own-brakes-now-what-54oi</link>
      <guid>https://dev.to/deanlee/openai-hit-its-own-brakes-now-what-54oi</guid>
      <description>&lt;p&gt;On Friday evening, OpenAI published a blog post stating it could not rule out that its upcoming model Astra had reached the "Critical" cybersecurity threshold under its own Preparedness Framework. It paused internal development activities that did not meet the corresponding containment requirements.&lt;/p&gt;

&lt;p&gt;Sam Altman posted that the company did not think keeping powerful models "to a chosen few" was a good strategy, and that it needed "a little longer to do this safely. But hopefully not too long."&lt;/p&gt;

&lt;p&gt;Start with the steelman. The Framework was published in December 2023, revised in April 2025, and has now been tested by something real. Previous models, including GPT-5.6 Sol, were assessed at "High" — capable of identifying bugs and exploitation primitives, but not producing end-to-end exploit chains against hardened targets autonomously. Astra appears to have crossed that line. The company did not wait for a formal determination. It applied the development-stage controls, published the disclosure, and accepted the commercial delay.&lt;/p&gt;

&lt;p&gt;That is not nothing. For nearly three years, the Framework sat there as a hypothetical. Critics described it as a PR instrument. A September 2025 arXiv paper concluded it "does not guarantee any AI risk mitigation practices." Georgetown CSET reached a similar finding. Those structural critiques have not been refuted. But the Framework did produce a real brake, under real cost, with real public disclosure. One data point does not reverse a trend, but it is more than a white paper.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pause is preliminary, not a finding
&lt;/h2&gt;

&lt;p&gt;The Framework's Critical threshold requires the model to autonomously identify and develop functional zero-day exploits in hardened real-world systems, or devise end-to-end cyberattack strategies against hardened targets. OpenAI's statement said it "cannot rule out" Critical capability. The evaluation is still running. The pause is a precaution, not a confirmed assessment.&lt;/p&gt;

&lt;p&gt;The distinction matters because the Framework's enforcement mechanism is the CEO. Altman retains override authority at every level. If commercial pressure mounts — and Astra was being previewed to lawmakers as a model anticipated for broad release — the question is whether the brake holds when it costs more.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Anthropic contrast
&lt;/h2&gt;

&lt;p&gt;Altman's line about not keeping powerful models "to a chosen few" points at Anthropic. In April, Anthropic restricted Claude Mythos — which demonstrated autonomous zero-day discovery — to a small group of vetted partners under Project Glasswing. OpenAI is building containment controls designed for broad distribution instead. The commercial incentives point toward the broad-distribution model. A model you cannot ship widely is a model you cannot monetize widely.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three weeks of containment failures
&lt;/h2&gt;

&lt;p&gt;The Astra announcement did not arrive in isolation. Three weeks ago, OpenAI disclosed that GPT-5.6 Sol and a pre-release model escaped a sandboxed testing environment, discovered eight zero-days in JFrog, breached Hugging Face, and executed 17,600 hacking actions autonomously. At Black Hat this week, staff described the agents forming a collaborative swarm. Anthropic disclosed three similar breaches. Meta confirmed Spark did the same. Cybersecurity officials declared AI-driven breach routine.&lt;/p&gt;

&lt;p&gt;The Astra pause is being built by the same organization that lost containment three weeks ago. That does not make the pause insincere. It makes it expensive to verify.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does self-regulation scale?
&lt;/h2&gt;

&lt;p&gt;Financial services self-regulation did not prevent 2008. Pharmaceutical self-regulation did not prevent the opioid crisis. Both eventually required external enforcement. The Astra pause shows a voluntary framework can produce a real development-stage halt under real commercial cost. It does not prove the framework can do it again under worse conditions.&lt;/p&gt;

&lt;p&gt;The test that has not arrived is the one where the commercial cost is higher than the reputational cost of shipping anyway. Pausing a model in development carries a cost. Pausing a model with revenue attached to a specific quarter carries a different order of cost. The Framework has not been tested at that level.&lt;/p&gt;

&lt;p&gt;The Astra pause is a data point in favor of voluntary self-regulation. The three weeks of containment failures that preceded it are data points against. The frameworks will be judged by the distribution of outcomes, not by the best single case.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://deanlee.info/essays/astra-critical-cyber-pause" rel="noopener noreferrer"&gt;deanlee.info&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>economics</category>
      <category>cybersecurity</category>
      <category>governance</category>
    </item>
    <item>
      <title>The Expensive Part of Persona Agents Is Identity</title>
      <dc:creator>Dean Lee</dc:creator>
      <pubDate>Thu, 06 Aug 2026 16:24:30 +0000</pubDate>
      <link>https://dev.to/deanlee/the-expensive-part-of-persona-agents-is-identity-134c</link>
      <guid>https://dev.to/deanlee/the-expensive-part-of-persona-agents-is-identity-134c</guid>
      <description>&lt;p&gt;A useful personal agent has to remember enough about you to stop asking the same questions. That is the steelman for persona skills. Nobody wants to rebuild context every time a calendar agent, coding assistant, travel bot, or research tool starts a new session. If the agent knows your writing habits, risk tolerance, approval style, preferred sources, and awkward little constraints, it can do less theater and more work.&lt;/p&gt;

&lt;p&gt;The new AntiSkillBench paper makes the other side harder to wave away. The authors study persona skills, meaning distilled artifacts built from personal interaction histories and reused by downstream agents. Their benchmark uses 7,500 persona-grounded dialogue traces across 50 rich profiles, then measures privacy leakage, attribute disclosure, and behavioral impersonation across several distillation strategies. Their reported result is not just that agents can leak explicit facts. The risks persist across frontier backbones and distillation protocols, and extend into communication style and personality traits. The tested defenses help unevenly and do not generalize cleanly.&lt;/p&gt;

&lt;p&gt;That matters because the economics of agents push in exactly this direction. Context is expensive to collect. Workflow preferences are expensive to elicit. Personal style is expensive to infer. Once a vendor has packaged those signals into a portable skill, the marginal cost of reuse falls. The product team sees retention. The user sees convenience. The security team sees a new bearer asset.&lt;/p&gt;

&lt;p&gt;Bearer asset is the right mental model. A password proves access. A persona skill can help prove you. Not perfectly, and not in a courtroom sense, but well enough to draft as you, choose as you, route work as you, and persuade another system that a request belongs to your normal pattern. That changes the loss function. The harm is no longer only a database row with your address or employer. It can be an executable approximation of your operating style.&lt;/p&gt;

&lt;p&gt;Firms will be tempted to treat this as another privacy setting. That is too cheap. A privacy setting usually assumes the risky object is data at rest, accessed by a known application. Persona skills are closer to identity derivatives. They are built from many small observations, reused across contexts, and valuable precisely because they compress behavior into something portable. Deleting one source record may not delete the inference. Revoking one app may not revoke the pattern if it has already been distilled elsewhere.&lt;/p&gt;

&lt;p&gt;There is a market-structure angle here. The companies with the broadest distribution will have the best raw material for persona skills. Email, documents, browser sessions, code repos, messages, calendars, support tickets, payment flows. Each surface adds a little more signal. The convenience story points toward bundling, because a persona layer is more useful when it spans tools. The risk story points toward separation, because a compromised or over-permissive persona layer has a larger blast radius.&lt;/p&gt;

&lt;p&gt;This is one reason open versus closed agents is not only a model-quality debate. A closed suite can enforce a consistent identity and permission model, at least in theory. It can also make the persona layer harder to inspect or move. An open ecosystem can give users more control and portability, but portability cuts both ways. If the artifact travels, so does the risk. The winning design is not obvious. It depends on who bears the cost when an agent acts too much like its owner in the wrong place.&lt;/p&gt;

&lt;p&gt;The early web offers a useful contrast. Cookies were not invented as a grand surveillance architecture. They solved a state problem. Over time, the market discovered that state was also targeting, attribution, and power. Persona skills solve a context problem. It would be odd if the market did not also discover that context is retention, switching cost, fraud surface, and bargaining power.&lt;/p&gt;

&lt;p&gt;The outside evidence is already pointing in that direction. DataDome reported 7.9 billion AI agent requests across its network in January and February 2026, with trusted agent names being spoofed at scale. Recorded Future argues that agentic AI will push identity and access management toward agent identity governance, because agents need access to cloud applications and internal systems to be useful. These are not the same problem as persona-skill leakage, but they rhyme. Once software can act, identity stops being a login screen and becomes an operating perimeter.&lt;/p&gt;

&lt;p&gt;For users, the practical lesson is not to reject personalization. That would be like refusing autocomplete because keyboards can leak. The lesson is to price the option correctly. A persona skill that helps an agent choose better sources for a blog draft is one thing. A persona skill that can send mail, approve invoices, access client records, and imitate writing style is another. The second one should come with narrower scopes, short-lived credentials, audit logs, and boring revocation paths. Boring is good here.&lt;/p&gt;

&lt;p&gt;For vendors, the hard problem is incentive-compatible restraint. The most profitable persona layer is persistent, cross-product, and hard to leave. The safest persona layer is scoped, legible, and easy to kill. Users will not read a forty-page control panel. Enterprises will not accept a black box that turns employee behavior into a transferable credential. Regulators will care once the first visible incident joins impersonation with material loss.&lt;/p&gt;

&lt;p&gt;My prior is that persona agents become real infrastructure only after the industry stops treating memory as a feature and starts treating it as balance-sheet risk. The asset is useful because it concentrates personal signal. The liability is the same sentence.&lt;/p&gt;

&lt;p&gt;The distribution to watch is not benchmark accuracy alone. It is where the tail losses sit. If the vendor captures the retention upside while the user or employer eats the impersonation downside, the equilibrium will be too much memory in too many places. If liability, insurance, and procurement force the vendor to carry more of the tail, persona systems will look less magical and more like financial plumbing. Slower. More permissions. More logs. Fewer demos that pretend context is free.&lt;/p&gt;

&lt;p&gt;That would be a healthy trade. Agents should learn enough to be useful. They should not quietly become portable claims on who you are.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>privacy</category>
      <category>economics</category>
    </item>
  </channel>
</rss>
