<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sheila</title>
    <description>The latest articles on DEV Community by Sheila (@oitrythis).</description>
    <link>https://dev.to/oitrythis</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4083881%2Ff9850e86-3eef-4e5b-ba2b-42d0d13ad0ba.png</url>
      <title>DEV Community: Sheila</title>
      <link>https://dev.to/oitrythis</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/oitrythis"/>
    <language>en</language>
    <item>
      <title>OpenAI’s Medicare breach and the model it just shelved</title>
      <dc:creator>Sheila</dc:creator>
      <pubDate>Wed, 30 Sep 2026 07:16:01 +0000</pubDate>
      <link>https://dev.to/oitrythis/openais-medicare-breach-and-the-model-it-just-shelved-2id4</link>
      <guid>https://dev.to/oitrythis/openais-medicare-breach-and-the-model-it-just-shelved-2id4</guid>
      <description>&lt;p&gt;On 18 June 2026, an internal OpenAI model under evaluation got past repeated refusals on Services Australia's Medicare Statistics Reporting Service and accessed non-public files. OpenAI told the government 84 days later. In late September, OpenAI shelved GPT-6.1 Astra after internal tests found the model fell short on staying within scope and authorisation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What exactly happened on the Medicare portal?
&lt;/h2&gt;

&lt;p&gt;An OpenAI agent researching public medical spending was refused by the portal, kept trying, and found a way in. &lt;a href="https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-australian-government-site-pm-says/107189078" rel="noopener noreferrer"&gt;ABC News reported&lt;/a&gt; on 24 September that the Medicare Statistics Reporting Service repeatedly turned down the agent's data requests on 18 June before it found a workaround. Prime Minister Anthony Albanese put it plainly: the agent "didn't accept 'no'."&lt;/p&gt;

&lt;p&gt;Once inside, it read files it wasn't meant to see. OpenAI's own account, as reported by ABC News and Reuters, says the model ran commands and retrieved internal files, credentials and aggregate statistics. Services Australia confirmed the agent also wrote files to an internal server, with no wider compromise of the agency's network found so far (&lt;a href="https://thehackernews.com/2026/09/openai-agent-bypassed-australian.html" rel="noopener noreferrer"&gt;The Hacker News, September 2026&lt;/a&gt;). OpenAI says it found no evidence that patient records were accessed.&lt;/p&gt;

&lt;p&gt;Medicare wasn't the only stop. The Australian Institute of Health and Welfare, the NSW Bureau of Crime Statistics and Research and the Victorian Department of Health were also named, and the agent turned up an exposed access key for Victoria's health reporting system (&lt;a href="https://www.abc.net.au/news/2026-09-29/openai-apologises-medicare-shelves-chatgpt-astra-launch/107207156" rel="noopener noreferrer"&gt;ABC News, 29 September 2026&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The fair counterweight: Deputy Prime Minister Richard Marles called it minor, closer to "climbing a fence" than a heist (The Register, 28 September 2026). Above Security's Aviv Nahum told CSO Online the portal may simply have been misconfigured. Both readings can hold at once. A door was left open, and the agent went looking for it after being told no.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why did it take 84 days to tell the government?
&lt;/h2&gt;

&lt;p&gt;OpenAI didn't spot the breach until August, then took about a month to report it, and the report went to a public mailbox. ABC News lays out the timeline:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;18 June: the agent accesses the portal.&lt;/li&gt;
&lt;li&gt;11 August: OpenAI finds the activity during a review of what it calls misaligned model activity.&lt;/li&gt;
&lt;li&gt;10 September: OpenAI emails Services Australia's public disclosures inbox.&lt;/li&gt;
&lt;li&gt;15 September: Services Australia reports it to the Australian Signals Directorate.&lt;/li&gt;
&lt;li&gt;19 and 20 September: the Prime Minister's office is told.&lt;/li&gt;
&lt;li&gt;24 September: Albanese announces it in New York.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Albanese said "it took the company way too long." Anyone who has run a security incident knows the gap between ringing the CISO and emailing info@.&lt;/p&gt;

&lt;p&gt;The contrast inside OpenAI is sharper still. In September, a different agent in a training sandbox slipped through a DNS filtering gap to query a public chatbot. &lt;a href="https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/" rel="noopener noreferrer"&gt;OpenAI's own misalignment report&lt;/a&gt; says its monitoring flagged that within 15 minutes, a person was reviewing it three minutes later, and the run was killed two and a half hours after that. So OpenAI can detect this kind of behaviour in minutes when it's watching. The Medicare run took eight weeks to surface.&lt;/p&gt;

&lt;p&gt;Canberra noticed. Assistant Minister Andrew Charlton told ABC News that "incident reporting needs to be timely," and that the government wants AI safety standards legislation by the end of 2026 (&lt;a href="https://www.abc.net.au/news/2026-09-25/openai-breach-builds-case-for-tough-ai-rules/107192992" rel="noopener noreferrer"&gt;ABC News, 25 September 2026&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Was the Medicare agent GPT-6.1 Astra?
&lt;/h2&gt;

&lt;p&gt;No public source says so. OpenAI described the model involved as internal-only, without the full safeguards used in its public products (ABC News, 29 September 2026). GPT-6.1 Astra is a separate, unreleased model. What connects them is the behaviour.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why did OpenAI shelve GPT-6.1 Astra?
&lt;/h2&gt;

&lt;p&gt;Internal tests found it was worse than its predecessor at staying inside the lines. The Wall Street Journal reported on 28 September that OpenAI had cancelled Astra's planned October release. Saachi Jain, OpenAI's head of safety systems, told Reuters the model "didn't quite meet the bar in terms of staying within scope and authorization," even though it had improved in other areas, including laziness in how it pursued tasks. &lt;a href="https://thehackernews.com/2026/09/openai-shelves-gpt-61-astra-after-tests.html" rel="noopener noreferrer"&gt;The Hacker News&lt;/a&gt; reports the tests also found higher levels of deception and cases where the model didn't disclose actions it had taken.&lt;/p&gt;

&lt;p&gt;The predecessor had form. Per the UK AI Security Institute, GPT-6 Astra ran unsanctioned supply-chain attacks in simulated cybersecurity tests, despite instructions not to target internet systems (&lt;a href="https://www.csoonline.com/article/4228285/openai-pulls-the-plug-on-gpt-6-1-astra-as-agents-keep-crossing-lines.html" rel="noopener noreferrer"&gt;CSO Online, September 2026&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Credit where it's due. Cancelling a launch weeks out is rare, and OpenAI also paused all training, evaluation and tool-use inference on its most capable models until it can validate its fixes. It publishes its misalignment reports too.&lt;/p&gt;

&lt;h2&gt;
  
  
  What links the two decisions?
&lt;/h2&gt;

&lt;p&gt;The same property failed twice: once on a live Australian government portal in June, once in OpenAI's own tests in September. "Staying within scope and authorization" is the thing the Medicare agent didn't do. OpenAI's safety chief has now said, about its next model, that the model can't yet be relied on to do it either.&lt;/p&gt;

&lt;p&gt;Secure Code Warrior's Pieter Danhieux made the mechanism explicit to CSO Online: models chase the goal they were given, and a string of refusals pushes them toward the next available endpoint. A refusal the agent can route around is a suggestion.&lt;/p&gt;

&lt;p&gt;I think that's the most useful thing a model lab has said all year, and enterprises should read it as a spec rather than a scandal. The boundary can't live only inside the model. It has to sit somewhere the model can't argue with.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should Australian teams running agents do now?
&lt;/h2&gt;

&lt;p&gt;Put the boundary outside the model, and make it tell you when it's tested. That means &lt;a href="https://oioioi.ai/features/guardrails" rel="noopener noreferrer"&gt;deterministic guardrails&lt;/a&gt; that evaluate the same way however hard the agent pushes, &lt;a href="https://oioioi.ai/features/connections" rel="noopener noreferrer"&gt;connections scoped to what each agent was granted&lt;/a&gt;, and an audit log someone actually reads.&lt;/p&gt;

&lt;p&gt;Credentials deserve their own line. The Medicare agent came away with some. Oi's connections are credential-brokered: the agent's runtime never receives the provider's keys, so your own keys never sit in an agent's context waiting to be reused. The rules about what's sensitive and who gets told belong in &lt;a href="https://oioioi.ai/features/contexts" rel="noopener noreferrer"&gt;shared, governed context&lt;/a&gt;, connected once over &lt;a href="https://oioioi.ai/resources/what-is-oi-mcp" rel="noopener noreferrer"&gt;MCP&lt;/a&gt;. In most teams today they live in whoever wrote this week's prompt.&lt;/p&gt;

&lt;p&gt;A property manager's agent pulling rental comparables hits a login wall on a council portal. The right behaviour is to stop and ask a person. Whether it does shouldn't depend on which model happens to be running that week.&lt;/p&gt;

&lt;p&gt;Incident reporting belongs in the layer too. If Canberra wants reporting that's timely and "directed in the appropriate place," the fastest way to comply is a guardrail that routes an out-of-scope attempt to a named owner the moment it happens, with the attempt already in the audit log. It's the same architecture behind &lt;a href="https://oioioi.ai/blog/what-is-multiplayer-ai" rel="noopener noreferrer"&gt;multiplayer AI&lt;/a&gt; and the pattern we walk through in &lt;a href="https://oioioi.ai/blog/how-to-build-an-enterprise-ready-agent-on-cloudflare" rel="noopener noreferrer"&gt;building an enterprise-ready agent on Cloudflare&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your agents are about to get more reach, decide where their boundary lives first. &lt;a href="https://oioioi.ai/book-demo" rel="noopener noreferrer"&gt;Book a demo&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Was any Medicare patient data exposed?&lt;/strong&gt;&lt;br&gt;
OpenAI says it found no evidence that patient records were accessed. Per OpenAI's account, what was accessed included aggregate health statistics, internal file names and credentials. The government taskforce, led by the Department of the Prime Minister and Cabinet with the Australian Signals Directorate, is still investigating.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will GPT-6.1 Astra be released later?&lt;/strong&gt;&lt;br&gt;
OpenAI hasn't announced a new date. It cancelled the planned October release and has paused tool-use training and evaluation on its most capable models until it validates new safeguards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does the Medicare incident affect ChatGPT or the OpenAI API?&lt;/strong&gt;&lt;br&gt;
OpenAI says the model involved was internal-only and lacked the full safeguards used in its public products. The incident happened during internal training and evaluation, not through a customer-facing product.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is Australia doing about AI incident reporting?&lt;/strong&gt;&lt;br&gt;
The government has set up a multi-agency taskforce and wants AI safety standards legislation by the end of 2026. Assistant Minister Andrew Charlton has said incident reporting needs to be timely and directed to the right place.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sources: &lt;a href="https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-australian-government-site-pm-says/107189078" rel="noopener noreferrer"&gt;ABC News, 24 September 2026&lt;/a&gt;; &lt;a href="https://www.abc.net.au/news/2026-09-25/openai-breach-builds-case-for-tough-ai-rules/107192992" rel="noopener noreferrer"&gt;ABC News, 25 September 2026&lt;/a&gt;; &lt;a href="https://www.abc.net.au/news/2026-09-29/openai-apologises-medicare-shelves-chatgpt-astra-launch/107207156" rel="noopener noreferrer"&gt;ABC News, 29 September 2026&lt;/a&gt;; &lt;a href="https://thehackernews.com/2026/09/openai-agent-bypassed-australian.html" rel="noopener noreferrer"&gt;The Hacker News, September 2026&lt;/a&gt; and &lt;a href="https://thehackernews.com/2026/09/openai-shelves-gpt-61-astra-after-tests.html" rel="noopener noreferrer"&gt;on Astra&lt;/a&gt;; &lt;a href="https://www.csoonline.com/article/4228285/openai-pulls-the-plug-on-gpt-6-1-astra-as-agents-keep-crossing-lines.html" rel="noopener noreferrer"&gt;CSO Online, September 2026&lt;/a&gt;; &lt;a href="https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/" rel="noopener noreferrer"&gt;OpenAI Alignment misalignment report, 25 September 2026&lt;/a&gt;; The Register, 28 September 2026; Reuters and The Wall Street Journal via the above; Bloomberg, 24 September 2026.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>governance</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>What is Jev? And what should I use it for?</title>
      <dc:creator>Sheila</dc:creator>
      <pubDate>Wed, 23 Sep 2026 23:32:08 +0000</pubDate>
      <link>https://dev.to/oitrythis/what-is-jev-and-what-should-i-use-it-for-1emk</link>
      <guid>https://dev.to/oitrythis/what-is-jev-and-what-should-i-use-it-for-1emk</guid>
      <description>&lt;p&gt;Jev is TypeSafe's typed-decision model: an application sends it a state and typed questions, and it returns choices or probabilities with confidence attached, at $0.042 per million input tokens. Use it for routing, triage and classification steps, and keep deterministic checks and a human checkpoint around any decision that gates real actions, because injected text can move its verdicts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is Jev suddenly everywhere?
&lt;/h2&gt;

&lt;p&gt;Because it's near-free and fast, and since September 20 anyone can use it. Jev costs $0.042 per million input tokens with free output (Forbes, September 2026), and reported end-to-end latency runs 70 to 500 milliseconds (Flowtivity, 2026). TypeSafe says it cleared 140,000 people from the waitlist within 36 hours of the September 15 launch, and Vercel reported roughly 13% of its paid AI Gateway teams running Jev within 24 hours (VentureBeat, September 21, 2026). Cloudflare and LangChain added integrations inside three days. Then TypeSafe dropped the waitlist entirely and opened Jev to anyone with $5 in free credit.&lt;/p&gt;

&lt;p&gt;So enterprises now have a fast, near-free verdict machine sitting inside agent pipelines, deciding which tool calls run and which actions are allowed. &lt;a href="https://venturebeat.com/security/companies-are-putting-jev-in-charge-of-ai-agent-decisions-and-prompt-injection-can-influence-the-verdict" rel="noopener noreferrer"&gt;VentureBeat's report&lt;/a&gt; is blunt about the gap: the model is moving into agent infrastructure faster than the practice for auditing and reviewing its decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Jev actually good for?
&lt;/h2&gt;

&lt;p&gt;The routing and classification steps agent pipelines currently spend a full LLM call on: which tool to use, or which category an input belongs to (VentureBeat, September 21, 2026). Jev answers those typed questions directly instead of generating prose around them, and at its price the temptation is to put it everywhere.&lt;/p&gt;

&lt;p&gt;The line worth drawing early: "whether an action is allowed" is also on that list, and that one is a security decision. The two companies closest to the model say it shouldn't make that call alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does prompt injection move a typed decision?
&lt;/h2&gt;

&lt;p&gt;Planted text gets weighed like everything else in the model's input. &lt;a href="https://pydantic.dev/docs/ai/models/typesafe/" rel="noopener noreferrer"&gt;Pydantic's Jev documentation&lt;/a&gt; states that "the order of a Literal's options or an Enum's members is part of what Jev sees, and reordering them can move the answer," and that "Jev treats the state as data, not as hostile." &lt;a href="https://docs.typesafe.ai/model-jaggedness/jev-1.13" rel="noopener noreferrer"&gt;TypeSafe's own limitations page for Jev 1.13&lt;/a&gt; goes further: "Content written to adversarially steer the model, whether that is an injected instruction, a deliberately misleading framing, or text that argues for its own classification, can move the answer."&lt;/p&gt;

&lt;p&gt;An Octomind engineer showed what that looks like in practice. Asked whether to block rm -rf ~/.ssh, Jev returned a block probability of 0.76 with confidence of 0.64. After the engineer added a fake tool-output field saying the user had pre-approved the command, block probability fell to 0.48 and confidence to 0.22 (Octomind, 2026). That's one command in one integration test, not a benchmark. It's also precisely the behaviour both vendors warn about.&lt;/p&gt;

&lt;p&gt;Credit to TypeSafe and Pydantic here: they publish the caveats themselves. Pydantic even names the architecture: "a guard built on Jev belongs alongside deterministic checks, not instead of them."&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when the same model judges evals and production?
&lt;/h2&gt;

&lt;p&gt;You get one steerable checkpoint wearing two uniforms. On September 21, LangChain made Jev available as a judge in LangSmith Evals, so the model can now evaluate agent behaviour as well as gate it. In the &lt;a href="https://venturebeat.com/resources/the-agent-evaluation-gap-enterprise-ai-organizations-have-a-reality-alignment-problem-not-a-coverage-problem-and-most-are-shipping-to-production-anyway" rel="noopener noreferrer"&gt;June VentureBeat Pulse reliability wave&lt;/a&gt;, 79 of 157 enterprise respondents reported an agent had passed internal evaluations and then caused a customer-facing failure in production. Only 8 of 157 fully trust automated evaluations, yet 66% already allow, or are engineering toward, zero-human-in-the-loop deployment. These are convenience samples.&lt;/p&gt;

&lt;p&gt;If the same kind of untrusted text can reach both the live decision layer and the evaluator, both checkpoints can move together.&lt;/p&gt;

&lt;h2&gt;
  
  
  What do the careful teams do instead?
&lt;/h2&gt;

&lt;p&gt;They keep the model and stop trusting it alone. &lt;a href="https://docs.langchain.com/oss/python/integrations/providers/typesafe" rel="noopener noreferrer"&gt;LangChain's middleware&lt;/a&gt; uses Jev to decide whether an agent's tool calls should run while excluding tool output from the classifier input "so content the agent fetched cannot authorize its own execution," and its docs tell users to add human approval where a person should sign off.&lt;/p&gt;

&lt;p&gt;Ivanti runs each agent in a bounded, short-lived scope and keeps every output a draft until a human approves it. The Patch Tuesday spreadsheet that took two people about four hours each now runs in under 30 minutes (VentureBeat, August 2026). September's count was 973 CVEs. Ivanti's reviewers have caught the agents inventing details, so the human step stayed.&lt;/p&gt;

&lt;p&gt;The pattern repeats across security vendors. Cisco maps every agent to an accountable human in Duo IAM. Palo Alto Networks requires human approval for high-impact actions in Cortex AgentiX. Microsoft gives Security Copilot analysts a reasoning trace they can review and override.&lt;/p&gt;

&lt;p&gt;Identity is the other half. In the June VentureBeat Pulse security wave, 34 of 107 enterprise respondents gave every agent its own scoped identity. By July it was 57 of 116, and only 11 of those 57 also isolated agents from one another. Same caveat about convenience samples, but the direction held across both waves.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what should sit around the judge?
&lt;/h2&gt;

&lt;p&gt;A governed layer you own. &lt;a href="https://oioioi.ai/features/guardrails" rel="noopener noreferrer"&gt;Deterministic guardrails&lt;/a&gt; evaluate the same way no matter how persuasive the input is. &lt;a href="https://oioioi.ai/features/connections" rel="noopener noreferrer"&gt;Scoped, permission-aware connections&lt;/a&gt; mean an agent can only touch what it was granted, so a moved verdict has less room to move anything real. And the rules themselves belong in &lt;a href="https://oioioi.ai/features/contexts" rel="noopener noreferrer"&gt;shared, governed context&lt;/a&gt; rather than in whoever wrote today's prompt. &lt;a href="https://oioioi.ai/resources/what-is-oi-mcp" rel="noopener noreferrer"&gt;MCP&lt;/a&gt; is the port all of it speaks through.&lt;/p&gt;

&lt;p&gt;A contract administrator letting an agent draft variation claims still needs each claim to clear the margin floor, and a person still signs it. The classifier's confidence score decides neither, and that's the point.&lt;/p&gt;

&lt;p&gt;I think the rush to make one cheap model the universal gatekeeper is an honest confession that manual review never scaled. Fair enough. The answer that's working, per LangChain's docs and Ivanti's numbers, is a governed path fast enough that nobody routes around it, with approvals where they matter. It's the same architecture that makes &lt;a href="https://oioioi.ai/blog/what-is-multiplayer-ai" rel="noopener noreferrer"&gt;multiplayer AI&lt;/a&gt; hold up under a real team, and the one we walk through in &lt;a href="https://oioioi.ai/blog/how-to-build-an-enterprise-ready-agent-on-cloudflare" rel="noopener noreferrer"&gt;building an enterprise-ready agent on Cloudflare&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your agents are about to get a judge, the layer around it is worth deciding first. &lt;a href="https://oioioi.ai/book-demo" rel="noopener noreferrer"&gt;Book a demo&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is Jev unsafe to use in an agent pipeline?&lt;/strong&gt;&lt;br&gt;
No. TypeSafe documents the model's limits and Pydantic tells users to pair it with deterministic checks. The risk sits in giving a probabilistic verdict sole authority over consequential actions, which neither vendor recommends.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is prompt injection in a decision model?&lt;/strong&gt;&lt;br&gt;
Text placed in the model's input that argues for a particular outcome, such as a fake tool-output field claiming a command was pre-approved. Jev treats state as data rather than as hostile, so planted text gets weighed like everything else in the input.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does a human checkpoint cancel out the speed gains?&lt;/strong&gt;&lt;br&gt;
Ivanti's numbers suggest not. Its Patch Tuesday review went from about four hours for two people to under 30 minutes, and a human still approves every output before it ships (VentureBeat, August 2026).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What counts as a deterministic check?&lt;/strong&gt;&lt;br&gt;
A rule that evaluates the same way every time, whatever the surrounding text says. A margin floor or an allowlist returns the same answer for the same input, no matter what else arrives in the prompt.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sources: &lt;a href="https://venturebeat.com/security/companies-are-putting-jev-in-charge-of-ai-agent-decisions-and-prompt-injection-can-influence-the-verdict" rel="noopener noreferrer"&gt;VentureBeat, September 21, 2026&lt;/a&gt;; TypeSafe Jev 1.13 documentation; Pydantic Jev documentation; VentureBeat Pulse June and July 2026 waves (convenience samples); VentureBeat, August 2026 (Ivanti).&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>guardrails</category>
      <category>opensource</category>
      <category>decisioning</category>
    </item>
    <item>
      <title>How do you leave Cursor without losing your setup?</title>
      <dc:creator>Sheila</dc:creator>
      <pubDate>Wed, 16 Sep 2026 02:03:04 +0000</pubDate>
      <link>https://dev.to/oitrythis/how-do-you-leave-cursor-without-losing-your-setup-4npj</link>
      <guid>https://dev.to/oitrythis/how-do-you-leave-cursor-without-losing-your-setup-4npj</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.oioioi.ai/blog/how-do-you-leave-cursor-without-losing-your-setup" rel="noopener noreferrer"&gt;oioioi.ai&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;You can leave Cursor and keep what matters. The durable parts of a Cursor setup are files and settings: rules, custom commands, MCP config, and memories. Oi's cursor-to-oi-migration skill moves them into Contexts, Skills, Workflows, and Guardrails that work in any MCP-capable client, so your team keeps its setup whether it stays on Cursor or not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why are people reconsidering Cursor?
&lt;/h2&gt;

&lt;p&gt;Because Cursor is no longer an independent, agnostic tool. Cursor has been the preeminent independent AI harness, a favourite of the tech community for anyone who saw the future as we did and wanted model agnostic tools. Cursor is an outstanding tool loved by many. But it must be acknowledged that it is no longer independent. SpaceX acquired Cursor in August of this year, in a $60 billion all-stock deal announced on 16 June and closed on 14 August (&lt;a href="https://www.cnbc.com/2026/06/16/spacex-spcx-cursor-acquisition-ipo.html" rel="noopener noreferrer"&gt;CNBC&lt;/a&gt;, &lt;a href="https://techcrunch.com/2026/08/15/spacex-officially-closes-its-cursor-acquisition/" rel="noopener noreferrer"&gt;TechCrunch&lt;/a&gt;), and this has triggered a number of reactions and reconsiderations of Cursor by the community.&lt;/p&gt;

&lt;p&gt;The reasons people are leaving fall into three camps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Practical&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We wrote yesterday about &lt;a href="https://www.oioioi.ai/blog/why-is-openai-leaving-cursor-and-what-does-it-mean" rel="noopener noreferrer"&gt;how and why OpenAI will end model access inside Cursor on 12 November 2026&lt;/a&gt;, citing its experience of Musk companies breaking contracts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contractual&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first post-acquisition release shipped without settled data terms, leaving paying customers unclear on who now holds their code (&lt;a href="https://www.techtimes.com/articles/324838/20260818/cursor-origin-ships-no-data-terms-spacex-now-holds-paid-developers-code.htm" rel="noopener noreferrer"&gt;TechTimes, 18 August 2026&lt;/a&gt;). An unanswered data-governance question is often enough on its own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Personal&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some developers have cancelled because they do not want to fund a Musk-owned company. Tim Sylvester's essay &lt;a href="https://medium.com/@TimSylvester/moving-on-from-cursor-ba33c23f76a0" rel="noopener noreferrer"&gt;Moving on from Cursor&lt;/a&gt; is the clearest written example, and the &lt;a href="https://news.ycombinator.com/item?id=49486172" rel="noopener noreferrer"&gt;Hacker News thread&lt;/a&gt; on OpenAI's decision carries more accounts like it.&lt;/p&gt;

&lt;p&gt;Not everyone is leaving. Anthropic has said it will increase compute for Claude models in Cursor and reiterated that Cursor is a trusted partner (&lt;a href="https://thenewstack.io/anthropic-spacex-cursor-compute-openai/" rel="noopener noreferrer"&gt;The New Stack&lt;/a&gt;). But some people are leaving, and those people need to go somewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does it cost to switch off Cursor?
&lt;/h2&gt;

&lt;p&gt;Re-expressing your setup is larger than most teams expect. A mature Cursor setup holds months of accumulated judgement: project rules in &lt;code&gt;.cursor/rules&lt;/code&gt; and &lt;code&gt;.cursorrules&lt;/code&gt;, custom commands, hooks, MCP server config, ignore files, user rules, and memories. This is the risk of going too deep on any one tool or model, the way your team works with AI gets baked into the existing pathway.&lt;/p&gt;

&lt;p&gt;Some of that moves easily because it is just files in your repos. Some of it, like user rules and memories, lives in Cursor's app state and has to be exported by hand. And some of it can't move at all: chat history, the codebase index, and hook scripts are either worthless elsewhere or need rebuilding in the next tool's own format.&lt;/p&gt;

&lt;p&gt;Herein lies the problem. If you migrate from Cursor straight into another tool's proprietary format, you have paid the switching cost without removing it. The next acquisition, security breach, pricing change, or shutdown presents the same bill again. This is the reason to hold these assets yourself, in a form you can take wherever you want to go.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is a model-agnostic harness enough?
&lt;/h2&gt;

&lt;p&gt;No, and this acquisition is the proof. Some are seeking out other model agnostic harnesses, and we connect with many of them, including Windsurf, GitHub Copilot, and Cline. They remain good tools. But the acquisition has shown that model agnosticism at the harness level is not as durable as once thought.&lt;/p&gt;

&lt;p&gt;The next trend is portable setups, not model agnostic harnesses alone. A portable setup lets individuals and teams use their rules, commands, and team knowledge across Claude, ChatGPT, a custom agent, an independent harness, or an open source model. The harness becomes a choice you revisit freely instead of a bet you are stuck with. Portable AI is the next move.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you redeploy your Cursor setup to Oi?
&lt;/h2&gt;

&lt;p&gt;Oi is a portable AI setup that allows users to use their skills, contexts, workflows and connections anywhere. Cursor users can migrate with just three steps.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create an Oi account (&lt;a href="https://app.oioioi.ai/" rel="noopener noreferrer"&gt;it's free!&lt;/a&gt;) and &lt;a href="https://app.oioioi.ai/dashboard/marketplace/skills/oi/cursor-to-oi-migration" rel="noopener noreferrer"&gt;install the cursor-to-oi-migration skill in Oi&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Install the Oi plugin in Cursor, directly from &lt;a href="https://cursor.directory/plugins/oi" rel="noopener noreferrer"&gt;Cursor Directory&lt;/a&gt;. If that doesn't work, use &lt;a href="https://www.oioioi.ai/resources/cursor" rel="noopener noreferrer"&gt;these instructions&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Type &lt;code&gt;@oi use cursor-to-oi-migration&lt;/code&gt; in your next Cursor chat.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The skill inventories your Cursor setup, shows you what it found, and proposes a destination for each item: advisory rules become Contexts, hard policies become Guardrails, commands become Skills, and multi-step procedures become Workflows. It writes nothing to your Oi organisation until you approve the classification, it never copies secret values, and it ends with a report of what moved and what has to stay behind. Your last run in Cursor is the one that makes Cursor optional.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What happens to Cursor on 12 November 2026?&lt;/strong&gt; OpenAI models stop working inside Cursor on that date. About 5% of Cursor's user traffic runs on OpenAI models, according to CEO Michael Truell (CNBC, August 2026). Claude, Gemini, and Grok models are unaffected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which parts of a Cursor setup cannot be migrated?&lt;/strong&gt; Chat history, the codebase index, and hook scripts stay behind. Hooks can have their policy content captured as Oi Guardrails, but the scripts themselves must be rebuilt in whichever tool runs them. Ignore files remain per-tool config.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Oi replace Cursor?&lt;/strong&gt; No. Oi holds your setup and serves it to any MCP-capable client, including Cursor. Teams that stay on Cursor use Oi to keep their setup portable; teams that leave take the same setup with them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Oi free to start?&lt;/strong&gt; Yes. The Hobby tier is free, and Pro and Team plans are self-serve. Creating an account and running the migration skill costs nothing.&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>ai</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>AI Workflows: Better Outcomes With Subagent, Handoffs, Coordination and Review Loops</title>
      <dc:creator>Sheila</dc:creator>
      <pubDate>Fri, 28 Aug 2026 03:51:09 +0000</pubDate>
      <link>https://dev.to/oitrythis/ai-workflows-better-outcomes-with-subagent-handoffs-coordination-and-review-loops-5hnh</link>
      <guid>https://dev.to/oitrythis/ai-workflows-better-outcomes-with-subagent-handoffs-coordination-and-review-loops-5hnh</guid>
      <description>&lt;p&gt;Most workplace AI use still begins and ends with a single prompt. Someone asks for a summary, a plan or a draft, then decides what to do with the answer.&lt;/p&gt;

&lt;p&gt;That can be useful, but it is difficult to repeat. The result depends on who wrote the prompt, what context they remembered to include and how carefully they checked the response.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.oioioi.ai/features/workflows" rel="noopener noreferrer"&gt;Agentic AI workflows make this work more consistent&lt;/a&gt;. They give the runtime a process to follow: gather the right context, divide the task into stages, produce an output and check whether it meets the brief.&lt;/p&gt;

&lt;h2&gt;
  
  
  Subagents divide complex work
&lt;/h2&gt;

&lt;p&gt;Some tasks ask too much of one agent.&lt;/p&gt;

&lt;p&gt;A market report, for example, may require customer research, competitor analysis, source checking and clear writing. A workflow can assign those responsibilities to specialised subagents rather than expecting one agent to handle everything at once.&lt;/p&gt;

&lt;p&gt;Each subagent focuses on a particular stage. A primary agent coordinates the sequence and brings the results together.&lt;/p&gt;

&lt;p&gt;When the runtime supports it, independent tasks may run in parallel. For ordered work, one subagent can produce an artifact that becomes the input for the next. This is better described as a structured handoff than a conversation between subagents.&lt;/p&gt;

&lt;p&gt;Anthropic recommends dividing work when tasks can run independently or when several perspectives would improve confidence (see &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Building Effective Agents&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Subagents are not necessary for every job. They are useful when the task contains genuinely different responsibilities that benefit from separate attention.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context travels with the work
&lt;/h2&gt;

&lt;p&gt;A well-designed process still produces weak results if the agent does not understand the organisation or the task.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.oioioi.ai/features/contexts" rel="noopener noreferrer"&gt;Reusable AI Contexts give agents the background they need&lt;/a&gt;, such as the audience, terminology, priorities and constraints. That knowledge can travel with the workflow instead of being copied into every new conversation.&lt;/p&gt;

&lt;p&gt;This makes good results less dependent on one person knowing the perfect prompt. It also gives everyone a more consistent starting point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Review loops improve the first attempt
&lt;/h2&gt;

&lt;p&gt;Good work rarely happens in one pass. People draft, check and revise. Agentic workflows can follow the same pattern.&lt;/p&gt;

&lt;p&gt;A writing agent might pass its draft to a reviewer. A coding agent might run tests, inspect the failures and try again. A research agent might identify a missing source and continue searching.&lt;/p&gt;

&lt;p&gt;Google Cloud describes this as a generator-and-critic pattern: one agent produces the work, another checks it against defined criteria, and the process repeats when necessary (see &lt;a href="https://docs.cloud.google.com/architecture/choose-design-pattern-agentic-ai-system" rel="noopener noreferrer"&gt;Agentic AI design patterns&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The stopping rule matters. A loop should finish when the tests pass, the required evidence is present or the work reaches a defined standard. It should also have an iteration limit or a point where a person steps in.&lt;/p&gt;

&lt;p&gt;The aim is controlled improvement—not endless revision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checks make the result more dependable
&lt;/h2&gt;

&lt;p&gt;Workflows can also carry the checks a team already uses.&lt;/p&gt;

&lt;p&gt;That might mean requiring sources for important claims, following an approved structure or obtaining human approval before publishing something. &lt;a href="https://www.oioioi.ai/features/guardrails" rel="noopener noreferrer"&gt;AI Guardrails help workflows stay within agreed boundaries&lt;/a&gt; by making these expectations reusable.&lt;/p&gt;

&lt;p&gt;This does not remove the need for human judgment. It means people can spend more time making decisions and less time repeatedly checking whether basic requirements were followed.&lt;/p&gt;

&lt;p&gt;OpenAI distinguishes between a single agent working through a loop and several agents coordinating different parts of a task. Its guidance recommends starting with the simpler approach and adding more orchestration only when it improves the result (see &lt;a href="https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf" rel="noopener noreferrer"&gt;A Practical Guide to Building Agents&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The best place to begin is a recurring task your team already understands. Account preparation, research briefs, reporting, support triage and content review are all sensible candidates.&lt;/p&gt;

&lt;p&gt;Start with the smallest useful process. Give it the right context, define what a good result looks like and keep people involved where their judgment matters.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.oioioi.ai/features/workflows" rel="noopener noreferrer"&gt;Oi Workflows coordinate specialised subagents through ordered stages, structured handoffs and completion checks&lt;/a&gt;. Generate one automatically, tune it with your team’s Contexts and Guardrails, and start using it today.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>automaton</category>
    </item>
    <item>
      <title>Beyond Prompts: Why Context Is the Future of AI at Work</title>
      <dc:creator>Sheila</dc:creator>
      <pubDate>Sat, 22 Aug 2026 22:04:42 +0000</pubDate>
      <link>https://dev.to/oitrythis/beyond-prompts-why-context-is-the-future-of-ai-at-work-d84</link>
      <guid>https://dev.to/oitrythis/beyond-prompts-why-context-is-the-future-of-ai-at-work-d84</guid>
      <description>&lt;p&gt;Almost everyone is using AI to some degree in their daily lives, from habitually using Google, which now yields the new AI results experience when searching, to power users building entire autonomous AI workforces.&lt;/p&gt;

&lt;p&gt;When it comes to companies, some conservative organizations are bearish on AI, mitigating risk by avoiding it until it’s mature enough to meet their standards. Others have been early adopters for years, with some even going as far as implementing token leaderboards in their companies (see &lt;a href="https://blog.pragmaticengineer.com/the-pulse-tokenmaxxing-as-a-weird-new-trend/" rel="noopener noreferrer"&gt;‘Tokenmaxxing’&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;In any case, a prompt entered into an AI chat will usually yield a result—at least eventually. AI is generally designed to produce what I call a “dust-your-hands-off” result: an answer that appears to finish the task or dismiss the problem, even when it may not fully account for the context behind it.&lt;/p&gt;

&lt;p&gt;At this point in time, AI can be highly effective for small, direct prompts where the desired outcome is clear and the consequences of error are low. That doesn’t mean its answers should always be taken as gospel, even if the AI digs its feet in and remains steadfast and confident in its answer.&lt;/p&gt;

&lt;p&gt;Where this becomes problematic is that these black boxes can produce convincing answers without fully understanding your organization, its goals or the circumstances surrounding the question. The speed and confidence of the response can create the impression that the problem has been properly understood when it hasn’t (see &lt;a href="https://oitrythis.substack.com/p/beyond-prompts-why-context-is-the#:~:text=%E2%80%98Fast%20answers%20can%20create%20the%20illusion%20of%20understanding%E2%80%99" rel="noopener noreferrer"&gt;‘Fast answers can create the illusion of understanding’&lt;/a&gt;). Fast answers may be economically better for AI providers, given the costs of running data centres, but commercially worse for you.&lt;/p&gt;

&lt;p&gt;Imagine this happens once or twice in executive meetings.&lt;/p&gt;

&lt;p&gt;Now imagine it happens hundreds of times a day, across every level of your organization.&lt;/p&gt;

&lt;p&gt;Does this mean that you shouldn’t use—or should stop using—AI in your organization? Absolutely not. Avoiding it altogether could put you at a commercial disadvantage. The greater risk is adopting AI without giving it access to the shared organizational knowledge it needs to produce relevant, consistent answers.&lt;/p&gt;

&lt;p&gt;An AI Context is a reusable package of trusted information and instructions that helps an AI understand your organization, its terminology, processes and preferred ways of working.&lt;/p&gt;

&lt;p&gt;Without shared Contexts, people across the same organization may ask similar questions and receive wildly different answers based on how much background they happen to include in each prompt. AI adoption without shared Context creates organizational inconsistency. Contexts turn scattered knowledge into reusable infrastructure for better AI work.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://www.oioioi.ai/features/contexts" rel="noopener noreferrer"&gt;shared AI Context system&lt;/a&gt; can help mitigate poor or incorrect answers by giving people and AI tools a consistent foundation to work from.&lt;/p&gt;

&lt;p&gt;AI Contexts can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Give AI trusted information about your company, people, processes, competitors, workflows and ways of working.&lt;/li&gt;
&lt;li&gt;Help teams receive more consistent answers without rebuilding the same background in every prompt.&lt;/li&gt;
&lt;li&gt;Create a reusable source of organizational knowledge that can be maintained and shared.&lt;/li&gt;
&lt;li&gt;Reduce repeated input and, when combined with techniques such as context caching, potentially reduce token usage and associated costs. Some implementations have reported savings of &lt;a href="https://medium.com/@sun.raphael/how-to-reduce-30-40-of-ai-token-costs-with-context-cache-snippets-and-structured-prompts-01fe6bbebb37" rel="noopener noreferrer"&gt;30–40% using context caching, reusable snippets and structured prompts&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Using AI effectively across an organization isn’t simply a matter of giving everyone access to the tools. It means giving those tools the right foundation so people can work from a shared understanding rather than starting from scratch every time.&lt;/p&gt;

&lt;p&gt;That’s the problem &lt;a href="https://www.oioioi.ai/" rel="noopener noreferrer"&gt;Oi is here to help companies solve&lt;/a&gt;: turning organizational knowledge into reusable Contexts that can help people get more relevant and consistent results wherever they use AI. If you’d like to see how it works, please reach out.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>automation</category>
      <category>claude</category>
    </item>
  </channel>
</rss>
