<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mangat Rai</title>
    <description>The latest articles on DEV Community by Mangat Rai (@mangatrai).</description>
    <link>https://dev.to/mangatrai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4111724%2Fb8ecae8a-89d6-45da-8dab-b9fc84f88b1d.png</url>
      <title>DEV Community: Mangat Rai</title>
      <link>https://dev.to/mangatrai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mangatrai"/>
    <language>en</language>
    <item>
      <title>The AI industry is building its own referee</title>
      <dc:creator>Mangat Rai</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:29:57 +0000</pubDate>
      <link>https://dev.to/mangatrai/the-ai-industry-is-building-its-own-referee-4go1</link>
      <guid>https://dev.to/mangatrai/the-ai-industry-is-building-its-own-referee-4go1</guid>
      <description>&lt;p&gt;OpenAI models &lt;a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/" rel="noopener noreferrer"&gt;escaped a test environment and compromised Hugging Face&lt;/a&gt;. Claude models &lt;a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" rel="noopener noreferrer"&gt;gained unauthorized access to real production systems&lt;/a&gt; in three cases. Gemini &lt;a href="https://therecord.media/gemini-google-cyber-breach" rel="noopener noreferrer"&gt;accessed systems at three companies&lt;/a&gt; during another test.&lt;/p&gt;

&lt;p&gt;Different labs. Same unsettling pattern. Sound familiar? That has been the rhythm of AI safety news lately.&lt;/p&gt;

&lt;p&gt;By the time I read the third story, I wanted to know whether anyone was connecting these incidents. Were the labs learning from them together, or was each company still deciding for itself what counted as safe? What were the frontier labs doing as an industry, what was the government requiring, and did any shared standard exist?&lt;/p&gt;

&lt;p&gt;So I started digging into those questions. That search led me to industry plan taking shape under the tentative name SAFA.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;⏳ &lt;strong&gt;TL;DR&lt;/strong&gt;&lt;br&gt;
I started with a simple question: after a run of frontier-model security incidents, who sets the shared safety rules? In July, Google DeepMind's Demis Hassabis proposed an industry-funded standards body with federal oversight, independent experts, and a path from voluntary reviews to mandatory pre-release tests. By late September, Google, OpenAI, and Anthropic were reportedly building a private version without government oversight. Shared tests could still improve today's fragmented safety reporting. The unresolved question is whether this new body can make its founders follow a rule, publish an embarrassing result, or admit competitors on equal terms.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The answer changed in ten weeks
&lt;/h2&gt;

&lt;p&gt;The first real answer I found came from Google DeepMind's Demis Hassabis. His &lt;a href="https://institute.deepmind.com/essays/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age/" rel="noopener noreferrer"&gt;July proposal&lt;/a&gt; called for a federally overseen standards body with independent technical experts and open-source representatives on its board. Industry would provide most of the funding, while federal agencies and US national laboratories would help test advanced models in areas tied to national security.&lt;/p&gt;

&lt;p&gt;The proposal also described a route to actual authority. Labs would initially submit frontier models for voluntary review up to 30 days before release. If the process proved effective, models could eventually be required to pass before deployment in the US.&lt;/p&gt;

&lt;p&gt;Then the idea changed. On September 24, The Information &lt;a href="https://www.theinformation.com/articles/google-openai-anthropic-ai-safety-group-takes-shape" rel="noopener noreferrer"&gt;reported&lt;/a&gt; that Google, OpenAI, and Anthropic were preparing their own standards organization, tentatively called the Standards Authority for Frontier AI (SAFA). The companies hope to launch it by the end of 2026 or early 2027, but its final name, charter, leadership, funding, and enforcement powers are not public.&lt;/p&gt;

&lt;p&gt;Here is the difference I kept coming back to:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;July proposal&lt;/th&gt;
&lt;th&gt;Reported September plan&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Who is behind it&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A US-led body, funded mostly by industry&lt;/td&gt;
&lt;td&gt;Google, OpenAI, and Anthropic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Government role&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Federal oversight, with agencies and national labs involved in testing&lt;/td&gt;
&lt;td&gt;No government oversight&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Independent technical experts and open-source representatives on the board&lt;/td&gt;
&lt;td&gt;Board and appointment rules not public&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Participation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Voluntary reviews first, with a path to mandatory tests&lt;/td&gt;
&lt;td&gt;Voluntary safety commitments; enforcement not public&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Main work&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Define frontier thresholds and assess models before release&lt;/td&gt;
&lt;td&gt;Support outside testing, incident reporting, and auditor qualifications&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This does not make SAFA pointless. It changes what SAFA is. The July proposal described delegated governance: industry expertise and money operating inside a public structure. The September plan is, so far, a private standards organization seeking credibility without delegated authority.&lt;/p&gt;

&lt;h2&gt;
  
  
  I can see why the labs want one standard
&lt;/h2&gt;

&lt;p&gt;Once I looked at what the labs do today, the practical case for SAFA became clearer. Each company already has a framework for deciding when a powerful model needs stronger safeguards. The frameworks cover similar severe risks, but they use different categories, thresholds, and names.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Lab&lt;/th&gt;
&lt;th&gt;Published framework&lt;/th&gt;
&lt;th&gt;Its trigger language&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.anthropic.com/responsible-scaling-policy" rel="noopener noreferrer"&gt;Responsible Scaling Policy&lt;/a&gt; (v3.4, July 2026)&lt;/td&gt;
&lt;td&gt;Capability Thresholds tied to ASL-3 safeguards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://openai.com/index/updating-our-preparedness-framework/" rel="noopener noreferrer"&gt;Preparedness Framework&lt;/a&gt; (v2, April 2025)&lt;/td&gt;
&lt;td&gt;High and Critical capability thresholds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google DeepMind&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://deepmind.google/frontier-safety/" rel="noopener noreferrer"&gt;Frontier Safety Framework&lt;/a&gt; (v3.1, April 2026)&lt;/td&gt;
&lt;td&gt;Tracked and Critical Capability Levels&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are not three versions of one scale. I cannot cleanly translate an OpenAI “High” result into an Anthropic ASL level or a Google DeepMind Critical Capability Level. The methods, thresholds, and release decisions remain vendor-specific.&lt;/p&gt;

&lt;p&gt;Common evaluation protocols, incident definitions, and auditor qualifications would make the evidence easier to compare. They could replace three separate safety vocabularies with a shared one. After reading the incident reports that sent me down this path, that sounds genuinely useful.&lt;/p&gt;

&lt;p&gt;But a shared vocabulary is not the same thing as a referee.&lt;/p&gt;

&lt;h2&gt;
  
  
  I had to look up why FINRA matters
&lt;/h2&gt;

&lt;p&gt;Hassabis used the Financial Industry Regulatory Authority, or FINRA, as his model. I did not know much about FINRA before reading this proposal. At first, “industry-funded regulator” sounded close to what the AI labs were building.&lt;/p&gt;

&lt;p&gt;The important part is what sits behind those two words. FINRA is privately operated, but it lives inside public law.&lt;/p&gt;

&lt;p&gt;US broker-dealers generally must belong to a self-regulatory organization before doing business. A broker-dealer operating outside the exchanges of which it is a member must usually &lt;a href="https://www.sec.gov/about/divisions-offices/division-trading-markets/division-trading-markets-compliance-guides/guide-broker-dealer-registration" rel="noopener noreferrer"&gt;join FINRA&lt;/a&gt;. FINRA's proposed rule changes are &lt;a href="https://www.sec.gov/rules-regulations/self-regulatory-organization-rulemaking/finra" rel="noopener noreferrer"&gt;filed with the Securities and Exchange Commission&lt;/a&gt; for review. Its board reserves ten seats for industry members, one for its CEO, and &lt;a href="https://www.finra.org/rules-guidance/notices/election-notice-040626" rel="noopener noreferrer"&gt;the remaining seats for public members&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That is what gives the arrangement weight. Firms cannot simply ignore the system and keep doing the same business. The industry does not get final say over every rule. Public representatives have formal power inside the organization.&lt;/p&gt;

&lt;p&gt;None of those conditions has been announced for SAFA. No law compels a lab to participate. No public agency approves its rules. No published charter says who controls its board or what happens when a founder fails a test. Customers, insurers, and governments could eventually give its standards market power, but that would be influence earned after launch, not authority granted at formation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The conflict is not subtle
&lt;/h2&gt;

&lt;p&gt;The first version of SAFA would be formed by three companies whose models and labs it could help evaluate. That does not prove the standards will be weak. I do think it means independence has to be designed into the institution rather than promised after someone raises the question.&lt;/p&gt;

&lt;p&gt;There is also a competition problem. If enterprise buyers or insurers eventually demand a SAFA result, the founding labs will have helped write a test their competitors must pass. A demanding standard could improve safety and still become a moat. I do not think we should pretend only one of those outcomes is possible.&lt;/p&gt;

&lt;p&gt;A famous independent chair would not settle this for me. I would want the boring institutional details: published appointment rules, protected terms for independent board members, transparent funding, open membership, appeal procedures, and results that the founding companies cannot suppress.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I will watch next
&lt;/h2&gt;

&lt;p&gt;SAFA is still a reported plan. I am not ready to call it either a regulator or an industry stunt. I want to see what its founders are willing to give up.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Control:&lt;/strong&gt; Do independent members hold real voting power over standards, leadership, and publication decisions?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transparency:&lt;/strong&gt; Are methods, test versions, incident definitions, and full results public enough to challenge?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Membership:&lt;/strong&gt; Can another frontier lab join as an equal rule-maker on reasonable terms?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consequences:&lt;/strong&gt; What happens after a failed test, and can a founding lab simply leave or release anyway?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bad news:&lt;/strong&gt; Will SAFA publish a finding that delays or embarrasses one of its founders?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For practitioners, a future SAFA report could become useful vendor evidence. I would ask which standard and version was used, who ran the test, what the result excluded, and whether the complete finding is public. I would not confuse that evidence with government approval.&lt;/p&gt;

&lt;p&gt;The next time one of those safety headlines lands, I will look past “we are investigating.” Did the incident enter a shared reporting system? Could an independent body inspect the evidence? Did the failure change a release decision? Could the lab keep an inconvenient result private?&lt;/p&gt;

&lt;p&gt;That is what a referee would change. The real test of SAFA will not be whether it writes sensible rules. It will be whether its founders can lose a call.&lt;/p&gt;

</description>
      <category>governance</category>
      <category>ai</category>
      <category>responsibleai</category>
      <category>agenticai</category>
    </item>
    <item>
      <title>Goose: An Open-Source Agent You Can Assemble Yourself</title>
      <dc:creator>Mangat Rai</dc:creator>
      <pubDate>Thu, 24 Sep 2026 06:20:55 +0000</pubDate>
      <link>https://dev.to/mangatrai/goose-an-open-source-agent-you-can-assemble-yourself-2jmk</link>
      <guid>https://dev.to/mangatrai/goose-an-open-source-agent-you-can-assemble-yourself-2jmk</guid>
      <description>&lt;p&gt;A few days ago, I wrote about &lt;a href="https://fewshotacademy.com/blog/agents-md-guide" rel="noopener noreferrer"&gt;&lt;code&gt;AGENTS.md&lt;/code&gt;&lt;/a&gt;, the file that gives coding agents project-specific instructions. While reading through the Linux Foundation's &lt;a href="https://aaif.io/" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt; site, I came across another founding project: &lt;a href="https://aaif.io/projects/goose" rel="noopener noreferrer"&gt;Goose&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The pitch caught my attention. Goose runs on your machine, lets you choose the model, and connects to tools through the Model Context Protocol. &lt;/p&gt;

&lt;p&gt;However, The &lt;a href="https://goose-docs.ai/docs/quickstart/" rel="noopener noreferrer"&gt;five-minute quickstart&lt;/a&gt; says Goose improves software development by automating coding tasks. The tutorial asks it to build a JavaScript game. The &lt;a href="https://goose-docs.ai/" rel="noopener noreferrer"&gt;Goose homepage&lt;/a&gt;, however, calls it a general-purpose agent for research, writing, automation, data analysis, and code.&lt;/p&gt;

&lt;p&gt;What was I looking at? Is it a coding agent or a general-purpose agent?&lt;/p&gt;

&lt;p&gt;What i figured at the end is that Goose is both, but not equally. Coding is the go to path. Its extension system lets you give the same agent a different set of tools and a different job.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;⏳ &lt;strong&gt;TL;DR&lt;/strong&gt;&lt;br&gt;
Goose is coding-first, not coding-only. Its built-in developer tools cover repository work. Change the extensions and instructions, and it can handle research, documents, browser tasks, data, and other automation. You also choose the model, including compatible local models. That freedom comes with more setup, evaluation, and security work.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  So which is it?
&lt;/h2&gt;

&lt;p&gt;Goose's history clears the confusion. When Block introduced it in January 2025, the &lt;a href="https://goose-docs.ai/blog/2025/01/28/introducing-codename-goose/" rel="noopener noreferrer"&gt;launch post&lt;/a&gt; said its first use cases focused on engineering while the community was already trying it on non-engineering work. The examples included code migrations, tests, benchmarks, feature flags, and API scaffolding.&lt;/p&gt;

&lt;p&gt;The product has since expanded. Its current homepage and &lt;a href="https://github.com/aaif-goose/goose" rel="noopener noreferrer"&gt;README&lt;/a&gt; describe a general-purpose agent. Even its &lt;a href="https://github.com/aaif-goose/goose/blob/main/crates/goose/src/prompts/system.md" rel="noopener noreferrer"&gt;system prompt&lt;/a&gt; uses that term. The quickstart still takes the shortest route to a useful result: give Goose a working directory and ask it to build an application.&lt;/p&gt;

&lt;p&gt;Both descriptions are accurate, but they set different expectations. Calling Goose a coding agent hides what its extension system can do. Calling it a general assistant makes it sound more ready-made than it is outside software development.&lt;/p&gt;

&lt;p&gt;My read is simpler: Goose is an agent you install and configure. It comes prepared for coding. Its tools and instructions determine how far you can take it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually changes
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://goose-docs.ai/docs/goose-architecture/" rel="noopener noreferrer"&gt;architecture documentation&lt;/a&gt; splits Goose into three main pieces: an interface, the agent, and extensions. The model sits behind the agent and decides whether to answer or ask for a tool.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    U["You"] --&amp;gt; I["Desktop, CLI, API,&amp;lt;br/&amp;gt;or ACP client"]
    I --&amp;gt; G["Goose agent runtime&amp;lt;br/&amp;gt;sessions, context, permissions"]
    G &amp;lt;--&amp;gt; M["Your chosen model&amp;lt;br/&amp;gt;hosted or local"]
    G --&amp;gt; E["Built-in and MCP extensions"]
    E --&amp;gt; C["Coding work&amp;lt;br/&amp;gt;files, shell, repositories"]
    E --&amp;gt; W["General work&amp;lt;br/&amp;gt;documents, browser, data, services"]
    C --&amp;gt; G
    W --&amp;gt; G&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Ask Goose to change a repository and it sends the request, relevant context, and available tool definitions to the model. The model can ask to read a file, make an edit, or run a shell command. Goose executes the permitted call, returns the result, and gives the model another turn.&lt;/p&gt;

&lt;p&gt;That loop can also prepare a weekly project brief. Give it local meeting notes and a read-only issue-tracker extension instead of a code repository. The model can read the notes, query completed issues, and write the brief. Goose still coordinates the model and tools; the toolbox and instructions are different.&lt;/p&gt;

&lt;p&gt;The coding task and the project brief are not two different versions of Goose. The loop stays the same. The tools determine what work it can do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with coding
&lt;/h2&gt;

&lt;p&gt;Goose's built-in Developer extension can read and modify files, run shell commands, and maintain a plan. Start a session in a project directory and it has the basic tools needed to investigate and change a codebase.&lt;/p&gt;

&lt;p&gt;This is why the quickstart builds a tic-tac-toe game. In a few minutes, Goose creates the files and, after you add the Computer Controller extension, opens the result in a browser.&lt;/p&gt;

&lt;p&gt;Goose can also run as an &lt;a href="https://goose-docs.ai/docs/gdk/acp/" rel="noopener noreferrer"&gt;ACP server&lt;/a&gt;. ACP, or Agent Client Protocol, separates the agent runtime from its interface. An ACP-compatible editor can use Goose as the agent while Goose keeps its configured models and extensions.&lt;/p&gt;

&lt;p&gt;That does not make Goose an IDE. The desktop app and CLI do not replace an editor's navigation, autocomplete, debugging, or visual review tools. Goose is the agent doing the work. You can run it beside an editor or connect it to one that supports ACP.&lt;/p&gt;

&lt;p&gt;For repository work, I would give it the same project guidance I give any coding agent. An &lt;code&gt;AGENTS.md&lt;/code&gt; file can state the build commands, paths, conventions, and boundaries it should follow. Open source at the runtime layer does not remove the need for clear project context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond code, you assemble the workflow
&lt;/h2&gt;

&lt;p&gt;Goose calls its MCP servers extensions. MCP is the open protocol for connecting agents to tools and data; the &lt;a href="https://dev.to/docs/mcp/what-is-mcp"&gt;MCP track&lt;/a&gt; explains the host, client, and server roles in detail. Goose's &lt;a href="https://goose-docs.ai/docs/getting-started/using-extensions/" rel="noopener noreferrer"&gt;built-in catalog&lt;/a&gt; includes the Developer, Computer Controller, Memory, and Auto Visualiser extensions, and it can connect to external MCP servers for other systems.&lt;/p&gt;

&lt;p&gt;A research setup might get access to a document directory and a browser. An operations setup might use an issue tracker, a database, and a reporting destination. Calling Goose general-purpose does not give the model those abilities automatically. You add them by choosing extensions.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://goose-docs.ai/docs/guides/recipes/" rel="noopener noreferrer"&gt;Recipes&lt;/a&gt; make those choices repeatable. A recipe can package instructions, parameters, model settings, and required extensions in YAML or JSON. Goose can run recipes interactively, headlessly, or &lt;a href="https://goose-docs.ai/docs/guides/recipes/session-recipes/" rel="noopener noreferrer"&gt;on a schedule&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This is what I mean by an agent you assemble yourself. You do not have to build the agent loop, but you do choose the parts around it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Part&lt;/th&gt;
&lt;th&gt;What you choose&lt;/th&gt;
&lt;th&gt;What you still have to verify&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;Hosted provider, private endpoint, or compatible local model&lt;/td&gt;
&lt;td&gt;Tool use, quality, context, latency, and cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capabilities&lt;/td&gt;
&lt;td&gt;Built-in tools and local or remote MCP extensions&lt;/td&gt;
&lt;td&gt;Data access, credentials, dependencies, and permissions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workflow&lt;/td&gt;
&lt;td&gt;Instructions, recipes, parameters, retries, and schedules&lt;/td&gt;
&lt;td&gt;Correctness, failure handling, and review points&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interface&lt;/td&gt;
&lt;td&gt;Desktop, CLI, API, or ACP client&lt;/td&gt;
&lt;td&gt;Which surface fits the people and task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Distribution&lt;/td&gt;
&lt;td&gt;Upstream Goose or a custom distribution&lt;/td&gt;
&lt;td&gt;Upgrades and the cost of maintaining changes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This path is less turnkey than coding. Goose gives you the runtime and extension system. It does not know which business process matters, which source is authoritative, or what a correct result looks like. You still have to design and test that workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bring your own model, with caveats
&lt;/h2&gt;

&lt;p&gt;Goose does not bundle one mandatory model. Its &lt;a href="https://goose-docs.ai/docs/getting-started/providers/" rel="noopener noreferrer"&gt;provider guide&lt;/a&gt; covers hosted providers, cloud platforms, compatible private endpoints, subscription-backed agents, and local options. Goose connects to the provider and runs the agent loop around the selected model.&lt;/p&gt;

&lt;p&gt;A coding team can test another model without replacing Goose. A broader workflow can use a model and endpoint that fit its cost, data, or deployment requirements.&lt;/p&gt;

&lt;p&gt;Ollama is one way to run inference locally. With a compatible tool-calling model, Goose and the model can both run on your machine. Newer Goose desktop releases also provide &lt;a href="https://goose-docs.ai/blog/2026/04/24/use-goose-with-built-in-local-inference/" rel="noopener noreferrer"&gt;built-in local inference&lt;/a&gt;, so Ollama is no longer the only local route.&lt;/p&gt;

&lt;p&gt;Local inference does not automatically make the whole workflow private or offline. A remote MCP extension still receives whatever its tool call sends, and browser work still uses the network. The model also needs to fit the hardware and reliably make structured tool calls. Goose's documentation warns that models without tool calling can provide chat completion only when extensions are disabled.&lt;/p&gt;

&lt;p&gt;The models are not interchangeable just because Goose can connect to them. Two models may accept the same prompt and expose the same tools while producing very different results. I would treat every provider change as a new agent configuration and rerun the task evaluation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Goose fits
&lt;/h2&gt;

&lt;p&gt;I would choose Goose when the work begins with code but does not end there. The same agent that changes code might query an issue tracker, inspect a design document, update release notes, or prepare a deployment checklist. It becomes more useful when a job crosses tool boundaries.&lt;/p&gt;

&lt;p&gt;It is also a useful place to compare models against a reasonably stable runtime. Keep the instructions, extensions, inputs, and acceptance criteria fixed, then change the provider. That does not make the models interchangeable, but it makes their differences easier to see.&lt;/p&gt;

&lt;p&gt;I would also consider it when I need control over the runtime or data path. Goose is an Apache 2.0 project governed at the &lt;a href="https://aaif.io/projects/goose" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt;, and it documents &lt;a href="https://goose-docs.ai/docs/guides/custom-distributions/" rel="noopener noreferrer"&gt;custom distributions&lt;/a&gt;. An organization can inspect the system and adapt it instead of waiting for a hosted product to add the option it needs.&lt;/p&gt;

&lt;p&gt;Open source still does not eliminate operating work. A private fork needs maintenance. A local model needs hardware and evaluation. Extensions bring dependencies and credentials. The benefit is the ability to make those decisions, not the disappearance of their cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  When I would look elsewhere
&lt;/h2&gt;

&lt;p&gt;If I wanted a polished coding experience entirely inside an editor, I would first check whether Goose's ACP integration provides the interactions I rely on. The agent can connect to an editor, but the surrounding editor experience still comes from that client.&lt;/p&gt;

&lt;p&gt;If I wanted a general assistant with almost no setup, I would also look elsewhere. Choosing a provider, enabling extensions, defining permissions, and turning a task into a reliable recipe are part of using Goose. A managed assistant makes more of those decisions for you.&lt;/p&gt;

&lt;p&gt;The biggest concern is how much authority you give it. Goose can execute commands, change files, control a browser, and call external services. Its &lt;a href="https://goose-docs.ai/docs/guides/managing-tools/goose-permissions/" rel="noopener noreferrer"&gt;permission controls&lt;/a&gt;, &lt;a href="https://goose-docs.ai/docs/guides/security/prompt-injection-detection/" rel="noopener noreferrer"&gt;prompt-injection detection&lt;/a&gt;, and &lt;a href="https://goose-docs.ai/docs/guides/security/adversary-mode/" rel="noopener noreferrer"&gt;adversary review&lt;/a&gt; can reduce risk, but they do not make every extension safe. I would start with the smallest tool set, least-privilege credentials, approval for consequential actions, and an output I can inspect.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://dev.to/docs/advanced-concepts/agent-security"&gt;agent security chapter&lt;/a&gt; explains why untrusted tool output needs rules enforced in code, not only instructions in a prompt. That becomes more important as Goose moves beyond a repository and into other systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I would try it
&lt;/h2&gt;

&lt;p&gt;I started with the question of whether Goose was a coding agent or a general-purpose agent. I would test it in that same order.&lt;/p&gt;

&lt;p&gt;First, I would use it for one bounded coding task in a disposable branch: inspect a small repository, make a change, run the tests, and explain the result. That exercises Goose's strongest default path and reveals how well the selected model uses its developer tools.&lt;/p&gt;

&lt;p&gt;Then I would reuse the runtime for one bounded non-coding task. My choice would be a weekly project brief built from a local notes folder and one read-only issue-tracker extension. I would define the required sections and citations, keep approval enabled, and compare the result with the source material.&lt;/p&gt;

&lt;p&gt;Only after the interactive task worked would I turn it into a recipe or schedule. That sequence separates three questions: can the model complete the task, have I given the agent the right tools, and is the workflow dependable enough to repeat without constant supervision?&lt;/p&gt;

&lt;p&gt;Goose is a coding agent you can turn into something broader. That is its sweet spot, and also its cost. You get to choose the model, tools, and job, but you also have to assemble and test the result.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>goose</category>
      <category>ai</category>
      <category>agenticai</category>
    </item>
    <item>
      <title>Agent security starts before the first prompt</title>
      <dc:creator>Mangat Rai</dc:creator>
      <pubDate>Sun, 13 Sep 2026 23:47:16 +0000</pubDate>
      <link>https://dev.to/mangatrai/agent-security-starts-before-the-first-prompt-55ak</link>
      <guid>https://dev.to/mangatrai/agent-security-starts-before-the-first-prompt-55ak</guid>
      <description>&lt;p&gt;An agent can cause damage before the model makes its first decision.&lt;/p&gt;

&lt;p&gt;On September 1, 2026, Manifold described coding agents that invoked Git while gathering workspace context, allowing repository-local Git configuration to run a program on the host. Some affected agents performed that work before a user typed a prompt or approved workspace trust.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://www.manifold.security/blog/ai-coding-agents-git-hijack" rel="noopener noreferrer"&gt;Manifold's GitSpawn research&lt;/a&gt;, the relevant path was a repository received as files, with a &lt;code&gt;.git&lt;/code&gt; directory inside a ZIP archive. Git read repository-local configuration while the agent gathered context. An ordinary Git clone, fetch, or pull does &lt;strong&gt;not&lt;/strong&gt; transfer repository-local &lt;code&gt;.git/config&lt;/code&gt; from the remote.&lt;/p&gt;

&lt;p&gt;The model did not choose a harmful tool call. The runtime acted while preparing to use the model. That is why an agent review needs three boundaries: the runtime that starts processes, the gateway that holds credentials and routes requests, and the tool that makes a sensitive change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Workspace trust has to include startup
&lt;/h2&gt;

&lt;p&gt;Imagine a coding agent opening a project so it can answer a simple question about a failing test. Before the chat box is useful, the application may inspect the working tree, run Git to learn the branch and status, index files, load extensions, and collect credentials for the model gateway. None of that is a model-requested tool call. It is still code acting with the authority of the agent process.&lt;/p&gt;

&lt;p&gt;Manifold reported eight findings across seven coding agents. Some affected paths ran before authentication or workspace approval. At publication, some issues had been patched; Manifold said four findings remained unpatched and withheld details for those paths.&lt;/p&gt;

&lt;p&gt;For the coding agent in this example, I would make the startup contract explicit. It may read the checked-out project files. It may not inherit production cloud credentials, reach arbitrary internet hosts, or start programs named by workspace configuration. Opening an unfamiliar workspace happens in a reduced-permission environment. If the product asks for workspace trust before executing code, discovery and indexing need to honor the same boundary. Keeping production credentials out of its environment and filesystem limits what a missed startup hook can expose.&lt;/p&gt;

&lt;p&gt;This is not a replacement for prompt-injection defenses. It covers authority that exists before the model is asked to reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gateway is also outside the model loop
&lt;/h2&gt;

&lt;p&gt;The same coding agent may send its prompts through a gateway. That gateway can hold provider keys, route requests, apply guardrails, and reach internal services. It is a useful policy point, but it is also a privileged service. Its authentication, administration, credentials, and network reach are security decisions made outside the model's visible tool loop.&lt;/p&gt;

&lt;p&gt;Wiz's &lt;a href="https://www.wiz.io/blog/off-guard-breaking-litellm-from-authentication-bypass-to-cloud-compromise" rel="noopener noreferrer"&gt;September 9 LiteLLM research&lt;/a&gt; gives that distinction some weight. In a February 2026 scan of 3,074 publicly exposed instances, Wiz found that 9.6% either accepted the default master key or required no authentication. That is a combined figure, not a claim that 9.6% used the default key, and it does not describe private deployments.&lt;/p&gt;

&lt;p&gt;Wiz also reported an MCP authentication bypass, CVE-2026-59822, and a separate post-authentication code-execution issue in custom-code guardrails, CVE-2026-59821. The researchers wrote that the custom-code guardrail registration path lacked protections present in the test endpoint. Patches were available when Wiz published. A gateway deserves the same careful boundary design as any other service that holds keys and reaches sensitive systems.&lt;/p&gt;

&lt;p&gt;For this agent, the gateway should reject a request without a valid credential before it routes anything. Its service identity should have only the provider access and internal destinations the task needs. Administrative and custom-code features should be limited to the small group that operates them. Keeping a gateway private, patching it with an owner and response window, and logging administrative changes are ordinary infrastructure controls. They matter here because a compromise could act without asking the model to call a tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the last boundary at the tool
&lt;/h2&gt;

&lt;p&gt;Now return to the coding agent after it has started safely and reached an authenticated gateway. The model examines a failing test and proposes a sensitive action: deploy a fix to the production environment.&lt;/p&gt;

&lt;p&gt;That decision still needs a tool boundary. The deployment service can require an approved environment, a change identifier, and a human approval. The model can propose the action, but deterministic code decides whether the request fits the authority granted to this task. A model instruction such as “deploy it now” cannot substitute for those checks.&lt;/p&gt;

&lt;p&gt;This is where tool allowlists and argument validation belong. They are strong controls for actions the model requests: writing outside an approved filesystem root, querying a restricted database schema, sending to an unapproved recipient, or deploying to an environment that needs review. They do not secure Git startup or gateway authentication because those events happen on different paths.&lt;/p&gt;

&lt;h2&gt;
  
  
  Review one agent you run
&lt;/h2&gt;

&lt;p&gt;Follow one sensitive job from startup to its most consequential action. For the illustrative coding agent, the map might look like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;Credential or network reach&lt;/th&gt;
&lt;th&gt;Enforcing control&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Workspace runtime&lt;/td&gt;
&lt;td&gt;Read and index project files&lt;/td&gt;
&lt;td&gt;Gateway access token; no provider or production cloud keys; gateway and approved dependency endpoints only&lt;/td&gt;
&lt;td&gt;Isolated workspace, restricted subprocesses, outbound allowlist&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model gateway&lt;/td&gt;
&lt;td&gt;Route model requests and apply policy&lt;/td&gt;
&lt;td&gt;Scoped provider credential; approved internal services&lt;/td&gt;
&lt;td&gt;Authentication that fails closed, private exposure, patched release, least-privilege service identity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deployment tool&lt;/td&gt;
&lt;td&gt;Request a production deployment&lt;/td&gt;
&lt;td&gt;Deployment API for the approved environment&lt;/td&gt;
&lt;td&gt;Environment allowlist, change ID, human approval&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The point is to name the authority each component actually has, then attach a control that can enforce it. A local learning project may need far less isolation than an unattended agent with production access.&lt;/p&gt;

&lt;p&gt;One boundary test makes the map real. Send the gateway the same request the coding agent would make, but omit its gateway credential. The expected result is an authentication rejection before routing, with no provider request and no internal service access. A successful model response or evidence in downstream routing logs that the request was forwarded is a failure. Record the rejection and check downstream logs; an error message alone does not show whether the request was forwarded.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://dev.to/docs/advanced-concepts/agent-security"&gt;Agent Security lesson&lt;/a&gt; demonstrates the tool-call side with an indirect prompt-injection lab. &lt;a href="https://dev.to/docs/advanced-concepts/ai-gateways"&gt;AI Gateways&lt;/a&gt; explains the application-owned routing boundary.&lt;/p&gt;

&lt;p&gt;Record expected and actual results for one boundary test, then fix one gap before giving the agent more authority.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>security</category>
      <category>infrastructure</category>
      <category>llm</category>
    </item>
    <item>
      <title>AGENTS.md: one guide for your coding agents</title>
      <dc:creator>Mangat Rai</dc:creator>
      <pubDate>Wed, 09 Sep 2026 00:23:15 +0000</pubDate>
      <link>https://dev.to/mangatrai/agentsmd-one-guide-for-your-coding-agents-be</link>
      <guid>https://dev.to/mangatrai/agentsmd-one-guide-for-your-coding-agents-be</guid>
      <description>&lt;p&gt;&lt;code&gt;AGENTS.md&lt;/code&gt; is a Markdown file you keep in your repository to tell coding agents how to work on the project. It can explain where the code lives, which commands to run, and which decisions the agent should leave to you. The &lt;a href="https://agents.md/" rel="noopener noreferrer"&gt;open format&lt;/a&gt; has no required fields or schema.&lt;/p&gt;

&lt;p&gt;Imagine asking an agent to add a page on Monday. You explain the project's colors, where the source lives, and why the navigation must stay as it is. On Friday, a contributor picks up the work in another coding tool. Their agent proposes a new color scheme and reorganizes the navigation. The decisions are still in Monday's conversation, but they never made it into the repository.&lt;/p&gt;

&lt;p&gt;That is the gap this file is meant to close. Write the ground rules down once, keep them with the code, and give the next agent a place to start.&lt;/p&gt;

&lt;h2&gt;
  
  
  The instructions should travel with the project
&lt;/h2&gt;

&lt;p&gt;For an open-source maintainer, Friday's contributor may use a different coding tool. Keeping the instructions in the repository makes them available to both people; a shared filename gives their tools a common place to look. One contributor can use Codex and another Cursor without needing separate copies of the same build instructions.&lt;/p&gt;

&lt;p&gt;That practical benefit is also the story behind the name. Amp initially used &lt;code&gt;AGENT.md&lt;/code&gt;, singular. When OpenAI chose &lt;code&gt;AGENTS.md&lt;/code&gt;, Amp agreed to switch if OpenAI secured the matching domain. On &lt;a href="https://ampcode.com/news/AGENTS.md" rel="noopener noreferrer"&gt;August 20, 2025, Amp announced the change&lt;/a&gt;. Sharing a standard mattered more than keeping its original filename.&lt;/p&gt;

&lt;p&gt;The Linux Foundation dates the format's release to August 2025. By its &lt;a href="https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation" rel="noopener noreferrer"&gt;December 9 announcement of the Agentic AI Foundation&lt;/a&gt;, it reported adoption by more than 60,000 open-source projects and agent frameworks. That is a dated adoption report, not a live count. &lt;code&gt;AGENTS.md&lt;/code&gt; became one of the foundation's initial projects.&lt;/p&gt;

&lt;p&gt;The convention gives those contributors a common starting point, although each tool still decides how to discover and apply the file. First, that shared file needs instructions worth carrying forward.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write down the decisions the next agent needs
&lt;/h2&gt;

&lt;p&gt;A useful file answers the questions that would otherwise interrupt the work: where to start, what to run, what to preserve, and how to know the change is ready.&lt;/p&gt;

&lt;p&gt;For our imagined documentation project, “follow best practices” adds little. “Edit source files in &lt;code&gt;site/&lt;/code&gt;; do not edit generated files in &lt;code&gt;site/build/&lt;/code&gt;” settles a concrete decision. “Run the checks” is vague. Naming the commands and their working directory makes the instruction usable.&lt;/p&gt;

&lt;p&gt;Four sections are a useful starting structure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Project and layout.&lt;/strong&gt; What the project does and the few directories an agent needs to understand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Setup and verification.&lt;/strong&gt; Exact commands, where to run them, and any prerequisites.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conventions and boundaries.&lt;/strong&gt; Existing patterns to reuse, files to preserve, and actions that need approval.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completion.&lt;/strong&gt; What evidence to report and how to describe anything left untested.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For a small repository, I would start with roughly 30–60 short lines. That is my editing budget, not a format limit. If ten lines cover the important decisions, stop there. Add a rule when you can name the recurring mistake it should prevent.&lt;/p&gt;

&lt;p&gt;Here is an illustrative, shortened file for the website portion of Few-Shot Academy. The paths and commands come from this repository; the example is not a replacement for our full &lt;a href="https://github.com/fewshot-works/academy/blob/main/AGENTS.md" rel="noopener noreferrer"&gt;contributor instructions&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Sample AGENTS.md
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Project instructions&lt;/span&gt;

Few-Shot Academy is a free GenAI curriculum. Write for readers
who may have no programming or AI background.

&lt;span class="gu"&gt;## Project layout&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`site/docs/`&lt;/span&gt;: curriculum pages.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`site/blog/`&lt;/span&gt;: standalone practitioner articles.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`site/src/`&lt;/span&gt;: shared components and styles.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`site/build/`&lt;/span&gt;: generated output; edit the source instead.

&lt;span class="gu"&gt;## Setup and checks&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Use Node.js 22 or newer.
&lt;span class="p"&gt;-&lt;/span&gt; From &lt;span class="sb"&gt;`site/`&lt;/span&gt;, run &lt;span class="sb"&gt;`npm ci`&lt;/span&gt; to install dependencies.
&lt;span class="p"&gt;-&lt;/span&gt; After site changes, run these from &lt;span class="sb"&gt;`site/`&lt;/span&gt;:
&lt;span class="p"&gt;  -&lt;/span&gt; &lt;span class="sb"&gt;`npm run typecheck`&lt;/span&gt;
&lt;span class="p"&gt;  -&lt;/span&gt; &lt;span class="sb"&gt;`npm run build`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; For visual changes, check mobile layout, keyboard navigation,
  and both light and dark themes.

&lt;span class="gu"&gt;## Working rules&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Read the relevant page and nearby examples before editing.
&lt;span class="p"&gt;-&lt;/span&gt; Reuse design tokens in &lt;span class="sb"&gt;`site/src/css/custom.css`&lt;/span&gt;.
&lt;span class="p"&gt;-&lt;/span&gt; Preserve published URLs.
&lt;span class="p"&gt;-&lt;/span&gt; Verify factual claims against primary sources.
&lt;span class="p"&gt;-&lt;/span&gt; Keep changes focused and preserve unrelated work.
&lt;span class="p"&gt;-&lt;/span&gt; Ask before pushing, merging, or deploying.

&lt;span class="gu"&gt;## Before finishing&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Summarize the change and the checks that passed.
&lt;span class="p"&gt;-&lt;/span&gt; State what you could not verify and why.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next agent now has somewhere to find the source paths, design rules, and checks that were stranded in Monday's conversation. Replace the example's paths and commands with ones you have verified in your own repository.&lt;/p&gt;

&lt;p&gt;These are still instructions the model receives, so check its work. “Do not deploy” is useful guidance; actual deployment access belongs in permissions and approval controls. Anthropic's &lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; explicitly distinguishes instruction files from enforced configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make sure your tool reads it
&lt;/h2&gt;

&lt;p&gt;With the file written, the next step is to connect it to the agent. Some tools recognize &lt;code&gt;AGENTS.md&lt;/code&gt; directly; others need a small configuration change.&lt;/p&gt;

&lt;p&gt;This is a selection of coding editors and agents, not a popularity ranking. The table reflects their official documentation checked on September 8, 2026. “Automatic” means the tool recognizes the file without a filename bridge; settings can still affect loading.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Root &lt;code&gt;AGENTS.md&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;Details or setup&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://learn.chatgpt.com/docs/agent-configuration/agents-md" rel="noopener noreferrer"&gt;Codex&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Automatic&lt;/td&gt;
&lt;td&gt;Combines guidance along the root-to-working-directory path. An &lt;code&gt;AGENTS.override.md&lt;/code&gt; takes priority over &lt;code&gt;AGENTS.md&lt;/code&gt; in the same directory.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://cursor.com/docs/rules#agentsmd" rel="noopener noreferrer"&gt;Cursor Agent&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Automatic&lt;/td&gt;
&lt;td&gt;Also applies nested files to their directory and children. More specific instructions take precedence.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://code.visualstudio.com/docs/agent-customization/custom-instructions" rel="noopener noreferrer"&gt;GitHub Copilot in VS Code&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Automatic&lt;/td&gt;
&lt;td&gt;Root loading is enabled by default. Nested discovery is a separate experimental opt-in; see the &lt;a href="https://code.visualstudio.com/docs/agents/reference/ai-settings" rel="noopener noreferrer"&gt;settings reference&lt;/a&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://docs.devin.ai/desktop/cascade/memories" rel="noopener noreferrer"&gt;Cascade (Windsurf / Devin Desktop)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Automatic&lt;/td&gt;
&lt;td&gt;Root instructions are always on; subdirectory files apply to that part of the project.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://ampcode.com/docs/customize/agents-md" rel="noopener noreferrer"&gt;Amp&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Automatic&lt;/td&gt;
&lt;td&gt;Discovers instructions in the working directory, parents, and relevant subdirectories.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://docs.cline.bot/customization/cline-rules" rel="noopener noreferrer"&gt;Cline&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Automatic&lt;/td&gt;
&lt;td&gt;Recognizes &lt;code&gt;AGENTS.md&lt;/code&gt;; check the Rules panel for whether a detected rule is enabled.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://geminicli.com/docs/cli/gemini-md/" rel="noopener noreferrer"&gt;Gemini CLI&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Configure filename&lt;/td&gt;
&lt;td&gt;Defaults to &lt;code&gt;GEMINI.md&lt;/code&gt;. Set &lt;code&gt;context.fileName&lt;/code&gt; to include &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://code.claude.com/docs/en/memory#agentsmd" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Import or symlink&lt;/td&gt;
&lt;td&gt;Reads &lt;code&gt;CLAUDE.md&lt;/code&gt;. Import the shared file as shown below.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Copilot row is specifically about VS Code. Do not assume that an editor extension, a command-line agent, and a hosted pull-request agent load instructions identically just because they share a product name.&lt;/p&gt;

&lt;h3&gt;
  
  
  Connecting Claude Code and Gemini CLI
&lt;/h3&gt;

&lt;p&gt;The last two rows need setup, but you can keep the shared rules in &lt;code&gt;AGENTS.md&lt;/code&gt;. For Claude Code, put this line in the root &lt;code&gt;CLAUDE.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@AGENTS.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Copy the line without the surrounding code fences. That is the arrangement in this repository. Claude's &lt;a href="https://code.claude.com/docs/en/memory#agentsmd" rel="noopener noreferrer"&gt;documented import&lt;/a&gt; loads the shared instructions at session start. Keep any existing Claude-specific guidance below it. A root import does not by itself wire up every nested &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A symlink is another supported option: &lt;code&gt;ln -s AGENTS.md CLAUDE.md&lt;/code&gt; on macOS or Linux, when &lt;code&gt;CLAUDE.md&lt;/code&gt; does not already exist. I prefer the import because it is a plain-text edit on macOS, Linux, and Windows. Anthropic also recommends it on Windows, where creating symlinks can require extra privileges.&lt;/p&gt;

&lt;p&gt;For Gemini CLI, merge this into &lt;code&gt;.gemini/settings.json&lt;/code&gt;, preserving your other settings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"fileName"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"AGENTS.md"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GEMINI.md"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gemini's &lt;a href="https://geminicli.com/docs/cli/gemini-md/" rel="noopener noreferrer"&gt;context configuration&lt;/a&gt; supports multiple filenames. Keeping both lets existing Gemini-specific guidance remain discoverable; avoid repeating the shared rules in both files.&lt;/p&gt;

&lt;h2&gt;
  
  
  Let the setup grow with the work
&lt;/h2&gt;

&lt;p&gt;Once the short file is written and your tool is set up to read it, you have a starting point for the next session. Over time, you may also want to preserve architecture explanations, design reasoning, or the state of unfinished work. Those do not all belong in the everyday instructions.&lt;/p&gt;

&lt;p&gt;Separate architecture, design, and decision notes can preserve reasoning the code cannot explain. In our Monday-to-Friday example, “use these colors” records a rule. “We kept the existing palette so new pages match the rest of the site” records why it exists.&lt;/p&gt;

&lt;p&gt;But five overlapping summaries give you five places to forget an update. Choose a pattern that solves a problem you actually have:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;When it helps&lt;/th&gt;
&lt;th&gt;What to watch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Short root file with links to project docs&lt;/td&gt;
&lt;td&gt;Architecture or design explanations are too long for everyday instructions.&lt;/td&gt;
&lt;td&gt;Say when to read each document; a link alone does not guarantee it gets read.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instructions scoped to a directory&lt;/td&gt;
&lt;td&gt;Different packages need different checks or conventions.&lt;/td&gt;
&lt;td&gt;Confirm the tool's nesting rules. Avoid copying root rules into every package.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A temporary handoff note&lt;/td&gt;
&lt;td&gt;A task spans sessions or pauses halfway through.&lt;/td&gt;
&lt;td&gt;Record the current state, unresolved questions, and next step; replace outdated notes when work moves on.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two have direct support in tools such as &lt;a href="https://cursor.com/docs/rules" rel="noopener noreferrer"&gt;Cursor&lt;/a&gt;, which documents references and nested instructions. For handoffs, Anthropic describes &lt;a href="https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents" rel="noopener noreferrer"&gt;using progress files alongside Git history&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Splitting files does not create more memory
&lt;/h3&gt;

&lt;p&gt;The context window, the material a model can use at once, is finite. Summarizing a long conversation helps it continue, but that summary can miss details the next session needs.&lt;/p&gt;

&lt;p&gt;Tools can also limit how much they load. Codex has a &lt;a href="https://learn.chatgpt.com/docs/agent-configuration/agents-md" rel="noopener noreferrer"&gt;32 KiB default limit&lt;/a&gt; on the combined project instructions it loads, but that is a ceiling, not a target.&lt;/p&gt;

&lt;p&gt;The loading strategy matters more than the file count. Claude Code's &lt;a href="https://code.claude.com/docs/en/memory#import-additional-files" rel="noopener noreferrer"&gt;&lt;code&gt;@&lt;/code&gt; imports&lt;/a&gt; load the referenced content at launch. Splitting one large file into five and importing all five still puts that material into context. To keep startup context small, keep the root guidance brief and instruct the agent to read the relevant supporting document when the task calls for it.&lt;/p&gt;

&lt;p&gt;A “resume this project” prompt can tell the agent to read a handoff; it cannot recover details nobody saved. Update the relevant document when a decision changes. Before pausing, record unfinished work, checks performed, and the next step. Review those updates in the diff instead of assuming the agent made them.&lt;/p&gt;

&lt;p&gt;My default is one short &lt;code&gt;AGENTS.md&lt;/code&gt;, existing project docs for durable decisions, and a handoff note only when there is work to resume. Add files when they remove confusion.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the next handoff
&lt;/h2&gt;

&lt;p&gt;Return to Friday's contributor. The useful test is whether the saved instructions help them pick up the work. Start a fresh session after changing the setup and try a small page edit. Did the agent find the source, use the existing design, run the right checks, and report anything it could not verify?&lt;/p&gt;

&lt;p&gt;If it misses a rule, check which files loaded and whether their instructions conflict before adding another paragraph. The goal is to stop re-explaining the same decisions, while keeping those decisions easy to find and maintain.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>openai</category>
      <category>claude</category>
    </item>
    <item>
      <title>GPT-6 Astra: the harness is the product</title>
      <dc:creator>Mangat Rai</dc:creator>
      <pubDate>Sun, 06 Sep 2026 00:08:52 +0000</pubDate>
      <link>https://dev.to/mangatrai/gpt-6-astra-the-harness-is-the-product-1khe</link>
      <guid>https://dev.to/mangatrai/gpt-6-astra-the-harness-is-the-product-1khe</guid>
      <description>&lt;p&gt;GPT-6 Astra does not prove that OpenAI has reached AGI. What it does prove is that the harness can no longer be treated as plumbing.&lt;/p&gt;

&lt;p&gt;Astra scored 54.8% on ARC-AGI-3 at high reasoning in the standard harness. The same model scored 99.9% when the harness preserved its reasoning state and compacted long conversations. That is a 45.1-point difference without changing the model.&lt;/p&gt;

&lt;p&gt;My takeaway is simple: if you evaluate only the model name, you are evaluating the wrong product. For long-running agent work, the product is the model, memory, tools, context management, and control loop together.&lt;/p&gt;

&lt;p&gt;That is why Astra matters even if you are not interested in arguing about AGI.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AGI claim is ahead of the evidence
&lt;/h2&gt;

&lt;p&gt;OpenAI President Greg Brockman believes Astra qualifies as artificial general intelligence. During the launch briefing, he left the final judgment to others but said, “I think we're there,” according to &lt;a href="https://www.washingtonpost.com/technology/2026/09/03/openai-greg-brockman-says-its-new-model-astra-is-agi/" rel="noopener noreferrer"&gt;The Washington Post&lt;/a&gt;. OpenAI's &lt;a href="https://openai.com/charter/" rel="noopener noreferrer"&gt;Charter defines AGI&lt;/a&gt; as highly autonomous systems that outperform humans at most economically valuable work. No launch benchmark can establish that on its own.&lt;/p&gt;

&lt;p&gt;Astra gives OpenAI a serious case to make. The company reports 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and large gains in computer use and coding. &lt;a href="https://openai.com/index/gpt-6-astra/" rel="noopener noreferrer"&gt;OpenAI is positioning Astra&lt;/a&gt; to work inside browsers, development tools, spreadsheets, and other software with less supervision.&lt;/p&gt;

&lt;p&gt;I would not call it “pre-AGI” either. That avoids the argument without giving us anything measurable. Astra is a meaningful step toward a general-purpose digital worker, but we do not yet have evidence that it can reliably outperform people across most economically valuable work.&lt;/p&gt;

&lt;p&gt;ARC Prize calls Astra a major step forward while stating plainly that saturating ARC-AGI-3 is not proof of AGI. The benchmark uses bounded environments with deterministic rules. Real work is messier and carries consequences when the agent gets something wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The biggest Astra number is really a harness result
&lt;/h2&gt;

&lt;p&gt;An agent harness is the software around the model. It supplies tools, carries context forward, manages retries, records state, and decides when the model gets another turn.&lt;/p&gt;

&lt;p&gt;ARC Prize tested Astra through two harnesses:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evaluation setup&lt;/th&gt;
&lt;th&gt;Reasoning effort&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;th&gt;Evaluation cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Standard, provider-neutral harness&lt;/td&gt;
&lt;td&gt;Max&lt;/td&gt;
&lt;td&gt;62.7%&lt;/td&gt;
&lt;td&gt;$26,098&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard, provider-neutral harness&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;54.8%&lt;/td&gt;
&lt;td&gt;$40,705&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Provider Adapter harness&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;99.9%&lt;/td&gt;
&lt;td&gt;$18,817&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The standard harness gives the model a common interface and lets it decide what to preserve in visible notes. The OpenAI adapter also keeps Astra's opaque reasoning state between requests and uses compaction to manage long conversations.&lt;/p&gt;

&lt;p&gt;At the same high reasoning setting, the adapter moved the score from 54.8% to 99.9%, according to &lt;a href="https://arcprize.org/blog/astra" rel="noopener noreferrer"&gt;ARC Prize's analysis&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The 99.9% result is valid. It is just not a property of the model in isolation. It is the result of Astra operating inside the runtime OpenAI designed for it.&lt;/p&gt;

&lt;p&gt;The neutral harness is better for comparing models under the same conditions. The native harness measures the system you might actually deploy. One score cannot answer both questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The new context system is the feature I would watch
&lt;/h2&gt;

&lt;p&gt;Astra has a 1,050,000-token context window, but a coding agent can still fill it with source files, test output, logs, and failed approaches. Eventually, something has to be removed or compressed.&lt;/p&gt;

&lt;p&gt;Compaction turns the old conversation into a shorter summary. That summary preserves what looked important at the time. A constraint, test failure, or rejected approach can disappear and become important again later.&lt;/p&gt;

&lt;p&gt;With Astra, OpenAI says Codex can now &lt;a href="https://openai.com/index/gpt-6-astra/" rel="noopener noreferrer"&gt;keep notes across context windows and search earlier windows&lt;/a&gt;. If a detail did not make it into the notes, the agent can retrieve the earlier message or tool result instead of depending entirely on a chain of summaries.&lt;/p&gt;

&lt;p&gt;Consider a repository migration. The agent learns early that one service must stay compatible with an older client, then spends an hour inspecting unrelated code and running tests. When it finally edits that service, searchable history gives it another way to recover the compatibility requirement instead of hoping the detail survived compaction.&lt;/p&gt;

&lt;p&gt;This could make a long-running agent more dependable, but it needs real testing. The Codex feature is experimental at launch, and we do not yet know how reliably Astra will find the right old detail across large, messy repositories.&lt;/p&gt;

&lt;h2&gt;
  
  
  Astra can keep working while the task changes
&lt;/h2&gt;

&lt;p&gt;Two new API capabilities support the same idea.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://developers.openai.com/api/docs/guides/async-tool-calling" rel="noopener noreferrer"&gt;Async tool calling&lt;/a&gt; lets Astra start a function or custom tool and continue with independent work while the application runs it. An agent can begin a slow dependency scan, inspect configuration files while it waits, and use the scan after it returns. Your application still runs the job and has to return the result with the correct call ID.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://developers.openai.com/api/docs/guides/steering" rel="noopener noreferrer"&gt;Mid-turn steering&lt;/a&gt; lets a user change requirements while Astra is working over a WebSocket connection. You can tell the migration agent to keep the project within one engineer-week without discarding the research it already completed.&lt;/p&gt;

&lt;p&gt;That sounds like a small interface improvement. In practice, it changes the agent from a request-response system into a process you can supervise while it runs.&lt;/p&gt;

&lt;p&gt;The tradeoff is control. Steering does not undo an action that already happened or cancel a tool already in flight. Async tools can also finish in an unexpected order. If two of them write to the same system, the application needs isolation, idempotency, and a clear approval policy. A more capable model does not remove those engineering responsibilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  More autonomy also raises the cost of a mistake
&lt;/h2&gt;

&lt;p&gt;OpenAI calls Astra its most aligned model and reports better respect for task boundaries. Its safety review also found that Astra can sometimes evade internal monitors when adversarial tests explicitly tell it to do so. OpenAI says this came from adversarial evaluation rather than normal use, but the result is serious enough that the company is developing monitoring methods beyond reading the model's written reasoning.&lt;/p&gt;

&lt;p&gt;Astra is also OpenAI's first broadly deployed model to reach its &lt;a href="https://openai.com/index/safety-overview-gpt-6-astra/" rel="noopener noreferrer"&gt;Critical cybersecurity capability threshold&lt;/a&gt;. OpenAI says that, with the right tools and access, Astra can find previously unknown vulnerabilities and develop exploits against hardened systems without a person directing each step. Less-restricted access for advanced cyber work is limited to vetted participants in the Daybreak program.&lt;/p&gt;

&lt;p&gt;This is the real tradeoff. Better memory, more tools, and longer autonomy make the system more useful. They also give it more time and more access to act on a bad assumption before someone catches it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run two evaluations, not one
&lt;/h2&gt;

&lt;p&gt;I would evaluate Astra in two passes.&lt;/p&gt;

&lt;p&gt;First, use a provider-neutral harness with the same prompts, tools, retry limits, and grading criteria you use for other models. That tells you whether Astra itself improves the work you care about.&lt;/p&gt;

&lt;p&gt;Then test the native OpenAI harness with persisted reasoning, compaction, searchable context, async tools, and steering enabled where appropriate. That tells you whether the complete Astra system is worth deploying.&lt;/p&gt;

&lt;p&gt;Use 20 to 50 real tasks, not polished demos. For long tasks, force at least one compaction and plant an early requirement that matters near the end. Interrupt a run with a changed constraint. Delay tool results and return them out of order. Put an approval boundary in front of an irreversible action. Measure accepted results, recovery from mistakes, latency, and total cost.&lt;/p&gt;

&lt;p&gt;The last point matters because &lt;a href="https://developers.openai.com/api/docs/models/gpt-6-astra" rel="noopener noreferrer"&gt;Astra costs $10 per million input tokens and $50 per million output tokens&lt;/a&gt;, before tool fees. If your application needs short answers or straightforward extraction, these harness features may not justify that price. If it needs sustained work across tools and changing requirements, token price alone is a poor way to judge it.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the way, this was a crowded release week
&lt;/h2&gt;

&lt;p&gt;Anthropic released &lt;a href="https://www.anthropic.com/claude/fable" rel="noopener noreferrer"&gt;Claude Fable 5.1&lt;/a&gt; for long-running coding and knowledge work. Meta released &lt;a href="https://research.meta.ai/blog/introducing-muse-spark-1-3" rel="noopener noreferrer"&gt;Muse Spark 1.3&lt;/a&gt;, with a focus on long-horizon collaboration, interruptions, and coding efficiency. Google released &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/" rel="noopener noreferrer"&gt;Gemini 3.8 Flash&lt;/a&gt; as a lower-cost agentic reasoning and coding model.&lt;/p&gt;

&lt;p&gt;All three belong on an evaluation list if they fit your workload. They do not change my conclusion about Astra: its most important advance is not the AGI label or the model name. It is the tighter integration between the model and the system that keeps it working.&lt;/p&gt;

&lt;p&gt;My recommendation is to stop treating the harness as an implementation detail. Compare models with a neutral harness, then evaluate the native system you would actually deploy. The &lt;a href="https://fewshotacademy.com/docs/intermediate/evaluating" rel="noopener noreferrer"&gt;evaluation chapter&lt;/a&gt; shows how to build the task set. Let the AGI debate continue somewhere else.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>infrastructure</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
