DEV Community

Cover image for AI Dev Weekly #28: GPT-6.1 Sol, Claude Sonnet 5.5, Dots and NVIDIA OpenShell
Joske Vermeulen
Joske Vermeulen

Posted on Originally published at aimadetools.com

AI Dev Weekly #28: GPT-6.1 Sol, Claude Sonnet 5.5, Dots and NVIDIA OpenShell

AI Dev Weekly is a Thursday series where I cover the week's most important AI developer news, with my take as someone who actually uses these tools daily.

Last week brought a crowded model launch cycle. This week turned those models into systems. OpenAI replaced the two-week-old Sol baseline with GPT-6.1 Sol, added native delegation and browser execution to its agent stack, and introduced an always-on ChatGPT agent. Anthropic answered with a Sonnet that reaches much closer to Opus while keeping mainstream pricing. NVIDIA, meanwhile, pushed agent security below the prompt and application layers.

The common thread is operational. The most important questions are no longer only which model writes the best code. They are how long an agent can work, what it can access, what happens when it delegates, and whether its limits still hold after the model encounters hostile input.

1. GPT-6.1 Sol arrives one week after GPT-6 Sol

OpenAI released GPT-6.1 Sol on September 29 with the API ID gpt-6.1-sol. It targets complex coding, computer use, and professional work while keeping the standard GPT-6 Sol token rates: $2 per million input tokens, $0.10 cached input, $2.50 cache writes, and $10 output for prompts up to 272K input tokens.

The headline is capability at the same broad price point, but the API details matter more:

  • 1,050,000-token context window and 128,000 maximum output tokens
  • text and image input with text output
  • reasoning levels from low through max, with medium as the default
  • tool calling through the Responses API
  • native multi-agent delegation in beta
  • US and EU data residency, although Fast mode is unavailable with EU residency

OpenAI positions 6.1 Sol as delivering near-Astra performance at one fifth of Astra's standard input and output prices. That is a vendor claim and a reason to run evaluations, not a reason to skip them. It also does not make the original Sol obsolete overnight. The two models share list pricing and headline context limits, so the migration case depends on accepted-task quality, latency, and behavioral compatibility.

The official OpenAI changelog also added multi-agent support to 6.1 Sol. A Responses request can let the model delegate work to subagents instead of forcing the application to implement all orchestration itself. That simplifies some agent designs, but it also makes traces, budgets, and permission inheritance more important. A parent agent should not be able to turn a narrow permission into a broad one merely by creating another worker.

My take: GPT-6.1 Sol is the new model to evaluate first for serious OpenAI coding workflows, but the release cadence itself is a warning. Do not scatter a model ID across prompts, tests, and production code. Put it behind one configuration boundary, keep a rollback path, and rerun repository-level evaluations before moving traffic. Our GPT-6 Sol guide covers the shared API shape, while the AI model comparison helps frame cross-provider testing.

2. Claude Sonnet 5.5 brings Opus-like work closer to Sonnet pricing

Anthropic launched Claude Sonnet 5.5 on September 28 with the model ID claude-sonnet-5-5. Standard pricing remains $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 and cache writes at $2.50 per million tokens.

Anthropic says the model generates output more than 30% faster than Sonnet 5 and can cost up to 30% less per task because it uses fewer tokens. Its launch table reports 70.6% on Terminal-Bench 4.0 and 55.5% on CursorBench 4.0. Those are first-party results produced under specific harness and effort settings. They are useful for selecting candidates, not declaring a universal winner.

The migration is not a simple model-name swap. Sonnet 5.5 changes several integration contracts:

  • applications that previously ran with thinking off need the new between_tools setting;
  • forced tool selection can return an error and needs a fallback path;
  • preserved thinking is tied to the model and account context;
  • some computer-use tool versions differ by cloud provider;
  • text between tool calls can appear inside thinking blocks rather than ordinary text blocks.

Sonnet 5.5 defaults to Medium effort in Claude products and High on the Claude Platform. That distinction makes benchmark and cost comparisons easy to distort. Always record the effort level, harness, tool set, and stopping conditions with the result.

My take: Sonnet 5.5 is the strongest default candidate in Anthropic's lineup for routine coding and professional workflows. Opus 5.5 remains the escalation tier for ambiguous work that needs sustained judgment. The interesting comparison is now Sonnet 5.5 versus GPT-6.1 Sol because their standard input and output prices match. Test completed work, review time, and failure recovery, not isolated prompt answers. See the Claude Sonnet 5.5 guide and Sonnet 5.5 versus GPT-6 Sol for the current integration differences.

3. OpenAI Dots turn ChatGPT agents into an ongoing responsibility

OpenAI used DevDay on September 29 to introduce Dots, always-on agents in ChatGPT that can keep working between conversations. A dot gets its own cloud computer and browser, can use connected apps, and returns results or asks for judgment when it reaches a decision point.

That is meaningfully different from a saved prompt or a scheduled message. The product is designed around an ongoing responsibility, such as monitoring a market, preparing a recurring report, or coordinating research and implementation. OpenAI says Dots are powered by GPT-6 Astra and can delegate deeper work to Work or Codex.

Availability is deliberately limited at launch. The official DevDay overview describes a gradual rollout. Pro access excludes the European Economic Area, the United Kingdom, and Switzerland at launch. Business Premium is rolling out worldwide, while Enterprise access is a beta that administrators must enable.

For developers, the programmable counterpart also moved forward. The Agents API gained computer use in an OpenAI-hosted browser, including website access approvals and application-managed sign-in. This closes part of the gap between building a tool-calling agent and giving it a usable browser environment.

The difficult part is not keeping an agent awake. It is controlling state and authority over time. Long-running agents accumulate authenticated sessions, remembered context, files, retries, and opportunities to act on changed information. Teams need clear expiration rules, audit trails, spend limits, approval gates, and a way to stop work without leaving half-completed external actions behind.

My take: Dots are the clearest sign yet that the product unit is shifting from conversation to responsibility. Start with read-heavy work that produces a reviewable artifact. Do not begin with publishing, payments, deletion, production deployment, or unsupervised external messaging. Our OpenAI Dots explainer covers the product boundaries, and the Agents API comparison covers the developer-facing alternative.

4. NVIDIA moves agent safety outside the agent

NVIDIA announced the Open Agent Safety Platform on September 28. It combines OpenShell, an open-source secure runtime, with Sentry, an out-of-band reference design built around BlueField-4 DPUs.

OpenShell places an enforceable boundary around an agent process. It can govern filesystem, network, process, tool, and credential access while recording allow and deny decisions. The important architectural point is that the model does not enforce its own limits. An injected instruction can ask for a forbidden action, but the runtime still decides whether that action is possible.

Sentry adds a separate monitoring and enforcement plane. NVIDIA says it can quarantine an agent in milliseconds when behavior moves outside policy. That hardware path is distinct from OpenShell itself. OpenShell is open-source software and NVIDIA says it can extend to third-party Arm and Intel platforms, so evaluating the runtime does not require adopting the complete Sentry architecture.

The launch does not make autonomous agents safe by default. An overly broad allowlist is still overly broad, an approved action can still be wrong, and permitted package registries can still serve malicious content. Runtime enforcement must sit alongside isolated workspaces, scoped credentials, external logs, dependency controls, code review, and human approval.

My take: This may be the week's most consequential infrastructure release. Prompt rules express intent; runtime rules establish capability. Every team letting agents touch repositories, cloud accounts, or internal systems should be able to answer which layer actually blocks a forbidden file read or network request. The NVIDIA OpenShell guide and our broader agent sandboxing guide provide a practical evaluation path.

Quick hits

  • GPT-6 Astra gets Ultrafast mode: OpenAI added service_tier: "ultrafast" in the Responses API. It targets lower inter-token latency, but supports global processing and US data residency rather than EU or other regional inference residency.
  • GPT-6.1 Sol lands in Copilot: GitHub added the model on September 29. Availability through Copilot does not imply identical context, tool, billing, or reasoning behavior to the direct API.
  • Copilot chat gets more context: Slack can now use supported files, attachments, and message links. Teams adds inline images, forwarded-message context, and channel or thread history.
  • GPT-6 vision fix: OpenAI fixed an image-encoding bug in GPT-6 Sol and Luna on September 25. If you evaluated either model on image input before the fix, rerun those tests.

That's it for this week. The release race is no longer just model versus model. It is model, harness, runtime, permissions, and the evidence a human gets before approving the result.

Want this in your inbox? Subscribe to AI Dev Weekly.

Previous issue: AI Dev Weekly #27

FAQ

Is GPT-6.1 Sol more expensive than GPT-6 Sol?

No at standard list rates. Both are listed at $2 per million input tokens and $10 per million output tokens, although cache pricing and service tiers should be checked for the exact request shape.

Is Claude Sonnet 5.5 a drop-in replacement for Sonnet 5?

No. Thinking behavior, forced tool use, preserved thinking, response blocks, and some computer-use integrations changed. Run complete agent-loop regression tests before migration.

Can developers build with OpenAI Dots through an API?

Dots are a ChatGPT product experience. Developers who need programmable persistent agents should evaluate the Agents API, which now includes hosted browser computer use.

Does NVIDIA OpenShell require NVIDIA hardware?

No. OpenShell is the portable open-source runtime layer. NVIDIA Sentry is the separate out-of-band design that uses BlueField-4 hardware.

Related Articles

- Agent Sandboxing at Scale

Originally published at https://www.aimadetools.com

Top comments (0)