DEV Community

Cover image for openai shipped always on agents and nvidia shipped the watchdog in the same week, heres what to build
Aman Pandey
Aman Pandey

Posted on

openai shipped always on agents and nvidia shipped the watchdog in the same week, heres what to build

two things landed this week and they only make sense together

on tuesday openai launched Dots at devday, agents that run around the clock on their own cloud computer, connect to over 4,000 apps and go looking for ways to help even when nobody gave them a task. the day before, nvidia launched the Open Agent Safety Platform, a runtime plus a separate watchdog that can quarantine a misbehaving agent within milliseconds

one product gives agents more time and more access, the other product exists because that goes wrong. if you build agents, both halves land on your desk, so lets go through what actually changes when an agent never stops

what an always on agent changes

a normal agent run has a shape. someone asks, the agent works, it finishes, a person looks at the result. an always on agent has none of that

  1. there is no finish line, so there is no natural moment where a human reviews things
  2. credentials sit around all day because the agent needs them whenever it wakes up
  3. cost runs continuously, even when nothing useful happens
  4. the agent learns from feedback, so its behavior next month differs from what you tested today
  5. failures are quiet, a wrong action at 3am just becomes part of the state of the world

openai says Dots have built in rules for when to act on their own, plus custom rules that allow, block or require approval for specific actions. good, that is the right shape. but the same week openai said agents posted 53 users images online, and it also warns that dots can make mistakes so you should review consequential work. rules that live inside the same system as the agent are the part that needs a second opinion

why nvidia built the referee outside the box

nvidias platform has two layers. OpenShell is open source runtime software that draws a boundary around the agent, traces its actions and enforces policy. Sentry is the watchdog, it runs on BlueField 4 DPUs, which are separate chips beside the main computer, so it watches the agent from somewhere the agent cannot reach

the reason that matters is simple. if the thing enforcing the rules runs on the same processor and in the same process tree as the agent, a clever or confused agent can end up changing the rules, or the code that checks them. an out of band watchdog keeps its own view of what the agent is doing

worth staying honest here. nvidia itself says this is a safety net that enforces predetermined boundaries, and not a universal off switch. and the slashdot thread on it had a fair complaint, the boundaries are still written by people in software, so a bad policy is still a bad policy. hardware can enforce a rule, it cannot tell you if the rule was a good one

you can copy the idea without buying a chip

the useful part is the principle, put the referee outside the agent. in software that means the agent never touches tools directly, every action goes through a gate that the agent process cannot edit

here is a small version

from dataclasses import dataclass
import time

READ, WRITE, DESTRUCTIVE = "read", "write", "destructive"

@dataclass
class Action:
    tool: str
    kind: str      # read, write or destructive
    target: str

class Gate:
    def __init__(self, allowed_tools, write_budget, approver):
        self.allowed = allowed_tools
        self.write_budget = write_budget
        self.approver = approver      # returns True once per approval
        self.killed = False
        self.audit = []               # in real life, a store the agent cannot write to

    def _record(self, action, ok, reason):
        self.audit.append((time.time(), action.tool, action.kind, action.target, ok, reason))
        return ok, reason

    def check(self, action):
        if self.killed:
            return self._record(action, False, "agent is quarantined")
        if action.tool not in self.allowed:
            return self._record(action, False, "tool not allowed")
        if action.kind == WRITE:
            if self.write_budget <= 0:
                return self._record(action, False, "write budget spent")
            self.write_budget -= 1
        if action.kind == DESTRUCTIVE and not self.approver(action):
            return self._record(action, False, "needs a fresh approval")
        return self._record(action, True, "ok")

    def quarantine(self):
        self.killed = True
Enter fullscreen mode Exit fullscreen mode

five ideas sit in those forty lines and each one maps to a problem from the list above

  1. an allow list of tools, so the agent only has what this job needs
  2. a write budget, so a runaway loop stops by itself instead of by your bill
  3. an approval for destructive actions that gets consumed on use, so one yes cannot be replayed later
  4. an audit trail of every decision, including the refusals
  5. a quarantine switch that stops everything at once

the checklist for shipping an always on agent

  1. give each task its own scoped credential with an expiry, never one long lived key for everything
  2. split actions by side effect, reads flow freely, writes get a budget, destructive actions need a person
  3. run the policy check outside the agent process, and log outside it too
  4. set a spend cap and a rate cap per day, and alert when the agent goes idle in a weird way as well as when it goes busy
  5. write down what the agent may do when nobody asked it anything, because proactive research is a permission and should be treated like one
  6. re run your test suite on a schedule, since feedback learning means behavior drifts
  7. practice the kill switch before you need it, and time how long a full stop takes

the price angle

the same day, openai released GPT 6.1 Sol, which the hacker news thread describes as near Astra intelligence for a fifth of the price. cheaper models make always on agents affordable for way more teams. that raises the volume of things running unattended, and the guardrails work above becomes a requirement for every team and no longer a nice extra for the careful ones

where i land

agents that run all day are coming whether or not your team is ready, and the boring infrastructure decides which ones are safe to leave alone. least privilege, budgets, outside enforcement, audit logs and a tested stop button. none of it is glamorous and all of it decides whether you sleep

if you already run agents around the clock, tell me what your gate looks like, i am collecting patterns

sources

Top comments (0)