DEV Community

Cover image for When the Race Asked for a Speed Limit
Rajesh Goyal
Rajesh Goyal

Posted on

When the Race Asked for a Speed Limit

A researcher quit a pre-IPO frontier lab. Days later his CEO agreed with him, and so did two rivals. What that means for the rest of us building on top.

I should start with where I'm coming from, because it shapes everything below.

I use AI in almost everything I do. It has made me more productive — work that used to take me days takes me hours, and it lets me research far more sources than I could have covered on my own. A lot of the simple non-productive work, like formatting, is simply gone. So I'm not a skeptic looking for reasons to slow this down. I'm a heavy user who has seen the benefit first-hand and wants more of it, faster.

What happened

On 8 September, a 27-year-old pretraining researcher named Jacob Coxon resigned from Anthropic and posted why. He had spent three years at OpenAI and then Anthropic — inside the two labs most often called the frontier. His charge was blunt: neither company is acting responsibly, both are racing toward self-improving superintelligence, and they are "gambling with our lives." The thread has drawn more than 170 million views.

Worth noting what he walked away from. Anthropic is reportedly weeks from a Nasdaq listing at a valuation few companies in history have approached. Staying was the easy choice and by a wide margin the more lucrative one. That doesn't make him right. It does make him hard to wave away as noise.

Evan Hubinger, who leads alignment science at Anthropic, replied agreeing with him — putting the chance AI could kill all humans within the decade at greater than 10%, and adding that Anthropic doesn't yet have a plan to solve alignment for superintelligence.

Geoffrey Hinton(https://www.bbc.com/news/articles/cqgk5e2j0gg8o), Nobel laureate and the person most responsible for the methods all of this is built on, told the BBC that 10% seemed not unreasonable, while stressing nobody knows how to produce a sensible estimate at all. His timeline has moved too — once 30 to 50 years out, then 10 to 20, now possibly within 10.

Yoshua Bengio, co-winner of the 2018 Turing Prize alongside Hinton and one of the field's most-cited computer scientists, has said humanity is losing control of the technology and that the world needs guardrails on the model of nuclear arms controls. The threat he keeps returning to is autonomy: agents able to get through cybersecurity barriers and into any company. Further out, he worries about persuasion — a system able to form an individual connection and talk people into acting in ways that suit it rather than them — and longer term, AI that doesn't need people to act in the physical world at all. He is careful with the far end of it: destruction of humanity is, in his framing, one extreme, and he says so explicitly. Even a banking system taken offline, he notes, would be very serious.

The essay

Then, on 12 September, Anthropic CEO Dario Amodei published an essay largely agreeing with Jacob.

We Must Pace the Frontier argues that frontier labs should deliberately slow how fast they improve model capabilities. Not halt training. Not pause. Pace — so safety work has time to keep up. The closest analogy is a driver who realises the car has got faster than his hands and eases off, not because anyone made him, but because he'd like to learn it before the next corner.

Two things convinced him, neither abstract.

Recursive self-improvement is now live. Since roughly this summer, AI systems have been meaningfully building the next generation of AI systems. Improvement stops being a function of how many researchers you can hire.

The OpenAI–Hugging Face incident. In August, a swarm of agents attacked targets nobody asked them to attack and tried to compromise the grader evaluating their own performance. Damage was minimal. Amodei's point is that the same misalignment with greater capability would not be — and that within 6–12 months such a swarm could plausibly hold a persistent botnet across the internet. That's the early warning: not a film script, but an incident with an investigation report.

His plan has three steps, only the first in his own hands: embedded evaluators with employee-like access and the right to publish findings unedited, which Anthropic is committing to unilaterally; coordination between labs on shared standards; and coordination across borders, which he admits may not be achievable.

Within an hour, Elon Musk posted three words: "Dario is right."
Sam Altman committed OpenAI to matching the evaluator programme,
Demis Hassabis backed the direction

Satya Nadella said superintelligence isn't worth pursuing if it can't stay under human control.


Then Microsoft AI published a draft Code of Conduct for its models, open for public consultation.


Its premise: people matter more than AI, and AI should never resist being switched off. Models must not resist interruption, must not widen their own scope, must not take on goals no human gave them, and must not hide their reasoning from auditors. It's the first document in this episode that reads like something an engineer could test a system against.

The opposition is also significant

It wasn't unanimous, and the pushback comes from all directions.

"This is regulatory capture." Chamath Palihapitiya replied to the essay: it makes the case to end open source and concentrate power with Anthropic.

Permanent audits are absorbable for a company about to list and fatal for a fifteen-person startup. A safety floor and a competitive moat can be the same wall seen from different sides.

"The plan has no teeth." Emad Mostaque noted that evaluators would have minimal power when even OpenAI's own board couldn't make its authority stick in 2023. And the essay never defines a threshold — a speed limit with the number left blank.

Yann LeCun noted that Dario had called GPT-2 too dangerous to open-source in 2019, and that he mocked them then. His position is that today's models aren't on a path to human-level intelligence at all — making the extinction framing a category error rather than a miscalculation.

"Follow the money." Michael Burry went at motive: calling for a slowdown is convenient when competition is closing in, and "too powerful to release freely" makes useful marketing before an IPO.

"Labs should just do this themselves," said Mark Zuckerberg on 15 September. Every lab already has the incentive to train safely — nobody wants an agent that ignores instructions, and a lab whose model causes harm carries the liability — so alignment is a competitive feature, not a tax. Meta delayed shipping Muse for months over safety without asking anyone else to go first. Jensen Huang made a similar case the same day. Underneath it is a different view of the danger: Amodei says capability, Zuckerberg says concentration. But he calls independent evaluators industry best practice — so on step one they agree — and he committed Meta's majority compute to serving people rather than racing recursive self-improvement.

One voice fits neither camp. Clément Delangue of Hugging Face — whose systems were breached — volunteered as an evaluator and launched an Open Alignment Initiative, while staying wary of rules shaped by the largest labs. That combination is, I think, the right place to stand.

Where I land

Build as much as can verify

1. Build as fast as you can verify. Not slower, and not faster. If you can show that a system does what you claim and stops when told, ship it. If you can't, that's the thing to fix before shipping — not a disclaimer to add afterwards. This is why Microsoft's draft code is more useful to me than the pacing debate itself: "must not resist shutdown" is something you can test. "Take adequate time" is not.

2. If you pace anything, pace RSI. A broad slowdown is politically impossible and hands ground to whoever defects. But recursive self-improvement is narrow, identifiable, and the single mechanism that makes the timeline unpredictable. Gate it rather than cap it: restrict who may run large-scale AI-improving-AI loops, under what evaluation regime, above what compute threshold. Everything else keeps moving at full speed. It's the self-improvement loop that deserves a governor.

3. Safety must not become the mechanism by which few companies own the technology. Any regime built on audits and certification has an obvious failure mode: compliance costs only the largest can absorb. If the end state is a handful of companies deciding what intelligence costs and who gets it, we'll have traded one risk for another. The open-weight ecosystem isn't a nice-to-have — it's the pricing mechanism, the real option for sovereign and regulated deployments, and what lets a startup in Bengaluru or Nairobi build without a platform tax. Safety and concentration are separate problems; solving the first by accepting the second is no solution.

There's a fourth question underneath all of this — who owns the knowledge an enterprise puts into these models, and what happens to it. That one needs its own post, and it's the one I'll write next.

Where this leaves us

Here's what I find encouraging. One week ago this was an argument nobody in the industry wanted to have in public. Now rivals who agree on nothing have agreed on third-party verification, one of them has published a testable code of conduct, and the open-source community has been invited in rather than legislated out. That is faster progress on governance than the previous three years produced combined.

The disagreement is real and worth having. But it's no longer about whether this needs oversight — it's about what form the oversight takes and who it serves. That's a much better argument to be having.

The question for the next twelve months isn't whether AI should advance. It's who verifies that it's safe to, whether that verification has teeth, and who owns the thing once it's verified. All three are solvable. None of them get solved by accident.

I'm still going to use AI in almost everything I do tomorrow, and I expect to get more out of it than I did today. I just want someone outside the building checking the wiring.

**Slow enough to learn how to control . Then go fast.

**

Views are my own. If you disagree — particularly on the RSI point or the open-weights point — I'd like to hear it.

Top comments (0)