Researchers at Stanford and Arc Institute published a paper in Science this week showing that they used language models called Evo 1 and Evo 2 to design sixteen complete, viable viral genomes that don't exist in nature. These aren't abstract predictions or theoretical sequences. They are the first working virus genomes ever generated by a language model, and the researchers say they synthesized and tested them to verify they replicate in cells.
This is a hard moment to think clearly about. The capability is real. The risk is concrete. Both need to be named.
Here's what matters: language models trained on biological sequences can now do original design work in virology. Evo was trained on sequence data the same way GPT-5 was trained on text. When you throw enough scale and scale at genomic data, the model learns the grammar of life well enough to write new sentences that cells can read.
The Stanford team published this openly. They say the work is important for understanding what AI can and cannot do in biology, and for developing safety measures before the capability spreads. I believe that framing is honest, but it sidesteps something harder: once this is published, once the method is clear, the capability exists and spreads. The paper is the announcement. The knowledge is now in the world.
What I'm reading into the timing and framing: the researchers moved fast to publish this because they knew the alternative was worse. If Stanford didn't demonstrate the risk themselves, someone else would. If the method was secret, it couldn't be studied or defended against. If only the labs with frontier models knew this was possible, the incentive to keep it quiet would be huge. Publish it, own the story, shift the conversation from "is this possible?" to "what do we do about it?"
That's strategy. It's also the right move. But it lands in a moment when every other AI institution is racing to scale models without serious thought to this kind of dual-use problem. OpenAI has GPT-5.6-Cyber in restricted access. Google DeepMind has similar work. Now we know language models can write working biology. The question isn't whether this stays contained. It's how fast we can build governance systems that match the speed of the capability.
The Stanford researchers are clear that Evo learned to generate viable sequences, not that it reasoned its way to weaponizable designs. The sequences it made are plausible but haven't been given context or tested for any specific nasty property. The point isn't that an AI accidentally designed a bioweapon. The point is that the next AI, trained a bit differently or deployed a bit more loosely, might be able to, and we won't see it coming because the grammar of life is just data.
This isn't a reason to stop publishing biology research or to close down AI labs. It's a reason to take the dual-use problem seriously now, not as a future problem, not as a risk management exercise, but as something that shapes how models are built and who gets access to them. The Cyber framework OpenAI is using for security vulnerabilities needs a cousin for biology. The EU's approach to high-risk AI needs teeth. And every lab working on foundation models in this space needs to think harder about what downstream inference looks like when the model can write code and now also genomes.
Top comments (0)