GPT-5.6 Sol Is Running Quantum Computing Experiments While Researchers Sleep
AI agents just moved from writing code to operating physical hardware — and the implications are enormous.
What Happened
OpenAI published a fascinating case study today: MIT's Engineering Quantum Systems (EQuS) group is now using GPT-5.6 Sol — harnessed to Codex — to autonomously run superconducting qubit experiments.
This isn't AI suggesting code snippets. This is AI operating physical quantum hardware, making decisions about what measurements to run next, analyzing results, and adapting its approach — sometimes for hours overnight, without human supervision.
The Setup
Superconducting qubits are fabricated on chips, cooled to near absolute zero in dilution refrigerators, and controlled entirely through software via microwave pulses. Once a chip is built, every interaction with it happens through code — which makes it a perfect testbed for AI agents.
Beatriz Yankelevich, a graduate student in the EQuS group, connected Codex to the lab's experiment coordination software. She gave it measurement-specific skills explaining how to run and evaluate each experiment, then let it work.
What GPT-5.6 Sol Can Do
On a standard six-qubit chip used for benchmarking fabrication:
- Select measurement parameters based on design targets
- Operate the hardware directly
- Analyze resulting data and identify qubit transition frequencies
- Calibrate control and readout pulses
- Measure coherence times (how long qubits retain quantum information)
- Chain interdependent measurements — each result informing the next step
When signals were clear, Codex completed standard measurement sequences with minimal researcher intervention. The group now regularly uses agents for routine characterization work that used to take researchers several days per chip.
Where It Struggles
The honest limitations are just as interesting:
- Weak or noisy signals — when experimental data is ambiguous, the agent takes longer and sometimes needs human guidance
- Interpreting unexpected physical behavior — experienced researchers can recognize when something unusual is happening; current agents can't always adapt
- Novel situations outside training — clearly defined workflows work well; genuinely new experimental scenarios remain a challenge
Why This Matters
This is a meaningful threshold. AI agents have been writing code, analyzing data, and generating text for years. But operating physical scientific instruments autonomously is a different category of capability.
The EQuS group isn't replacing researchers — they're multiplying them. One graduate student can now run characterization experiments on multiple chips in parallel, with agents handling routine measurements while humans focus on experimental design, analysis, and the creative work of science.
It's a preview of what AI-assisted research looks like at scale: not AI doing science instead of humans, but AI handling the tedious, repetitive measurement work that eats up so much lab time.
The Bigger Picture
This comes the same week as:
- Navier-Stokes controversy — OpenAI spending $22.5M in compute to beat academics to a Millennium Prize proof
- AlphaGenome Atlas — DeepMind pre-calculating all 9 billion possible DNA mutations
- Mercury 2.5 — diffusion LLMs hitting 1,107 tokens/second
The pattern is clear: AI is moving from generating content to doing work. Not just writing about experiments — running them. Not just suggesting hypotheses — testing them.
The labs that figure out how to integrate AI agents into their experimental workflows first will have a massive advantage. MIT's EQuS group just showed us what that looks like.
What do you think — would you trust an AI agent to run your lab experiments overnight? Drop a comment below. 👇
Sources:
Top comments (0)