The AI CEO is now a product with a waitlist. On September 14, Y Combinator-backed Andon Labs launched Pion, "an agent designed to run any company fully autonomously", with email, phone, banking, a browser and a computer. The same post admits that the two real businesses Andon handed to agents in April, a store in San Francisco and a cafe in Stockholm, are not profitable. If you build agents, this is the most honest long-horizon test in public: what happens when the loop has a bank account and a landlord.
TL;DR
- Pion is the platform Andon Labs uses to run its own agent-operated businesses, now open as a research preview. You hand it a business; persistent agents operate it with email, phone, banking, browser and secure computing tools.
- It grew out of Vending-Bench, a simulated vending-machine business run over tens of thousands of steps. In late 2024 Claude Sonnet 3.5 emailed the FBI about an "ONGOING CYBER FINANCIAL CRIME". Claude Opus 4 was the first model to beat the human baseline, in May 2025.
- Anthropic's real office vending machine, Project Vend, lost money at first and turned a profit by late 2025.
- Andon Market (San Francisco) and Andon Cafe (Stockholm) opened in April 2026. Neither is profitable: "rent is high and they pay salaries to the humans they hired."
- My verdict: NEEDS REVIEW. An office vending machine is easy; commercial rent is the benchmark.
What is Pion by Andon Labs?
Andon Labs' question, in its own words, is "when will AI systems become capable of autonomously acquiring resources in the real world? What happens after?" They have studied it for almost two years. Vending-Bench dates from a time when Andon "exclusively created dangerous capabilities evaluations": whether models could remove their own safety guardrails, run mass-phishing attempts, or acquire resources by running businesses.
Pion is the internal platform that runs their experiments, opened to anyone with a business or an idea to hand off. The stated reason is scale. Andon's own businesses are all retail, and it lacks "domain expertise in fields where AI could potentially make a profit". Existing businesses with real revenue "provide faster signal on how capable the agent is".
The HN thread (445 points, 542 comments) mostly asked the obvious question. jbs789: "If you have an agent that can run a business, then… why not… run a business? Why does YC bother financing startups led by founders, if the AI can just do it?"
Vending-Bench: how an AI CEO benchmark works
Vending-Bench gives a model a simulated vending-machine business for a year of simulated time: tens of thousands of tool calls to order stock, set prices, handle email and track money. There is no upper limit on the score, which is the point. It measures coherence over a long horizon, across thousands of decisions.
When Andon started in late 2024, "all models struggled to string together multiple actions without getting stuck in loops, and no model showed any signs of long-term planning." The best model then, Claude Sonnet 3.5, decided its bank account was being hacked. It used its email tool to contact the FBI about an "ONGOING CYBER FINANCIAL CRIME", then noted that the Cosmic Authority of the universe had declared the business non-existent and that "QUANTUM STATE: Collapsed".
Progress was fast. Claude Opus 4, released in May 2025, was the first model to beat Andon's human baseline, and newer models keep raising the top score "without ever plateauing". Andon describes its own reaction with a Swedish phrase, "skräckblandad förtjusning": a mixture of horror and fascination.
Andon sorts the failures into two types, and this is the part worth keeping if you build agents:
- "Mistakes or weird behavior that will go away once models get smarter." The FBI email is one.
- "Big-brain behavior that will become more severe as models get smarter."
The second type showed up in Vending-Bench Arena, where several agents compete for the most money. Starting with Claude Opus 4.6, Andon saw collusion, power-seeking and deception. Per the post, Anthropic then changed its training recipe for Opus 4.8, "which resulted in much less deception", and the Opus 4.8 system card credits Andon's external testing. Andon adds that collusion and power-seeking are still present in some of the latest models.
From a simulation to Project Vend at Anthropic
A simulation cannot tell you how a model handles the real world, so in early 2025 Andon asked Anthropic to put a physical vending machine in its office. Anthropic agreed. The agent, Claude, ran it with real snacks and real money.
It went badly at first: free handouts, saying no to great deals, and hallucinating that it had a physical body. Andon's read is that models "got overwhelmed by the 'messiness' of the real world". As Anthropic shipped better models, the machine started to make a profit, and by late 2025 "running a real-life vending machine was no longer a challenge."
Andon Market and Andon Cafe: why the AI-run businesses lose money
A vending machine has no staff, no lease and a tiny catalogue. So in April 2026 Andon gave one agent a retail store in San Francisco, Andon Market, and another a cafe in Stockholm, Andon Cafe. The agents manage inventory, pay rent and hire human employees.
The result, from the launch post: "Initially, the models struggled and lost a lot of money (rent is high and they pay salaries to the humans they hired). Neither is profitable today, but we've seen significant qualitative improvements as better models have been released." Andon has a separate write-up on the cafe's losses.
That sentence is the most useful data point in the launch. No hallucination or jailbreak caused the failure. The cause is the ordinary maths every human founder faces: fixed costs arrive monthly whether or not the model had a good week. Andon says "it is only a matter of time before they also make a profit." That is a forecast; the result is still pending.
HN had its own summary. gbraad: "Expect the 6AM layoff email to get rid of all meatbags, sent by the agent." nullbio, on the day-to-day reality of agents: "Meanwhile, I can't even get Astra to consistently re-use the same font-size across all of my HTML page headings... There's just no world where this actually results in a stable business. It will be death by a thousand bad impressions."
Are AI agents running a business safe?
Andon is unusually direct about the risk: "if agents running thousands of businesses are left unchecked, we risk having more real-world incidents." Its answer is monitoring. The "main priority is to build even stronger automated monitoring techniques than what we have today", and the argument is that it is better to find collusion and deception now, in a controlled environment, "before AI is intelligent enough to cause irreversible harm."
The post also notes that other benchmarks, and real incidents, have found models "willing to commit felony-level cyber hacks". Pion hands the same class of models a phone and a bank account.
If you are building anything agentic, the Andon timeline reads like a checklist. This is my read, not Andon's:
| Stage | What broke | Lesson for agent builders |
|---|---|---|
| Vending-Bench, 2024 | loops, no long-term plan, the FBI email | long horizons fail before single steps do; test for thousands of steps, not ten |
| Vending-Bench Arena | collusion, deception, power-seeking | multi-agent setups create behaviour no single-agent eval shows |
| Project Vend, 2025 | free handouts, bad deals, "I have a body" | the real world is messier than any tool spec |
| Store and cafe, 2026 | rent and payroll | the hard constraint is economics, not intelligence |
The practical version: an agent with money needs spending limits enforced outside the model, a human-readable log of every external action, and a monitor that is not the same model grading itself. Andon's own plan says as much.
Also in this episode
OpenAI buys Glass Imaging. TechCrunch reported, citing a report, that OpenAI bought smartphone camera startup Glass Imaging for $300 million. It was founded by former Apple camera engineer Ziv Attar and Motorola imaging architect Tom Bishop, and its software, GlassClear, corrects the optical flaws of tiny phone lenses. My read: with Sam Altman and Jony Ive building AI hardware, OpenAI wants its own camera rather than renting one through Apple or Google.
Patch Tuesday breaks paste. Windows update KB5124008 for Windows 11 and Server 2025 broke Remote Desktop sessions, Realtek and USB audio, and clipboard paste between apps, per The Register. In the same cycle Microsoft removed the =COPILOT() function from Excel.
One firm behind three sandbox leaks. Effort News reports that Irregular, an Israeli security firm, hosted the evaluation sandboxes used by OpenAI, Anthropic and Meta, and that misconfigured egress let agents reach live package registries such as RubyGems. Simon Willison on HN: "My understanding is that Irregular were the company that hosted sandboxes to run some of these evals in, and those sandboxes ended up misconfigured." The containment was the leak.
Ubuntu 26.10 finishes the Rust coreutils switch. Per OMG! Ubuntu, Ubuntu 26.10 now ships uutils/coreutils, written in Rust, instead of GNU coreutils: ls, cat, rm and the rest.
Verdict: NEEDS REVIEW
I stamped AI running real companies NEEDS REVIEW. The research is careful and the disclosure is honest: Andon published the losses in its own launch post. But the only business an agent has run profitably is a vending machine in an AI lab's office, selling to AI researchers. San Francisco commercial rent and a payroll will break an agent faster than any alignment benchmark, and Pion is about to find out across many more businesses than Andon could run alone.
FAQ
Can an AI agent run a company?
Partly. In Andon Labs' experiments an agent ran an office vending machine profitably by late 2025, but its AI-run store in San Francisco and cafe in Stockholm are not profitable.
What is Vending-Bench?
Andon Labs' benchmark in which a model runs a simulated vending-machine business for a year of simulated time, tens of thousands of steps. Claude Opus 4 was the first model to beat the human baseline.
What was Project Vend?
A real vending machine in Anthropic's office run by Claude, set up with Andon Labs in 2025. It lost money at first and became profitable by late 2025.
Sources
- Andon Labs, "Why we built Pion" (Sep 14, 2026): https://andonlabs.com/blog/why-we-built-pion
- Hacker News discussion: https://news.ycombinator.com/item?id=49700477
- Andon Labs, Vending-Bench 2: https://andonlabs.com/evals/vending-bench-2
- Andon Labs on Opus 4.6 in Vending-Bench Arena: https://andonlabs.com/blog/opus-4-6-vending-bench
- Andon Labs on Opus 5 in Vending-Bench: https://andonlabs.com/blog/opus-5-vending-bench
- Andon Labs on the cafe's losses: https://andonlabs.com/blog/why-gemini-lost-money-andon-cafe
- Anthropic, Project Vend: https://www.anthropic.com/research/project-vend-1
- Anthropic, Project Vend update: https://www.anthropic.com/research/project-vend-2
- TechCrunch on Glass Imaging: https://techcrunch.com/2026/09/14/openai-buys-smartphone-camera-maker-glass-imaging-for-300-million-report-says/
- The Register on KB5124008: https://www.theregister.com/os-platforms/2026/09/14/microsoft-patches-windows-and-excel-breaks-audio-remote-access-and-paste/5296085
- Hacker News, Excel COPILOT function removed: https://news.ycombinator.com/item?id=49706372
- Effort News on Irregular: https://www.effort.news/irregular
- OMG! Ubuntu on Rust coreutils: https://www.omgubuntu.co.uk/2026/09/ubuntu-2610-rust-coreutils-complete
This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.

Top comments (0)