DEV Community

Evgenii Arsentev
Evgenii Arsentev

Posted on Originally published at arsentev.ai

My company has no full-time developers, DevOps or designers. Here's what three $200 Claude subscriptions do instead

At night the agent shifts run by the clock. New builds of educational games get built, and a separate fresh agent checks each one by screenshots. The mail agent of one business line starts every hour, night too. In the morning the results are ready. In the businesses I run there are no full-time developers, no DevOps, no designers, no analysts and no product manager. At US median pay a developer, a product manager, an analyst and a designer come to about $38,500 a month (Bureau of Labor Statistics, see the note at the end). The agents doing that work run on three $200 Claude subscriptions, so $600 a month (about 64 times less).

There is me, there is a product builder, and there are AI agents that run on a schedule on separate servers (not on my laptop). I'm a medical doctor by education and a CEO. This is not a forecast, it's what runs today, what I kept for people and what didn't work.

The order we did it in

Designers went first, as soon as agents got good enough at design. Analysts went a long time ago too. The product manager role is closed too, and Claude does its tasks now (tasks that took days now take minutes). The developer was the last one.

We kept our former CTO on a small retainer as a safety net, in case an agent breaks something serious. So far we haven't needed him once. Everything our DevOps used to do, the agents do.

What is on the schedule

Each business line has its own agent shift. It starts by the clock, does its job and stops. I only start things, the agents do the rest.

Deploys. Releases to all our projects and zones are done by agents.

Our own CRM. We didn't buy one. The agents built a CRM from scratch, and the agents also work inside it. So why buy a CRM if we built our own?

Email. Every business line has its own mail shift, and the agent sorts the mailbox of that line only.

Meetings. Scheduling meetings is on the agents too.

Technical support. Technical support is done by agents.

Content. On a news and learning site I founded, agents publish news four times a day and scan new open-source AI releases three times a day. The games project I mentioned works overnight, and the builder's opinion about its own work doesn't count - only the fresh agent's check does.

And not everything is an agent. Backups for example are plain scripts on a timer.

The supervisor of the supervisor

This is the part I like most. In one of our US business lines we have a call-center team. The operators and their supervisor are real people. But the supervisor gets tasks from an agent.

Every day the agent goes through the operators' work - the calls, their transcripts, the results in the CRM. Then it writes to the human supervisor what to do today and tomorrow (with a number and a deadline), who needs help, who should be replaced and what the plan is for each person. Each operator gets a separate message with 3-4 small skills to fix, taken from the transcripts of their last calls. Every evening each skill gets a ✅ or ❌, and after three ✅ in a row the skill is closed.

The agent asks three things in every shift. Is each operator better or worse than last time? Did the supervisor improve the work of their people (by the people's numbers, not their own)? What should the supervisor do today and tomorrow? If the team is stuck, that is written as the supervisor's failure. The agent only suggests. Decisions about people are made by people and are not posted in the team channel, and the operators were told openly that calls are checked every day.

Subscription, not tokens

The agent work runs on subscription, not on paying per token through the API. On a flat price the students on my course experiment more boldly. The limits get eaten by re-reading context - in my own measurement more than 85% of the modeled spend was context work (DOI 10.5281/zenodo.22759216). And every shift is short, it does one thing and exits.

Where people stayed

An agent can do the work, but it can't be accountable for it. Every agent workflow has a named person who owns the result and answers for it as if a colleague had done the work.

Checks stay mandatory. Agents don't show doubt (a wrong result comes with the same confidence as a correct one). Nothing ships just because an agent finished it. It ships because it passed a check. And an agent gets access for a task, not a standing role.

What didn't work for us

First, vague tasks. A vague task given to a person gets clarified over coffee. Given to an agent, it gets done confidently and wrong. So the main management skill for me now is writing the task - the goal, the limits, what "done" means and how we check it.

Second, measuring the wrong thing. The first AI judge of calls in our CRM scored them against a script checklist. It measured if the operator followed the script. I think we should judge an operator by whether the partner gets to a purchase, not by the checklist.

Third, screens nobody opens. Our CRM has supervisor pages, and in 30 days they were opened 9 times. What works is the plain daily message to each person.

Fourth, AI summaries garbled a supervisor's name. So for names we trust only the spelling a person confirmed.

A case I documented

A company whose case I documented replaced a video studio with Claude agents - 1,497 videos across 67 channels in 101 days, at $2.88 per video. With people it would take 42 to 78 staff and cost 138 to 935 times the AI subscriptions (DOI 10.5281/zenodo.22802229). It's not our company, and I only wrote the analysis.

What I would tell another CEO

Just start with one function where the result is easy to check. Name a person who owns the result and track the cost per result from the first week. I look at the cost of a task that was done and accepted (with reruns and review time), not the price of the tool.

Keep the shifts short. But don't clear the context after every single task either. In my own small experiment that was about a third more expensive than clearing it every three tasks.

Where the limit moved

Before, we were limited by people and the job market. Now the limit is Claude tokens and computing power - memory and GPUs. That's why I rent separate servers, including GPU machines, so the agents don't run on a laptop that chokes.

I think memory is becoming the new gold. And I think for everyday agent work we'll need something like what ASICs did for bitcoin, a class of chips made for this kind of load.

Note on the numbers: US median annual pay from the U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025 (bls.gov/oes) - software developers (15-1252) $135,980, project management specialists (13-1082, the closest BLS category to a product manager) $102,320, data scientists (15-2051) $120,230, web and digital interface designers (15-1255) $104,000. Monthly = annual / 12. These are market numbers, not our salaries.


Evgenii Arsentev, MD, PhD — CEO, digital health; founder, arsentev.ai
Evgenii Arsentev is a medical doctor and a digital health CEO whose businesses run almost entirely on AI agents. He is the founder of arsentev.ai.

Originally published at arsentev.ai.

Top comments (1)

Collapse
 
hamid_ahmadian_3570449f72 profile image
Hamid Ahmadian •

The "vague tasks get done confidently and wrong" point is the one people underestimate most. A person who gets a fuzzy task will usually go ask a clarifying question before burning an afternoon on it; an agent will just pick the most plausible interpretation and execute it with full confidence, and you don't find out it guessed wrong until you review the output. The other detail I'd underline: "an agent gets access for a task, not a standing role" — that's a real security posture, not just a phrasing choice. Most of the horror stories I've seen (agent wipes the wrong table, agent emails the wrong list) trace back to a standing credential that outlived the one task it was scoped for, not a bad model decision in the moment. Curious how you handle credential lifetime across the overnight shifts specifically — do the agents get fresh short-lived access per run, or a longer-lived key that's just narrowly scoped?