DEV Community

Kacper Włodarczyk
Kacper Włodarczyk

Posted on

AgenticOS: self-hosted AI agents your whole team can improve

AgenticOS: self-hosted AI agents your whole team can improve

Vstorm AgenticOS is an open-source, self-hosted platform for company AI agents. It lets you build your own AI agents in a browser, give them access to your documents, and choose the tools they can use, on infrastructure you control. This week it has a new README and two decks built on the real product screens, and this article is the longer version of them.

I'm Principal Engineer and Open Source Tech Lead at Vstorm, and I build AgenticOS, so read this as the maintainer's account rather than a review. The article was drafted with AI assistance. Every product fact in it was checked against the repository at release v0.0.516 or the README pull request that ships this week.

The problem it is built around

Most agent frameworks give you a library. You write Python, you deploy it, and every change to an agent's behaviour is a pull request, a review and a release. That is the right shape for a product feature. It is a poor shape for the many small agents a company actually wants, and the AgenticOS docs put the reason in one sentence: the person who knows what the agent should say is not the person with commit access.

So AgenticOS splits the work along that line. Engineers add capabilities in typed Python. The people who know the work write the instructions, upload the documents and maintain the procedures, in the browser. Colleagues use the published agent in web chat, Slack, Mattermost or Telegram, and developers call it through the API.

The promise on the README is the short form of that: AI agents your whole team can use and improve.

What is new this week, and what is not

There is no 0.1.0. AgenticOS is still 0.0.x: the public history starts on 28 July 2026, and releases v0.0.1 to v0.0.516 followed between 2 August and 1 October. Treat the interfaces as young.

What changed is how the project explains itself:

  • The README now opens with what AgenticOS is and who it is for, then what you can do with it, what ships today and where it stops.
  • A short introduction deck (14 slides at launch) covers the problem, the idea, the product, its controls and limits, and how to start. A longer product tour (44 slides at launch) goes through every screen in detail.
  • The README and the decks include a list of where it stops today. The parts that matter most for an evaluation are at the end of this article.

Building an agent in the browser

An agent starts as instructions, a model and a set of tools. You choose them in the builder, publish a version, then test it in chat. Earlier versions stay readable, and you can roll back.

Three kinds of company knowledge sit beside the agent:

  • Skills are reusable procedures, such as how to review code or write a report. You maintain one once and attach it to several agents.
  • Context is standing knowledge, such as a glossary or a policy. The agent gets it in the prompt or reads it on demand.
  • Knowledge bases make uploaded documents searchable, so an answer can come from your own files and cite the passage it used.

With a container sandbox configured and command execution switched on, an agent can also read and edit files, run shell commands and execute Python or JavaScript. If you have used Claude Code or Codex, that part will feel familiar. What AgenticOS adds is the shared workspace around it: company knowledge, published agents and organisation access controls. What any agent achieves still depends on its model, tools and instructions.

If you'd rather watch than read, the README has a 37-second recording of one run from 1 October: an agent reads a brief in Notion, researches our GitHub repositories and publishes an interactive page with its sources. The recording is edited, with the waiting cut out, and it runs on our own projects, so it is an integration example, not a customer story.

Three controls an IT or security reviewer will ask about

Putting agents in front of a whole team raises three questions before anyone asks about features. What can it spend? What can it do without asking? What happened yesterday?

The budget is checked before every model request

Each agent can have a monthly budget, and a new organisation starts with one ($100 by default). AgenticOS checks both before each model request, not after. The docs explain why: checking afterwards means the request that broke the budget was already paid for.

There is a limit worth knowing before you rely on it. A run's cost lands on its record when the run finishes, so runs that start together cannot see one another. Fifty parallel runs against a cap that is one call short of full can each read the same under-cap total and overshoot it by up to their combined cost. If you need a strict cap, run those agents through one queue.

Sensitive tools wait for a person

A tool that acts on the outside world parks the run and waits. A person sees the call, approves or rejects it, and the decision is recorded with who asked and who decided. You set this per tool or per capability. In the screenshot the same test account asked and decided. The permission to decide approvals is its own, and the built-in Operator role holds it without being able to edit agents.

The limit here is coverage. The gate applies to the platform's own capability tools. Tools from an MCP server (the Model Context Protocol, an open standard for connecting agents to external tools) pass through it only when a web-chat session is set to Ask about everything, and a tool that runs inside the model provider never reaches the local gate. Test the approval policy for the tools you actually enable.

Every run is recorded

Every run lands in run history with the agent version that executed, the tools it called, the tokens and the cost. Because the version is recorded, a run from last week still makes sense after someone rewrites the agent.

Governance actions go into a separate audit log: approval decisions, sharing and membership changes, secret rotation, exports and more. Each entry carries a hash over its own contents with the previous entry's hash folded in, per organisation, so editing, reordering, inserting or deleting an entry shows up when the chain is verified. That is detection, not prevention. It does not protect against someone holding the database's own credentials, who can rewrite a row and recompute the chain. And not every write is audited. Knowledge-base edits, for one, are not.

For developers: code defines, configuration composes

AgenticOS is built on Pydantic AI, the Python agent framework from the Pydantic team, and pydantic-ai-harness, a set of ready-made capabilities for Pydantic AI agents such as planning and tool-output limits. The backend is FastAPI with PostgreSQL (pgvector for the document search), Redis and Prefect for background and scheduled jobs, and the console is Next.js.

The built-in capabilities are typed Python that an engineer writes, tests and registers. The console can only switch on what code registered, which the docs call "code defines, configuration composes". That is the boundary that lets the person who owns an agent's instructions change them without a release.

MCP servers are the exception, and it is worth knowing before you choose a platform for a team. A builder or an admin connects one by pasting its URL, with no code change, and its tools sit outside the approval gate described above unless a chat asks about everything. Decide who holds the Builder role with that in mind.

Running it yourself

The quick start is a Docker Compose deployment through an installer script, on macOS, Linux or Windows under WSL2. You can run it with --check first to test the prerequisites without installing anything, or read the script, or follow the manual installation guide. You need access to a model: a hosted provider, or a local model through Ollama or an OpenAI-compatible server such as vLLM, both of which serve open-weight models on your own hardware.

Self-hosting the console does not make everything local. A hosted model, a cloud document parser or an external tool still receives the data you send it, so review the destinations you configure.

Two things from the repository that a self-hoster may find more useful than a feature list:

The first bottleneck was the database pool, not CPU. On 16 September we load-tested one laptop: one worker, a stub model instead of a paid provider, 4,740 requests across a ramp, a sustained rate, a burst and a recovery. With the default pool of 5 connections plus 10 overflow, 1,537 requests failed. With 20 plus 30, 13 failed. CPU peaked at about 90 percent of one core in both runs. The shipped defaults are still 5 plus 10, so if you expect parallel load, raise DB_POOL_SIZE and DB_MAX_OVERFLOW first. This is one machine and a stub model, a request rate rather than a user count, so it says where the first wall is, not what a production deployment can carry.

There is a HIPAA configuration profile, and it is not a certification. It covers the technical safeguards of 45 CFR §164.312 and nothing else: TLS to Postgres, Redis and the browser, a long vault key, a local model, local traces, single sign-on (OIDC or Kerberos), invite-only or closed signup, six years of audit retention and the audit chain. A doctor command checks the profile and exits non-zero when something is missing. The administrative and physical safeguards, encryption of content at rest and any business associate agreement remain the operator's job, and the profile says so.

What it does not do yet

This is the short list for an evaluation:

  • No ACL mirroring. Permissions from SharePoint or Google Drive are not copied per user into a knowledge base. Scope the credential and share the collection with the right group.
  • No SAML, no SCIM, no native MFA. OIDC, LDAP and Kerberos sign-in work. A SAML-only identity provider can come in through a broker such as Keycloak, and multi-factor is your identity provider's job.
  • No Kubernetes manifests. It is Docker Compose on one host today.
  • No reranker. Retrieval is vector search with options.
  • No visual workflow builder. It is in development and not in any release.
  • Budgets can be overshot by parallel runs, as described above.

The README and the decks add a few more, such as no built-in Microsoft 365 trigger.

Try it

Start with the introduction deck if you want the shape of it: https://vstorm-co.github.io/agenticos/presentation/

For every screen in detail, the product tour: https://vstorm-co.github.io/agenticos/presentation/tour/

Then the quick start and the document-assistant walkthrough, which uploads a handbook, asks questions and checks the answers against cited sources: https://github.com/vstorm-co/agenticos

It is Apache-2.0. Issues and pull requests are open, and I'd especially like to hear what would stop you from running it: a missing sign-in method, a deployment target, a control your security team would ask for.


I'm Kacper, Principal Engineer and Open Source Tech Lead at Vstorm. I build AI agents and open-source tools in the Pydantic AI ecosystem. Find me on LinkedIn or GitHub.

Top comments (0)