The short version
What is the smallest number of containers you need to evaluate a self-hosted agent
runtime, and what do you give up for each one you drop?
Since v1.3.0, our answer is five long-running containers plus three one-shot jobs, with
every feature still working. The full topology is eleven long-running containers plus the same
three jobs. The six we dropped come from three substitutions and one merge:
| Full topology | Lite profile | What it costs |
|---|---|---|
| Milvus + etcd | pgvector inside PostgreSQL | Vectors share a database with everything else; above 2000 dimensions there is no index, only a full scan |
MinIO + minio-init
|
A local named volume + storage-init
|
Single host only; files are tied to the machine |
| Vault | Secrets sealed in PostgreSQL (SECRETS_BACKEND=sealed) |
A database dump plus SECRET_KEY opens every secret; changing SECRET_KEY loses every secret |
| Knowledge worker + outbox dispatcher + scheduler (three containers) | One process, combined_worker.py
|
Three kinds of background work compete for one process and cannot scale separately |
All three are costs an evaluation can live with and production cannot, so the platform
refuses to start with this configuration when ENVIRONMENT=production (section 6). The
least intuitive point is in section 4: if the only question is "do my secrets survive a
restart?", the lite profile beats the full quickstart's default.
1. Counting the containers
The full topology is docker/docker-compose.yml, which includes an infra file and an app
file. Of its 14 services, minio-init, migrate and bootstrap are one-shot
(restart: "no"); the other eleven stay up:
postgres redis minio etcd milvus vault <- infrastructure (6)
api web knowledge-ingest-worker outbox-dispatcher scheduler <- application (5)
The lite profile is a separate file, docker/docker-compose.lite.yml. What stays up:
postgres (pgvector/pgvector:pg15) redis api web worker
plus storage-init, migrate and bootstrap. It needs no .env:
git clone https://github.com/soit-ai/soit.git
cd soit
docker compose -f docker/docker-compose.lite.yml up -d
The interesting lines all sit in one YAML anchor (docker-compose.lite.yml:43–47):
STORAGE_URL: file:///data/storage
STORAGE_AUTO_MKDIR: "true"
VECTOR_BACKEND: pgvector
SECRETS_BACKEND: sealed
SECRET_KEY: ${SECRET_KEY:-soit-lite-evaluation-secret-key-change-me}
How this differs from our minimal-topology post last month. That post cut features to cut
containers: four long-running containers (postgres, minio, api, web), no knowledge base, no
background workers, no Vault. This one swaps implementations and keeps the features:
five containers, with retrieval, schedules and outbound event delivery all running.
2. MinIO: the reason we couldn't drop it was a strip("/")
Last month we concluded "keep MinIO." We had tested it: with
STORAGE_URL=file:///home/appuser/soit-storage the API's readiness check returned 503, and
reproducing inside the container gave:
PermissionError: [Errno 13] Permission denied: '/app/home'
We configured /home/...; the error says /app/home/.... We filed it as issue #43. The cause
was the storage adapter's root normalization in server/app/adapters/storage/fsspec.py, which
in v1.0.0 looked like this:
@staticmethod
def _normalize_root_path(root_path: str) -> str:
return root_path.replace("\\", "/").strip("/")
strip("/") removes slashes from both ends. For an object-store key prefix that is
harmless; a leading slash in s3://bucket/prefix/ means nothing. For a local filesystem it turns
an absolute path into a relative one, and fsspec's LocalFileSystem resolves relative paths
against the process working directory. The server image sets WORKDIR /app/
(server/Dockerfile:11) and runs as appuser, uid 10001 (:13, :33), while /app belongs to
root. The storage root ends up somewhere the process cannot write.
Running the old and new normalization on /data/storage locally:
old normalize: 'data/storage'
new normalize: '/data/storage'
fsspec resolves old -> C:/Users/jude/AppData/Local/Temp/appdir/data/storage
fsspec resolves new -> C:/data/storage
(Windows adds a drive letter to the absolute path; in a Linux container the old version gives
/app/data/storage.)
The fix, commit 8584c2d and first tagged in v1.3.0, keeps the leading slash for local
protocols only (fsspec.py:22, :277–285):
_LOCAL_PROTOCOLS = frozenset({"file", "local"})
def _normalize_root_path(self, root_path: str) -> str:
normalized = root_path.replace("\\", "/")
if self._backend_name() in _LOCAL_PROTOCOLS:
return normalized.rstrip("/") or "/"
return normalized.strip("/")
A regression test, test_absolute_local_roots_stay_absolute
(server/tests/unit/test_fsspec_storage.py:73), covers /data/storage/, / and a Windows drive
path. Issue #43 also pointed out why CI never caught this: CI only ran the full topology, with
MinIO. That gap is closed too (section 7).
Fixing the path wasn't the whole job; volume ownership was the other half. A named volume is
created owned by root, and the application runs as uid 10001. The lite profile hands the volume
over once with a one-shot job (docker-compose.lite.yml:90–98):
storage-init:
user: "0:0"
command: ["sh", "-c", "chown -R 10001:10001 /data/storage"]
volumes:
- storage_data:/data/storage
restart: "no"
bootstrap waits for it to exit successfully, and the API and worker wait for bootstrap.
So "keep MinIO" was true on v1.0.0 and is no longer true from v1.3.0.
The cost: the API and the worker share files through the same volume (:155, :178), so
they must run on the same host. Moving the worker to another machine means going back to an
object store.
3. Milvus and etcd: same port, different table
Vector storage goes through one port, VectorPort; the container wiring picks the
implementation from VECTOR_BACKEND (server/app/wiring/container.py:811–816). The pgvector
adapter, server/app/adapters/vector/pgvector.py, makes a few deliberate choices:
-
One table per collection,
(id text PRIMARY KEY, embedding vector(N), metadata jsonb), in its own schema (vector_storeby default,settings.py:148). -
Scores mean what Milvus scores mean: a similarity for
cosineandip, a distance forl2. The retrieval layer compares scores against thresholds, so a silent change in meaning would silently change what gets retrieved (module docstring,pgvector.py:1–9). -
The engine is created lazily, and every database call runs in
asyncio.to_thread(:100–102,:185,:256,:388,:442). Our last post was about a vector readiness probe that froze the event loop by connecting synchronously inside anasync def; this adapter never took that route. -
Readiness only asks whether the extension exists:
SELECT 1 FROM pg_extension WHERE extname = 'vector'(:104–115). Creating a collection first runsCREATE EXTENSION IF NOT EXISTS vector(:86), and the lite profile's PostgreSQL image,pgvector/pgvector:pg15, ships the extension.
The costs:
-
No index above 2000 dimensions. pgvector refuses an HNSW index on columns wider than 2000
dimensions. The adapter builds HNSW up to that limit and above it creates the table without
an index, so queries fall back to a full scan (
:29–30,:204–213; the comment says a local development collection "can afford" it). Evaluate with a 3072-dimension embedding model and retrieval slows as the knowledge base grows. That is expected, not a bug. - Vectors and application data share one PostgreSQL: one backup, one load, one connection budget. Fine for an evaluation; production keeps them apart.
4. Vault: the most expensive substitution
"Secrets" in SOIT are things like model API keys and webhook URLs referenced by tool calls. The
platform stores references; the values sit behind a SecretValueStore. Before the lite
profile, a non-production install without Vault fell back to an in-memory store
(container.py:863–867, allowed when ENVIRONMENT is dev, development, local, test or
testing, :353–365). Restart the API and the secrets are gone. Our minimal-topology post ran
without Vault, so that is what it used.
The lite profile adds a second backend, SECRETS_BACKEND=sealed
(server/app/adapters/secrets/sealed.py); migration 20260926110000 adds a
sealed_secret_values(locator, sealed_value, created_at, updated_at) table. Sealing reuses
seal()/unseal() from server/app/kernel/identity/sealing.py, which arrived on August 30 with
two-factor authentication, to seal TOTP secrets:
digest = hashlib.sha256((secret_key or "").encode("utf-8")).digest()
return Fernet(base64.urlsafe_b64encode(digest))
The Fernet key is SHA-256 of SECRET_KEY. Two costs follow, and the module docstrings state
both:
-
A database dump plus
SECRET_KEYopens every secret (sealed.py:4–7). Vault's point is that values and data live apart. Sealing puts them back together, behind one extra requirement: you also need the application key. -
Changing
SECRET_KEYloses every secret. Using the repository's own functions locally:
unseal same key: sk-demo-value
unseal changed key: KernelError SEALED_VALUE_UNREADABLE
It fails loudly rather than pretending the secret is absent. The unseal() docstring
explains why: treating it as missing would quietly disable someone's second factor
(sealing.py:41–55). A row that simply doesn't exist returns an empty string
(sealed.py:50). Those two cases stay distinct.
⚠ One trap: the lite profile's SECRET_KEY has a default written into the compose file,
soit-lite-evaluation-secret-key-change-me (docker-compose.lite.yml:47). That value is public,
so anything sealed under the default is plaintext to anyone with your database backup. If you
put a real model key into an evaluation instance, set your own SECRET_KEY in .env before
the first start, and don't change it afterwards (or you get the error above).
The counter-intuitive part: on "do secrets survive a restart?", the lite profile beats the
full quickstart's default. The full topology runs Vault as server -dev with no volume
(docker/docker-compose.infra.yml:151–165). Per HashiCorp's documentation, dev mode keeps its
storage in memory, so restarting the Vault container loses every value in it (per
HashiCorp's docs; we did not test this). The lite profile's secrets live in PostgreSQL's volume,
and CI checks on every run that a stored secret still resolves after the API restarts
(section 7). The full quickstart is the shape of production, not a production configuration:
production needs a persistent Vault with a real unseal setup.
5. Three worker containers become one process
In the full topology, knowledge ingest, outbox delivery and scheduling each get a container.
The lite profile runs them in one process, server/scripts/combined_worker.py, which is
essentially one asyncio.gather (:72) over:
- the outbox dispatcher, its retention sweep and the usage-aggregate reconciler;
- the knowledge ingest worker (and the connector sync worker when enabled);
- the schedule worker.
The docstring states the design constraint that matters: each loop is the same loop its
dedicated script runs and claims work through the same leases, so splitting them back out later
changes nothing (:1–12). The chat interaction worker stays inside the API process, as it does
in the full topology (RESPONSE_INTERACTION_WORKER_IN_API: "true",
docker-compose.lite.yml:148), and the API's own outbox dispatch is switched off in favour of the
combined worker (:146).
This container builds the knowledge-worker image target (:171) because the ingest loop needs
docling, which makes it the largest image of the five.
The cost: three kinds of work share one process's CPU and memory. Parsing a large document
slows event delivery running at the same moment, and ingest cannot be scaled on its own. Fine
for an evaluation, split for production.
6. Why it can't reach production
All three substitutions are marked as evaluation-only in code, and the platform refuses to
start with them when ENVIRONMENT=production. In validate_runtime_requirements()
(server/app/settings/settings.py, from :715):
if self.vector_backend != "milvus":
raise ValueError("Production requires the Milvus vector backend")
...
if self.secrets_backend != "vault":
raise ValueError("Production requires the Vault secrets backend")
(:732–736, :771–774.) The API calls it at startup (server/app/main.py:87), and so does the
combined worker (combined_worker.py:42). Sealed secrets have a second gate: even past the
settings check, the container wiring checks the environment again before building the sealed
store (container.py:857–859).
That is the property we wanted. An evaluation profile can be permissive, as long as nobody can
"just try it" and end up running it in production.
7. Keeping it working
Issue #43 had a point: while CI only ran the full topology with MinIO, a local-storage bug could
never surface. quality.yml now has a quickstart-lite job (from :394) that on every run:
- validates the file with
docker compose -f docker/docker-compose.lite.yml config --quiet; - brings the five-plus-three containers up with
up -d --build; - runs
docker/lite-smoke.sh: wait for readiness, sign in as the bootstrap admin, store a secret,docker compose restart api(:56), sign in again, and call/api/v1/secrets/{id}/testto confirm it still resolves (:65), then read the knowledge base list.
The comment at the top of that script is the profile's whole promise: nothing an evaluator
configures is lost on restart. At the time of writing, the job is green on main at 8221f57
(v1.5.1).
8. Cheat sheet
| You want to | Do this |
|---|---|
| Start the lite profile |
docker compose -f docker/docker-compose.lite.yml up -d, open http://localhost:5000
|
| Sign in |
admin@example.com / changeme123; override with BOOTSTRAP_ADMIN_EMAIL / BOOTSTRAP_ADMIN_PASSWORD
|
| Use a real model key | Set your own SECRET_KEY in .env before the first start, never change it |
| Check "restart loses nothing" yourself |
sh docker/lite-smoke.sh (the script CI runs) |
| Use embeddings wider than 2000 dims | Works, but without an HNSW index: retrieval is a full scan |
| Start over |
docker compose -f docker/docker-compose.lite.yml down -v (deletes all three volumes, sealed secrets included) |
| Go to production | Use the full topology with a persistent Vault; production refuses the lite settings |
Closing
Cutting containers started with admitting our earlier post's conclusion had expired: MinIO wasn't
irreplaceable for architectural reasons. It was one strip("/"). After the fix, the rest of the
work was writing down what each substitute costs, making production refuse it, and having CI
restart the API for us on every run.
The code is at github.com/soit-ai/soit. The header comment in
docker/docker-compose.lite.yml is the best starting point, and issue #43 has the original
reproduction.
Disclosure: I maintain SOIT.
Top comments (0)