v0.5.0 is the first feature release since v0.4.0: 30 commits since v0.4.3, 19 of them changing product code. Two themes run through it. Knowledge in motion — a package you build can be mounted straight into an MCP host as a one-line config entry, or stacked with other packages into a single database your agent queries. And measurement — two benchmark boards that score local models pass/fail without a judge model, a /leaderboard page that blends them into one Overall score, and model pickers in the app that read the same numbers.
chaoscypher mount — a package, served over MCP, in one command
The gap this closes: you build a knowledge graph on one machine, export it as a .ccx package, and then have to write a setup script on the other machine before an assistant can use it.
chaoscypher mount is one idempotent command. It resolves a .ccx — a local file, or a Lexicon Hub owner/name — imports it into a database of its own (templates, knowledge, workflows and sources, chunks and citations), indexes it for search, and hands that database to the stdio MCP server. A mount.json SHA-256 marker makes every later launch skip the import, which is the property that matters: the command belongs directly in an MCP host's config rather than in a setup script that runs once and is never thought about again.
It is first-run-safe. No chat or extraction model is needed on the serving side, because the host's model asks the questions. All output goes to stderr, since stdout is the JSON-RPC channel.
In plain English: you can paste one command into Claude Desktop's config and the graph is there, with its citations, on every launch.
chaoscypher compose now works on CCX 3.0
compose was a pre-3.0 prototype, and it failed quietly in three separate ways: the merger read a data/*.json layout that no package has, so every build produced 0 entities; up launched a Cortex entrypoint that exited 0 at import time; and down could not stop a detached server.
The merger now imports each resolved package into a real runtime database through CcxImporter — templates, knowledge, sources, chunks, citations and IRIs intact — counts totals from that database, records its inputs in composition.json, and runs a best-effort embedding and vector indexing pass afterwards. up launches Cortex the way serve does, logs to <output>/server.log, watches the child through a startup grace period and fails with the log tail if it dies. --detach records compose.pid and refuses to stack a second live server. down does SIGTERM then SIGKILL with the configured grace, cleans stale records, and says "nothing running" instead of claiming success.
Two new subcommands come with it. compose init [--name] [PACKAGES…] writes a starter axiomatize.yaml and refuses to overwrite one without --force. compose mcp builds the composition if its database is missing, then replaces the process with chaoscypher mcp on the composed database — stdio passes straight through and the host sees one pid.
In plain English: several knowledge packages, one database, one server your agent queries — and the commands now tell you the truth about whether it is running.
Citations into audio and video now carry timestamps
The audio and video loaders joined Whisper's transcript and threw the segment times away, so a citation into a recording could say which file but never when. Chunks now carry the seconds of the recording they came from, the way PDF chunks carry a page number, and the position travels everywhere a chunk goes: GraphRAG provenance and vector chunks, search_chunks, summarize, the search-engine dicts and chat citation resolution. The interface renders them as time badges — "12:34" beside a chunk number, "At 1:27–2:13" on a chat citation chip.
Two benchmark boards, scored without a judge
chaoscypher benchmark gains an instruction-probe tier: 65 probes across five sections, one per extraction instruction, each with a passage constructed so the correct extraction is unambiguous and checkable by eye. Probes call the real extractor with the real prompts, settings pins and line parser, so a pass means the product could use the output. A "carrier" tier splices each probe into a real ~3,000-character extraction group from the corpus, testing production density rather than production length.
The second board is grounded chat. Every labelled question is scored by pure functions: the answer names the terms a correct concise answer must contain, carries no names or numbers the retrieved context does not support, leads with the right entity when two are confusable, declines when the sources cannot answer, and passes a completion gate. A judge model is now optional rather than required. The fixture is War and Peace Book One — 113,934 characters against a ~7.7k retrieval window, so a question is answered from retrieval and not from the whole corpus — with 80 labelled questions covering single-hop, paraphrase, multi-hop across chapters, fine-grained confusable pairs, and out-of-scope questions with decoys.
/leaderboard blends the two: Model, Overall, Extraction, Chat, VRAM and Speed. Overall is 0.6 × extraction + 0.4 × chat, and the weights are written into the exported data rather than recomputed per page, so the leaderboard, the homepage teaser and the app's model pickers agree. Each score shows its counts, its pass rule, the retrieval floor and the run time.
In plain English: the numbers are produced by checks you can read, and every page that shows them shows the same ones.
The measurement work found two things that made every earlier number suspect
This is the part worth reading even if you never run a benchmark.
Extraction truncation was invisible four layers deep — and a truncated extraction scores higher than a complete one. Measured twice, independent of thinking mode: gemma4:31b scored 72.09 truncated against 67.02 complete, and qwen3.8:27b under a forced 300-token budget scored 76.49 against 71.25. Core computed the signal correctly the whole time; it died on the way to anyone who could act on it. The CLI ingest path received the metrics dict from extract_single_chunk and discarded it, so every chaoscypher load dropped truncation on the floor and chaoscypher source get reported a clean ingest for a source whose chunks were cut off. A new get_source_counters validates against the same allowlist the increment uses, so the read and write sets cannot drift, and the leaderboard warns by name about any model whose extraction did not run to completion.
Benchmark runs advertised determinism they never applied. Temperature and seed were pinned behind hasattr guards keyed on field names LLMSettings does not have, so both assignments silently no-opped while the result row recorded seed=42, temperature=0.0. Five identical runs spanned 3.5 score points before the fix and are byte-identical after it. Two product defects surfaced with it: there was no seed plumbing anywhere (LLMSettings.seed is new and applied by the Ollama provider; the default stays None, so nothing changes unless you pin it), and thinking_for_extraction was read by the schema extractor but never passed by the path every chaoscypher load takes.
In plain English: if you compared local models on an earlier release, compare them again.
Preset chat defaults that can actually call tools
phi4:14b does not advertise tool calling in its Ollama manifest, and the product chat loop always passes the graph tools — so anyone on the 16 or 20 GB preset failed on the first turn with ToolCallingNotSupportedError. Those tiers now ship qwen3.5:9b, which scores 57/80 on the Book One chat board and is the strongest model under 14B on it. The 24 GB preset moves to qwen3.8:27b (62/80) and 32/48 GB to qwen3.6:35b-a3b (63/80) — the top scorers that fit each tier with 4 GB of context headroom. The previous default, qwen3:30b, had never been benchmarked; the board measured the instruct variant.
An existing install keeps whatever its settings.yaml already names. Both model pickers in the app are now generated from the leaderboard data, and a model that cannot call tools, or whose chat run timed out, lands in a "Not recommended" group.
Other fixes worth knowing about
-
Exported packages carried zero node templates.
_templates()dropped everyis_systemtemplate, and extraction files entities under system templates — so every mounted, loaded or composed entity arrived untyped. -
A CLI-built package silently lost every citation across a load or mount:
package exporthard-codedinclude_sources=Falseon the false premise that the CLI has no sources repository. -
Ollama ignored
llm_request_timeout. A model that produced zero bytes sat 75 minutes at 100% GPU. The timeout is now the httpx read timeout, so silence fails on both the streaming and non-streaming paths while a progressing generation never trips it. -
The
\|escape the extraction prompt documents never worked — the parser split on every|and unescaped afterwards, so an entity namedRostov \| Bolkonsky & Sonsproduced one field too many and the whole line was dropped. -
Entity lines survived neither a missing nor a doubled
aliasesfield — 740 + 18 such lines across the saved probe runs.parse_entity_linenow anchors on the confidence and sentence-reference pair instead of counting fields from the left. -
python -m chaoscypher_cortex.main startbuilt the app and exited 0 — noif __name__ == "__main__"guard, and that is the exact command bothchaoscypher serveandcompose uprun.
Upgrading
One migration, 0007_chunk_media_timestamps, is additive only — two nullable REAL columns on document_chunks, classified safe_auto, so it applies on startup with automatic migrations enabled. There are no breaking changes to the product API.
docker pull ghcr.io/chaoscypherinc/chaoscypher:0.5.0
Or, if you run the Python packages directly:
pip install -U chaoscypher-core chaoscypher-cortex chaoscypher-neuron chaoscypher-cli
Starting fresh:
docker run -d --name chaoscypher \
-p 80:80 \
-p 443:443 \
-v chaoscypher-data:/data \
--add-host=host.docker.internal:host-gateway \
ghcr.io/chaoscypherinc/chaoscypher:latest
Three things to read before you upgrade
-
Two defaults flipped toward carrying more data.
chaoscypher package exportandgraph package loadnow include sources, chunks and citations unless you pass--no-sources. Exported packages are correspondingly larger, and loading one costs an inline search-indexing pass. This is the fix for silently losing every citation across a package round-trip, but it does change the size and timing of both commands. - Packages exported before this release carry no node templates, and if they were built by the CLI, no sources. Re-export them to get typed entities and citations on the far side.
-
The MCP benchmark tool payloads changed shape late in this release. The benchmark tools are new in v0.5.0, so nothing released is broken — but if you built a client against
mainwhile the extraction-probe track was landing, re-read the payloads.get_benchmark_tasknow returnstask_id;probe_idis a deprecated alias kept for one release.submit_benchmark_outputno longer returns a verdict — nothing is scored untilfinish_benchmark. And each stage is answered exactly once: re-submitting is refused withTASK_FINAL.
Next steps
- If you keep knowledge packages, re-export them and try
chaoscypher mountagainst an MCP host — that is the shortest path from a built graph to an assistant that can query it. - If you picked a local model from the app's dropdown before this release, check it against
/leaderboard. The truncation and determinism fixes above changed what the numbers mean.
Full details in the changelog. Chaos Cypher is AGPL-3.0 and local-first: the graph, the chat, and the import and export paths all run on your own machine. Repo: https://github.com/chaoscypherinc/chaoscypher


Top comments (0)