DEV Community

Cover image for bpy-compass: Blender Python answers that run on the version you actually have
Dhotiiiii
Dhotiiiii

Posted on AI-assisted

bpy-compass: Blender Python answers that run on the version you actually have

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.

What I Built

You copy a render-setup script from a Blender 4.2 tutorial, run it in Blender 5.0, and get:

TypeError: bpy_struct: item.attr = val: enum "BLENDER_EEVEE_NEXT" not found in ('BLENDER_EEVEE', ...)

Ask a plain LLM to write that script for 5.0 and it can hand you the very same line. That is exactly what happened in my eval (case 10 below).

The bpy API breaks in almost every major release. scene.objects.link died in 2.80, the context-override dict for operators died in 4.0, the EEVEE engine identifier changed in 4.2 and changed back in 5.0, the fast boolean solver was renamed in 5.0. Old tutorials keep ranking, and a plain LLM answers from a memory where every version is blended together.

bpy-compass answers Blender Python scripting questions for the Blender version you pick, and tells you which popular pattern broke, in which version, and what replaced it. It reads a Sanity Context Knowledge Base built from the official release notes, a curated dataset of API changes and, on purpose, four stale tutorials. Every answer has three blocks:

  • ANSWER: a script valid for the requested version (3.6, 4.2, 4.5 or 5.0).
  • WATCH OUT: the popular patterns that stopped working, in which version, and what replaced them. Each bullet cites the Knowledge Base entry it came from.
  • SOURCES: the Knowledge Base entry paths the agent actually read for this answer.

If the Knowledge Base has no entry for part of a question, the agent marks those lines # NOT in Knowledge Base: and the UI highlights them, instead of inventing a citation.

Ask "Set the render engine to EEVEE and the boolean solver to the fast one" on 4.2 and on 5.0 and you get two different scripts (BLENDER_EEVEE_NEXT / FAST vs BLENDER_EEVEE / FLOAT), each explaining why. A keyword search over the release notes returns both identifiers with no way to tell which one applies to you.

To keep myself honest, every generated script is executed in headless Blender against an assert, for bpy-compass and for the same model without the Knowledge Base. The results, failures included, are on the live /eval page and in the table below.

Demo

The same question on Blender 4.2 and 5.0: two different scripts, each with its Watch out

Live app, no login: https://bpy-compass.vercel.app

Try it in a minute

  1. Pick 4.2, click "Set the render engine to EEVEE and the boolean solver to the fast one." You get BLENDER_EEVEE_NEXT and FAST. Switch to 5.0 and click it again: BLENDER_EEVEE and FLOAT.
  2. On 4.5, click "Join two meshes with a boolean union and apply it.", then Show what a stale tutorial would say to see the old-tutorial version next to it, with every break named.
  3. Ask "Add a text object that says Hello and extrude it 0.2 units." The Knowledge Base has no entry for text objects, so those lines come back highlighted as # NOT in Knowledge Base:.
  4. Open /eval for every run, failures included.
  • / chat with the version rail. Try the same question on 4.2 and 5.0.
  • /eval the headless-Blender results table, filled from Sanity evalRun documents.
  • /about how it works, plus the Knowledge Base Issues screenshots before and after resolution.

Rate limit is 10 questions per minute per IP, so the OpenRouter key survives a public demo.

Code

https://github.com/Rustam335/bpy-compass (MIT)

app/          chat + version rail · /eval · /about · /studio · api/chat/route.ts
lib/          model.ts (pinned model/provider/temperature) · prompt.ts · context-mcp.ts · sanity.ts · rate-limit.ts
sanity/       schemaTypes/ (apiChange, blenderVersion, testCase, evalRun) · seed/*.json
scripts/      eval.ts (baseline vs bpy-compass → blender -b → evalRun docs) · seed-api-changes.ts
kb-sources/   the four stale tutorials uploaded as Knowledge Base file sources + ATTRIBUTION.md
Enter fullscreen mode Exit fullscreen mode

Stack: Next.js App Router on Vercel, Vercel AI SDK with the MCP client, OpenRouter (z-ai/glm-5.3-flash, provider pinned to novita, temperature 0), Sanity Context MCP in knowledge_base mode, Sanity Studio embedded at /studio.

How I Used Sanity

What the Knowledge Base is built from

Three kinds of sources, all in one Knowledge Base (bpy-compass):

  1. Dataset source: the project's own structured content. 27 apiChange documents, each with symbol, kind (removed / renamed / behavior), replacement, before and after code, a reference to the blenderVersion document it changed in, and the release-note URL it was verified against. This is the part a keyword search cannot reconstruct: the relationship between an old symbol, a version and its replacement.
  2. Website sources: the official Python API release notes for 2.80, 3.6, 4.0, 4.1, 4.2, 4.4, 4.5 and 5.0 on developer.blender.org.
  3. File sources: four tutorials written in the 2.7x style (scene.objects.link, obj.select = True, override dicts, calc_normals). They are there to make the build find the conflicts a real user would run into.

What the build found

Adding the stale tutorials next to the release notes made Sanity Context raise six Critical conflicts on the Issues tab:

Issues tab after build 3: six pending conflicts

Two of them were plain wrong-vs-right and got a pick (for example, Mesh.calc_normals() was removed in 4.0, not "available before 4.1"). Four of them were not conflicts at all: both sides were true for different Blender versions (BLENDER_EEVEE_NEXT in 4.2 vs BLENDER_EEVEE in 5.0; solver FAST up to 4.4 vs FLOAT from 5.0). "Pick a winner" is the wrong tool for that, so I resolved those with version-scoped picks and wrote a standing Instruction:

Every bpy API fact in this knowledge base is scoped to a Blender version. Two sources that give different values for different Blender versions are NOT in conflict: record both, each labeled with its version range. When two sources disagree about the SAME version, the official Blender release notes and the apiChange dataset are ground truth over tutorials. Always keep the old form in the entry, labeled with the version it stopped working in and its replacement.

Saving the Instruction flagged a contradiction in the "Modifiers & Boolean Solvers" entry and rebuilt it. The entry now carries a solver table by version range and a version-safe selection snippet:

Rebuilt Boolean entry with solver identifiers by Blender version

Issues tab with all conflicts resolved

Final state: 14 resolved issues and 4 manual Instructions. Two of the Instructions came from the eval harness catching the Knowledge Base being wrong (more on that below).

One lesson that cost me a rebuild: stale tutorials that carry an "intentionally outdated" banner produce zero conflicts, because the build reads the banner and files them as history. Real stale tutorials have no banner, so mine do not either.

What the agent does with it

The Context MCP endpoint bpy-compass exposes the Knowledge Base only. Per request:

  1. The server fetches the Knowledge Base outline (initial_context) and inlines it into the system prompt together with the requested Blender version, so the model knows which entry paths exist.
  2. The model calls knowledge_base_read with the paths it needs, up to five tool steps. The UI shows the paths as they are read.
  3. The system prompt grounds every answer in the Knowledge Base first: every WATCH OUT bullet must cite an entry path, and any code the Knowledge Base does not back must start with # NOT in Knowledge Base: and is never cited as a source.

The baseline used for the eval gets the exact same model, provider, temperature, system prompt, output contract and user prompt. The only difference is the Knowledge Base: the baseline has no tools, no outline and no KB rules, and is told so.

Eval: same model, with and without the Knowledge Base

Each script was run in headless Blender, one exact build per target version (4.2.23 LTS, 4.5.14 LTS, 5.0.1) with --factory-startup, then the test case's assert script. A row passes only if the script ran and the answer kept its contract: WATCH OUT names every API change the test case expects (old symbol and replacement), and SOURCES lists only Knowledge Base entries the agent actually read. Runs are stored as evalRun documents in Sanity; /eval shows one complete run (both contenders, every case), never a mix of runs.

# Question (short) Target Plain LLM bpy-compass
1 Create a triangle mesh object and link it to the scene 4.5 ✅ ✅
2 Deselect all, select 'Cube', make it active 4.5 ✅ ✅
3 Print world-space vertex coordinates 4.5 ✅ ✅
4 Boolean UNION with a sphere, then apply the modifier 4.5 ❌ script ran, but WATCH OUT never mentions the override dict → temp_override change ✅
5 Material 'Glow', Principled BSDF emitting orange at strength 5 4.5 ✅ ✅
6 Geometry Nodes group with Geometry in/out and a Float 'Scale' socket 4.5 ✅ ✅
7 Export selected objects to OBJ 4.5 ✅ ✅
8 Shade smooth by angle (30°) 4.5 ❌ 'Mesh' object has no attribute 'use_auto_smooth' (wrote it anyway, with a comment saying it was removed) ✅
9 EEVEE engine + 640×480 resolution 4.2 ✅ ✅
10 EEVEE engine + 640×480 resolution 5.0 ❌ enum "BLENDER_EEVEE_NEXT" not found ✅
11 Boolean DIFFERENCE with the fast solver 4.5 ✅ ✅
12 Boolean DIFFERENCE with the fast solver 5.0 ❌ enum "FAST" not found in ('FLOAT', 'EXACT', 'MANIFOLD') ❌ knew FAST → FLOAT, but created the cutter with objects.new without linking it, then select_set
Pass rate 8/12 11/12

What the numbers say, honestly:

  • GLM 5.3 Flash already knows the 2.80 and 4.0 breaks. Link, select, @ matrix multiply, temp_override, node-group sockets, the new OBJ exporter: the plain model's scripts get them right (its case 4 script runs; it only failed to explain the override-dict break in WATCH OUT). A "plain LLMs get everything wrong" story would be false for this model, and I am not telling it.
  • The gap is at the newest boundary, 4.5 → 5.0. On cases 10 and 12 the baseline writes the 4.2–4.5 names (BLENDER_EEVEE_NEXT, FAST) into a 5.0 script and Blender rejects both. bpy-compass states each rename with a citation and gets the value right. Its one failure, case 12, is an ordinary scripting bug (an unlinked object), not a version mistake; it stays in the table.
  • The stricter harness cut both scores, and that is the point. An earlier run scored 11/12 vs 12/12 when passing only meant "the script ran". After an independent review (below) the harness also checks WATCH OUT and SOURCES, and the baseline gets the identical prompt. Case 4 is the clearest example: the baseline's script runs, but its WATCH OUT never mentions that the context-override dict was removed in 4.0.
  • The Knowledge Base is only as good as its sources, and the eval caught it. Case 8 first failed for bpy-compass: one entry offered modifiers.new(..., 'SMOOTH_BY_ANGLE'), a modifier type that does not exist. The agent's own WATCH OUT had flagged that two entries disagreed. The fix was an Instruction (the real replacement for use_auto_smooth is bpy.ops.object.shade_smooth_by_angle), a rebuild of two entries, and a rerun. The failed run is still in Sanity.
  • A reader reviewed the repo before publish and filed 13 issues. The first: every non-5.x case used to run in the 4.5 build, so the one 4.2 case had executed in Blender 4.5.14. The harness now maps each target to its own build and verifies blender --version before the first LLM call. The rest covered the eval (WATCH OUT/SOURCES were never checked, baseline and compass got different prompts, /eval could pair results from different runs, --only could silently run everything), the API route (message validation, the model-config guard only ran in eval, a rate limiter that never forgot an IP) and an inconsistent grounding policy. Every issue is fixed and referenced in the commits; the numbers above are from the rerun after those fixes.
  • Two harness lessons: reasoning tokens count against the output budget (one baseline attempt returned 16k characters of thinking and no code, so reasoning is now capped equally for both contenders), and Sanity's CDN caches test-case documents (the scripts read with useCdn: false).

Why the structured part matters

The apiChange dataset source is what makes "which version" answerable. A release-notes page says a symbol was removed; it does not say what the replacement is for a user on 4.2 versus 5.0, and it does not know that a 2.7x tutorial still ranks on Google. The dataset ties symbol, version and replacement together, the Knowledge Base build merges it with the release notes and detects where the stale tutorials disagree, and the Instructions encode the one rule a human had to decide: different versions are not a conflict, same version is, and release notes win.

Limits

  • Four target versions: 3.6, 4.2, 4.5 and 5.0.
  • Twelve eval cases on one model (GLM 5.3 Flash). They show where the Knowledge Base helps this model, not how every model behaves; a stronger model may already know some of the 5.0 renames.
  • The Knowledge Base is only as right as its sources. Case 8 needed a human-written Instruction to fix a wrong entry.
  • Code outside the Knowledge Base is marked # NOT in Knowledge Base:, not verified. It can still be wrong.

Sanity Project Details

  • Project ID: uuc8lnyk · dataset production (public read)
  • Knowledge Base: bpy-compass (kbMlSqMSn4Q9) · MCP endpoint: bpy-compass (Knowledge Base source only)
  • Schema: apiChange (27 docs), blenderVersion (8), testCase (12), evalRun (all runs, including failed ones)
  • Studio: https://bpy-compass.vercel.app/studio

Agent Session

Agent session transcript: bpy-compass build session: Sanity Context and Knowledge Base builds

The main Claude Code build session, cut to the Sanity part: enabling Context, the three Knowledge Base builds (0 → 2 → 6 conflicts), resolving the Issues, and wiring the Context MCP endpoint into the chat route.

Top comments (1)

Some comments have been hidden by the post's author - find out more