Most LLM products in rule-heavy domains break the same way: they let the model do the arithmetic.
I build an app that reads Chinese metaphysics charts (Zi Wei Dou Shu / BaZi). It's a good stress test for this problem, because the domain splits cleanly in two. The calculation side is fully deterministic — given birth date, birth time and gender there is exactly one correct chart, and the rules that produce it are fixed tables that haven't moved in centuries. The interpretation side is prose. Those two jobs want completely different machinery, and the bugs show up the moment you let them blur.
Here's the separation I landed on, plus the parts that were not obvious.
The failure mode: plausible instead of correct
An LLM asked to "generate a chart" will produce something that looks like a chart. The star names are right. The palace structure is right. The formatting is beautiful. And the placements are wrong, because placement depends on calendar facts the model has no way to compute reliably.
The nasty part is that nobody can tell by looking. A user who doesn't already know their chart can't audit it. So wrong output and right output are indistinguishable at the UI level — which means the whole product is unverifiable, no matter how polished the prose on top is.
That's the actual defect. Not "the AI is inaccurate" — the output is unauditable.
Stage 1: compute, don't generate
Stage 1 takes birth data and returns a chart. Every value comes from a fixed rule table or a calendar conversion. No model call, no sampling, no temperature.
Two things make this stage harder than it sounds:
Calendar conversion is the real trap. Converting a Gregorian birth date into the calendar system the rules operate on is not a formatting problem. You're dealing with historical timezone offsets, daylight-saving periods that existed in some regions and not others, and the difference between clock time and true solar time — which is a function of longitude and the equation of time. Get any of those wrong and every downstream placement shifts by one palace. I stopped writing this myself and pinned a well-tested calendar library instead, because the edge cases are historical data, not logic.
Rule tables have to be data, not prompt text. Star placement is a set of lookup rules keyed on things like the value and parity of specific cycle indices. If those tables live in a prompt, they're re-interpreted on every call and drift. If they live in code as tables, they're diffable, testable, and reviewable — and when a disputed rule shows up, you can point at one line.
The output of Stage 1 is a plain data structure. That matters for what comes next.
Stage 2: the model explains what already exists
Stage 2 receives the computed chart and produces human language. It explains what a star sitting in a particular palace means in this system. It does not decide what's in the palace — that's already settled.
This split buys a specific property: the reading can be traced back to a fixed rule. When someone asks "why does it say that?", the chain is: this sentence interpreted this fact, and this fact came from this table at this line. Not "the model felt that way."
I'm deliberately not claiming the interpretation is objective — it's a traditional reading system, and reasonable practitioners read the same chart differently. What I can claim is that the mechanical layer is reproducible and the interpretation layer is constrained to it. Those are separate promises and it's worth keeping them separate in how you talk about them.
Testing is unusually easy here, and that's the tell
Because Stage 1 is deterministic, tests are exact-match fixtures: input birth data → expected chart. No fuzzy scoring, no LLM-as-judge, no "does this look reasonable."
That property is a good smell test for architecture. If your core layer needs a model to grade it, you don't have a deterministic core — you have a second model, and now you need to validate that.
The interpretation layer is different and I test it differently: regression fixtures on the facts handed to the model, plus checks that it never contradicts the computed chart. Checking "did the model stay inside the box it was given" is a much better-behaved test than trying to grade prose.
Boring infrastructure on purpose
The whole thing runs as a small Python service behind a tunnel:
- one application process, bound to
127.0.0.1only — the tunnel is the only way in, so there's no directly exposed port to misconfigure - SQLite for state, single file, included in the daily backup rather than run as a server
- a process supervisor that restarts the app if it dies, plus a separate watchdog that checks the app answers before declaring anything healthy
Two operational lessons that cost me real time:
"The service responded 200" is not the same as "the supervisor is running the service." You can end up with a stray process serving traffic while the managed one is dead, and then the next restart takes down a service you thought was supervised. Compare the supervisor's PID against the PID actually holding the port. If they differ, you have a ghost.
A watchdog that auto-restarts on a timer will happily restart you mid-deploy. Pause it during deployments and restore it after, with a trap so it comes back even if the deploy fails.
What I'd pass on
- If your domain has one correct answer, compute it in code and never let the model near it.
- Make the mechanical layer auditable before you invest in prose quality. Prose on top of an unauditable core is a confidence trick, even when it's accidental.
- Exact-match fixtures are a gift. If your core can't have them, that's worth investigating.
- Keep your claims as narrow as what you can actually verify. "Reproducible computation" and "meaningful interpretation" are different promises; conflating them is how products end up overclaiming.
The two-stage split is the only structural decision in this project I'd defend without caveats. Everything else is details.
I write about this build in more detail — architecture, the bilingual pipeline, and the calendar edge cases — at askziwei.com. If you're working on a rule-heavy LLM product, I'd genuinely like to hear how you drew the line between computed and generated.
Top comments (0)