I built a tool that reads modded Minecraft crash reports and tells you which mod killed the server. It passed every test I wrote and was still wrong about most real crashes.
Here is what the real data taught me, in the order it hurt.
Your own samples will lie to you
I wrote the tool first and the test fixtures second. Six synthetic logs, one per failure mode, each one crafted to contain exactly the string my pattern looked for. All six passed.
Then I pointed it at an actual crash-reports/ folder from a live 234-mod server. It diagnosed three of eight.
This is the test-writing equivalent of grading your own homework. I had written both the question and the answer, so of course they matched. The patterns were not wrong exactly — they were narrow, tuned to a tidy example of each error rather than the shape those errors take in the wild. Five real crashes used wrapper exceptions, stale-jar NoClassDefFoundErrors and malformed resource IDs that my neat little samples never contained.
The fix was not cleverness. It was reading thirteen real crash reports and writing rules for what was actually in them.
The filename that contains two versions
Forge ships as forge-1.20.1-47.4.10-universal.jar. That is Minecraft 1.20.1, Forge 47.4.10.
My regex to find the loader version was:
r"(?:neo)?forge[ \-](\d+\.\d+[\.\d]*)"
Reasonable looking. It reported every server's Forge version as 1.20.1, because that is the first number after forge-.
Worse, it was plausibly wrong. 1.20.1 is a real version string that appears everywhere in the log, so nothing looked broken. It took reading the output next to a crash report I already understood to notice the loader and the game had suspiciously identical versions.
The fix is to prefer the authoritative line and only fall back to the filename:
for pat in (
r"Forge:\s*net\.(?:neo)?(?:minecraft)?forge:(?:forge:)?(\d+\.\d+[\.\d]*)",
r"(?:neo)?forge-\d+\.\d+(?:\.\d+)?-(\d+\.\d+[\.\d]*)-universal",
r"(?:neo)?forge[^\n]{0,24}?version[:\s]+(\d+\.\d+[\.\d]*)",
):
m = re.search(pat, joined, re.I)
if m:
break
Crash reports carry Forge: net.minecraftforge:47.4.10 in their system details. Use the thing that is unambiguous.
One line of HTML, 107,445 characters long
The tool took 59 seconds on a real debug.log. The file was 20,000 lines and 3.4 MB — that should be well under a second.
My first theory was the obvious one. I measured the longest line in the file:
107,445 characters. It was a complete GitHub HTML page. Some mod had fetched a URL and logged the entire response body on one line.
So I capped line length for matching at 4,000 characters, re-ran, and the time went from 58.84s to 58.84s. Zero change. The theory was wrong, and the only reason I knew is that I measured before and after.
So I stopped guessing and instrumented each stage:
read_log 0.04s
environment 0.01s
blame_mod 0.00s
diagnose >115s <- there it is
Then timed each rule individually. The first rule printed in 0.04s. The second never printed at all, because a single re.search() call on one line had not returned.
The pattern:
r"(?P<mod>\S+).*is client ?-?only"
An unanchored \S+ followed by .*. On a long line with no match, the engine tries every split of \S+ against every split of .*. That is catastrophic backtracking, and it does not finish.
r"(?P<mod>\S{1,80}) is client[ \-]?only"
58.84s to 1.97s.
The lesson I actually took from this is not "avoid nested quantifiers". It is that I had a theory, it was wrong, and I only found out because I measured instead of declaring victory. The HTML line was real, interesting, and completely irrelevant.
Now there is a test that fails loudly rather than hanging:
def test_no_pattern_has_unbounded_quantifier_after_a_capture(self):
bad = re.compile(r"\(\?P<\w+>\\S\+\)\.\*")
for rule in RULES:
for pat in rule.patterns:
self.assertIsNone(bad.search(pat), f"{rule.id}: {pat}")
Blaming the right jar
The genuinely useful trick, and the reason the tool is worth anything: Forge annotates every stack frame with the jar it came from.
at net.mcreator.grimfauna.procedures.StalkerProcedure.execute
(StalkerProcedure.java:250) ~[grimfauna-1.20.1-2.1.11.jar%23205!/:?]
So you can walk the trace, skip anything belonging to Minecraft, Forge or the JDK, and the first jar left is almost always the mod at fault.
The whole feature lives in that skip list. A naive version blames netty-common on half of all reports, because Netty sits near the top of any network-related trace. Then you fix that and it blames fmlloader. Then modlauncher. Every one of those was a real bug I shipped and caught by running it against reports whose cause I already knew.
The same mistake, a second time, on a different tool
I built a second tool that validates modpack IDs — the typo'd minecraft:diamond_swrd that never crashes and simply gives the player nothing.
Its first run on a real 233-mod pack reported 169 typos. Most were noise:
apocalypsenow:alicepack
did you mean: apocalypsenow:icepick
That is not a typo. That is fuzzy matching with no idea what it is doing.
The cause was the same shape as the Forge version bug: I was comparing the whole ID, so the shared apocalypsenow: prefix — fourteen identical characters — inflated every score before the interesting part was even reached.
Measured:
| compared | full ID | path only |
|---|---|---|
cooked_caned_fish vs cooked_canned_fish
|
0.984 | 0.971 |
alicepack vs icepick
|
0.909 | 0.750 |
Comparing paths alone separates a real typo from nonsense cleanly. 169 candidates became 63.
And three of those are real, still live in a pack thousands of people play, confirmed against the mod's own jar:
-
cooked_caned_fish→ the item iscooked_canned_fish -
cooked_canned_rabit_soup→ the item iscooked_canned_rabbit_soup -
stampler→ the item isstapler
Two of those sit in the pack's diet tags, so those foods silently have no diet category. Nothing crashes. Nothing logs. You find out when a player complains, if ever.
What I would tell myself at the start
Write the fixtures from real data, or accept that your tests prove nothing. A green suite built on invented inputs measures your imagination, not your code.
Revert each fix and confirm the test fails. I found one test that passed with its fix removed entirely. It had been quietly measuring nothing for days.
When you have a theory about a performance problem, measure before and after. My plausible, interesting, well-researched theory moved the number by 0.00 seconds.
A tool that cries wolf gets ignored. Findings are now tiered by confidence, and the low-confidence tier is hidden behind a flag. Sixty-three real candidates beats a hundred and sixty-nine you learn to scroll past.
Both tools are single Python files with no dependencies. The free editions are on GitHub:
- forge-server-doctor — crash report diagnosis
- pack-doctor — modpack ID validation
There are paid versions with the full rule sets at jaakoby.github.io, but the free ones are standalone and the debugging stories above are the real point.
Top comments (0)