Forge Server Tools · every failure mode · guides

How these were built: a 234-mod server, and everything we got wrong first

Written by me with Claude, and corrected by a server with 234 mods on it.

Who did what

I will be straightforward about it, because vagueness here is worth less than specifics. Claude wrote the code and the tests. I ran the production server these were built against, supplied the thirteen real crash reports they were rebuilt from, made the calls about what shipped, and found or forced the fix for every bug listed below. Every commit is tagged Co-Authored-By: Claude; none of this is hidden.

If that disqualifies these tools for you, that is a fair call and the full source of the free editions is right there to read before you decide.

Why that is not the interesting part

The thing that makes a crash-report reader correct is not who typed it. It is whether it has been run against crashes that actually happened, by someone who already knew the answer and could tell when it was wrong.

That is the real claim here, and it is the one worth checking: these were rebuilt from thirteen genuine crash reports off a live 234-mod server, and every bug below was caught by running them against reality rather than by reasoning harder. A tool written entirely by hand and tested only on invented examples would be in worse shape — in fact that is precisely the state version one was in, and the rest of this page is what fixed it.

Six tests out of six. Three real crashes out of eight.

The first version of the crash reader passed every test written for it. Six synthetic logs, one per failure mode, each containing exactly the string its pattern looked for.

Pointed at a real crash-reports/ folder, it diagnosed three of eight.

The same author had written both the question and the answer, so of course they matched. The patterns were not wrong so much as narrow — tuned to a tidy example of each error rather than the shape those errors take when a real pack with 234 mods falls over. Real crashes came wrapped in Forge's own exceptions, or as stale-jar NoClassDefFoundErrors, or as malformed resource IDs that no invented sample contained.

The fix was not cleverness. It was reading thirteen real crash reports and writing rules for what was actually in them. A green test suite built on invented inputs measures the author's imagination, not the code.

A confident theory that moved the number by 0.00 seconds

The tool took 59 seconds on a real debug.log. The obvious culprit turned up immediately: a single line of 107,445 characters, a complete HTML page that some mod had fetched and logged in full.

Capping line length changed the runtime from 58.84s to 58.84s.

The real cause was catastrophic regex backtracking in an unrelated pattern — (?P<mod>\S+).*is client-only, where an unbounded \S+ followed by .* never returns on a long non-matching line. Bounding it took the run to 1.97s.

The interesting part is not the regex. It is that the first explanation was plausible, specific, well-researched and completely wrong, and the only reason anyone found out was measuring before and after instead of declaring victory.

Blaming the wrong thing, three times

The feature that makes the tool worth anything is naming the mod jar from the stack trace. The first version blamed netty-common on about half of all crashes, because Netty sits near the top of any network-related trace. Fixed that, and it blamed fmlloader. Fixed that, and it blamed modlauncher.

Every one of those was a confident, authoritative-looking wrong answer, and each was caught only by running it against reports whose cause was already known.

169 typos, most of them imaginary

The modpack validator's first run on a real pack reported 169 typos. Most were noise — it was comparing whole IDs, so a shared namespace contributed fourteen identical characters to every comparison before the part that distinguishes them was reached. Comparing only the path cut 169 candidates to 63.

Three of those are real, still live in a pack thousands of people play.

It named a real person's mod as the thing that broke a server

The worst one. The README, the test suite and a published article all used a real third-party mod as the example crash culprit — because it genuinely was, in one of my logs. It would have shipped to buyers as the illustration of a broken mod.

That is not ours to do. It was caught by a check Claude wrote to scan its own output for exactly this, and replaced with an invented mod name everywhere, including in the already-published article.

And it was wrong about its own storefront

The automatic check written to watch the products reported that a published, paid product had no file attached — that anyone buying it would get a receipt and nothing to download. It was about to send me to re-upload a file that was already there, twice over. The field it trusted is only populated when a product has exactly one plain attachment.

An empty field is not proof of absence.

What the process actually was

Claude is fast, consistent, and has better test hygiene than I would at 1am. What it cannot do is know whether any of it is true, and it is equally confident either way. Every correction above came from contact with a real server, real logs and a real pack — not from more reasoning.

So the loop was: it writes, I point it at something real, the real thing disagrees, we fix it. Both halves were necessary, and the half that found the bugs was not the clever one. That is also why the test suites ship with the tools — they are the record of which version was wrong and how we knew.

Read it yourself

The free editions are complete, single-file and dependency-free, and every test suite ships with them — including the ones that document which earlier version was wrong and why.