C gets hard to trust in three predictable ways: deep nesting, pointers to
pointers, and a->b->c chains where nobody null-checks the middle pointer.
So I wrote a tiny scorer that squashes that into one number per function:
score(f) = nesting × pointer depth × deref chain
Each factor is one sentence:
-
nesting — the depth of
if/for/whilelevels -
pointer depth —
int *,int **,int ***in params and locals (deeper is closer to a memory bug) -
deref chain —
a->b->cruns, i.e. null-derefs where the middle pointer came from somewhere else
It's one file, zero dependencies, no parser worth mentioning.
Does the number mean anything?
I checked it against maintenance churn — how often a function actually gets
touched — on two very different codebases (libXt, a 40-year-old X11 toolkit,
and libtiff, a format parser):
- Spearman(Score, Churn) = 0.52 (libXt), 0.38 (libtiff)
- better than raw line count (0.50 / 0.33) and cyclomatic complexity (0.41 / 0.32)
- the top-15 functions see roughly 3–5× the churn of the bottom-15
I also ran it across 20+ years of git history for libtiff, curl, redis and
OpenMotif. None of them trend toward smaller functions — the median stays flat
and the biggest function only ever grows. Which is exactly why a cheap pointer
to the hot spots helps.
The part I actually use it for
LLM-generated C has the same tell. The loop is:
generate → c-score file.c → "rewrite the top 3" → re-score
One round visibly flattens the output. A dumb, explainable number is enough to
steer an LLM — you don't need a real static analyzer for this.
Honest limits
It's a triage tool, not a bug detector. It won't find semantic bugs, and it
can't see a deep call stack full of side effects. The parser is not perfect,
and the Python is LLM-generated too. But the idea is the sound part — before
you pick it apart, run it on a real codebase.
pip install c-code-score
for f in src/*.c; do c-score "$f"; done \
| awk '/^[^ ]/{n=$1} /^ score:/{print n,$2}' \
| sort -k2 -rn | head
Repo: https://github.com/xtforever/c-score · https://codeberg.org/au1064/c-score
Top comments (1)
The Spearman correlation against maintenance churn is the right way to validate a metric like this — not "does it correlate with known bugs" (too hard to measure) but "does it predict which functions people actually have to touch." 0.52 on libXt beating raw line count and cyclomatic complexity is a meaningful result for something this simple.
The LLM feedback loop is the part I'd highlight more. The generate → score → "rewrite the top 3" → re-score loop is useful precisely because it's dumb and explainable. LLMs respond well to concrete, numeric feedback they can reason about. A high c-score on a function gives the model something specific to optimize against — much more actionable than "make this less complex."
The historical flatness finding (median score stays flat across 20+ years, biggest functions only grow) is also worth sitting with. It means this isn't a problem that gets better over time without deliberate intervention. Tools like this are the intervention.