DEV Community

Aiepco
Aiepco

Posted on Originally published at plc.aiepco.com

A linter that silently passes is worse than no linter at all

Our static analyzer had been reporting no issues found on our PLC codebase for months.

Then we built a fixture stuffed with deliberate bugs — and found that two of its rules had never fired on a single line of code. Not once, since the day they were written.

This is a post-mortem of our own verification toolchain failing, and the eight checks we now run before we trust any checker.


Context: what we verify, and why it matters

We develop and verify Siemens S7-1200/1500 PLC programs (SCL, IEC 61131-3). The code runs production lines. A defect is not a bad user experience — it's a stopped line, scrapped product, or a safety incident.

So we built a verification chain: static analysis rules, a self-test suite, and mutation testing. This article is about that toolchain failing at its job.

Finding #1 — a rule that could never fire

Here is what a normal run looked like:

[motor_control_ref.scl]
  LOC=86 SLOC=61 comment_rate=27% McCabe≈9 nesting=3 magic_numbers=0
  POU: FUNCTION_BLOCK FB200_MotorControl (line 11)
  No static issues found
Enter fullscreen mode Exit fullscreen mode

One green line. Looks like good news.

Meanwhile we were writing a fixture — a file full of deliberately planted defects — to check whether the rules actually work. The result was unsettling: the fixture contained #m_rB used as a divisor, and unguarded array access like #m_Buffer[#i_Index]. The rules for division-by-zero risk and array-out-of-bounds risk reported nothing.

The cause was in the regex:

/(?![/\\*])\s*([A-Za-z_][A-Za-z0-9_]*)
Enter fullscreen mode Exit fullscreen mode

It looks for a variable name following a slash. But SCL local variables carry a # prefix — the real code is #m_rA / #m_rB. After the slash comes a space, then #, then the name. The # broke the match.

Two rules, written months earlier, had never matched anything.

Finding #2 — the scary part is its siblings

Fixing the regex took five minutes. The whole next day went into a different question: how many other rules are like this?

We ran every rule against the fixture. Three more problems surfaced:

Problem Symptom Nature
VAR CONSTANT treated as a static variable All 6 constants flagged with "non-standard naming prefix" False positive — rule too broad
Safety-comment check too aggressive One function block produced 17 warnings of the same kind Noise flood — real issues get buried
Call-graph parsed ELSIF (...) as an FB call Call-depth analysis distorted Misparse — parser boundary error

Three rules, three different diseases: too broad, too noisy, misparsed. Together with the two that never fired, that's five failure modes.

The common property: none of them raised an error. The tool didn't crash, didn't throw, didn't warn. It just quietly produced a wrong conclusion.

Finding #3 — "passes" is not one thing

We then ran mutation testing: inject a defect into clean code and see whether the chain catches it. Twelve mutation classes:

  • Syntax / structural (missing END_IF, missing semicolon, unbalanced parentheses, naming violations, magic numbers, stripped comments) → caught 6/6 by static analysis.
  • Semantic (AND → OR, deleted NOT, inverted comparison, altered initial value, removed interlock) → caught 0/6.

The most instructive number: we flipped the AND in a start condition to OR — the logic is now exactly backwards — and compared the text against the reference implementation:

BLEU-4 cumulative : 0.9981
Line-level match  : 97.6%
Enter fullscreen mode Exit fullscreen mode

Nearly a perfect score. Semantic mutations can make code look more like the reference, not less. No text- or structure-based tool can see them.

So we put that limitation into the report rather than tuning the numbers:

Static analysis passing ≠ compiling ≠ being logically correct.
Semantic correctness is the job of dynamic verification (simulation, on-machine trace comparison) and human review.

Finding #4 — the docs referenced scripts that didn't exist

One more, and it is embarrassing rather than interesting: our framework documentation listed a set of tooling scripts. At that moment, none of them existed yet. The documentation was complete; the implementation was zero.

Same disease: interface declarations divorced from implementations. We added an assertion that scans the docs and fails if any referenced script or template is missing. Dead links don't get a second chance.

What we changed

Three rules, all cheap:

  1. Every rule needs a fixture that must trigger it. Write the rule, immediately build a sample containing that defect, assert it is reported. Freeze the sample as a regression test.
  2. Every rule needs a counter-fixture that must not trigger it. Run it against compliant code, assert zero warnings. Without this you get the worse failure: the team learns to ignore all warnings, which is identical to having no checker.
  3. The toolchain must prove itself first. Ours is 45 self-test assertions, a few seconds to run:
$ selftest_tools.py
assertions 45, failures 0
result: OK — toolchain behaves as declared
Enter fullscreen mode Exit fullscreen mode

The rule is blunt: if the toolchain self-test isn't fully green, you may not use it to judge code. An unverified tool's "pass" is not evidence.

The eight checks we now apply to any checker

Abstracted, this is a checklist you can run against your own linters, analyzers or CI gates:

  1. Does every rule have a sample that must trigger it? (No → you don't know whether it runs at all.)
  2. Does every rule have a counter-sample that must not trigger it? (No → it may be killing valid code.)
  3. Is the warning count sane? 17 identical warnings in one file means the rule is noise, and the team will ignore it.
  4. When the tool fails, does it error or does it go silent? Silent failure is the dangerous kind.
  5. Is there a meta-test guarding the tool itself? Changing A and breaking B is the norm, not the exception.
  6. Have you written down what the tool deliberately does not do? Unstated boundaries are read as omniscience.
  7. Does every file referenced in the documentation actually exist?
  8. Before the word "passed", is there a line saying which layer passed?

Why publish this

There's no growth chart or ROI curve here. It's a record of our own toolchain failing and being fixed.

We publish it because of a fairly plain belief: if a supplier won't tell you whether their own checkers have been verified, you should trust what they hand you even less.

We prefer to say three separate sentences — static checks passed / it compiles / the logic is correct — rather than one vague "quality assured". Keeping them separate is more useful to the person paying for the work.


This is a real internal post-mortem. All figures (14 static rules, 45 self-test assertions, 6/6 vs 0/6 mutation detection, BLEU 0.9981) come from actual tool output, and contain no client information.

We develop and verify Siemens S7-1200/1500 PLC programs — and we build the tooling that proves they're correct. How we work: plc.aiepco.com

Top comments (0)