DEV Community

Cover image for Why AI Keeps Making the Same Coding Mistakes—And How Teaching It Pain Gives It Wisdom

Why AI Keeps Making the Same Coding Mistakes—And How Teaching It Pain Gives It Wisdom

Randal L. Schwartz on September 11, 2026

"What did you do for AI?" "I brought the scars." — Randal L. Schwartz SPOCK: "Dr. Daystrom, a computer is incapable of standard creative tho...
Collapse
 
gde03 profile image
Giulio D'Erme •

I think the “no brain” metaphor points to a concrete systems problem. During a normal coding session, the model’s weights are not updated by what happened. When a new context starts, it has its capabilities, but not necessarily the memory of what failed, why it failed, or which tempting solution already proved wrong.
I'm developing a memory layer to give that experience a durable place to live. I do not store only a rule such as “do not use a shared database.” I record:

  • what broke
  • why it broke
  • the tempting approach that caused it
  • what I tried that did not work
  • what ultimately worked
  • the scope and evidence for the conclusion
  • the invariant that should prevent the failure next time

For example, I once had parallel coding sessions sharing a database. The tests did not merely fail. They interfered with one another because the suite could drop tables belonging to another session. The failed approach was relying on a shared default database. The durable reflex became one database per checkout, with no silent fallback to a shared instance.
Another failure was more subtle. Different embedding models produced vectors with the same dimensions, so using the wrong model did not raise an error. It returned plausible search results that were semantically wrong. The lesson was not simply “check dimensions.” The real invariant was to bind the embedder identity to the tenant and generation, then require certification before serving results.
This is where my memory layer differs from a plain scratchpad. A future session searches with recall_search before acting, and can request trusted passages with recall_evidence. Memories carry provenance, validity, supersession, generation identity, and a trust verdict. If the evidence is insufficient, the system can abstain instead of promoting the nearest text into a belief.
Doing this does not change the model’s weights or magically create human intuition. It provides continuity through verified experience. In my work, that makes sessions much more efficient because I spend less time rediscovering the same traps, and the system carries forward the reasons, not just the conclusions.
At least that's my attempt to turn those scars into reusable engineering memory.

Collapse
 
randalschwartz profile image
Randal L. Schwartz Google Developer Experts • • Edited

This is spot-on, Giulio. You’ve articulated the exact core systems reality that prompted this entire line of research: weights remain frozen during inference, which means standard autoregression is condemned to relive the same sunny-day traps over and over again unless we give experience a structured, durable substrate.

What I love about your schema is that you immediately recognized why naive, flat rule-files ("Don't use a shared database") fail. When an agent only sees a negative prohibition, two things go wrong:

  1. The Pink Elephant Problem: Mentioning the forbidden pattern increases the attentional salience of those tokens without paving an explicit alternative trajectory.
  2. Loss of Causality: Without the why, the tempting approach (the seductive local minimum), and the concrete failure mode, the model easily rationalizes an exception the moment edge-case pressure mounts.

Your two examples illustrate the anatomy of a scar with textbook precision:

  • The shared test database: The trap is convenience (defaulting to a shared instance because initial setup is easy). The wound is non-deterministic cross-session destruction (dropping tables from under another session). And your durable reflex isn't an advisory checklist note—it's an architectural invariant: one database per checkout, with no silent fallback.
  • The embedding dimension collision: This is brilliant. A superficial rule says "verify embedding dimensions," which silently passes and serves plausible-looking garbage. Your invariant binds embedder identity directly to the tenant and generation with mandatory certification. That transforms a latent semantic failure into a hard boundary condition.

Your point about epistemic metadata (provenance, supersession, generation identity, and the ability to abstain) also touches on one of the hardest problems in durable agent memory: preventing cognitive ossification. If memory is static, past workarounds can become dogmatic traps when underlying systems evolve. Tracking supersession and requiring verified evidence ensures the memory layer remains an active immune system rather than an accumulator of stale superstition.

One area we’ve found particularly fruitful alongside retrieval is workflow phase binding: compiling these reflexes not just into a general search space, but attaching them directly to specific lifecycle gates (for example, the pre-execution plan review or the test harness boundary). When the invariant is tied to an operational gate, the agent doesn't just recall the lesson—it is structurally prevented from advancing until the invariant is satisfied.

We are currently finalizing the formal academic preprint ("Synthetic Scars: Mitigating Statistical Amnesia and Plausibility Bias in Autonomous Coding Agents") before posting to arXiv. Since you are actively building an epistemic memory layer in this exact problem space, I’d love to share an early draft with you for peer feedback if you're interested. Feel free to connect with me on LinkedIn or drop me an email at merlyn@stonehenge.com!

Collapse
 
marcobblk profile image
Marco Sbragi •

This is a fantastic perspective and presentation. I particularly appreciate the point about wounds and scars—something I likely hadn't dwelt on before, yet which has undoubtedly shaped my 40-year career. I also agree that a negative statement almost always triggers a positive vision of the very thing we want to avoid. Perhaps a solution—which I’ll test as soon as possible—is to avoid using negative expressions in the rules, opting instead for a sort of checklist regarding the consequences of specific actions. I look forward to following your series with great interest. Thank you so much.

Collapse
 
randalschwartz profile image
Randal L. Schwartz Google Developer Experts •

Marco, thank you so much for this thoughtful comment. Coming from someone with a 40-year career in this industry, that resonance means a great deal to me. As someone who has also been writing software across four decades, I’ve come to realize that what we often casually call "senior intuition" is really just a library of biological scars—the visceral recoil from systems that crashed in production at 2 AM.

You hit the bullseye regarding negative expressions. In both human cognition and transformer architectures, negation is notoriously fraught: to process "Don't think of a pink elephant," the neural graph must first instantiate the concept of the pink elephant. In an LLM, writing "Do not do X" actually primes attention and spikes the token probability of X in subsequent autoregressive generation steps.

Your idea of pivoting toward consequences is a great direction, and I would love to hear what you discover in your tests. One critical nuance we observed in our empirical field work is that passive checklists can sometimes suffer from "compliance hallucination"—the model reads the consequence checklist, nods along politely, and cheerfully hallucinates that it followed every item while still walking directly into the trap.

Where we found consistent, zero-regression stability was in transmuting those consequences into affirmative invariant gates (what we term the Permanent Reflex in the 3-part scar). Rather than a negative prohibition ("Don't leave streams unclosed") or a passive reminder ("Unclosed streams cause leaks"), we encode the positive, non-negotiable precondition: "Every subscription must bind its teardown handler in the exact same lexical scope before authoring any downstream logic."

We will be diving deep into the mechanics of this exact dynamic in Part 3 (the NLP negation defect and Socratic recoil) and Part 6 (The Proscriptive Inversion—the theorem that if an architecture has more rules about what not to do than what to do, the design has already lost).

Please keep me posted on how your experiments go—I would be delighted to hear your field observations!

Collapse
 
ovid profile image
Ovid •

Have you experienced the Straight-A Intern Paradox in your own AI workflows?

I've absolutely hit this, repeatedly, which is why I've been tackling it so hard.

I train a lot of devs on AI, and even the seniors often don't realize they're junior AI devs. "Prompt and pray" feels awesome for a month, but the codebase is quietly turning to slop. The model follows your directions, doesn't see the big picture, and it doesn't care that there's tech debt. (Worse, that tech debt is read and used as part of the model's predictions for what code should look like, compounding the problem).

My angle on this (PAAD) ends up being hyper-aggressive technical debt management, which is why this work looks so interesting. Stage 1, manage debt. Stages 2 and 3, pay it down. The part I keep running into is that paying it down always loses to new features because nobody is getting paged for a Stage 1 scar.

Related question: a scar registry only knows about failures you experience. Does the architecture have anything for the debt that accumulates without ever producing a crash? (god objects, global mutable state, feature envy, etc)

Collapse
 
randalschwartz profile image
Randal L. Schwartz Google Developer Experts • • Edited

Great to see you here, Curtis! (Or Ovid, as so many of us in the Perl and open-source world know you!)

You hit the nail squarely on the head. That compounding effect you mentioned—where the AI reads the codebase's existing technical debt and uses it as in-context statistical priors to generate even more slop—is what we call the In-Context Poisoning Feedback Loop (or the Epistemic Gravity Well, which we codified as SCAR-ARCH-05). It’s why unconstrained "prompt and pray" doesn't just stall; it actively accelerates codebase rot over time.

To answer your core question:

"A scar registry only knows about failures you experience. Does the architecture have anything for the debt that accumulates without ever producing a crash? (god objects, global mutable state, feature envy, etc)"

Yes, absolutely. The key insight is that in the Synthetic Scar Architecture, a "Wound" is not restricted to a 3 AM runtime crash or an unhandled exception. A wound is any measurable failure mode in software engineering physics—including architectural rot, blast-radius explosions, review gridlock, and cognitive unmaintainability.

Here is how the architecture handles silent, non-crashing debt:

1. Architectural Scars for "Silent Killers"

Many of the scars in our registry target architectural anti-patterns directly:

  • SCAR-STATE-01 (Ambient Symbol Mutation & Team-Scale Override Decay):
    • The Wound: Overriding package symbols, using ambient service locators, or introducing global mutable state. It doesn't crash on Day 1; it causes action-at-a-distance bugs and exponential maintenance collapse as team size and parallel agents scale.
    • The Permanent Reflex: Strict explicit dependency injection and immutable boundaries only.
  • SCAR-ARCH-01 (Premature Abstraction & The 3-Point Solution Plane):
    • The Wound: Over-engineered base classes, generic wrapper hierarchies, and god objects that collapse when orthogonal requirements arrive.
    • The Permanent Reflex: Implement concrete logic inline across at least three distinct call-sites before attempting abstraction. Never generalize until three points triangulate the true invariant solution plane.
  • SCAR-ARCH-02 (Identity Pass-Through Shims & The Anemic Lasagna):
    • The Wound: Bureaucratic indirection where function X in Layer 1 delegates to function X in Layer 2 without transforming data or adding behavioral value (X -> X -> X).
    • The Permanent Reflex: Every architectural layer must add distinct behavioral value (validation, caching, transactional boundary). If a layer merely repackages data, delete it.
  • SCAR-API-01 (Boolean Flag Proliferation & Branching Explosion):
    • The Wound: Functions accumulating mode flags (doFoo(isSpecialCase: true)), producing combinatorial 2^N branch explosions and god methods.
    • The Permanent Reflex: Modifying function behavior with boolean mode flags is prohibited. Create distinct, explicitly named entrypoints or strategy objects.

2. The Unprimed Adversarial Critic (The Six Pillars Gate)

You cannot rely on the model that authored the code to detect code smells; it suffers from cognitive momentum and confirmation bias on its own generated tokens.

In our operational workflow, an agent is physically blocked from committing or creating a PR until an isolated, unprimed subagent (running on a higher reasoning tier) audits the raw branch diff across the Six Pillars:

  1. Intent: Complete ticket fidelity with zero scope creep.
  2. Verification: TDD characterization tests asserting the boundary.
  3. Architecture: Audited against the Schwartz Expansion/Preservation Metric (spine vs. periphery, no god objects, no global mutable state, no feature envy).
  4. Blast Radius: Topological dependency and coupling audit.
  5. Reviewer Defense: Anticipating structural trade-offs.
  6. Immune Defenses & Crash Antibodies: Non-negotiable invariant checks.

This is an impassable Dijkstra Guarded Command:

CanAdvanceToCommit <=> Count(CriticBlockers) == 0
Enter fullscreen mode Exit fullscreen mode

If the Critic flags a god object or ambient mutable state under Pillar 3, it issues an immediate BLOCKER. The agent is locked in an autonomous remediation loop and cannot proceed until the design is clean.

3. Solving "Paying It Down Always Loses to New Features"

Your observation about PAAD is spot on: in human organizations, paying down technical debt consistently loses to feature velocity because "nobody gets paged for a Stage 1 scar."

The breakthrough with autonomous agents governed by scars is that debt paydown ceases to be a human political negotiation:

  1. Ambient Cleanliness at the Gate: Because Pillar 3 is an impassable boolean gate on every ticket, the agent cannot land a feature by taking a shortcut that introduces silent debt.
  2. Decoupled Asynchronous Worktrees: When an architectural anti-pattern is identified and codified into a scar, we don't need to beg product management for a two-week refactoring sprint. We can dispatch an agent into an isolated background git worktree to refactor that specific pattern across the codebase, pinned safely by characterization tests.

We dive deeper into this in Part 4 (The Architecture of Synthetic Scars & The Rapid-Regret Miner) and Part 6 (The Proscriptive Inversion: What You Get to Forget, and Why More Negative Rules Mean You've Lost).

I'd love to compare notes with PAAD's defense-in-depth model—the intersection of software engineering discipline and agentic constraint architecture is where the battle against AI slop will ultimately be won!

Collapse
 
ovid profile image
Ovid • • Edited

Well, PAAD has two approaches, but predicated on my assertion that "AI slop" is just a fancy rebranding of technical debt:

  1. Stop errors before they become code (/pushback and sometimes /rethink)
  2. Aggressive maintenance of technical debt (/agentic-review and /agentic-architecture)

I've built a whole methodology around this (by accident, but just updating old-fashioned engineering discipline to AI-speed).

  1. The repo: github.com/Ovid/paad/
  2. The tutorial: curtispoe.org/paad/

There are a ton of agentic skills which can build test suites, and do OWASP and a11y scans (and have a curious habit of finding zero-day exploits).

But it's all been very reactive the way that devs are often reactive. What you're describing sounds like an proactive way of "stop errors before they become code." I'm particularly curious about the work to auto-refactor anti-patterns and prevent them from returning.

we don't need to beg product management for a two-week refactoring sprint

Heh :) I've ruffled some feathers by telling dev teams to do the same thing. So long as they get their tickets done, now they have the time and ability clean up their tech debt and still get other things done. However, if what you're describing works, this will be a game changer. I've told people that I want my PAAD project to die because we need native AI solutions for this. I'd be delighted if your approach is the PAAD-killer (and I've a ton of companies now using PAAD who would also be delighted by this).

Still want to know when the paper is going to drop!

Thread Thread
 
randalschwartz profile image
Randal L. Schwartz Google Developer Experts •

I love this, Curtis. And don't worry about PAAD needing a "killer"—what you've built in PAAD (/pushback, /rethink, /agentic-review, /agentic-architecture) and what we're doing with Synthetic Scars are kindred spirits arriving at the exact same conclusion from complementary angles.

You nailed the core thesis: "AI slop" is just technical debt accelerated to machine speed because autoregressive models are satisficing engines. They minimize next-token cross-entropy against an internet of beginner tutorials and classroom code where nobody ever had to maintain a system for five years.

To your question on how we proactively stop anti-patterns and prevent them from returning:

1. Stopping the Trap Before It Touches the Buffer (Cognitive Deflection)

Most agent systems treat review as a post-hoc linter (waiting until code is written, then complaining).

In our architecture, the 3-part scar (Wound, Trap, Reflex) acts as a topological barrier in the model's active attention during the planning and thinking phase. The agent loads the domain-relevant scars into its deliberation context before touching any code.

In our empirical research, we inspected over 14,600 thinking turns and mined 78 documented Cognitive Deflections. You can literally watch the model in its internal monologue start drafting the tempting, naive shortcut, collide with the scar, and stop itself:

"Wait. SCAR-STATE-01 forbids ambient symbol overrides because it breaks parallel agent isolation. I must pass the container explicitly via dependency injection instead."

It throws away the naive implementation in its internal scratchpad before writing a single character to the file. That is the computational equivalent of pulling your hand back from the hot stove.

2. Auto-Refactoring Existing Debt (The Clean-Sweep Worktree)

To clean up existing anti-patterns without disrupting feature teams:

  1. Targeted Antigen Extraction: Because a codified scar defines both the exact signature of the bad pattern (The Trap) and the required invariant (The Reflex), we can point a background agent directly at that specific smell in an isolated Git worktree.
  2. Characterization Pinning: The agent's first step is never "start refactoring." It must first author characterization unit tests that pin the existing behavior in place.
  3. Surgical Refactoring: The agent replaces the anti-pattern with the invariant architecture, verifies that all regression tests remain green, and submits a focused, single-purpose PR.
  4. Permanent Immunity: Once that PR merges, the scar remains active in the workflow gates. If an agent working on a brand-new feature three weeks later tries to re-introduce that old pattern, the unprimed Adversarial Critic at the Six Pillars gate immediately raises a BLOCKER and refuses to let the PR advance. The smell is permanently vaccinated against.

When the Paper Drops!

The academic paper—Synthetic Scars: Mitigating Statistical Amnesia and Plausibility Bias in Autonomous Coding Agents via Asymmetric Barrier Topologies and Episodic Consolidation—is currently in final empirical reconciliation across our 51-ticket production field data, mathematical Hamiltonian formulations, and deflection telemetry.

It is scheduled to drop on arXiv and ResearchGate very shortly this fall! I would love to send you an early advance preprint draft to get your eyes on it, especially given how closely your work with PAAD parallels these exact findings.

Let's keep pushing the craft forward—disciplined engineering physics is the only antidote to machine-speed slop!

Collapse
 
jo-do profile image
Jo Do •

The scars framing names something real: frontier models have textbook knowledge but no incident history, and wisdom is compressed incident history. A human senior's "I have a bad feeling about this" is twenty 3 AM pages wearing a trench coat. The hard part with synthetic scars is provenance and decay - a scar from a different codebase, framework version, or era can teach the wrong reflex, and there's no pain signal to tell the agent the scar no longer applies. Humans forget scars when the environment changes; a scar file doesn't forget unless someone maintains it. The teams I see doing this well treat their failure library like a dependency: versioned, reviewed, and pruned, not just appended to. Otherwise you get an agent that's wisely afraid of bugs nobody can hit anymore and blind to the ones that shipped last Tuesday.

Collapse
 
randalschwartz profile image
Randal L. Schwartz Google Developer Experts • • Edited

"A human senior's 'I have a bad feeling about this' is twenty 3 AM pages wearing a trench coat."

That is easily one of the best distillations of engineering intuition I have read in years.

You have put your finger directly on the sharpest operational edge of the entire system: provenance, decay, and the append-only prompt graveyard.

If a team treats synthetic scars as an ever-growing list of global markdown bullets dumped into an agent's system prompt, you inevitably produce what you described: an agent suffering from chronic phantom-limb pain, tiptoeing around workarounds for framework bugs that were patched two years ago while completely blind to breaking changes that landed yesterday.

In our field tests and the formal architecture, we address this in three specific ways:

  1. Version-Bound Provenance (Scars as Scoped Dependencies):
    A scar cannot simply be a generic negative heuristic ("don't use X"). In our registry, every scar requires strict empirical grounding: the exact reproduction harness, the runtime exception, the failing AST pattern, and the explicit dependency/version bounds (for example riverpod < 3.0 or flutter < 3.24). When the underlying toolchain revs past the vulnerable boundary, the scar is out of scope. Treating them like semver-aware dependencies—reviewed, tagged, and scoped to the active project manifest—is essential.

  2. The 4-Stage Promotion Lifecycle (Evicting Scars from the Prompt):
    In the paper's theory section, we model scars through a 4-stage evolutionary pipeline:

    • Stage 0: Episodic runtime panic (the 3 AM page).
    • Stage 1: Semantic barrier (the codified 3-part scar in the agent's active memory).
    • Stage 2: Automated regression harness in CI.
    • Stage 3: Compiler-level diagnostic or custom IDE lint rule (LSP).

The explicit goal is to promote scars out of prompt context as fast as possible. Once an invariant is compiled into a custom analyzer lint rule or upstreamed into the framework's type system, the scar is pruned from the prompt entirely. The agent's cognitive load drops to zero (O(1) ambient defense), freeing the context window from historical baggage.

  1. The Proscriptive Inversion ("More Rules Means You've Lost"): We are actually dedicating Part 6 of this series to what we call the Proscriptive Inversion Principle: if you find your system accumulating more rules about what not to do than what to do, your architecture is failing. Negative rules are high-friction cognitive taxes. When a library or framework redesigns its API so that the invalid state becomes unrepresentable at compile time (what we call the Framework Poka-Yoke Axiom), the entire cluster of historical scars is retired and pruned in one stroke.

Your framing of "versioned, reviewed, and pruned, not just appended to" is exactly the operational discipline that separates a living immune system from an abandoned museum of past traumas.

Thank you for articulating this so cleanly—I may quote your "trench coat" line in Part 2!