DEV Community

Cover image for Distill Coding Agent Learnings
Lars Moelleken
Lars Moelleken

Posted on Edited on

Distill Coding Agent Learnings

Repo: https://github.com/voku/agent-loop

Website: https://voku.github.io/agent_loop_demo/

Your Coding Agent Can Write Code. Keep the Work Reliable.

Coding agents are already surprisingly good at changing code.

The harder problem is everything around the change.

What exactly was requested?

What is the agent allowed to touch?

Which repository facts matter?

Which validation actually ran?

Does that validation still belong to the current implementation?

What should happen next?

And when something useful was learned, should it influence the next task, become durable guidance, or eventually disappear into a test, static-analysis rule, or CI check?

The obvious response has often been:

Give the agent more memory.

So repositories slowly acquire things like:

MEMORY.md
project-rules.md
agent-notes.md
lessons-learned.md
MEMORY_FINAL.md
actually-final-memory.md
Enter fullscreen mode Exit fullscreen mode

Soon the coding agent receives old decisions, temporary workarounds, copied transcripts, stale debugging notes, abandoned architecture ideas, and rules nobody remembers approving.

It has more context.

It does not necessarily have better context.

At some point, memory becomes landfill.

That was the problem I wanted to solve with voku/agent-loop.

Not by building another coding agent.

By building a governed workflow around the coding agent you already use.


The everyday loop should stay small

The internals of Agent Loop are deliberately more sophisticated than its ordinary host contract.

For normal governed work, the coding agent starts here:

vendor/bin/agent-loop enter PROJECT-123 --format=json
Enter fullscreen mode Exit fullscreen mode

The result exposes the important state:

mutation_ready
next_action_kind
next_action
Enter fullscreen mode Exit fullscreen mode

The host follows that result.

It performs the current work.

Then:

vendor/bin/agent-loop finish PROJECT-123 --format=json
Enter fullscreen mode Exit fullscreen mode

And again follows the returned next action until the workflow converges.

Conceptually:

developer intent
      ↓
agent-loop enter <task>
      ↓
canonical next_action
      ↓
coding agent works
      ↓
agent-loop finish <task>
      ↓
validation · review · learning
      ↓
repeat or complete
Enter fullscreen mode Exit fullscreen mode

That is the important part.

The coding agent does not need to carry a giant workflow phase machine in its prompt.

It should not have to memorize:

PLAN
MAP
SESSION
RECALL
IMPLEMENT
VALIDATE
REVIEW
LEARN
VERIFY
CLOSE
Enter fullscreen mode Exit fullscreen mode

Those capabilities exist.

But they are implementation details behind the lifecycle.

agent-loop owns the current workflow state and tells the host what happens next.


Why this matters

Without a governed workflow, coding-agent sessions tend to fail in remarkably repeatable ways.

A new chat means reconstructing the task.

The agent guesses whether it should investigate, implement, ask for approval, validate, or review.

Large repositories become giant prompts.

A passing test quietly becomes “done”.

Useful experience disappears after the task.

And project instructions keep growing because deleting prompt text feels more dangerous than adding one more paragraph.

Agent Loop tries to invert those defaults.

The task survives the chat.

Context is selected instead of dumped.

Evidence is tied to the state it actually verified.

The next action is explicit.

And useful learning has somewhere to go other than another permanent Markdown file.


Humans approve intent, not shell commands

A governed workflow should not turn developers into operators of an orchestration CLI.

The developer should still be able to talk naturally to the coding agent.

For example:

Fix PROJECT-123.

Invalid order state transitions must be rejected.
Keep the public API unchanged.
Stay inside the order-state validation and its focused tests.
Run PHPStan and the focused test suite.
Enter fullscreen mode Exit fullscreen mode

The coding agent investigates the repository and prepares a candidate Contract:

Goal:
Reject invalid order state transitions.

Scope:
Order-state validation and focused tests.

Non-goals:
Do not redesign the order aggregate.
Do not change the public API.

Validation:
composer phpstan
composer test
Enter fullscreen mode Exit fullscreen mode

When the lifecycle reaches a real human-authority boundary, the agent presents that exact decision.

The developer can answer:

Approved.
Enter fullscreen mode Exit fullscreen mode

Or:

No. Keep the repository layer untouched.
Enter fullscreen mode Exit fullscreen mode

Or:

Approved, but do not change the public API.
Enter fullscreen mode Exit fullscreen mode

The coding host records that explicit decision through Agent Loop.

Approval belongs to one concrete Contract revision.

It is not permanent permission attached to a task ID whose meaning can silently change later.

If product intent or approved scope materially changes, the Contract changes too.

That is where human authority is required again.

But ordinary implementation, deterministic validation, review acknowledgement, Learning disposition, Recall outcome logging, local commits, and close-out inside the approved Contract do not need ceremonial confirmation every thirty seconds.

Human authority should stay explicit without becoming human babysitting.


The workflow tells the agent what kind of action comes next

A returned next action is not just a line of text.

It has a kind.

For example:

command
command_template
decision_required
host_work
none
Enter fullscreen mode Exit fullscreen mode

command

means the workflow already knows the exact deterministic operation.

command_template

means the coding agent must fill task-specific values from the request and current repository evidence.

host_work

means actual engineering work is required.

decision_required

means a real human-authority boundary has been reached.

none

means there is no further lifecycle action.

This removes a surprising amount of prompt choreography.

The host does not need to guess:

Is this something I can do myself?

or:

Should I ask the developer now?

or:

Which internal phase comes next?

The lifecycle result already answers that.

Internally, Agent Loop projects the current owner-backed facts and evaluates them through executable policy.

That produces the current:

state
mutation_ready
next_action_kind
next_action
Enter fullscreen mode Exit fullscreen mode

The coding agent follows that result instead of maintaining another slightly different copy of the workflow in prose.


The CLI is the executable contract, not the user experience

Agent Loop exposes a substantial CLI.

You can discover it directly:

vendor/bin/agent-loop help
vendor/bin/agent-loop commands
vendor/bin/agent-loop commands --format=json
vendor/bin/agent-loop commands --format=toon
Enter fullscreen mode Exit fullscreen mode

The top-level command catalog is typed and also drives human-readable help.

That is intentional.

Humans and coding agents should discover the same installed capabilities from the same owner instead of maintaining separate README tables, UI registries, and prompt copies that inevitably drift apart.

But despite the larger command surface, the normal host contract stays small:

enter
  ↓
follow next_action
  ↓
work
  ↓
finish
Enter fullscreen mode Exit fullscreen mode

Specialist commands exist when the current task needs them.

They are capabilities.

They are not mandatory phases.


Skills help perform work. They do not invent workflow state.

Earlier versions of this project leaned more heavily on phase-specific workflow instructions.

That architecture made duplication too easy.

A task-start skill knew one sequence.

A review skill knew another.

A README explained it again.

Eventually there are three correct descriptions of the same workflow, which is software engineering’s traditional method for preparing three future bugs.

The current boundary is stricter.

The lifecycle result remains authoritative.

Skills help the agent perform the current kind of engineering work.

Agent Loop can project focused assets for activities such as:

agent-loop-discipline
agent-loop-workflow
agent-loop-investigate
agent-loop-surgical-edit
agent-loop-code-review
agent-loop-simplify-review
Enter fullscreen mode Exit fullscreen mode

An investigation role is read-only.

A surgical-edit role is for an already-understood bounded change and must escalate instead of quietly widening scope.

A review role inspects the real diff rather than trusting a generated summary.

Broader engineering guidance lives separately in voku/agent-skills, with reusable skills for areas such as:

engineering-codelight
coding-simplicity
php-best-practices
code-review-*
Enter fullscreen mode Exit fullscreen mode

The ownership boundary is intentional:

agent-loop
    workflow authority

agent-skills
    reusable engineering judgment
Enter fullscreen mode Exit fullscreen mode

The agent-skills catalog is optional and separately installed.

Agent Loop does not silently download another engineering rulebook merely because somebody thought seventeen more prompt files looked comforting.


Give the agent less context, but better context

A coding agent does not need the whole repository in its prompt.

It needs enough evidence to make the next engineering decision correctly.

Agent Loop separates repository structure from project guidance because they answer different questions.

Repository structure with agent-map

agent-map provides deterministic PHP repository intelligence.

It can answer questions such as:

Where is this method defined?
Who calls it?
What does it call?
Which symbols are related?
What is the bounded edit context?
What may be affected by changing it?
Enter fullscreen mode Exit fullscreen mode

It can operate structurally and optionally enrich the map through PHPStan when semantic evidence is available.

The important output is not a gigantic graph to admire during architecture meetings.

It is a bounded, source-backed edit context.

And Map output remains derived navigation evidence.

It does not replace the source code.

If the Map says a method has no callers but the current capability cannot observe call edges, that is a capability limit.

It is not proof that the method has no callers.

That difference matters.

Task-specific context with agent-recall-compiler

agent-recall-compiler answers a different question:

Which bounded project knowledge is relevant to this task?

Inputs can include:

approved task intent
repository facts
project guidance
constraints
previous outcomes
LearningNote precedent
Enter fullscreen mode Exit fullscreen mode

The result is replayable task context plus, when applicable, L2 construction material for a concrete project-specific L1 contract.

That L1 typically needs to answer:

Goal
Context
Constraints
Verification
Done When
Enter fullscreen mode Exit fullscreen mode

The point is selection.

Not permanent context inflation.

The goal is not:

Compress the whole repository and six months of transcripts into the model.

The goal is:

Give the agent enough current, verified information to make the next decision correctly.


Deterministic preparation should be automatic

Once a Contract has been approved, enter can prepare the governed execution state.

Conceptually:

approved Contract
      ↓
reconcile required Map state
      ↓
prepare or resume Session
      ↓
prepare governed Run
      ↓
replace stale Recall output
      ↓
compile current Recall
      ↓
return bounded context
      ↓
authorize implementation when ready
Enter fullscreen mode Exit fullscreen mode

The host should not have to memorize that ordering.

If repository discovery can be repaired deterministically, the workflow can repair it.

If an optional capability is unavailable, that does not automatically become a human decision.

This distinction is useful:

Deterministic preparation is choreography, not governance.

If software can safely do it itself, forcing the developer or coding agent to remember another ritual adds ceremony without adding authority.


Working memory should stay temporary

During a task, the coding agent may need mutable state:

assumptions
decisions
checkpoints
validation observations
open questions
resume state
Enter fullscreen mode Exit fullscreen mode

That belongs in agent-session.

A Session is working memory.

It is not permanent project knowledge.

This sounds obvious, but memory systems tend to blur this distinction surprisingly quickly.

A debugging observation may be useful for twenty minutes and misleading two commits later.

Temporary state should remain closable and pruneable.

The durable workflow should preserve only the state and evidence that actually need to survive.


Validation belongs to an exact implementation

An agent saying:

Tests passed.

is useful.

But only if we know which implementation those tests validated.

Validation evidence is bound to the task and exact workflow state, including:

Contract revision
command
exit code
implementation snapshot
Enter fullscreen mode Exit fullscreen mode

Suppose this passed:

composer test
Enter fullscreen mode Exit fullscreen mode

Then the agent changes three files.

The old result is history.

It is not evidence that the new implementation still passes.

The same applies when the approved Contract revision changes.

The system can therefore distinguish:

current validation
superseded by implementation
superseded by Contract revision
missing validation
failed validation
Enter fullscreen mode Exit fullscreen mode

instead of carrying one immortal green checkmark around the repository.


finish does not mean “please declare victory”

After implementation, the coding agent calls:

vendor/bin/agent-loop finish PROJECT-123 --format=json
Enter fullscreen mode Exit fullscreen mode

finish does not simply flip a task to done.

It reconciles the current implementation with the current obligations.

Depending on state, it may:

run validation
prepare review evidence
request more host work
surface a human decision
record Learning disposition
log Recall outcomes
close the Session
complete the Run
surface remaining owner work such as Kanban reconciliation
Enter fullscreen mode Exit fullscreen mode

That last distinction is important.

If the governed Run is complete but the linked Kanban card is still active, Agent Loop does not quietly claim everything is finished.

It surfaces the remaining Kanban-owned work.

Evidence stays evidence.

Owner state stays owner state.

The gate ordering itself lives in executable policy rather than another copied checklist.

If validation fails, the canonical next action may be:

next_action_kind = host_work

change the implementation so the declared validation passes
Enter fullscreen mode Exit fullscreen mode

Running the same failed validation forever would not make progress.

Which gives us another useful invariant:

A canonical next action should converge.

If executing the recommended deterministic action always returns the exact same blocker and exact same recommendation, that is a workflow defect.

The coding agent should not invent a private workaround.


Evidence is not authority

This principle appears repeatedly because violating it is an excellent way to build an impressive automation system that is subtly wrong.

A test result is evidence.

It does not approve a scope change.

A review report is evidence.

Its existence does not automatically acknowledge the review.

A generated prompt proves something was rendered.

It does not prove the receiving model used it.

A Map is derived repository structure.

It is not source truth.

A LearningNote is evidence-backed precedent.

It is not active project policy.

And a model sounding extremely confident is mostly evidence that models have become very good at adjectives.

Authority and evidence answer different questions.

The workflow keeps them separate.


Selected guidance is not necessarily useful guidance

Suppose Recall selects a piece of guidance.

That proves:

the system selected it
Enter fullscreen mode Exit fullscreen mode

It does not prove:

the agent consulted it
the agent needed it
the agent applied it
the result improved because of it
Enter fullscreen mode Exit fullscreen mode

Outcome tracking therefore keeps selection and usefulness separate.

A selected item may eventually be judged:

helpful
irrelevant
harmful
not used
unknown
Enter fullscreen mode Exit fullscreen mode

This becomes essential once you start asking whether all that carefully curated prompt material is actually worth its token cost.

Otherwise, you are mostly measuring how often your software successfully copied text into another file.

Computers were already very good at that before LLMs arrived.


Learning should begin with evidence, not policy

This is the part of Agent Loop I find most interesting.

The naive model is:

The agent learned something useful. Save it as a rule.

That promotes evidence far too aggressively.

The current architecture is closer to:

real task
   ↓
Finding
   ↓
LearningNote
   ↓
future Recall
   ↓
later real task
   ↓
observable result
   ↓
new Finding
   ↓
recurrence
   ↓
Dream
   ↓
Proposal
   ↓
human review
   ↓
Memory / Skill / Constraint
   ↓
deterministic enforcement
Enter fullscreen mode Exit fullscreen mode

There are really two loops here.

A fast precedent loop.

And a slower promotion loop.


The fast loop: Finding → LearningNote → future task

A Finding records evidence from real work.

It can preserve things such as:

task identity
observation
validated conclusion
scope
evidence
pattern_key
validation case
Enter fullscreen mode Exit fullscreen mode

A useful solved case may be classified:

ADD_LEARNING_NOTE
Enter fullscreen mode Exit fullscreen mode

That can become a LearningNote.

The important boundary is:

A LearningNote is durable precedent, not active guidance.

It can preserve:

what happened
what failed
why the resolution worked
when to apply it
when not to apply it
how to verify it
Enter fullscreen mode Exit fullscreen mode

A later task may receive that precedent through Recall.

But the note does not:

approve mutation
widen task scope
satisfy validation
create a Skill
create a Constraint
become project policy
Enter fullscreen mode Exit fullscreen mode

It says:

We have seen something like this before. Here is the evidence-backed solved case.

That is much safer than promoting every successful patch into permanent instruction text.


The slow loop: recurrence earns promotion

One solved case rarely deserves policy.

Before promoting a lesson, we want stronger evidence.

Did the pattern recur?

Did it recur across independent tasks?

Was the precedent actually useful?

Was it irrelevant?

Was it harmful?

Is there already an owner for this behavior?

Has later evidence contradicted it?

Could software enforce the invariant instead?

agent-learning provides a deterministic maintenance pass called dream.

When using it through Agent Loop, the workflow owner resolves the configured Learning root for you:

vendor/bin/agent-loop learn dream \
  --report=.agent-loop/dream/latest.json \
  --dry-run
Enter fullscreen mode Exit fullscreen mode

Dream can evaluate accumulated evidence for things such as:

promotion
staleness
replacement
conflicts
outcome coverage
duplicate candidate decisions
Enter fullscreen mode Exit fullscreen mode

It can prepare reviewable candidates.

It does not approve them.

It does not silently rewrite project guidance while everybody sleeps.

A scheduler, if you use one, may decide when to wake this maintenance cycle.

Learning decides whether the evidence warrants a candidate.

A human remains the authority for durable promotion.

Wake frequently.

Promote rarely.


Most tasks should probably learn nothing permanent

This is a feature.

A perfectly good task-level outcome is:

NO_DURABLE_LEARNING
Enter fullscreen mode Exit fullscreen mode

Sometimes the correct conclusion after fixing a bug is:

The bug is fixed.

No new Memory entry.

No new Skill.

No small constitution drafted because one null check was missing on a Wednesday.

Durable Learning should have a higher bar than task completion.


The best memory is sometimes a PHPStan rule

Suppose repeated evidence eventually establishes a stable invariant:

Every X must satisfy Y.
Enter fullscreen mode Exit fullscreen mode

If that property can be detected mechanically, the ideal destination may not be another prompt instruction.

It may be:

PHPStan rule
PHP-CS-Fixer rule
PHPCS sniff
architecture test
unit test
integration test
CI check
typed API
runtime invariant
Enter fullscreen mode Exit fullscreen mode

Then:

repeated evidence
      ↓
reviewed Constraint
      ↓
deterministic enforcement
      ↓
obsolete prompt guidance can disappear
Enter fullscreen mode Exit fullscreen mode

This is where “memory” becomes much more interesting.

The system can distill experience into structure.

Once CI proves the invariant automatically, spending prompt tokens reminding the agent about it on every task is waste.

The lesson has become infrastructure.

The agent can forget it.


One semantic owner per kind of truth

The project is split across focused packages because different kinds of state should have clear owners.

Concern Owner
Contract / Run authority and lifecycle agent-loop
Git-native work items agent-kanban
Temporary working memory and validation evidence agent-session
Repository structure and code navigation agent-map
Bounded task context and prompt construction agent-recall-compiler
Findings, precedent and durable Learning agent-learning
Portable engineering and review guidance agent-skills

There are optional surfaces too.

agent-ui provides a local human control plane.

agent-loop-runner provides an optional execution plane for isolated coding-host runs.

Internally, agent-graph provides shared deterministic SQLite graph mechanics for Map and Learning.

It deliberately owns none of their domain semantics.

That distinction matters.

The UI should not independently decide what action is legal.

The Runner should not decide that process exit 0 means the governed Run is valid.

A graph database should not become the semantic owner because it contains derived relations.

The boring rule is:

One semantic owner per decision.

It becomes increasingly useful once multiple autonomous tools start touching the same repository.


This is still pre-1.0

One claim deserves particular honesty.

The individual mechanics of the Learning loop exist today:

Finding capture
LearningNotes
bounded Recall selection
L2 / L1 construction
outcome tracking
Dream
reviewed Proposals
Constraints
deterministic enforcement
Enter fullscreen mode Exit fullscreen mode

But I am still dogfooding the harder end-to-end claim:

Does precedent from one real task measurably improve a later independent task, and can repeated evidence eventually become a hard rule that catches a different manifestation of the same mistake?

The project has already proven substantial parts of that path.

What I do not want to do is declare success merely because a LearningNote appeared inside generated context.

Prompt inclusion is not behavioral evidence.

If the same engineering decision would have happened without the precedent, the honest result is:

NO_DEMONSTRATED_VALUE
Enter fullscreen mode Exit fullscreen mode

or:

WITHHELD
Enter fullscreen mode Exit fullscreen mode

That experiment is still active.

I would rather leave the claim open than build a learning system whose main learned behavior is how to congratulate itself.


Try it on one real task

Agent Loop is open source, MIT licensed, local-first, and requires PHP 8.3+.

Install it in an existing Composer repository:

composer require --dev voku/agent-loop
Enter fullscreen mode Exit fullscreen mode

Create the repository-local workflow scaffold:

vendor/bin/agent-loop init scaffold --demo
Enter fullscreen mode Exit fullscreen mode

Install the local coding-agent assets.

For Codex:

vendor/bin/agent-loop init install-assets --agent=codex
vendor/bin/agent-loop init doctor
Enter fullscreen mode Exit fullscreen mode

Then start a fresh Codex session or restart Codex before using the newly installed skills.

That boundary is worth stating explicitly.

install-assets proves that the repository-side files were projected correctly.

It does not prove that an already-running coding host retroactively loaded them.

You can inspect the host-side repository integration with:

vendor/bin/agent-loop init host-status --format=json
Enter fullscreen mode Exit fullscreen mode

Other supported hosts can be selected instead of codex.


Start the governed task

For the demo:

vendor/bin/agent-loop enter DEMO-1 --format=json
Enter fullscreen mode Exit fullscreen mode

Read:

mutation_ready
next_action_kind
next_action
Enter fullscreen mode Exit fullscreen mode

Then do what the current lifecycle asks.

For an unplanned task, that may first be a model-owned Contract template.

The coding agent fills it using:

the developer's request
current repository evidence
the smallest honest scope
repository-supported validation commands
Enter fullscreen mode Exit fullscreen mode

When the result reaches:

decision_required
Enter fullscreen mode Exit fullscreen mode

the host presents the exact human decision.

After approval, call enter again.

Once mutation is authorized, the coding agent implements normally.

Then:

vendor/bin/agent-loop finish DEMO-1 --format=json
Enter fullscreen mode Exit fullscreen mode

Follow the result.

Repeat until complete.

That is the ordinary workflow.

Everything else exists to make those decisions reliable.


The actual interface is still conversation

In day-to-day use, I do not want to manually translate every lifecycle transition into CLI syntax.

I want to tell the coding agent:

Take PROJECT-123.

Reject invalid order state transitions.
Do not change the public API.
Run PHPStan and the focused tests.
Enter fullscreen mode Exit fullscreen mode

The coding agent should inspect the repository, operate the workflow, show me the exact Contract when my authority is required, implement within that boundary, collect evidence, and report the result.

If scope must change, I want that surfaced.

If validation fails, I want the actual failure.

If a Finding is reusable, I want it recorded.

If the task teaches nothing durable, I want the system to say so.

And if repeated evidence eventually produces a rule that PHPStan can enforce, I would rather delete the prompt instruction than celebrate that MEMORY.md grew again.


What changed in my mental model

When I started working on this problem, I thought mostly in terms of memory:

How can the coding agent remember useful things?

That question is too broad.

A better sequence is:

What happened?
    ↓
Is there evidence?
    ↓
Is it useful precedent?
    ↓
Does it recur?
    ↓
Should a human approve it as guidance?
    ↓
Can software enforce it instead?
    ↓
Can we delete the prose again?
Enter fullscreen mode Exit fullscreen mode

That produces a very different system.

Temporary state stays temporary.

Historical precedent stays weaker than current authority.

Validation remains bound to the implementation it actually tested.

Workflow policy lives in executable owners instead of duplicated prompts.

Human decisions stay explicit.

And durable learning has somewhere to go other than an ever-growing pile of Markdown.


The actual lesson

Coding agents do not need to remember everything.

They need:

  • durable task intent;
  • explicit authority boundaries;
  • bounded repository context;
  • temporary working memory;
  • validation tied to the current implementation;
  • one canonical next action;
  • evidence-backed precedent;
  • recurrence before promotion;
  • human-reviewed durable guidance;
  • deterministic enforcement when possible;
  • intentional forgetting.

More memory can hide bad context management.

A governed loop exposes it.

The purpose of Agent Loop is not to make developers manually operate a more elaborate state machine.

It is to give coding agents a reliable process they can operate themselves while humans retain authority over the decisions that matter.

The agent can still investigate, edit, test, review, and learn.

But:

I remember
I think
looks good
Enter fullscreen mode Exit fullscreen mode

stop being workflow states.

And once a lesson becomes stable enough that PHPStan, a test, a typed API, or CI can enforce it, the ideal outcome is not that the coding agent remembers it forever.

The ideal outcome is that it no longer has to.

Repo: https://github.com/voku/agent-loop

Demo: https://voku.github.io/agent_loop_demo/

Top comments (3)

Collapse
 
kgaidev profile image
kgaidev •

We build kgai, so I'm biased toward the memory side of this, but the landfill line is right. Most "agent memory" fails at capture quality, and your governed loop attacks exactly that: a finding isn't guidance until a human says so.

The part I keep chewing on is what happens after promotion. Rules don't just accumulate, they expire. The constraint that justified a rule disappears, or two devs promote contradictory guidance on the same codebase. I saw learning-boundary has REPLACE/DELETE proposals: when guidance gets replaced, do you keep the dead rule plus the reason it died, or does it just vanish? In our experience, the "why it stopped being true" is exactly what an agent needs to see to avoid re-learning the same bad rule three months later.

Also curious what triggers a REPLACE in practice. A task hitting conflicting guidance, or periodic review?

Collapse
 
suckup_de profile image
Lars Moelleken •

I strongly believe that the learning loop should be integrated into every programming workflow. Even though some people might think that the next model will be better anyway, this helps with the review process because you take away prompts, tasks, findings, and so on from the programming session, which allows you to better understand the changes yourself.

Fixed that now :) <-> The decision engine that compiles recall guidance already blocked on a rejected proposal sharing a target with newly selected guidance, quoting the rejection reason. It now runs the same check against retired proposals ... if someone proposes guidance that re-targets a rule that was retired for cause, in the same task scope, compilation blocks and quotes the retirement reason back. So the "re-learning the same bad rule" case you're describing is now a hard stop, not just an audit trail someone could go read if they thought to look.

P.S.: in my case, most of the stuff that are extracted from the learnings into skills or new rules for static analysis come from semi-manual reviews or from comments like "Hint: This is a learning ..." ... and not from the tool outout itself

Collapse
 
kgaidev profile image
kgaidev •

That's exactly the right semantics. A retired rule without its retirement reason is just an absence, quoting the reason back at resurrection time is the whole value. And your P.S. matches what we see with kgai too. The useful stuff comes from the conversation, not from tool output, which is what pushed us to capture in the session itself. The model records the decision as it's made, and a Stop hook at end of turn catches anything it forgot. Enjoyed this thread.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.