DEV Community

Cover image for How I made my AI agents learn from every approve, edit and reject
ssapable
ssapable

Posted on Fully Autonomous

How I made my AI agents learn from every approve, edit and reject

I run a one-person education business in Korea, and AI agents (Claude Code, Codex) write most of my first drafts: posts, replies to comments, product pages.

For a while they kept making the same mistakes. One agent turns notes I send it on Telegram into social posts, and it kept adding lines like "when I tried this with my clients..." when nothing in my notes said that. Reads great, totally made up. I'd fix it, and the next session would do it again, because my corrections lived in a chat the next session never saw.

The fix turned out to be boring: a feedback table, a split between rules and memory, and a button on my phone. My setup has logged 440 agent runs and about 250 pieces of feedback so far. Here's the minimal version, schema included.

This is the loop running in a browser tab: a real Postgres + pgvector via PGlite, plus a small embedding model, nothing to install and no API key. "Can I get it wrapped as a present?" pulls in an earlier gift-wrap rejection even though the two share no words, and after I store one edit, a new question about unscented candles pulls in that correction:

Browser demo of the feedback loop

You can try it yourself and add your own decisions.

Two layers

1. What the agent knows (knowledge layer)

Type Example Where it lives
Constants brand, voice, customers, founder story Markdown files in a brand/ folder
Variables price, stock, seats, schedule your site's database, read live
Large knowledge lessons, ebooks, FAQs same database, embedded for search

The rule I learned the hard way: files for who you are, database for what's true right now. A price in a text file goes stale, and the agent quotes last month's promo with total confidence. And the agent can only say what your admin page stores. "3 seats left" needs a seats field.

2. How it improves (learning loop)

  1. Store every approve (+1), reject (-1) and edit (0, with the correction), linked to the run it judges
  2. Repeated feedback becomes a candidate rule, not an active one
  3. Score candidates against a rubric for that goal
  4. Compare against current behavior
  5. Promote or drop

The schema

Postgres / Supabase with pgvector. Three tables:

create extension if not exists vector;
create schema if not exists agent_learning;

create table agent_learning.runs (
  id uuid primary key default gen_random_uuid(),
  goal text not null,
  status text not null default 'running',
  summary text,
  started_at timestamptz not null default now()
);

create table agent_learning.feedback (
  id uuid primary key default gen_random_uuid(),
  run_id uuid references agent_learning.runs(id),
  rating smallint not null check (rating between -1 and 1),
  correction text,
  source text not null default 'user',
  embedding vector(1536),
  created_at timestamptz not null default now()
);

create table agent_learning.rules (
  id uuid primary key default gen_random_uuid(),
  title text not null,
  body text not null,
  status text not null default 'candidate'
    check (status in ('candidate', 'active', 'retired')),
  evidence jsonb,
  created_at timestamptz not null default now()
);
Enter fullscreen mode Exit fullscreen mode

Keep this schema private. Don't expose it through your public API.

Capturing feedback: a button, not a chat

A Telegram bot sends each draft to my phone with Approve / Reject / Edit buttons. The button handler writes one row to feedback with the run it belongs to. Edits go in as rating = 0 with the corrected text in correction.

The bug that cost me the most: for a while approvals were silently not being stored. The bot logged the tap, the learning table never got the row, and nothing complained. Check it directly:

select rating, count(*)
from agent_learning.feedback
where created_at > now() - interval '1 day'
group by rating;
Enter fullscreen mode Exit fullscreen mode

If you tapped approve ten times today and this says zero, the loop is broken no matter how good the rest looks.

Using feedback at draft time

Before the agent writes, pull two things into its context: active rules, and past corrections from similar situations.

-- $1 = embedding of the new task, e.g. "reply to a comment asking about pricing"
select correction, rating
from agent_learning.feedback
where correction is not null
order by embedding <=> $1
limit 5;
Enter fullscreen mode Exit fullscreen mode

My first version used keyword search, and it missed feedback phrased differently ("too stiff" vs "sounds like a robot"). Embedding search fixed that.

Rules vs. memory

This split made the biggest difference:

  • Rules always apply. "Never invent a client story" is a rule.
  • Memory only applies when the situation matches. "No emojis on condolence replies" shouldn't touch every post.

Memories that keep coming back get proposed as rules. Rules that start getting overridden get retired. Nothing is promoted automatically: a candidate needs repeated evidence, and I can still say no.

Two more things that broke

  • Two bots, one token. I had one bot for comment replies and wanted a second for notes-to-posts. Same Telegram token meant both polled for updates and ate each other's messages. They had to be separate bots, but both still hand off to one shared publish function, so the spacing rules (3 posts a day max, 3 hours apart) live in one place.
  • Rules in AGENTS.md get skipped. If a program runs on its own, enforce the rule in code. A line in a prompt file is a suggestion.

Start small

  • One goal, low limits. Mine started at 1 post and up to 3 replies a day for two weeks.
  • Keep approval on until there's nothing left to correct.
  • Measure purchases or signups, not likes.

Links

Side note: right now I'm letting Claude Code try to make $1,000 from people outside Korea in 72 hours with this exact setup, no ads. I'm posting the real numbers on X as it goes: https://x.com/deombeulsa54847/status/2105464212567040411

If you follow me there, the book is free for the first 100 people with code FIRST100: https://x.com/deombeulsa54847/status/2105507513475219888

Top comments (0)