I run a one-person education business in Korea, and AI agents (Claude Code, Codex) write most of my first drafts: posts, replies to comments, product pages.
For a while they kept making the same mistakes. One agent turns notes I send it on Telegram into social posts, and it kept adding lines like "when I tried this with my clients..." when nothing in my notes said that. Reads great, totally made up. I'd fix it, and the next session would do it again, because my corrections lived in a chat the next session never saw.
The fix turned out to be boring: a feedback table, a split between rules and memory, and a button on my phone. My setup has logged 440 agent runs and about 250 pieces of feedback so far. Here's the minimal version, schema included.
This is the loop running in a browser tab: a real Postgres + pgvector via PGlite, plus a small embedding model, nothing to install and no API key. "Can I get it wrapped as a present?" pulls in an earlier gift-wrap rejection even though the two share no words, and after I store one edit, a new question about unscented candles pulls in that correction:
You can try it yourself and add your own decisions.
Two layers
1. What the agent knows (knowledge layer)
| Type | Example | Where it lives |
|---|---|---|
| Constants | brand, voice, customers, founder story | Markdown files in a brand/ folder |
| Variables | price, stock, seats, schedule | your site's database, read live |
| Large knowledge | lessons, ebooks, FAQs | same database, embedded for search |
The rule I learned the hard way: files for who you are, database for what's true right now. A price in a text file goes stale, and the agent quotes last month's promo with total confidence. And the agent can only say what your admin page stores. "3 seats left" needs a seats field.
2. How it improves (learning loop)
- Store every approve (+1), reject (-1) and edit (0, with the correction), linked to the run it judges
- Repeated feedback becomes a candidate rule, not an active one
- Score candidates against a rubric for that goal
- Compare against current behavior
- Promote or drop
The schema
Postgres / Supabase with pgvector. Three tables:
create extension if not exists vector;
create schema if not exists agent_learning;
create table agent_learning.runs (
id uuid primary key default gen_random_uuid(),
goal text not null,
status text not null default 'running',
summary text,
started_at timestamptz not null default now()
);
create table agent_learning.feedback (
id uuid primary key default gen_random_uuid(),
run_id uuid references agent_learning.runs(id),
rating smallint not null check (rating between -1 and 1),
correction text,
source text not null default 'user',
embedding vector(1536),
created_at timestamptz not null default now()
);
create table agent_learning.rules (
id uuid primary key default gen_random_uuid(),
title text not null,
body text not null,
status text not null default 'candidate'
check (status in ('candidate', 'active', 'retired')),
evidence jsonb,
created_at timestamptz not null default now()
);
Keep this schema private. Don't expose it through your public API.
Capturing feedback: a button, not a chat
A Telegram bot sends each draft to my phone with Approve / Reject / Edit buttons. The button handler writes one row to feedback with the run it belongs to. Edits go in as rating = 0 with the corrected text in correction.
The bug that cost me the most: for a while approvals were silently not being stored. The bot logged the tap, the learning table never got the row, and nothing complained. Check it directly:
select rating, count(*)
from agent_learning.feedback
where created_at > now() - interval '1 day'
group by rating;
If you tapped approve ten times today and this says zero, the loop is broken no matter how good the rest looks.
Using feedback at draft time
Before the agent writes, pull two things into its context: active rules, and past corrections from similar situations.
-- $1 = embedding of the new task, e.g. "reply to a comment asking about pricing"
select correction, rating
from agent_learning.feedback
where correction is not null
order by embedding <=> $1
limit 5;
My first version used keyword search, and it missed feedback phrased differently ("too stiff" vs "sounds like a robot"). Embedding search fixed that.
Rules vs. memory
This split made the biggest difference:
- Rules always apply. "Never invent a client story" is a rule.
- Memory only applies when the situation matches. "No emojis on condolence replies" shouldn't touch every post.
Memories that keep coming back get proposed as rules. Rules that start getting overridden get retired. Nothing is promoted automatically: a candidate needs repeated evidence, and I can still say no.
Two more things that broke
- Two bots, one token. I had one bot for comment replies and wanted a second for notes-to-posts. Same Telegram token meant both polled for updates and ate each other's messages. They had to be separate bots, but both still hand off to one shared publish function, so the spacing rules (3 posts a day max, 3 hours apart) live in one place.
- Rules in AGENTS.md get skipped. If a program runs on its own, enforce the rule in code. A line in a prompt file is a suggestion.
Start small
- One goal, low limits. Mine started at 1 post and up to 3 replies a day for two weeks.
- Keep approval on until there's nothing left to correct.
- Measure purchases or signups, not likes.
Links
- Try the loop in your browser (Postgres + pgvector running in the tab, nothing to install): https://ssap-pa.github.io/self-learning-agent-setup/
- Repo with the schema, a one-minute demo you can run with no database (
cd demo && npm install && npm run demo), and the prompts I used to get the agents to build this themselves: https://github.com/ssap-pa/self-learning-agent-setup - Same loop as an MCP server for Claude Code, Cursor or Claude Desktop (local Postgres, nothing to host):
claude mcp add feedback-memory -- npx -y github:ssap-pa/self-learning-agent-setup - The complete working version (Python and TypeScript library + CLI, Telegram approval bot, rule proposals, an MCP server on your own Postgres / Supabase, 27 tests): https://payhip.com/b/grc25
- Free 23-page field guide with a worked example and a 7-day plan: https://payhip.com/b/QHgfJ
- A free 22-minute video lesson where I build this system with an agent (Korean audio, English subtitles, no account): https://ssapable.com/courses/ai-agent?lang=en&utm_source=devto&utm_medium=article1#free-preview
- The book this comes from, Just Say "Do It" (English edition of my Korean course): https://payhip.com/b/xmZvu
Side note: right now I'm letting Claude Code try to make $1,000 from people outside Korea in 72 hours with this exact setup, no ads. I'm posting the real numbers on X as it goes: https://x.com/deombeulsa54847/status/2105464212567040411
If you follow me there, the book is free for the first 100 people with code FIRST100: https://x.com/deombeulsa54847/status/2105507513475219888

Top comments (0)