Every Monday, my content-planning agent gave me the same advice: "Post consistently, lead with value, try more video." It had read 49 Instagram posts the week before. It remembered none of them.
That's the failure mode nobody warns you about with LLM agents. They're great at reasoning over whatever you paste into the prompt, and they forget everything the moment the call returns. For a content strategist, forgetting is the job failing. The whole value of a human strategist is that they remember which posts flopped, which ones took off, and what the editor keeps crossing out.
So I rebuilt the agent around memory. This is what that looked like, what broke, and the pattern I'd reuse.
What the system does
The agent plans next week's posts for a brand, drafts them, and learns from how they perform and how I edit them. It's a planning workspace, not a chatbot:
- Import: I upload a post-level analytics export (Instagram, LinkedIn, X, TikTok, YouTube, or any CSV/Excel/JSON). Known platforms are detected from their column headers.
- Plan: it proposes three ideas, each with the past posts that justify it.
- Write: it drafts the post. I edit and approve it, and it learns from the diff.
- Agent's Brain: it shows what memory has learned, including a learning curve.
Under the hood it's deliberately boring: Streamlit, SQLite, Groq's gpt-oss-120b for generation, and Hindsight, an open-source agent memory system, for everything the agent needs to remember. Each brand gets exactly one Hindsight memory bank, created the first time you import data and updated in place after that.
The core idea: remember judgments, count in SQL
My first instinct was to dump every post into memory and let the model figure it out. That's the wrong split.
LLM-synthesized memory is excellent at judgment: "how-to posts consistently beat listicles," "readers keep asking about security." It's the wrong tool for arithmetic. In an early version, I asked memory how many posts we'd published in one content pillar, and the synthesized answer drifted from the real count. That's not a flaw in the memory layer; I was asking a summarizer to be a database.
So the rule became: SQLite holds every exact number; Hindsight holds what those numbers mean. Each post is ranked against other posts on the same channel, deterministically, before it's retained. The memory never stores "engagement 0.0826"; it stores "top quarter of Instagram posts by engagement," alongside the post itself.
Here's what a single post looks like on its way into Hindsight:
def memory_item(row: dict, ws: Workspace, text: str | None = None) -> dict:
tags = [f"channel:{row['channel']}"]
tags += [f"{k}:{row[k]}" for k in ("format", "pillar") if row.get(k)]
return {
"content": text or memory_text(row, ws),
"context": "published content and its performance",
"timestamp": datetime.fromisoformat(row["publish_date"]).replace(hour=9, tzinfo=timezone.utc),
"document_id": f"content:{row['content_id']}",
"tags": tags,
"metadata": {k: str(row.get(k) or "") for k in ("content_id", "title", "channel", "format", "pillar", "url")},
"observation_scopes": "shared",
}
Three details in there carry most of the weight:
-
timestampis the publish date, not the import date. That's what lets Hindsight notice that a topic "used to work and doesn't anymore." -
document_idis stable. Re-importing next month's export upserts the same document instead of creating a duplicate. -
observation_scopes: "shared"makes Hindsight consolidate across all posts. With per-tag scopes, "question hooks win" got split into dozens of tiny observations nobody could use.
Once posts are in, Hindsight's consolidation does the part I couldn't have written by hand: it turns 49 separate facts into durable observations, and maintains three mental models per brand: Brand Voice Guide, What Works / What Doesn't, and Topic Coverage Map. The Hindsight docs on memory banks and mental models are worth reading before you design your own; the configuration surface is bigger than it first looks.
Where it got expensive, and the fix
The first working version was correct and wasteful. Building one plan made about 30 recall calls (one per topic in the taxonomy) plus a structured reflect call. That single reflect used 17,260 tokens. Worse, every tiny write (a plan, an edit, a saved setting) triggered consolidation, and consolidation re-ran all three mental models.
The fix was to stop asking memory to think on every click, and start reading what it had already thought:
def _mental_model(bank: str, model_id: str, limit: int) -> str:
"""A mental model's current text. Hindsight pre-computes it on consolidation, so
reading it costs no tokens, unlike a fresh reflect (~17k tokens per call)."""
try:
model = memory.call("get_mental_model", bank_id=bank, mental_model_id=model_id, detail="content")
except Exception:
return ""
text = " ".join((getattr(model, "content", None) or "").split())
return text[:limit] + ("…" if len(text) > limit else "")
The planner now builds a compact packet: exact pillar and format numbers from SQLite, the What Works mental model, and one recall for reader requests. It makes a single LLM call. A plan went from roughly 17k tokens to about 2k.
The mental models themselves got a cheaper refresh policy: edit the existing text instead of regenerating it, and coalesce bursts of writes.
MODEL_TRIGGER = {
"refresh_after_consolidation": True,
"mode": "delta",
"min_refresh_interval_seconds": REFRESH_INTERVAL_SECONDS, # 300
"recall_max_tokens": 3000,
}
The same thinking applied to imports. Every upload used to re-send the entire history. Now each post's memory text is hashed, and only new or changed posts are sent. One more lesson hid in there: a local hash only means "already in memory" if the bank still has the document. If someone deletes a bank, stale hashes silently leave memory half-empty. So before skipping anything, the agent checks what the bank actually holds:
known = db.memory_hashes(ws.db_path) if tracked and not force else {}
if known:
# A hash only counts if the bank still holds that document: if the bank was
# deleted or recreated elsewhere, the missing posts are sent again.
present = memory.document_ids(bank, prefix="content:")
known = {cid: digest for cid, digest in known.items() if f"content:{cid}" in present}
When the next month's export arrives, the preview reads 27 new, 7 updated, 42 already up to date, and 36 of 76 posts are sent. Same bank, no duplicates.
Proving it actually learns
"It has memory" is an easy claim. I wanted a number.
I tested against a sample dataset for a fictional yoga studio, where I knew the true patterns in advance. Reels beat single images. Mindful Movement is the strongest pillar. Nutrition & Recipes has never been posted about. Question hooks win.
Then I replayed the history in date order into a fresh Hindsight bank. At four checkpoints I asked the same question, "What should we publish next week and why?", and had a judge count how many of the real patterns the answer identified:
| Posts in memory | Patterns found |
|---|---|
| 0 | 0 of 5 |
| 5 | 3 of 5 |
| 20 | 4 of 5 |
| 49 | 5 of 5 |
At zero posts, the answer is the generic advice from my opening paragraph. At 49, it's specific.
That same learning shows up in the planner. Its top idea for this brand was a smoothie reel: "Reels achieve the highest engagement at 8.4% and nutrition content is a complete coverage gap." Every card lists the actual past posts behind it. When I say the agent has receipts, that's what I mean.
One edit, remembered
The part that surprised me most was voice.
The first draft came back with emojis, exclamation marks, and a generic call to action. I edited it the way I'd actually publish it: no emojis, no exclamation marks, and a closing line with a real class time. On approval, the agent diffs the draft against my edit, extracts reusable rules, and retains them in Hindsight:
memory.retain(
memory_text,
bank_id=bank,
context="voice feedback",
document_id=f"voice-feedback:{draft_id}",
tags=["voice-feedback", "draft-review", f"workspace:{ws.key}"],
metadata={"draft_id": draft_id, "topic": record["topic"], "decision": status},
)
The next draft, for a different idea, had zero emojis, zero exclamation marks, and ended with "Come try it at the 7am Tuesday class." I never repeated the instruction. It came back from memory.
Lessons I'd reuse
- Split memory by job. Put exact numbers in a database and meaning in agent memory. Retain judgments ("top quarter on Instagram"), not raw metrics.
- Read your mental models; don't re-reflect on every request. Consolidated models are pre-computed. Treating them as a cache cut my planning cost by roughly 8x without losing the learned context.
- Stable document IDs make memory idempotent. Upserts beat duplicates. Use the publish date as the timestamp, so time-based patterns can form.
- Verify what memory actually holds. Local bookkeeping drifts. A cheap document listing before skipping writes keeps memory and your database in sync.
- Measure learning, don't assert it. A replay against a dataset with known answers turned "it has memory" into a curve I could show anyone.
If you're building an agent that should get better the longer you use it, start with what agent memory actually is and the Hindsight GitHub repository. My agent didn't need a bigger prompt. It needed to remember last Monday.
The code for this project is on GitHub: Content Strategy Agent.

Top comments (0)