The first thing my extraction prompt got wrong was the word "I."
Someone would say "I'll send the pricing deck by Friday," and my agent would faithfully record: owner "I", due date "Friday". Neither is useful a week later. This article is about the small prompt changes that fixed that, and about why the agent's memory in Hindsight made a much simpler prompt design possible.
I own the prompts for Promise-Keeper, an agent that tracks what you've promised people, briefs you before you meet them, and closes the loop when a promise is kept.
What the system does
You paste a meeting transcript. The agent does three things:
- Extracts the new commitments made in it: who promised what, to whom, by when.
- Checks which earlier open promises were fulfilled in this same transcript.
- Saves and recalls everything through Hindsight agent memory, so the next brief starts from what the agent already knows.
I didn't want to stuff raw chat history into a context window and hope. Hindsight is where the agent's knowledge lives, and the UI shows which memories were recalled for each brief. My teammate covers the memory and backend design in their article.
The prompt
Here is the whole thing. It lives in app.py, and the values in braces are filled in by the code before each call:
The user is {me}. The other person is {name}. The meeting date is {date};
convert relative dates like 'Friday' into YYYY-MM-DD based on it.
Previously open promises:
{open_text}
New transcript:
{transcript}
Return ONLY JSON with this shape:
{"fulfilled": [{"promise": "...", "evidence": "..."}],
"commitments": [{"owner": "", "recipient": "", "task": "",
"due_date": "YYYY-MM-DD or Unspecified"}]}.
'fulfilled' lists only earlier promises clearly completed in this transcript.
'commitments' lists NEW promises made in this transcript; use real names
instead of 'I'.
It's short on purpose. Four decisions in it did most of the work.
Decision one: two jobs in one call
Extracting new commitments and checking old ones are two separate tasks, and the obvious design is two separate model calls. Merging them into one prompt with two output fields halved the number of requests. That mattered because the model API has a daily request cap.
The merge only works because both jobs read the same input. The transcript is the evidence for new promises and for fulfilled ones, so there's no reason to send it twice.
Decision two: JSON only, with a fixed shape
The code needs to act on the output: save new commitments, mark old ones done, render the brief. So the prompt demands JSON only, and states the exact shape. I also set response_format: json_object on the request, so the API enforces valid JSON rather than relying on the model to behave.
A fixed shape also makes the output easy to inspect. If fulfilled is empty, I know the model found nothing. It can't hide a guess inside a paragraph.
Decision three: ground the model in who and when
This was the change that fixed the "I" problem. Two lines of context at the top of the prompt:
Who the user is. With
The user is {me}. The other person is {name}, the model can resolve "I'll send" to a real name, and the prompt tells it to use real names instead of "I". The stored commitment says who owes what, and that stays true a month later, in a brief a different person might read.This doesn't always hold. The Memory Inspector sometimes shows the same
promise saved with the owner as "User" or "Speaker" instead of the real
name, which means the model isn't reading{me}consistently on every
call. It's a grounding problem, not a formatting one — worth fixing
before I trust the owner field completely.When the meeting happened. With the meeting date in the prompt, "Friday" becomes a real date. For a Tuesday meeting on 2026-09-29, it becomes
2026-10-02. Without a date, the model has nothing to anchor a relative word to, so it copies it.
Neither needed a bigger model or a longer prompt. The model was already capable of doing this and simply lacked the context.
Decision four: "only clearly completed"
The fulfilled field is the risky one. A wrong "done" is worse than a stale "open": a stale promise gets noticed and corrected, while a promise wrongly marked complete quietly disappears from the brief.
So the prompt says 'fulfilled' lists only earlier promises clearly completed in this transcript, and asks for an evidence string with each match. That biases the model toward leaving things open when the transcript is ambiguous. I'd rather the agent remind me about something that's already done than stay silent about something that isn't.
Where Hindsight fits
The Previously open promises block is what makes the fulfilled check possible. Without it, "were any promises fulfilled?" is an open-ended question and the model has nothing to compare against. With it, the model gets a closed list to match the transcript against.
That list comes from Hindsight. The two calls that talk to it are small:
def retain(bank, text):
r = requests.post(f"{BASE}/banks/{bank}/memories",
json={"items": [{"content": text}]},
headers=HEADERS, timeout=60)
return r.status_code == 200
def recall(bank, query):
r = requests.post(f"{BASE}/banks/{bank}/memories/recall",
json={"query": query}, headers=HEADERS, timeout=60)
if r.status_code == 404:
return [] # bank doesn't exist yet
return [m.get("text", "") for m in r.json().get("results", [])]
retain writes a commitment. recall reads back what's relevant to a query. Note the 404 branch: a bank that doesn't exist yet returns an empty list, so the very first meeting with a new contact works with no special-casing. The prompt just sees an empty Previously open promises block, and fulfilled comes back empty.
What gets saved is plain text, not a database row. A stored memory looks like this:
Tony promised to email the pricing deck to Priya Sharma by October 2, 2026. | When: 2026-09-29 | Involving: Tony, Priya Sharma
The sentence, the meeting date, and the people involved all live in one string, so recall can return it as a readable line and the Live Memory Inspector can show it as-is.
Example interaction
Meeting note on 2026-09-29, with Tony as the user and Priya Sharma as the other person:
"I'll email the pricing deck to Priya by Friday."
The model returns a commitment with the owner as Tony, the recipient as Priya Sharma, the task as emailing the pricing deck, and the due date as 2026-10-02. That's saved to Hindsight as the memory shown above. The next time I prep for a meeting with Priya, the promise comes back as an open item in the brief.
The near-miss matters too. If a later transcript says "I still haven't sent Priya the deck," the promise must stay open. That's what the "clearly completed" wording in the prompt is for.
What I'd still improve
The prompt has two states for a promise: completed or not mentioned. It has no explicit "unclear" state. When a transcript hints that something was done but doesn't say so, the model has to choose, and the "clearly completed" wording pushes it toward leaving the promise open. That's the safe failure, but a third state would let the brief say "possibly done, confirm?" instead. That's the next change I want to make.
The second gap is duplicates. Nothing checks whether a promise is already stored before retain writes it, so the same promise can end up in memory more than once, and you can see it in the Memory Inspector. The fix belongs before the write: recall first, and skip anything that already matches.
Name resolution isn't fully reliable.** Even with {me} and {name}
in the prompt, some saved memories still say "User" or "Speaker" instead
of the real name. I haven't isolated why — whether it's a specific
phrasing that confuses the model, a missing value on some calls, or
something else — but it's the reason the owner field in the Memory
Inspector isn't consistent yet.
Lessons learned
1. Context beats cleverness. Two lines telling the model who the user is and what date it is fixed more than any rewording did.
2. Merge jobs that read the same input. Extraction and the fulfilled check share a transcript, so one call halves the request count.
3. Fix the output shape. JSON only, an exact schema, and response_format: json_object let the code parse the result reliably.
4. Bias toward the safe failure. A stale "open" is recoverable and a false "done" isn't, so the prompt says "only clearly completed."
5. Give the model a list to match against. Recalling open promises from Hindsight first turns an open-ended question into a closed one.
If you're building an agent that acts on what it remembers, start with the Hindsight documentation and decide what gets written to memory before you tune the prompt.
This project was built as part of a build "Sponsored by Code.in"



Top comments (0)