<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Tanjiro-90</title>
    <description>The latest articles on DEV Community by Tanjiro-90 (@tanjiro90).</description>
    <link>https://dev.to/tanjiro90</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148461%2Fce199175-b0df-44ba-a6b6-ec0cd9a95654.png</url>
      <title>DEV Community: Tanjiro-90</title>
      <link>https://dev.to/tanjiro90</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tanjiro90"/>
    <language>en</language>
    <item>
      <title>What 404s Taught Us About Building on a Memory API</title>
      <dc:creator>Tanjiro-90</dc:creator>
      <pubDate>Tue, 29 Sep 2026 04:52:44 +0000</pubDate>
      <link>https://dev.to/tanjiro90/what-404s-taught-us-about-building-on-a-memory-api-l1g</link>
      <guid>https://dev.to/tanjiro90/what-404s-taught-us-about-building-on-a-memory-api-l1g</guid>
      <description>&lt;p&gt;My first call to Hindsight failed with a 404, and so did my second. I'd guessed a REST path that looked reasonable and wasn't. That set the tone for most of this build: less "design a clever architecture," more "read the docs carefully and handle what a real API actually does."&lt;br&gt;
I'm Antony Sebastian, and I built the backend for Promise-Keeper: a Streamlit app that tracks what you've promised people, briefs you before meetings, and closes the loop when a promise is kept. This is the story of wiring it to Hindsight for memory, and the specific things that broke along the way.&lt;br&gt;
The shape of the system&lt;br&gt;
Promise-Keeper is a Streamlit app that talks to Hindsight's REST API directly, no SDK in between. The flow for one meeting looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Paste in meeting text.&lt;/li&gt;
&lt;li&gt;Gemini extracts the promises made in it.&lt;/li&gt;
&lt;li&gt;Each promise is saved to Hindsight agent memory with a &lt;code&gt;retain&lt;/code&gt; call.&lt;/li&gt;
&lt;li&gt;Before the next meeting, a &lt;code&gt;recall&lt;/code&gt; call pulls the relevant memories back, and Gemini turns them into a brief.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each contact gets their own memory bank, named something like &lt;code&gt;contact_priya_sharma&lt;/code&gt;, created automatically the first time something is saved for them. Nothing about the bank has to be provisioned in advance.&lt;br&gt;
What actually gets written is plain text, one sentence per promise:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;On 2026-09-29, Tony promised Priya Sharma to email the pricing
deck (due: 2026-10-02). Status: OPEN.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When a promise is later fulfilled, that isn't an edit to the old memory; it's a new memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FULFILLED: Tony sent the pricing deck to Priya Sharma. Evidence:
"Thanks for the deck, it's great" (2026-10-02).
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recall handles the merging. A query like "What promises are open or overdue?" run before saving, or "open promises, overdue items, recent interactions" for the brief, returns both the original and the fulfillment note, and Gemini reconciles them into the current status.&lt;/p&gt;

&lt;h2&gt;
  
  
  Problem one: Guessed REST paths
&lt;/h2&gt;

&lt;p&gt;The first version of the retain function pointed at &lt;code&gt;/v1/banks/...&lt;/code&gt;. It looked like a sensible REST path, and it was wrong. The real route is &lt;code&gt;/v1/default/banks/{bank_id}/memories&lt;/code&gt;, and the request body has a specific shape too: &lt;code&gt;{"items": [{"content": ...}]}&lt;/code&gt;.&lt;br&gt;
There's not much of a lesson here beyond the obvious one: read the actual API documentation before writing the client code, rather than pattern-matching from other REST APIs you've used before. A wrong path fails immediately and loudly, which is at least easy to debug. The next few problems were quieter.&lt;/p&gt;
&lt;h2&gt;
  
  
  Problem two: a 404 that isn't an error
&lt;/h2&gt;

&lt;p&gt;The very first recall for a brand-new contact returns a 404 because their bank doesn't exist yet; nobody has saved anything for them. My first instinct was to treat that as a failure. It isn't. It's the expected response for "this person has no memories yet."&lt;br&gt;
The fix was to catch that specific case and return an empty list instead of raising:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;BASE&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/banks/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;bank&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/memories/recall&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                      &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HEADERS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;404&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;  &lt;span class="c1"&gt;# bank doesn't exist yet
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;results&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This matters more than it looks. Without it, the very first meeting with any new contact would crash the app. With it, a new contact just starts with an empty brief, which is exactly correct.&lt;/p&gt;

&lt;h2&gt;
  
  
  Problem three: retains are slow, and recall can miss a fresh save
&lt;/h2&gt;

&lt;p&gt;Hindsight doesn't just store the text you send it; it processes it with an LLM on the way in. That means a &lt;code&gt;retain&lt;/code&gt; call can take several seconds, and a &lt;code&gt;recall&lt;/code&gt; run immediately after a save can occasionally miss something that hasn't finished indexing yet.&lt;br&gt;
Two adjustments came out of this. First, the request timeout on both calls needs to be generous, not the default a typical HTTP client ships with. Second, I stopped assuming retain-then-recall is instant; the UI shows the save completing before offering to generate a brief, rather than chaining the two calls back to back with no gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Problem four: the LLM provider was the flakiest part of the system
&lt;/h2&gt;

&lt;p&gt;Hindsight was, honestly, the reliable half of this stack. The LLM calls were where things kept breaking:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Groq's API returned 403s, blocked at the network level.&lt;/li&gt;
&lt;li&gt;Gemini retired the model I'd started with for new users partway through.&lt;/li&gt;
&lt;li&gt;A 503 "high demand" spike took the API down for a stretch.&lt;/li&gt;
&lt;li&gt;The free tier caps out at 20 requests a day, which is nothing once you're testing.
None of this is about Hindsight, but it directly affected how the backend had to be written, since retain and recall both depend on an LLM call succeeding first (to extract promises, or to write the brief). The fixes:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;code&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;code&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;502&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;503&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;504&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;   &lt;span class="c1"&gt;# busy server: wait and retry
&lt;/span&gt;                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Retrying on 5xx errors with backoff absorbed the transient outages. Merging extraction and fulfillment-checking into a single model call, instead of two, cut the request count roughly in half, which mattered a lot against a 20-request daily cap. And the provider and model are read from a &lt;code&gt;.env&lt;/code&gt; file rather than hardcoded, so switching providers didn't mean editing code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Problem five: the setup itself
&lt;/h2&gt;

&lt;p&gt;Some of the most time lost wasn't API-related at all. A Windows machine needs its PowerShell execution policy adjusted before a virtual environment will activate. A &lt;code&gt;.env&lt;/code&gt; file saved from a text editor can silently become &lt;code&gt;.env.txt&lt;/code&gt;, and the app fails with confusing "missing API key" errors that have nothing to do with the key being wrong. More than once, the version of &lt;code&gt;app.py&lt;/code&gt; actually running wasn't the one I had just edited. None of this is glamorous, but it ate real hours, and it's worth mentioning because it's the kind of friction every teammate hits and nobody writes up.&lt;br&gt;
Lessons learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Don't guess REST paths. A wrong endpoint fails fast, which is the easy case. Read the docs first, and save debugging time for failures that aren't your fault.&lt;/li&gt;
&lt;li&gt;A 404 isn't always an error. An empty memory bank for a new contact is a normal, expected state. Design for it instead of crashing on it.&lt;/li&gt;
&lt;li&gt;Account for processing time. If the memory layer runs an LLM on write, your timeouts and your assumptions about read-after-write both need to account for that.&lt;/li&gt;
&lt;li&gt;The model provider will be your least reliable dependency. Retries, backoff, and making the provider swappable through config cost little and saved the project more than once.&lt;/li&gt;
&lt;li&gt;Environment friction is a real cost. Budget time for it, especially across different operating systems on a team.
If you're integrating an external memory API for the first time, start with its actual documentation rather than assuming its shape, and build your retry and timeout handling before you need it, not after the demo where you needed it most.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>backend</category>
    </item>
  </channel>
</rss>
