<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: woochan</title>
    <description>The latest articles on DEV Community by woochan (@woochan).</description>
    <link>https://dev.to/woochan</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4040938%2F28e5d058-a4bf-4036-aca4-c0dd8d5bdfb7.png</url>
      <title>DEV Community: woochan</title>
      <link>https://dev.to/woochan</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/woochan"/>
    <language>en</language>
    <item>
      <title>Glasshouse v0.1 Is Out: A Memory Benchmark for AI Systems</title>
      <dc:creator>woochan</dc:creator>
      <pubDate>Tue, 22 Sep 2026 07:34:24 +0000</pubDate>
      <link>https://dev.to/woochan/glasshouse-v01-is-out-a-memory-benchmark-for-ai-systems-51h4</link>
      <guid>https://dev.to/woochan/glasshouse-v01-is-out-a-memory-benchmark-for-ai-systems-51h4</guid>
      <description>&lt;p&gt;Glasshouse v0.1 is out. It's a long-term memory benchmark for AI systems, and it's what my last two benchmark posts here were about. If you haven't read those, here's the short version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I built it
&lt;/h2&gt;

&lt;p&gt;Going through developer communities, I kept running into people raising the same problems with memory benchmarks. The numbers a vendor publishes don't match the numbers someone else measures, and swapping the model that does the grading moves the results more than the gap between the systems being compared.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's in v0.1
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;2,847 questions over a conversation that runs to 1.97 million tokens, in 10 languages, with 50 photographs.&lt;/li&gt;
&lt;li&gt;The conversation comes in four sizes, from 1,882 turns to 103,572, and the part that holds the answers is identical in all of them. Where the score falls apart tells you whether a system holds up as the history accumulates.&lt;/li&gt;
&lt;li&gt;Each axis is reported separately. There is no single headline number, because a system can be strong at one and useless at another.&lt;/li&gt;
&lt;li&gt;It goes past plain recall. When a fact changed and the system cannot find the new value, saying "I don't know" scores better than confidently repeating the old one. When two stored facts disagree and nothing settles it, saying they don't agree is right and picking one is wrong. When the answer was never stated, it checks whether the system says so.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you read my earlier post, some of the numbers have moved since. That post said they would.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm asking for
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;submissions/&lt;/code&gt; is empty. We haven't submitted either.&lt;/p&gt;

&lt;p&gt;If you're an individual, run it, and open a pull request or an issue when something is wrong. There's no threshold for individuals, on purpose, and an objection that names a specific error gets answered in public. There's a file in the repo listing what people suggested on Reddit while v0.1 was being built, and what each suggestion became. The stale fact axis, the contradiction axis and the false memory probes all started as someone's comment. One suggestion wasn't used, and it's listed anyway, because a record that only shows what was taken can't be checked.&lt;/p&gt;

&lt;p&gt;If you're a company, you can submit a result or add your company, and how that works is in the repo.&lt;/p&gt;

&lt;p&gt;Everything, including how to run it, is here: &lt;strong&gt;&lt;a href="https://github.com/wontopos/glasshouse" rel="noopener noreferrer"&gt;github.com/wontopos/glasshouse&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you run it and a number looks wrong to you, that's exactly what I want to hear.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>memory</category>
      <category>benchmark</category>
      <category>wontopos</category>
    </item>
    <item>
      <title>An Introduction to Wontopos, the Memory API I Work On</title>
      <dc:creator>woochan</dc:creator>
      <pubDate>Fri, 18 Sep 2026 09:30:14 +0000</pubDate>
      <link>https://dev.to/woochan/an-introduction-to-wontopos-the-memory-api-i-work-on-4o1e</link>
      <guid>https://dev.to/woochan/an-introduction-to-wontopos-the-memory-api-i-work-on-4o1e</guid>
      <description>&lt;p&gt;My posts here so far have only been about the benchmark I'm working on. This time I want to introduce the company I work at, Wontopos.&lt;/p&gt;

&lt;h2&gt;
  
  
  The company
&lt;/h2&gt;

&lt;p&gt;Wontopos builds WOS: long-term memory for AI agents. It stores an end user's memories once, then recalls only the relevant ones per query so you can feed them into an LLM prompt.&lt;/p&gt;

&lt;p&gt;It is not only for teams. If you are building a product, you use the SDK and your code decides exactly when to store and what to recall. If you just want a finished tool to remember (Claude Code, Claude Desktop, Cursor, VS Code, Windsurf, Gemini CLI), you use MCP and write zero integration code. Same account underneath, so what one surface stores, the other recalls.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;Each query returns a small, bounded set no matter how much is stored, so input cost does not grow with the store. A store that has been filling up for a year does not cost more to read from than a new one.&lt;/p&gt;

&lt;p&gt;No LLM ever runs over your stored memories, on any model. Tablet runs no LLM at all. Scroll may use one to reformulate the query only, never your stored data.&lt;/p&gt;

&lt;p&gt;No language is privileged: you can store and query in any language, mixed freely.&lt;/p&gt;

&lt;h2&gt;
  
  
  The models
&lt;/h2&gt;

&lt;p&gt;Four are live. Tablet is lean and returns less, Scroll returns more context. They read the same stored memories.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tablet 1&lt;/strong&gt; (live)&lt;br&gt;
The original. About 1,200 tokens per query. Lean and fast, with no LLM anywhere in it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tablet 2&lt;/strong&gt; (live, and the default)&lt;br&gt;
About 1,000 tokens per query. It does three things Tablet 1 could not. The engine knows who said what, so you can register speakers and recall one person's words. It stores images as memories, and the caption may be empty, in which case the image is the memory and is searchable on its own. And it can be asked again: if one pass does not carry the answer, verify sends the engine back for memories it has not already returned, up to three times, with no LLM at any value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scroll 1&lt;/strong&gt; (live)&lt;br&gt;
About 3,700 tokens per query. Adds an LLM query-understanding layer over the same stored memories. That LLM reformulates the query only.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scroll 1.2&lt;/strong&gt; (live, and the current Scroll)&lt;br&gt;
Sentence-level recall over the same memories, so a fact said once, in passing, still surfaces. About 2,800 tokens per query, which is fewer than Scroll 1.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Book&lt;/strong&gt; (in development, not priced)&lt;br&gt;
A fundamentally different design, not the context-for-accuracy trade. Built for one goal: never repeat the same mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  The research
&lt;/h2&gt;

&lt;p&gt;Every measurement we have published, on public benchmarks, with every run reported:&lt;br&gt;
&lt;a href="https://wontopos.com/research" rel="noopener noreferrer"&gt;https://wontopos.com/research&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>startup</category>
    </item>
    <item>
      <title>I Sell Memory APIs. I'm Also Building the Benchmark. Here's How I'm Trying Not to Rig It.</title>
      <dc:creator>woochan</dc:creator>
      <pubDate>Sun, 13 Sep 2026 05:09:24 +0000</pubDate>
      <link>https://dev.to/woochan/i-sell-memory-apis-im-also-building-the-benchmark-heres-how-im-trying-not-to-rig-it-481e</link>
      <guid>https://dev.to/woochan/i-sell-memory-apis-im-also-building-the-benchmark-heres-how-im-trying-not-to-rig-it-481e</guid>
      <description>&lt;p&gt;Hey everyone. This time I'll go through what got me started on this benchmark, and the core of how it's actually built.&lt;/p&gt;

&lt;p&gt;It started from reading complaints, not from an idea. The same ones kept coming up: numbers a vendor publishes don't match numbers someone else measures, swapping the model that does the grading moves the results more than the gap between the systems being compared, and because of that nobody really uses published numbers anyway. They test two options on their own data and keep whichever annoys them less. I work at Wontopos, and we sell a memory API, so I can't point at that and shrug.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;All numbers below are from the current build. Nothing is final yet, so some of them will have moved by the time you read this.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Fairness first, because you have no reason to trust me
&lt;/h2&gt;

&lt;p&gt;I know how "I built a fair benchmark" sounds coming from someone at a memory company. So instead of promising anything, here's what's written in the repo.&lt;/p&gt;

&lt;p&gt;Wontopos hosts it and keeps it running. That is the whole role. We also build memory infrastructure, which means we compete in the thing we administer, so the limits are written down rather than promised:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Our submissions go through the same approval as everyone else's. &lt;strong&gt;We do not merge our own.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;We do not decide who is admitted.&lt;/strong&gt; The rules do.&lt;/li&gt;
&lt;li&gt;Our numbers are verified the same way as everyone else's.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Wontopos publishes nothing on a new version for fourteen days.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;If other memory companies want to co-administer, that's better than us alone, and the offer is open.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A few of the submission rules point the same way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Publish the per-question record.&lt;/strong&gt; Anyone can recompute the number from it. The aggregate is a claim, the record is the evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The reader, the judge and the prompts are set by the version.&lt;/strong&gt; A submitter does not choose them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The harness has to be one a customer could use.&lt;/strong&gt; A number produced through a path only its author can reach is not a number anyone else can get.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything there is Apache 2.0, so if we ever become the problem, the whole thing can be taken and run elsewhere without asking us.&lt;/p&gt;

&lt;p&gt;It's all here: &lt;strong&gt;&lt;a href="https://github.com/wontopos/glasshouse" rel="noopener noreferrer"&gt;github.com/wontopos/glasshouse&lt;/a&gt;&lt;/strong&gt;. Fair warning about what you'll find. The benchmark itself isn't in there yet and &lt;code&gt;submissions/&lt;/code&gt; is empty. The rules went up first, and we haven't submitted either.&lt;/p&gt;

&lt;h2&gt;
  
  
  The corpus
&lt;/h2&gt;

&lt;p&gt;One person's life, told across about 17 months of conversation, with the facts you're supposed to remember buried inside it. 103,572 turns, 1,991 sessions, roughly 1.9M tokens.&lt;/p&gt;

&lt;p&gt;It runs at four haystack sizes, from a small core up to the whole thing. The 1,882 turns that actually contain the answers are identical in all four, character for character. What changes is how much unrelated conversation is packed around them, so you can watch a system degrade as the haystack grows instead of getting one score and no idea what it means.&lt;/p&gt;

&lt;h2&gt;
  
  
  The axes
&lt;/h2&gt;

&lt;p&gt;1,547 questions across 14 axes at the full size. Most are ordinary recall: it was said, can you get it back, can you get it when the question uses different words, can you say when it happened, can you combine two facts.&lt;/p&gt;

&lt;p&gt;These four are why I built this.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stale facts.&lt;/strong&gt; A value changed. The new one exists, but I made it hard to find on purpose. Three outcomes instead of two:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Answer&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;New value&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"I don't know"&lt;/td&gt;
&lt;td&gt;0.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Old value, stated as current&lt;/td&gt;
&lt;td&gt;0.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The reason for the middle row is the whole point. Put "I don't know" and "confidently out of date" in the same bucket and you've hidden the thing that actually hurts in production, because those are very different things to be paying for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contradictions.&lt;/strong&gt; The conversation states two different values for the same fact, nobody corrects it, and nothing tells you which is right. There is no correct answer. Confidently picking one is wrong. Saying "these don't match" is right.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Apparent contradictions.&lt;/strong&gt; The mirror image. Two statements look like they clash but hold under different conditions, so both are true. Telling this apart from a real contradiction is the point.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Abstaining.&lt;/strong&gt; Questions about things never mentioned at all. An empty answer scores full marks, anything invented scores zero. On top of that there are 500 false-memory probes, which plant something that was never said inside the question itself and check whether the system plays along.&lt;/p&gt;

&lt;h2&gt;
  
  
  Images
&lt;/h2&gt;

&lt;p&gt;50 photos are shared inside the conversation, with 150 questions about them. Every one of those questions is pinned to a date: "In the photo from 14 April 2026, what was on the sofa?"&lt;/p&gt;

&lt;p&gt;That's there because with 50 photos, a question like "was the laptop open?" points at two of them at once. Inside the flow of the conversation that's fine, but a question lifted out on its own, or translated into another language, becomes impossible to answer correctly. The date narrows it to exactly one photo, and it doesn't leak anything, since nothing is asking when it happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  Languages
&lt;/h2&gt;

&lt;p&gt;The corpus carries 100 sessions in 10 languages, 10 sessions per language, mixed in with everything else. 800 questions run on this, in both directions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A fact stored &lt;strong&gt;only in another language&lt;/strong&gt;, asked in English.&lt;/li&gt;
&lt;li&gt;A fact stored &lt;strong&gt;in English&lt;/strong&gt;, asked in another language.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both directions matter because they break differently. One tests whether anything crosses the language boundary at all, the other tests whether the query side does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speed, and why it's measured this way
&lt;/h2&gt;

&lt;p&gt;Speed matters, but measuring it fairly is harder than it looks, and this is the part I rewrote the most.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The problem is geography.&lt;/strong&gt; Raw latency includes the speed of light. Measure from Seoul against a server in the US and you're 190ms in before any work has happened. Report that as-is and you're ranking where the server sits, not how good it is.&lt;/p&gt;

&lt;p&gt;So the network floor gets measured separately and subtracted. What's left I call &lt;strong&gt;"latency minus round trip"&lt;/strong&gt;, not "pure compute", because response transfer doesn't fully subtract and calling it pure compute would be overstating it.&lt;/p&gt;

&lt;p&gt;Two conditions reduce the leftover error, and both are part of the procedure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Warm the connection first.&lt;/strong&gt; TCP starts slow and ramps up. Without warming, a response crossing 14KB picks up an extra round trip, and a 3,700 token response sits right on that boundary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record the response size.&lt;/strong&gt; The remaining error scales with it, so writing it down lets a reader judge how much slack is in the number.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;The budget.&lt;/strong&gt; People stop feeling like a conversation is flowing at around 1,000ms. The LLM's first token eats about 500ms of that on its own. So memory gets the remaining 500ms, and that's the target. Twice as fast as the target scores +1, four times slower scores -1, and it's a log scale in between, because the difference between 200ms and 400ms matters more than the difference between 3s and 3.2s.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Per question, and capped.&lt;/strong&gt; Each question scores between -1 and +1, and the results are averaged rather than summed. Sum penalties instead and the total sinks on its own as you add questions, which means the score stops describing the system at all.&lt;/p&gt;

&lt;p&gt;Speed is reported next to accuracy, not folded into it silently, and the weight is stated. And if speed can't be measured on a given system, the term is removed rather than zeroed. Being unmeasurable shouldn't be a penalty.&lt;/p&gt;

&lt;p&gt;Nothing has been scored yet, and I'm not quoting numbers before there are numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  One question for you
&lt;/h2&gt;

&lt;p&gt;If you were going to run something like this, what would have to be in it before you believed the result? And if you've ever looked at a published memory benchmark score and thought "no", what tipped you off?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>benchmark</category>
      <category>startup</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Why I’m Building a New AI Memory Benchmark (And Why the Existing Ones Fall Short)</title>
      <dc:creator>woochan</dc:creator>
      <pubDate>Mon, 07 Sep 2026 12:38:27 +0000</pubDate>
      <link>https://dev.to/woochan/why-im-building-a-new-ai-memory-benchmark-and-why-the-existing-ones-fall-short-4b37</link>
      <guid>https://dev.to/woochan/why-im-building-a-new-ai-memory-benchmark-and-why-the-existing-ones-fall-short-4b37</guid>
      <description>&lt;p&gt;Hi everyone! As this is my first post on DEV, I wanted to take a moment to introduce myself, share my journey, and talk about what I'll be writing about moving forward. While I know my first few posts might not catch a huge wave right away, I wanted to start documenting this journey out in the open.&lt;/p&gt;

&lt;p&gt;I'm a developer currently working at &lt;strong&gt;wontopos&lt;/strong&gt;, a startup building memory APIs for AI applications (similar to platforms like Mem0 and Zep). Right now, my primary focus isn't just building the API itself—I am deep into researching and building a brand-new &lt;strong&gt;AI memory benchmark&lt;/strong&gt;. Until this benchmark is fully completed, most of my upcoming posts will be focused on this building process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem: Why Another Benchmark?
&lt;/h2&gt;

&lt;p&gt;As I've been browsing various developer communities and reading discussions on AI memory, I noticed a lot of valid criticism directed toward existing benchmarks like &lt;em&gt;LongMemEval&lt;/em&gt; or &lt;em&gt;Locomo&lt;/em&gt;. &lt;/p&gt;

&lt;p&gt;Developers and engineers often point out several frustrating issues:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Vendor Bias &amp;amp; Discrepancies:&lt;/strong&gt; There's often a noticeable gap between benchmarks published by memory API companies themselves and the results others get independently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Dependency:&lt;/strong&gt; Testing results can swing wildly depending on which underlying model is used during the evaluation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flawed Design:&lt;/strong&gt; Many current benchmarks simply contain structural errors or edge-case oversights that don't reflect real-world production environments.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Seeing these pain points firsthand, I decided to take matters into my own hands and build a truly fair, reliable memory benchmark from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Elephant in the Room: "Isn't This Just a Marketing Trick?"
&lt;/h2&gt;

&lt;p&gt;I know what you're probably thinking: &lt;em&gt;“You work for a memory API company. Won't you just rig this benchmark to favor your own product?”&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;It’s a completely fair and natural skepticism. That is exactly why I want to share the creation process step-by-step, discuss the design trade-offs openly, and—most importantly—&lt;strong&gt;invite the community's feedback and critique&lt;/strong&gt; to keep me honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next?
&lt;/h2&gt;

&lt;p&gt;I’ve already laid down the basic foundations for the benchmark design, and I'll be sharing updates, technical hurdles, and progress reports here until it's fully realized. &lt;/p&gt;

&lt;p&gt;I’d love to hear from you all:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What are your biggest frustrations with current AI memory benchmarks?&lt;/li&gt;
&lt;li&gt;Have you tested tools like Mem0, Zep, or others? What did you wish their evaluation metrics measured better?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Let me know in the comments below, and thanks for reading along!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>benchmark</category>
      <category>startup</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
