<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Angelo Pantano</title>
    <description>The latest articles on DEV Community by Angelo Pantano (@ghilteras).</description>
    <link>https://dev.to/ghilteras</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4130533%2Fb299cd0a-6a2e-4f97-bd76-9f645e330acd.jpg</url>
      <title>DEV Community: Angelo Pantano</title>
      <link>https://dev.to/ghilteras</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ghilteras"/>
    <language>en</language>
    <item>
      <title>A decision proxy in front of SearXNG: speculative execution, per-engine circuit breakers, and keyless fallback</title>
      <dc:creator>Angelo Pantano</dc:creator>
      <pubDate>Fri, 18 Sep 2026 19:03:50 +0000</pubDate>
      <link>https://dev.to/ghilteras/a-decision-proxy-in-front-of-searxng-speculative-execution-per-engine-circuit-breakers-and-4knj</link>
      <guid>https://dev.to/ghilteras/a-decision-proxy-in-front-of-searxng-speculative-execution-per-engine-circuit-breakers-and-4knj</guid>
      <description>&lt;h1&gt;
  
  
  TL;DR
&lt;/h1&gt;

&lt;p&gt;&lt;code&gt;searxng-gateway&lt;/code&gt; is a Go HTTP proxy that sits in front of SearXNG and adds multi-provider fallback (Brave, Exa, Jina, Tavily), a per-engine circuit breaker, an in-memory LRU cache, Prometheus metrics, and a zero-key mode. It preserves SearXNG's JSON response shape, so clients do not change. The interesting parts are the execution model, the breaker semantics, and the trade-offs — including a couple that are easy to get wrong.&lt;/p&gt;

&lt;p&gt;What it adds, in one list:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Speculative execution&lt;/strong&gt; — SearXNG and a configurable number of premium providers start together; premium calls are serial within that pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bounded fallback&lt;/strong&gt; — if the merged result count is below a threshold, remaining providers are tried round-robin until the time budget expires.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-engine circuit breakers&lt;/strong&gt; — one 4xx-class error opens that engine for five minutes, then a single probe decides recovery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability&lt;/strong&gt; — Prometheus metrics plus an importable Grafana dashboard and example alert rules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keyless mode&lt;/strong&gt; — runs on SearXNG's free engines with no API keys at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rest of this post covers why the pieces are shaped this way, and where the design still costs you.&lt;/p&gt;

&lt;h1&gt;
  
  
  Why a proxy in front of SearXNG
&lt;/h1&gt;

&lt;p&gt;Self-hosted search looks simple from the outside: run SearXNG, ask it for JSON, get results. In practice the free engines behind it do not fail in a coordinated way. One engine starts returning 403, another gets rate-limited, a third serves a captcha challenge, and a fourth simply gets slower until it times out. SearXNG does report which engines were unresponsive, but it does not make a per-request decision about what to do next. The aggregate answer degrades quietly: fewer results, more latency, and no clear owner for the fix.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;searxng-gateway&lt;/code&gt; is a small Go service that sits in front of a SearXNG instance and makes those decisions. It is open source under MIT, currently at v2.6.3, maintained at github.com/Ghilteras/searxng-gateway. It was originally derived from sx and adds an HTTP server, a per-engine circuit breaker, Prometheus metrics, an in-memory LRU cache, and Docker packaging.&lt;/p&gt;

&lt;h1&gt;
  
  
  A SearXNG-shape-compatible layer
&lt;/h1&gt;

&lt;p&gt;The gateway is not a replacement for SearXNG; it is a layer in front of it. Clients keep talking to a single endpoint and keep receiving the same JSON shape SearXNG produces, with fields such as &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;url&lt;/code&gt;, &lt;code&gt;content&lt;/code&gt;, &lt;code&gt;engine&lt;/code&gt;, and &lt;code&gt;engines&lt;/code&gt;. Premium provider responses are normalised into that shape at merge time, so a caller does not need to know whether a result came from SearXNG or from the Brave API. The client never talks to SearXNG directly.&lt;/p&gt;

&lt;p&gt;Three endpoints are exposed: &lt;code&gt;GET /search?q=&amp;lt;query&amp;gt;&amp;amp;format=json&lt;/code&gt; for queries, &lt;code&gt;GET /healthz&lt;/code&gt; for liveness, and &lt;code&gt;GET /metrics&lt;/code&gt; for Prometheus exposition, with the metrics path configurable through &lt;code&gt;METRICS_PATH&lt;/code&gt;.&lt;/p&gt;

&lt;h1&gt;
  
  
  The execution model, precisely
&lt;/h1&gt;

&lt;p&gt;This is the part most write-ups get wrong, so it is worth stating carefully. The pipeline runs in stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The query is normalised (lowercased, whitespace collapsed) and checked against the LRU cache. A cache hit returns immediately without calling any backend.&lt;/li&gt;
&lt;li&gt;SearXNG's cooldown state is checked. If SearXNG has hit its consecutive-failure threshold, it is skipped for the whole request.&lt;/li&gt;
&lt;li&gt;SearXNG starts concurrently with the premium pass. The gateway picks &lt;code&gt;T1_PREMIUM_COUNT&lt;/code&gt; providers by atomic round-robin and invokes them serially within that pass. Premium providers are not called in parallel with each other; the concurrency is between SearXNG and the premium pass as a whole.&lt;/li&gt;
&lt;li&gt;Results from SearXNG and the premium pass are merged and deduplicated by URL.&lt;/li&gt;
&lt;li&gt;If the merged count is below &lt;code&gt;SUFFICIENT_MIN_RESULTS&lt;/code&gt;, a bounded fallback loop tries the remaining providers by round-robin, one at a time, until the threshold is met, the providers are exhausted, or &lt;code&gt;FALLBACK_TIMEOUT_SECONDS&lt;/code&gt; expires.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The outcome of each request is recorded under a fixed set of labels on &lt;code&gt;searxng_gateway_requests_total&lt;/code&gt;: &lt;code&gt;cache_hit&lt;/code&gt; (served from cache, no backend called), &lt;code&gt;searxng_ok&lt;/code&gt; (only SearXNG contributed results), &lt;code&gt;premium_ok&lt;/code&gt; (only premium providers contributed, meaning SearXNG was skipped, errored, or returned nothing), &lt;code&gt;searxng_plus_premium_ok&lt;/code&gt; (both contributed), and &lt;code&gt;fallback_fail&lt;/code&gt; (everything failed with no results). There is also a &lt;code&gt;timeout&lt;/code&gt; label, incremented when SearXNG returns a deadline-exceeded error; unlike the others it is recorded in addition to the final outcome rather than replacing it, so a single request can increment both.&lt;/p&gt;

&lt;p&gt;One subtlety worth knowing before you build dashboards on these labels: "contributed" means a provider's results actually survived URL dedup into the merged set. The rule is symmetric across providers, but it is not order-free. If two providers return an identical URL set, whichever one is merged first supplies the URLs and the other is recorded as contributing nothing.&lt;/p&gt;

&lt;h1&gt;
  
  
  Per-engine circuit breakers
&lt;/h1&gt;

&lt;p&gt;Each premium provider and each SearXNG engine gets its own circuit breaker, built on &lt;code&gt;sony/gobreaker&lt;/code&gt;. The trip rule is deliberately aggressive: a single 4xx-class client error opens the circuit immediately. Here a "4xx-class" error is broader than the HTTP status code. It includes 403, 429, access-denied, "too many requests", "blocked by", and captcha messages. The reasoning is that these signals usually mean the server is telling you to stop, so retrying makes things worse.&lt;/p&gt;

&lt;p&gt;Once open, the engine is excluded from subsequent requests. After a five-minute timeout the circuit goes half-open and sends a single probe; success closes it and increments the recovery counter, while failure reopens it for another five minutes. SearXNG itself is handled separately with a binary cooldown counter rather than &lt;code&gt;gobreaker&lt;/code&gt;: after &lt;code&gt;SEARXNG_FAIL_THRESHOLD&lt;/code&gt; consecutive failures it is skipped entirely for &lt;code&gt;SEARXNG_FAIL_COOLDOWN_SECONDS&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;SearXNG calls also get retries with exponential backoff, up to three attempts. Notably, that retry path makes no 4xx/5xx distinction: all error classes are retried, and the per-attempt metrics (&lt;code&gt;searxng_gateway_retry_attempts_total&lt;/code&gt;, &lt;code&gt;searxng_gateway_retry_exhausted_total&lt;/code&gt;) exist so the retry path stays visible even when the first attempt succeeds.&lt;/p&gt;

&lt;h1&gt;
  
  
  Cache, metrics, and a Grafana dashboard
&lt;/h1&gt;

&lt;p&gt;The cache is an in-memory LRU with 1000 entries and a one-hour TTL by default, both configurable. It only helps repeated queries, and it is per-process, so it does not survive a restart and is not shared across replicas.&lt;/p&gt;

&lt;p&gt;All metrics are prefixed &lt;code&gt;searxng_gateway_&lt;/code&gt;. Beyond request outcomes, the interesting ones are per-engine result counts (&lt;code&gt;searxng_gateway_engine_results_total&lt;/code&gt;), unresponsive reasons reported by SearXNG (&lt;code&gt;searxng_gateway_engine_unresponsive_total&lt;/code&gt;), last-seen engine status (&lt;code&gt;searxng_gateway_engine_status&lt;/code&gt;), retry counters, cache size, and the circuit-breaker family: state (0 closed, 1 half-open, 2 open), trips by reason, rejections, and recovery events. There is also a set of request-window quota gauges. For Brave specifically these are &lt;code&gt;searxng_gateway_brave_rate_limit_remaining&lt;/code&gt;, &lt;code&gt;searxng_gateway_brave_rate_limit_limit&lt;/code&gt;, and &lt;code&gt;searxng_gateway_brave_rate_limit_reset_seconds&lt;/code&gt;; they are parsed from the &lt;code&gt;X-RateLimit-Remaining&lt;/code&gt;, &lt;code&gt;X-RateLimit-Limit&lt;/code&gt;, and &lt;code&gt;X-RateLimit-Reset&lt;/code&gt; response headers, so they describe the request window the API reports, not an account credit or dollar balance.&lt;/p&gt;

&lt;p&gt;A reference Grafana dashboard ships in the repository at &lt;code&gt;examples/grafana/searxng-gateway-dashboard.json&lt;/code&gt;, with panels for circuit-breaker state per engine, cumulative trips coloured by reason, recoveries, cache hit rate, and cache size. The README also includes example alert rules, for instance firing when a circuit stays open and when fallback API usage spikes. One caveat on the bundled screenshots: they date from July 2026 and predate an outcome-label rename, so a few legends show older label names. The dashboard JSON is the current source of truth.&lt;/p&gt;

&lt;h1&gt;
  
  
  Keyless mode
&lt;/h1&gt;

&lt;p&gt;The gateway runs with no API keys at all. In keyless mode it uses SearXNG's free engines, including Bing, Wikipedia, Wikidata, GitHub, StackOverflow, ArXiv, PyPI, Docker Hub, Mwmbl, and Marginalia, and still applies the circuit breaker, retry, and caching. What you lose is Google results, since Serper needs a key, and the Brave fallback. Adding keys later is just a matter of setting environment variables.&lt;/p&gt;

&lt;h1&gt;
  
  
  Quickstart
&lt;/h1&gt;

&lt;p&gt;The repository ships a reference stack that runs SearXNG and the gateway together. The gateway image tag is pinned, so upgrade it deliberately rather than tracking latest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;searxng&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;searxng/searxng:latest&lt;/span&gt;
    &lt;span class="na"&gt;container_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;searxng&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;unless-stopped&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./examples/searxng/settings.example.yml:/etc/searxng/settings.yml:ro&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./examples/searxng-engines/serper.py:/usr/local/searxng/searx/engines/serper.py:ro&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./examples/searxng-engines/mojeek_api.py:/usr/local/searxng/searx/engines/mojeek_api.py:ro&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;SERPER_API_KEY=${SERPER_API_KEY:-}&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;127.0.0.1:8081:8080"&lt;/span&gt;

  &lt;span class="na"&gt;searxng-gateway&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io/ghilteras/searxng-gateway:v2.6.3&lt;/span&gt;
    &lt;span class="na"&gt;container_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;searxng-gateway&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;unless-stopped&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8080:8080"&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;SEARXNG_BACKEND_URL=http://searxng:8080&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then query it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s1"&gt;'http://localhost:8080/search?q=hello+world&amp;amp;format=json'&lt;/span&gt;
curl &lt;span class="s1"&gt;'http://localhost:8080/metrics'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note that the SearXNG image above tracks an upstream tag and is not digest-pinned, so a quickstart run is not byte-for-byte reproducible — the gateway image is the pinned half.&lt;/p&gt;

&lt;h1&gt;
  
  
  Configuration
&lt;/h1&gt;

&lt;p&gt;Everything is configured through environment variables. The most consequential ones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;SEARXNG_BACKEND_URL&lt;/code&gt;: where the SearXNG instance lives.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;FALLBACK_PROVIDERS&lt;/code&gt;: comma-separated premium provider names (default &lt;code&gt;brave&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;T1_PREMIUM_COUNT&lt;/code&gt;: how many providers run in the hot path alongside SearXNG; zero means none.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;SUFFICIENT_MIN_RESULTS&lt;/code&gt;: the merged-result target that stops the fallback loop.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;FALLBACK_TIMEOUT_SECONDS&lt;/code&gt;: the overall budget for the speculative pass plus the fallback loop.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;SEARXNG_FAIL_THRESHOLD&lt;/code&gt; and &lt;code&gt;SEARXNG_FAIL_COOLDOWN_SECONDS&lt;/code&gt;: the SearXNG cooldown.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;CACHE_SIZE&lt;/code&gt; and &lt;code&gt;CACHE_TTL_SECONDS&lt;/code&gt;: cache tuning.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;LOG_LEVEL&lt;/code&gt; and &lt;code&gt;METRICS_PATH&lt;/code&gt;: operational knobs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Premium providers need their own keys (&lt;code&gt;BRAVE_API_KEY&lt;/code&gt;, &lt;code&gt;EXA_API_KEY&lt;/code&gt;, &lt;code&gt;JINA_API_KEY&lt;/code&gt;, &lt;code&gt;TAVILY_API_KEY&lt;/code&gt;), and each consumes that provider's quota.&lt;/p&gt;

&lt;h1&gt;
  
  
  Trade-offs and limitations
&lt;/h1&gt;

&lt;p&gt;This is a decision layer, not a free lunch, and some of the costs are structural.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Premium providers require their own API keys and consume real quota. The fallback billing alert exists precisely because a loop over providers can spend money.&lt;/li&gt;
&lt;li&gt;The premium pass is serial, and the fallback loop is serial too. Concurrency exists only between SearXNG and the whole premium pass, so enabling more providers in the hot path adds latency rather than hiding it.&lt;/li&gt;
&lt;li&gt;The circuit-breaker thresholds are heuristics. Tripping on the first 4xx is aggressive and will occasionally sideline an engine that returned a transient error, and the five-minute cooldown is a fixed guess. These values are tunable, not auto-tuned.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;timeout&lt;/code&gt; outcome is not mutually exclusive with the terminal outcome labels, which can be surprising when building dashboards.&lt;/li&gt;
&lt;li&gt;Attribution between SearXNG and premium contributors is not order-free: when both return an identical URL set, the label reflects whichever was merged first.&lt;/li&gt;
&lt;li&gt;The cache is in-memory and per-process. It does not persist and does not coordinate across instances.&lt;/li&gt;
&lt;li&gt;Keyless mode is genuinely usable but drops Google and Brave, which changes result quality noticeably for general web queries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nothing here solves engine-level blocking at the source. If every free engine is blocked from your IP, the gateway can only fail over to a provider you hold a key for, or fail cleanly.&lt;/p&gt;

&lt;h1&gt;
  
  
  AI-assistance disclosure
&lt;/h1&gt;

&lt;p&gt;Most of this project's implementation was written by AI, working from the author's specifications. The author reviewed the code, ran the tests, and deployed and operated the gateway; the design decisions, the trade-offs above, and the choice of what to ship are the author's. The project has been AI-heavy in its implementation history, and this article is part of that same practice. No benchmarks or performance numbers are claimed beyond what the repository's code and documentation state.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>go</category>
      <category>selfhosted</category>
      <category>ai</category>
    </item>
    <item>
      <title>Giving an OpenCode Coding Agent Persistent, Editable Memory</title>
      <dc:creator>Angelo Pantano</dc:creator>
      <pubDate>Thu, 17 Sep 2026 22:24:01 +0000</pubDate>
      <link>https://dev.to/ghilteras/giving-an-opencode-coding-agent-persistent-editable-memory-458n</link>
      <guid>https://dev.to/ghilteras/giving-an-opencode-coding-agent-persistent-editable-memory-458n</guid>
      <description>&lt;h1&gt;
  
  
  Giving an OpenCode Coding Agent Persistent, Editable Memory
&lt;/h1&gt;

&lt;p&gt;AI coding agents are good at using the context currently in front of them. They are much less reliable at carrying useful context from one session to the next.&lt;/p&gt;

&lt;p&gt;Project conventions, decisions made during a debugging session, and preferences that were obvious yesterday often need to be re-explained after a restart or after context compaction.&lt;/p&gt;

&lt;p&gt;I maintain &lt;code&gt;@ghilteras/opencode-agent-memory&lt;/code&gt;, an experimental plugin for OpenCode that explores an alternative: treating agent memory as editable, scoped Markdown state on disk rather than as an opaque remote service.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with always-in-context instructions
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;AGENTS.md&lt;/code&gt; and custom instruction files are a good start. I rely on them heavily. But they are a single flat document with no notion of scope, no size enforcement, and no dedicated operations for the agent to maintain it.&lt;/p&gt;

&lt;p&gt;What I actually wanted was structured state:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;some facts belong to me globally, across every project&lt;/li&gt;
&lt;li&gt;some facts belong only to one codebase&lt;/li&gt;
&lt;li&gt;each piece should have a description telling the agent how to use it&lt;/li&gt;
&lt;li&gt;each piece should have a size limit so it cannot silently grow without bound&lt;/li&gt;
&lt;li&gt;the agent should be able to read and rewrite that state with explicit tools&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Memory blocks
&lt;/h2&gt;

&lt;p&gt;The plugin gives the agent three tools: &lt;code&gt;memory_list&lt;/code&gt;, &lt;code&gt;memory_set&lt;/code&gt;, and &lt;code&gt;memory_replace&lt;/code&gt;. Blocks are plain Markdown files with YAML frontmatter.&lt;/p&gt;

&lt;p&gt;Global blocks live in &lt;code&gt;~/.config/opencode/memory/*.md&lt;/code&gt; and are shared across projects. Project blocks live in &lt;code&gt;.opencode/memory/*.md&lt;/code&gt; and are shared across sessions in that codebase, and are gitignored automatically.&lt;/p&gt;

&lt;p&gt;Each block has:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;label&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;filename&lt;/td&gt;
&lt;td&gt;unique identifier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;description&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;generic fallback&lt;/td&gt;
&lt;td&gt;tells the agent how to use the block&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;limit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;integer&lt;/td&gt;
&lt;td&gt;5000&lt;/td&gt;
&lt;td&gt;maximum characters&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;read_only&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;prevents agent edits&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code&gt;description&lt;/code&gt; field carries more weight than it looks like it should. Without it, the agent gets a generic fallback and does not know when the block is relevant. This mirrors the emphasis Letta puts on describing memory blocks well.&lt;/p&gt;

&lt;p&gt;Three blocks are seeded on first run: &lt;code&gt;persona&lt;/code&gt; and &lt;code&gt;human&lt;/code&gt; globally, &lt;code&gt;project&lt;/code&gt; for the current codebase.&lt;/p&gt;

&lt;h2&gt;
  
  
  The journal
&lt;/h2&gt;

&lt;p&gt;Memory blocks are curated state. Some things are not: observations, dead ends, discoveries, decisions, and the reasoning behind them.&lt;/p&gt;

&lt;p&gt;For that the plugin adds an optional append-only journal with &lt;code&gt;journal_write&lt;/code&gt;, &lt;code&gt;journal_search&lt;/code&gt;, and &lt;code&gt;journal_read&lt;/code&gt;. Entries are Markdown files with YAML frontmatter stored under &lt;code&gt;~/.config/opencode/journal/&lt;/code&gt;, and each entry records which project, model, provider, agent, and session it came from.&lt;/p&gt;

&lt;p&gt;The journal is deliberately opt-in. Enabling it is a line in &lt;code&gt;~/.config/opencode/agent-memory.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"journal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Append-only plus retrieval is what makes it useful. I do not want the agent rewriting history; I want it to be able to find what happened last month.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local semantic search
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;journal_search&lt;/code&gt; uses local embeddings rather than a hosted API. Entries are embedded with &lt;code&gt;paraphrase-multilingual-MiniLM-L12-v2&lt;/code&gt; (384 dimensions, multilingual) through Transformers.js, and the model is cached locally. Journal content does not leave the machine for search.&lt;/p&gt;

&lt;p&gt;Two implementation details were worth the effort:&lt;/p&gt;

&lt;p&gt;Embedding files are versioned. Each &lt;code&gt;.embedding&lt;/code&gt; sidecar stores &lt;code&gt;{ v, model, dimension, vector }&lt;/code&gt;. If the stored dimension does not match the current model, the entry degrades to text matching instead of failing the search. Legacy bare-array embeddings from earlier versions remain readable.&lt;/p&gt;

&lt;p&gt;The in-memory index is fingerprinted. &lt;code&gt;journal_search&lt;/code&gt; keeps an index per store instance and re-reads an entry only when the entry &lt;code&gt;.md&lt;/code&gt; or its &lt;code&gt;.embedding&lt;/code&gt; sidecar changed (mtime plus size). Regenerating or deleting a sidecar is picked up on the next search without a restart, and the embedding model is warmed up in the background at plugin init so the first search after a restart does not pay the cold model-load cost.&lt;/p&gt;

&lt;p&gt;There is also a deliberate retrieval floor: since v0.4.2, a query whose text matches an entry title (in either direction, case-insensitive) is guaranteed a high score, so title-based pointers stay retrievable even for entries with long bodies.&lt;/p&gt;

&lt;h2&gt;
  
  
  Operational trade-offs
&lt;/h2&gt;

&lt;p&gt;I want to be clear about the failure modes, because memory systems fail quietly.&lt;/p&gt;

&lt;p&gt;Stale memory is worse than no memory. A block that says something true six weeks ago will be confidently applied today. The &lt;code&gt;description&lt;/code&gt; field and size limits help the agent judge relevance, but they do not make the content true. Having the state as editable Markdown on disk means you can inspect and fix it directly; that is a feature, not a workaround.&lt;/p&gt;

&lt;p&gt;Automatic persistence is not the same as truth. The journal records what the agent observed, including wrong conclusions. Treat it as a log, not a knowledge base.&lt;/p&gt;

&lt;p&gt;The journal is opt-in because silent collection is worse than explicit collection.&lt;/p&gt;

&lt;h2&gt;
  
  
  Installation
&lt;/h2&gt;

&lt;p&gt;Requires OpenCode v1.0.115 or later.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"plugin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"@ghilteras/opencode-agent-memory@0.4.3"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OpenCode fetches unpinned plugins from npm on each startup; pinned versions are cached and need a manual bump. Restart OpenCode after changing plugin configuration; editing the config file alone is not enough to load a new plugin version.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;cacheDir&lt;/code&gt; key at the top level of &lt;code&gt;agent-memory.json&lt;/code&gt; relocates the Transformers.js model cache if you prefer not to use the default Hugging Face cache location.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would like feedback on
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What belongs in global memory versus project memory in your workflow?&lt;/li&gt;
&lt;li&gt;How should stale memories be surfaced rather than silently applied?&lt;/li&gt;
&lt;li&gt;How useful is semantic retrieval over the journal compared with plain text search, in practice?&lt;/li&gt;
&lt;li&gt;What should survive context compaction, and what should be allowed to decay?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Status
&lt;/h2&gt;

&lt;p&gt;This is experimental software, MIT licensed. It is maintained at &lt;a href="https://github.com/Ghilteras/opencode-agent-memory" rel="noopener noreferrer"&gt;https://github.com/Ghilteras/opencode-agent-memory&lt;/a&gt; and published as &lt;code&gt;@ghilteras/opencode-agent-memory&lt;/code&gt;. It is not built by or affiliated with the OpenCode team.&lt;/p&gt;

&lt;p&gt;If you try it, issues and concrete workflow reports are welcome.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>agentskills</category>
      <category>opencode</category>
    </item>
  </channel>
</rss>
