<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Royal Simpson Pinto</title>
    <description>The latest articles on DEV Community by Royal Simpson Pinto (@royalpinto007).</description>
    <link>https://dev.to/royalpinto007</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F947695%2F8007e48f-21ef-4490-8290-df21931e329a.jpg</url>
      <title>DEV Community: Royal Simpson Pinto</title>
      <link>https://dev.to/royalpinto007</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/royalpinto007"/>
    <language>en</language>
    <item>
      <title>Signing and Verifying Data with Ed25519 in Python</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:30:29 +0000</pubDate>
      <link>https://dev.to/royalpinto007/signing-and-verifying-data-with-ed25519-in-python-3e0b</link>
      <guid>https://dev.to/royalpinto007/signing-and-verifying-data-with-ed25519-in-python-3e0b</guid>
      <description>&lt;p&gt;I got into Ed25519 while building answerproof, a small tool that signs the receipts my RAG pipeline emits so a reader can later prove the answer came from a specific set of retrieved chunks and was not tampered with afterward. The moment I needed "prove this bytes blob is exactly what I produced," a signature scheme was the right tool, and Ed25519 turned out to be the friendliest one to reach for. In this post I want to teach the technique itself, using those RAG receipts as a running example.&lt;/p&gt;

&lt;p&gt;Ed25519 is an elliptic-curve signature scheme. You hold a private (signing) key, you hand out a public (verifying) key, and anyone with the public key can check that a piece of data was signed by whoever holds the private key. It is fast, the keys and signatures are small (32-byte keys, 64-byte signatures), and the API is hard to misuse. That last point matters more than it sounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generating a key pair
&lt;/h2&gt;

&lt;p&gt;I use the &lt;code&gt;cryptography&lt;/code&gt; library because it ships good Ed25519 primitives. Install it first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;cryptography
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Generating a key pair is one call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;cryptography.hazmat.primitives.asymmetric.ed25519&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;Ed25519PrivateKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Ed25519PublicKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;private_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Ed25519PrivateKey&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;public_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;private_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;public_key&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The private key is your secret. The public key is what you publish so others can verify your signatures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Signing data (detached signatures)
&lt;/h2&gt;

&lt;p&gt;Ed25519 signatures are detached by design: signing produces a 64-byte signature that lives separately from the message. You keep the message as it is and attach the signature alongside it. This fits my receipts perfectly, because the receipt JSON stays human-readable and the signature rides in a separate field.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="n"&gt;receipt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What is the refund window?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk_ids&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doc12#p3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doc12#p4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer_hash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sha256:9f2a...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Serialize deterministically so the exact same bytes are what you sign and verify.
&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;receipt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sort_keys&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;separators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;,&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;signature&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;private_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sign&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# 64 raw bytes
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The single most important habit here: sign bytes, not objects. Whatever you sign, you must be able to reproduce those exact bytes at verification time. That is why I serialize with &lt;code&gt;sort_keys=True&lt;/code&gt; and fixed separators. If the byte layout changes, verification fails, and it should.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verifying a signature
&lt;/h2&gt;

&lt;p&gt;Verification uses the public key. The &lt;code&gt;cryptography&lt;/code&gt; library does not return &lt;code&gt;True&lt;/code&gt; or &lt;code&gt;False&lt;/code&gt;. Instead it raises &lt;code&gt;InvalidSignature&lt;/code&gt; on failure and returns &lt;code&gt;None&lt;/code&gt; on success. Treat the exception as the failure path.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;cryptography.exceptions&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;InvalidSignature&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;public_key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Ed25519PublicKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;public_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;InvalidSignature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;public_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;  &lt;span class="c1"&gt;# True
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Change a single byte of &lt;code&gt;message&lt;/code&gt; or &lt;code&gt;signature&lt;/code&gt; and it flips to &lt;code&gt;False&lt;/code&gt;. That is the whole guarantee: the signature is valid only for the exact bytes it was made over, made by the exact key it claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sharing keys: base64url encoding
&lt;/h2&gt;

&lt;p&gt;Raw keys are bytes, and bytes do not travel well through JSON, HTTP headers, or config files. I encode them as base64url, which is URL and filename safe and avoids the &lt;code&gt;+&lt;/code&gt; and &lt;code&gt;/&lt;/code&gt; characters that plain base64 uses.&lt;/p&gt;

&lt;p&gt;To export the public key so a verifier can import it later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;cryptography.hazmat.primitives&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;serialization&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;b64url_encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlsafe_b64encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;rstrip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ascii&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;b64url_decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;padding&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlsafe_b64decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;padding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Public key to a shareable string
&lt;/span&gt;&lt;span class="n"&gt;public_raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;public_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;public_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;serialization&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Encoding&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Raw&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nb"&gt;format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;serialization&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PublicFormat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Raw&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;public_b64&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;b64url_encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;public_raw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I strip the &lt;code&gt;=&lt;/code&gt; padding on the way out and re-add it on the way in, which keeps the string clean in URLs. Reconstructing the public key on the verifier side:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;public_key_again&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Ed25519PublicKey&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_public_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;b64url_decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;public_b64&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The signature travels the same way. In my receipts it lands as a field:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;receipt_signed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;receipt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sig&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;b64url_encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pubkey&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;public_b64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A verifier reads the receipt, re-serializes the signed fields into the same canonical bytes, decodes &lt;code&gt;sig&lt;/code&gt; and &lt;code&gt;pubkey&lt;/code&gt;, and calls &lt;code&gt;verify&lt;/code&gt;. If it passes, the receipt is authentic and unmodified.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pynacl alternative
&lt;/h2&gt;

&lt;p&gt;If you prefer PyNaCl, the same flow is a few lines. It returns a combined &lt;code&gt;SignedMessage&lt;/code&gt;, but it also supports detached signatures:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;nacl.signing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SigningKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;VerifyKey&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;nacl.exceptions&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BadSignatureError&lt;/span&gt;

&lt;span class="n"&gt;signing_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;SigningKey&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;verify_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;signing_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;verify_key&lt;/span&gt;

&lt;span class="n"&gt;signature&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;signing_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sign&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;signature&lt;/span&gt;  &lt;span class="c1"&gt;# detached 64 bytes
&lt;/span&gt;
&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;verify_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;ok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;BadSignatureError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;ok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both libraries implement the same standard, so a signature made by one verifies under the other as long as you feed identical bytes and use the raw key encoding.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest caveat: key management
&lt;/h2&gt;

&lt;p&gt;Signing is the easy part. The hard part, the part no library solves for you, is protecting the private key. A signature only means something if the private key stayed secret. If it leaks, anyone can forge receipts that verify perfectly, and you have no way to tell forgeries from the real thing after the fact.&lt;/p&gt;

&lt;p&gt;So do not hardcode signing keys in source, and do not commit them. Load them from a secret manager or an environment variable at runtime, keep them off disk where you can, and give yourself a rotation story: publish a key identifier alongside each signature so you can retire a compromised key and move to a new one without invalidating your whole history. In answerproof I treat the private key as the single most sensitive thing in the system, because it is. The cryptography is only as trustworthy as the secret behind it.&lt;/p&gt;

&lt;p&gt;That is the full loop: generate, sign detached, verify, and exchange keys as base64url strings. Once you have it wired up, signing becomes a boring, reliable primitive you can drop into receipts, audit logs, webhook payloads, or anything else where "this is exactly what I produced" needs to be provable. If you want to see it grounded in a real RAG receipt pipeline, the code lives at github.com/AgentPostmortem/answerproof.&lt;/p&gt;

</description>
      <category>python</category>
      <category>cryptography</category>
      <category>security</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Merkle Trees and Inclusion Proofs in Python From Scratch</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Sun, 27 Sep 2026 09:30:31 +0000</pubDate>
      <link>https://dev.to/royalpinto007/merkle-trees-and-inclusion-proofs-in-python-from-scratch-25lo</link>
      <guid>https://dev.to/royalpinto007/merkle-trees-and-inclusion-proofs-in-python-from-scratch-25lo</guid>
      <description>&lt;p&gt;I kept running into the same question while building AI retrieval systems: how do you prove, after the fact, that a particular source document was actually part of the set the model saw? You want a compact receipt that anyone can check without re-running your whole pipeline. Merkle trees are the classic answer, and once I built one by hand I realized the idea is much simpler than the cryptography reputation suggests.&lt;/p&gt;

&lt;p&gt;In this tutorial I will build a Merkle tree from scratch in Python, generate an inclusion proof for a single leaf, and verify that proof standalone. I will use retrieval provenance as the running example, but the technique applies to any situation where you want to commit to a set of items and later prove membership: transaction logs, file backups, certificate transparency, and so on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core idea
&lt;/h2&gt;

&lt;p&gt;A Merkle tree hashes your data in a binary tree. The leaves are hashes of your items. Each internal node is the hash of its two children concatenated. The single hash at the top, the Merkle root, is a fingerprint of the entire set. Change any item, and the root changes.&lt;/p&gt;

&lt;p&gt;The magic is the inclusion proof. To prove a leaf is in the tree, you do not need the whole tree. You only need the sibling hashes along the path from that leaf up to the root. That is O(log n) hashes. A verifier who trusts the root can recompute their way up and confirm they arrive at the same root.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two details that matter
&lt;/h2&gt;

&lt;p&gt;Before any code, two things separate a toy Merkle tree from a correct one.&lt;/p&gt;

&lt;p&gt;First, domain separation. If you hash leaves and internal nodes the same way, an attacker can present an internal node as if it were a leaf. This is the classic second-preimage weakness. The fix is to prefix a different byte before hashing leaves versus internal nodes. I use &lt;code&gt;0x00&lt;/code&gt; for leaves and &lt;code&gt;0x01&lt;/code&gt; for internal nodes.&lt;/p&gt;

&lt;p&gt;Second, odd-node handling. When a level has an odd number of nodes, one node has no sibling to pair with. Different systems handle this differently. I promote the lone node unchanged to the next level. This is simple and deterministic, though I will flag a caveat about it later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building the tree
&lt;/h2&gt;

&lt;p&gt;Let me start with the two hashing primitives.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;

&lt;span class="n"&gt;LEAF_PREFIX&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\x00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;NODE_PREFIX&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\x01&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;hash_leaf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LEAF_PREFIX&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;hash_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;left&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;right&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;NODE_PREFIX&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;left&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;right&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the tree. I build it level by level, keeping every level so I can generate proofs later.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;MerkleTree&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cannot build a Merkle tree over zero items&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;leaves&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;hash_leaf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;items&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;levels&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_build&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;leaves&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_build&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;leaves&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;]]:&lt;/span&gt;
        &lt;span class="n"&gt;levels&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;leaves&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;leaves&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;nxt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                    &lt;span class="n"&gt;nxt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;hash_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
                &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="c1"&gt;# odd node: promote it unchanged
&lt;/span&gt;                    &lt;span class="n"&gt;nxt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
            &lt;span class="n"&gt;levels&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nxt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nxt&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;levels&lt;/span&gt;

    &lt;span class="nd"&gt;@property&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;root&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;levels&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;levels&lt;/code&gt; list holds the leaf hashes at index 0, then each successive parent level, ending with a single-element level that holds the root.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generating an inclusion proof
&lt;/h2&gt;

&lt;p&gt;A proof is the list of sibling hashes on the way up, each tagged with whether it sits on the left or right. The verifier needs the side so it concatenates in the correct order.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;proof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;]]:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;leaves&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;IndexError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;leaf index out of range&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;level&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;levels&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="n"&gt;is_right_node&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;is_right_node&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;sibling_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
                &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;left&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;sibling_index&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;sibling_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;sibling_index&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                    &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;right&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;sibling_index&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
                &lt;span class="c1"&gt;# else: promoted odd node, no sibling at this level
&lt;/span&gt;            &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;//=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When our node is on the right, its sibling is to the left, and vice versa. If our node is a left node at the end of an odd level, it has no sibling and was promoted, so we add nothing for that level. We then move up by halving the index.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verifying a proof standalone
&lt;/h2&gt;

&lt;p&gt;This is the part that makes Merkle trees useful. Verification needs only three things: the original item, the proof, and the trusted root. No tree, no other items.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verify_proof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;proof&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;hash_leaf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;side&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sibling&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;proof&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;side&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;left&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;hash_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sibling&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;computed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;hash_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;computed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;sibling&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let me clean that verify up, since the ternary is noise:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verify_proof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;proof&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;hash_leaf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;side&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sibling&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;proof&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;side&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;left&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;hash_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sibling&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;computed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;hash_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;computed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sibling&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;computed&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The verifier hashes the item as a leaf, then walks the proof, folding in each sibling on the correct side, and checks whether it lands on the root.&lt;/p&gt;

&lt;h2&gt;
  
  
  The retrieval provenance example
&lt;/h2&gt;

&lt;p&gt;Now the concrete use case. Say a RAG system retrieved five source chunks to answer a question. We want a receipt proving that one specific chunk, say the one the answer cited, was in that retrieval set.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sources&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doc:handbook#p12 :: refunds within 30 days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doc:handbook#p13 :: return shipping is prepaid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doc:policy#p4    :: no refunds on final-sale items&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doc:faq#q9       :: exchanges allowed within 60 days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doc:handbook#p14 :: store credit never expires&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;tree&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MerkleTree&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;root&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tree&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;root&lt;/span&gt;  &lt;span class="c1"&gt;# publish or log this as the retrieval commitment
&lt;/span&gt;
&lt;span class="n"&gt;cited_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
&lt;span class="n"&gt;cited_source&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;cited_index&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;receipt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tree&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;proof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cited_index&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;root:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;hex&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verified:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;verify_proof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cited_source&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;receipt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="c1"&gt;# Tamper check: a source that was never retrieved must fail
&lt;/span&gt;&lt;span class="n"&gt;fake&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doc:policy#p4    :: refunds on final-sale items are fine&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tampered:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;verify_proof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fake&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;receipt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Running this prints &lt;code&gt;verified: True&lt;/code&gt; and &lt;code&gt;tampered: False&lt;/code&gt;. The tampered line proves the point: flip a single word in the source and the proof no longer reconstructs the root. If you log the root at retrieval time, anyone holding a source and its receipt can later confirm it was genuinely part of that set, without access to your database or the other four chunks. That is exactly the property I wanted for retrieval provenance in my project answerproof.&lt;/p&gt;

&lt;h2&gt;
  
  
  One honest caveat
&lt;/h2&gt;

&lt;p&gt;The odd-node promotion I used is simple, but it is worth knowing its weakness. Because a promoted node is carried up unchanged, a tree over an odd count can, in adversarial constructions, share structure with a differently shaped tree, which is the root of the well known CVE-2012-2459 style ambiguity in some Merkle implementations. For a trusted producer logging its own retrieval sets this is fine. If you need to defend against a malicious tree builder, use a scheme that removes the ambiguity: duplicate the lone node instead of promoting it, or better, prepend the leaf count to what you hash and reject proofs whose length does not match the committed size. Domain separation, which I did include, already blocks the leaf-versus-node confusion, so we are only talking about the shape ambiguity here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;That is a complete, correct Merkle tree in well under a hundred lines: domain-separated leaf and node hashing, level-by-level construction with odd-node promotion, logarithmic inclusion proofs, and standalone verification that needs nothing but the item, the receipt, and the root. The technique is general, but retrieval provenance made it click for me: a tiny receipt that survives long after the pipeline that produced it.&lt;/p&gt;

&lt;p&gt;If you want to see this idea wired into a real retrieval-answer flow, the provenance work lives in github.com/AgentPostmortem/answerproof. Clone the snippets above, break something on purpose, and watch the root refuse to lie.&lt;/p&gt;

</description>
      <category>python</category>
      <category>cryptography</category>
      <category>ai</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Combining Vector and Full-Text Search with Reciprocal Rank Fusion</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Fri, 25 Sep 2026 09:30:29 +0000</pubDate>
      <link>https://dev.to/royalpinto007/combining-vector-and-full-text-search-with-reciprocal-rank-fusion-16mj</link>
      <guid>https://dev.to/royalpinto007/combining-vector-and-full-text-search-with-reciprocal-rank-fusion-16mj</guid>
      <description>&lt;p&gt;Every time I build retrieval for a RAG system, I run into the same wall. Vector search is wonderful at understanding meaning. Ask it "what is our policy on working from home" and it happily finds the paragraph titled "Remote Work Guidelines" even though the words do not match. But ask it for "ERR_4021" or "Policy 7.3" and it flounders, because an error code has no meaningful embedding neighborhood. Keyword search is the mirror image: it nails the exact token and misses the paraphrase.&lt;/p&gt;

&lt;p&gt;Real questions contain both kinds of signal. So the honest answer is to run both searches and combine them. The problem is that the two searches produce scores on completely different scales. A cosine distance of 0.18 and a BM25 score of 7.4 are not comparable numbers. You cannot just add them.&lt;/p&gt;

&lt;p&gt;This is exactly the problem Reciprocal Rank Fusion solves, and I want to teach you the technique here. I will use my own project, vaultrag, as the running example, but the method transfers to any two rankers you have.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core idea: throw away the scores, keep the ranks
&lt;/h2&gt;

&lt;p&gt;The insight behind RRF is almost rude in its simplicity. Ignore the raw scores entirely. They are not comparable, so stop trying to compare them. Instead, look only at the position a document holds in each list. Rank 1 is rank 1 whether it came from a vector index or a keyword index, and those you can combine.&lt;/p&gt;

&lt;p&gt;Here is the formula. For a document &lt;code&gt;d&lt;/code&gt;, its fused score is the sum over every ranked list of one divided by a constant &lt;code&gt;k&lt;/code&gt; plus the document's rank in that list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RRF(d) = sum over lists L of  1 / (k + rank_L(d))
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rank is 1-based (best result is rank 1). If a document does not appear in a given list at all, it simply contributes nothing from that list. The constant &lt;code&gt;k&lt;/code&gt; is a damping term. A larger &lt;code&gt;k&lt;/code&gt; flattens the curve so the top result of any single list does not dominate; a smaller &lt;code&gt;k&lt;/code&gt; lets the top hits win harder. The value from the original 2009 paper by Cormack, Clarke, and Buettcher is &lt;code&gt;k = 60&lt;/code&gt;, and it is a perfectly reasonable default that I have never had a strong reason to change.&lt;/p&gt;

&lt;p&gt;Why &lt;code&gt;1 / (k + rank)&lt;/code&gt;? Because it is steeply decreasing but never zero. Moving from rank 1 to rank 2 costs a lot; moving from rank 40 to rank 41 costs almost nothing. That matches intuition: the difference between the best and second-best result matters far more than the difference between the fortieth and forty-first.&lt;/p&gt;

&lt;h2&gt;
  
  
  A pure-Python implementation
&lt;/h2&gt;

&lt;p&gt;Before wiring it into a database, here is the whole technique in a few lines you can drop into a notebook and reason about:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;reciprocal_rank_fusion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ranked_lists&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Fuse several ranked lists of ids into one.

    ranked_lists: list of lists, each already ordered best-first.
    Returns: list of (id, score) sorted best-first.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ranked&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ranked_lists&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ranked&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;kv&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;kv&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;reverse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice what this does and does not need. It never sees a cosine distance or a BM25 score. It only needs each list in the correct order. That is the entire contract, which is why RRF works with any rankers you can name, including a lexical index, a dense retriever, and a reranker all at once.&lt;/p&gt;

&lt;p&gt;A quick sanity check:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;vec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;c&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;     &lt;span class="c1"&gt;# vector search order
&lt;/span&gt;&lt;span class="n"&gt;kw&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;c&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;d&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;     &lt;span class="c1"&gt;# keyword search order
&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;reciprocal_rank_fusion&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;vec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kw&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Document &lt;code&gt;a&lt;/code&gt; appears at rank 1 in vec and rank 2 in kw, so it scores &lt;code&gt;1/61 + 1/62&lt;/code&gt;. Document &lt;code&gt;c&lt;/code&gt; is rank 3 in vec and rank 1 in kw. Both landing near the top of some list beats &lt;code&gt;b&lt;/code&gt; and &lt;code&gt;d&lt;/code&gt;, which each show up in only one list. Agreement between the two searches is rewarded, which is precisely the behavior we want.&lt;/p&gt;

&lt;h2&gt;
  
  
  Doing it inside the database
&lt;/h2&gt;

&lt;p&gt;In vaultrag I do not fuse in Python. I fuse in the same SQL query that runs both searches, so Postgres hands me one already-fused list. The two arms are a vector search using pgvector's &lt;code&gt;&amp;lt;=&amp;gt;&lt;/code&gt; distance operator and a full-text search using &lt;code&gt;ts_rank_cd&lt;/code&gt;. &lt;code&gt;ROW_NUMBER()&lt;/code&gt; turns each arm's ordering into an explicit rank, and then a &lt;code&gt;FULL OUTER JOIN&lt;/code&gt; lets me add the two reciprocal terms:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;vec&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ROW_NUMBER&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="n"&gt;OVER&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;visible&lt;/span&gt;
    &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
    &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt;
    &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="n"&gt;kw&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ROW_NUMBER&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="n"&gt;OVER&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
               &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;ts_rank_cd&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tsv&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;websearch_to_tsquery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'english'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
           &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;visible&lt;/span&gt;
    &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;tsv&lt;/span&gt; &lt;span class="o"&gt;@@&lt;/span&gt; &lt;span class="n"&gt;websearch_to_tsquery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'english'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="n"&gt;fused&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
           &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;vec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
         &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;kw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;vec&lt;/span&gt;
    &lt;span class="k"&gt;FULL&lt;/span&gt; &lt;span class="k"&gt;OUTER&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;kw&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;kw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;vec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details are load-bearing. The &lt;code&gt;FULL OUTER JOIN&lt;/code&gt; is what allows a document to appear in one arm but not the other, and &lt;code&gt;COALESCE(..., 0)&lt;/code&gt; is the "contributes nothing when absent" rule from the formula made literal. A document found only by keyword search has a NULL vector rank, so its vector term collapses to zero, and only its keyword term counts. That is RRF working exactly as designed.&lt;/p&gt;

&lt;p&gt;I also limit each arm to a candidate pool (I use 50) before fusing, then return the top few. Fusing the whole corpus would be pointless work; anything ranked fiftieth in both arms is not going to win.&lt;/p&gt;

&lt;h2&gt;
  
  
  One honest caveat
&lt;/h2&gt;

&lt;p&gt;RRF is deliberately blind to how confident each search was. A document that is a near-perfect keyword match at rank 1 and a document that is a lukewarm match at rank 1 contribute the identical &lt;code&gt;1/(k+1)&lt;/code&gt;, because the rank is the same and the score was discarded. Usually that robustness is a feature, since it stops one loud arm from steamrolling the other. But when one of your rankers is genuinely much more trustworthy than the other for a given query, RRF cannot express that. If you need to weight the arms or preserve calibrated confidence, you will have to reach for weighted fusion or a learned reranker instead. RRF is the strong, simple baseline, not the ceiling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;Reciprocal Rank Fusion earns its keep because it demands so little: no shared score scale, no training, no tuning beyond one constant that has a sensible default. Give it two lists in the right order and it gives you back one better list. For hybrid search that is very often all you need.&lt;/p&gt;

&lt;p&gt;If you want to see the full ACL-scoped hybrid query this snippet came from, including how both search arms start from the same authorized candidate set, the code is at &lt;a href="https://github.com/AgentPostmortem/vaultrag" rel="noopener noreferrer"&gt;github.com/AgentPostmortem/vaultrag&lt;/a&gt;. Clone it, read &lt;code&gt;app/retrieval.py&lt;/code&gt;, and try changing &lt;code&gt;k&lt;/code&gt; to see the ranking shift for yourself.&lt;/p&gt;

</description>
      <category>rag</category>
      <category>search</category>
      <category>python</category>
      <category>ai</category>
    </item>
    <item>
      <title>Enforce Permissions Inside the pgvector Query, Not After It</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Wed, 23 Sep 2026 09:30:32 +0000</pubDate>
      <link>https://dev.to/royalpinto007/enforce-permissions-inside-the-pgvector-query-not-after-it-439</link>
      <guid>https://dev.to/royalpinto007/enforce-permissions-inside-the-pgvector-query-not-after-it-439</guid>
      <description>&lt;p&gt;Most permission bugs in RAG systems look harmless in review. You retrieve the top-k nearest chunks, then you drop the ones the user is not allowed to see. It reads as correct. It is not, and the failure is quiet.&lt;/p&gt;

&lt;p&gt;I want to show you why the post-filter is the wrong place, and how to push the access-control decision inside the retrieval query itself so that an unauthorized chunk is never selected, never ranked, never counted. I will ground this in a small permission-aware RAG project of mine called vaultrag, but the technique is portable to any Postgres plus pgvector stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why post-filtering leaks and lies
&lt;/h2&gt;

&lt;p&gt;Here is the pattern I want you to stop writing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# fetch top 10 by vector distance
&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;fetch_nearest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query_embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# then remove what the user cannot see
&lt;/span&gt;&lt;span class="n"&gt;visible&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;user_can_see&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things are wrong here.&lt;/p&gt;

&lt;p&gt;The first is a correctness bug. Your &lt;code&gt;LIMIT 10&lt;/code&gt; runs against every chunk in the table. If eight of the ten nearest chunks belong to documents this user cannot read, you filter them out and hand back two results. The user experiences this as a broken search, not as a security boundary. Relevant material they are allowed to see sat at rank 11 and never made it into the candidate set.&lt;/p&gt;

&lt;p&gt;The second is worse, and it is the reason to care. Every code path that touches the raw result set is now a place a leak can happen. Logging the pre-filter rows, a metrics counter, a debug endpoint, a caching layer that memoizes by query text: each one can see chunks the user cannot. The filter protects exactly one exit. Everything upstream of it is holding forbidden data.&lt;/p&gt;

&lt;p&gt;The fix is to make it structurally impossible to hold that data in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model the access control as data
&lt;/h2&gt;

&lt;p&gt;Start with an explicit access-control list per document. In vaultrag a document has an ACL made of principals, where a principal is either a user id or a group name.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;          &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;title&lt;/span&gt;       &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;deleted_at&lt;/span&gt;  &lt;span class="n"&gt;TIMESTAMPTZ&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;doc_acl&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;doc_id&lt;/span&gt;      &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;DELETE&lt;/span&gt; &lt;span class="k"&gt;CASCADE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;principal&lt;/span&gt;   &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;          &lt;span class="n"&gt;BIGSERIAL&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;doc_id&lt;/span&gt;      &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;DELETE&lt;/span&gt; &lt;span class="k"&gt;CASCADE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nb"&gt;text&lt;/span&gt;        &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;embedding&lt;/span&gt;   &lt;span class="n"&gt;VECTOR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1536&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The critical design choice is resolving the caller's principals on the server, never from the request. If a client can assert its own group membership, the ACL is decorative.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;resolve_principal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;row_factory&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;dict_row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT id, groups FROM users WHERE id = %s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,))&lt;/span&gt;
        &lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetchone&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="c1"&gt;# principals we match against: the user id plus every group they belong to
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;groups&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;())]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Put the boundary in a CTE that everything reads from
&lt;/h2&gt;

&lt;p&gt;The technique is a single common table expression, &lt;code&gt;visible&lt;/code&gt;, that defines the universe of chunks this caller may see. Every other part of the query reads from &lt;code&gt;visible&lt;/code&gt;, not from &lt;code&gt;chunks&lt;/code&gt;. There is no path that starts anywhere else.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="n"&gt;visible&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
    &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;
    &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deleted_at&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
      &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
          &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;doc_acl&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;
          &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;
            &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;principal&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;ANY&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;principals&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="n"&gt;ranked&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;visible&lt;/span&gt;
    &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
    &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt;
    &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;ranked&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read the order of operations, because it is the whole point. The &lt;code&gt;EXISTS&lt;/code&gt; against &lt;code&gt;doc_acl&lt;/code&gt; runs before the &lt;code&gt;ORDER BY ... &amp;lt;=&amp;gt; ...&lt;/code&gt;. The nearest-neighbor ranking and the &lt;code&gt;LIMIT&lt;/code&gt; operate on &lt;code&gt;visible&lt;/code&gt;, which already excludes forbidden chunks. Your top-k is now the top-k of what the user is allowed to see. The correctness bug from earlier is gone, and so is the leak surface, because the query planner never materializes a forbidden row into the candidate set.&lt;/p&gt;

&lt;p&gt;A detail worth stealing: use &lt;code&gt;EXISTS&lt;/code&gt; rather than a plain &lt;code&gt;JOIN doc_acl&lt;/code&gt;. A document with three matching ACL rows would otherwise multiply into three copies of each chunk. &lt;code&gt;EXISTS&lt;/code&gt; short-circuits on the first match, so each visible chunk appears exactly once.&lt;/p&gt;

&lt;p&gt;The same &lt;code&gt;visible&lt;/code&gt; CTE composes cleanly with hybrid search. In vaultrag both the vector arm and the full-text arm select from &lt;code&gt;visible&lt;/code&gt; before they are fused with reciprocal rank fusion, so neither arm can surface a chunk the other could not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="n"&gt;visible&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt; &lt;span class="p"&gt;),&lt;/span&gt;                 &lt;span class="c1"&gt;-- the boundary, defined once&lt;/span&gt;
&lt;span class="n"&gt;vec&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ROW_NUMBER&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="n"&gt;OVER&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;visible&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
    &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt; &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="n"&gt;kw&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ROW_NUMBER&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="n"&gt;OVER&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;ts_rank_cd&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tsv&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;websearch_to_tsquery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'english'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;visible&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;tsv&lt;/span&gt; &lt;span class="o"&gt;@@&lt;/span&gt; &lt;span class="n"&gt;websearch_to_tsquery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'english'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;-- fuse vec and kw, both already scoped to visible&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because both arms read from &lt;code&gt;visible&lt;/code&gt;, there is no code path in the function that can rank an unauthorized chunk. That is a property you can state about the query, not a behavior you hope the filter enforces.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest caveat
&lt;/h2&gt;

&lt;p&gt;Pushing the ACL into the query moves the boundary, it does not make the boundary free. The &lt;code&gt;EXISTS&lt;/code&gt; subquery runs per candidate document, and on a large table the planner's choices matter. You want an index that makes the ACL check cheap:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;doc_acl&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;principal&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is real tension here with the pgvector index. An HNSW or IVFFlat index gives you approximate nearest neighbors fast, but the approximation happens before your &lt;code&gt;WHERE&lt;/code&gt; filter is applied inside the scan. When a filter is very selective, meaning the user can see only a tiny fraction of documents, the vector index can return a page of candidates that are then almost entirely filtered out, and you get fewer results than your &lt;code&gt;LIMIT&lt;/code&gt; asked for. This is the well-known filtered-search problem. Mitigations exist: raise &lt;code&gt;hnsw.ef_search&lt;/code&gt;, use partitioning, or on recent pgvector use iterative index scans. Measure it on your data. Do not assume the CTE is free just because it is correct.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Correctness and security both improve when the permission check is a precondition of ranking rather than a cleanup step after it. Define the visible set once, in a CTE, and make every arm of your search read from it. The guarantee you get is structural: an unauthorized chunk is not filtered out late, it is never selected.&lt;/p&gt;

&lt;p&gt;If you want a full working reference, the retrieval query, the ACL schema, and the server-side principal resolution live in my project at &lt;a href="https://github.com/AgentPostmortem/vaultrag" rel="noopener noreferrer"&gt;github.com/AgentPostmortem/vaultrag&lt;/a&gt;. Borrow the CTE and adapt the ACL model to your own tenancy rules.&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>rag</category>
      <category>security</category>
      <category>ai</category>
    </item>
    <item>
      <title>The boring half of AI: verification is harder than generation</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Mon, 21 Sep 2026 09:30:29 +0000</pubDate>
      <link>https://dev.to/royalpinto007/the-boring-half-of-ai-verification-is-harder-than-generation-2cn9</link>
      <guid>https://dev.to/royalpinto007/the-boring-half-of-ai-verification-is-harder-than-generation-2cn9</guid>
      <description>&lt;p&gt;I have spent a long stretch building small tools around AI systems, and people keep asking me what ties them together. The honest answer is a single word that gets overused: trust. But I do not mean trust as a feeling. I do not mean a model that sounds confident, or a demo that goes well on stage. I mean something I can define operationally, test, and point at in code. This essay is my attempt to say what "trustworthy AI infrastructure" actually means to me, using the things I have already built as evidence.&lt;/p&gt;

&lt;p&gt;Here is the short version. A system earns trust when you enforce constraints at the boundary, verify what happened after the fact, keep a human gate in front of anything irreversible, and make regressions fail loudly instead of silently. None of these make a model correct. They make a model checkable. That distinction is the whole point, and I will come back to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enforce at the boundary
&lt;/h2&gt;

&lt;p&gt;The first principle is that you do not ask the model to behave. You constrain what it can reach. If a component can touch data it should never see, then a good prompt is the only thing standing between you and a leak, and prompts are not a security control.&lt;/p&gt;

&lt;p&gt;This is why I built vaultrag, a permission-aware retrieval layer. The idea is simple and, to me, non-negotiable: retrieval should respect the same permissions the rest of your system respects. A user's query should only ever be able to pull context that user is allowed to read. The enforcement lives at the retrieval boundary, not in a hopeful instruction telling the model to be careful. The same instinct runs through Bridgekit, a scoped MCP server. An agent connected through it gets a deliberately narrow surface, not the whole machine. Scope is the feature.&lt;/p&gt;

&lt;p&gt;Boundaries are also something you have to inspect, because they drift. mcp-audit is a scanner for MCP setups: it looks at what a server actually exposes rather than what the README claims. And injection-arena, a prompt-injection challenge game, is really a teaching tool for this same lesson. You play it and you feel, viscerally, how quickly a system that trusts its input gets walked straight past its own rules. Once you have lost that round a few times, "enforce at the boundary" stops being a slogan.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verify after the fact
&lt;/h2&gt;

&lt;p&gt;Boundaries stop the obvious harm. They do not tell you what the system actually did. For that you need a record, and the record has to be trustworthy on its own terms.&lt;/p&gt;

&lt;p&gt;answerproof is my answer to this: signed RAG receipts. When the system produces an answer, it also produces a verifiable artifact of what sources went into that answer. You are not taking the pipeline's word for it later. You can check the receipt. The signing matters because an unsigned log is just another thing that can be edited to tell a comfortable story.&lt;/p&gt;

&lt;p&gt;agentrace comes at verification from the behavior side, attaching trust flags to what an agent does as it runs, so a review is not an exercise in re-reading raw transcripts and guessing. And ctxlens, a context profiler, answers a question that sounds boring but is central: what was actually in the context window? So much unexplained model behavior turns out to be explainable the moment you can see the real assembled context rather than the tidy version you imagined you sent. Verification, across all three, means the same thing: reconstruct what happened from evidence, not from trust in the narrator.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep a human gate for irreversible actions
&lt;/h2&gt;

&lt;p&gt;Some actions can be undone. Some cannot. Sending money, deleting records, shipping a message to a customer, calling an external side effect that the world then reacts to. My rule is that anything in the second category gets a human gate, on purpose, by default.&lt;/p&gt;

&lt;p&gt;This is the whole reason the human-in-the-loop tools exist: Greenlite, Webhands, and relayg. They are built around the assumption that an agent will propose and a person will approve before the irreversible thing happens. I know the current fashion is full autonomy, and I understand the appeal. But a system that can take an unrecoverable action without a checkpoint is not more advanced, it is just less careful. The gate is not a lack of ambition. It is where I decided the risk was not worth the convenience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make regressions fail loudly
&lt;/h2&gt;

&lt;p&gt;The last principle is about time. A system that is trustworthy today can rot quietly, because prompts, models, and data all move underneath you. If a regression can slip in without anyone noticing, then all the earlier work has a short shelf life.&lt;/p&gt;

&lt;p&gt;evalgate is prompt regression CI: it treats prompt behavior like code, so a change that degrades quality fails the build instead of shipping. voiceeval does the equivalent for voice agents, where the failure modes are harder to eyeball and therefore easier to miss. The shared belief is that quality you do not continuously test is quality you are slowly losing. Loud failure is a feature. Silence is the bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not buy you
&lt;/h2&gt;

&lt;p&gt;I want to be honest about the ceiling here, because overclaiming would undercut the entire argument. None of these tools make a model correct. A permission-aware retriever will faithfully return the wrong-but-authorized document. A signed receipt will faithfully sign a bad answer. A regression test only catches the regressions you thought to write. What this infrastructure buys is not correctness. It is checkability: the ability to constrain, to inspect, to gate, and to notice. A checkable wrong answer is one you can catch. An unattributable, unconstrained, silent wrong answer is one that ships.&lt;/p&gt;

&lt;p&gt;I have made my peace with that ceiling. I would rather build systems that are honest about what they cannot guarantee than systems that feel trustworthy and are not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I am headed
&lt;/h2&gt;

&lt;p&gt;The direction I care about now is making these principles compose instead of standing alone. A boundary that emits a receipt. A receipt a human gate can act on. A gate whose decisions feed the next regression test. Each tool proves a single point today; the work ahead is the connective tissue that turns them into one posture rather than a shelf of parts. That is what I mean by trustworthy AI infrastructure. Not a model I believe. A system I can check.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>agents</category>
      <category>security</category>
    </item>
    <item>
      <title>From Open-Source Programs to Shipping My Own AI Tools</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Sat, 19 Sep 2026 09:30:29 +0000</pubDate>
      <link>https://dev.to/royalpinto007/from-open-source-programs-to-shipping-my-own-ai-tools-2gf6</link>
      <guid>https://dev.to/royalpinto007/from-open-source-programs-to-shipping-my-own-ai-tools-2gf6</guid>
      <description>&lt;p&gt;I did not start out knowing how to build things. I started out reading other people's code, getting confused, and slowly learning how real software actually gets made. A lot of that learning happened inside open-source mentorship programs, and looking back, those programs shaped almost everything about how I work today.&lt;/p&gt;

&lt;h2&gt;
  
  
  The programs that raised me as an engineer
&lt;/h2&gt;

&lt;p&gt;Over the last few years I took part in three structured open-source programs: Google Summer of Code, the Linux Foundation mentorship program (LFX), and Symmetry Autumn of Code. Through them I contributed to open-source compilers and networking systems. That was a big jump for me. Compilers and networking are not the friendliest places to learn. They are large, they are old, and they assume you already understand a lot before you touch anything.&lt;/p&gt;

&lt;p&gt;That difficulty turned out to be the point.&lt;/p&gt;

&lt;p&gt;The first thing these programs taught me was how to work inside a real codebase I did not write. Not a tutorial project, not a fresh repo I could shape however I liked, but a living system with history, conventions, and reasons behind decisions that were not always obvious. I had to read before I could write. I had to figure out why a piece of code existed before I could safely change it. That skill, reading unfamiliar code with patience instead of panic, is probably the most useful thing I own.&lt;/p&gt;

&lt;p&gt;The second thing was mentorship. Having someone more experienced look at my work and tell me honestly what was wrong changed how fast I grew. Not vague encouragement, but specific feedback: this approach will not scale, this edge case is unhandled, this is not how the project does things. It stung sometimes. It also made me better in weeks in ways that would have taken me months alone.&lt;/p&gt;

&lt;p&gt;The third thing was shipping reviewed work. In these programs you do not just write code and walk away. You open it up, you defend it, you revise it, and eventually it gets merged into something people actually use. Learning to take a change all the way from idea to reviewed, accepted contribution is a complete skill on its own, and it is very different from just making something work on your own machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The habit that carried over
&lt;/h2&gt;

&lt;p&gt;Somewhere along the way, contributing to open source stopped being a program I signed up for and became a daily habit. I build open source almost every day now. Over the course of a year that added up to more than three thousand contributions, though honestly the number matters less than the routine behind it. Showing up consistently, in small pieces, is what compounds.&lt;/p&gt;

&lt;p&gt;That habit is what eventually pushed me from contributing to other people's projects toward building my own. Recently I shipped a suite of AI-infrastructure tools: vaultrag, mcp-audit, agentrace, evalgate, voiceeval, answerproof, ctxlens, and injection-arena. They sit around the problems that show up when you actually try to run AI systems in the real world, things like retrieval, auditing, evaluation, and observing what agents are doing.&lt;/p&gt;

&lt;p&gt;Here is the part I want to be honest about. Building these tools did not feel like a heroic leap. It felt like the same muscles I built in those programs, pointed at problems I cared about. Read the landscape first. Understand the existing systems before adding to them. Ship something small and real instead of something big and imaginary. Treat feedback as fuel, not as an attack. None of that is specific to AI. It is just how you build software that other people can trust and use.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest hard part
&lt;/h2&gt;

&lt;p&gt;I do not want to make this sound smooth, because it was not. The hardest thing for me was the gap between feeling behind and being behind. In big, serious codebases, especially compiler and networking work, it is very easy to look around and assume everyone else understands everything and you are the only one lost. For a long time I let that feeling slow me down. I hesitated to ask questions because I did not want to look like I did not belong.&lt;/p&gt;

&lt;p&gt;What eventually helped was realizing that the confusion was not a sign I was in the wrong place. It was the normal texture of doing hard work. The experienced people were not confused less. They were just more comfortable being confused, because they had learned that confusion is where the actual learning lives. Getting comfortable with not knowing, and asking anyway, was harder for me than any technical concept.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to start
&lt;/h2&gt;

&lt;p&gt;A few things I would tell someone standing where I stood a few years ago.&lt;/p&gt;

&lt;p&gt;Pick a real project and read it before you try to change it. Do not rush to your first contribution. Spend time understanding how the thing is put together. Your early value is often in small, careful fixes, not big rewrites.&lt;/p&gt;

&lt;p&gt;Apply to the structured programs. GSoC, LFX, and others like them give you something hard to get on your own: a mentor whose job is to help you, and a real deadline to ship against. That combination is rare and worth a lot.&lt;/p&gt;

&lt;p&gt;Get used to feedback early. The sooner you stop taking code review personally, the faster you grow. A reviewer pointing out a flaw is giving you a gift, even when it does not feel like one in the moment.&lt;/p&gt;

&lt;p&gt;Build the habit, not the highlight. Small contributions almost every day will take you further than occasional bursts. Consistency is quietly the whole game.&lt;/p&gt;

&lt;p&gt;And when you feel lost, keep going anyway. That feeling is not proof you do not belong. It is usually proof you are learning something real.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;I am genuinely grateful to the open-source programs and the people who mentored me through them. They took someone who could barely navigate a large codebase and taught me how to read, contribute, and eventually build. If you are early in this and it feels overwhelming, that is normal, and it does get better. Start small, stay consistent, and let the work compound. I am still doing exactly that, and I hope to keep doing it for a long time.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>career</category>
      <category>ai</category>
      <category>beginners</category>
    </item>
    <item>
      <title>How Do You Actually Test an AI System? A Layered Strategy From Five Tools I Built</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Thu, 17 Sep 2026 09:30:31 +0000</pubDate>
      <link>https://dev.to/royalpinto007/how-do-you-actually-test-an-ai-system-a-layered-strategy-from-five-tools-i-built-2hhb</link>
      <guid>https://dev.to/royalpinto007/how-do-you-actually-test-an-ai-system-a-layered-strategy-from-five-tools-i-built-2hhb</guid>
      <description>&lt;p&gt;Every engineer who has shipped an LLM feature eventually hits the same wall. Your unit tests are green. Nothing throws. And yet the thing is quietly, obviously worse than it was last week. Someone swapped a model, someone tweaked a prompt, someone added a tool, and the output did not error, it just got dumber. No stack trace ever tells you that.&lt;/p&gt;

&lt;p&gt;I have spent a while building tools that try to answer one question honestly: how do you actually test an AI system? Not demo it. Test it. Below is the layered strategy I landed on, mapped to five open-source projects I wrote to make each layer real.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why testing AI is different from testing code
&lt;/h2&gt;

&lt;p&gt;Two things break the assumptions you carry over from normal software.&lt;/p&gt;

&lt;p&gt;The first is nondeterminism. The same input can produce different output on two consecutive calls. A test that asserts equality against a golden string is either flaky or a lie.&lt;/p&gt;

&lt;p&gt;The second is deeper: there is often no single right answer. "Summarize this ticket" has a thousand acceptable outputs and a thousand bad ones, and no &lt;code&gt;assertEqual&lt;/code&gt; can separate them. The output is not correct or incorrect, it is better or worse, more grounded or less, safe or unsafe. Traditional testing asks "did it match?" AI testing has to ask "did it regress?" and "can I trust this particular answer?"&lt;/p&gt;

&lt;p&gt;That reframing is the whole game. You stop trying to prove correctness and start building signals: layered, honest, mostly comparative. Here is how I stacked them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 1: Evals as a build artifact
&lt;/h2&gt;

&lt;p&gt;The base layer is evals, and the trick is to treat quality like something CI can measure, not vibes you check by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;evalgate&lt;/strong&gt; is prompt and agent regression CI. You write a declarative eval suite that lives in version control next to your code, it runs the suite, scores it, and stores a baseline. On every pull request it re-runs, computes the quality delta against the base branch, and fails the build when the score drops, then posts the delta table as a PR comment.&lt;/p&gt;

&lt;p&gt;The important design choice: it does not ask "is this good?" It asks "is this worse than it was?" That is the only question CI can answer objectively. It ships ten scorers, including exact match, regex, JSON-schema, embedding similarity, LLM-as-judge, latency and cost budgets, and weighted rubrics, so you can score the fuzzy stuff and the strict stuff in the same pass. And because it runs on a deterministic mock provider, the whole thing works with zero API keys, offline. A quality gate that needs a paid key to run is a quality gate people turn off.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 2: The modality your text evals cannot see
&lt;/h2&gt;

&lt;p&gt;Evals over text are necessary and completely blind to whole classes of failure. Voice is the sharpest example.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;voiceeval&lt;/strong&gt; exists because if you evaluate a voice agent by reading its transcript, you are evaluating a text agent that happens to have been spoken. The demo failure that made this click for me: a caller asks for a refund, the agent refunds, the transcript agrees with itself perfectly. The caller said "fifteen." The agent heard "fifty." Read the transcript and the call is flawless.&lt;/p&gt;

&lt;p&gt;The fix is a &lt;code&gt;truth&lt;/code&gt; field in the input: what the caller actually said, alongside what the STT heard. Without it, mis-hearing is undetectable by construction, and voiceeval has a test that documents exactly that limitation. From there it catches things a text eval scores as a perfect call: misheard numbers, consequential actions taken without a confirmation, dead air where callers hang up, responses so slow the call already failed. The lesson generalizes past voice: every modality has failures that are invisible in the representation you happen to be evaluating.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 3: Record and replay
&lt;/h2&gt;

&lt;p&gt;Once you have scorers, you need something to score against that reflects real behavior over time. That is replay.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tracecase&lt;/strong&gt; is CI for agents built on record and replay. You define a suite of agent test cases, your CI runs them against the current agent config and POSTs the results, and Tracecase diffs the new run against the previous run of the same suite. It computes what regressed (passed before, fails now) and what got flagged (any safety flag such as a tool call that was not allowed), and returns a &lt;code&gt;shouldFail&lt;/code&gt; signal you wire straight into your CI exit code. The dashboard shows per-case REGRESSED and FIXED diffs with the offending output and tool calls, so a prompt bump that quietly breaks one case cannot slip through as an aggregate that still looks fine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 4: Trust individual answers with receipts
&lt;/h2&gt;

&lt;p&gt;Regression gates protect the system over time. They say nothing about whether one specific answer, right now, can be trusted by someone who was not there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;answerproof&lt;/strong&gt; attaches a cryptographically signed, tamper-evident receipt to every generated answer: which sources were retrieved, which the answer actually used, under whose permissions, with which model and parameters, plus a content hash of each source and a Merkle root over the retrieval set. Anyone can later verify a receipt independently, with nothing but the receipt and the library. It turns "trust our logs" into "verify it yourself," which matters when the leaked document or the compliance question shows up three months later.&lt;/p&gt;

&lt;p&gt;I am careful about what it claims. A receipt proves integrity, source authenticity, set membership, and a transparent record of which claims are grounded in which sources. It does not prove truth. A well-grounded claim can still be wrong if the source is wrong, and the citation binding is n-gram overlap, an auditable signal, not a judge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 5: Observability over what agents actually did
&lt;/h2&gt;

&lt;p&gt;The last layer is watching the runs you already have. When you fan out ten subagents and each returns a confident wall of text, generation is not the bottleneck, verification is, and you cannot verify what you cannot see.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;agentrace&lt;/strong&gt; reads Claude Code session transcripts, shows what your agents actually did, and flags the results you should not trust. No instrumentation, no SDK, because the sessions are already on disk. Every check comes from a real failure across roughly 150 research subagents: an agent concluding a company was not hiring because an API returned an empty list (that API returns empty with HTTP 200 for accounts that do not exist), hedged claims that hardened into facts downstream, twenty URLs cited and none opened, agents dying on session limits mid-sweep. Run against the session that motivated it, it flags 36 of 152 runs, and the most useful number is the 12 where I forgot to specify an output shape. Most agent tooling assumes the model is the problem. Often the prompt was.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one honest caveat
&lt;/h2&gt;

&lt;p&gt;None of these tools tell you what is true. That is not a gap I plan to close, it is the nature of the problem. evalgate tells you the score dropped, not that the new answer is wrong. agentrace and voiceeval are heuristics over text: they tell you what to go read, not what is true, and severity is deliberately conservative because a checker that cries wolf gets switched off, which is worse than no checker. answerproof proves an answer was grounded and unaltered, not that it was correct. Testing an AI system does not give you a green checkmark that means "correct." It gives you layered signals that catch getting worse, catch the invisible failure, and let someone else verify what happened. That is a smaller promise than unit tests make, and it is the honest one.&lt;/p&gt;

&lt;p&gt;If any of this is useful, all five are open source at &lt;a href="https://github.com/royalpinto007" rel="noopener noreferrer"&gt;github.com/royalpinto007&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Shipping an embeddable widget behind one script tag on Cloudflare's free tier</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Tue, 15 Sep 2026 09:30:29 +0000</pubDate>
      <link>https://dev.to/royalpinto007/shipping-an-embeddable-widget-behind-one-script-tag-on-cloudflares-free-tier-53ad</link>
      <guid>https://dev.to/royalpinto007/shipping-an-embeddable-widget-behind-one-script-tag-on-cloudflares-free-tier-53ad</guid>
      <description>&lt;p&gt;Collecting testimonials is easy. Showing them is where everything falls apart.&lt;/p&gt;

&lt;p&gt;You get a nice DM. Someone leaves a glowing comment. A customer emails you something kind. And then it just sits there. To actually put it on your site you screenshot it, crop it, drop the image into your page, and hope the alignment holds up on mobile. Six testimonials later you have six images of six different sizes, no way to moderate what shows, and no idea whether anyone who saw them clicked through. The next time you want to add one, you do the whole dance again.&lt;/p&gt;

&lt;p&gt;I wanted the boring version of this to be one step: paste a script tag, and a wall of approved testimonials appears. That is ProofClip.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core idea
&lt;/h2&gt;

&lt;p&gt;ProofClip is a testimonial wall and social-proof generator for creators and small SaaS. There is a shareable form to collect testimonials, a dashboard to approve or hide them, a public wall of love, and an embeddable widget that drops into any site with a single tag:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;data-proofclip=&lt;/span&gt;&lt;span class="s"&gt;"your-space-slug"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/div&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;script &lt;/span&gt;&lt;span class="na"&gt;src=&lt;/span&gt;&lt;span class="s"&gt;"https://your-proofclip.example/widget.js"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/script&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole integration. No React, no iframe, no build step on your end. The &lt;code&gt;data-proofclip&lt;/code&gt; attribute names which space to render, and the script does the rest.&lt;/p&gt;

&lt;p&gt;The whole thing runs on Cloudflare's free tier: Workers plus &lt;a href="https://hono.dev" rel="noopener noreferrer"&gt;Hono&lt;/a&gt; for the API and the server-rendered pages, D1 (SQLite) for relational data, and R2 for uploaded images. There is no separate frontend build. The widget and card scripts are served as plain strings straight from the Worker.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Collection.&lt;/strong&gt; Each workspace gets a public form at &lt;code&gt;/c/&amp;lt;slug&amp;gt;&lt;/code&gt;. It takes text, a star rating, an optional photo, and a permission flag so the person is explicitly agreeing to be shown. You can also import proof you already have: &lt;code&gt;POST /app/import&lt;/code&gt; lets you upload a screenshot of a DM, a comment, or a review, so testimonials that were never going to fill out a form still make it onto the wall.&lt;/p&gt;

&lt;p&gt;Every testimonial lands in D1 with a status. The schema is blunt about it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="s1"&gt;'pending'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;-- pending | approved | hidden&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So nothing is public by default. New submissions sit in &lt;code&gt;pending&lt;/code&gt; until you say otherwise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Moderation.&lt;/strong&gt; The dashboard at &lt;code&gt;/app&lt;/code&gt; is where you approve, hide, or delete, all through &lt;code&gt;POST /app/testimonial/:action&lt;/code&gt;. Because the public wall and the widget only ever read approved rows, moderation is the gate for everything downstream. There is a settings route too (&lt;code&gt;/app/settings&lt;/code&gt;) for branding: name, accent color, logo, and a toggle for the "Collected with ProofClip" credit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The wall and the widget.&lt;/strong&gt; There is a hosted public wall at &lt;code&gt;/w/&amp;lt;slug&amp;gt;&lt;/code&gt;, and there is the embed. The widget script finds every &lt;code&gt;[data-proofclip]&lt;/code&gt; node on the page, fetches &lt;code&gt;/api/wall/&amp;lt;slug&amp;gt;&lt;/code&gt; for that space, and renders a masonry-style column layout using CSS &lt;code&gt;column-width&lt;/code&gt;, so cards flow into however many columns fit. All user text is escaped through a throwaway DOM node before it touches &lt;code&gt;innerHTML&lt;/code&gt;, so a testimonial cannot inject markup into someone else's page. If branding is enabled for that plan, a small ProofClip credit is appended.&lt;/p&gt;

&lt;p&gt;The widget also does light analytics. On render it fires a &lt;code&gt;view&lt;/code&gt; event via &lt;code&gt;navigator.sendBeacon&lt;/code&gt;, and it records a &lt;code&gt;click&lt;/code&gt; event on interaction, both to &lt;code&gt;POST /api/event&lt;/code&gt;. That is how a workspace sees widget views and clicks without any third-party tracker.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cards.&lt;/strong&gt; For turning a single testimonial into a share-ready image, &lt;code&gt;/card.js&lt;/code&gt; runs a client-side canvas that exports a PNG in 9:16, 1:1, or 16:9. No paid image API is in the loop; the rendering happens in the browser. The card studio is gated to Pro and above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Plans.&lt;/strong&gt; Limits live in &lt;code&gt;src/plans.ts&lt;/code&gt; and are enforced server-side, not just hidden in the UI. Free allows one space, ten testimonials, and one widget, with the branding credit on. Starter ($19) raises the testimonial and widget caps and lets you remove branding. Pro ($39) unlocks unlimited testimonials, multiple spaces, and the card generator. Agency ($79) goes wider still. The &lt;code&gt;accounts.plan&lt;/code&gt; column is the single source of truth, and a protected webhook (&lt;code&gt;POST /api/billing/activate&lt;/code&gt;, plus a Gumroad-shaped route) flips it after a purchase.&lt;/p&gt;

&lt;h2&gt;
  
  
  One honest limitation
&lt;/h2&gt;

&lt;p&gt;The paid-tier roadmap is real work that is not done yet. Video testimonials, custom-domain wiring, team seats, and multiple widget configs per space are listed as not built. The data model was designed to support them, but "the schema has a column for it" is not the same as "it ships." There is also no provider-specific webhook signature verification yet. Billing activation is guarded by a shared secret header, which is fine for a hosted payment link plus a webhook, but it is not the same as verifying a signed payload from a specific provider. If you deployed this to take real money, that is the first thing I would harden.&lt;/p&gt;

&lt;p&gt;I would rather say that plainly than imply the feature grid is fuller than it is. The collect, moderate, embed, and card loop is the part that actually works end to end, and that is the part most people needed anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this shape
&lt;/h2&gt;

&lt;p&gt;Most testimonial tools are either a heavyweight SaaS with a monthly floor that does not make sense for a solo creator, or a pile of manual screenshots. I wanted the middle: something a single person can self-host on a free tier, that treats moderation as a first-class step, and that embeds with the least possible ceremony. Serving the client scripts as strings from the Worker, keeping the data in D1, and putting images in R2 means there is no infrastructure to babysit and nothing that bills you while it sits idle.&lt;/p&gt;

&lt;p&gt;If you want to walk the flow yourself, the README covers local dev: sign up, grab the API key, open the collection link, submit a testimonial, approve it, and view the wall and the embed.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/royalpinto007/ProofClip" rel="noopener noreferrer"&gt;https://github.com/royalpinto007/ProofClip&lt;/a&gt;&lt;/p&gt;

</description>
      <category>cloudflare</category>
      <category>webdev</category>
      <category>typescript</category>
      <category>saas</category>
    </item>
    <item>
      <title>I built BoardEject: an open-source Apple Freeform Excalidraw converter</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Sun, 13 Sep 2026 20:05:31 +0000</pubDate>
      <link>https://dev.to/royalpinto007/i-built-boardeject-an-open-source-apple-freeform-excalidraw-converter-36ka</link>
      <guid>https://dev.to/royalpinto007/i-built-boardeject-an-open-source-apple-freeform-excalidraw-converter-36ka</guid>
      <description>&lt;p&gt;I wanted a way to get &lt;strong&gt;editable content&lt;/strong&gt; out of Apple Freeform instead of only ending up with a flattened PDF.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;BoardEject&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The basic flow is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Freeform → BoardEject → editable &lt;code&gt;.excalidraw&lt;/code&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Where supported, BoardEject preserves things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;shapes&lt;/li&gt;
&lt;li&gt;text&lt;/li&gt;
&lt;li&gt;connectors&lt;/li&gt;
&lt;li&gt;groups&lt;/li&gt;
&lt;li&gt;drawings&lt;/li&gt;
&lt;li&gt;tables&lt;/li&gt;
&lt;li&gt;images&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flbyjcogqy9ck4veat8xe.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flbyjcogqy9ck4veat8xe.gif" alt="boardeject-demo.gif" width="720" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Unsupported or unverified structures are reported instead of silently pretending the conversion worked.&lt;/p&gt;

&lt;p&gt;The interesting part has been validating behavior against &lt;strong&gt;real macOS Freeform captures&lt;/strong&gt; rather than only synthetic fixtures.&lt;/p&gt;

&lt;p&gt;It’s local-first, open source, and the board content isn’t uploaded to a server.&lt;/p&gt;

&lt;p&gt;Try it:&lt;br&gt;
&lt;a href="https://boardeject.dev" rel="noopener noreferrer"&gt;https://boardeject.dev&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Source:&lt;br&gt;
&lt;a href="https://github.com/royalpinto007/boardeject" rel="noopener noreferrer"&gt;https://github.com/royalpinto007/boardeject&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I’m still working through native edge cases, so feedback from people who use Freeform or Excalidraw heavily would be useful.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>webdev</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Parsing bank statements in memory: keep the fields, store nothing else</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Sun, 13 Sep 2026 09:30:29 +0000</pubDate>
      <link>https://dev.to/royalpinto007/parsing-bank-statements-in-memory-keep-the-fields-store-nothing-else-1hj5</link>
      <guid>https://dev.to/royalpinto007/parsing-bank-statements-in-memory-keep-the-fields-store-nothing-else-1hj5</guid>
      <description>&lt;p&gt;Most personal finance apps ask you to hand over a lot. You connect a bank aggregator, or you upload a PDF statement, or you snap a photo of a receipt, and then that raw artifact lives somewhere on a server. The statement that lists every place you spent money for a month. The receipt image with your card's last four digits on it. Even when a company means well, the raw file is now a liability it has to protect, and a thing that can leak.&lt;/p&gt;

&lt;p&gt;I kept coming back to a simpler question while building PennyRush: what does a money tracker actually need to keep? Not the file. It needs the numbers. It needs to know that on a given date you spent an amount at a merchant, and roughly what category that falls under. The original document is just a transport format for those few fields. So the design goal became: extract the fields, then throw the file away before it ever touches disk or a server.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core idea
&lt;/h2&gt;

&lt;p&gt;PennyRush is a native Android app (Kotlin and Jetpack Compose) plus a Next.js web companion, both talking to one Supabase backend. The privacy contract is deliberately narrow and written down in the repo: there is no object storage in v1. When you import a statement or scan a receipt, the file is read into memory, parsed into candidate transactions, and then dropped by the client flow. It is never written to disk, object storage, analytics, or logs. What persists is only the saved activity fields: amount, date, merchant, note, type, and category.&lt;/p&gt;

&lt;p&gt;That constraint shapes everything else. If you are not allowed to keep the file, you have to be good at pulling structure out of it on the spot.&lt;/p&gt;

&lt;h2&gt;
  
  
  How statement import works
&lt;/h2&gt;

&lt;p&gt;CSV import is fully on-device and rules-based. The parser (&lt;code&gt;StatementParser&lt;/code&gt;) takes the file text and does a few defensive things before it trusts anything. It rejects blank files, and it runs a quick binary check: if more than 5% of the first 4KB is non-printable, it assumes you handed it a PDF or an image by mistake and tells you to download a spreadsheet version instead. There are hard caps too, a 5MB byte limit and a 20,000 line limit, so a malformed or huge file cannot hang the import.&lt;/p&gt;

&lt;p&gt;Then it hunts for the header row instead of assuming row zero. It scores each line on whether it contains a date-like column, a money-like column (amount, debit, credit, withdrawal, deposit), and a description-like column (narration, particulars, payee, memo, and so on). Banks all name these differently, so the matching is fuzzy on purpose. Once it locks onto the header, it maps the specific column indices, handles the amount-versus-separate-debit-and-credit case (a debit becomes a negative amount, a credit stays positive), and parses each row.&lt;/p&gt;

&lt;p&gt;Amount parsing has to survive real-world messiness: currency symbols, thousands separators, &lt;code&gt;Rs.&lt;/code&gt;, &lt;code&gt;INR&lt;/code&gt;, &lt;code&gt;USD&lt;/code&gt; prefixes, and accounting-style parentheses for negatives all get stripped or interpreted. Dates are tried against a list of common formats before falling back to ISO parsing. Rows with out-of-range dates (more than 50 years past or 10 years future) or absurd amounts are skipped and counted, so if nothing imports the app can tell you why rather than just failing silently.&lt;/p&gt;

&lt;h2&gt;
  
  
  How receipt scanning works
&lt;/h2&gt;

&lt;p&gt;Receipt scanning uses on-device OCR through ML Kit's text recognizer. You pick an image or take a photo, the app runs &lt;code&gt;TextRecognition&lt;/code&gt; locally to pull the text out, and then it builds a candidate transaction from the recognized lines for you to review before saving. The recognizer is closed as soon as it succeeds, fails, or the coroutine is cancelled. If the read is low-confidence or empty, you still get an editable draft with a warning rather than a dead end. The important part is that the OCR happens on the phone, and the only thing that can end up in the database is the handful of fields you confirm.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning a description into something readable
&lt;/h2&gt;

&lt;p&gt;A raw bank line like &lt;code&gt;UPI/1234567890/SWIGGY/PAYTM&lt;/code&gt; is noise to a human. &lt;code&gt;MerchantExtractor&lt;/code&gt; classifies the transaction kind first (UPI, card, transfer, ATM, salary, bill) using keyword patterns, then tries to recover the actual merchant. It splits the string on separators, drops tokens that look like reference numbers, VPAs, pure digits, or known noise words (&lt;code&gt;UPI&lt;/code&gt;, &lt;code&gt;NEFT&lt;/code&gt;, &lt;code&gt;REF&lt;/code&gt;, &lt;code&gt;PAYTM&lt;/code&gt;, and friends), and scores the survivors, rewarding letters and penalizing digits, to pick the most name-like token. From there &lt;code&gt;CategorizationRules&lt;/code&gt; maps merchants to categories with a plain keyword table: Swiggy and Zomato become Food, Uber and Ola become Transport, Netflix and Spotify become Entertainment, and so on. Income and transfers are handled by sign and kind rather than text.&lt;/p&gt;

&lt;p&gt;All of this is intentionally simple and deterministic. The README is honest that categorization and insights are rules-based today. If server-side AI is added later, the stated rule is that clients never receive model API keys and payloads stay minimized, but none of that is required for the current app to work.&lt;/p&gt;

&lt;h2&gt;
  
  
  One honest limitation
&lt;/h2&gt;

&lt;p&gt;Rules-based extraction is transparent and private, but it is not magic. The keyword tables cover common Indian banks and merchants well and degrade gracefully to an "Other" category everywhere else, which means the first import from an unusual bank or a foreign merchant list will need some manual correction. Because I refuse to keep the raw file, I also cannot silently re-parse it later with a better algorithm. The tradeoff is deliberate: I would rather ship an extractor you can read in one sitting and correct by hand than a black box that keeps your statements around to improve itself. Better merchant and category coverage is ongoing work, not a solved problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the constraint is the feature
&lt;/h2&gt;

&lt;p&gt;The satisfying thing about the in-memory rule is how much it removes. There is no upload pipeline to secure, no storage bucket to scope, no background job that might log a filename. The privacy story is not a policy you have to trust, it is a shape the code physically has: the file exists only for the milliseconds it takes to read the fields, and then it is gone. Everything downstream, dashboards, plans, spending insights, only ever sees the small structured records you chose to save.&lt;/p&gt;

&lt;p&gt;PennyRush is MIT licensed and the full source (Android app, web companion, docs, and privacy contract) is on GitHub: &lt;a href="https://github.com/royalpinto007/PennyRush" rel="noopener noreferrer"&gt;https://github.com/royalpinto007/PennyRush&lt;/a&gt;&lt;/p&gt;

</description>
      <category>android</category>
      <category>kotlin</category>
      <category>privacy</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Turning a Company Website into a Sales-Ready Prospect Record with a Browser Extension</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Fri, 11 Sep 2026 09:30:29 +0000</pubDate>
      <link>https://dev.to/royalpinto007/turning-a-company-website-into-a-sales-ready-prospect-record-with-a-browser-extension-2h77</link>
      <guid>https://dev.to/royalpinto007/turning-a-company-website-into-a-sales-ready-prospect-record-with-a-browser-extension-2h77</guid>
      <description>&lt;p&gt;Sales research is mostly reading. You land on a company's website, click around the homepage, the about page, the product pages, and try to answer a few boring but essential questions: what does this company actually do, who is their customer, are they a fit for what I sell, and who should I email. Then you write the email. Then you do it again for the next company, and the next, until the tab count is a personal insult.&lt;/p&gt;

&lt;p&gt;I built SignalizeAI to collapse that loop. It is a browser extension, shipped on both the Chrome Web Store and Firefox Add-ons, that turns a public company website into a usable prospect record. It has real users today, and this post is about how the pieces fit together.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core idea
&lt;/h2&gt;

&lt;p&gt;The unit of value is a prospect record. Point the extension at the site you are looking at (or paste any URL) and it runs a Quick Website Check, then presents the results in a tabbed insights flow: Strategy, Emails, and Snapshot.&lt;/p&gt;

&lt;p&gt;The Strategy side generates the research a rep would otherwise assemble by hand: what the company does, a company overview, their value proposition, their target customer, a sales-readiness read, a recommended buyer persona, a likely goal, and an outreach angle. The Emails side turns that into something you can actually send: three different outreach approaches, one recommended email, and follow-ups.&lt;/p&gt;

&lt;p&gt;The point is not "AI wrote an email." The point is that the research and the message come out of the same pass over the same source, so the email is grounded in what the company's own site says rather than a generic template with the company name pasted in.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works: the extension
&lt;/h2&gt;

&lt;p&gt;The extension is Manifest V3, written in TypeScript, and bundled with esbuild. Manifest V3 shapes a lot of the architecture. There is no persistent background page, so the flow is built around message passing and a backend that holds the heavy logic rather than the extension trying to be smart on its own.&lt;/p&gt;

&lt;p&gt;One detail I care about is that the extension does not blindly re-analyze. If you have already saved a prospect for a website, it shows the saved analysis by default instead of burning a fresh run on a site you have already researched. Every saved or unsaved analysis can be opened on the web app through an Open in website action, so the extension is a fast capture surface and the web app is the durable workspace.&lt;/p&gt;

&lt;p&gt;Batch Prospecting is where it stops being a toy. You can upload a CSV or paste a list of URLs, run them, multi-select which results to save, and generate outreach and follow-ups in bulk, then export the whole thing to CSV or Excel. For large runs there is a compact batch analysis mode that trades some depth for speed, and there is a resilient fallback path that still produces email content when an AI response fails mid-run, so one bad response does not sink a batch of a hundred.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works: the web app and sync
&lt;/h2&gt;

&lt;p&gt;The extension is paired with the web app at &lt;code&gt;signalizeai.org&lt;/code&gt;, and the two stay in sync rather than being separate products that happen to share a login. Saved prospects live in a workspace at &lt;code&gt;/prospects&lt;/code&gt; with search and filtering, status tracking, inline status editing, and copy, open, and delete actions.&lt;/p&gt;

&lt;p&gt;The interesting part is that the extension and the website share the same prospect data and reflect each other live. The extension listens to the site for auth state, sign-out, theme changes, prospect status updates, prospect content refreshes, and install detection. Change a prospect's status on the website and the extension sees it; save from the extension and it appears in the web workspace. To keep that bridge safe, the production manifests only inject the website bridge script on &lt;code&gt;signalizeai.org&lt;/code&gt;. Dev manifests additionally allow localhost so I can run the web app locally against the extension without opening that channel up in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works: auth, storage, and billing
&lt;/h2&gt;

&lt;p&gt;Auth and storage run on Supabase, with Google sign-in. Supabase is also where saved prospects are stored, which is what makes the shared-data model between extension and website straightforward: both talk to the same backend and the same user identity.&lt;/p&gt;

&lt;p&gt;The analysis and generation work is served by a Cloudflare Workers backend. Workers is a good fit here because prospecting is bursty and latency-sensitive, and running the request-handling logic on the edge keeps it responsive without me babysitting servers. The extension is configured per environment: a dev build points at a dev API host and localhost, while production points at the live API. Plan limits are enforced on that backend, so the free and paid tiers are a real boundary in the service rather than something the client politely respects.&lt;/p&gt;

&lt;h2&gt;
  
  
  An honest limitation
&lt;/h2&gt;

&lt;p&gt;SignalizeAI reads what a company chooses to publish. If a website is thin, vague, or mostly a login screen, the research is only as good as that surface. A sparse marketing site produces a sparser prospect record, and the sales-readiness read reflects the signal that is actually there, not private data the company never put online. In practice this means the tool shines on companies with real content and is weakest exactly where a human would also struggle. The batch fallback helps keep a run from failing outright, but a fallback email is a safety net, not a match for a clean generation on a rich site. I would rather be upfront that this is a research accelerator over public information than pretend it manufactures signal that does not exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;The version of SignalizeAI users are running today is the result of a lot of unglamorous work: Manifest V3 message passing, a two-way sync bridge that has to be locked to a single origin in production, batch runs that survive individual failures, and plan limits enforced where they cannot be bypassed. The payoff is simple to describe. Open a website, get a prospect record and an email grounded in what that company actually says about itself, and keep the whole thing in a workspace that the extension and the site share.&lt;/p&gt;

&lt;p&gt;It is live at &lt;a href="https://signalizeai.org" rel="noopener noreferrer"&gt;https://signalizeai.org&lt;/a&gt;, on both the Chrome and Firefox stores, with real users. If you do outbound and are tired of the tab pile, that is exactly who I built it for.&lt;/p&gt;

</description>
      <category>browserextension</category>
      <category>typescript</category>
      <category>cloudflare</category>
      <category>supabase</category>
    </item>
    <item>
      <title>Designing an offline-first planner that repairs your day instead of shaming you</title>
      <dc:creator>Royal Simpson Pinto</dc:creator>
      <pubDate>Wed, 09 Sep 2026 09:30:29 +0000</pubDate>
      <link>https://dev.to/royalpinto007/designing-an-offline-first-planner-that-repairs-your-day-instead-of-shaming-you-4b88</link>
      <guid>https://dev.to/royalpinto007/designing-an-offline-first-planner-that-repairs-your-day-instead-of-shaming-you-4b88</guid>
      <description>&lt;p&gt;Most planners I have tried ask for too much before they give anything back. You make an account, you sync to a cloud you did not ask for, and then you get judged: streaks broken, tasks glowing red, a little counter reminding you how many days you missed. The tools that are supposed to reduce the friction of a day end up adding a layer of guilt on top of it.&lt;/p&gt;

&lt;p&gt;I wanted the opposite. So I built &lt;strong&gt;Tiny Day&lt;/strong&gt;, a cozy daily planner for Android. Your day, made manageable. No account, no cloud, no analytics.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core idea: gentle and local
&lt;/h2&gt;

&lt;p&gt;Two commitments shaped every decision.&lt;/p&gt;

&lt;p&gt;The first is &lt;strong&gt;offline-first&lt;/strong&gt;. Tiny Day is a React Native app built with Expo (SDK 57). There is no application backend at all. You brain-dump your day in plain language, the app shapes it into a timeline, and everything you type lives on your device. That is not a marketing angle bolted on afterward; it is the architecture. Expo Router owns navigation, Zustand stores own the state, and AsyncStorage persists it. When I say planning is deterministic and local, I mean the scheduler runs on your phone with your data and produces the same result every time. No round trip, no server, works on a plane.&lt;/p&gt;

&lt;p&gt;The second is &lt;strong&gt;gentleness&lt;/strong&gt;. Tiny Day never vibrates. It has no streaks. When you do not finish something, it never auto-carries it to tomorrow behind your back. The whole emotional posture of the app is captured in two lines it actually says to you: &lt;em&gt;"Your day has been repaired,"&lt;/em&gt; and at night, &lt;em&gt;"You did enough for today."&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;The flow starts in the morning. You do not fill in a form field by field. You brain-dump, in plain language, and &lt;code&gt;lib/parse.ts&lt;/code&gt; interprets that text into task cards, pulling out a name, a category, a duration, a priority, and time heuristics. From there &lt;code&gt;lib/schedule.ts&lt;/code&gt; does a priority and energy pass and lays out a timeline. Crucially, it protects meals and deliberately keeps free time open. A planner that fills every minute is just a different kind of stress, so leaving gaps on purpose is a feature, not an oversight.&lt;/p&gt;

&lt;p&gt;Priority in Tiny Day is three tiers, and they are shown as a glyph plus a label, never color alone:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;▲ &lt;strong&gt;must&lt;/strong&gt; gets a reminder plus one follow-up&lt;/li&gt;
&lt;li&gt;● &lt;strong&gt;should&lt;/strong&gt; gets a single reminder&lt;/li&gt;
&lt;li&gt;○ &lt;strong&gt;optional&lt;/strong&gt; stays silent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Notifications are local only, scheduled on the device through expo-notifications. There are quiet hours, and a privacy mode that shows a generic "Important reminder due now" instead of leaking the task name onto your lock screen. At most one gentle replan prompt. That is the ceiling on how much the app is allowed to nag you.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Room&lt;/strong&gt; is the part I am fondest of. It is a layered SVG scene (built with react-native-svg) with a window sky, a lamp glow, a tint overlay, and a tiny character. It has four time states plus a rain variant, it follows the real clock, and it brightens as you complete things and dims into lamplight at night. It is ambient feedback that carries no numbers and no pressure.&lt;/p&gt;

&lt;p&gt;Then there is the feature that named the whole design philosophy: &lt;strong&gt;"My day went wrong."&lt;/strong&gt; Real days fall apart. You wake late, a task runs long, you get too tired, an urgent thing lands. Instead of watching the plan turn into a wall of overdue items, you tap what changed. The repair engine in &lt;code&gt;lib/repair.ts&lt;/code&gt; moves optionals, shortens flexibles, protects your ▲ musts and any fixed appointments, and inserts rest. Then it shows you exactly what it is about to change before it applies anything. You stay in control; the app just does the tedious reshuffling. And because repair is deterministic and local like the scheduler, it is fast and predictable.&lt;/p&gt;

&lt;p&gt;The day closes with an &lt;strong&gt;evening&lt;/strong&gt; review: a mood check-in, gentle stats, and leftovers triage where you decide, per item, whether something goes to tomorrow, to the backlog, or you simply let it go. Nothing moves on its own.&lt;/p&gt;

&lt;p&gt;A few things I made sure not to skip. Accessibility is real: 44px and larger touch targets, dynamic type, reduce-motion and high-contrast toggles, and screen-reader labels. There is a Plan tab with tomorrow, an expandable week, a backlog, and a routine builder. And the Profile can load and remove reversible sample data so you can explore the app without polluting your own tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  One honest limitation
&lt;/h2&gt;

&lt;p&gt;The natural-language parsing is a heuristic, not a language model. &lt;code&gt;lib/parse.ts&lt;/code&gt; reads your brain-dump with rules for durations, categories, priority cues, and time hints. That is exactly why it can run fully offline and stay deterministic, which I consider a fair trade. But it means phrasing the parser has not seen can land in the wrong category or miss a time you clearly implied. The app leans on this by letting you edit any task in place or reschedule it to an explicit future date and time, so a miss is a quick correction rather than a dead end. Still, if you are expecting the free-form understanding of a cloud AI assistant, this is a simpler, more predictable engine, and I would rather be honest about that than oversell it.&lt;/p&gt;

&lt;p&gt;There is also a practical packaging note: the current v1.1.0 APK is ARM64-only, which covers most modern phones and tablets but means a 32-bit-only device needs a separately configured build. Tiny Day targets Android 7.0 (API 24) and newer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Tiny Day is open source under the MIT license. It is a fully offline React Native app: no account, no cloud, no analytics, and no vibration. If a planner that repairs your day instead of shaming you sounds like the kind of tool you have been missing, the code and the latest Android release are here:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/royalpinto007/Tiny-Day" rel="noopener noreferrer"&gt;https://github.com/royalpinto007/Tiny-Day&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You did enough for today.&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>expo</category>
      <category>typescript</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
