<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Athreya aka Maneshwar</title>
    <description>The latest articles on DEV Community by Athreya aka Maneshwar (@lovestaco).</description>
    <link>https://dev.to/lovestaco</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1002302%2F9dbe5057-f6da-4c08-9b5d-37fe9d281476.png</url>
      <title>DEV Community: Athreya aka Maneshwar</title>
      <link>https://dev.to/lovestaco</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lovestaco"/>
    <language>en</language>
    <item>
      <title>GFS Cost More Than 50 Cents a Run. Now Each Run Costs 4 Cents</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Sat, 26 Sep 2026 18:02:42 +0000</pubDate>
      <link>https://dev.to/lovestaco/gfs-cost-more-than-50-cents-a-run-now-it-costs-4-cents-3lf8</link>
      <guid>https://dev.to/lovestaco/gfs-cost-more-than-50-cents-a-run-now-it-costs-4-cents-3lf8</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The bill was 50 cents a run.&lt;/p&gt;

&lt;p&gt;Not a month. Not a thousand runs. One run.&lt;/p&gt;

&lt;p&gt;Paste a design proposal in, get back a grounded review with precedent from our shelf of books and internal postmortems, and watch half a dollar evaporate.&lt;/p&gt;


&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/lovestaco/cheap-rag-in-go-with-gemini-file-search-no-vector-db-two-calls-one-hosted-store-4kb5" class="crayons-story__hidden-navigation-link"&gt;Cheap RAG in Go with Gemini File Search: no vector DB, two calls, one hosted store&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/lovestaco/cheap-rag-in-go-with-gemini-file-search-no-vector-db-two-calls-one-hosted-store-4kb5" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Comments warn about API key store isolation quirks&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/lovestaco" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1002302%2F9dbe5057-f6da-4c08-9b5d-37fe9d281476.png" alt="lovestaco profile" class="crayons-avatar__image" width="800" height="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/lovestaco" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Athreya aka Maneshwar
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Athreya aka Maneshwar
                
                
              
              &lt;div id="story-author-preview-content-4710440" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/lovestaco" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1002302%2F9dbe5057-f6da-4c08-9b5d-37fe9d281476.png" class="crayons-avatar__image" alt="" width="800" height="800"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Athreya aka Maneshwar&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/lovestaco/cheap-rag-in-go-with-gemini-file-search-no-vector-db-two-calls-one-hosted-store-4kb5" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 22&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/lovestaco/cheap-rag-in-go-with-gemini-file-search-no-vector-db-two-calls-one-hosted-store-4kb5" id="article-link-4710440"&gt;
          Cheap RAG in Go with Gemini File Search: no vector DB, two calls, one hosted store
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/go"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;go&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/rag"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;rag&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/gemini"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;gemini&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/lovestaco/cheap-rag-in-go-with-gemini-file-search-no-vector-db-two-calls-one-hosted-store-4kb5" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;47&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/lovestaco/cheap-rag-in-go-with-gemini-file-search-no-vector-db-two-calls-one-hosted-store-4kb5#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              9&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            14 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


&lt;p&gt;Part 1 was the happy version of this story: a Go binary, Gemini's hosted File Search, no vector database, and a free tier that made the whole thing look like a free lunch.&lt;/p&gt;

&lt;p&gt;Then the free tier ran out and the paid rates arrived, and the free lunch turned out to be a tasting menu.&lt;/p&gt;

&lt;p&gt;Five people on the team, several runs a day each, and a number that scaled with how hard the tool was thinking.&lt;/p&gt;

&lt;p&gt;This post is what we replaced it with, what we measured, and the one component we built, benchmarked, and then deleted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The word "free" was doing a lot of work
&lt;/h2&gt;

&lt;p&gt;Here is the thing the pricing page tells you about File Search, and it is all true.&lt;/p&gt;

&lt;p&gt;Storage is free. Query-time embeddings are free. You pay once at indexing time.&lt;/p&gt;

&lt;p&gt;Here is the thing it does not put in bold.&lt;/p&gt;

&lt;p&gt;The search is performed &lt;strong&gt;by a model&lt;/strong&gt;. It is a tool the model calls mid-answer. Which means the retrieved chunks land in that model's context as ordinary input tokens, and the model reasons its way through them before replying.&lt;/p&gt;

&lt;p&gt;And on &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Gemini 3.5 Flash&lt;/a&gt;, reasoning costs $9.00 per million output tokens.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ferpi3ttmyvlkia96ued0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ferpi3ttmyvlkia96ued0.png" alt="Diagram: anatomy of one File Search call. 22,053 uncached input tokens at $1.50 per million is $0.033, 7,903 cached input tokens at $0.15 is $0.001, and 9,529 output tokens at $9.00 is $0.086, for about $0.12 on one search call. 72 percent of the cost is output and 8,493 of the 9,529 output tokens were the model thinking. Below, cost per run as the pipeline changed: $0.70 when every draft ran its own search, $0.12 with one shared search, $0.045 with local retrieval and DeepSeek" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Look at that output row.&lt;/p&gt;

&lt;p&gt;9,529 tokens out, of which 8,493 were the model thinking. For a call whose entire job is "find the relevant passages and hand them over."&lt;/p&gt;

&lt;p&gt;We paid a reasoning model to reason about which paragraphs to copy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz068hxdpykoazk3c3uxe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz068hxdpykoazk3c3uxe.png" alt="Inigo Montoya meme: free retrieval, you keep using that word, I do not think it means what you think it means" width="360" height="195"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The first fix was structural and it helped a lot. The draft step used to do its own search, which meant three drafts meant three searches. Splitting search out so one search feeds every draft took a run from about $0.70 to about $0.12.&lt;/p&gt;

&lt;p&gt;Six times cheaper, and still the wrong shape.&lt;/p&gt;

&lt;p&gt;Because the expensive part was never the searching. It was the thinking attached to it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval does not need a brain.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It needs an index, a similarity function, and a tiebreaker. Every one of those is a thing you can run on a laptop while it charges.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we swapped in
&lt;/h2&gt;

&lt;p&gt;A local &lt;a href="https://www.trychroma.com/" rel="noopener noreferrer"&gt;Chroma&lt;/a&gt; store, embedded with &lt;a href="https://huggingface.co/Qwen/Qwen3-Embedding-0.6B" rel="noopener noreferrer"&gt;Qwen3-Embedding-0.6B&lt;/a&gt;, searched with dense vectors plus BM25, fused with Reciprocal Rank Fusion.&lt;/p&gt;

&lt;p&gt;Every model call moved to DeepSeek V4 Flash on Atlas Cloud at $0.14 in and $0.28 out per million. &lt;/p&gt;

&lt;p&gt;Roughly ten times cheaper per token than what we were on, with a 1M context window and JSON mode.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqe4i7syv8za6iyll3yvm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqe4i7syv8za6iyll3yvm.png" alt="Diagram: what runs where. Billed on Atlas with DeepSeek V4 Flash: call 1 analyze producing a pattern and 3 queries, call 2 drafting three times at once at temperature 0.7, call 3 reviewing. Free on your own machine: the commentor Go binary handling checks, citations and SQLite, rag/serve.py as a child process on port 8791 holding Qwen3-Embedding and BM25, and db/chroma with 3,639 chunks in 69 MB committed to git. Queries come down from call 1, 12 passages go up to the draft verbatim. Retrieval takes 0.5s and costs nothing" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The retrieval service is Python, running as a child process of the Go binary, on a port nobody else talks to. It starts with &lt;code&gt;make run&lt;/code&gt; and dies with it.&lt;/p&gt;

&lt;p&gt;I did try to talk myself out of that boundary.&lt;/p&gt;

&lt;p&gt;Go has no mature CUDA story. The honest route is exporting the model to ONNX and linking a Go runtime through cgo, matching driver and CUDA and cuDNN versions exactly. PyTorch is the path a million people have already debugged, including on WSL2, where CUDA passthrough breaks in creative ways.&lt;/p&gt;

&lt;p&gt;Adding a process boundary was cheaper than adding a whole new class of "works on my machine."&lt;/p&gt;

&lt;h2&gt;
  
  
  Chunking is where the accuracy actually lives
&lt;/h2&gt;

&lt;p&gt;Everyone talks about which embedding model to pick. Almost nobody talks about what you feed it, and that is where the wins were.&lt;/p&gt;

&lt;p&gt;Our corpus is 71 markdown files, and most of them are books that used to be PDFs. PDFs converted to markdown are full of things that are not text.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00mqg1az1l0byn30c3j1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00mqg1az1l0byn30c3j1.png" alt="Diagram: one page of a converted book showing a page number, a running header repeated on every page, a picture-text block and a hyphen line break, all marked for removal, leaving the real sentence. The stored chunk is about 300 words with a hard cap of 450, two sentences of overlap carried forward, and a real section heading starts a new one. The embedded string is the title, kind and heading path followed by the text, so a chunk that only says " width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The cleaner strips the converter's own footer, lines that are only a page number, picture-text blocks, and hyphenated line breaks that split a word across two lines.&lt;/p&gt;

&lt;p&gt;The one that surprised me is running headers.&lt;/p&gt;

&lt;p&gt;A converted book repeats the chapter title at the top of every single page, usually promoted to a markdown heading. &lt;/p&gt;

&lt;p&gt;Left in, it shreds the text into confetti, and a passage about paper reactors comes back with "Ship Project and Civilian Power" wedged into the middle of a sentence.&lt;/p&gt;

&lt;p&gt;The rule that fixed it is embarrassingly simple: a short line that appears five or more times in one file is page furniture, not prose.&lt;/p&gt;

&lt;p&gt;Then the chunking. About 300 words, hard cap 450, split on sentence boundaries, with two sentences carried into the next chunk so a quote that straddles a boundary survives in at least one piece. A real section heading starts a new chunk, once the current one has enough in it to stand alone.&lt;/p&gt;

&lt;p&gt;And then the part I would steal even if you take nothing else from this post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Embed more than you store.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The chunk you store is the text the draft is allowed to quote. The string you embed has a header glued on top of it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;EMBED_MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen/Qwen3-Embedding-0.6B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# what goes into the embedding, per chunk
&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;kind&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;) › &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;heading_path&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# and the query side gets an instruction, because Qwen3-Embedding is
# instruction-tuned on queries only. documents are embedded as they are.
&lt;/span&gt;&lt;span class="n"&gt;QUERY_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Instruct: Given an abstract pattern or principle, retrieve historical cases, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;documented examples, and named principles from books and essays that show &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;the same pattern&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Query: &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;A chunk in the middle of chapter nine might only say "he decided otherwise."&lt;/p&gt;

&lt;p&gt;With the header, that vector still knows it is Rickover, in a book about Rickover, in a section about paper reactors. Without it, it is a pronoun floating in space.&lt;/p&gt;

&lt;p&gt;The query instruction is the Part 1 idea, "search with the pattern, not the post," pushed one layer down into the embedding itself. The corpus is cases. The queries are patterns. Saying so out loud to a model that was trained to listen costs nothing.&lt;/p&gt;
&lt;h2&gt;
  
  
  Dense finds the paraphrase, BM25 finds the name
&lt;/h2&gt;

&lt;p&gt;Vector search is great at "these two paragraphs mean the same thing" and oddly bad at "this paragraph contains the word Rickover."&lt;/p&gt;

&lt;p&gt;Keyword search is the reverse.&lt;/p&gt;

&lt;p&gt;So we run both, take 40 candidates each, and fuse them.&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    Q[one query from call 1] --&amp;gt; DN[dense top 40, Qwen3-Embedding]
    Q --&amp;gt; BM[BM25 top 40, title + text]
    DN --&amp;gt; RRF[Reciprocal Rank Fusion, k=60]
    BM --&amp;gt; RRF
    RRF --&amp;gt; SEL[keep the best 4 per query]
    SEL --&amp;gt; CAP{2 chunks from this file already?}
    CAP -- yes --&amp;gt; SKIP[skip, so no book dominates]
    CAP -- no --&amp;gt; KEEP[keep it]
    KEEP --&amp;gt; RR[round-robin merge, 3 queries]
    SKIP --&amp;gt; RR
    RR --&amp;gt; DUP{near-duplicate of one picked?}
    DUP -- yes --&amp;gt; DROP[drop: a post synced twice]
    DUP -- no --&amp;gt; OUT[12 passages go to the draft]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start    fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef dense    fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a
    classDef lex      fill:#6ea8ff,stroke:#2f5fc4,color:#1a1a1a
    classDef good     fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef bad      fill:#ff9a5c,stroke:#c4602a,color:#1a1a1a

    class CAP,DUP decision
    class Q start
    class DN,RRF dense
    class BM,SEL,RR lex
    class KEEP,OUT good
    class SKIP,DROP bad&lt;/code&gt;&lt;/pre&gt;




&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;hybrid&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CANDIDATES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]]:&lt;/span&gt;
    &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ranked&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dense&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lexical&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ranked&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;RRF_K&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])[:&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That is the whole fusion. Six lines.&lt;/p&gt;

&lt;p&gt;The reason it works is that it never compares the two scores. A cosine similarity of 0.82 and a BM25 score of 14.3 have nothing to say to each other. Ranks do. &lt;/p&gt;

&lt;p&gt;A document at position 3 in both lists beats one that is first in a single list, and the constant &lt;code&gt;k&lt;/code&gt; (60 is the number the literature settled on) keeps the top of each list from steamrolling everything else.&lt;/p&gt;

&lt;p&gt;BM25 indexes the title and heading alongside the text, same as the embedding header does, so naming a book in your query actually finds pages from that book.&lt;/p&gt;

&lt;p&gt;A few rules keep the final twelve honest. At most two chunks per file per query, so one 400-page book cannot fill every slot. Queries merge round robin, so each of the three contributes its best passage before any of them gets its fourth. &lt;/p&gt;

&lt;p&gt;And anything with a word-overlap above 0.6 against something already picked gets dropped, because our blog corpus has a couple of posts that got synced twice and they were politely returning themselves as two independent sources.&lt;/p&gt;
&lt;h2&gt;
  
  
  The reranker that got fired
&lt;/h2&gt;

&lt;p&gt;Standard advice says: retrieve broadly, then rerank with a cross-encoder. So we did that, with &lt;code&gt;bge-reranker-v2-m3&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Then a run took six and a half minutes and I went looking.&lt;/p&gt;

&lt;p&gt;The retrieval service's own log had it in one line: &lt;code&gt;search: 3 queries -&amp;gt; 8 chunks in 117.5s&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Three queries times 40 candidates is 120 cross-encoder forward passes on a 4 GB GTX 1650 that is also drawing the desktop. It was not thrashing. It was just honest work on unfit hardware.&lt;/p&gt;

&lt;p&gt;Before ripping it out, we measured. The golden set builds itself out of finished sessions: take Call 1's retrieval queries, pair them with the files the accepted draft actually cited, and you have a retrieval test built from real usage rather than from vibes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fac1t0y78w4ggignhpj9g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fac1t0y78w4ggignhpj9g.png" alt="Diagram: four retrieval stacks measured on the same golden session. Dense only scores recall@12 of 1.00, hit@4 of 0.00, MRR 0.12 in 4.4s. BM25 only scores 1.00, 1.00, 0.33 in 0.0s. Hybrid fused with RRF scores 1.00, 0.00, 0.17 in 0.5s. Hybrid plus a cross-encoder reranker scores 1.00, 1.00, 0.50 in 108.3s. recall@12 is 1.00 in every row, and the draft is handed all 12 passages anyway, so a better order inside those 12 buys nothing" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The reranker won on MRR and &lt;a href="mailto:hit@4"&gt;hit@4&lt;/a&gt;. It genuinely put better passages nearer the top.&lt;/p&gt;

&lt;p&gt;It also did not change recall@12 at all, and recall@12 is the only number with a consumer.&lt;/p&gt;

&lt;p&gt;The draft call gets all twelve passages in its prompt. It reads all twelve. There is no top-4 cutoff downstream, no truncation, nothing that treats passage 1 differently from passage 9.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsha2y0i6qv4qti6d9zvh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsha2y0i6qv4qti6d9zvh.png" alt="Obi-Wan meme: you were supposed to improve recall, you reordered 12 chunks nobody was ranking and added 108 seconds" width="360" height="205"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So the reranker was spending 108 seconds improving an ordering that nothing downstream reads. It was optimising a metric we had accidentally chosen because it appears in every retrieval paper, not because our pipeline consumed it.&lt;/p&gt;

&lt;p&gt;Out it went, with the reasoning written into the top of &lt;code&gt;retriever.py&lt;/code&gt; so the next person does not "fix" its absence.&lt;/p&gt;

&lt;p&gt;And the honest caveat, which lives there too: this is one golden session. A harder query might genuinely need reranking to pull the right passage into the top twelve. The code is in git history, the eval is a make target, and when the golden set is fat enough to mean something we will run it again.&lt;/p&gt;

&lt;p&gt;Measure before you delete. Also measure before you keep.&lt;/p&gt;
&lt;h2&gt;
  
  
  Two hours of embedding, four laptops
&lt;/h2&gt;

&lt;p&gt;3,639 chunks at roughly 1.8 seconds each on a shared 4 GB GPU is about two hours, which is about one hour and fifty minutes more than anyone wants to wait.&lt;/p&gt;

&lt;p&gt;But embedding is deterministic and the corpus splits cleanly by file. So it parallelises across people, not just across cores.&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    C[71 files, 3,639 chunks] --&amp;gt; S[shard_files: disjoint quarters]
    S --&amp;gt; M[four laptops, --shard i/4]
    M --&amp;gt; R{sha256 AND chunk count match?}
    R -- yes --&amp;gt; SK[already embedded, skip]
    R -- no --&amp;gt; EM[embed, halve the batch on OOM]
    EM --&amp;gt; G[commit the shard, hand it back]
    SK --&amp;gt; G
    G --&amp;gt; I[rag/integrate.py]
    I --&amp;gt; V{same embedding model everywhere?}
    V -- no --&amp;gt; X[refuse: mixed vectors lie quietly]
    V -- yes --&amp;gt; O{every file in exactly one shard?}
    O -- in two --&amp;gt; X
    O -- in none --&amp;gt; W[warn, merge what arrived]
    O -- yes --&amp;gt; F[copy vectors into db/chroma]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start    fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef work     fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a
    classDef box      fill:#6ea8ff,stroke:#2f5fc4,color:#1a1a1a
    classDef good     fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef bad      fill:#ff9a5c,stroke:#c4602a,color:#1a1a1a

    class R,V,O decision
    class C start
    class M,EM work
    class S,G,I,SK box
    class F good
    class X,W bad&lt;/code&gt;&lt;/pre&gt;




&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# on four different machines, one quarter each&lt;/span&gt;
uv run rag/ingest.py &lt;span class="nt"&gt;--shard&lt;/span&gt; 1/4 &lt;span class="nt"&gt;--out&lt;/span&gt; db/chroma-shards/1
uv run rag/ingest.py &lt;span class="nt"&gt;--shard&lt;/span&gt; 2/4 &lt;span class="nt"&gt;--out&lt;/span&gt; db/chroma-shards/2
&lt;span class="c"&gt;# ... then, once everyone commits their shard back&lt;/span&gt;
uv run rag/integrate.py db/chroma-shards/&lt;span class="k"&gt;*&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The merge does no embedding at all. It checks that every shard used the same embedding model, that no file landed in two shards, and that no file landed in none, then copies the vectors into one store.&lt;/p&gt;

&lt;p&gt;Those checks are not paranoia. Mixed embedding models do not crash. They return confidently wrong neighbours forever, which is a far worse failure than a stack trace.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52zaej8meo8altwu3oox.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52zaej8meo8altwu3oox.png" alt="Oprah meme: you get a shard, and you get a shard, everybody gets a shard" width="360" height="269"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Building this also shook out a real bug in the incremental logic. Resume was deciding "already done" by comparing the file's sha256 against what was in the store.&lt;/p&gt;

&lt;p&gt;A run killed halfway through a file leaves that file's sha perfectly correct and its chunk count short, so it would have been marked done forever, silently missing half a book.&lt;/p&gt;

&lt;p&gt;The fix is one &lt;code&gt;AND&lt;/code&gt;: sha256 and chunk count both have to match.&lt;/p&gt;

&lt;p&gt;Then we committed the finished store. &lt;code&gt;db/chroma&lt;/code&gt; is 69 MB in git, which is nothing, and it means nobody else on the team ever embeds anything. Clone, run, search.&lt;/p&gt;
&lt;h2&gt;
  
  
  Now that retrieval is free, spend it on drafts
&lt;/h2&gt;

&lt;p&gt;The nice thing about killing your most expensive call is that the cheap calls get interesting.&lt;/p&gt;

&lt;p&gt;A draft is now a few tenths of a cent. So instead of drafting once and retrying on failure, the pipeline fires three or four drafts at once at temperature 0.7 and lets them race.&lt;/p&gt;

&lt;p&gt;Each reply gets resolved and checked in Go as it lands. Drafts that fail the checks are rejected on the spot. The first one that passes goes on to the review call. Only if all of them fail does the batch retry, with the closest draft's failures appended to the prompt.&lt;/p&gt;

&lt;p&gt;Every draft is kept and shown, including the rejected ones with the checks they failed, because "here are four attempts and why three of them were bad" is more useful to the person reading than one draft and a shrug.&lt;/p&gt;

&lt;p&gt;This did produce one genuinely dumb bug, which parallelism made much more likely.&lt;/p&gt;

&lt;p&gt;In one run the first batch produced a draft that passed every check. The reviewer then asked for a revision. Eight redrafts later, none of them passing, the pipeline shipped the closest failing redraft.&lt;/p&gt;

&lt;p&gt;It had a passing draft in hand and threw it away for a worse one. The fix is the obvious fallback: if no redraft passes, ship the draft that already did, with the reviewer's notes attached.&lt;/p&gt;

&lt;p&gt;Retries are cheap. Losing work you already paid for is not.&lt;/p&gt;
&lt;h2&gt;
  
  
  The check that replaced trusting the grounding metadata
&lt;/h2&gt;

&lt;p&gt;Handing the passages in ourselves unlocked the thing I actually care about.&lt;/p&gt;

&lt;p&gt;When a hosted search tool returns grounding metadata, you know which chunks the model looked at. You do not know that the words it put in quotation marks are in any of them.&lt;/p&gt;

&lt;p&gt;Now the pipeline can prove it.&lt;/p&gt;

&lt;p&gt;Every &lt;code&gt;history.sources[].passage&lt;/code&gt; in the output has to appear word for word in a chunk from the file it names. Same for an authority's &lt;code&gt;exact_words&lt;/code&gt; when it cites one of the given files. Markdown, line wrapping and quote style are normalised away first, and a &lt;code&gt;...&lt;/code&gt; marks an omission, with the remaining pieces required to appear in order.&lt;/p&gt;

&lt;p&gt;A failure is not a warning. It is a check failure, exactly like a length violation or an unresolved law citation, and it feeds the same retry loop that everything else does.&lt;/p&gt;

&lt;p&gt;This is the difference between "the model had access to the right book" and "the model quoted the right book correctly," and only one of those is worth showing a reviewer.&lt;/p&gt;
&lt;h2&gt;
  
  
  What the bill looks like now
&lt;/h2&gt;

&lt;p&gt;One real run, end to end, with 15 model calls including 12 drafts across three batches:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;210,976 input tokens, 55,285 output tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$0.045&lt;/strong&gt;, about ₹4.3&lt;/li&gt;
&lt;li&gt;retrieval: 0.5 seconds of it, and none of the money&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Against $0.70 for the same pipeline shape on hosted search. Fifteen times cheaper, and the expensive part is now the part that does the actual writing, which is how it should be.&lt;/p&gt;

&lt;p&gt;Two honest footnotes on that number.&lt;/p&gt;

&lt;p&gt;Atlas reports most of each draft's input as cached, and we charge every input token at the full rate, so $0.045 is a ceiling, not an estimate.&lt;/p&gt;

&lt;p&gt;And latency went up, not down. Retrieval dropped from 117 seconds to half a second, but DeepSeek thinks hard before each draft, so a full run is still two to four minutes. &lt;/p&gt;

&lt;p&gt;We bought cost, not speed. For "paste a design doc, come back with a coffee," that is the right trade. For a chat box it would be the wrong one.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I would steal
&lt;/h2&gt;

&lt;p&gt;If your retrieval is a model call, you are paying reasoning prices for a lookup.&lt;/p&gt;

&lt;p&gt;Pull it onto your own machine, spend the effort on cleaning and chunking rather than on model selection, embed a contextual header you never show anyone, fuse dense and lexical on ranks instead of scores, and check the quotes rather than trusting them.&lt;/p&gt;

&lt;p&gt;Then build the eval before you build the clever part. Ours told us to throw the clever part away, which saved 108 seconds a search and a permanent dependency we did not need.&lt;/p&gt;

&lt;p&gt;The best component in this system is the one that is not in it.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>ai</category>
      <category>go</category>
      <category>python</category>
    </item>
    <item>
      <title>Cheap RAG in Go with Gemini File Search: no vector DB, two calls, one hosted store</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Tue, 22 Sep 2026 07:12:46 +0000</pubDate>
      <link>https://dev.to/lovestaco/cheap-rag-in-go-with-gemini-file-search-no-vector-db-two-calls-one-hosted-store-4kb5</link>
      <guid>https://dev.to/lovestaco/cheap-rag-in-go-with-gemini-file-search-no-vector-db-two-calls-one-hosted-store-4kb5</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every team I have been on has the same shelf.&lt;/p&gt;

&lt;p&gt;Nine engineering books somebody swears by, a couple of long essays that get linked in every third design review, and sixty-odd internal blog posts and postmortems that nobody rereads.&lt;/p&gt;

&lt;p&gt;About 1.4 million tokens of "we already learned this once."&lt;/p&gt;

&lt;p&gt;I wanted a tool where you paste a design proposal, an ADR, a postmortem draft, and it comes back with: here is what is actually being proposed, here is the pattern underneath it, and here are the three places on the shelf where we, or someone smarter, already ran into this exact shape.&lt;/p&gt;

&lt;p&gt;Grounded. With citations. Not "an LLM read your doc and had feelings about it."&lt;/p&gt;

&lt;p&gt;The obvious build is a vector database, an embedding pipeline, a chunker, a reranker, and a weekend.&lt;/p&gt;

&lt;p&gt;I built it with none of those, in Go, on Gemini's free tier, and the retrieval side of it costs nothing to run.&lt;/p&gt;

&lt;p&gt;This post is about how, and about the four things that bit me on the way.&lt;/p&gt;

&lt;h2&gt;
  
  
  RAG without owning the R
&lt;/h2&gt;

&lt;p&gt;Gemini has a thing called &lt;a href="https://ai.google.dev/gemini-api/docs/file-search" rel="noopener noreferrer"&gt;File Search&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;You create a "store", upload documents into it, and Google chunks them, embeds them, and indexes them.&lt;/p&gt;

&lt;p&gt;Then you attach that store to a normal &lt;code&gt;generateContent&lt;/code&gt; call as a tool, and the model searches it by itself, mid-answer, and hands you back the chunks it used as &lt;code&gt;groundingMetadata&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;No pgvector. No Pinecone bill. No embedding model to pick and then regret.&lt;/p&gt;

&lt;p&gt;The pricing is the part that made me sit up.&lt;/p&gt;

&lt;p&gt;Storage is free. Query-time embeddings are free. You pay once at indexing time, at &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;embedding prices&lt;/a&gt;, and then the chunks the model pulls in are billed as ordinary context tokens on the call you were already making.&lt;/p&gt;

&lt;p&gt;On the free tier the store caps at 1 GB. My entire shelf, every book and every post converted to markdown, is 5.3 MB.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh09pkzx4z6kx7dd8adud.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh09pkzx4z6kx7dd8adud.png" alt="Diagram: the whole pipeline. Once: 71 markdown files go through a sha256 diff into an ingest step that uploads to a Gemini File Search store. Every question: a design proposal goes to call 1 which analyzes it without tools and emits a generalization plus retrieval queries, then call 2 drafts with the File Search tool attached and returns grounding chunks, then deterministic checks with up to two retries produce a review JSON with sources and cited laws" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So the architecture is embarrassingly short.&lt;/p&gt;

&lt;p&gt;One Go binary. SQLite via &lt;code&gt;modernc.org/sqlite&lt;/code&gt;, so no cgo. A REST client for Gemini with no SDK, because I wanted to log the exact request and response bodies verbatim, and SDKs love to hide those.&lt;/p&gt;

&lt;p&gt;Two model calls per question. One hosted store. That is the whole thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step zero: getting books into markdown without losing the spaces
&lt;/h2&gt;

&lt;p&gt;Before any of the clever parts, the corpus has to exist as text, and this cost me an evening.&lt;/p&gt;

&lt;p&gt;The first tool I reached for was &lt;a href="https://github.com/microsoft/markitdown" rel="noopener noreferrer"&gt;markitdown&lt;/a&gt;. It gave me headings, sort of, and it also gave me this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Theubiquityoffrustrating,unhelpfulsoftwareinterfaceshasmotivateddecadesofresearch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Every word in the PDF glued to its neighbour. The embedding model does not know what "Theubiquityoffrustrating" is, and neither does the retrieval.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pdftotext&lt;/code&gt; fixed the spacing and threw away every heading, so a 300-page book became one undifferentiated scroll.&lt;/p&gt;

&lt;p&gt;The one that won was &lt;a href="https://pymupdf.readthedocs.io/en/latest/pymupdf4llm/" rel="noopener noreferrer"&gt;pymupdf4llm&lt;/a&gt;. It looks at font sizes to decide what is a heading, keeps bold and italics, and its text extraction handled the spacing correctly on the same PDF.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run &lt;span class="nt"&gt;--with&lt;/span&gt; pymupdf4llm python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"
import pymupdf4llm, pathlib
pathlib.Path('out.md').write_text(pymupdf4llm.to_markdown('in.pdf'))
"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Forty-four real headings out of one essay, versus zero.&lt;/p&gt;

&lt;p&gt;Headings matter more than they look like they should here, because File Search chunks on whitespace with a token budget. A chunk that starts at a heading is a chunk that means something on its own. A chunk that starts mid-sentence in a wall of text is noise with an embedding attached.&lt;/p&gt;

&lt;p&gt;I put the conversion behind &lt;code&gt;make process-data&lt;/code&gt; so nobody on the team has to rediscover this. EPUBs, PDFs, and a sync of our blog repos, all into one &lt;code&gt;post_processed_data/&lt;/code&gt; tree of markdown.&lt;/p&gt;
&lt;h2&gt;
  
  
  The store belongs to the key, not to you
&lt;/h2&gt;

&lt;p&gt;This is the first thing that bit me, and it is not in the big print.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3b15zaljtwr5tdnwwmte.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3b15zaljtwr5tdnwwmte.png" alt="Matrix Morpheus meme: what if I told you the File Search store belongs to the API key, not to your app" width="360" height="218"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A File Search store lives inside the Google Cloud project behind the API key that created it.&lt;/p&gt;

&lt;p&gt;Create a store with key A, try to search it with key B from a different project, and it does not error in a helpful way. It just is not there.&lt;/p&gt;

&lt;p&gt;That matters the moment you have more than one key, which on the free tier you will, because the per-minute caps are real and the fix everyone reaches for is "rotate across a few keys."&lt;/p&gt;

&lt;p&gt;You cannot rotate a tool-attached call across keys. The store pins you.&lt;/p&gt;

&lt;p&gt;So keys got roles.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F37p4y7ze6v9klzeiam7o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F37p4y7ze6v9klzeiam7o.png" alt="Diagram: keys split by role. keys/store.txt has key 0 owning store 0 (primary) and key 1 owning store 1 (fallback), and call 2 tries store 0 with key 0 then falls to store 1 with key 1, never key 1 on store 0. keys/rotate.txt has keys 2 through 5 rotated on 429, 401, 403 and 5xx for call 1 and corpus index summaries which have no tool attached" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;keys/store.txt&lt;/code&gt; is the short list. Each key in it owns a complete copy of the corpus in its own store. Ingest uploads to every one of them. The search call tries store 0 with key 0, and only on a quota or auth failure falls through to store 1 with key 1.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;keys/rotate.txt&lt;/code&gt; is the long list, for any call with no tool attached: the analysis call, the corpus-index summaries. Those genuinely do not care which key answers, so they rotate on 429, 401, 403, and anything 5xx.&lt;/p&gt;

&lt;p&gt;Two things I got wrong before I got them right:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;503 is not a key problem.&lt;/strong&gt; Gemini returns 503 "high demand" when Google is busy, and switching keys does nothing except burn another key's quota. So the client waits and retries the same key a couple of times before falling over.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The key goes in the &lt;code&gt;x-goog-api-key&lt;/code&gt; header, never the URL.&lt;/strong&gt; Put it in the query string and the first transport error prints your key into your own logs. Ask me how I know.&lt;/p&gt;

&lt;p&gt;Yes, two stores means the one-time indexing cost happens twice. It is the price of the search call never dying on a single quota, and at embedding prices for 5 MB, it is coffee money.&lt;/p&gt;
&lt;h2&gt;
  
  
  Ingest is a diff, not an upload
&lt;/h2&gt;

&lt;p&gt;The ingest step is where "just upload the folder" turns into an actual program.&lt;/p&gt;

&lt;p&gt;Every file gets a sha256. A manifest in SQLite records &lt;code&gt;path -&amp;gt; sha -&amp;gt; document id&lt;/code&gt; for each store. Then:&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[walk data/**/*.md, sha256 each] --&amp;gt; B{path in manifest?}
    B -- no --&amp;gt; U[upload to store]
    B -- yes --&amp;gt; C{sha changed?}
    C -- no --&amp;gt; S[skip, unchanged]
    C -- yes --&amp;gt; D[delete old doc id] --&amp;gt; U
    U --&amp;gt; P[poll the indexing operation until done]
    P --&amp;gt; M[upsert manifest: path, sha, doc id]
    A --&amp;gt; R{manifest path missing on disk?}
    R -- yes --&amp;gt; X[delete doc from store, drop manifest row]
    R -- no --&amp;gt; S
    M --&amp;gt; I{set of path+sha changed?}
    S --&amp;gt; I
    X --&amp;gt; I
    I -- no --&amp;gt; K[corpus index unchanged]
    I -- yes --&amp;gt; G[summarize new files with flash-lite, one line each] --&amp;gt; J[store corpus_index]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start    fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef net      fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef local    fill:#6ea8ff,stroke:#2f5fc4,color:#1a1a1a
    classDef idx      fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a

    class B,C,R,I decision
    class A start
    class U,D,P,X net
    class S,M,K local
    class G,J idx&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Unchanged files are skipped. Changed files have their old document deleted first, then get re-uploaded. Files that vanished from disk get deleted from the store.&lt;/p&gt;

&lt;p&gt;Each upload is a multipart POST that returns a long-running operation, and you poll it until &lt;code&gt;done&lt;/code&gt; before you trust the document id. Four uploads run in parallel. Serial, a 71-file corpus takes about ten minutes. Parallel, a few.&lt;/p&gt;

&lt;p&gt;The chunking is set per upload, and I landed on 400 tokens with 60 of overlap. The docs' example is 200 and 20, which for a book felt like reading through a letterbox.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;meta&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="k"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="n"&gt;any&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="s"&gt;"displayName"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;displayName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s"&gt;"chunkingConfig"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="k"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="n"&gt;any&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="s"&gt;"whiteSpaceConfig"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="k"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="s"&gt;"maxTokensPerChunk"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="m"&gt;400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="s"&gt;"maxOverlapTokens"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;  &lt;span class="m"&gt;60&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c"&gt;// multipart/related: part 1 is this JSON, part 2 is the markdown bytes&lt;/span&gt;
&lt;span class="c"&gt;// POST /upload/v1beta/{store}:uploadToFileSearchStore&lt;/span&gt;
&lt;span class="c"&gt;// with X-Goog-Upload-Protocol: multipart&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The last box in that flowchart is the corpus index, and it is small but it is the thing that makes the next section work.&lt;/p&gt;

&lt;p&gt;For every file, a one-sentence summary from the cheapest model available (&lt;code&gt;gemini-3.5-flash-lite&lt;/code&gt;), keyed by the file's sha so it is only ever generated once. All the lines get joined into one block, "path, summary", that gets pasted into the analysis prompt.&lt;/p&gt;

&lt;p&gt;That way the model deciding what to search for knows what is actually on the shelf. It aims at sources that exist instead of guessing.&lt;/p&gt;
&lt;h2&gt;
  
  
  Search with the pattern, not the post
&lt;/h2&gt;

&lt;p&gt;This is the idea in the post that I would keep if I had to throw everything else away.&lt;/p&gt;

&lt;p&gt;Embedding search finds things that sound alike.&lt;/p&gt;

&lt;p&gt;Paste a proposal that says "we should standardize on one MCP transport across all our agents" straight into retrieval, and you get back every chunk that contains the words "MCP", "agents", and "standardize". Which is your own recent blog posts about MCP and agents.&lt;/p&gt;

&lt;p&gt;That is not precedent. That is a mirror.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3hdajakprhf4roek2qjq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3hdajakprhf4roek2qjq.png" alt="Diagram: the same proposal searched two ways. Raw text as the query returns blog posts about MCP and agents, chunks about the subject with no precedent. The generalization as the query, fragmented incompatible implementations to one open standard to mass adoption, returns the browser wars, Rickover on standardization, and Han Fei on uniform law, the same pattern in other clothes" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What you actually want is the shape of the situation with the nouns removed.&lt;/p&gt;

&lt;p&gt;"Fragmented, incompatible implementations, then one open standard, then mass adoption."&lt;/p&gt;

&lt;p&gt;Search the shelf with that, and the browser wars come back. Rickover on standardizing the nuclear navy comes back. A 2,300-year-old Legalist essay on uniform law comes back.&lt;/p&gt;

&lt;p&gt;Same pattern, different clothes. That is what a reviewer with thirty years of reading brings, and it is what the model cannot do if you let it search with the raw text.&lt;/p&gt;

&lt;p&gt;So the pipeline never lets it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Call 1&lt;/strong&gt; gets the lawbook, the corpus index, and the proposal, with no tools attached. It returns JSON: the motive, what is actually happening, who the actors are, the generalization, and two or three retrieval queries derived from that generalization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Call 2&lt;/strong&gt; gets the analysis from call 1, the File Search tool pointed at the store, and the instruction to search with those queries and then draft.&lt;/p&gt;

&lt;p&gt;The raw document is in call 2's context, so the draft can quote it. But the search terms came from call 1, and they are about the pattern, not the subject.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"systemInstruction"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"parts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"...the lawbook..."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"parts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"PROPOSAL:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;...&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;ANALYSIS FROM CALL 1:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;{...}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;Search the corpus with the retrieval queries above, then reply with the JSON contract."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"fileSearch"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"fileSearchStoreNames"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"fileSearchStores/abc123"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"generationConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"responseMimeType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"application/json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"temperature"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The response carries &lt;code&gt;candidates[0].groundingMetadata.groundingChunks&lt;/code&gt;, each with the file's display name and the retrieved text. Those get shown to the user, in full, next to the draft. If the "authority" the review leans on is not in that list, it did not come from the shelf, and the checks catch that.&lt;/p&gt;
&lt;h2&gt;
  
  
  Every instruction is a law, every reply cites its laws
&lt;/h2&gt;

&lt;p&gt;There is no free-text system prompt anywhere in this thing.&lt;/p&gt;

&lt;p&gt;The prompts are an &lt;a href="https://github.com/shrsv/AgentLaws" rel="noopener noreferrer"&gt;AgentLaws&lt;/a&gt; lawbook: a folder of markdown where every instruction the model sees is a numbered law, grouped into chapters, versioned in git, and compiled once at startup.&lt;/p&gt;

&lt;p&gt;The model is required to return &lt;code&gt;applied_laws&lt;/code&gt;, the numbers it relied on, and the pipeline resolves each one back to &lt;code&gt;file:line&lt;/code&gt; in the lawbook.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxi82fvjo5fhuqcqfr1xi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxi82fvjo5fhuqcqfr1xi.png" alt="Star Wars checkpoint officer meme: applied_laws 2.4.1 resolves to analysis/generalization.md:14, an older law, sir, but it resolves" width="360" height="194"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A number that does not resolve fails the run. Which sounds pedantic until the first time the model confidently cites law 7.2.3 in a lawbook with six chapters.&lt;/p&gt;

&lt;p&gt;The practical win is that "why did it do that" has an answer that is a file and a line number, and "make it stop doing that" is a pull request against a markdown file, not archaeology in a Go string.&lt;/p&gt;

&lt;p&gt;The output contract is the last thing in the prompt, because that is where the model weights it most. Laws about tone and structure go first, the JSON schema goes last.&lt;/p&gt;
&lt;h2&gt;
  
  
  Trust, but run the checks
&lt;/h2&gt;

&lt;p&gt;The model also fills a &lt;code&gt;checks&lt;/code&gt; block in its own reply, a self-audit. I do not trust it, but it is useful as a second signal.&lt;/p&gt;

&lt;p&gt;Before that gets read, a small pile of deterministic checks in plain Go runs on every draft: length bounds, no markdown, at most one question, no "see point N" references, every cited law resolved.&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[call 2: search + draft, JSON mode] --&amp;gt; B{reply parses as JSON?}
    B -- no --&amp;gt; F[resend once without JSON mode, extract first object] --&amp;gt; C
    B -- yes --&amp;gt; C[resolve every cited law number to file:line]
    C --&amp;gt; D{all citations resolve?}
    D -- no --&amp;gt; R
    D -- yes --&amp;gt; E[deterministic checks: length, no markdown, one question max, distinct openers]
    E --&amp;gt; G{checks pass?}
    G -- no --&amp;gt; R{attempts left?}
    R -- yes --&amp;gt; H[append the failures to the user message] --&amp;gt; A
    R -- no --&amp;gt; S[ship the draft with failures listed]
    G -- yes --&amp;gt; V[call 3: reviewer pass, verdict + objections]
    V --&amp;gt; W{verdict?}
    W -- ship --&amp;gt; O[return JSON with sources and applied laws]
    W -- revise, rounds left --&amp;gt; H
    W -- revise, no rounds --&amp;gt; S

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef llm      fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef local    fill:#6ea8ff,stroke:#2f5fc4,color:#1a1a1a
    classDef bad      fill:#ff9a5c,stroke:#c4602a,color:#1a1a1a
    classDef good     fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a

    class B,D,G,R,W decision
    class A,F,V llm
    class C,E,H local
    class S bad
    class O good&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;A failure gets appended to the user message as a plain list, "the previous attempt failed these checks, fix them", and call 2 runs again. Two retries, then the draft ships anyway with the failures listed, because a draft with a warning is more useful than a spinner that gave up.&lt;/p&gt;

&lt;p&gt;One box in there deserves its own sentence: &lt;strong&gt;JSON mode plus a tool is not guaranteed&lt;/strong&gt;. &lt;code&gt;responseMimeType: application/json&lt;/code&gt; together with File Search worked most of the time and then, occasionally, returned prose with a JSON object somewhere in it. The fix is boring: if the reply does not parse, resend once without JSON mode and pull the first &lt;code&gt;{...}&lt;/code&gt; out of the text. Both attempts get logged.&lt;/p&gt;
&lt;h2&gt;
  
  
  The bug that made me split the database
&lt;/h2&gt;

&lt;p&gt;The tool is used by five people, and each of them runs it against their own SQLite file, so history is per person.&lt;/p&gt;

&lt;p&gt;The first version put everything in that one file: sessions, steps, and the ingest manifest.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9sk2s29l1a6ynv6wjawn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9sk2s29l1a6ynv6wjawn.png" alt="Bad Luck Brian meme: creates a fresh SQLite DB for a teammate, re-uploads all 71 files and duplicates every doc in the store" width="360" height="360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Spot it?&lt;/p&gt;

&lt;p&gt;A new teammate makes a new database. New database, empty manifest. Empty manifest means every file on disk is "new", so all 71 get uploaded again. And because the manifest is empty, ingest does not know the old document ids, so nothing gets deleted.&lt;/p&gt;

&lt;p&gt;The store now has two of everything. Do it five times and search quality quietly rots, because every query returns the same chunk five times and crowds out the second-best hit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqlxpga3o08jpzyp59dan.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqlxpga3o08jpzyp59dan.png" alt="Diagram: three per-person session databases, alice.db, bob.db, you.db, each holding sessions and steps, all reading from one shared corpus.db that holds the store names per key, the manifest of path to sha256 to doc id, and the corpus index. corpus.db mirrors the Google File Search store of 71 documents chunked at 400 tokens with 60 overlap. The bug this fixed: manifest lived in the session db, so a new teammate got an empty manifest, re-uploaded all 71 files, deleted nothing, and the store filled with duplicates" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The fix is one sentence: the manifest is a property of the store, not of whoever is asking questions today.&lt;/p&gt;

&lt;p&gt;So there are two databases now. &lt;code&gt;corpus.db&lt;/code&gt; holds the store names, the manifest, and the corpus index, and every session database reads from it. &lt;code&gt;alice.db&lt;/code&gt; holds Alice's sessions and nothing else.&lt;/p&gt;

&lt;p&gt;Delete &lt;code&gt;corpus.db&lt;/code&gt; and you get a full re-ingest. Delete &lt;code&gt;alice.db&lt;/code&gt; and nothing about the corpus changes.&lt;/p&gt;
&lt;h2&gt;
  
  
  What the free tier actually gives you
&lt;/h2&gt;

&lt;p&gt;Since "cheap" is in the title, the honest numbers, as of when I built this.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Models.&lt;/strong&gt; Flash only. &lt;code&gt;gemini-3.5-flash&lt;/code&gt; for the real calls, &lt;code&gt;gemini-3.5-flash-lite&lt;/code&gt; for the throwaway summaries. Pro returns a 429 with &lt;code&gt;limit: 0&lt;/code&gt; on free keys. A paid key unlocks it with one env var, and I have not needed to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rate limits.&lt;/strong&gt; Per-minute caps are the thing you hit, not per-day. Ingest backs off with a growing wait when every key is rate limited, and the UI shows each key fallover as it happens instead of hiding it behind a spinner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Storage.&lt;/strong&gt; 1 GB per store. My 5.3 MB corpus does not register.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost per question.&lt;/strong&gt; Two Flash calls, sometimes three with the review pass, around 10k tokens in and 2k out. The retrieved chunks are already counted in that. On the free tier, zero. On a paid key, well under a cent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Latency.&lt;/strong&gt; 30 to 90 seconds end to end, most of it call 2 doing the search and the draft. That is fine for "paste a design doc, get a review" and would be wrong for a chat box.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The catch.&lt;/strong&gt; On the free tier, Google may train on what you upload. My shelf is published books and public blog posts, so I do not care. If your corpus is your customer contracts, read &lt;a href="https://ai.google.dev/gemini-api/terms" rel="noopener noreferrer"&gt;the terms&lt;/a&gt; before you &lt;code&gt;make ingest&lt;/code&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I would tell you to steal
&lt;/h2&gt;

&lt;p&gt;If you have a shelf, and a question shaped like "have we seen this before", you do not need a vector database to answer it.&lt;/p&gt;

&lt;p&gt;You need markdown with real headings, a hosted store you diff against instead of re-uploading, keys with roles because the store pins you to one, and a first model call whose only job is to turn the question into the pattern underneath it.&lt;/p&gt;

&lt;p&gt;The rest is checks and logging, and the checks are the part you will thank yourself for.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>go</category>
      <category>ai</category>
      <category>rag</category>
      <category>gemini</category>
    </item>
    <item>
      <title>How to Count 100 Billion Things in 12 Kilobytes</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Wed, 16 Sep 2026 12:23:43 +0000</pubDate>
      <link>https://dev.to/lovestaco/hyperloglog-how-to-count-100-billion-things-in-12-kilobytes-5aae</link>
      <guid>https://dev.to/lovestaco/hyperloglog-how-to-count-100-billion-things-in-12-kilobytes-5aae</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;An interviewer asks you a question that sounds almost insultingly easy.&lt;/p&gt;

&lt;p&gt;How many unique visitors did the site have today?&lt;/p&gt;

&lt;p&gt;You say: easy, I'll hash every request ID into a set. &lt;code&gt;HashSet&amp;lt;String&amp;gt;&lt;/code&gt;, one entry per unique visitor, done before the coffee's cold.&lt;/p&gt;

&lt;p&gt;Then they say: it's a hundred billion requests.&lt;/p&gt;

&lt;p&gt;Your hash set just quietly asked the OS for about a terabyte of RAM, and the OS just as quietly said no.&lt;/p&gt;

&lt;p&gt;You could shard it. Spin up forty machines, split the ID space, merge partial sets at query time.&lt;/p&gt;

&lt;p&gt;That works. It is also an entire distributed system you now have to run, just to answer "how many different people showed up."&lt;/p&gt;

&lt;p&gt;There's a much smaller way to answer this, and it doesn't even need to remember the visitors.&lt;/p&gt;

&lt;h2&gt;
  
  
  The naive fix, and the bill for it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2trn58fvgwj96s5s2hsx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2trn58fvgwj96s5s2hsx.png" alt="Diagram: the same 100 billion request IDs going into a hash set that needs about a terabyte of RAM, versus going into HyperLogLog which needs about 12 kilobytes regardless of how many IDs there are" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A hash set is exact because it remembers everything.&lt;/p&gt;

&lt;p&gt;Every single unique ID gets a slot, forever, because that's the only way to know for certain you haven't seen it before.&lt;/p&gt;

&lt;p&gt;That's also exactly why it doesn't scale. Memory grows linearly with the number of uniques, no matter how you slice it, shard it, or compress the keys.&lt;/p&gt;

&lt;p&gt;If your product only ever has a few hundred thousand daily uniques, this is a complete non-problem. Use the hash set, go home.&lt;/p&gt;

&lt;p&gt;But once you're talking billions, "remember everything" stops being an engineering decision and starts being a bet against your cloud bill.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxd02sscr0vlhuziajtt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxd02sscr0vlhuziajtt.png" alt="Midwit meme about counting exact uniques with a hash table on the left, an overcomplicated sharded hash table in the middle, and going back to one register array with a harmonic mean on the right" width="360" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The way out isn't a bigger hash set. It's giving up on "exact" entirely, on purpose, in exchange for something that fits in your CPU's L2 cache.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trick: rare things are evidence of scale
&lt;/h2&gt;

&lt;p&gt;Here's the actual insight, stripped of the mechanism around it.&lt;/p&gt;

&lt;p&gt;Hash every ID into a long, effectively random bit string.&lt;/p&gt;

&lt;p&gt;Now look at how many zeros it has at the start, before the first &lt;code&gt;1&lt;/code&gt; bit shows up.&lt;/p&gt;

&lt;p&gt;Since each bit is a coin flip, the odds are simple: 50% of hashes start with at least one zero, 25% start with at least two, 12.5% with at least three, and so on. Each extra leading zero you demand halves how many hashes will have it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8azuemj5p3wi9uq33hk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8azuemj5p3wi9uq33hk.png" alt="Diagram: four example hashes with increasingly long runs of leading zeros highlighted, each row noting how rare that run is, from 1-in-2 odds up to 1-in-64 odds" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Flip that around and it becomes a counting trick.&lt;/p&gt;

&lt;p&gt;If you've hashed a handful of IDs and none of them start with three zeros, that's unremarkable, you'd expect that from 8 items. But if you've seen a hash that starts with six zeros, that's a 1-in-64 event. Seeing it &lt;em&gt;once&lt;/em&gt; is weak evidence you've hashed something on the order of 64 items, because you'd need to try roughly that many random hashes before one that rare shows up.&lt;/p&gt;

&lt;p&gt;So: hash every incoming ID, and keep a running max of the longest leading-zero streak you have ever seen. That single number, "longest streak seen so far," gives you a rough estimate of &lt;code&gt;2^streak&lt;/code&gt; unique items.&lt;/p&gt;

&lt;p&gt;This idea goes back to Flajolet and Martin's original 1985 probabilistic counting paper, and the modern form is called &lt;a href="https://en.wikipedia.org/wiki/HyperLogLog" rel="noopener noreferrer"&gt;HyperLogLog&lt;/a&gt;, from &lt;a href="http://algo.inria.fr/flajolet/Publications/FlFuGaMe07.pdf" rel="noopener noreferrer"&gt;Flajolet, Fusy, Gandouet and Meunier's 2007 paper&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Here's the single-counter version, in about ten lines, so you can see exactly how weak it is on its own:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;leading_zeros&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bits&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;bits&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;bits&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;bit_length&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;single_counter_estimate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_items&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;max_streak&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_items&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;h&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getrandbits&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;max_streak&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_streak&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;leading_zeros&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="n"&gt;max_streak&lt;/span&gt;

&lt;span class="c1"&gt;# run it a few times on the same n and watch the estimate swing wildly
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;single_counter_estimate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10_000&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Run that a handful of times against the same &lt;code&gt;n_items&lt;/code&gt; and the estimates will swing by 2x, 4x, sometimes more in either direction. One lucky hash with an extra-long streak, or one unlucky run without one, and the whole estimate moves.&lt;/p&gt;

&lt;p&gt;A single coin flip streak is just too noisy to trust. HyperLogLog doesn't use one counter, it uses thousands, and averages them in a very specific way.&lt;/p&gt;
&lt;h2&gt;
  
  
  Averaging away the noise
&lt;/h2&gt;

&lt;p&gt;Instead of one counter, take the first few bits of each hash and use them to pick one of &lt;code&gt;m&lt;/code&gt; buckets, say &lt;code&gt;m = 16384&lt;/code&gt;. The remaining bits of the hash get the leading-zero treatment from before, and the result updates that bucket's own running max.&lt;/p&gt;

&lt;p&gt;You now have thousands of tiny, independent "longest streak I've seen" estimators, each looking at a different slice of the ID space. Average them, and the noise from any one unlucky or lucky streak gets smoothed out by the other 16,383 buckets.&lt;/p&gt;

&lt;p&gt;There's one more wrinkle worth knowing. HyperLogLog uses the &lt;em&gt;harmonic mean&lt;/em&gt; across buckets, not the arithmetic mean. A plain average gets wrecked by a single bucket that got a freakishly long streak, since &lt;code&gt;2^streak&lt;/code&gt; grows exponentially, one outlier bucket can dominate a normal average completely. The harmonic mean punishes large outliers far more than small ones, which is exactly the failure mode you need to guard against here.&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[ID arrives] --&amp;gt; B[Hash it to a uniform bitstring]
    B --&amp;gt; C[First p bits pick a bucket]
    C --&amp;gt; D[Count leading zeros in the rest]
    D --&amp;gt; E{Longer streak than&amp;lt;br/&amp;gt;this bucket has seen?}
    E --&amp;gt;|Yes| F[Store it as the bucket's max]
    E --&amp;gt;|No| G[Discard, bucket keeps its max]
    F --&amp;gt; H[m buckets, each holding one max streak]
    G --&amp;gt; H
    H --&amp;gt; I[Harmonic mean across all buckets]
    I --&amp;gt; J[Bias-correct for small and huge counts]
    J --&amp;gt; K[Cardinality estimate, ~2% error]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef step fill:#6ea8ff,stroke:#2f5fbf,color:#1a1a1a
    classDef store fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef result fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a

    class A start
    class E decision
    class B,C,D step
    class F,G,H store
    class I,J,K result&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;With &lt;code&gt;m&lt;/code&gt; buckets, the standard error works out to roughly &lt;code&gt;1.04 / sqrt(m)&lt;/code&gt;. Plug in 16,384 buckets and you land at just over 0.8% expected error, which is where Redis's implementation sits.&lt;/p&gt;

&lt;p&gt;That error bound doesn't depend on whether you counted a million items or a hundred billion. It only depends on &lt;code&gt;m&lt;/code&gt;, the number of buckets, which is a number you pick up front.&lt;/p&gt;

&lt;p&gt;That's the whole trade. You fix your error tolerance once, at design time, by choosing &lt;code&gt;m&lt;/code&gt;, and the memory cost stays flat forever after. A few small corrections handle the edges: &lt;a href="https://en.wikipedia.org/wiki/HyperLogLog#Practical_considerations" rel="noopener noreferrer"&gt;linear counting&lt;/a&gt; kicks in when very few buckets have been touched yet, so a tiny actual count doesn't get wildly overestimated, and a large-range correction avoids hash collisions skewing things once you approach 2^32 distinct items.&lt;/p&gt;
&lt;h2&gt;
  
  
  Somebody already built this into your database
&lt;/h2&gt;

&lt;p&gt;The best part of HyperLogLog isn't the estimate, it's that two of them merge for free.&lt;/p&gt;

&lt;p&gt;Since each bucket is just "the max streak seen," merging two HyperLogLogs means taking the elementwise max of their buckets. No replay, no recomputation, no access to the original IDs at all. That's why it works so well for things like "daily uniques" that you also want to roll up into "weekly uniques."&lt;/p&gt;

&lt;p&gt;Redis has had this built in for years, as three commands:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# add IDs as they arrive, one PFADD per event&lt;/span&gt;
PFADD visitors:2026-09-16 user_881 user_204 user_991

&lt;span class="c"&gt;# get the estimated cardinality, ~12KB per key no matter how big it gets&lt;/span&gt;
PFCOUNT visitors:2026-09-16

&lt;span class="c"&gt;# merge daily counters into a weekly one, no replaying the day's events&lt;/span&gt;
PFMERGE visitors:week-38 visitors:2026-09-12 visitors:2026-09-13 visitors:2026-09-14 visitors:2026-09-15 visitors:2026-09-16
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://redis.io/docs/latest/commands/pfadd/" rel="noopener noreferrer"&gt;Redis's PFADD docs&lt;/a&gt; put the standard error at 0.81%, using &lt;a href="https://research.google/pubs/hyperloglog-in-practice-algorithmic-engineering-of-a-state-of-the-art-cardinality-estimation-algorithm/" rel="noopener noreferrer"&gt;Google's HyperLogLog++ paper&lt;/a&gt; for the small-and-large range corrections on top of the original algorithm.&lt;/p&gt;

&lt;p&gt;It's not just Redis either. &lt;a href="https://cloud.google.com/bigquery/docs/reference/standard-sql/approximate_aggregate_functions" rel="noopener noreferrer"&gt;BigQuery's &lt;code&gt;APPROX_COUNT_DISTINCT&lt;/code&gt;&lt;/a&gt;, Presto and Trino's &lt;code&gt;approx_distinct&lt;/code&gt;, and Spark's &lt;code&gt;approx_count_distinct&lt;/code&gt; are all the same idea wearing a SQL function's clothes. Every "unique visitors" widget on every analytics dashboard you've ever glanced at is very possibly running this exact algorithm under the hood, right now, while you read this sentence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz1n1up3b22z4uyk3tl4w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz1n1up3b22z4uyk3tl4w.png" alt="Woman yelling at a cat meme: one side demanding the exact unique count, the cat calmly responding with a 2% error estimate in 12 kilobytes" width="360" height="231"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  When you actually still want the hash set
&lt;/h2&gt;

&lt;p&gt;None of this means throw away exact counting. It means know which question you're answering.&lt;/p&gt;

&lt;p&gt;If the number needs to be exact, billing a customer per API call, deduplicating a payment, deciding whether a user already redeemed a coupon, HyperLogLog is the wrong tool. It cannot tell you whether one specific ID has been seen, only roughly how many distinct ones have.&lt;/p&gt;

&lt;p&gt;If your cardinality is small anyway, a few thousand, a few hundred thousand, a plain hash set fits in memory with room to spare and gives you an exact answer for free. Reach for HyperLogLog when the count itself is the product, "how many uniques," "how many distinct IPs," "how many distinct search terms," and the scale is big enough that remembering every item stops being realistic.&lt;/p&gt;

&lt;p&gt;That's the whole pitch. You trade a guarantee you never actually needed for a memory bill you can actually afford, with error bounds you get to choose in advance.&lt;/p&gt;

&lt;p&gt;Next time somebody says "a hundred billion" in an interview, you now have a much better answer than "shard it."&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>algorithms</category>
      <category>redis</category>
      <category>computerscience</category>
    </item>
    <item>
      <title>MVC, MVP, MVVM, MVVM-C, VIPER: one pattern wearing five outfits</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Tue, 15 Sep 2026 17:47:49 +0000</pubDate>
      <link>https://dev.to/lovestaco/mvc-mvp-mvvm-mvvm-c-viper-one-pattern-wearing-five-outfits-3b73</link>
      <guid>https://dev.to/lovestaco/mvc-mvp-mvvm-mvvm-c-viper-one-pattern-wearing-five-outfits-3b73</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every mobile team eventually has the same argument.&lt;/p&gt;

&lt;p&gt;Someone says the &lt;a href="https://en.wikipedia.org/wiki/Model%E2%80%93view%E2%80%93controller" rel="noopener noreferrer"&gt;ViewController&lt;/a&gt; is 3,000 lines long and needs to be broken up.&lt;/p&gt;

&lt;p&gt;Someone else says "just use MVVM," like that phrase alone fixes anything.&lt;/p&gt;

&lt;p&gt;A third person mentions &lt;a href="https://christiantietze.de/wiki/viper/" rel="noopener noreferrer"&gt;VIPER&lt;/a&gt; and the room goes quiet, because everyone has a VIPER war story.&lt;/p&gt;

&lt;p&gt;I went down this rabbit hole recently and realized something that should have been obvious years ago: MVC, &lt;a href="https://en.wikipedia.org/wiki/Model%E2%80%93view%E2%80%93presenter" rel="noopener noreferrer"&gt;MVP&lt;/a&gt;, MVVM, MVVM-C and VIPER are not five different ideas.&lt;/p&gt;

&lt;p&gt;They're one idea, wearing five different amounts of clothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one shape underneath all of them
&lt;/h2&gt;

&lt;p&gt;Strip away the acronyms and every pattern is arguing about the same three things.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;View&lt;/strong&gt;, which is the face of the app. It renders pixels and captures your taps.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;Model&lt;/strong&gt;, the brain, which owns the business logic and the data.&lt;/p&gt;

&lt;p&gt;And a translator in between, whose entire job is making sure the View and the Model never talk to each other directly.&lt;/p&gt;

&lt;p&gt;That translator is the part that keeps getting renamed. Controller. Presenter. View-Model. Everything else is a debate about how much power to give it, and who else gets to share the job.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4tau3y53yzevtqnmo1iy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4tau3y53yzevtqnmo1iy.png" alt="Diagram: View, a translator layer, and Model, the shape every one of these patterns reuses" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Model%E2%80%93view%E2%80%93controller" rel="noopener noreferrer"&gt;Model-View-Controller&lt;/a&gt; is genuinely old, coming out of Smalltalk work at Xerox PARC in the late 1970s, which makes it nearly as old as the personal computer itself.&lt;/p&gt;

&lt;p&gt;It was built to separate "what the data is" from "what's on screen," and for a single-window desktop app in 1978, that was already a big improvement over one giant blob of code that did everything.&lt;/p&gt;

&lt;p&gt;To see where each pattern actually differs, let's run the exact same tiny feature through all five: a user tapping their profile picture to pick a new one.&lt;/p&gt;

&lt;h2&gt;
  
  
  MVC: everything reports to one controller
&lt;/h2&gt;

&lt;p&gt;MVC connects the View and the Model through a Controller that does all the coordinating.&lt;/p&gt;

&lt;p&gt;You tap the picture, the View tells the Controller, the Controller updates the Model, and once that's done the Controller turns around and tells the View to refresh.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5pz292k3rc1i21ndcu3o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5pz292k3rc1i21ndcu3o.png" alt="Diagram: View notifies Controller, Controller updates Model, Controller refreshes View" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It reads clean on a whiteboard, and honestly it's a fine choice for a small app. That's not a knock, it's the actual selling point.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ProfileController&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;onPhotoPicked&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setAvatar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;// update the model&lt;/span&gt;
    &lt;span class="nx"&gt;view&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;render&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;avatar&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;// tell the view to refresh&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The catch is that this Controller doesn't stay this small. Every new feature wants to talk to the Model somehow, and the Controller is the only door in the building.&lt;/p&gt;

&lt;p&gt;Six months later it's not a controller, it's a lobby that every department in the company has to walk through.&lt;/p&gt;
&lt;h2&gt;
  
  
  MVP: the Presenter takes the homework off the View
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Model%E2%80%93view%E2%80%93presenter" rel="noopener noreferrer"&gt;MVP&lt;/a&gt; answers the bloated-controller problem by handing UI logic to a dedicated Presenter, and telling the View to do nothing but draw.&lt;/p&gt;

&lt;p&gt;Tap the photo, the View notifies the Presenter, the Presenter updates the Model, formats whatever comes back into something display-ready, and pushes that formatted result straight into the View.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8mvklcq1sd89cjf08wk2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8mvklcq1sd89cjf08wk2.png" alt="Diagram: View notifies Presenter, Presenter updates Model and formats data back to View" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The View in MVP is deliberately dumb, which is a compliment here. A dumb View is a View you can unit test the Presenter against without spinning up any UI at all, since the Presenter only ever talks to a thin interface the View implements.&lt;/p&gt;

&lt;p&gt;That testability is the entire reason teams reach for MVP. Not because it's fancier than MVC, but because "does the right thing happen" becomes a question you can answer without a simulator or a device.&lt;/p&gt;
&lt;h2&gt;
  
  
  MVVM: the View stops asking, and starts just knowing
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://learn.microsoft.com/en-us/dotnet/architecture/maui/mvvm" rel="noopener noreferrer"&gt;MVVM&lt;/a&gt; swaps the Presenter's manual "update the view" call for two-way data binding between the View and a View-Model.&lt;/p&gt;

&lt;p&gt;Pick a new photo, the View pushes that change into the View-Model through binding, the View-Model persists it to the Model, and when the Model's data changes, that change flows back to the bound View property automatically. No explicit refresh call, anywhere.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qj4qbcj11ssn7jv9cxy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qj4qbcj11ssn7jv9cxy.png" alt="Diagram: View and ViewModel bound both ways, ViewModel and Model bound both ways" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the pattern that plays nicest with reactive frameworks, because reactive frameworks are basically data binding with better marketing.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="kt"&gt;ProfileViewModel&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;@Observable&lt;/span&gt; &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;avatar&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Image&lt;/span&gt;   &lt;span class="c1"&gt;// View binds to this directly&lt;/span&gt;

  &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;onPhotoPicked&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;avatar&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;image&lt;/span&gt;        &lt;span class="c1"&gt;// View updates instantly via binding&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;persist&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;// no explicit "refresh the view" call anywhere&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Less boilerplate, definitely. The tradeoff shows up later, when a bug means tracing a chain of automatic notifications instead of reading a linear function call, which is a genuinely different debugging skill.&lt;/p&gt;
&lt;h2&gt;
  
  
  MVVM-C: somebody still has to decide where the app goes next
&lt;/h2&gt;

&lt;p&gt;MVVM never says who's in charge of moving between screens, and by default that job quietly lands on the View-Model, which is exactly the thing MVVM was trying to keep simple.&lt;/p&gt;

&lt;p&gt;MVVM-C fixes that by adding a Coordinator that sits above the binding triangle and owns navigation on its own. In our example, that's the move from the profile screen to the image picker and back, including whatever save logic happens in between.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjxyq98h2jruw5neht8fj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjxyq98h2jruw5neht8fj.png" alt="Diagram: Coordinator controls the ViewModel, same binding triangle as MVVM underneath" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Same View, View-Model, Model relationship as plain MVVM. The only new thing is that "where do we go next" has its own dedicated home, instead of leaking into whichever View-Model happened to trigger the transition.&lt;/p&gt;
&lt;h2&gt;
  
  
  VIPER: five jobs, five files, no exceptions
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.objc.io/issues/13-architecture/viper/" rel="noopener noreferrer"&gt;VIPER&lt;/a&gt; takes the "give everyone exactly one job" idea about as far as it can go.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;V&lt;/strong&gt;iew displays and forwards user actions. &lt;strong&gt;I&lt;/strong&gt;nteractor owns the business logic. &lt;strong&gt;P&lt;/strong&gt;resenter prepares data for the View and talks to the Interactor. &lt;strong&gt;E&lt;/strong&gt;ntity is the raw data, the Model equivalent. &lt;strong&gt;R&lt;/strong&gt;outer owns navigation.&lt;/p&gt;

&lt;p&gt;Tap the photo: View tells Presenter, Presenter tells Interactor to do the actual work, Interactor manages the Entity, Presenter formats the result back for the View, and if a screen change is needed, Router handles it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5xypmxwe6q5938m1sdke.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5xypmxwe6q5938m1sdke.png" alt="Diagram: View, Presenter, Interactor, Entity and Router, five components each with one job" width="800" height="609"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;protocol&lt;/span&gt; &lt;span class="kt"&gt;PhotoInteractorProtocol&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;updateAvatar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="kd"&gt;protocol&lt;/span&gt; &lt;span class="kt"&gt;PhotoPresenterProtocol&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;didPickPhoto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;didUpdateAvatar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="kd"&gt;protocol&lt;/span&gt; &lt;span class="kt"&gt;PhotoRouterProtocol&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;closePhotoPicker&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Notice there's no shared base class holding this together, just protocols. That's the whole VIPER pitch: maximum separation, maximum testability, and yes, maximum file count for updating a single profile picture.&lt;/p&gt;

&lt;p&gt;It earns its keep on genuinely large apps with multiple teams working the same codebase. On a five-screen app it's mostly ceremony.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpvh2j15afk0ffechx8ik.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpvh2j15afk0ffechx8ik.png" alt=" " width="360" height="216"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  So which one do you actually pick
&lt;/h2&gt;

&lt;p&gt;None of them are wrong, they're just answers to different problems.&lt;/p&gt;

&lt;p&gt;MVC for a small app where shipping fast matters more than architecture purity. &lt;/p&gt;

&lt;p&gt;MVP once you need the mediator to be testable in isolation but don't need data binding. &lt;/p&gt;

&lt;p&gt;MVVM once your framework is reactive and binding removes real boilerplate. &lt;/p&gt;

&lt;p&gt;MVVM-C once that same app also has navigation flows worth centralizing. &lt;/p&gt;

&lt;p&gt;VIPER once the app and the team are both big enough that "everyone has one job" stops being overhead and starts being the thing that keeps ten engineers from stepping on each other.&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Starting a new client app] --&amp;gt; B{Small app, small team?}
    B --&amp;gt;|Yes| MVC[MVC: simplicity wins, ship it]
    B --&amp;gt;|No| C{Need the mediator unit-testable?}
    C --&amp;gt;|Yes, no data binding needed| MVP[MVP: Presenter tests clean in isolation]
    C --&amp;gt;|Need reactive data binding| D{Complex navigation flows too?}
    D --&amp;gt;|No| MVVM[MVVM: binding kills boilerplate]
    D --&amp;gt;|Yes| E{Large app, many teams?}
    E --&amp;gt;|Not yet| MVVMC[MVVM-C: Coordinator owns navigation]
    E --&amp;gt;|Yes| VIPER[VIPER: five single-job pieces, max modularity]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start    fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef pick     fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef heavy    fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a

    class A start
    class B,C,D,E decision
    class MVC,MVP,MVVM pick
    class MVVMC,VIPER heavy&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The pattern isn't the architecture. It's just how much ceremony your team is willing to pay for, in exchange for how much chaos it's trying to avoid.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5rxbsljm0d9oth7ls2me.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5rxbsljm0d9oth7ls2me.png" alt=" " width="360" height="486"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Pick the smallest one that still lets you sleep at night, and upgrade only when the pain is real, not hypothetical.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>mobile</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>The ffmpeg Pipeline Nobody Explains</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Mon, 14 Sep 2026 11:34:19 +0000</pubDate>
      <link>https://dev.to/lovestaco/the-ffmpeg-pipeline-nobody-explains-7d8</link>
      <guid>https://dev.to/lovestaco/the-ffmpeg-pipeline-nobody-explains-7d8</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;ffmpeg is the one CLI tool everyone uses and most doesn't understand&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every developer has run this exact incantation at some point in their life, copy-pasted from a Stack Overflow answer from 2014, with zero idea what it does:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; input.mov &lt;span class="nt"&gt;-vcodec&lt;/span&gt; h264 &lt;span class="nt"&gt;-acodec&lt;/span&gt; mp2 output.mp4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;It works. You move on. You never ask what &lt;code&gt;-vcodec&lt;/code&gt; actually does, or why some conversions finish in a blink and others chew your CPU for ten minutes.&lt;/p&gt;

&lt;p&gt;Turns out there's a genuinely elegant pipeline hiding under the hood, and once you see it, half of ffmpeg's flag soup starts making sense on its own.&lt;/p&gt;

&lt;p&gt;ffmpeg was created by &lt;a href="https://bellard.org/" rel="noopener noreferrer"&gt;Fabrice Bellard&lt;/a&gt; back in the year 2000, and the name is a mashup of "fast forward" and MPEG, the Moving Picture Experts Group, the folks behind most of the video formats you've heard of.&lt;/p&gt;

&lt;p&gt;It's not just a CLI toy either.&lt;/p&gt;

&lt;p&gt;It's the encode/decode engine quietly running inside Chrome, Blender, YouTube, Vimeo, and roughly half of every video-adjacent tool you've ever used.&lt;br&gt;
Source: &lt;a href="https://ffmpeg.org/about.html" rel="noopener noreferrer"&gt;ffmpeg.org&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  The pipeline nobody tells you about
&lt;/h2&gt;

&lt;p&gt;Here's the thing that clicked for me: ffmpeg isn't one operation, it's a &lt;em&gt;pipeline&lt;/em&gt;, and every flag you pass just tweaks one stage of it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftgv6j72c9rpbm7srxjp3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftgv6j72c9rpbm7srxjp3.png" alt="ffmpeg pipeline: input to demux to decode to filter to encode to mux to output" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Walk through it left to right:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Input&lt;/strong&gt; — your file lands, &lt;code&gt;in.mp4&lt;/code&gt;, whatever.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demux&lt;/strong&gt; — a demultiplexer splits the container into its separate streams. Your "video file" was never one thing, it's a video track, an audio track, maybe subtitles, all interleaved into one file for convenience.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decode&lt;/strong&gt; — each stream's compressed packets get decoded into raw, uncompressed frames. This is the CPU-hungry part.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Filter&lt;/strong&gt; — &lt;em&gt;optional.&lt;/em&gt; This is the only stage that actually touches pixels or audio samples: brightness, contrast, scaling, adding subtitles, drawing a waveform. Skip it and nothing changes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Encode&lt;/strong&gt; — raw frames get compressed back down into packets, in whatever codec you asked for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mux&lt;/strong&gt; — a multiplexer interleaves the encoded streams back into one output container.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output&lt;/strong&gt; — &lt;code&gt;out.mkv&lt;/code&gt;, done.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;code&gt;ffmpeg -i in.mp4 out.mp4&lt;/code&gt; runs the whole thing with sane defaults. &lt;code&gt;ffprobe&lt;/code&gt; just runs the first couple of stages and prints what it finds, no encode required, which is why it's instant.&lt;/p&gt;

&lt;p&gt;With over 100 codecs supported, that same seven-stop pipeline is how ffmpeg decodes, encodes, transcodes, filters, and plays basically any multimedia file that exists. Source: &lt;a href="https://ffmpeg.org/about.html" rel="noopener noreferrer"&gt;ffmpeg.org/about.html&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  The flag that changes everything: &lt;code&gt;-c copy&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Once you see the pipeline, one thing jumps out immediately: decode and encode are the only expensive stages. Demux and mux are just bookkeeping, shuffling bytes around.&lt;/p&gt;

&lt;p&gt;So what happens if you skip decode and encode entirely?&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; input.mp4 &lt;span class="nt"&gt;-c&lt;/span&gt; copy output.mkv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1618cb0qa29kk4mhlbhg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1618cb0qa29kk4mhlbhg.png" alt="-c copy skips decode and encode entirely, vs -c:v libx264 which runs the full pipeline" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;-c copy&lt;/code&gt; tells ffmpeg "don't touch the streams, just repackage them." It demuxes, then immediately muxes into the new container. &lt;/p&gt;

&lt;p&gt;No decode, no encode, no quality loss, because the actual video bytes never change, they just get moved into a different box.&lt;/p&gt;

&lt;p&gt;This is why converting &lt;code&gt;.mp4&lt;/code&gt; to &lt;code&gt;.mkv&lt;/code&gt; is instant, but converting &lt;code&gt;.mp4&lt;/code&gt; to &lt;code&gt;.webm&lt;/code&gt; is not: mp4 and mkv can both hold H.264 video, so it's a pure repackage. &lt;/p&gt;

&lt;p&gt;webm wants VP8/VP9/AV1, a codec mp4 usually isn't carrying, so ffmpeg has no choice but to actually decode and re-encode every frame.&lt;/p&gt;

&lt;p&gt;Once you've internalized that, most of ffmpeg's common recipes stop being magic incantations and start being obvious:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; input.mov &lt;span class="nt"&gt;-c&lt;/span&gt; copy &lt;span class="nt"&gt;-ss&lt;/span&gt; 00:00:30 &lt;span class="nt"&gt;-t&lt;/span&gt; 00:00:10 clip.mov
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Trim ten seconds starting at 0:30. Since we're not changing codecs, &lt;code&gt;-c copy&lt;/code&gt; keeps it instant. Same file, smaller slice.&lt;/p&gt;

&lt;p&gt;Need to glue several clips together? List them in a text file and hand it to the concat demuxer:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-f&lt;/span&gt; concat &lt;span class="nt"&gt;-safe&lt;/span&gt; 0 &lt;span class="nt"&gt;-i&lt;/span&gt; list.txt &lt;span class="nt"&gt;-c&lt;/span&gt; copy joined.mp4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Still &lt;code&gt;-c copy&lt;/code&gt;. Still no re-encode. It's the same pipeline principle, just applied to multiple inputs instead of one.&lt;/p&gt;
&lt;h2&gt;
  
  
  When you actually need the expensive path
&lt;/h2&gt;

&lt;p&gt;Sometimes copy isn't an option, because you're deliberately changing the pixels or the codec, not just the wrapper. That's when &lt;code&gt;-vf&lt;/code&gt; (video filter) and real encode flags come in:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; input.mp4 &lt;span class="nt"&gt;-vf&lt;/span&gt; &lt;span class="s2"&gt;"scale=1280:720"&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; 30 &lt;span class="nt"&gt;-b&lt;/span&gt;:v 2M &lt;span class="nt"&gt;-c&lt;/span&gt;:v libx264 output.mp4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;-vf scale&lt;/code&gt; resizes, &lt;code&gt;-r&lt;/code&gt; sets the frame rate, &lt;code&gt;-b:v&lt;/code&gt; sets the video bitrate, &lt;code&gt;-c:v libx264&lt;/code&gt; picks the encoder. Every one of these forces a real decode → filter → encode pass, which is exactly why this version is slow and &lt;code&gt;-c copy&lt;/code&gt; isn't.&lt;/p&gt;

&lt;p&gt;Subtitles go through the same filter stage. Got an &lt;code&gt;.srt&lt;/code&gt; file? Convert it to &lt;code&gt;.ass&lt;/code&gt; first, then burn it in with &lt;code&gt;-vf subtitles&lt;/code&gt;, since &lt;code&gt;-vf&lt;/code&gt; is the only stage in the whole pipeline that's allowed to touch a frame.&lt;/p&gt;

&lt;p&gt;Here's the decision tree I actually keep in my head now, roughly:&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[I have a media file and ffmpeg] --&amp;gt; B{Just changing container, same codecs?}
    B --&amp;gt;|yes| B1[-c copy]
    B --&amp;gt;|no| C{Need a smaller file or a different codec?}
    C --&amp;gt;|yes| C1[-c:v libx264 -c:a aac]
    C --&amp;gt;|no| D{Just trimming a section?}
    D --&amp;gt;|yes| D1[-ss start -t dur -c copy]
    D --&amp;gt;|no| E{Joining multiple clips together?}
    E --&amp;gt;|yes| E1[concat demuxer + -c copy]
    E --&amp;gt;|no| F{Changing resolution, framerate or bitrate?}
    F --&amp;gt;|yes| F1[-vf scale, -r, -b:v]
    F --&amp;gt;|no| G{Burning in subtitles?}
    G --&amp;gt;|yes| G1[srt to ass, then -vf subtitles]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start    fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef action    fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a

    class A start
    class B,C,D,E,F,G decision
    class B1,C1,D1,E1,F1,G1 action&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Every branch of that tree is the same seven-stage pipeline, just with a different subset of stages actually doing work.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fut0mbl3totsnbbvctvgm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fut0mbl3totsnbbvctvgm.png" alt="Gru's Plan meme about -c copy being the twist that actually works in your favor" width="360" height="230"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  The rest of the toolbox
&lt;/h2&gt;

&lt;p&gt;Two more binaries ship alongside &lt;code&gt;ffmpeg&lt;/code&gt; and are worth knowing exist:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ffprobe&lt;/code&gt;&lt;/strong&gt; — inspects a file and dumps its metadata: codecs, resolution, duration, bitrate, stream count. It's the tool for "wait, what actually &lt;em&gt;is&lt;/em&gt; this file" before you commit to a slow encode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ffplay&lt;/code&gt;&lt;/strong&gt; — a minimal media player built on the same libraries, for when you just want to preview something from the terminal without opening a full video app.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And under all three sits a stack of libraries, &lt;code&gt;libavcodec&lt;/code&gt;, &lt;code&gt;libavformat&lt;/code&gt;, &lt;code&gt;libavfilter&lt;/code&gt;, and friends, that other software links against directly rather than shelling out to the CLI. That's the actual reason ffmpeg ended up powering Chrome's media playback and Blender's video editor: it was never really "a command line tool," it was a media engine that happened to ship a command line tool as its front door.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4kdx9v8ofg0b7xuoti6q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4kdx9v8ofg0b7xuoti6q.png" alt="This Is Fine meme about starting a real libx264 re-encode and watching your CPU catch fire" width="360" height="349"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Why this is worth knowing
&lt;/h2&gt;

&lt;p&gt;None of this is trivia for its own sake.&lt;br&gt;
Once the pipeline is in your head, ffmpeg's entire flag surface stops being a wall of options to memorize.&lt;/p&gt;

&lt;p&gt;You start asking "which stage am I actually touching" before you type a command, and the answer tells you whether it'll take a second or a coffee break, and whether you even need to.&lt;/p&gt;

&lt;p&gt;That's the whole trick.&lt;br&gt;
The tool has a hundred flags, but only one shape underneath all of them.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ffmpeg</category>
      <category>cli</category>
      <category>video</category>
      <category>linux</category>
    </item>
    <item>
      <title>Seven Patterns That Decide If Your AI App Survives 10,000 Users</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Sat, 12 Sep 2026 20:30:43 +0000</pubDate>
      <link>https://dev.to/lovestaco/seven-patterns-that-decide-if-your-ai-app-survives-10000-users-2e0b</link>
      <guid>https://dev.to/lovestaco/seven-patterns-that-decide-if-your-ai-app-survives-10000-users-2e0b</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here's a fun fact that nobody puts on the slide at the AI meetup.&lt;/p&gt;

&lt;p&gt;Your agent can be perfect. The prompt can be tuned within an inch of its life. The eval score can be a smug 94%.&lt;/p&gt;

&lt;p&gt;None of that matters the moment your product gets popular.&lt;/p&gt;

&lt;p&gt;One user hits your FastAPI endpoint, the agent runs, the model answers, everybody's happy.&lt;/p&gt;

&lt;p&gt;Ten thousand users hit it in the same five minutes, and suddenly the model is the &lt;em&gt;least&lt;/em&gt; interesting part of your outage.&lt;/p&gt;

&lt;p&gt;The bottleneck moved. It always moves. It moved from "can the model answer this" to "can the system around the model survive being asked."&lt;/p&gt;

&lt;p&gt;That system is what this post is about. Seven patterns, one request, and the boring infrastructure work that decides whether your AI product works for a demo or for a Tuesday afternoon in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Zoom out one level
&lt;/h2&gt;

&lt;p&gt;Picture the AI service you already built: FastAPI route, an agent workflow, a vector store, tracing, evals, the whole thing. &lt;/p&gt;

&lt;p&gt;That entire box becomes one small rectangle in a bigger picture starting today.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6v0oezm6jwh4rnrguqhz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6v0oezm6jwh4rnrguqhz.png" alt="Diagram: one request flowing through API gateway, rate limiter, load balancer, cache, queue and workers, and circuit breaker, with the autoscaler watching all of it, before reaching the same core AI service" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nothing inside that rectangle changes. FastAPI still handles the route, the agent still calls the model, tracing still tells you what happened.&lt;/p&gt;

&lt;p&gt;What's new is everything &lt;em&gt;around&lt;/em&gt; it, deciding who gets in, what waits, what fails safely, and how much capacity exists in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pattern 1: the API gateway is your reception desk
&lt;/h2&gt;

&lt;p&gt;An API gateway sits in front of your application and handles traffic before it ever reaches your route handlers.&lt;/p&gt;

&lt;p&gt;It checks who's calling, rejects garbage requests before they cost you compute, tags each one with a request ID, and routes it to the right internal service.&lt;/p&gt;

&lt;p&gt;Think of it as reception for a big office building. One door in, one person checking where you need to go, and you never have to know which floor anything is actually on.&lt;/p&gt;

&lt;p&gt;One distinction that trips people up: this is not the same as a &lt;em&gt;model&lt;/em&gt; gateway. An API gateway manages traffic coming &lt;strong&gt;into&lt;/strong&gt; your app from users. A model gateway manages calls going &lt;strong&gt;out&lt;/strong&gt; to OpenAI, Anthropic, or whoever's serving your model. Same word, opposite direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pattern 2: rate limiting decides how much work gets to start
&lt;/h2&gt;

&lt;p&gt;A gateway controlling who can knock on the door still lets in more valid traffic than you can serve. The next question isn't "who's allowed in," it's "how much work do we let start."&lt;/p&gt;

&lt;p&gt;Plain rate limiting counts requests per user per minute. Fine for a CRUD API. Not fine for an LLM app, where request count barely correlates with cost.&lt;/p&gt;

&lt;p&gt;One call classifies "yes" or "no" in ten tokens. Another retrieves eight documents, calls three tools, and streams back four paragraphs. Same "one request" on your dashboard, wildly different bill.&lt;/p&gt;

&lt;p&gt;So AI systems widen the definition to &lt;strong&gt;admission control&lt;/strong&gt;: cap requests, cap input tokens, cap output tokens, cap concurrent model calls, whatever actually predicts cost in your system.&lt;/p&gt;

&lt;p&gt;When you're over budget, the honest answer is an HTTP 429 with a &lt;code&gt;Retry-After&lt;/code&gt; header, &lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/429" rel="noopener noreferrer"&gt;the standard response for telling a client to back off&lt;/a&gt; instead of quietly queuing infinite work. That's called back pressure: push the slowdown back to whoever's asking, before it becomes your problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pattern 3: cache what's actually safe to reuse
&lt;/h2&gt;

&lt;p&gt;A gateway and a rate limiter, and you've still got a pile of accepted requests all asking suspiciously similar questions.&lt;/p&gt;

&lt;p&gt;If five hundred people ask about the same public doc today, the first request should do the retrieval and embedding work, and the other 499 should get it basically for free.&lt;/p&gt;

&lt;p&gt;Redis in front of embeddings, retrieval results, or common model responses buys you real latency and cost wins.&lt;/p&gt;

&lt;p&gt;The catch, and it's a big one for anything touching an LLM: caching an answer is a claim that it's still correct and still safe for whoever's asking.&lt;/p&gt;

&lt;p&gt;A private answer generated for one user is not a cache entry, it's a leak, if it ever gets served to somebody else. A cached fact from three months ago about pricing or an API limit is not "reused work," it's a wrong answer with good latency.&lt;/p&gt;

&lt;p&gt;Every cache entry needs an expiry, a scope (user, tenant, permission level), and an invalidation path tied to the source changing. Skip any of those three and the cache stops saving you money and starts costing you incidents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pattern 4: let the work that can wait, wait
&lt;/h2&gt;

&lt;p&gt;Caching removes duplicate work. It doesn't remove the &lt;em&gt;different&lt;/em&gt; work, and there's a lot of AI-adjacent work that genuinely doesn't need to finish before you respond to the user: document ingestion, embedding generation on upload, a long report, an email.&lt;/p&gt;

&lt;p&gt;That's what a durable queue is for. Take the order, hand the customer a ticket number, let the kitchen cook at the kitchen's pace instead of the counter's.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcey4n2ry0qu0mr9yy41d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcey4n2ry0qu0mr9yy41d.png" alt="Diagram: background jobs flowing into a bounded durable queue, a worker pool draining it at a controlled rate, a dead letter queue catching the jobs that keep failing, with a note about idempotency and separate lanes per job type" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A few details that separate "we added a queue" from "we added a queue that actually holds":&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bound the backlog.&lt;/strong&gt; An unbounded queue doesn't prevent an outage, it just delays it and makes it bigger when it finally lands.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Route failures somewhere visible.&lt;/strong&gt; A job that keeps failing should land in a &lt;a href="https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html" rel="noopener noreferrer"&gt;dead letter queue&lt;/a&gt; for a human to look at, not retry silently forever and quietly eat your worker capacity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make retried actions idempotent.&lt;/strong&gt; At-least-once delivery means a worker can see the same job twice. If "process this job" means "charge this card" or "send this email," processing it twice is a bug wearing a queue's clothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Separate lanes for separate work.&lt;/strong&gt; One giant document-ingestion job parked in the same queue as your live chat's background tasks will happily starve the thing your paying customers are waiting on.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;process_job&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setnx&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;processed:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expire&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;processed:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;86400&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;do_the_actual_work&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# already processed once, safely a no-op the second time
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  Pattern 5: stop one broken dependency from taking the site with it
&lt;/h2&gt;

&lt;p&gt;Now the harder failure mode. Not "the dependency is down," which fails fast and cleanly, but "the dependency is &lt;em&gt;slow&lt;/em&gt;," which is much worse.&lt;/p&gt;

&lt;p&gt;A dependency that's fully down refuses your call in milliseconds. Your worker's free again instantly, checkout fails cleanly, the rest of the site keeps serving.&lt;/p&gt;

&lt;p&gt;A dependency that's slow holds that same worker for thirty seconds while it decides whether to answer. Do that across your whole pool and a random product page dies for a problem it never even called.&lt;/p&gt;

&lt;p&gt;Three tools handle this, in order:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Timeouts.&lt;/strong&gt; Decide up front how long you're willing to wait. No timeout means every slow dependency gets to set your latency for you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retries, with backoff and jitter.&lt;/strong&gt; A brief retry after a random-ish delay handles transient blips. &lt;a href="https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/" rel="noopener noreferrer"&gt;AWS has the definitive writeup on why the jitter matters&lt;/a&gt;: without it, every client backs off on the same clock and retries in a synchronized wave that looks exactly like the outage you were trying to survive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The circuit breaker&lt;/strong&gt;, for when it's not transient. &lt;a href="https://martinfowler.com/bliki/CircuitBreaker.html" rel="noopener noreferrer"&gt;Martin Fowler's original writeup&lt;/a&gt; frames it exactly like the breaker in your fuse box: too many failures and the circuit trips open, calls stop reaching the broken dependency, and everyone's spared the wait.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftilpxnya8tibkn30ckns.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftilpxnya8tibkn30ckns.png" alt="Circuit breaker state diagram: closed state where calls flow normally transitioning to open when failures cross a threshold, open transitioning to half-open after a cooldown timer, half-open transitioning back to closed on success or back to open on failure" width="663" height="547"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Closed is normal operation. Open is refusing everything instantly while the failing dependency recovers off the hook. Half-open is the cautious part: let a handful of test calls through, and only reopen the floodgates if they succeed.&lt;/p&gt;

&lt;p&gt;When the circuit's open, don't fill the gap with a guess. Serve a tested fallback model, a verified cached answer, or an honest "this feature's briefly unavailable." &lt;a href="https://learn.microsoft.com/en-us/azure/well-architected/reliability/handle-transient-faults" rel="noopener noreferrer"&gt;Graceful degradation&lt;/a&gt; means reduced functionality, never a confident wrong answer.&lt;/p&gt;

&lt;p&gt;And keep failing dependencies from sharing a resource pool with healthy ones — separate connection pools per dependency, a pattern usually called a &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/patterns/bulkhead" rel="noopener noreferrer"&gt;bulkhead&lt;/a&gt;, named for the same reason ships have them: one flooded compartment shouldn't sink the whole hull.&lt;/p&gt;
&lt;h2&gt;
  
  
  Pattern 6: spread the load across more than one copy
&lt;/h2&gt;

&lt;p&gt;Every one of those live requests is still landing on a single instance of your core AI service, and one healthy server still has a ceiling.&lt;/p&gt;

&lt;p&gt;Run several copies, put a load balancer in front, and it hands each request to whichever healthy instance has room.&lt;/p&gt;

&lt;p&gt;Worth separating from the gateway, since people conflate them: the &lt;strong&gt;gateway&lt;/strong&gt; decides which service a request should reach. The &lt;strong&gt;load balancer&lt;/strong&gt; decides which &lt;em&gt;copy&lt;/em&gt; of that service handles it. Some managed products bundle both jobs into one offering, but they're still two different decisions.&lt;/p&gt;

&lt;p&gt;The thing that makes running several copies safe is keeping them stateless. Durable conversation history, workflow checkpoints, job status — put those in Postgres. Short-lived shared state goes in something like Redis. Then any instance can pick up any request, and a retried call resumes from a saved checkpoint instead of restarting the whole agent loop from scratch.&lt;/p&gt;
&lt;h2&gt;
  
  
  Pattern 7: autoscaling adds the capacity the rest bought you room to add
&lt;/h2&gt;

&lt;p&gt;Everything above buys time. Eventually, if traffic keeps growing, distributing the load across existing copies stops being enough, and you actually need more of them.&lt;/p&gt;

&lt;p&gt;Autoscaling adds or removes instances as demand shifts, so you're not paying for a night's worth of GPU capacity at 3am for zero users.&lt;/p&gt;

&lt;p&gt;The part worth getting right is &lt;em&gt;what signal drives it&lt;/em&gt;. CPU alone is a bad proxy for a GPU-backed model server, since you can be maxed out on the accelerator while CPU idles. Better signals: queue depth, how long the oldest task's been waiting, active requests, or GPU utilization directly.&lt;/p&gt;

&lt;p&gt;New capacity also isn't instant. A stateless API pod might come up in seconds. A model server loading weights into GPU memory can take minutes. If your traffic pattern is predictable (the daily 9am spike, say), keep some capacity warm ahead of it rather than reacting cold.&lt;/p&gt;

&lt;p&gt;This is also why autoscaling comes &lt;em&gt;last&lt;/em&gt; in the list, not first. More servers don't fix a cache serving stale answers, a retry storm with no jitter, or a queue with no bound. Control the work first. Add capacity second.&lt;/p&gt;
&lt;h2&gt;
  
  
  Walking one request through the whole thing
&lt;/h2&gt;

&lt;p&gt;A user's request hits the API gateway, which gives it one controlled entrance and routes it toward the right service.&lt;/p&gt;

&lt;p&gt;The rate limiter checks whether this user, and this kind of request, is within budget.&lt;/p&gt;

&lt;p&gt;The load balancer picks a healthy instance of the core AI service.&lt;/p&gt;

&lt;p&gt;That instance checks the cache first: if the answer is safe to reuse, it comes back immediately.&lt;/p&gt;

&lt;p&gt;If the request needs an answer now, it stays on the live path. If it's kicking off longer background work, it drops into the queue for a worker to pick up.&lt;/p&gt;

&lt;p&gt;Timeouts and circuit breakers guard every call the instance makes outward, to the model provider, the vector store, any tool. The autoscaler is watching the whole system and adjusting instance count as it goes.&lt;/p&gt;

&lt;p&gt;And inside that one small rectangle, nothing about the actual AI work changed: FastAPI handles the route, the agent runs its loop, the model generates the answer, tracing records what happened, evals score whether it was any good.&lt;/p&gt;

&lt;p&gt;That last point is the one worth sitting with. Metrics tell you if the system's fast. Traces tell you where a request spent its time. Evals tell you if the answer was actually worth sending.&lt;/p&gt;

&lt;p&gt;None of these seven patterns touch that last question, and they were never meant to. Scaling a wrong answer faster is not a win, it's just a faster wrong answer reaching more people.&lt;/p&gt;

&lt;p&gt;If you ever get a system design question about this in an interview, resist the urge to start naming seven cloud products. Start with the request: how much traffic, what needs to answer immediately versus what can wait, and which dependency is going to be the first one to have a bad day. Add a pattern only when it's solving a problem you can point to.&lt;/p&gt;

&lt;p&gt;The model was never the bug. The system that decides who gets to ask it something, and what happens while it's thinking, always was.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>systemdesign</category>
      <category>backend</category>
      <category>scaling</category>
    </item>
    <item>
      <title>How Uber Knows Your Driver Is 7 Minutes Away</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Fri, 11 Sep 2026 20:24:22 +0000</pubDate>
      <link>https://dev.to/lovestaco/how-uber-knows-your-driver-is-7-minutes-away-ao3</link>
      <guid>https://dev.to/lovestaco/how-uber-knows-your-driver-is-7-minutes-away-ao3</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Your Uber says the driver is 7 minutes away.&lt;/p&gt;

&lt;p&gt;They show up in 7 minutes.&lt;/p&gt;

&lt;p&gt;That is not a lucky guess, that is one of the more quietly insane systems in consumer tech, and it is worth taking apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: chop the world into segments
&lt;/h2&gt;

&lt;p&gt;Uber does not think about roads, it thinks about road &lt;em&gt;segments&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;A single road gets cut into a handful of pieces, and globally Uber is tracking around 100 million of these segments.&lt;/p&gt;

&lt;p&gt;Each segment has a number attached to it: how long it takes to cross, right now.&lt;/p&gt;

&lt;p&gt;To get from your driver to you, the routing engine finds the fastest path through these segments and adds up the crossing times along the way.&lt;/p&gt;

&lt;p&gt;That sum is your ETA.&lt;/p&gt;

&lt;p&gt;Simple enough, except that number, "how long it takes to cross this segment", is not a constant.&lt;/p&gt;

&lt;p&gt;It is 20 seconds at 2am and 2 minutes at 6pm on a Friday, on the exact same 200 meters of road.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1a3ceruocx7khedfbprv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1a3ceruocx7khedfbprv.png" alt="Diagram: a road split into four numbered segments, each with a crossing-time label that changes between a 2am value and a 6pm Friday value" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: don't ask Google, ask your drivers
&lt;/h2&gt;

&lt;p&gt;The obvious move is to buy this traffic data from someone who already maps the whole planet.&lt;/p&gt;

&lt;p&gt;Uber doesn't. Uber measures it.&lt;/p&gt;

&lt;p&gt;Every driver on the platform is already pinging their location every 4 seconds, because that's what the app needs to do anyway.&lt;/p&gt;

&lt;p&gt;That stream is basically free traffic telemetry. Every time a driver crosses a segment, Uber now knows, to the second, how long that segment just took.&lt;/p&gt;

&lt;p&gt;Multiply that by every active driver and you get a live traffic sensor network that nobody had to install a single camera for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: the part that actually needs machine learning
&lt;/h2&gt;

&lt;p&gt;Here's the catch. Live measurements only tell you what a segment did in the past few minutes.&lt;/p&gt;

&lt;p&gt;Your ETA needs to know what it's going to do while your driver is still en route to you, which could be 15 minutes from now.&lt;/p&gt;

&lt;p&gt;So in 2022, Uber shipped a deep learning system called &lt;strong&gt;DeepETA&lt;/strong&gt;, and it does not just average recent history.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If a normally fast road has a fresh accident on it, DeepETA trusts the live signal over the historical pattern.&lt;/li&gt;
&lt;li&gt;If a quiet backstreet has no recent driver on it at all, DeepETA infers its state from the busier roads around it instead of guessing blind.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It refreshes these forecasts for every segment, for the next 3 hours, every few minutes, and answers roughly 2 million forecast requests a second, making it one of the busiest models running inside Uber. (&lt;a href="https://www.uber.com/us/en/blog/deepeta-how-uber-predicts-arrival-times/" rel="noopener noreferrer"&gt;Uber Engineering: DeepETA&lt;/a&gt;)&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Driver GPS pings&amp;lt;br/&amp;gt;every 4s] --&amp;gt; B[Live segment&amp;lt;br/&amp;gt;crossing times]
    B --&amp;gt; C[DeepETA&amp;lt;br/&amp;gt;forecasts 3h ahead]
    C --&amp;gt; D[Routing engine sums&amp;lt;br/&amp;gt;segments on your path]
    D --&amp;gt; E[Correction model&amp;lt;br/&amp;gt;trained on real trips]
    E --&amp;gt; F[The ETA on&amp;lt;br/&amp;gt;your screen]

    classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef chip   fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef accel  fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a

    class A start
    class B,D chip
    class C,E accel
    class F start&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;And there's one more pass after all that. The routing engine's segment-summed number goes through a second model, trained on millions of completed real trips, whose only job is to catch and correct the places where the physics-based sum tends to be systematically wrong (a stop sign nobody accounted for, a left turn that always takes longer than it looks).&lt;/p&gt;

&lt;h2&gt;
  
  
  The payoff
&lt;/h2&gt;

&lt;p&gt;Shipping DeepETA improved long-trip arrival accuracy by 6%.&lt;/p&gt;

&lt;p&gt;Uber estimates that alone is worth around $100 million a year in gross bookings, because an ETA people trust is an ETA people don't cancel on.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4qfctr1gwoom332ov6zo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4qfctr1gwoom332ov6zo.png" alt="Meme: a guy at a debate table with a sign reading " width="360" height="286"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnvhow7ox1hu03ey91d8n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnvhow7ox1hu03ey91d8n.png" alt="Meme: expanding brain meme escalating from cameras and satellites to just reading the GPS pings drivers already send every 4 seconds" width="360" height="504"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Next time your ETA ticks down without drama, that's 100 million road segments, a live sensor network made of other people's cars, and a model answering 2 million questions a second so a number on your screen can be boring.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;

  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The exact math, not a black box&lt;/th&gt;
&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;
&lt;th&gt;Every factor that feeds the score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

How does Blast Radius scoring work? (a more technical explanation)
&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;
&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>machinelearning</category>
      <category>backend</category>
      <category>uber</category>
    </item>
    <item>
      <title>Down Is Kind, Slow Is Fatal: Circuit Breakers and the Three Links Around Them</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Thu, 10 Sep 2026 20:15:05 +0000</pubDate>
      <link>https://dev.to/lovestaco/down-is-kind-slow-is-fatal-circuit-breakers-and-the-three-links-around-them-3acn</link>
      <guid>https://dev.to/lovestaco/down-is-kind-slow-is-fatal-circuit-breakers-and-the-three-links-around-them-3acn</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Picture an e-commerce store.&lt;/p&gt;

&lt;p&gt;Checkout calls the payment service on every order. One call, one arrow on the architecture diagram, the least interesting line in the codebase.&lt;/p&gt;

&lt;p&gt;Now break it in two different ways and watch which one hurts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Take the payment service completely down.&lt;/strong&gt; Every call to it is refused in about three milliseconds. Checkout returns a clean failure for that order, and every other page on the site keeps serving.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Now leave it running and answering every call correctly, only slower.&lt;/strong&gt; 40 milliseconds becomes 30 seconds.&lt;/p&gt;

&lt;p&gt;Checkout has nothing to react to, because from its point of view nothing failed.&lt;/p&gt;

&lt;p&gt;On paper the second case is the healthier one. Every call is answered, correctly, with the right data.&lt;/p&gt;

&lt;p&gt;It is also the one that takes the whole site down.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5s7u3mdinqieq8bghf3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5s7u3mdinqieq8bghf3.png" alt="Diagram comparing a payment service that is down, where workers are held for 3ms and the site keeps serving, against one that is slow, where every worker is held for 30 seconds and the product page dies" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The reason is not subtle once you see it. Your service holds a worker for the entire length of every call it makes.&lt;/p&gt;

&lt;p&gt;When the dependency is down, that worker comes back in milliseconds and moves straight on to the next request.&lt;/p&gt;

&lt;p&gt;A slow dependency keeps the same worker for 30 seconds, and you only have so many of them.&lt;/p&gt;

&lt;p&gt;So this whole article is about a pattern that protects &lt;strong&gt;you&lt;/strong&gt;, not the service you are calling.&lt;/p&gt;

&lt;p&gt;It works by stopping your own service from waiting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bill for a slow call, itemised
&lt;/h2&gt;

&lt;p&gt;Every call your service makes takes a small pile of things with it while it runs.&lt;/p&gt;

&lt;p&gt;A thread. A database connection. Some memory. A port.&lt;/p&gt;

&lt;p&gt;It holds all of that until the call finishes, one way or the other. A slow call just holds it for longer.&lt;/p&gt;

&lt;p&gt;You have a fixed number of each, and nothing new can start once they are all in use.&lt;/p&gt;

&lt;p&gt;That is why the product page died in the opening scene, even though the product page never calls payment. The blocked requests were sitting on threads and connections that the rest of your system shares.&lt;/p&gt;

&lt;p&gt;One slow dependency can saturate every resource on every one of your services in seconds.&lt;/p&gt;

&lt;p&gt;And this is not a rare situation, because dependencies multiply. Say your service calls 30 of them, and each one is up 99.99% of the time. You need all 30.&lt;/p&gt;

&lt;p&gt;Multiply those together and you land at about 99.7%.&lt;/p&gt;

&lt;p&gt;That is 3 million failed requests in every billion, and over two hours of downtime a month, with all 30 dependencies behaving exactly as advertised.&lt;/p&gt;

&lt;p&gt;The fix is to &lt;strong&gt;fail fast&lt;/strong&gt;. Turn down a call you already expect to fail, instead of waiting for it to time out.&lt;/p&gt;

&lt;p&gt;But nothing frees a slot until a call ends. So something has to decide when a call has taken too long.&lt;/p&gt;

&lt;h2&gt;
  
  
  Link one: the timeout
&lt;/h2&gt;

&lt;p&gt;That decision is the timeout, and it is the one rule here with no exception.&lt;/p&gt;

&lt;p&gt;Set a timeout on every remote call. Honestly, on any call that leaves your process, even one going to something on the same machine.&lt;/p&gt;

&lt;p&gt;And there are two of them, not one. The &lt;strong&gt;connection timeout&lt;/strong&gt; covers getting the socket open. The &lt;strong&gt;request timeout&lt;/strong&gt; covers waiting for the answer.&lt;/p&gt;

&lt;p&gt;Picking the number is not a vibe. &lt;a href="https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/" rel="noopener noreferrer"&gt;AWS's Builders' Library has the method&lt;/a&gt;: decide what rate of false timeouts you can live with, meaning calls you cut off that would have succeeded, then set the timeout at the matching latency percentile of the service you are calling.&lt;/p&gt;

&lt;p&gt;Their worked example accepts one call in a thousand being cut off wrongly, so they use the P99.9 latency.&lt;/p&gt;

&lt;p&gt;Set it too high and it barely counts as a timeout at all. You are still holding the thread, the connection and the memory for the entire time you wait for a number in a config file to tell you that you are protected.&lt;/p&gt;

&lt;p&gt;The percentile method has two places it breaks down.&lt;/p&gt;

&lt;p&gt;It struggles with clients that carry real network latency, like calls coming in over the internet.&lt;/p&gt;

&lt;p&gt;And it struggles with services where P99.9 sits close to P50, because then the timeout lands almost on top of the normal case.&lt;/p&gt;

&lt;p&gt;Both of those need padding on top of the measured number.&lt;/p&gt;

&lt;p&gt;There is one case a timeout alone does not cover, and it is the sneaky one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A call that is slow and still succeeds never counts as a failure.&lt;/strong&gt; It never moves a failure counter anywhere, so nothing downstream ever learns that the dependency has gone bad.&lt;/p&gt;

&lt;p&gt;That is why a breaker needs a separate slow call rule. &lt;a href="https://resilience4j.readme.io/docs/circuitbreaker" rel="noopener noreferrer"&gt;Resilience4j ships one&lt;/a&gt;: a &lt;code&gt;slowCallDurationThreshold&lt;/code&gt; and a &lt;code&gt;slowCallRateThreshold&lt;/code&gt;, where a call only counts toward opening the breaker once it crosses that duration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Link two: retries, the powerful medicine
&lt;/h2&gt;

&lt;p&gt;A timeout turns a hang into an error. And the first thing anyone does with an error is try it again.&lt;/p&gt;

&lt;p&gt;That second attempt costs the server another slice of its time, and it buys that time for exactly one client.&lt;/p&gt;

&lt;p&gt;AWS has a blunt word for this: retries are &lt;strong&gt;selfish&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That trade is a good one when the failure was a one-off, because the server had time to spare anyway.&lt;/p&gt;

&lt;p&gt;It is a terrible one when the failure came from overload in the first place. You are adding load to something that is already buried, which makes it worse and holds it down long after the original problem is gone.&lt;/p&gt;

&lt;p&gt;AWS calls retries a powerful medicine, which is a good way to hold it in your head. The right dose helps. The same substance in a larger dose does the real damage.&lt;/p&gt;

&lt;p&gt;Backing off between attempts is the first fix, and on its own it does not go far enough.&lt;/p&gt;

&lt;p&gt;All the calls that failed, failed at the same moment, because the dependency went bad at one instant. Back every one of them off by the same amount and they come back together too, overloading the thing a second time.&lt;/p&gt;

&lt;p&gt;What breaks the pattern is &lt;strong&gt;jitter&lt;/strong&gt;, a random delay on top of the backoff, so the attempts spread out instead of arriving in a block.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkts2wjszuuh3tnorpho0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkts2wjszuuh3tnorpho0.png" alt="American Chopper argument meme about retrying a failed call into an overloaded dependency and discovering that equal backoff makes every client return at the same moment" width="360" height="299"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Notice what has happened to the original cause here.&lt;/p&gt;

&lt;p&gt;Whatever knocked the dependency over is gone, and the system is still down, held there by its own traffic.&lt;/p&gt;

&lt;p&gt;Researchers call this a &lt;a href="https://sigops.org/s/conferences/hotos/2021/papers/hotos21-s11-bronson.pdf" rel="noopener noreferrer"&gt;metastable failure&lt;/a&gt;. A trigger pushes the system into a bad state, the bad state feeds itself, and removing the trigger changes nothing.&lt;/p&gt;

&lt;p&gt;Retries make you more vulnerable to it, because they turn a small outage into an internal storm.&lt;/p&gt;

&lt;p&gt;The part worth remembering is what this looks like from inside your own monitoring.&lt;/p&gt;

&lt;p&gt;Clients give up waiting before the answer arrives, then send the request again.&lt;/p&gt;

&lt;p&gt;Meanwhile your system is still finishing old calls that nobody is listening for any more.&lt;/p&gt;

&lt;p&gt;Your &lt;strong&gt;throughput&lt;/strong&gt; graph looks excellent, because work genuinely is being completed.&lt;/p&gt;

&lt;p&gt;The number that matters is &lt;strong&gt;goodput&lt;/strong&gt;, the work that reaches a caller who is still waiting, and that has gone to zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  Link three: the breaker itself
&lt;/h2&gt;

&lt;p&gt;Backoff and jitter slow the storm down, and they are worth having.&lt;/p&gt;

&lt;p&gt;But nothing we have put on that wire so far actually stops the call.&lt;/p&gt;

&lt;p&gt;That is the breaker's job, and the pattern is not new. It comes from Michael Nygard's &lt;em&gt;Release It!&lt;/em&gt;, popularised precisely to prevent the cascade we have spent three sections building.&lt;/p&gt;

&lt;p&gt;Everyone can draw the three states. Far fewer can say what number theirs trips at, or whether it can trip at all.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;stateDiagram-v2
    [*] --&amp;gt; Closed
    Closed --&amp;gt; Open: failure rate crosses the threshold&amp;lt;br/&amp;gt;over a real number of calls
    Open --&amp;gt; HalfOpen: wait duration has elapsed&amp;lt;br/&amp;gt;AND a call actually arrives
    HalfOpen --&amp;gt; Closed: the permitted probe calls succeed
    HalfOpen --&amp;gt; Open: ONE probe fails, timer restarts
    Closed: Closed&amp;lt;br/&amp;gt;calls pass through, failures counted
    Open: Open&amp;lt;br/&amp;gt;call not attempted, fails in microseconds
    HalfOpen: Half-Open&amp;lt;br/&amp;gt;a few real calls allowed through as probes&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;&lt;strong&gt;Closed&lt;/strong&gt; is normal, and where the breaker spends almost all of its life. Calls pass straight through, failures are counted as they happen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open&lt;/strong&gt; means it has tripped. The call is not attempted at all, it fails on the spot, and your application gets an exception back in microseconds instead of waiting 30 seconds. That is the fail fast idea from earlier, turned into something you can point at.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Half-open&lt;/strong&gt; is the trial run, and it is how the breaker finds its way back. Once the reset timeout expires, it lets a limited number of real requests through to see whether the problem is fixed.&lt;/p&gt;

&lt;p&gt;People mix this up with the retry pattern constantly, and the two are doing opposite jobs.&lt;/p&gt;

&lt;p&gt;A retry assumes the call will succeed eventually, so it keeps trying.&lt;/p&gt;

&lt;p&gt;A breaker assumes it will not, and stops you making a call that is likely to fail.&lt;/p&gt;

&lt;p&gt;Three states is the teaching model. Resilience4j actually exposes six, adding &lt;code&gt;METRICS_ONLY&lt;/code&gt;, &lt;code&gt;DISABLED&lt;/code&gt; and &lt;code&gt;FORCED_OPEN&lt;/code&gt;, which a person puts it into rather than a failure rate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually pulls the lever
&lt;/h2&gt;

&lt;p&gt;Ask most engineers what trips a breaker and they will say a count. Something like five failures in a row.&lt;/p&gt;

&lt;p&gt;The libraries you are most likely to have installed do not work that way at all.&lt;/p&gt;

&lt;p&gt;They measure a &lt;strong&gt;failure rate&lt;/strong&gt;, and when that rate reaches your threshold, the breaker opens.&lt;/p&gt;

&lt;p&gt;That rate is measured over a sliding window, and the window comes in two shapes. A &lt;strong&gt;count-based&lt;/strong&gt; window covers the last N calls, whatever length of time those took. A &lt;strong&gt;time-based&lt;/strong&gt; window covers the calls from the last N seconds, however many arrived.&lt;/p&gt;

&lt;p&gt;A rate over a handful of calls is meaningless, because two failures out of three is 67% and also nothing.&lt;/p&gt;

&lt;p&gt;So the rate is gated. It can only be calculated once a minimum number of calls has been recorded.&lt;/p&gt;

&lt;p&gt;Hystrix called that gate the &lt;strong&gt;request volume threshold&lt;/strong&gt;: the minimum number of requests that have to arrive in the rolling window before the circuit is allowed to trip at all.&lt;/p&gt;

&lt;p&gt;Set it to 20. Nineteen requests come in over 10 seconds and every one of them fails. The circuit does not trip, because 19 is short of the 20 it was told to wait for.&lt;/p&gt;

&lt;p&gt;Counting breakers do still exist. &lt;a href="https://github.com/sony/gobreaker" rel="noopener noreferrer"&gt;gobreaker&lt;/a&gt; trips by default once consecutive failures go above five. And a counter in the closed state does not pile up forever either. &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/patterns/circuit-breaker" rel="noopener noreferrer"&gt;Azure's guidance&lt;/a&gt; is that the closed-state failure count is time based and resets at intervals, so the odd failure scattered across a quiet afternoon never adds up into a trip.&lt;/p&gt;

&lt;p&gt;So four separate settings decide whether your breaker ever opens: the threshold, the shape of the window, the size of the window, and the minimum call volume.&lt;/p&gt;

&lt;p&gt;Every one of those is a value somebody chose for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The defaults do not agree with each other
&lt;/h2&gt;

&lt;p&gt;Start with the one most tutorials still reach for.&lt;/p&gt;

&lt;p&gt;Hystrix is no longer in active development. It is in maintenance mode, which in Netflix's own words means they will not review issues, merge pull requests or release new versions. The last release was 1.5.18.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Netflix/Hystrix#hystrix-status" rel="noopener noreferrer"&gt;Netflix's own recommendation&lt;/a&gt; for new work is to use an active project instead, and they name Resilience4j. Internally they moved toward adaptive implementations that react to real-time performance rather than to thresholds you pick in advance.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2md3xo7dvpkq9pi0hkw7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2md3xo7dvpkq9pi0hkw7.png" alt="Imperial officer meme: this tutorial uses Hystrix, an older code sir, but it has been in maintenance mode since 2018" width="360" height="194"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here are four sets of stock defaults, read the same way each time. What trips it, how long it stays open, how many calls it lets back through.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd9plq97wyb36nwinhe7e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd9plq97wyb36nwinhe7e.png" alt="Diagram table comparing the stock circuit breaker defaults of Resilience4j, Hystrix, Polly and gobreaker across trip condition, open duration and half-open probe count" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The disagreement is material rather than cosmetic.&lt;/p&gt;

&lt;p&gt;The trip threshold runs from 10% in Polly, up to 50% in Hystrix and Resilience4j, over to six consecutive failures in gobreaker.&lt;/p&gt;

&lt;p&gt;Time in the open state runs from 5 seconds at one end to 60 at the other.&lt;/p&gt;

&lt;p&gt;And the half-open probe count runs from 1 to 10, which is a tenfold difference in how hard each library leans on a service that is still recovering.&lt;/p&gt;

&lt;p&gt;Two of those defaults quietly switch the breaker off, and this is the part nobody checks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Resilience4j needs 100 recorded calls before it evaluates a rate at all&lt;/strong&gt;, measured over a window of the last 100 calls. Put that on a low traffic internal endpoint and it may never accumulate enough calls to evaluate anything. That breaker simply never opens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Resilience4j also does not move from open to half-open on a timer.&lt;/strong&gt; The transition happens when a call arrives after the wait duration has elapsed, not at the moment it elapses. On an endpoint that has gone quiet, the breaker sits open until something finally asks.&lt;/p&gt;

&lt;p&gt;Reading your own config out loud is a genuinely useful five minutes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;resilience4j&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;circuitbreaker&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;instances&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;payment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;slidingWindowType&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;COUNT_BASED&lt;/span&gt;
        &lt;span class="na"&gt;slidingWindowSize&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;          &lt;span class="c1"&gt;# the last 100 calls&lt;/span&gt;
        &lt;span class="na"&gt;minimumNumberOfCalls&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;       &lt;span class="c1"&gt;# &amp;lt;- can your endpoint even reach this?&lt;/span&gt;
        &lt;span class="na"&gt;failureRateThreshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;50&lt;/span&gt;        &lt;span class="c1"&gt;# percent, not a count&lt;/span&gt;
        &lt;span class="na"&gt;slowCallDurationThreshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;2s&lt;/span&gt;   &lt;span class="c1"&gt;# slow successes count as failures&lt;/span&gt;
        &lt;span class="na"&gt;slowCallRateThreshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;      &lt;span class="c1"&gt;# percent&lt;/span&gt;
        &lt;span class="na"&gt;waitDurationInOpenState&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;60s&lt;/span&gt;
        &lt;span class="na"&gt;permittedNumberOfCallsInHalfOpenState&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;   &lt;span class="c1"&gt;# &amp;lt;- the recovery dial&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  Half-open is where systems break a second time
&lt;/h2&gt;

&lt;p&gt;Half-open exists for one reason. A service that has just come back is not at full strength, and if you hand it everything at once it falls over again.&lt;/p&gt;

&lt;p&gt;So the breaker does not go straight from open back to closed. It opens a gate only a few calls fit through, and &lt;strong&gt;you&lt;/strong&gt; choose how many.&lt;/p&gt;

&lt;p&gt;Resilience4j describes those calls as &lt;em&gt;permitted&lt;/em&gt; calls, and their job is to find out whether the backend is still unavailable or has become available again.&lt;/p&gt;

&lt;p&gt;Watch what happens if one of them comes back red.&lt;/p&gt;

&lt;p&gt;If any probe fails, the breaker assumes the fault is still there, drops straight back to open, and restarts its timer.&lt;/p&gt;

&lt;p&gt;A single bad probe undoes the whole cooldown and puts you back at the start.&lt;/p&gt;

&lt;p&gt;There is a way to get this wrong in the other direction too. Set the cooldown too short and the breaker keeps trying before anything has changed, so it swings between open and half-open over and over. That flapping shows up in your application as worse response times.&lt;/p&gt;

&lt;p&gt;Put those together and you can see why half-open is where systems break twice. The first incident knocked the dependency down. The second one is your breaker's doing, because a dependency that is only halfway back got handed full traffic before it could carry it.&lt;/p&gt;
&lt;h2&gt;
  
  
  Link four: what do you actually return?
&lt;/h2&gt;

&lt;p&gt;Every call the breaker turns away has to return something to whoever asked.&lt;/p&gt;

&lt;p&gt;That is the fourth link, and this one is your job rather than the breaker's. The breaker only decides whether to make the call. What comes back when it says no is application logic you write.&lt;/p&gt;

&lt;p&gt;Martin Fowler's &lt;a href="https://martinfowler.com/bliki/CircuitBreaker.html" rel="noopener noreferrer"&gt;two examples&lt;/a&gt; are the ones worth keeping in your head. A credit card authorization does not have to happen right now, so it can go onto a queue and be dealt with later. And missing data can often be covered by showing something slightly out of date.&lt;/p&gt;

&lt;p&gt;Azure lists four realistic options while the breaker is open:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Degrade the feature&lt;/strong&gt;, so the page still loads without that section.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Call something else&lt;/strong&gt;, an alternative operation that does not depend on the broken thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Return a default value&lt;/strong&gt; that means something to your application.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Report the exception to the user&lt;/strong&gt;, which is what you get if you write no fallback at all, and is a real option rather than a failure to decide.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Think about what actually changed when you added the breaker. Your caller was going to get an error either way. The breaker changed &lt;strong&gt;which&lt;/strong&gt; error it is, and &lt;strong&gt;how long&lt;/strong&gt; the caller waited to receive it.&lt;/p&gt;

&lt;p&gt;Amazon's position runs the other way, and it is worth hearing.&lt;/p&gt;

&lt;p&gt;They avoid fallbacks for two reasons. The effectiveness of a fallback is hard to prove and hard to test. And a fallback is a mode a system only enters at the most chaotic possible moment, when things are already breaking, and switching modes right then increases the chaos.&lt;/p&gt;

&lt;p&gt;The mechanism behind that objection is staleness in the code path itself.&lt;/p&gt;

&lt;p&gt;Think about a fallback you wrote two years ago that almost never triggers. If there is a bug in it, or a side effect that makes the whole problem worse, nobody has looked since, and the people who wrote it have long forgotten how it worked.&lt;/p&gt;

&lt;p&gt;Their preference turns into something you can act on this week: &lt;strong&gt;favour code paths that run in production continuously over ones that run rarely.&lt;/strong&gt; If a fallback really is essential, exercise it in production as often as you can, so it behaves as predictably as the primary path.&lt;/p&gt;
&lt;h2&gt;
  
  
  The uncomfortable part: AWS is not sold on breakers
&lt;/h2&gt;

&lt;p&gt;If fallback code is that risky, it is fair to ask whether the breaker in front of it is worth having at all.&lt;/p&gt;

&lt;p&gt;AWS puts the objection in writing, in the Builders' Library, word for word:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Circuit breakers, where calls to a downstream service are stopped entirely when an error threshold is exceeded, are widely promoted to solve this problem. Unfortunately, circuit breakers introduce modal behavior into systems that can be difficult to test, and can introduce significant additional time to recovery.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There are two separate complaints packed in there.&lt;/p&gt;

&lt;p&gt;The first is &lt;strong&gt;modal behavior&lt;/strong&gt;, and it is everything we just spent this whole article building, described as a cost instead of a feature. Closed, open and half-open are three modes. Two of them are states your system is almost never in, which makes them exactly the states hardest to test.&lt;/p&gt;

&lt;p&gt;The second is &lt;strong&gt;added recovery time&lt;/strong&gt;, which is the cooldown seen from the other side. The dependency can be perfectly healthy again while your breaker is still refusing to call it, because its timer has not finished running.&lt;/p&gt;

&lt;p&gt;What AWS ships instead, inside its own SDK, is a &lt;a href="https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/" rel="noopener noreferrer"&gt;retry token bucket&lt;/a&gt;. Retries cost tokens, the bucket refills over time, and while there are tokens in it everything retries freely. Once the tokens run out, retries do not stop, they continue at a fixed rate. That behaviour went into the AWS SDK back in 2016.&lt;/p&gt;

&lt;p&gt;The structural difference is that the bucket is continuous where the breaker is modal.&lt;/p&gt;

&lt;p&gt;No tripped state, no cooldown, no trial run, so no rarely exercised mode for you to test. Retry capacity gets thinner as the bucket empties instead of switching off.&lt;/p&gt;

&lt;p&gt;The bucket also never stops the original call. It limits the retries only, so the caller keeps giving the dependency a chance on every request rather than blocking a whole dependency at once.&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Your call to a dependency&amp;lt;br/&amp;gt;is failing] --&amp;gt; B{What are you&amp;lt;br/&amp;gt;actually protecting&amp;lt;br/&amp;gt;against?}
    B --&amp;gt;|My own retries are&amp;lt;br/&amp;gt;amplifying the outage| C[Retry token bucket]
    B --&amp;gt;|My workers pile up on a&amp;lt;br/&amp;gt;dependency that stopped&amp;lt;br/&amp;gt;being useful| D[Circuit breaker]
    C --&amp;gt; E[Continuous, no modes,&amp;lt;br/&amp;gt;easy to test&amp;lt;br/&amp;gt;never blocks the first call]
    D --&amp;gt; F{Can it actually trip&amp;lt;br/&amp;gt;on your traffic volume?}
    F --&amp;gt;|No, too few calls&amp;lt;br/&amp;gt;for minimumNumberOfCalls| G[You have a decoration,&amp;lt;br/&amp;gt;not a breaker]
    F --&amp;gt;|Yes| H[You now own a mode&amp;lt;br/&amp;gt;you must test and alarm on]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef good fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef warn fill:#ff9a5c,stroke:#c25b23,color:#1a1a1a

    class B,F decision
    class A start
    class C,D,E good
    class G,H warn&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;So the rule comes down to what you are actually protecting against.&lt;/p&gt;

&lt;p&gt;If the worry is amplification from your own retries, the bucket handles that continuously and is easier to test.&lt;/p&gt;

&lt;p&gt;If the worry is your workers piling up against a dependency that has stopped being useful, the breaker is the thing that stops the waiting.&lt;/p&gt;

&lt;p&gt;Choose it knowing you have added a mode you now have to test and alarm on.&lt;/p&gt;
&lt;h2&gt;
  
  
  Failures neither of them should ever act on
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Your own bugs.&lt;/strong&gt; A 400 is the clearest case. That is a broken request, and it will fail identically on every host you send it to. Count those and you can trip a breaker against a dependency that is completely healthy. A good breaker sorts failures by type, and can require more timeouts before tripping than outright "service unavailable" responses, because those two mean different things.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Writes with side effects.&lt;/strong&gt; A write is only safe to retry if it is idempotent, meaning the side effect happens once no matter how many times the same request arrives. Read-only APIs usually get that for free. Resource creation APIs often do not, and that is what makes retries and fallbacks dangerous on a write path with no idempotency key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scoping&lt;/strong&gt;, which is the one people get wrong quietly. One breaker per resource type stops working the moment that resource has independent providers behind it. Take a sharded data store. One shard can be completely fine while another has a temporary problem, and if you merge their errors into one breaker, it blocks calls to healthy shards while still letting calls through to the failing one.&lt;/p&gt;

&lt;p&gt;Azure also lists five cases where a breaker only adds overhead: local in-memory resources with no network to protect, anything an ordinary retry already handles, cases where waiting out a reset introduces a delay you cannot accept, event-driven systems where failed work already lands in a dead letter queue, and systems where the platform underneath already handles recovery.&lt;/p&gt;

&lt;p&gt;That last one deserves its own section.&lt;/p&gt;
&lt;h2&gt;
  
  
  Something may already be doing this for you
&lt;/h2&gt;

&lt;p&gt;Everything so far has lived inside your own process, and that is not the only place this happens.&lt;/p&gt;

&lt;p&gt;A service mesh can run a breaker for you, as a sidecar or as a capability of the platform underneath.&lt;/p&gt;

&lt;p&gt;What it does is not quite the same thing though. Your library breaker stops calling a dependency. &lt;a href="https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/outlier" rel="noopener noreferrer"&gt;Envoy ejects a host&lt;/a&gt;, which means it pulls one bad instance out of the load balancing pool and lets the other four keep serving.&lt;/p&gt;

&lt;p&gt;It also arrives with its own defaults, and they look nothing like the library numbers.&lt;/p&gt;

&lt;p&gt;Envoy ejects after five consecutive 5xx responses, sweeping and re-evaluating every 10 seconds. It ejects a host for a base of 30 seconds, and it will never eject more than 10% of the fleet.&lt;/p&gt;

&lt;p&gt;That 30 seconds is not fixed, because ejection backs itself off. The duration is the base time multiplied by how many times that host has been ejected in a row, so 30 seconds becomes 60, then 90.&lt;/p&gt;

&lt;p&gt;Coming back is not only a matter of waiting either. A successful active health check un-ejects the host and clears its outlier counters, and clearing those resets the backoff too.&lt;/p&gt;

&lt;p&gt;The setting that matters most mid-incident is that &lt;strong&gt;maximum ejection percentage&lt;/strong&gt;. Once enough of the fleet is already out, the mesh stops ejecting, which is why you can watch a host that is clearly failing stay in the pool.&lt;/p&gt;

&lt;p&gt;So before you tune anything in your own process, find out which of these layers is already acting on your traffic.&lt;/p&gt;

&lt;p&gt;Possibly more than one. And all of them can be doing their job without telling you.&lt;/p&gt;
&lt;h2&gt;
  
  
  The part everyone skips: alarm on the state change
&lt;/h2&gt;

&lt;p&gt;Fowler's requirement is straightforward. Any change in breaker state gets logged, and the breaker reveals its state so something can monitor it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The state change is the thing you alarm on, not the errors underneath it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Skip that wire and an open breaker becomes a silent partial outage.&lt;/p&gt;

&lt;p&gt;Think about what your dashboards see. Every call the breaker turns away comes back fast and looks like a clean response. Your latency graph improves. Your error rate stays flat.&lt;/p&gt;

&lt;p&gt;A good fallback does exactly the same thing to your monitoring.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhtv3o46znakvmmly0t7q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhtv3o46znakvmmly0t7q.png" alt="Nothing To Do Here meme: the breaker is open and every call is failing fast, while the latency graph is green, the error rate is flat and nothing alarms" width="360" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The device needs a handle a person can reach, too. Operations staff should be able to trip or reset a breaker by hand, so you can force one closed when you know the dependency is back, or force it open when you know it is down.&lt;/p&gt;

&lt;p&gt;And the person on call needs to tell those two apart. Polly makes that visible by throwing a different exception when a breaker was deliberately isolated, so a human decision never looks like the system tripping on its own.&lt;/p&gt;
&lt;h2&gt;
  
  
  The whole chain, one question per link
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz7fac7w7kisztf8ifrzd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz7fac7w7kisztf8ifrzd.png" alt="Diagram of the four links on the call path from checkout to payment: timeout, retry, breaker and fallback, each labelled with the question it answers" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Timeout.&lt;/strong&gt; How long are you willing to wait before you call it a failure?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retry.&lt;/strong&gt; How many times do you try again, and with how much jitter so your retries do not all arrive together?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Breaker&lt;/strong&gt;, which is really three questions, because that is what a breaker is. What share of a real number of calls has to fail before you stop trying? How long do you stay stopped? How many probes do you let through on the way back?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fallback.&lt;/strong&gt; What do you actually return while you are stopped?&lt;/p&gt;

&lt;p&gt;Two of those numbers carry more weight than the rest, and both are usually sitting at whatever the library shipped.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;minimum call volume&lt;/strong&gt; decides whether your failure rate means anything at all.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;half-open probe count&lt;/strong&gt; decides whether your recovery survives contact with real traffic.&lt;/p&gt;

&lt;p&gt;The most useful thing you can do after reading this is open your own config and read those two lines. Then check that something alarms when the breaker changes state.&lt;/p&gt;

&lt;p&gt;Set your timeout from a percentile rather than a round number. Jitter your retries so they do not all arrive together. Trip on a rate measured over a real number of calls. And decide what you return before the breaker ever opens.&lt;/p&gt;

&lt;p&gt;Those four choices are the whole chain, and every number on it is yours rather than the library's.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>backend</category>
      <category>architecture</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How Google Stores a Planet: The GFS, Explained</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Wed, 09 Sep 2026 17:32:43 +0000</pubDate>
      <link>https://dev.to/lovestaco/how-google-stores-a-planet-the-gfs-explained-1fcp</link>
      <guid>https://dev.to/lovestaco/how-google-stores-a-planet-the-gfs-explained-1fcp</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Roughly 720,000 hours of video get uploaded to YouTube every day.&lt;/p&gt;

&lt;p&gt;Call it 1,000 terabytes. Tomorrow, another 1,000. The day after, another.&lt;/p&gt;

&lt;p&gt;To hold that you need hundreds of thousands of machines with millions of drives spinning inside them.&lt;/p&gt;

&lt;p&gt;And here is the uncomfortable arithmetic: in a fleet that size, a drive is dying right now. Another one will die before you finish this post.&lt;/p&gt;

&lt;p&gt;Yet not one second of anybody's cat video goes missing.&lt;/p&gt;

&lt;p&gt;So how?&lt;/p&gt;

&lt;p&gt;The obvious answer is money. &lt;/p&gt;

&lt;p&gt;Google is worth a few trillion dollars, so surely they just buy the good computers, the ones that do not break.&lt;/p&gt;

&lt;p&gt;Wrong. That computer does not exist.&lt;/p&gt;

&lt;p&gt;Physics does not offer an enterprise tier. &lt;/p&gt;

&lt;p&gt;Any machine you run will eventually fail, and once you have hundreds of thousands of them, failure stops being an event and becomes a background hum.&lt;/p&gt;

&lt;p&gt;That is almost exactly how Google opened &lt;a href="https://static.googleusercontent.com/media/research.google.com/en//archive/gfs-sosp2003.pdf" rel="noopener noreferrer"&gt;the 2003 Google File System paper&lt;/a&gt;: component failures are the norm, not the exception.&lt;/p&gt;

&lt;p&gt;What YouTube runs on today is a descendant of that design. &lt;/p&gt;

&lt;p&gt;The open source clone, &lt;a href="https://hadoop.apache.org/docs/stable/hadoop-project-dist/hadoop-hdfs/HdfsDesign.html" rel="noopener noreferrer"&gt;HDFS&lt;/a&gt;, became the storage layer the entire big data industry stood on for a decade.&lt;/p&gt;

&lt;p&gt;Let's build it up from scratch, one broken assumption at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, the file system you already have
&lt;/h2&gt;

&lt;p&gt;Before we scale to a planet, look at your laptop.&lt;/p&gt;

&lt;p&gt;Your operating system ships with a file system whose whole job is organizing bytes on a disk.&lt;/p&gt;

&lt;p&gt;It carves your storage into equal sized blocks. Usually 4 KB each.&lt;/p&gt;

&lt;p&gt;A 1 TB drive is therefore something like 268 million blocks, numbered from 0 all the way up.&lt;/p&gt;

&lt;p&gt;Now save &lt;code&gt;cat.png&lt;/code&gt;, a 12 KB masterpiece.&lt;/p&gt;

&lt;p&gt;The file system chops it into three 4 KB chunks and drops each chunk into whichever block happens to be free. Not neatly in a row. Wherever there is space.&lt;/p&gt;

&lt;p&gt;To ever see your cat again, it records where the pieces went in an index. Every file is a row: here are its chunks, here are the block numbers.&lt;/p&gt;

&lt;p&gt;Click the file, the index is read, the chunks are gathered, the cat appears.&lt;/p&gt;

&lt;p&gt;Hold that picture, because the rest of this post is the same idea with the blocks replaced by entire computers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5kp4o2k1l7ll7mjcxuz9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5kp4o2k1l7ll7mjcxuz9.png" alt="Diagram: cat.png split into three 4KB chunks scattered across numbered blocks, with the master index table mapping file to block numbers" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The file that does not fit
&lt;/h2&gt;

&lt;p&gt;Now try to store YouTube.&lt;/p&gt;

&lt;p&gt;The biggest enterprise drive you can buy today tops out around a few hundred terabytes and costs about as much as a car you would be nervous to park outside.&lt;/p&gt;

&lt;p&gt;One day of uploads already exceeds it.&lt;/p&gt;

&lt;p&gt;The naive fix is to build one gigantic machine and jam drives into it until the ingest fits.&lt;/p&gt;

&lt;p&gt;Two problems, and neither is subtle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is one power cable away from oblivion.&lt;/strong&gt; One outage, one fire, one clumsy technician, and every video ever uploaded is gone at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It does not scale.&lt;/strong&gt; There is a hard ceiling on how many drives you can hang off a single box, and a much lower ceiling on how many requests it can serve.&lt;/p&gt;

&lt;p&gt;So take the local file system idea and stretch it. Instead of many blocks on one machine, use many machines.&lt;/p&gt;

&lt;p&gt;Separate boxes, separate storage, separate buildings, separate power, talking over a network.&lt;/p&gt;

&lt;p&gt;A file comes in, gets chopped into chunks, each chunk lands on one of those machines.&lt;/p&gt;

&lt;p&gt;Call them &lt;strong&gt;chunk servers&lt;/strong&gt;, because they store chunks and later serve them. Each one holds many chunks from many different files.&lt;/p&gt;

&lt;p&gt;Then one &lt;strong&gt;master&lt;/strong&gt; plays the role the index table played. It knows every file, its chunk list, and the IP of the chunk server holding each chunk.&lt;/p&gt;

&lt;p&gt;A client asks for a cat video. The master hands back a list of chunks and addresses. The client fetches them itself and reassembles the video.&lt;/p&gt;

&lt;p&gt;Notice what the master is not doing: it never touches the video bytes. It hands out a map, then gets out of the way. &lt;/p&gt;

&lt;p&gt;That single decision is why one master can serve thousands of machines without melting.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqjt7x1tkg7atym6muk9u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqjt7x1tkg7atym6muk9u.png" alt="Diagram: client asking the master for a file, master returning chunk handles plus chunk server IPs, client fetching chunks directly from three chunk servers" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part where everything dies
&lt;/h2&gt;

&lt;p&gt;Now kill a chunk server.&lt;/p&gt;

&lt;p&gt;A piece of the cat video just became unreachable, and a video missing a chunk is not a video. It is a buffering spinner with commitment issues.&lt;/p&gt;

&lt;p&gt;Your instinct says this is rare. And for one machine your instinct is right. A single server might fail once a year.&lt;/p&gt;

&lt;p&gt;But run the numbers across a fleet.&lt;/p&gt;

&lt;p&gt;A year is about 31 million seconds. Spread one failure per server per year across a million servers and you get a failure roughly every 30 seconds, forever.&lt;/p&gt;

&lt;p&gt;If your system needs every machine online to work, then your system is broken every 30 seconds.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffq9xwjrfxwjjp7dlzjc5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffq9xwjrfxwjjp7dlzjc5.png" alt="This Is Fine meme, engineer calmly sitting in a room on fire because constant server death is the design assumption" width="360" height="349"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the real GFS insight, and it is more philosophical than technical.&lt;/p&gt;

&lt;p&gt;You do not build a reliable system by buying reliable parts. You build it by assuming the parts are garbage and designing around their funerals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Replication buys you time
&lt;/h2&gt;

&lt;p&gt;The first move is the obvious one. Keep more than one copy.&lt;/p&gt;

&lt;p&gt;Every chunk is written to multiple chunk servers. Each copy is a &lt;strong&gt;replica&lt;/strong&gt;, and the number of copies is the &lt;strong&gt;replication factor&lt;/strong&gt;, which GFS sets to 3 by default.&lt;/p&gt;

&lt;p&gt;The master now tracks, per chunk, the desired replication factor and every server holding a copy.&lt;/p&gt;

&lt;p&gt;One server goes dark, the client just asks a different one. No drama.&lt;/p&gt;

&lt;p&gt;But be honest about what this actually bought you.&lt;/p&gt;

&lt;p&gt;Nothing was solved. Time was purchased.&lt;/p&gt;

&lt;p&gt;Run long enough and all three replicas of some chunk will eventually be dead at the same time, and that chunk is gone for good. Three copies of a decaying thing is still a decaying thing.&lt;/p&gt;

&lt;p&gt;You need the system to notice the loss and react faster than the losses accumulate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Heartbeats, and a system that heals itself
&lt;/h2&gt;

&lt;p&gt;Every chunk server sends the master a &lt;strong&gt;heartbeat&lt;/strong&gt;, say every few seconds. It means nothing more than "still here."&lt;/p&gt;

&lt;p&gt;Miss a few in a row and the master declares that server dead. It strikes it from the replica list of every chunk it was holding.&lt;/p&gt;

&lt;p&gt;Which means those chunks now have two replicas instead of three. Under target.&lt;/p&gt;

&lt;p&gt;So the master picks a fresh chunk server that does not already hold that chunk and tells it to copy the chunk from a server that still has a good replica.&lt;/p&gt;

&lt;p&gt;Three again.&lt;/p&gt;

&lt;p&gt;That loop, running constantly, is the whole trick. &lt;/p&gt;

&lt;p&gt;Machines die at a steady rate and re-replication runs at a faster rate, so the system sits in equilibrium while the hardware underneath it quietly rots.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Chunk server heartbeats to master] --&amp;gt; B{Heartbeat received?}
    B --&amp;gt;|Yes| C[Mark server alive, refresh chunk map]
    C --&amp;gt; A
    B --&amp;gt;|No, 3 misses in a row| D[Declare server dead]
    D --&amp;gt; E[Remove it from replica list of every chunk it held]
    E --&amp;gt; F{Replicas below replication factor?}
    F --&amp;gt;|No| A
    F --&amp;gt;|Yes| G[Pick a chunk server without this chunk]
    G --&amp;gt; H[Copy chunk from a healthy replica]
    H --&amp;gt; I[Replica count restored to 3]
    I --&amp;gt; A

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef chip fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef bad fill:#ff9a5c,stroke:#c25c1f,color:#1a1a1a

    class B,F decision
    class A start
    class C,G,H,I chip
    class D,E bad&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The nice property here is that nobody pages a human at 3am for a dead disk. The dead disk is a routine input to a loop, not an incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  But who watches the master?
&lt;/h2&gt;

&lt;p&gt;You have probably spotted the hole.&lt;/p&gt;

&lt;p&gt;Every replica of every chunk is tracked by exactly one master, and that master holds the state of the world in its memory.&lt;/p&gt;

&lt;p&gt;Congratulations, you built a fleet of disposable machines and then hung its entire availability off one very important box.&lt;/p&gt;

&lt;p&gt;The fix rhymes with what you already did. &lt;/p&gt;

&lt;p&gt;The master streams every change it makes to a backup master, which sits there receiving updates and doing absolutely nothing else. This is the &lt;strong&gt;failover&lt;/strong&gt; master.&lt;/p&gt;

&lt;p&gt;Both masters heartbeat to a health check service.&lt;/p&gt;

&lt;p&gt;Clients never hardcode a master IP. They resolve a DNS name, say &lt;code&gt;master.internal&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;When the health check service stops hearing from the primary master, it flips that DNS record to point at the failover, which already holds a near current copy of the state and simply takes over.&lt;/p&gt;

&lt;p&gt;The pattern repeats at every layer: detect death with heartbeats, keep a warm copy, redirect traffic. Same song, different instrument.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F27xn4gk6qqn4w3grurkm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F27xn4gk6qqn4w3grurkm.png" alt="Diagram: primary master streaming state changes to a failover master, both heartbeating to a health check server, DNS record being flipped when the primary goes silent" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading is easy. Writing is where it gets fun.
&lt;/h2&gt;

&lt;p&gt;So far everything has been about pulling data out. Downloads are pleasant because nothing changes underneath you.&lt;/p&gt;

&lt;p&gt;Writes are where distributed systems earn their reputation.&lt;/p&gt;

&lt;p&gt;Here is the scenario the paper cares about, dressed in something familiar.&lt;/p&gt;

&lt;p&gt;You share a spreadsheet with your neighbours for booking apartment parking spots. Each row is a slot. Reserving means appending your name.&lt;/p&gt;

&lt;p&gt;Say the whole thing is one chunk, replicated across three chunk servers.&lt;/p&gt;

&lt;p&gt;You want the spot. You ask the master for the chunk server addresses, get all three, and send your update to each of them. They confirm. All three copies are identical. Your car is parked. Beautiful.&lt;/p&gt;

&lt;p&gt;Now your neighbour wants the same slot at the same moment.&lt;/p&gt;

&lt;p&gt;You both get the same three addresses. You both fire off your updates.&lt;/p&gt;

&lt;p&gt;Chunk server A receives yours first, then your neighbour's. It appends you, then them.&lt;/p&gt;

&lt;p&gt;Chunk server B receives your neighbour's first. It appends them, then you.&lt;/p&gt;

&lt;p&gt;The three replicas of a chunk that are supposed to be byte identical now disagree about reality, and there is no way to tell which one is right.&lt;/p&gt;

&lt;p&gt;The paper has a word for this. &lt;strong&gt;Inconsistent.&lt;/strong&gt; Data that should be the same on every server is not.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9jaqpq2omc86kpchxu0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9jaqpq2omc86kpchxu0.png" alt="Who Killed Hannibal meme: the chunk server applies writes in whatever order they arrive, then asks why the replicas disagree" width="360" height="404"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Elect somebody to be right
&lt;/h2&gt;

&lt;p&gt;The root cause is that each chunk server orders updates from its own point of view. It applies what it sees in the order it sees it, cheerfully unaware that its peers saw something else.&lt;/p&gt;

&lt;p&gt;Local time is not global truth.&lt;/p&gt;

&lt;p&gt;The fix is to stop asking three machines to independently agree and instead appoint one of them to decide. GFS calls that server the &lt;strong&gt;primary&lt;/strong&gt; for the chunk.&lt;/p&gt;

&lt;p&gt;The master chooses it, guarantees there is exactly one primary per chunk at any moment, and remembers who it is.&lt;/p&gt;

&lt;p&gt;Now the write flow changes shape:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You and your neighbour both ask the master for the chunk servers, and the master also tells you which one is the primary.&lt;/li&gt;
&lt;li&gt;You both push your data to all three replicas, which buffer it without applying anything yet.&lt;/li&gt;
&lt;li&gt;You both send a write request to the primary.&lt;/li&gt;
&lt;li&gt;The primary picks a serial order for every mutation it received and applies them in that order locally.&lt;/li&gt;
&lt;li&gt;It forwards that same order to the other replicas, which apply it exactly as told, discarding their own opinion about who came first.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Every replica ends up byte identical, no matter how many clients wrote at once.&lt;/p&gt;

&lt;p&gt;Data flows to whoever is closest on the network. Order flows from a single point of authority. Separating those two is the elegant bit.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;sequenceDiagram
    participant You
    participant Neighbour
    participant M as Master
    participant P as Primary replica
    participant S as Secondary replicas
    You-&amp;gt;&amp;gt;M: Where is this chunk?
    M--&amp;gt;&amp;gt;You: 3 addresses, plus who is primary
    Neighbour-&amp;gt;&amp;gt;M: Where is this chunk?
    M--&amp;gt;&amp;gt;Neighbour: same 3 addresses, same primary
    You-&amp;gt;&amp;gt;S: push data (buffered, not applied)
    Neighbour-&amp;gt;&amp;gt;S: push data (buffered, not applied)
    You-&amp;gt;&amp;gt;P: write request
    Neighbour-&amp;gt;&amp;gt;P: write request
    P-&amp;gt;&amp;gt;P: assign serial order: You, then Neighbour
    P-&amp;gt;&amp;gt;S: apply in this exact order
    S--&amp;gt;&amp;gt;P: applied
    P--&amp;gt;&amp;gt;You: done
    P--&amp;gt;&amp;gt;Neighbour: done&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Note who does not appear in the hot path there. The master hands out a map at the start and then vanishes. All the heavy lifting is client to chunk server.&lt;/p&gt;

&lt;h2&gt;
  
  
  What GFS deliberately gave up
&lt;/h2&gt;

&lt;p&gt;Every design decision above optimises the same thing: moving enormous amounts of data to enormous numbers of clients.&lt;/p&gt;

&lt;p&gt;It is a bandwidth machine, not a latency machine.&lt;/p&gt;

&lt;p&gt;Reading a 1 GB chunk stream is glorious. Reading one 200 byte record with a tight deadline is not what this was built for, and the designers knew it.&lt;/p&gt;

&lt;p&gt;That is the actual lesson hiding inside GFS, and it generalises far beyond storage.&lt;/p&gt;

&lt;p&gt;You pick the two or three properties your system genuinely needs, you optimise those relentlessly, and you take the others off the table on purpose.&lt;/p&gt;

&lt;p&gt;A system that refuses to choose is a system that is mediocre at everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  A question to sit with
&lt;/h2&gt;

&lt;p&gt;Here is the one the paper leaves you with, and it is worth chewing on before you look it up.&lt;/p&gt;

&lt;p&gt;Imagine one file in your GFS cluster goes viral. &lt;/p&gt;

&lt;p&gt;A single post, a single video, and suddenly a large slice of all traffic is hammering the handful of chunk servers holding its chunks.&lt;/p&gt;

&lt;p&gt;Those machines are drowning while the rest of the fleet is idle. The paper calls this a &lt;strong&gt;hotspot&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;How would you spread that load?&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://static.googleusercontent.com/media/research.google.com/en//archive/gfs-sosp2003.pdf" rel="noopener noreferrer"&gt;chunk size section of the paper&lt;/a&gt; has the answer Google reached for, and it is a smaller change than you would expect.&lt;/p&gt;

&lt;p&gt;Go read it. It is fifteen pages, it is written in plain English, and it is one of the few papers that reads like somebody explaining a thing they actually built rather than a thing they wanted funded.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;

  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The exact math, not a black box&lt;/th&gt;
&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;
&lt;th&gt;Every factor that feeds the score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

How does Blast Radius scoring work? (a more technical explanation)
&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;
&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>distributedsystems</category>
      <category>systemdesign</category>
      <category>googlecloud</category>
      <category>architecture</category>
    </item>
    <item>
      <title>7 Ways to Make Your API Faster</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Tue, 08 Sep 2026 16:42:29 +0000</pubDate>
      <link>https://dev.to/lovestaco/7-ways-to-make-your-api-faster-4020</link>
      <guid>https://dev.to/lovestaco/7-ways-to-make-your-api-faster-4020</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Your API is slow.&lt;/p&gt;

&lt;p&gt;You know it is slow because somebody in the product channel posted a screenshot of a spinner with the caption "is this normal".&lt;/p&gt;

&lt;p&gt;So you open the codebase, and within ninety seconds you have three theories, two refactors planned, and a strong urge to swap the JSON library.&lt;/p&gt;

&lt;p&gt;Stop.&lt;/p&gt;

&lt;p&gt;Put the keyboard down.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule before all the other rules
&lt;/h2&gt;

&lt;p&gt;Optimization is not step one. Optimization is step four, after measurement, after confirmation, and after you have found something that is actually slow.&lt;/p&gt;

&lt;p&gt;Every optimization in this post buys you speed with &lt;strong&gt;complexity&lt;/strong&gt;. &lt;/p&gt;

&lt;p&gt;Cache invalidation, connection pool tuning, pagination cursors, async log buffers, none of that is free. You pay for it forever, in every future debugging session.&lt;/p&gt;

&lt;p&gt;So the entry fee is a profile. &lt;/p&gt;

&lt;p&gt;Load test the endpoint, look at where the time actually goes, and only then pick a technique from the list below.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faha1inyrvlgafza0369i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faha1inyrvlgafza0369i.png" alt="The optimization loop: load test, then profile, then confirm the endpoint is actually slow, and only then pick one of the seven techniques" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The number of times I have watched somebody spend a week optimizing serialization for an endpoint whose real problem was one unindexed query is not a small number.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Measure. Confirm. Then optimize. In that order, every time.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwwxxxdazd05f22ud3me2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwwxxxdazd05f22ud3me2.png" alt="Left Exit 12 Off Ramp meme about swerving away from profiling and straight into a full rewrite" width="360" height="322"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Okay. Assume you measured. Here are the seven things worth reaching for.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Caching, or how to not do the work twice
&lt;/h2&gt;

&lt;p&gt;Caching is the highest leverage trick on this list, because the fastest database query is the one you never send.&lt;/p&gt;

&lt;p&gt;The shape is simple. An expensive computation runs once, the result goes into Redis or Memcached, and the next N callers asking the same question get the stored answer.&lt;/p&gt;

&lt;p&gt;The catch is that people think caching is a big architectural commitment. It usually is not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_top_products&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;top_products:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;hit&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query_expensive_top_products&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setex&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;   &lt;span class="c1"&gt;# thirty seconds. that is it.
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Look at that TTL. Thirty seconds.&lt;/p&gt;

&lt;p&gt;That feels almost insultingly short, and it is exactly the point. &lt;/p&gt;

&lt;p&gt;If an endpoint takes 400ms and gets hit 200 times a minute, a thirty second cache removes something like 99% of those database hits, and nobody downstream ever notices data that is half a minute stale.&lt;/p&gt;

&lt;p&gt;Short TTLs are underrated because they give you most of the win with almost none of the invalidation pain. &lt;/p&gt;

&lt;p&gt;You are not maintaining a cache, you are just refusing to answer the same question 200 times in a row.&lt;/p&gt;

&lt;p&gt;Where it gets genuinely hard is when the data must be fresh, and then you are in invalidation territory, which is famously &lt;a href="https://martinfowler.com/bliki/TwoHardThings.html" rel="noopener noreferrer"&gt;one of the two hard things in computer science&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Start with the boring TTL version. Graduate to invalidation only when the TTL version is provably wrong for your use case.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. Connection pooling, or stop reintroducing yourself
&lt;/h2&gt;

&lt;p&gt;Opening a database connection is not free.&lt;/p&gt;

&lt;p&gt;There is a TCP handshake, usually a TLS handshake, then authentication, then session setup. &lt;/p&gt;

&lt;p&gt;You can easily spend more time saying hello to Postgres than you spend querying it.&lt;/p&gt;

&lt;p&gt;Connection pooling keeps a set of connections open and hands them out. Your request borrows one, runs its query, and gives it back.&lt;/p&gt;

&lt;p&gt;Most frameworks do this by default and you never think about it. Which is fine, right up until the day you go serverless.&lt;/p&gt;

&lt;p&gt;Serverless breaks the assumption underneath pooling. Each function instance is its own little process with its own little pool, and the platform will happily spin up 500 of them during a traffic spike. &lt;/p&gt;

&lt;p&gt;Now your database, which is configured for maybe 100 connections, is getting introduced to 500 strangers at once.&lt;/p&gt;

&lt;p&gt;Postgres in particular does not degrade gracefully here. It forks a process per connection, so connection exhaustion is not a slowdown, it is a wall.&lt;/p&gt;

&lt;p&gt;That is the entire reason &lt;a href="https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-proxy.html" rel="noopener noreferrer"&gt;AWS RDS Proxy&lt;/a&gt; exists, along with pgbouncer and friends. &lt;/p&gt;

&lt;p&gt;They sit between your ephemeral functions and your very much non-ephemeral database, and multiplex a small pool of real connections across a large number of callers.&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    subgraph Serverless
      F1[Fn instance 1]
      F2[Fn instance 2]
      F3[Fn instance ...500]
    end

    F1 --&amp;gt; P[Connection Proxy]
    F2 --&amp;gt; P
    F3 --&amp;gt; P
    P --&amp;gt;|small, reused pool| DB[(Postgres)]

    F1 -.-&amp;gt;|without a proxy| DB
    F2 -.-&amp;gt;|500 handshakes| DB
    F3 -.-&amp;gt;|database says no| DB

    classDef fn fill:#6ea8ff,stroke:#2f5fb8,color:#1a1a1a
    classDef proxy fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef db fill:#ff9a5c,stroke:#c0632c,color:#1a1a1a

    class F1,F2,F3 fn
    class P proxy
    class DB db&lt;/code&gt;&lt;/pre&gt;


&lt;h2&gt;
  
  
  3. The N+1 query, the bug that looks like clean code
&lt;/h2&gt;

&lt;p&gt;This is my favourite one, because the N+1 query is almost always the result of writing &lt;em&gt;nicer&lt;/em&gt; code.&lt;/p&gt;

&lt;p&gt;You have posts. Each post has comments. So you write the obvious thing:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;posts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM posts LIMIT 20&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# one extra round trip. every. single. time.
&lt;/span&gt;    &lt;span class="n"&gt;post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;comments&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM comments WHERE post_id = %s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That is 21 queries to render one page. Twenty of them are identical in shape and differ only in an integer.&lt;/p&gt;

&lt;p&gt;On your laptop the database is a localhost away, each query costs 0.2ms, and the whole loop finishes before you can blink. &lt;/p&gt;

&lt;p&gt;In production the database is across a network, each round trip costs 3ms, and you have just spent 60ms doing nothing but waiting.&lt;/p&gt;

&lt;p&gt;Scale that page to 200 posts and you have an endpoint that is somehow slow without a single slow query in the logs. &lt;/p&gt;

&lt;p&gt;That is what makes N+1 so nasty. Every individual query looks fine. The slow query log has nothing to say. Only the count is wrong.&lt;/p&gt;

&lt;p&gt;The fix is to stop asking one at a time:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;posts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM posts LIMIT 20&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;ids&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM comments WHERE post_id = ANY(%s)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ids&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;by_post&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;defaultdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;by_post&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;post_id&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;comments&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;by_post&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Two queries. Constant, regardless of how many posts you fetch.&lt;/p&gt;

&lt;p&gt;If you are on an ORM, this is what &lt;code&gt;select_related&lt;/code&gt; and &lt;code&gt;prefetch_related&lt;/code&gt; in Django, &lt;code&gt;joinedload&lt;/code&gt; in SQLAlchemy, and &lt;code&gt;include&lt;/code&gt; in Prisma exist for. &lt;/p&gt;

&lt;p&gt;The tooling is there, it is just off by default, because the ORM cannot know whether you wanted the related rows.&lt;/p&gt;

&lt;p&gt;The single most useful habit here is to log your query count per request in development. An endpoint that fires 47 queries will tell on itself immediately.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe6iifguz11svnkcz3gn4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe6iifguz11svnkcz3gn4.png" alt="N+1 versus a batched fetch: twenty-one round trips and 63ms on the left, two round trips and 6ms on the right" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  4. Pagination, because nobody is reading 40,000 rows
&lt;/h2&gt;

&lt;p&gt;Somewhere in every codebase there is an endpoint that started life returning 12 records and now returns 40,000, because the table grew and nobody revisited the handler.&lt;/p&gt;

&lt;p&gt;The database has to fetch it. Your app has to serialize it. The network has to ship it. &lt;/p&gt;

&lt;p&gt;The client has to parse it, and then render precisely the first twenty of them.&lt;/p&gt;

&lt;p&gt;Pagination is the fix and everyone knows it. What everyone does not know is that &lt;code&gt;LIMIT 20 OFFSET 100000&lt;/code&gt; is not actually fast.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;OFFSET&lt;/code&gt; does not skip work. The database still walks all 100,000 rows and throws them away before handing you twenty. &lt;/p&gt;

&lt;p&gt;Deep pages get linearly slower, and your "optimization" quietly becomes the new bottleneck.&lt;/p&gt;

&lt;p&gt;Cursor based pagination avoids this by asking the question differently. Instead of "give me page 5000", you ask "give me the twenty rows after this one":&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- offset:  gets slower the deeper you go&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="k"&gt;OFFSET&lt;/span&gt; &lt;span class="mi"&gt;100000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- cursor:  uses the index, same cost on page 1 and page 5000&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;100000&lt;/span&gt; &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The second one is an index seek. It costs the same at any depth.&lt;/p&gt;

&lt;p&gt;The tradeoff is that you lose random access to page numbers, which is why offset pagination survives in admin panels and cursor pagination is what you find in &lt;a href="https://docs.stripe.com/api/pagination" rel="noopener noreferrer"&gt;Stripe's API&lt;/a&gt; and every infinite scroll feed you have ever used.&lt;/p&gt;
&lt;h2&gt;
  
  
  5. Serialization, the tax you pay per response
&lt;/h2&gt;

&lt;p&gt;Once the data is in memory, something has to turn it into JSON, and that something is running on your CPU for every single response.&lt;/p&gt;

&lt;p&gt;For small payloads this is noise. For an endpoint returning a few thousand objects, serialization can genuinely become the dominant cost, and the profiler will point right at it.&lt;/p&gt;

&lt;p&gt;The good news is that this is the cheapest fix on the entire list, because it is usually a library swap. In Python, &lt;code&gt;orjson&lt;/code&gt; is meaningfully faster than the standard library &lt;code&gt;json&lt;/code&gt;. &lt;/p&gt;

&lt;p&gt;In Node, the JSON serializer is native but schema based approaches like &lt;a href="https://github.com/fastify/fast-json-stringify" rel="noopener noreferrer"&gt;fast-json-stringify&lt;/a&gt; beat it by knowing the shape in advance. Serializers that get told the schema upfront can skip all the runtime type sniffing.&lt;/p&gt;

&lt;p&gt;But please, actually profile first. Swapping serializers on an endpoint that spends 95% of its time in the database is a lovely way to spend an afternoon achieving nothing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr9sl9vgxzqkd1rk43mib.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr9sl9vgxzqkd1rk43mib.png" alt="Meme of a man hammering nails into wet sand, labelled all seven optimization tricks applied to an endpoint you never profiled" width="360" height="352"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  6. Compression, the one your CDN can do for you
&lt;/h2&gt;

&lt;p&gt;JSON compresses beautifully, because JSON is mostly repeated key names and whitespace. Compression ratios of 5x to 10x on API responses are completely normal.&lt;/p&gt;

&lt;p&gt;That is 5x to 10x less data crossing the network, which matters enormously for anyone on mobile, and matters for everyone once payloads get big.&lt;/p&gt;

&lt;p&gt;gzip is the safe default that every client on earth supports. &lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Content-Encoding" rel="noopener noreferrer"&gt;Brotli&lt;/a&gt; generally compresses smaller at comparable speed and is supported by every modern browser, so it is worth enabling where you can.&lt;/p&gt;

&lt;p&gt;Two things to keep in mind.&lt;/p&gt;

&lt;p&gt;Compression costs CPU. You are trading processor time for network time, which is nearly always a good trade, but it is still a trade.&lt;/p&gt;

&lt;p&gt;Small responses can genuinely come out slower once you add the compression overhead, so most servers have a minimum size threshold, and you should leave it on.&lt;/p&gt;

&lt;p&gt;And you very likely should not be doing this yourself. Cloudflare, Fastly and friends will compress at the edge for you, which moves the CPU cost off your servers entirely and applies it uniformly to everything you serve. &lt;/p&gt;

&lt;p&gt;If you are already behind a CDN, this optimization is a checkbox.&lt;/p&gt;
&lt;h2&gt;
  
  
  7. Asynchronous logging, for when microseconds matter
&lt;/h2&gt;

&lt;p&gt;This one is last for a reason. Most services should not care.&lt;/p&gt;

&lt;p&gt;But in a high throughput path, writing a log line is a syscall, and if that write is synchronous and blocking, your request thread is sitting there waiting on a disk or a network socket while doing nothing useful.&lt;/p&gt;

&lt;p&gt;Async logging fixes this by making the request thread's job trivial. It drops the log entry into an in memory ring buffer and moves on immediately. A separate thread drains the buffer and does the actual writing.&lt;/p&gt;

&lt;p&gt;The request path goes from "wait for the write" to "append to a queue".&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    R[Request thread] --&amp;gt;|append, microseconds| B[In-memory buffer]
    B --&amp;gt; W[Logger thread]
    W --&amp;gt; D[(Disk / log service)]

    R --&amp;gt; RESP[Response sent]

    C{App crashes&amp;lt;br/&amp;gt;before flush?}
    B -.-&amp;gt; C
    C --&amp;gt;|yes| L[Buffered logs lost]
    C --&amp;gt;|no| D

    classDef thread fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef buf fill:#6ea8ff,stroke:#2f5fb8,color:#1a1a1a
    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef bad fill:#ff9a5c,stroke:#c0632c,color:#1a1a1a

    class R,W thread
    class B,RESP buf
    class C decision
    class L,D bad&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;That dotted branch is the whole tradeoff, and you should stare at it before enabling this.&lt;/p&gt;

&lt;p&gt;Anything sitting in the buffer when the process dies is gone. Which means the logs describing the crash are exactly the logs most likely to be lost.&lt;/p&gt;

&lt;p&gt;That is a genuinely bad trade for audit logs, payment records, or anything you would need in an incident review. &lt;/p&gt;

&lt;p&gt;It is a perfectly fine trade for high volume access logs where losing the last few hundred lines costs you nothing.&lt;/p&gt;

&lt;p&gt;Pick per log stream, not per application.&lt;/p&gt;
&lt;h2&gt;
  
  
  The actual takeaway
&lt;/h2&gt;

&lt;p&gt;Look at the seven again and notice how they cluster.&lt;/p&gt;

&lt;p&gt;Caching, pooling and N+1 are all about &lt;strong&gt;not talking to the database&lt;/strong&gt;, whether by skipping the question, skipping the handshake, or asking once instead of twenty times.&lt;/p&gt;

&lt;p&gt;Pagination, serialization and compression are all about &lt;strong&gt;moving less data&lt;/strong&gt;, at the query, at the CPU, and on the wire.&lt;/p&gt;

&lt;p&gt;Async logging is about &lt;strong&gt;getting out of the request path&lt;/strong&gt;, which is the same idea as background jobs, applied to something small.&lt;/p&gt;

&lt;p&gt;None of them are exotic. All of them are boring, well understood, and sitting one library call away.&lt;/p&gt;

&lt;p&gt;The hard part was never knowing the techniques. The hard part is having the discipline to find out which one your endpoint actually needs, instead of applying all seven and calling it architecture.&lt;/p&gt;

&lt;p&gt;Profile first. Fix the thing the profile points at. Then go do something more interesting.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fysboc2l5ruyyke0ymvba.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fysboc2l5ruyyke0ymvba.png" alt=" " width="360" height="312"&gt;&lt;/a&gt;&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>performance</category>
      <category>backend</category>
      <category>api</category>
    </item>
    <item>
      <title>SSH, Actually Explained: Handshakes, Keys, and the Tunnel Trick</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Mon, 07 Sep 2026 17:17:38 +0000</pubDate>
      <link>https://dev.to/lovestaco/ssh-actually-explained-handshakes-keys-and-the-tunnel-trick-48bf</link>
      <guid>https://dev.to/lovestaco/ssh-actually-explained-handshakes-keys-and-the-tunnel-trick-48bf</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;You type &lt;code&gt;ssh ubuntu@1.2.3.4&lt;/code&gt;, you get a prompt, and a server in another continent starts obeying your keyboard.&lt;/p&gt;

&lt;p&gt;Most of us learned that incantation on day one and never looked under it again.&lt;/p&gt;

&lt;p&gt;Which is fair. It works. It has worked for thirty years. Nothing about it demands your attention.&lt;/p&gt;

&lt;p&gt;But SSH is doing something genuinely clever in the half second before that prompt appears, and once you have seen it, a whole category of confusing errors stops being confusing.&lt;/p&gt;

&lt;p&gt;So let's open it up.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thing SSH replaced
&lt;/h2&gt;

&lt;p&gt;Before SSH there was telnet, and telnet had exactly one flaw.&lt;/p&gt;

&lt;p&gt;It sent everything in plain text. Your username, your password, every command, every byte of output.&lt;/p&gt;

&lt;p&gt;On a shared network that is not a subtle problem. Anyone sitting between you and the server was reading your session like a group chat they had been added to.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbj3xds0fl4lhzd8doqoy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbj3xds0fl4lhzd8doqoy.png" alt="Diagram: the same login shown twice on the wire, once as readable telnet text and once as SSH ciphertext, with an eavesdropper tapping both" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Telnet is a postcard. SSH is a locked briefcase.&lt;/p&gt;

&lt;p&gt;The postman is the same postman in both cases. That is the entire point.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fktsrt6kbalc5vm8fq3ki.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fktsrt6kbalc5vm8fq3ki.png" alt="It's A Trap meme about free cafe wifi and a telnet login" width="360" height="466"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Telnet is still installed on your machine, by the way. It survives as a debugging tool, because &lt;code&gt;telnet host 443&lt;/code&gt; is a quick way to ask "is this port even open". &lt;/p&gt;

&lt;p&gt;As a login protocol it is dead, and it deserved it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The handshake, which is the actually interesting part
&lt;/h2&gt;

&lt;p&gt;Here is the problem SSH has to solve in its first few milliseconds.&lt;/p&gt;

&lt;p&gt;Two machines that have never met need to agree on a secret encryption key, while shouting at each other across a network where everything they say is being recorded.&lt;/p&gt;

&lt;p&gt;That sounds impossible. It is not, and the reason is the neatest trick in applied cryptography.&lt;/p&gt;

&lt;p&gt;SSH uses two kinds of encryption, not one, and people conflate them constantly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Asymmetric crypto sets up the conversation.&lt;/strong&gt; It is slow, it involves public and private key pairs, and it is used only for the opening handshake.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Symmetric crypto carries the conversation.&lt;/strong&gt; One shared key, fast, and it does all the actual work of encrypting your keystrokes.&lt;/p&gt;

&lt;p&gt;The handshake exists purely to get both sides holding the same symmetric key.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7pcdclvpllk9tdkunnng.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7pcdclvpllk9tdkunnng.png" alt="Diagram: client and server exchanging public values, then each independently computing the identical session key that never crosses the wire" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The move is called a &lt;a href="https://en.wikipedia.org/wiki/Diffie%E2%80%93Hellman_key_exchange" rel="noopener noreferrer"&gt;Diffie-Hellman key exchange&lt;/a&gt;, and it goes like this.&lt;/p&gt;

&lt;p&gt;Each side generates a private value and derives a public value from it.&lt;/p&gt;

&lt;p&gt;They swap public values in the clear, where anyone can see them.&lt;/p&gt;

&lt;p&gt;Then each side combines its own private value with the other side's public value, and the maths works out such that both arrive at the same number.&lt;/p&gt;

&lt;p&gt;An observer who saw both public values cannot get there. They are missing either private half, and going backwards from a public value to a private one is the hard problem the whole thing rests on.&lt;/p&gt;

&lt;p&gt;So the session key is never transmitted. It is independently derived, twice, in two different countries.&lt;/p&gt;

&lt;p&gt;That key is also &lt;strong&gt;ephemeral&lt;/strong&gt;. It exists for this session and is thrown away when you disconnect. Record the traffic today, steal the server's private key next year, and you still cannot decrypt what you captured. That property is called &lt;a href="https://en.wikipedia.org/wiki/Forward_secrecy" rel="noopener noreferrer"&gt;forward secrecy&lt;/a&gt; and it is worth knowing the name of.&lt;/p&gt;

&lt;p&gt;Want to actually watch this happen? SSH will narrate it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;ssh &lt;span class="nt"&gt;-vv&lt;/span&gt; ubuntu@example.com
&lt;span class="go"&gt;debug1: SSH2_MSG_KEXINIT sent
debug1: kex: algorithm: curve25519-sha256
debug1: kex: host key algorithm: ssh-ed25519
debug1: Server host key: ssh-ed25519 SHA256:qN8Xb2...
debug1: SSH2_MSG_NEWKEYS sent
debug1: Authenticating to example.com:22 as 'ubuntu'
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;NEWKEYS&lt;/code&gt; is the moment the tunnel goes live. Notice that authentication happens on the line &lt;em&gt;after&lt;/em&gt; it.&lt;/p&gt;

&lt;p&gt;That ordering matters more than it looks.&lt;/p&gt;

&lt;p&gt;The encryption is set up first, and only then does SSH ask who you are.&lt;/p&gt;

&lt;p&gt;Your password, if you use one, is already inside the tunnel by the time you type it.&lt;/p&gt;
&lt;h2&gt;
  
  
  The host key, and that warning you always accept
&lt;/h2&gt;

&lt;p&gt;There is a gap in the story above.&lt;/p&gt;

&lt;p&gt;Key exchange gives you a secure channel to &lt;em&gt;somebody&lt;/em&gt;. It does not prove that somebody is the server you meant to reach.&lt;/p&gt;

&lt;p&gt;That is what the server's host key is for, and it is why the first connection asks you this:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The authenticity of host 'example.com' can't be established.
ED25519 key fingerprint is SHA256:qN8Xb2fV+3Kx9...
Are you sure you want to continue connecting (yes/no)?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;You have typed &lt;code&gt;yes&lt;/code&gt; there a thousand times without reading it. I have too.&lt;/p&gt;

&lt;p&gt;What you are being asked is whether that fingerprint belongs to the machine you think you are talking to. SSH cannot know. There is no certificate authority here, unlike the web.&lt;/p&gt;

&lt;p&gt;So it does the next best thing. It writes the fingerprint into &lt;code&gt;~/.ssh/known_hosts&lt;/code&gt; and screams if it ever changes.&lt;/p&gt;

&lt;p&gt;That scream is the &lt;code&gt;REMOTE HOST IDENTIFICATION HAS CHANGED&lt;/code&gt; block, and it is not being dramatic for fun.&lt;/p&gt;

&lt;p&gt;Either the server was genuinely rebuilt, or someone is sitting in the middle pretending to be it.&lt;/p&gt;

&lt;p&gt;Nine times out of ten it is a rebuilt server. The tenth time is the reason for the warning.&lt;/p&gt;
&lt;h2&gt;
  
  
  Keys beat passwords, and here is why
&lt;/h2&gt;

&lt;p&gt;Once the tunnel is up, SSH still needs to know who you are.&lt;/p&gt;

&lt;p&gt;You can use a password. You should not.&lt;/p&gt;

&lt;p&gt;The better way is a key pair you generate once, on your own machine.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh-keygen &lt;span class="nt"&gt;-t&lt;/span&gt; ed25519 &lt;span class="nt"&gt;-C&lt;/span&gt; &lt;span class="s2"&gt;"laptop"&lt;/span&gt;
&lt;span class="c"&gt;# ~/.ssh/id_ed25519       &amp;lt;- private. never leaves this machine.&lt;/span&gt;
&lt;span class="c"&gt;# ~/.ssh/id_ed25519.pub   &amp;lt;- public. paste it anywhere.&lt;/span&gt;

ssh-copy-id ubuntu@example.com   &lt;span class="c"&gt;# appends the .pub to the server's authorized_keys&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Use &lt;code&gt;ed25519&lt;/code&gt;, not RSA. It is smaller, faster, and has fewer ways to configure it badly. RSA is still fine at 4096 bits, but there is no reason to pick it for a new key.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn7a7w2p141ljobsee0zn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn7a7w2p141ljobsee0zn.png" alt="Diagram: server sends a random challenge, client signs it with the private key, server verifies with the stored public key" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The login is a challenge and response.&lt;/p&gt;

&lt;p&gt;You tell the server which public key you are claiming. The server checks &lt;code&gt;authorized_keys&lt;/code&gt;, finds it, and sends back a random chunk of data.&lt;/p&gt;

&lt;p&gt;Your client signs that data with the private key. The server verifies the signature against the public key it already had.&lt;/p&gt;

&lt;p&gt;Nothing secret ever crosses the network. Not on the first login, not on the thousandth.&lt;/p&gt;

&lt;p&gt;Compare that to a password, which crosses the wire on every single login, lives in someone's head, is probably reused, and is short enough to guess. A 256-bit key is not short enough to guess. That is not a marginal improvement, it is a different category.&lt;/p&gt;

&lt;p&gt;This is also why keys are the only sane option for automation. A CI pipeline cannot type a password, but it can hold a key.&lt;/p&gt;
&lt;h3&gt;
  
  
  You have already been doing this
&lt;/h3&gt;

&lt;p&gt;Two places where this is already in your muscle memory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;EC2.&lt;/strong&gt; When you launch an instance, AWS puts your public key into the image and hands you the &lt;code&gt;.pem&lt;/code&gt;, which is the private half. That is why &lt;code&gt;ssh -i mykey.pem ec2-user@1.2.3.4&lt;/code&gt; works with no password. AWS never had your private key and cannot recover it for you, which is &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-key-pairs.html" rel="noopener noreferrer"&gt;stated plainly in their docs&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub.&lt;/strong&gt; &lt;code&gt;git push&lt;/code&gt; over SSH is the exact same challenge and response. You put your public key in your account settings, and every push signs a challenge. GitHub &lt;a href="https://github.blog/security/application-security/removing-support-for-password-authentication/" rel="noopener noreferrer"&gt;removed password authentication for Git entirely in 2021&lt;/a&gt;, so it is keys or a token, nothing else.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flwt84hncyg90e59a3jhe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flwt84hncyg90e59a3jhe.png" alt="And It's Gone meme about committing the .pem file to a public repo" width="360" height="282"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A private key in a git repo is a server someone else owns now. GitHub scans public repos for exactly this, and bots scan them faster.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;chmod 600&lt;/code&gt; your keys, keep them out of repos, and put a passphrase on them so a stolen laptop is not a stolen server.&lt;/p&gt;
&lt;h2&gt;
  
  
  Local port forwarding, the underrated one
&lt;/h2&gt;

&lt;p&gt;Here is the SSH feature that feels like cheating the first time it works.&lt;/p&gt;

&lt;p&gt;You need to reach a MySQL database in a private subnet. It has no public IP. It has no route from the internet, on purpose, because it is a database.&lt;/p&gt;

&lt;p&gt;But there is a &lt;strong&gt;bastion host&lt;/strong&gt; in front of that network. One small public machine whose entire job is to be the single door, with port 22 open and nothing else.&lt;/p&gt;

&lt;p&gt;You can SSH into the bastion. So you can reach the database.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh &lt;span class="nt"&gt;-L&lt;/span&gt; 3306:db.internal:3306 ec2-user@bastion.example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Read the &lt;code&gt;-L&lt;/code&gt; argument as three parts: the local port you want to open, then the host and port to reach &lt;em&gt;from the bastion's point of view&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0i7m11l5gt8zwkvpqgd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0i7m11l5gt8zwkvpqgd.png" alt="Diagram: laptop to bastion over SSH, bastion to private MySQL inside the VPC, with the illusion that the app is just talking to localhost" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now &lt;code&gt;mysql -h 127.0.0.1 -P 3306&lt;/code&gt; on your laptop talks to that private database.&lt;/p&gt;

&lt;p&gt;Your MySQL client believes it is connected to something local. It has no idea a tunnel exists. Every byte goes through the encrypted SSH session and comes out on the far side of the firewall.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwm6scxbj9umdgsi55h2y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwm6scxbj9umdgsi55h2y.png" alt="Kramer meme about the firewall asking what is going on in port 22" width="360" height="438"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The same trick works for an internal dashboard, a Redis instance, an admin panel that should never be public, or a staging service someone forgot to expose.&lt;/p&gt;

&lt;p&gt;Two directions worth knowing, since the flags are easy to mix up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;-L&lt;/code&gt; brings a &lt;strong&gt;remote&lt;/strong&gt; service to your &lt;strong&gt;local&lt;/strong&gt; machine. This is the common one.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;-R&lt;/code&gt; pushes a &lt;strong&gt;local&lt;/strong&gt; service out to the &lt;strong&gt;remote&lt;/strong&gt; machine. Useful for showing someone your dev server, and also the flag you should audit on production boxes, since it lets an insider open a path inward.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you find yourself typing long tunnel commands daily, put them in &lt;code&gt;~/.ssh/config&lt;/code&gt; and forget them:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ssh"&gt;&lt;code&gt;&lt;span class="k"&gt;Host&lt;/span&gt; db-tunnel
    &lt;span class="k"&gt;HostName&lt;/span&gt; bastion.example.com
    &lt;span class="k"&gt;User&lt;/span&gt; ec2-user
    &lt;span class="k"&gt;IdentityFile&lt;/span&gt; ~/.ssh/prod.pem
    &lt;span class="k"&gt;LocalForward&lt;/span&gt; &lt;span class="m"&gt;3306&lt;/span&gt; db.internal:3306
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Then it is just &lt;code&gt;ssh db-tunnel&lt;/code&gt;. Your fingers will thank you.&lt;/p&gt;
&lt;h2&gt;
  
  
  When it refuses to let you in
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;Permission denied (publickey)&lt;/code&gt; is the least helpful error message in networking, because it covers about six different mistakes.&lt;/p&gt;

&lt;p&gt;Here is the order I actually check them in:&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Permission denied publickey] --&amp;gt; B{Did you pass the right key?}
    B --&amp;gt;|No| K[ssh -i mykey.pem or add to ~/.ssh/config]
    B --&amp;gt;|Yes| C{chmod 600 on the key file?}
    C --&amp;gt;|No| P[Fix permissions, SSH ignores world-readable keys]
    C --&amp;gt;|Yes| D{Public key in authorized_keys on the server?}
    D --&amp;gt;|No| U[ssh-copy-id, or paste it in]
    D --&amp;gt;|Yes| E{Right username for the image?}
    E --&amp;gt;|No| N[ec2-user, ubuntu, admin, not root]
    E --&amp;gt;|Yes| F{Home dir or .ssh perms too open?}
    F --&amp;gt;|Yes| H[chmod 700 ~/.ssh, 755 the home dir]
    F --&amp;gt;|No| L[Read the server log: journalctl -u sshd]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef fix fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef last fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a

    class B,C,D,E,F decision
    class A start
    class K,P,U,N,H fix
    class L last&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The permissions one catches everybody at least once. SSH will silently ignore a private key that other users can read, which is protective and infuriating in equal measure.&lt;/p&gt;

&lt;p&gt;The wrong-username one is the other classic. Amazon Linux wants &lt;code&gt;ec2-user&lt;/code&gt;, Ubuntu images want &lt;code&gt;ubuntu&lt;/code&gt;, and almost nothing wants &lt;code&gt;root&lt;/code&gt; anymore.&lt;/p&gt;

&lt;p&gt;When you have exhausted the client side, &lt;code&gt;ssh -vv&lt;/code&gt; tells you which keys were offered, and &lt;code&gt;journalctl -u sshd&lt;/code&gt; on the server tells you why each one was rejected. Between those two you will find it.&lt;/p&gt;
&lt;h2&gt;
  
  
  If you run the server side
&lt;/h2&gt;

&lt;p&gt;A short list, because most SSH hardening advice is longer than it needs to be.&lt;/p&gt;

&lt;p&gt;Turn off password authentication once your key works. &lt;code&gt;PasswordAuthentication no&lt;/code&gt; in &lt;code&gt;sshd_config&lt;/code&gt; deletes the entire brute-force attack surface in one line.&lt;/p&gt;

&lt;p&gt;Turn off direct root login. &lt;code&gt;PermitRootLogin no&lt;/code&gt;, then &lt;code&gt;sudo&lt;/code&gt; from a normal user, so the audit log has a name in it.&lt;/p&gt;

&lt;p&gt;Do not bother moving off port 22. It stops log noise from untargeted scanners, and nothing else. A real attacker runs a port scan.&lt;/p&gt;

&lt;p&gt;Test config changes in a second terminal while your first session is still open. Locking yourself out of a box by reloading a broken &lt;code&gt;sshd_config&lt;/code&gt; is a rite of passage best skipped.&lt;/p&gt;
&lt;h2&gt;
  
  
  What to take away
&lt;/h2&gt;

&lt;p&gt;SSH is not one idea, it is three stacked on each other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A key exchange&lt;/strong&gt; that lets two strangers agree on a secret in public, then hands the session to fast symmetric encryption.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An authentication step&lt;/strong&gt; that proves who you are by signing a challenge, so nothing worth stealing ever crosses the network.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A tunnel&lt;/strong&gt; that, once you have all that, will carry whatever other protocol you point at it.&lt;/p&gt;

&lt;p&gt;That third layer is the one most people never touch, and it is the one that turns SSH from a login tool into a general-purpose way to reach things that are deliberately unreachable.&lt;/p&gt;

&lt;p&gt;Go generate an ed25519 key, delete a password login, and forward a port. It is a good afternoon.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ssh</category>
      <category>security</category>
      <category>devops</category>
      <category>networking</category>
    </item>
    <item>
      <title>Markov Chain Monte Carlo: the 1953 algorithm hiding under modern AI</title>
      <dc:creator>Athreya aka Maneshwar</dc:creator>
      <pubDate>Sun, 06 Sep 2026 04:31:27 +0000</pubDate>
      <link>https://dev.to/lovestaco/markov-chain-monte-carlo-the-1953-algorithm-hiding-under-modern-ai-5cb4</link>
      <guid>https://dev.to/lovestaco/markov-chain-monte-carlo-the-1953-algorithm-hiding-under-modern-ai-5cb4</guid>
      <description>&lt;p&gt;&lt;em&gt;Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. &lt;a href="https://github.com/HexmosTech/LiveReview/" rel="noopener noreferrer"&gt;Star us&lt;/a&gt; to help devs discover the project, give it a try, and share your feedback to help improve the product.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;There is an algorithm that got invented in 1953 on a machine with less memory than the favicon of this page.&lt;/p&gt;

&lt;p&gt;It is used to forecast the weather.&lt;/p&gt;

&lt;p&gt;It is used to fit models of black hole mergers.&lt;/p&gt;

&lt;p&gt;It sits under the hood of every serious &lt;a href="https://en.wikipedia.org/wiki/Bayesian_statistics" rel="noopener noreferrer"&gt;Bayesian statistics&lt;/a&gt; library you have ever imported.&lt;/p&gt;

&lt;p&gt;And most working developers have never heard its name.&lt;/p&gt;

&lt;p&gt;It is called &lt;a href="https://en.wikipedia.org/wiki/Markov_chain_Monte_Carlo" rel="noopener noreferrer"&gt;Markov chain Monte Carlo&lt;/a&gt;, MCMC to its friends, and the reason it feels like a black box is that people usually explain it backwards. &lt;/p&gt;

&lt;p&gt;They start with detailed balance and ergodicity and stationary distributions, and by the time they get to the part that is actually clever you have closed the tab.&lt;/p&gt;

&lt;p&gt;So let's do it forwards.&lt;/p&gt;

&lt;p&gt;MCMC is two ideas bolted together, and both of them are simple enough to explain at a bar.&lt;/p&gt;

&lt;h2&gt;
  
  
  Idea one: you can measure things by throwing stuff at them
&lt;/h2&gt;

&lt;p&gt;Suppose I ask you for the value of pi and take away your calculator.&lt;/p&gt;

&lt;p&gt;You could derive it. People have. It is unpleasant.&lt;/p&gt;

&lt;p&gt;Or you could draw a square of side 2R, draw a circle of radius R inside it, and start throwing darts at the square while blindfolded.&lt;/p&gt;

&lt;p&gt;A dart that lands uniformly in the square has some probability of landing inside the circle, and that probability is just the ratio of the areas.&lt;/p&gt;

&lt;p&gt;Circle is &lt;code&gt;pi * R^2&lt;/code&gt;. Square is &lt;code&gt;4 * R^2&lt;/code&gt;. Ratio is &lt;code&gt;pi / 4&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;So throw a lot of darts, count how many landed inside, multiply by four, and you have pi.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsb57e2w689s69l90gnq2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsb57e2w689s69l90gnq2.png" alt="Diagram: a square with an inscribed circle, 150 random darts thrown at it, and the arithmetic showing pi is roughly four times the hit rate" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is the entire Monte Carlo method. It is named after the casino, because the people who invented it at Los Alamos were doing neutron diffusion calculations and needed a codename, and &lt;a href="https://en.wikipedia.org/wiki/Monte_Carlo_method#History" rel="noopener noreferrer"&gt;one of them had an uncle who kept borrowing money to gamble in Monte Carlo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In five lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;

&lt;span class="n"&gt;hits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
           &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;hits&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;10_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# 3.1417...
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Nobody in that snippet solved an integral. They counted.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyycn32ebeesy86muxhbk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyycn32ebeesy86muxhbk.png" alt="It ain't much but it's honest work meme, about approximating an integral by throwing ten million random darts at it" width="360" height="239"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The catch, and this is the catch that the whole rest of the article exists to fix, is that &lt;code&gt;random.random()&lt;/code&gt; gave us independent samples for free.&lt;/p&gt;

&lt;p&gt;We knew the shape we were sampling from. It was a square. Sampling uniformly from a square is easy.&lt;/p&gt;

&lt;p&gt;In every problem you actually care about, you do not know how to sample from the shape. That is the problem.&lt;/p&gt;
&lt;h2&gt;
  
  
  Idea two: a process that only remembers where it is right now
&lt;/h2&gt;

&lt;p&gt;A Markov chain is a sequence of states where the next state depends only on the current one.&lt;/p&gt;

&lt;p&gt;Not on how you got here. Not on the previous forty steps. Just on here.&lt;/p&gt;

&lt;p&gt;This is called the Markov property, and the shorthand for it is that the chain is memoryless.&lt;/p&gt;

&lt;p&gt;The textbook example is weather, so let's use the textbook example. Three states: rainy, cloudy, sunny. Fixed probabilities for hopping between them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxgco9nh3c4j6u9poncqv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxgco9nh3c4j6u9poncqv.png" alt="Diagram: a three state Markov chain between rainy, cloudy and sunny with transition probabilities, one sample run, and the Markov property written out" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;From rainy there is a 60% chance you go to cloudy and a 40% chance you stay put. From cloudy there is a 50% chance you go to sunny. And so on.&lt;/p&gt;

&lt;p&gt;Now run it. Rainy, cloudy, sunny, sunny, cloudy.&lt;/p&gt;

&lt;p&gt;Step five consulted step four and nothing else. It has no idea that the run started in the rain.&lt;/p&gt;

&lt;p&gt;Here is the property that makes this useful instead of merely cute.&lt;/p&gt;

&lt;p&gt;If the chain can reach every state and does not get permanently trapped anywhere, then the fraction of time it spends in each state converges to a fixed set of numbers.&lt;/p&gt;

&lt;p&gt;Run it a thousand steps and you might be rainy 30% of the time. Run it a million and it is still 30%. It has settled.&lt;/p&gt;

&lt;p&gt;That set of numbers is the &lt;strong&gt;stationary distribution&lt;/strong&gt;, and it does not depend on where you started.&lt;/p&gt;

&lt;p&gt;Sit with that for a second, because it is the whole trick.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A Markov chain, left alone, generates samples from a distribution.&lt;/strong&gt; Not a distribution you chose. Just whatever distribution falls out of the transition rules you happened to write down.&lt;/p&gt;

&lt;p&gt;So what if you ran that backwards?&lt;/p&gt;

&lt;p&gt;What if you had a distribution you wanted, and you designed the transition rules so that its stationary distribution was exactly that thing?&lt;/p&gt;

&lt;p&gt;Then you would have a machine that spits out samples from a distribution you were never able to sample from directly.&lt;/p&gt;

&lt;p&gt;That is MCMC. That is the entire idea. Everything else is engineering.&lt;/p&gt;
&lt;h2&gt;
  
  
  The distribution nobody can compute
&lt;/h2&gt;

&lt;p&gt;Time to be concrete about what "a distribution you cannot sample from" actually means, because otherwise this is all very abstract.&lt;/p&gt;

&lt;p&gt;Say you are doing Bayesian inference. You have data, you have a model with some parameters, and you want to know which parameter values are consistent with what you observed.&lt;/p&gt;

&lt;p&gt;A non-Bayesian method hands you one best fit number and a standard error. &lt;/p&gt;

&lt;p&gt;Bayesian inference hands you a whole distribution over parameters, called the posterior, which tells you both what is likely and how confident you are allowed to be about it.&lt;/p&gt;

&lt;p&gt;You get it from Bayes' theorem, which is four symbols and one enormous problem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8vivf0e4e2w757nk5i3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8vivf0e4e2w757nk5i3.png" alt="Diagram: Bayes theorem with the evidence integral in the denominator, a table showing grid evaluation cost exploding from 100 to 10^20, and the ratio trick that cancels the evidence" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The numerator is fine. The likelihood is "how well do these parameters explain my data", which you can evaluate. &lt;/p&gt;

&lt;p&gt;The prior is "what did I believe before I saw the data", which you wrote down yourself.&lt;/p&gt;

&lt;p&gt;The denominator is where it falls apart.&lt;/p&gt;

&lt;p&gt;It is called the evidence, and it is an integral over every possible combination of every parameter. &lt;/p&gt;

&lt;p&gt;It exists purely to make the whole thing sum to one.&lt;/p&gt;

&lt;p&gt;With one parameter, grid it at 100 points, evaluate 100 times, done before your coffee lands.&lt;/p&gt;

&lt;p&gt;With five parameters, that is ten billion evaluations.&lt;/p&gt;

&lt;p&gt;With ten parameters, which is a small model by any modern standard, you are at 10^20 and the sun has opinions about your timeline.&lt;/p&gt;

&lt;p&gt;This is &lt;a href="https://en.wikipedia.org/wiki/Curse_of_dimensionality" rel="noopener noreferrer"&gt;the curse of dimensionality&lt;/a&gt;, and it is not a performance problem you can engineer around. Grid methods die here. Every time.&lt;/p&gt;

&lt;p&gt;So you are stuck holding a distribution that you want to draw samples from, and you cannot even evaluate it, because evaluating it requires a normalising constant you cannot compute.&lt;/p&gt;

&lt;p&gt;Which sounds terminal.&lt;/p&gt;
&lt;h2&gt;
  
  
  The cancellation that saves everything
&lt;/h2&gt;

&lt;p&gt;Here is the move.&lt;/p&gt;

&lt;p&gt;MCMC never asks "what is the posterior probability at this point".&lt;/p&gt;

&lt;p&gt;It only ever asks "&lt;strong&gt;is this new point better or worse than the one I am standing on, and by how much&lt;/strong&gt;".&lt;/p&gt;

&lt;p&gt;That is a ratio. And in that ratio, the evidence appears on the top and on the bottom.&lt;/p&gt;

&lt;p&gt;So it cancels.&lt;/p&gt;

&lt;p&gt;You never compute the impossible thing. You just arrange never to need it.&lt;/p&gt;

&lt;p&gt;The consequence is that you only need the posterior &lt;strong&gt;up to a constant&lt;/strong&gt;, which is just likelihood times prior, which you can always evaluate.&lt;/p&gt;

&lt;p&gt;That is the load-bearing insight of the entire field.&lt;/p&gt;
&lt;h2&gt;
  
  
  Metropolis-Hastings, the whole thing
&lt;/h2&gt;

&lt;p&gt;The oldest MCMC algorithm is &lt;a href="https://bayes.wustl.edu/Manual/EquationOfState.pdf" rel="noopener noreferrer"&gt;Metropolis et al., 1953&lt;/a&gt;, later generalised by &lt;a href="https://www.jstor.org/stable/2334940" rel="noopener noreferrer"&gt;Hastings in 1970&lt;/a&gt;. It is short enough to hold in your head.&lt;/p&gt;

&lt;p&gt;You are standing at some parameter value. You want to take a step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Propose.&lt;/strong&gt; Draw a candidate near where you are, usually from a Gaussian centred on your current position.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Score.&lt;/strong&gt; Compute &lt;code&gt;R = p(proposed) / p(current)&lt;/code&gt;, using the unnormalised posterior, because that is all you have and all you need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decide.&lt;/strong&gt; If &lt;code&gt;R &amp;gt;= 1&lt;/code&gt;, the new spot is better. Move there. Always.&lt;/p&gt;

&lt;p&gt;If &lt;code&gt;R &amp;lt; 1&lt;/code&gt;, the new spot is worse. Move there anyway, with probability &lt;code&gt;R&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That last line is the one people skip past, and it is the one that makes the algorithm work.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fib8n4j0njvuhtcyr43y6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fib8n4j0njvuhtcyr43y6.png" alt="Diagram: a two peaked posterior with an uphill move always accepted, a downhill move accepted with probability R, and the dashed path showing how the chain crosses the valley to the second peak" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you only ever accepted uphill moves, you would have written a hill climber. &lt;/p&gt;

&lt;p&gt;It would sprint to the nearest peak, sit on it, and report that peak as the answer with total confidence, having never seen the rest of the distribution.&lt;/p&gt;

&lt;p&gt;The occasional deliberately bad move is what lets the chain roll down one hill and find another. It is what turns a greedy optimiser into a sampler.&lt;/p&gt;

&lt;p&gt;And the acceptance rule is not arbitrary. It is &lt;a href="https://en.wikipedia.org/wiki/Detailed_balance" rel="noopener noreferrer"&gt;constructed&lt;/a&gt; so that the chain's stationary distribution is exactly the posterior you handed it. &lt;/p&gt;

&lt;p&gt;The chain visits high probability regions often and low probability regions rarely, in precisely the right proportion.&lt;/p&gt;

&lt;p&gt;So it looks like a random walk, and it kind of is, but it is a random walk with a rigged floor.&lt;/p&gt;

&lt;p&gt;Here is the whole algorithm, and I mean the whole algorithm:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;metropolis&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;log_post&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_steps&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;step_size&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;log_post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;chain&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_steps&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;candidate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;normal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;step_size&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="n"&gt;lp_new&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;log_post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="c1"&gt;# log space, so the ratio is a subtraction and nothing overflows
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rand&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;lp_new&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;lp&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;candidate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lp_new&lt;/span&gt;
        &lt;span class="n"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="c1"&gt;# note: appended even when we rejected
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Two things in there that trip people up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;We work in log space.&lt;/strong&gt; Posteriors underflow to zero fast in float64, so the ratio becomes a subtraction of log densities. &lt;code&gt;np.log(rand()) &amp;lt; lp_new - lp&lt;/code&gt; is the same rule, numerically survivable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;We append &lt;code&gt;x&lt;/code&gt; even on rejection.&lt;/strong&gt; Staying put is a real outcome. If you only recorded accepted moves you would systematically under-count the sharp peaks, which are exactly the regions where most proposals get rejected.&lt;/p&gt;

&lt;p&gt;As a flowchart:&lt;br&gt;
&lt;/p&gt;
&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Pick a starting value] --&amp;gt; B[Propose a new value&amp;lt;br/&amp;gt;from a Gaussian around it]
    B --&amp;gt; C[Compute R = p_new / p_current]
    C --&amp;gt; D{R &amp;gt;= 1?}
    D -- yes, uphill --&amp;gt; E[Accept the move]
    D -- no, downhill --&amp;gt; F{Coin flip lands&amp;lt;br/&amp;gt;under R?}
    F -- yes --&amp;gt; E
    F -- no --&amp;gt; G[Reject, stay put&amp;lt;br/&amp;gt;and record the old value again]
    E --&amp;gt; H[Record the sample]
    G --&amp;gt; H
    H --&amp;gt; I{Still in burn-in?}
    I -- yes --&amp;gt; J[Throw this sample away]
    I -- no --&amp;gt; K[Keep it]
    J --&amp;gt; B
    K --&amp;gt; L{Enough samples?}
    L -- no --&amp;gt; B
    L -- yes --&amp;gt; M[Average the samples&amp;lt;br/&amp;gt;to get anything you want]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef good fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef bad fill:#ff9a5c,stroke:#c65f22,color:#1a1a1a
    classDef work fill:#6ea8ff,stroke:#2f5fbf,color:#1a1a1a

    class D,F,I,L decision
    class A,M start
    class E,K,H good
    class G,J bad
    class B,C work&lt;/code&gt;&lt;/pre&gt;


&lt;h2&gt;
  
  
  The step size is the one knob, and it will bite you
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;step_size&lt;/code&gt; looks like a tuning detail. It is not. It is the difference between a chain that works and a chain that lies to you.&lt;/p&gt;

&lt;p&gt;Make the proposals too narrow and almost everything gets accepted, because you barely moved. The chain shuffles across the distribution at a glacial pace and you will need millions of samples to see anything.&lt;/p&gt;

&lt;p&gt;Make them too wide and almost everything lands somewhere terrible and gets rejected. The chain stands still for hundreds of steps at a time, and your "10,000 samples" are actually about forty distinct values repeated.&lt;/p&gt;

&lt;p&gt;Both failures look like success from the outside. You get your array of 10,000 numbers either way.&lt;/p&gt;

&lt;p&gt;The folk rule for the simple random walk version is to aim for an acceptance rate around 20 to 25%, which comes from &lt;a href="https://projecteuclid.org/journals/annals-of-applied-probability/volume-7/issue-1/Weak-convergence-and-optimal-scaling-of-random-walk-Metropolis-algorithms/10.1214/aoap/1034625254.full" rel="noopener noreferrer"&gt;some genuinely lovely asymptotic work by Roberts, Gelman and Gilks&lt;/a&gt;. Print the acceptance rate. Always print the acceptance rate.&lt;/p&gt;

&lt;p&gt;This is also why nobody writes the loop above in production. Modern samplers like &lt;a href="https://mc-stan.org/" rel="noopener noreferrer"&gt;Stan&lt;/a&gt; and &lt;a href="https://www.pymc.io/" rel="noopener noreferrer"&gt;PyMC&lt;/a&gt; use Hamiltonian Monte Carlo and NUTS, which use gradients of the posterior to propose smart, distant moves instead of blind local wobbles, and tune themselves. The idea is identical. The proposal is just far less stupid.&lt;/p&gt;
&lt;h2&gt;
  
  
  Burn-in, or: throwing away the work you paid for
&lt;/h2&gt;

&lt;p&gt;You have to start the chain somewhere, and your somewhere is probably wrong.&lt;/p&gt;

&lt;p&gt;If your initial guess lands in a region of terrible posterior probability, the chain will spend a while wandering out of the wilderness before it finds the part of parameter space that matters.&lt;/p&gt;

&lt;p&gt;Those early samples are real samples, they cost real compute, and they are garbage. They describe your bad guess, not the posterior.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpvathgw5cce2bl892l3d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpvathgw5cce2bl892l3d.png" alt="Life finds a way meme, about starting the chain at a wildly wrong parameter value" width="360" height="202"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But remember the Markov property. The chain has no memory of where it started. Once it reaches the high probability region, it stays there, and nothing about its future behaviour is contaminated by the trek it took to get there.&lt;/p&gt;

&lt;p&gt;So the fix is embarrassingly blunt. Delete the first few hundred or few thousand samples.&lt;/p&gt;

&lt;p&gt;That is burn-in. It is not a hack, it is a direct consequence of memorylessness.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6694kprmfamnd63eyt7r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6694kprmfamnd63eyt7r.png" alt="Diagram: a trace plot with a red burn-in section walking in from a bad starting value, a cut line, the green stationary section after it, and the histogram the kept samples become" width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5j1po67h6u5c7xqoq6tr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5j1po67h6u5c7xqoq6tr.png" alt="Bilbo meme about keeping the first five thousand burn-in samples" width="360" height="359"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The way you check is the trace plot on the left of that diagram. Plot parameter value against iteration. A healthy chain looks like a fuzzy horizontal caterpillar. A chain still climbing in, or drifting, or sitting flat for long stretches, is telling you something and it is not good news.&lt;/p&gt;

&lt;p&gt;In practice people run several chains from different starting points and check they all converge to the same place, which is what &lt;a href="https://arxiv.org/abs/1903.08008" rel="noopener noreferrer"&gt;the R-hat statistic&lt;/a&gt; measures.&lt;/p&gt;
&lt;h2&gt;
  
  
  What you do with a pile of samples
&lt;/h2&gt;

&lt;p&gt;Once you have samples from the posterior, the hard part is over and everything downstream is embarrassingly easy.&lt;/p&gt;

&lt;p&gt;Want the mean of a parameter? Average the samples.&lt;/p&gt;

&lt;p&gt;Want a 95% credible interval? Sort them and take the middle 95%.&lt;/p&gt;

&lt;p&gt;Want the probability that a parameter is bigger than 10? Count how many are, divide by how many you have.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;chain&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;metropolis&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;log_post&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_steps&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;50_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;step_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;samples&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;5_000&lt;/span&gt;&lt;span class="p"&gt;:]&lt;/span&gt;                       &lt;span class="c1"&gt;# drop burn-in
&lt;/span&gt;
&lt;span class="n"&gt;samples&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                                &lt;span class="c1"&gt;# point estimate
&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;percentile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;samples&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;2.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;97.5&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;           &lt;span class="c1"&gt;# 95% credible interval
&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;samples&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                         &lt;span class="c1"&gt;# P(theta &amp;gt; 10), directly
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Every question you can ask about a distribution is an expectation, and every expectation is approximated by an average over samples. That is the Monte Carlo half, quietly doing its job at the end.&lt;/p&gt;

&lt;p&gt;Which is where the name finally makes sense. The Markov chain gives you the samples. The Monte Carlo gives you the answers. Neither half works alone.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where the two halves meet
&lt;/h2&gt;


&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    P[A posterior you cannot integrate] --&amp;gt; Q{Can you draw&amp;lt;br/&amp;gt;independent samples?}
    Q -- yes --&amp;gt; MC[Plain Monte Carlo&amp;lt;br/&amp;gt;throw darts, average them]
    Q -- no --&amp;gt; R{Can you at least&amp;lt;br/&amp;gt;evaluate it up to a constant?}
    R -- no --&amp;gt; STUCK[You are genuinely stuck]
    R -- yes --&amp;gt; CHAIN[Build a Markov chain&amp;lt;br/&amp;gt;whose stationary distribution&amp;lt;br/&amp;gt;is that posterior]
    CHAIN --&amp;gt; WALK[Walk it for a long time]
    WALK --&amp;gt; S[Correlated samples that still&amp;lt;br/&amp;gt;average to the right answer]
    MC --&amp;gt; ANS[Means, intervals, probabilities]
    S --&amp;gt; ANS

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef start fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef chip fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef accel fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a
    classDef bad fill:#ff9a5c,stroke:#c65f22,color:#1a1a1a

    class Q,R decision
    class P start
    class MC,S,ANS chip
    class CHAIN,WALK accel
    class STUCK bad&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;One honest caveat, because I do not want to oversell this.&lt;/p&gt;

&lt;p&gt;MCMC samples are correlated. Consecutive steps are near each other by construction, so 10,000 MCMC samples carry less information than 10,000 independent ones. The quantity that matters is the effective sample size, and it can be brutally smaller than the number you generated. Every decent library reports it. Look at it.&lt;/p&gt;

&lt;p&gt;MCMC also does not automatically work. It converges eventually, and "eventually" is doing real work in that sentence. A posterior with two well separated modes and a deep valley between them can keep a random walk chain trapped on one mode for longer than you are willing to wait, and the chain will look perfectly healthy the whole time. Convergence diagnostics are not paranoia, they are the job.&lt;/p&gt;
&lt;h2&gt;
  
  
  So why does this matter to you
&lt;/h2&gt;

&lt;p&gt;Because the pattern generalises well past statistics.&lt;/p&gt;

&lt;p&gt;You have an object you cannot enumerate. You cannot compute its total. But you can compare two candidates cheaply, and you can take a random step.&lt;/p&gt;

&lt;p&gt;That is enough. That is the whole precondition.&lt;/p&gt;

&lt;p&gt;Simulated annealing is this. So is a large chunk of statistical physics. So is the &lt;a href="https://en.wikipedia.org/wiki/PageRank" rel="noopener noreferrer"&gt;PageRank&lt;/a&gt; random surfer, which is a Markov chain whose stationary distribution is the ranking. So is every diffusion model generating an image right now, iteratively walking noise toward a distribution it learned rather than one you wrote down.&lt;/p&gt;

&lt;p&gt;The 1953 paper was about hard spheres in a box.&lt;/p&gt;

&lt;p&gt;SIAM later put the Metropolis algorithm on its &lt;a href="https://archive.siam.org/pdf/news/637.pdf" rel="noopener noreferrer"&gt;list of the top ten algorithms of the twentieth century&lt;/a&gt;, next to the FFT and QR decomposition.&lt;/p&gt;

&lt;p&gt;So next time someone asks how you do inference over a distribution you cannot compute, you have an answer.&lt;/p&gt;

&lt;p&gt;You do not compute it.&lt;/p&gt;

&lt;p&gt;You build a chain that walks through it for you, and you let the walking do the arithmetic.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;&lt;br&gt;
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.&lt;/em&gt;&lt;/p&gt;
&lt;em&gt;

&lt;p&gt;I'm building &lt;strong&gt;LiveReview&lt;/strong&gt;, a blast-radius aware AI code review built for your business-critical systems.&lt;/p&gt;

&lt;p&gt;Instead of presenting every diff with equal emphasis, &lt;strong&gt;LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Spend code review effort where business risk is highest — not spread evenly across every diff.&lt;/p&gt;

&lt;p&gt;⭐ Star it on GitHub: &lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/HexmosTech" rel="noopener noreferrer"&gt;
        HexmosTech
      &lt;/a&gt; / &lt;a href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;
        LiveReview
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Blast-Radius Aware AI Code Review for Business-Critical Systems
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/png/logo-with-text.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fpng%2Flogo-with-text.png" alt="LiveReview" height="80"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml" rel="noopener noreferrer"&gt;&lt;img alt="gitleaks.yml" title="gitleaks.yml: Secret scanning workflow" src="https://github.com/HexmosTech/LiveReview/actions/workflows/gitleaks.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml" rel="noopener noreferrer"&gt;&lt;img alt="osv-scanner.yml" title="osv-scanner.yml: Dependency vulnerability scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/osv-scanner.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml" rel="noopener noreferrer"&gt;&lt;img alt="govulncheck.yml" title="govulncheck.yml: Go vulnerability check" src="https://github.com/HexmosTech/LiveReview/actions/workflows/govulncheck.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml" rel="noopener noreferrer"&gt;&lt;img alt="semgrep.yml" title="semgrep.yml: Static analysis security scan" src="https://github.com/HexmosTech/LiveReview/actions/workflows/semgrep.yml/badge.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/gfx/dependabot-enabled.svg"&gt;&lt;img alt="dependabot-enabled" title="dependabot-enabled: Automated dependency updates are enabled" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fgfx%2Fdependabot-enabled.svg"&gt;&lt;/a&gt;&amp;nbsp;&lt;a href="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml" rel="noopener noreferrer"&gt;&lt;img alt="mcp-testcases.yml" title="mcp-testcases.yml: MCP integration test suite" src="https://github.com/HexmosTech/LiveReview/actions/workflows/mcp-testcases.yml/badge.svg"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;LiveReview is an AI code reviewer that scores every hunk of a diff by &lt;strong&gt;blast radius&lt;/strong&gt;: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.&lt;/p&gt;


  
    
    &lt;span class="m-1"&gt;blast-radius-demo.mp4&lt;/span&gt;
  

  

  


&lt;p&gt;&lt;i&gt;LiveReview's Blast Radius &amp;amp; Review Priority scoring, live in the diff viewer.&lt;/i&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;div class="table-wrapper-paragraph"&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;table&gt;

&lt;thead&gt;

&lt;tr&gt;

&lt;th&gt;The exact math, not a black box&lt;/th&gt;

&lt;th&gt;Visualize blast radius at a glance&lt;/th&gt;

&lt;th&gt;Every factor that feeds the score&lt;/th&gt;

&lt;/tr&gt;

&lt;/thead&gt;

&lt;tbody&gt;

&lt;tr&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-3.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-3.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-4.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-4.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/HexmosTech/LiveReview/./assets/screenshots/blast-radius/new-risk-score-2.webp"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FHexmosTech%2FLiveReview%2FHEAD%2F.%2Fassets%2Fscreenshots%2Fblast-radius%2Fnew-risk-score-2.webp" width="280"&gt;&lt;/a&gt;&lt;/td&gt;

&lt;/tr&gt;

&lt;/tbody&gt;

&lt;/table&gt;&lt;/div&gt;&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

How does Blast Radius scoring work? (a more technical explanation)

&lt;p&gt;&lt;strong&gt;Here's the goal:&lt;/strong&gt;&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.&lt;/li&gt;

&lt;li&gt;A 300-line UI change in one file, fully covered by…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/HexmosTech/LiveReview" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;&lt;b&gt;Click below to try LiveReview with your codebase:&lt;/b&gt;&lt;/p&gt;

&lt;/em&gt;&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexmos.com/livereview" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvls0pq7nymbrll98je6s.png" alt="LiveReview Banner" width="800" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>python</category>
      <category>datascience</category>
      <category>computerscience</category>
    </item>
  </channel>
</rss>
