<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sarthak Rawat</title>
    <description>The latest articles on DEV Community by Sarthak Rawat (@shogun_the_grt).</description>
    <link>https://dev.to/shogun_the_grt</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3648244%2F97b4e029-5833-4a8c-b9a0-6872183ab3a5.jpg</url>
      <title>DEV Community: Sarthak Rawat</title>
      <link>https://dev.to/shogun_the_grt</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shogun_the_grt"/>
    <language>en</language>
    <item>
      <title>Retrieval-Augmented Generation, Explained From First Principles</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Mon, 28 Sep 2026 12:43:52 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/retrieval-augmented-generation-explained-from-first-principles-199f</link>
      <guid>https://dev.to/shogun_the_grt/retrieval-augmented-generation-explained-from-first-principles-199f</guid>
      <description>&lt;p&gt;Large language models are remarkably good at producing answers. Ask one to explain a difficult idea, summarize a document, write code, or reason through a problem, and it can often do it in seconds.&lt;/p&gt;

&lt;p&gt;But there is a fundamental limitation hiding underneath all of that fluency: &lt;strong&gt;the model's knowledge is not the same thing as access to the information you need right now.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A model may know a great deal about a subject and still know nothing about your company's latest policy, yesterday's support tickets, this month's internal report, or the newest version of a regulation. It may also produce a confident answer when the information it needs is missing.&lt;/p&gt;

&lt;p&gt;Retrieval-Augmented Generation, or &lt;strong&gt;RAG&lt;/strong&gt;, is one of the most practical ways of dealing with that problem.&lt;/p&gt;

&lt;p&gt;The basic idea is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Instead of asking the model to remember everything, give it the&lt;br&gt;
relevant information at the moment it needs to answer.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The simplicity of that sentence is deceptive. A production RAG system is a chain of decisions about documents, representations, search, ranking, context, generation, evaluation, security, latency, and cost. Most of the interesting engineering happens in those decisions.&lt;/p&gt;

&lt;p&gt;This article builds RAG from the ground up, then follows the system into the problems that appear when a simple prototype becomes a real application.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem RAG is trying to solve
&lt;/h2&gt;

&lt;p&gt;Imagine a brilliant consultant who has read an enormous library of books, papers, manuals, and websites. They can reason extremely well, explain difficult ideas, and synthesize information quickly.&lt;/p&gt;

&lt;p&gt;There is one catch: they were locked in a room after a particular date.&lt;/p&gt;

&lt;p&gt;They have no idea what happened afterward.&lt;/p&gt;

&lt;p&gt;There is another problem. When they do not know something, they may still try to give you an answer. Because they are optimized to produce plausible language, the result can sound convincing even when it is wrong.&lt;/p&gt;

&lt;p&gt;That is a useful mental model for an LLM.&lt;/p&gt;

&lt;p&gt;Now imagine giving the consultant a research assistant. The assistant can search a filing cabinet, find the relevant pages, and put them on the consultant's desk before the consultant answers.&lt;/p&gt;

&lt;p&gt;The consultant still does the reasoning. The assistant provides the evidence.&lt;/p&gt;

&lt;p&gt;That is the central idea behind RAG.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval&lt;/strong&gt; finds relevant external information.&lt;br&gt;
&lt;strong&gt;Augmentation&lt;/strong&gt; puts that information into the model's context.&lt;br&gt;
&lt;strong&gt;Generation&lt;/strong&gt; produces an answer using that context.&lt;/p&gt;

&lt;p&gt;The model's reasoning ability and its external knowledge have been separated.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;RAG decouples what a model can reason about from what it currently has access to.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That distinction is more important than memorizing a particular RAG framework or vector database.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6v2vde29nwyfq2k9hawl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6v2vde29nwyfq2k9hawl.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  RAG, fine-tuning, and long context solve different problems
&lt;/h2&gt;

&lt;p&gt;A common mistake is to treat RAG and fine-tuning as competing ways of doing the same thing.&lt;/p&gt;

&lt;p&gt;They are not.&lt;/p&gt;

&lt;p&gt;RAG is primarily a way to give a model &lt;strong&gt;access to information&lt;/strong&gt; at inference time. Fine-tuning changes the model's learned behavior. A larger context window gives the model more room to receive information, but does not by itself create a retrieval mechanism.&lt;/p&gt;

&lt;p&gt;A useful shorthand is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Fine-tuning teaches the model how to behave; RAG gives the model what to know.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In practice, the techniques can be combined. A specialized model can learn domain vocabulary, formatting, or communication style through fine-tuning while RAG supplies current private knowledge.&lt;/p&gt;

&lt;p&gt;Long context changes the tradeoff again. If the complete information needed for a task is small enough to fit comfortably into context, retrieval may be unnecessary. RAG becomes more useful when the corpus is large, changing, private, or expensive to stuff into every prompt.&lt;/p&gt;

&lt;p&gt;There is no universal winner. The choice depends on the information, latency, cost, update frequency, and reliability requirements of the application.&lt;/p&gt;

&lt;h2&gt;
  
  
  The RAG pipeline
&lt;/h2&gt;

&lt;p&gt;At its simplest, RAG can be understood as four stages:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Index → Retrieve → Augment → Generate&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first stage prepares the knowledge base. The other stages happen when a user asks a question.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz4qyncbfcdiavoqetp27.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz4qyncbfcdiavoqetp27.png" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Indexing: prepare the knowledge
&lt;/h2&gt;

&lt;p&gt;Before anyone asks a question, the system needs to turn its source material into something searchable.&lt;/p&gt;

&lt;p&gt;A typical indexing pipeline looks like:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Load → Parse → Chunk → Embed → Store&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Load
&lt;/h3&gt;

&lt;p&gt;The source might be almost anything:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  PDFs&lt;/li&gt;
&lt;li&gt;  Word documents&lt;/li&gt;
&lt;li&gt;  web pages&lt;/li&gt;
&lt;li&gt;  database records&lt;/li&gt;
&lt;li&gt;  Confluence pages&lt;/li&gt;
&lt;li&gt;  Slack messages&lt;/li&gt;
&lt;li&gt;  support tickets&lt;/li&gt;
&lt;li&gt;  internal manuals&lt;/li&gt;
&lt;li&gt;  code or technical documentation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important point is that RAG does not require the knowledge to have originated as neatly structured text.&lt;/p&gt;

&lt;h3&gt;
  
  
  Parse and clean
&lt;/h3&gt;

&lt;p&gt;Raw documents are rarely ready for retrieval.&lt;/p&gt;

&lt;p&gt;A PDF may contain headers and footers repeated on every page. A webpage may contain navigation menus. A table may be extracted in the wrong order. OCR may introduce errors.&lt;/p&gt;

&lt;p&gt;If the parser turns a useful document into poor text, every later stage inherits the damage.&lt;/p&gt;

&lt;p&gt;This is one of the least glamorous parts of RAG and one of the most consequential:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Garbage in becomes searchable garbage.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ingestion also has to deal with practical questions such as&lt;br&gt;
deduplication, deletions, document versions, permissions, and what happens when the embedding model changes. These are not separate from RAG quality; they are part of it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Chunking
&lt;/h3&gt;

&lt;p&gt;A 200-page document is rarely a useful retrieval unit.&lt;/p&gt;

&lt;p&gt;Instead, it is divided into smaller passages called &lt;strong&gt;chunks&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The goal is to find a useful balance.&lt;/p&gt;

&lt;p&gt;Chunks that are too small can lose the context needed to understand a statement. Chunks that are too large contain too much unrelated information and become less precise retrieval units.&lt;/p&gt;

&lt;p&gt;A common starting point is a few hundred tokens with some overlap, but there is no universally correct chunk size. A legal contract, a support ticket, a code file, and a research paper may all want different structures.&lt;/p&gt;

&lt;p&gt;Overlap exists for a simple reason: boundaries are artificial.&lt;/p&gt;

&lt;p&gt;If a sentence begins at the end of one chunk and finishes at the&lt;br&gt;
beginning of another, retrieving either chunk alone could leave the model with an incomplete idea.&lt;/p&gt;

&lt;p&gt;More advanced approaches include &lt;strong&gt;parent-child retrieval&lt;/strong&gt;, where small child chunks are indexed for precise retrieval but a larger parent passage is supplied to the model after a match is found.&lt;/p&gt;

&lt;p&gt;The important principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Chunk for retrieval, not merely for storage.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Embeddings
&lt;/h3&gt;

&lt;p&gt;Once chunks exist, the system needs a way to represent their meaning numerically.&lt;/p&gt;

&lt;p&gt;An &lt;strong&gt;embedding&lt;/strong&gt; is a vector: a sequence of numbers produced by an embedding model. Text with related meanings should generally occupy nearby regions of the embedding space.&lt;/p&gt;

&lt;p&gt;For example, a user might ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Can I work from home?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;while the document says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Employees may work remotely up to three days per week."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The wording is different, but the meaning is related. Dense semantic retrieval can recognize that relationship.&lt;/p&gt;

&lt;p&gt;The embedding model is not the same thing as the generative LLM. The embedding model creates representations for search. The LLM reads context and generates language.&lt;/p&gt;

&lt;p&gt;One important engineering rule follows from this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The query and the indexed documents need compatible&lt;br&gt;
representations.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If documents were embedded using one model and queries are embedded using an incompatible model, the vectors are not necessarily meaningful to compare.&lt;/p&gt;

&lt;h3&gt;
  
  
  Dense and sparse representations
&lt;/h3&gt;

&lt;p&gt;Dense embeddings are excellent at semantic similarity, but they are not the only useful representation.&lt;/p&gt;

&lt;p&gt;Sparse retrieval methods such as BM25 are strongly tied to terms&lt;br&gt;
appearing in the text. That makes them particularly useful for exact names, product codes, error messages, rare technical identifiers, and other cases where the literal words matter.&lt;/p&gt;

&lt;p&gt;Dense and sparse retrieval therefore have complementary strengths.&lt;/p&gt;

&lt;p&gt;That is why production systems often combine them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens inside vector search?
&lt;/h2&gt;

&lt;p&gt;Once chunks have embeddings, they need to live somewhere that can search them efficiently.&lt;/p&gt;

&lt;p&gt;A vector database or vector index stores the vector alongside the&lt;br&gt;
original text and useful metadata such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  source&lt;/li&gt;
&lt;li&gt;  author&lt;/li&gt;
&lt;li&gt;  department&lt;/li&gt;
&lt;li&gt;  date&lt;/li&gt;
&lt;li&gt;  document version&lt;/li&gt;
&lt;li&gt;  access permissions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When the user asks a question, the query is embedded into the same representation space and compared with stored vectors.&lt;/p&gt;

&lt;p&gt;A common similarity measure is &lt;strong&gt;cosine similarity&lt;/strong&gt;, which focuses on the angle between vectors. Dot product is another common option, while L2 distance measures geometric distance directly. Normalization matters because it changes the relationship between these measures.&lt;/p&gt;

&lt;p&gt;The important idea is less about memorizing one formula than&lt;br&gt;
understanding what the system is doing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;It is trying to find stored representations that are close to the query under the chosen similarity function.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Searching every vector exactly becomes expensive as the corpus grows.&lt;br&gt;
&lt;strong&gt;Approximate nearest-neighbor (ANN)&lt;/strong&gt; methods trade some exactness for much faster search.&lt;/p&gt;

&lt;p&gt;The important word is &lt;em&gt;approximate&lt;/em&gt;. Exact nearest-neighbor search asks the system to compare the query against every candidate and identify the true closest vectors. That can become too expensive at large scale.&lt;br&gt;
ANN methods instead organize the search space so that the system can explore a much smaller set of promising candidates.&lt;/p&gt;

&lt;h3&gt;
  
  
  HNSW: navigating a graph
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;HNSW (Hierarchical Navigable Small World)&lt;/strong&gt; is a popular graph-based ANN method.&lt;/p&gt;

&lt;p&gt;A useful intuition is a city map with several levels. The upper levels contain a small number of long connections that let you move quickly across the map. Lower levels contain increasingly detailed local connections.&lt;/p&gt;

&lt;p&gt;A query starts at a higher level, moves toward promising regions, then descends through the hierarchy until it is navigating the detailed neighborhood containing the closest candidates.&lt;/p&gt;

&lt;p&gt;The result is that the system does not need to examine every stored vector.&lt;/p&gt;

&lt;p&gt;HNSW also exposes an important engineering tradeoff through parameters such as &lt;strong&gt;&lt;code&gt;ef_search&lt;/code&gt;&lt;/strong&gt;. A larger search effort means the algorithm explores more candidate nodes. That can improve recall, but it also increases search work and therefore latency. A smaller value can make search faster while increasing the chance that a genuinely good neighbor is missed.&lt;/p&gt;

&lt;p&gt;So even within the same vector database, ANN quality is not simply “on” or “off.” You are choosing where to sit on a speed-versus-recall curve.&lt;/p&gt;

&lt;h3&gt;
  
  
  IVF: search the relevant regions
&lt;/h3&gt;

&lt;p&gt;Another common approach is &lt;strong&gt;inverted file indexing (IVF)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Instead of treating the entire vector space as one enormous search region, IVF partitions vectors into clusters. During indexing, each vector is assigned to an appropriate cluster.&lt;/p&gt;

&lt;p&gt;At query time, the system first determines which clusters are closest to the query, then searches only those regions.&lt;/p&gt;

&lt;p&gt;The important tuning concept is &lt;strong&gt;&lt;code&gt;nprobe&lt;/code&gt;&lt;/strong&gt;: how many clusters should the query inspect?&lt;/p&gt;

&lt;p&gt;A low &lt;code&gt;nprobe&lt;/code&gt; means less work and lower latency, but it can miss&lt;br&gt;
relevant vectors that live in a cluster the search never examined. A higher &lt;code&gt;nprobe&lt;/code&gt; searches more regions and can improve recall at the cost of additional computation.&lt;/p&gt;

&lt;p&gt;HNSW and IVF therefore use different structures to make the same broad tradeoff: avoid exhaustive search while keeping the probability of missing useful neighbors acceptably low.&lt;/p&gt;

&lt;h3&gt;
  
  
  Product quantization: make the vectors smaller
&lt;/h3&gt;

&lt;p&gt;There is another problem at scale: memory.&lt;/p&gt;

&lt;p&gt;A large corpus may contain hundreds of millions of vectors, each with hundreds or thousands of dimensions. Keeping every vector in full precision can become expensive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Product Quantization (PQ)&lt;/strong&gt; compresses vectors into smaller&lt;br&gt;
representations. Instead of storing every original value directly, the system represents portions of the vector using compact codes.&lt;/p&gt;

&lt;p&gt;The benefit is substantially lower memory usage and potentially faster search. The cost is that compression introduces approximation, which can reduce retrieval accuracy.&lt;/p&gt;

&lt;p&gt;This is another example of the same RAG engineering principle: you are trading one resource for another.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why ANN recall becomes an engineering concern
&lt;/h3&gt;

&lt;p&gt;It is tempting to describe ANN as simply “less accurate search.” That is too simplistic.&lt;/p&gt;

&lt;p&gt;The real issue is that as the corpus grows, maintaining very high recall under a fixed latency and memory budget becomes harder. The system has more possible neighbors to consider, more index structure to maintain, and more pressure to limit how much work each query performs.&lt;/p&gt;

&lt;p&gt;That is why ANN systems expose tuning parameters and why retrieval should be evaluated empirically.&lt;/p&gt;

&lt;p&gt;In practice, you care about questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How much recall@k do we lose at our latency target?&lt;/li&gt;
&lt;li&gt;What happens when the corpus grows tenfold?&lt;/li&gt;
&lt;li&gt;How much memory does the index require?&lt;/li&gt;
&lt;li&gt;Which parameters improve recall without making latency unacceptable?&lt;/li&gt;
&lt;li&gt;Does compression change the ranking enough to affect downstream answer
quality?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The useful mental model is therefore not “ANN is an approximation.”&lt;/p&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;ANN is a controlled tradeoff between search cost and the probability of finding the best neighbors.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Other ANN approaches include inverted-file methods and vector&lt;br&gt;
compression techniques such as product quantization. The specific choice depends on corpus size, latency requirements, available memory, and the recall you need.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi0rne7fzbj8smp3gk0qx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi0rne7fzbj8smp3gk0qx.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Metadata filtering is another important capability. You may want&lt;br&gt;
semantic similarity &lt;strong&gt;only among documents from the legal department&lt;/strong&gt;, or &lt;strong&gt;only among documents newer than a certain date&lt;/strong&gt;, or &lt;strong&gt;only among documents the current user is allowed to see&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That last requirement becomes critical in multi-tenant and&lt;br&gt;
permission-sensitive systems. A retrieval system that finds the right answer but exposes information the user is not authorized to see is still a broken system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval: why "just use vector search" is not enough
&lt;/h2&gt;

&lt;p&gt;A basic RAG prototype might embed the query, perform a similarity&lt;br&gt;
search, take the top five chunks, and send them to the LLM.&lt;/p&gt;

&lt;p&gt;Sometimes that works beautifully.&lt;/p&gt;

&lt;p&gt;Sometimes it fails in ways that are difficult to diagnose.&lt;/p&gt;

&lt;p&gt;There are several recurring reasons.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vocabulary mismatch
&lt;/h3&gt;

&lt;p&gt;A user may say "heart attack" while the document says "myocardial&lt;br&gt;
infarction."&lt;/p&gt;

&lt;p&gt;Semantic retrieval can help bridge that gap, but no retrieval method is perfect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Exact terms matter too
&lt;/h2&gt;

&lt;p&gt;Now imagine the query contains a product ID such as &lt;code&gt;XR-4729&lt;/code&gt;, a legal clause number, or a specific error code.&lt;/p&gt;

&lt;p&gt;Semantic similarity is not necessarily the best tool for that.&lt;/p&gt;

&lt;p&gt;This is where sparse keyword retrieval can be extremely valuable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Query and document lengths differ
&lt;/h3&gt;

&lt;p&gt;A user query might be six words while the retrieved chunk is several hundred words. Comparing the two is not exactly the same problem as comparing two similarly sized documents.&lt;/p&gt;

&lt;h3&gt;
  
  
  Redundancy
&lt;/h3&gt;

&lt;p&gt;The top ten results may all say essentially the same thing.&lt;/p&gt;

&lt;p&gt;Sending ten versions of one idea to the LLM does not necessarily make the answer better.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hybrid retrieval: combine complementary searches
&lt;/h2&gt;

&lt;p&gt;A common production baseline is &lt;strong&gt;hybrid search&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Run a dense semantic search and a sparse lexical search in parallel, then combine their rankings.&lt;/p&gt;

&lt;p&gt;BM25 is a classic sparse retrieval method. It rewards useful term&lt;br&gt;
matches while accounting for how common a term is across the collection.&lt;/p&gt;

&lt;p&gt;Dense retrieval contributes semantic matching.&lt;/p&gt;

&lt;p&gt;BM25 contributes exact lexical matching.&lt;/p&gt;

&lt;p&gt;The results can be combined using a ranking method such as &lt;strong&gt;Reciprocal Rank Fusion (RRF)&lt;/strong&gt;, where a document receives contributions based on where it appeared in each ranked list.&lt;/p&gt;

&lt;p&gt;The important insight is not the formula itself. It is that two&lt;br&gt;
imperfect retrieval mechanisms can complement one another.&lt;/p&gt;

&lt;h2&gt;
  
  
  Query rewriting and decomposition
&lt;/h2&gt;

&lt;p&gt;The user's raw question is not always the best search query.&lt;/p&gt;

&lt;p&gt;A conversational question such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What about the policy for contractors?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;may depend on several earlier turns.&lt;/p&gt;

&lt;p&gt;A query rewriting step can turn conversational language into a&lt;br&gt;
standalone retrieval query.&lt;/p&gt;

&lt;p&gt;More complex questions can also be &lt;strong&gt;decomposed&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Suppose someone asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How did our Q3 revenue compare with the previous quarter, and what were the main reasons for the change?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That may require separate retrieval for the numbers and for the&lt;br&gt;
explanation.&lt;/p&gt;

&lt;p&gt;Instead of treating the whole thing as one search query, the system can create subqueries, retrieve evidence for each, and combine the results.&lt;/p&gt;

&lt;p&gt;This adds complexity and sometimes extra model calls, so it should be used when the query actually benefits from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  HyDE and multi-query retrieval
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;HyDE&lt;/strong&gt;, or Hypothetical Document Embeddings, takes another approach.&lt;/p&gt;

&lt;p&gt;Instead of embedding the user's short question directly, the system first asks an LLM to generate a hypothetical answer or document. That hypothetical text is then embedded and used for retrieval.&lt;/p&gt;

&lt;p&gt;The hypothetical text does not have to be factually correct. Its purpose is to provide a richer representation of what a relevant document might look like.&lt;/p&gt;

&lt;p&gt;The tradeoff is straightforward: another model call means additional latency and cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-query retrieval&lt;/strong&gt; is simpler conceptually. The system generates several alternative formulations of the same question, retrieves for each, merges and deduplicates the results, and then ranks the candidates.&lt;/p&gt;

&lt;p&gt;These techniques are useful when the original query is a poor&lt;br&gt;
representation of the information the system actually needs.&lt;/p&gt;

&lt;p&gt;They are not mandatory ingredients of every RAG system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reranking: retrieve broadly, then judge carefully
&lt;/h2&gt;

&lt;p&gt;One of the most useful patterns in practical RAG is to separate&lt;br&gt;
&lt;strong&gt;candidate generation&lt;/strong&gt; from &lt;strong&gt;precise ranking&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The first retrieval stage should be fast. It might return 20, 30, or 50 candidates.&lt;/p&gt;

&lt;p&gt;Then a more expensive &lt;strong&gt;cross-encoder reranker&lt;/strong&gt; examines the query and each candidate together and scores how well they match.&lt;/p&gt;

&lt;p&gt;This works because the two stages have different jobs.&lt;/p&gt;

&lt;p&gt;A bi-encoder or vector search system is efficient because the document representations can be computed ahead of time.&lt;/p&gt;

&lt;p&gt;A cross-encoder can make a more detailed query-document comparison, but doing that against millions of chunks would be too expensive.&lt;/p&gt;

&lt;p&gt;So the architecture becomes:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;large corpus → fast retrieval → small candidate set → expensive&lt;br&gt;
reranking → tiny final context&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is one of the most important patterns to understand in RAG&lt;br&gt;
engineering.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dm26t4v0oisjcinj59d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dm26t4v0oisjcinj59d.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Diversity matters too
&lt;/h3&gt;

&lt;p&gt;Sometimes the five highest-ranked chunks all repeat the same&lt;br&gt;
information.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Maximal Marginal Relevance (MMR)&lt;/strong&gt; provides a way to balance relevance with diversity.&lt;/p&gt;

&lt;p&gt;Conceptually, MMR rewards a chunk for being relevant to the query while penalizing it for being too similar to chunks already selected.&lt;/p&gt;

&lt;p&gt;The result is a final context containing several complementary pieces of evidence instead of five near-duplicates.&lt;/p&gt;

&lt;p&gt;The broader lesson is important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Retrieval is not about finding the most similar text. It is about constructing the most useful evidence set.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The generation stage: context is evidence, not decoration
&lt;/h2&gt;

&lt;p&gt;Once the system has retrieved its evidence, it needs to assemble the prompt that the LLM will see.&lt;/p&gt;

&lt;p&gt;A useful RAG prompt typically contains:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; clear system instructions&lt;/li&gt;
&lt;li&gt; labelled source chunks&lt;/li&gt;
&lt;li&gt; grounding rules&lt;/li&gt;
&lt;li&gt; the user's question&lt;/li&gt;
&lt;li&gt; a defined output format when necessary&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The structure matters because the model is now being asked to reason over evidence.&lt;/p&gt;

&lt;p&gt;A good instruction might say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Answer using only the provided context. If the answer is not supported by the context, say that the information is insufficient. Do not speculate beyond the sources.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That does not magically make hallucination impossible. But it gives the model a clear operating rule.&lt;/p&gt;

&lt;h3&gt;
  
  
  More context is not automatically better
&lt;/h3&gt;

&lt;p&gt;It is tempting to think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"If five chunks are useful, twenty must be better."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Usually, it is not that simple.&lt;/p&gt;

&lt;p&gt;Every additional chunk increases token usage and latency. Irrelevant chunks can introduce noise. Conflicting chunks can make the model uncertain. And long contexts can create attention problems.&lt;/p&gt;

&lt;p&gt;This is related to the &lt;strong&gt;lost-in-the-middle&lt;/strong&gt; phenomenon: when relevant information is buried among a long sequence of context, models may use information at the beginning and end more effectively than information buried in the middle.&lt;/p&gt;

&lt;p&gt;A practical strategy is therefore:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Retrieve more, send less.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Retrieve a broad candidate set. Rerank it. Select a small, high-quality context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe0jpc6z5zxt4hmlvf7o2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe0jpc6z5zxt4hmlvf7o2.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Citations
&lt;/h3&gt;

&lt;p&gt;One of RAG's practical advantages is that the system knows which sources it retrieved.&lt;/p&gt;

&lt;p&gt;Citations can be rendered inline:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Employees may work remotely up to three days per week [Source 1].&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Or as a source list:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Sources: HR Policy, page 4; Manager Handbook, page 12.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In production, structured citation data can be even more useful. The system can return an answer together with source IDs, pages, claims, and links so the interface can make evidence clickable and auditable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when the answer isn't in the documents?
&lt;/h2&gt;

&lt;p&gt;A trustworthy RAG system needs an explicit answer to this question.&lt;/p&gt;

&lt;p&gt;If the retriever finds nothing relevant, the model should not feel obligated to invent an answer.&lt;/p&gt;

&lt;p&gt;A simple design is to combine grounding instructions with retrieval thresholds. If none of the retrieved candidates passes an appropriate relevance threshold, the system can decline to answer from the knowledge base.&lt;/p&gt;

&lt;p&gt;The exact threshold is application-specific. A score from one embedding system cannot automatically be treated as a universal measure of relevance.&lt;/p&gt;

&lt;p&gt;The deeper principle is more important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A good RAG system must have a graceful failure mode.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;"I don't have enough information in the retrieved sources" is often much better than a confident fabrication.&lt;/p&gt;

&lt;h3&gt;
  
  
  Conversational RAG: the query is not always the query
&lt;/h3&gt;

&lt;p&gt;Chat makes retrieval harder.&lt;/p&gt;

&lt;p&gt;A user might say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What's the refund policy?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What about digital products?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"And if I used a gift card?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The final message is not a complete retrieval query.&lt;/p&gt;

&lt;p&gt;A common solution is &lt;strong&gt;query condensation&lt;/strong&gt;: use the conversation&lt;br&gt;
history to rewrite the latest message into a standalone question before retrieval.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What is the refund policy for digital products purchased with a gift&lt;br&gt;
card?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now the retrieval system has the information it needs to search&lt;br&gt;
effectively.&lt;/p&gt;

&lt;p&gt;This adds another processing step, but it solves a fundamental mismatch between how humans converse and how retrieval systems search.&lt;/p&gt;

&lt;h1&gt;
  
  
  When RAG gives a bad answer: debug the pipeline, not just the model
&lt;/h1&gt;

&lt;p&gt;One of the most useful ways to think about RAG is that a bad answer does not necessarily mean the LLM is the problem.&lt;/p&gt;

&lt;p&gt;A failure can happen at several layers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ingestion failure&lt;/strong&gt;\&lt;br&gt;
The relevant information was parsed incorrectly, omitted, duplicated, or indexed incorrectly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Chunking failure&lt;/strong&gt;\&lt;br&gt;
The information exists, but it was divided into retrieval units that destroyed the useful context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval failure&lt;/strong&gt;\&lt;br&gt;
The right information exists in the index but was never retrieved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ranking failure&lt;/strong&gt;\&lt;br&gt;
The right information was retrieved but buried below less useful&lt;br&gt;
candidates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context assembly failure&lt;/strong&gt;\&lt;br&gt;
The correct evidence was found but was poorly ordered, excessively long, or mixed with conflicting material.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generation failure&lt;/strong&gt;\&lt;br&gt;
The model received useful evidence but failed to follow it or answer the question correctly.&lt;/p&gt;

&lt;p&gt;This gives you a practical debugging order:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Check the source → check the retrieved evidence → check the &lt;br&gt;
ranking → check the assembled context → check the generation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Do not immediately swap the LLM when the real problem is retrieval.&lt;/p&gt;

&lt;h1&gt;
  
  
  How do you know whether a RAG system actually works?
&lt;/h1&gt;

&lt;p&gt;A RAG system can produce answers that sound excellent while failing underneath.&lt;/p&gt;

&lt;p&gt;Evaluation needs to separate the pipeline into understandable&lt;br&gt;
relationships.&lt;/p&gt;

&lt;h2&gt;
  
  
  The RAG triad
&lt;/h2&gt;

&lt;p&gt;Three questions form a useful mental model:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Context relevance
&lt;/h3&gt;

&lt;p&gt;Did the retriever find information that was actually related to the question?&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Faithfulness
&lt;/h3&gt;

&lt;p&gt;Are the claims in the generated answer supported by the retrieved&lt;br&gt;
context?&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Answer relevance
&lt;/h3&gt;

&lt;p&gt;Does the answer actually address what the user asked?&lt;/p&gt;

&lt;p&gt;These dimensions catch different failures.&lt;/p&gt;

&lt;p&gt;A system can retrieve excellent evidence and still hallucinate.&lt;/p&gt;

&lt;p&gt;It can produce a faithful answer that does not answer the user's&lt;br&gt;
question.&lt;/p&gt;

&lt;p&gt;It can generate a relevant answer because the model already knew the topic, even though retrieval failed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxuyzw9tb9jotrtmlerj5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxuyzw9tb9jotrtmlerj5.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Retrieval metrics
&lt;/h3&gt;

&lt;p&gt;The RAG triad is useful for end-to-end reasoning, but retrieval also benefits from traditional information-retrieval metrics.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recall@k&lt;/strong&gt; asks whether the relevant evidence appears somewhere in the top &lt;em&gt;k&lt;/em&gt; results.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MRR (Mean Reciprocal Rank)&lt;/strong&gt; cares about how early the first relevant result appears.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;nDCG&lt;/strong&gt; evaluates ranked results while giving more credit to highly relevant items near the top.&lt;/p&gt;

&lt;p&gt;These metrics answer a different question from faithfulness. They tell you whether the retrieval system is finding and ordering useful&lt;br&gt;
evidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  RAGAS and LLM-as-judge
&lt;/h3&gt;

&lt;p&gt;Frameworks such as RAGAS operationalize several RAG evaluation concepts, often using an LLM as a judge.&lt;/p&gt;

&lt;p&gt;One useful technique is to break an answer into atomic claims and verify those claims against the retrieved context.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Claim 1 → supported\&lt;br&gt;
Claim 2 → supported\&lt;br&gt;
Claim 3 → unsupported&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That gives a more meaningful faithfulness assessment than asking whether an answer "looks good."&lt;/p&gt;

&lt;p&gt;LLM-as-judge is powerful because it scales, but it is not an&lt;br&gt;
unquestionable authority. Judges can have biases, prefer verbose&lt;br&gt;
answers, or make mistakes of their own.&lt;/p&gt;

&lt;p&gt;A strong evaluation system therefore combines automated evaluation with carefully constructed golden datasets and periodic human validation.&lt;/p&gt;

&lt;p&gt;A practical test set can combine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  expert-written questions&lt;/li&gt;
&lt;li&gt;  synthetic questions generated from the corpus&lt;/li&gt;
&lt;li&gt;  sampled real production queries&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important part is establishing a baseline and using it to detect regressions when the RAG pipeline changes.&lt;/p&gt;

&lt;h1&gt;
  
  
  When basic RAG isn't enough
&lt;/h1&gt;

&lt;p&gt;Basic RAG makes a strong assumption:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Retrieve something, use it, and generate an answer.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That assumption breaks in several ways.&lt;/p&gt;

&lt;p&gt;Sometimes the question does not require retrieval.&lt;br&gt;
Sometimes the retrieved information is poor.&lt;br&gt;
Sometimes answering requires relationships across many documents.&lt;/p&gt;

&lt;p&gt;Sometimes the system needs to decide what tool or retrieval strategy to use next.&lt;/p&gt;

&lt;p&gt;Several advanced RAG patterns address these problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-RAG: should I retrieve, and did retrieval help?
&lt;/h2&gt;

&lt;p&gt;Basic RAG assumes retrieval should happen for every query.&lt;/p&gt;

&lt;p&gt;That is wasteful for questions that the model can answer directly, and it can actively hurt when retrieval returns poor evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Self-RAG&lt;/strong&gt; changes that assumption by giving the model a role in deciding when retrieval is needed and in evaluating whether retrieved information is useful.&lt;/p&gt;

&lt;p&gt;The original approach uses special reflection signals generated by a model that has been trained for this behavior. In other words, the reflection mechanism is part of the model's learned generation process; it is not simply a separate classifier bolted onto an ordinary LLM.&lt;/p&gt;

&lt;p&gt;Conceptually, the loop looks like:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;question → decide whether retrieval is useful → retrieve → assess the evidence → generate&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This makes Self-RAG attractive for systems that receive a mixture of questions: some needing external knowledge and others that do not.&lt;/p&gt;

&lt;p&gt;The tradeoff is important, though. The approach requires a model capable of the specialized reflection behavior, which makes it more involved than simply adding another retrieval step to an existing RAG pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Corrective RAG: don't trust the retriever blindly
&lt;/h2&gt;

&lt;p&gt;A standard RAG pipeline can make a dangerous assumption:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If the retriever returned something, it must be useful.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Corrective RAG (CRAG)&lt;/strong&gt; explicitly challenges that assumption.&lt;/p&gt;

&lt;p&gt;CRAG introduces a quality assessment around retrieved documents. If the retrieved material is judged weak, the system can take corrective action instead of blindly passing the results to the generator.&lt;/p&gt;

&lt;p&gt;One part of the original CRAG idea is &lt;strong&gt;knowledge refinement&lt;/strong&gt;. Retrieved chunks can be broken into smaller units, irrelevant material can be discarded, and the remaining high-signal information can be recombined before generation.&lt;/p&gt;

&lt;p&gt;The conceptual shift is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Basic RAG:&lt;/strong&gt; retrieve → trust → generate&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CRAG:&lt;/strong&gt; retrieve → evaluate → refine/correct → generate&lt;/p&gt;

&lt;p&gt;This is particularly useful when document quality is uneven or the retrieval system cannot be assumed to return reliable evidence every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  GraphRAG: when relationships matter
&lt;/h2&gt;

&lt;p&gt;Vector search answers a question like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Which passages are most similar to this query?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is powerful, but some questions are fundamentally about&lt;br&gt;
&lt;strong&gt;relationships&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Imagine a question that requires connecting an employee to a department, that department to a project, and that project to a supplier. The useful information may be spread across many documents and may not appear as one semantically similar passage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GraphRAG&lt;/strong&gt; addresses this by representing entities and relationships explicitly.&lt;/p&gt;

&lt;p&gt;A typical graph-oriented pipeline can involve:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;extracting entities from source documents&lt;/li&gt;
&lt;li&gt;identifying relationships between those entities&lt;/li&gt;
&lt;li&gt;constructing a knowledge graph&lt;/li&gt;
&lt;li&gt;detecting communities or groups of closely connected entities&lt;/li&gt;
&lt;li&gt;generating summaries of those communities&lt;/li&gt;
&lt;li&gt;retrieving through graph structure at query time&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This creates two useful modes of reasoning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Local search&lt;/strong&gt; focuses on a particular entity and its surrounding neighborhood.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Global search&lt;/strong&gt; can reason over broader themes represented by&lt;br&gt;
community-level summaries.&lt;/p&gt;

&lt;p&gt;GraphRAG therefore becomes attractive when the question depends on connections rather than simply semantic similarity: organizational structures, citation networks, supply chains, interconnected technical systems, and similar domains.&lt;/p&gt;

&lt;p&gt;The cost is substantial. Building and maintaining the graph can require many extraction steps and additional LLM calls, and graph-based infrastructure is more complex than a straightforward vector index.&lt;/p&gt;

&lt;p&gt;The important lesson is not that graphs replace vectors.&lt;/p&gt;

&lt;p&gt;It is that &lt;strong&gt;the retrieval structure should match the structure of the questions you need to answer.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Agentic RAG: what should the system do next?
&lt;/h2&gt;

&lt;p&gt;A conventional RAG pipeline follows a predefined sequence.&lt;/p&gt;

&lt;p&gt;An &lt;strong&gt;agentic RAG&lt;/strong&gt; system gives the model more control over that sequence.&lt;/p&gt;

&lt;p&gt;Instead of always performing exactly one retrieval call, the system can plan and decide among tools or actions. Depending on the task, those tools might include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;vector search&lt;/li&gt;
&lt;li&gt;SQL&lt;/li&gt;
&lt;li&gt;web search&lt;/li&gt;
&lt;li&gt;APIs&lt;/li&gt;
&lt;li&gt;code execution&lt;/li&gt;
&lt;li&gt;graph traversal&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A complex question might therefore become a sequence of actions:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;understand the question → retrieve internal data → identify a missing piece → search another source → calculate a result → retrieve supporting evidence → synthesize the answer&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is useful when the problem is inherently multi-step or the required information lives in heterogeneous systems.&lt;/p&gt;

&lt;p&gt;But agency introduces costs.&lt;/p&gt;

&lt;p&gt;More tool calls mean more latency. More branches make behavior harder to predict. More moving parts make debugging harder. And an agent can make the wrong decision about what to do next.&lt;/p&gt;

&lt;p&gt;The goal is therefore not to make every RAG system agentic.&lt;/p&gt;

&lt;p&gt;It is to use agency when the problem actually requires adaptive,&lt;br&gt;
multi-step behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four different responses to four different problems
&lt;/h2&gt;

&lt;p&gt;The four patterns are easier to remember when you connect each one to the assumption it changes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Self-RAG:&lt;/strong&gt; Should I retrieve?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Corrective RAG:&lt;/strong&gt; Is what I retrieved useful enough to trust?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GraphRAG:&lt;/strong&gt; Do I need explicit relationships between pieces of knowledge?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic RAG:&lt;/strong&gt; What should I do next to solve this problem?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They can also be combined. A system might use an agent to decide whether to retrieve, use hybrid search to find candidates, apply corrective evaluation to those candidates, and use a graph when the question turns out to be relational.&lt;/p&gt;

&lt;p&gt;The point is not to collect advanced-RAG names.&lt;/p&gt;

&lt;p&gt;The point is to understand &lt;strong&gt;which limitation each architecture is trying to remove.&lt;/strong&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  RAG in production is a different problem
&lt;/h1&gt;

&lt;p&gt;A prototype can work beautifully on a laptop with a few documents.&lt;/p&gt;

&lt;p&gt;Production introduces a different set of constraints.&lt;/p&gt;

&lt;h3&gt;
  
  
  Freshness
&lt;/h3&gt;

&lt;p&gt;A knowledge base changes.&lt;/p&gt;

&lt;p&gt;Documents are edited. Policies are replaced. New tickets arrive. Old information becomes invalid.&lt;/p&gt;

&lt;p&gt;A production system therefore needs a strategy for freshness:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  periodic re-indexing&lt;/li&gt;
&lt;li&gt;  incremental indexing triggered by document changes&lt;/li&gt;
&lt;li&gt;  timestamps and document versions&lt;/li&gt;
&lt;li&gt;  temporal ranking when recency matters&lt;/li&gt;
&lt;li&gt;  live retrieval for especially fast-changing sources&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The system should know not just &lt;em&gt;what&lt;/em&gt; a chunk says, but often &lt;em&gt;when&lt;/em&gt; that information was valid.&lt;/p&gt;

&lt;h3&gt;
  
  
  Latency
&lt;/h3&gt;

&lt;p&gt;A RAG request may involve query rewriting, retrieval, reranking, prompt assembly, and generation.&lt;/p&gt;

&lt;p&gt;Every extra step adds latency.&lt;/p&gt;

&lt;p&gt;Useful techniques include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  streaming the generated answer&lt;/li&gt;
&lt;li&gt;  caching repeated or semantically similar requests&lt;/li&gt;
&lt;li&gt;  running independent retrieval operations in parallel&lt;/li&gt;
&lt;li&gt;  using smaller models for lightweight tasks such as query rewriting&lt;/li&gt;
&lt;li&gt;  reducing the amount of context sent to the generator&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Streaming is particularly important for perceived latency: users can begin reading while the model is still generating.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost
&lt;/h3&gt;

&lt;p&gt;Every token, model call, and embedding operation has a cost.&lt;/p&gt;

&lt;p&gt;The major levers are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  context size&lt;/li&gt;
&lt;li&gt;  model choice&lt;/li&gt;
&lt;li&gt;  number of LLM calls&lt;/li&gt;
&lt;li&gt;  cache hit rate&lt;/li&gt;
&lt;li&gt;  embedding infrastructure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A common mistake is using the most expensive model for every stage. A smaller model may be perfectly adequate for query rewriting or classification while a stronger model is reserved for the final answer.&lt;/p&gt;

&lt;p&gt;The same principle applies to retrieval. Spending more computation on every candidate is wasteful if a fast first-stage search can narrow millions of chunks to a few dozen.&lt;/p&gt;

&lt;h3&gt;
  
  
  Security: the documents themselves can attack the system
&lt;/h3&gt;

&lt;p&gt;Prompt injection in RAG has an unusual property.&lt;/p&gt;

&lt;p&gt;The attacker does not necessarily need to control the user's question.&lt;/p&gt;

&lt;p&gt;They may control a &lt;strong&gt;document that gets retrieved&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Imagine someone uploads a malicious support ticket containing&lt;br&gt;
instructions aimed at the model. If that ticket is indexed and later retrieved, its contents are inserted into the context automatically.&lt;/p&gt;

&lt;p&gt;The user asking the next question may be completely innocent.&lt;/p&gt;

&lt;p&gt;This makes document trust a security boundary.&lt;/p&gt;

&lt;p&gt;Defenses should be layered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  index content only from appropriate sources&lt;/li&gt;
&lt;li&gt;  treat user-uploaded documents as untrusted&lt;/li&gt;
&lt;li&gt;  isolate untrusted content where appropriate&lt;/li&gt;
&lt;li&gt;  preserve a clear instruction hierarchy&lt;/li&gt;
&lt;li&gt;  monitor outputs for anomalous behavior&lt;/li&gt;
&lt;li&gt;  enforce authorization before retrieval, not only after generation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Permission-aware retrieval is particularly important in multi-tenant systems.&lt;/p&gt;

&lt;p&gt;A cache must also respect tenant and user boundaries. A perfectly cached answer is a security failure if it is returned to someone who was never authorized to see the underlying information.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multimodal RAG
&lt;/h2&gt;

&lt;p&gt;Real knowledge bases are not made entirely of paragraphs.&lt;/p&gt;

&lt;p&gt;PDFs contain tables and charts. Wikis contain images. Technical&lt;br&gt;
documentation contains diagrams.&lt;/p&gt;

&lt;p&gt;There are several ways to handle this.&lt;/p&gt;

&lt;p&gt;One approach is to extract or transcribe visual content into text. It is simple, but potentially lossy.&lt;/p&gt;

&lt;p&gt;Another is to use multimodal embeddings so visual and textual content can be represented in a compatible space.&lt;/p&gt;

&lt;p&gt;A third approach is to use a vision-language model during ingestion to generate rich descriptions of tables and images, then index those descriptions using standard text retrieval.&lt;/p&gt;

&lt;p&gt;The practical choice depends on the application. The important point is that &lt;strong&gt;retrieval quality depends on preserving the information that matters&lt;/strong&gt;, not merely converting everything into plain text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observability: if you can't see the pipeline, you can't debug it
&lt;/h2&gt;

&lt;p&gt;A production RAG system should make its decisions inspectable.&lt;/p&gt;

&lt;p&gt;For each query, useful telemetry can include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  the query after privacy/PII handling&lt;/li&gt;
&lt;li&gt;  retrieved chunk IDs&lt;/li&gt;
&lt;li&gt;  retrieval scores&lt;/li&gt;
&lt;li&gt;  reranking scores&lt;/li&gt;
&lt;li&gt;  prompt or prompt hash&lt;/li&gt;
&lt;li&gt;  generated answer&lt;/li&gt;
&lt;li&gt;  latency by stage&lt;/li&gt;
&lt;li&gt;  token counts&lt;/li&gt;
&lt;li&gt;  evaluation signals&lt;/li&gt;
&lt;li&gt;  user feedback&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tools such as LangSmith, TruLens, Weights &amp;amp; Biases, Arize AI, or a custom data warehouse can provide this visibility.&lt;/p&gt;

&lt;p&gt;The goal is not to collect logs for their own sake.&lt;/p&gt;

&lt;p&gt;It is to make questions like this answerable:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Show me the queries this week where faithfulness dropped below &lt;br&gt;
our baseline and tell me which retrieved chunks were involved."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That turns RAG debugging from guesswork into investigation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmpio4y7tzdlzfykjq1o5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmpio4y7tzdlzfykjq1o5.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  RAG is not always the right answer
&lt;/h2&gt;

&lt;p&gt;Understanding when &lt;strong&gt;not&lt;/strong&gt; to use RAG is just as important as knowing how to build it.&lt;/p&gt;

&lt;p&gt;If the complete knowledge base is ten pages and comfortably fits in context, retrieval may add unnecessary complexity.&lt;/p&gt;

&lt;p&gt;If the task is purely creative, retrieval may add noise rather than useful evidence.&lt;/p&gt;

&lt;p&gt;If the primary source is structured data, a direct database query may be better. Asking "What was Q3 revenue in APAC?" is fundamentally different from asking a question about an unstructured policy document.&lt;/p&gt;

&lt;p&gt;And if retrieval quality is consistently poor, adding more retrieval machinery does not automatically solve the problem. Sometimes the right answer is to improve the source data or rethink the architecture.&lt;/p&gt;

&lt;p&gt;The correct question is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Should I use RAG?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Where does the information needed for this task live, how often does it change, and what is the best way to give the model reliable access to it?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Designing a RAG system from requirements
&lt;/h2&gt;

&lt;p&gt;When designing a RAG system, it is tempting to immediately draw a vector database and an LLM.&lt;/p&gt;

&lt;p&gt;Start somewhere else.&lt;/p&gt;

&lt;p&gt;Start with the requirements.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Understand the use case
&lt;/h3&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  What kinds of documents are involved?&lt;/li&gt;
&lt;li&gt;  How large is the corpus?&lt;/li&gt;
&lt;li&gt;  How frequently does it change?&lt;/li&gt;
&lt;li&gt;  Are queries simple or multi-step?&lt;/li&gt;
&lt;li&gt;  Is the application single-turn or conversational?&lt;/li&gt;
&lt;li&gt;  What is the latency budget?&lt;/li&gt;
&lt;li&gt;  What happens if the answer is wrong?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last question changes the architecture dramatically.&lt;/p&gt;

&lt;p&gt;A wrong answer in a casual internal search tool is different from a wrong answer in a legal or medical workflow.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Design ingestion
&lt;/h3&gt;

&lt;p&gt;Choose parsing and chunking based on the actual document types.&lt;/p&gt;

&lt;p&gt;Think about deduplication, versions, deletions, metadata, permissions, and re-indexing.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Design retrieval
&lt;/h3&gt;

&lt;p&gt;Start with a simple baseline.&lt;/p&gt;

&lt;p&gt;Dense retrieval may be enough for some corpora.&lt;/p&gt;

&lt;p&gt;Hybrid retrieval is often a strong general-purpose starting point when exact terms and semantic matching both matter.&lt;/p&gt;

&lt;p&gt;Add reranking when precision matters.&lt;/p&gt;

&lt;p&gt;Add query rewriting or decomposition when the query itself is the&lt;br&gt;
bottleneck.&lt;/p&gt;

&lt;p&gt;Do not add every technique simply because it exists.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Design generation
&lt;/h3&gt;

&lt;p&gt;Choose the model based on the quality, latency, and cost requirements.&lt;/p&gt;

&lt;p&gt;Define grounding behavior.&lt;/p&gt;

&lt;p&gt;Decide how citations will work.&lt;/p&gt;

&lt;p&gt;Define what happens when evidence is missing.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Design evaluation and operations
&lt;/h3&gt;

&lt;p&gt;Decide how success will be measured before declaring the system&lt;br&gt;
finished.&lt;/p&gt;

&lt;p&gt;Then design for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  freshness&lt;/li&gt;
&lt;li&gt;  security&lt;/li&gt;
&lt;li&gt;  observability&lt;/li&gt;
&lt;li&gt;  caching&lt;/li&gt;
&lt;li&gt;  latency&lt;/li&gt;
&lt;li&gt;  cost&lt;/li&gt;
&lt;li&gt;  regression testing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The architecture should follow the requirements, not the other way around.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical example: designing RAG for a legal firm
&lt;/h2&gt;

&lt;p&gt;Imagine a legal firm wants an internal assistant that answers questions about contracts, case materials, policies, and research documents.&lt;/p&gt;

&lt;p&gt;The design immediately raises constraints.&lt;/p&gt;

&lt;p&gt;The documents may be long and structurally complex. Access permissions matter. Citations matter. Document versions matter. An incorrect answer can have serious consequences.&lt;/p&gt;

&lt;p&gt;That suggests several architectural choices.&lt;/p&gt;

&lt;p&gt;Documents need careful parsing and metadata.&lt;/p&gt;

&lt;p&gt;Chunking should preserve sections and surrounding legal context rather than blindly splitting every fixed number of tokens.&lt;/p&gt;

&lt;p&gt;Retrieval may benefit from hybrid search because legal questions can contain both semantic concepts and exact clause numbers, names, dates, and terminology.&lt;/p&gt;

&lt;p&gt;Reranking becomes valuable because returning the wrong clause can be worse than returning no clause.&lt;/p&gt;

&lt;p&gt;The generation layer should be strongly grounded and citation-oriented.&lt;/p&gt;

&lt;p&gt;Permission checks must happen before evidence reaches the model.&lt;/p&gt;

&lt;p&gt;Evaluation should include a carefully curated golden set, not only generic synthetic questions.&lt;/p&gt;

&lt;p&gt;And freshness/versioning cannot be an afterthought. A superseded&lt;br&gt;
contract clause should not silently compete with the current version.&lt;/p&gt;

&lt;p&gt;The interesting part of the design is not the component list.&lt;/p&gt;

&lt;p&gt;It is the reasoning:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Each architectural decision exists because of a property of &lt;br&gt;
the problem.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the transferable skill.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes when the system grows?
&lt;/h2&gt;

&lt;p&gt;Scale exposes different bottlenecks.&lt;/p&gt;

&lt;p&gt;At millions of queries, vector search throughput may become important.&lt;br&gt;
At hundreds of millions of chunks, index size and memory become serious concerns.&lt;/p&gt;

&lt;p&gt;ANN parameters can be tuned for the desired speed/recall balance. Vector compression can reduce memory usage at some accuracy cost. Large embedding pipelines need batching and parallel processing.&lt;/p&gt;

&lt;p&gt;And cost changes character at scale.&lt;/p&gt;

&lt;p&gt;At low traffic, an unnecessary model call is annoying.&lt;/p&gt;

&lt;p&gt;At very high traffic, the same unnecessary model call becomes a major line item.&lt;/p&gt;

&lt;p&gt;That is why caching, smaller models for simple tasks, efficient&lt;br&gt;
retrieval, and context reduction are not merely optimization tricks.&lt;br&gt;
They can determine whether the architecture is economically viable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The deeper lesson: RAG is a system of tradeoffs
&lt;/h2&gt;

&lt;p&gt;It is easy to collect RAG techniques as a list:&lt;/p&gt;

&lt;p&gt;Chunking.&lt;/p&gt;

&lt;p&gt;Embeddings.&lt;/p&gt;

&lt;p&gt;Vector databases.&lt;/p&gt;

&lt;p&gt;Hybrid search.&lt;/p&gt;

&lt;p&gt;Reranking.&lt;/p&gt;

&lt;p&gt;HyDE.&lt;/p&gt;

&lt;p&gt;MMR.&lt;/p&gt;

&lt;p&gt;GraphRAG.&lt;/p&gt;

&lt;p&gt;Agents.&lt;/p&gt;

&lt;p&gt;RAGAS.&lt;/p&gt;

&lt;p&gt;And so on.&lt;/p&gt;

&lt;p&gt;But knowing the names is not the same as understanding the system.&lt;/p&gt;

&lt;p&gt;The more useful mental model is that every decision changes a tradeoff.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Smaller chunks&lt;/strong&gt; can improve retrieval precision but may lose context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Larger chunks&lt;/strong&gt; preserve context but may dilute the relevant signal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More retrieved documents&lt;/strong&gt; can improve recall but increase noise and&lt;br&gt;
cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reranking&lt;/strong&gt; can improve precision but adds computation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Query rewriting&lt;/strong&gt; can improve search quality but adds latency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stronger grounding&lt;/strong&gt; can reduce unsupported claims but may increase "I don't know" responses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Larger models&lt;/strong&gt; can improve generation quality but increase cost and latency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More context&lt;/strong&gt; can provide more evidence but can also produce&lt;br&gt;
attention degradation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More sophisticated agents&lt;/strong&gt; can solve more complex problems but are harder to predict and debug.&lt;/p&gt;

&lt;p&gt;There is no magical configuration that wins everywhere.&lt;/p&gt;

&lt;p&gt;The engineering question is always:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What problem are we experiencing, and which change addresses &lt;br&gt;
that problem without creating a worse one somewhere else?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the difference between assembling a RAG demo and designing a RAG system.&lt;/p&gt;

&lt;h2&gt;
  
  
  RAG in one picture
&lt;/h2&gt;

&lt;p&gt;At the highest level, the whole system can be reduced to a simple idea.&lt;/p&gt;

&lt;p&gt;Documents become searchable representations.&lt;/p&gt;

&lt;p&gt;A user's question becomes a search request.&lt;/p&gt;

&lt;p&gt;Retrieval finds candidate evidence.&lt;/p&gt;

&lt;p&gt;Ranking decides which evidence matters most.&lt;/p&gt;

&lt;p&gt;The context is assembled carefully.&lt;/p&gt;

&lt;p&gt;The LLM reasons over that evidence.&lt;/p&gt;

&lt;p&gt;Evaluation checks whether the pipeline actually worked.&lt;/p&gt;

&lt;p&gt;Production infrastructure keeps it fresh, fast, secure, observable, and affordable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2sfs6ver0jyff26h98jh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2sfs6ver0jyff26h98jh.png" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Final takeaway
&lt;/h2&gt;

&lt;p&gt;RAG began with a simple observation:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A model does not have to contain all the knowledge it needs to answer a question. It needs a reliable way to access the right knowledge when the question arrives.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;From that idea follows the entire architecture.&lt;/p&gt;

&lt;p&gt;You need to ingest and parse the source material.&lt;/p&gt;

&lt;p&gt;You need to chunk it without destroying meaning.&lt;/p&gt;

&lt;p&gt;You need representations that make useful information searchable.&lt;/p&gt;

&lt;p&gt;You need retrieval that handles both semantic similarity and exact terms.&lt;/p&gt;

&lt;p&gt;You need ranking that separates good candidates from merely plausible ones.&lt;/p&gt;

&lt;p&gt;You need context assembly that gives the LLM enough evidence without drowning it.&lt;/p&gt;

&lt;p&gt;You need generation that stays grounded.&lt;/p&gt;

&lt;p&gt;You need evaluation that tells you whether the failure happened in retrieval or generation.&lt;/p&gt;

&lt;p&gt;And once the system becomes real, you need freshness, permissions, security, latency controls, caching, observability, and cost discipline.&lt;/p&gt;

&lt;p&gt;The most important lesson is therefore not a particular vector database, embedding model, reranker, or RAG framework.&lt;/p&gt;

&lt;p&gt;It is a way of thinking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When a RAG system fails, ask where the information was lost.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Was it lost during ingestion?&lt;/p&gt;

&lt;p&gt;During chunking?&lt;/p&gt;

&lt;p&gt;During retrieval?&lt;/p&gt;

&lt;p&gt;During ranking?&lt;/p&gt;

&lt;p&gt;During context assembly?&lt;/p&gt;

&lt;p&gt;Or during generation?&lt;/p&gt;

&lt;p&gt;Once you can answer that question, the enormous collection of techniques around RAG becomes much easier to understand.&lt;/p&gt;

&lt;p&gt;RAG stops looking like a bag of AI buzzwords and starts looking like what it really is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;an information pipeline designed to connect a reasoning model with the knowledge it needs.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>tutorial</category>
      <category>programming</category>
    </item>
    <item>
      <title>Multithreading in Java: a guide to actually understanding it</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Sun, 27 Sep 2026 22:56:58 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/multithreading-in-java-a-guide-to-actually-understanding-it-1pg1</link>
      <guid>https://dev.to/shogun_the_grt/multithreading-in-java-a-guide-to-actually-understanding-it-1pg1</guid>
      <description>&lt;p&gt;&lt;em&gt;The concepts that make concurrent code make sense, with enough real code that each one actually clicks.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq6d95omghgt5el1swvuz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq6d95omghgt5el1swvuz.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Multithreading has a specific texture that's different from most programming topics. You can know exactly what &lt;code&gt;synchronized&lt;/code&gt; does on paper and still get tripped up the moment you're asked to trace through what happens when two threads run the same three lines at the same time. A lot of the difficulty isn't the syntax, it's building the habit of imagining two (or ten) things happening at once, and noticing exactly where that stops being safe.&lt;/p&gt;

&lt;p&gt;This post goes through multithreading as a concept, using Java to express it, since Java's concurrency vocabulary (&lt;code&gt;synchronized&lt;/code&gt;, &lt;code&gt;volatile&lt;/code&gt;, the &lt;code&gt;java.util.concurrent&lt;/code&gt; package) is close to the vocabulary most languages end up reaching for anyway. It covers what a thread actually is, the lifecycle and the methods you'll actually use, race conditions, locks and monitors, the memory model, atomics, the executor framework, futures, concurrent collections, coordination utilities, deadlocks, virtual threads, and a way of reasoning through a stuck or misbehaving concurrent program when you're staring at one. The code targets Java 21, the current long-term support release, which matters here specifically because it's the release that made virtual threads a real, non-preview feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  What multithreading actually is
&lt;/h2&gt;

&lt;p&gt;Start simple. Multithreading means one program does more than one thing at the same time, by running several independent paths of execution instead of one. Each of those paths is a thread. Without threads, a program is a single line of instructions executing one after another; a slow file read blocks everything else, a slow network call blocks everything else, nothing overlaps. Threads let a program keep making progress on one piece of work while another piece is waiting on something slow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Process vs thread.&lt;/strong&gt; A process is a running program with its own private memory space, completely isolated from every other process. A thread is a path of execution &lt;em&gt;inside&lt;/em&gt; a process, and every thread in the same process shares that process's memory. That sharing is the entire reason threads exist (they're cheap to create and can pass data to each other just by reading and writing the same variables) and it's also the entire reason multithreading is hard: two threads reading and writing the same variable at the same time is where nearly every concurrency bug traces back to.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fufoz9trxaj38tmgq12fk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fufoz9trxaj38tmgq12fk.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each thread gets its own private call stack (its own local variables and method-call history) but reads and writes the same heap as every other thread in the process. That's the picture in the cover image: shared counter, private cutting boards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How the JVM actually runs a thread.&lt;/strong&gt; For an ordinary Java &lt;code&gt;Thread&lt;/code&gt; (called a "platform thread" from here on, to distinguish it from virtual threads, which work differently), the JVM maps it directly to a real operating system thread, one-to-one. The OS, not the JVM, decides when each thread actually gets CPU time and pays the cost of switching between them. That one-to-one mapping is simple and has worked for decades, but it means each thread carries real OS overhead (typically around a megabyte of stack memory, plus the cost of a context switch), which is why a server handling ten thousand concurrent connections by spinning up ten thousand platform threads runs into trouble. Virtual threads change exactly this mapping, not the language semantics around it, which is worth holding onto for now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Creating a thread: &lt;code&gt;Thread&lt;/code&gt; vs &lt;code&gt;Runnable&lt;/code&gt;.&lt;/strong&gt; Java gives you two ways to define what a thread should run.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// HelloThread.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;HelloThread&lt;/span&gt; &lt;span class="kd"&gt;extends&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nd"&gt;@Override&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Hello from "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getName&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// HelloTask.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;HelloTask&lt;/span&gt; &lt;span class="kd"&gt;implements&lt;/span&gt; &lt;span class="nc"&gt;Runnable&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nd"&gt;@Override&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Hello from "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getName&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// Option 1: extending Thread directly&lt;/span&gt;
        &lt;span class="nc"&gt;HelloThread&lt;/span&gt; &lt;span class="n"&gt;helloThread&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;HelloThread&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;helloThread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="c1"&gt;// Option 2: implementing Runnable, then handing it to a Thread&lt;/span&gt;
        &lt;span class="nc"&gt;HelloTask&lt;/span&gt; &lt;span class="n"&gt;task&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;HelloTask&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="c1"&gt;// Runnable also works fine as a lambda, since it's a single-method interface&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;lambdaThread&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
            &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Hello from "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getName&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
        &lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;lambdaThread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Runnable&lt;/code&gt; is the one to default to, and it's worth understanding why. Java doesn't allow multiple inheritance, so &lt;code&gt;HelloThread&lt;/code&gt;, which extends &lt;code&gt;Thread&lt;/code&gt;, can never extend anything else. &lt;code&gt;HelloTask&lt;/code&gt;, which implements &lt;code&gt;Runnable&lt;/code&gt;, stays free to extend whatever it needs to, and just as importantly, a &lt;code&gt;Runnable&lt;/code&gt; is a plain unit of work, not tied to any particular thread, which means the exact same &lt;code&gt;HelloTask&lt;/code&gt; object can be handed to a thread pool without any changes at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Thread lifecycle, states, and the methods that matter
&lt;/h2&gt;

&lt;p&gt;A thread moves through a small set of states, and Java exposes them directly through &lt;code&gt;Thread.getState()&lt;/code&gt;, which makes this an easy thing to check yourself if you're ever staring at a program that seems stuck.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk2d3oefqkhixwbmwjk3x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk2d3oefqkhixwbmwjk3x.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NEW&lt;/strong&gt;: the &lt;code&gt;Thread&lt;/code&gt; object exists but &lt;code&gt;start()&lt;/code&gt; hasn't been called yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RUNNABLE&lt;/strong&gt;: the thread is either actually running or ready and waiting for the OS scheduler to give it CPU time. Java doesn't distinguish these two as separate states, both count as RUNNABLE.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BLOCKED&lt;/strong&gt;: the thread wants a lock that another thread currently holds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WAITING&lt;/strong&gt;: the thread is waiting indefinitely for another thread to do something, typically via &lt;code&gt;wait()&lt;/code&gt;, &lt;code&gt;join()&lt;/code&gt;, or &lt;code&gt;LockSupport.park()&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TIMED_WAITING&lt;/strong&gt;: same as WAITING, but with a timeout, via &lt;code&gt;sleep(ms)&lt;/code&gt;, &lt;code&gt;wait(ms)&lt;/code&gt;, or &lt;code&gt;join(ms)&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TERMINATED&lt;/strong&gt;: &lt;code&gt;run()&lt;/code&gt; has finished, normally or because of an uncaught exception.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;start()&lt;/code&gt; vs &lt;code&gt;run()&lt;/code&gt; is worth getting exactly right, because it's an easy thing to get quietly wrong.&lt;/strong&gt; Calling &lt;code&gt;start()&lt;/code&gt; asks the JVM to create a new OS thread and have it execute &lt;code&gt;run()&lt;/code&gt; on that new thread. Calling &lt;code&gt;run()&lt;/code&gt; directly just calls it like any ordinary method, on the current thread, synchronously, with no new thread created at all. The compiler will not stop you from writing &lt;code&gt;myThread.run()&lt;/code&gt; instead of &lt;code&gt;myThread.start()&lt;/code&gt;, your code will compile and even run correctly in terms of output, and it will silently not be multithreaded at all.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// StartVsRun.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;StartVsRun&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
            &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"running on: "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getName&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
        &lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;    &lt;span class="c1"&gt;// prints "main" — ran on the calling thread, no new thread created&lt;/span&gt;
        &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;  &lt;span class="c1"&gt;// prints "Thread-0" (or similar) — actually runs on a new thread&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The methods worth knowing beyond &lt;code&gt;start()&lt;/code&gt;.&lt;/strong&gt; A handful of &lt;code&gt;Thread&lt;/code&gt; methods come up constantly once you're actually writing concurrent code, and knowing their exact behaviour, not just their names, is what makes the difference between code that looks right and code that actually is right.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;join()&lt;/code&gt; makes the calling thread wait until the target thread finishes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// JoinExample.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;JoinExample&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sleep&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// simulate slow work&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"worker finished its work"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="o"&gt;});&lt;/span&gt;

        &lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;join&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;          &lt;span class="c1"&gt;// main thread blocks here until worker finishes&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"worker is done, main can continue"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without that &lt;code&gt;join()&lt;/code&gt;, &lt;code&gt;main&lt;/code&gt; has no guarantee it runs after the worker, it might print first, might print after, the order is undefined. &lt;code&gt;join()&lt;/code&gt; also takes an optional timeout in milliseconds, after which the caller stops waiting regardless of whether the target thread finished.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;sleep(ms)&lt;/code&gt; is a static method that pauses the &lt;em&gt;current&lt;/em&gt; thread for at least the given duration. Two details are easy to miss here. First, it's static, so &lt;code&gt;someOtherThread.sleep(1000)&lt;/code&gt; doesn't pause &lt;code&gt;someOtherThread&lt;/code&gt;, it pauses whatever thread that line of code is actually running on, which trips people up more often than you'd expect. Second, and this one's genuinely important: &lt;code&gt;sleep()&lt;/code&gt; does not release any lock the thread currently holds. A thread sleeping inside a &lt;code&gt;synchronized&lt;/code&gt; block keeps that lock the entire time it's asleep, blocking every other thread that wants it, for the full duration.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;interrupt()&lt;/code&gt; is Java's only real mechanism for asking a thread to stop, and it's cooperative, not forceful. Calling &lt;code&gt;t.interrupt()&lt;/code&gt; doesn't stop the thread itself, it just sets an internal flag and, if the thread happens to be blocked in something like &lt;code&gt;sleep()&lt;/code&gt;, &lt;code&gt;wait()&lt;/code&gt;, or &lt;code&gt;join()&lt;/code&gt; at that moment, wakes it up with an &lt;code&gt;InterruptedException&lt;/code&gt;. If the thread is doing ordinary CPU work and never checks for the flag, &lt;code&gt;interrupt()&lt;/code&gt; does nothing at all. That's why a well-behaved long-running task periodically checks whether it's been asked to stop and exits cleanly if so:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// InterruptibleTask.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;InterruptibleTask&lt;/span&gt; &lt;span class="kd"&gt;implements&lt;/span&gt; &lt;span class="nc"&gt;Runnable&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nd"&gt;@Override&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(!&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;isInterrupted&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="c1"&gt;// do one unit of work&lt;/span&gt;
            &lt;span class="n"&gt;performOneStep&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"cleanly stopped"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;performOneStep&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// stand-in for real work&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;InterruptibleTask&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
        &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sleep&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;   &lt;span class="c1"&gt;// let it run for a second&lt;/span&gt;
        &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;        &lt;span class="c1"&gt;// ask it to stop&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no safe way to force a thread to stop from the outside in modern Java. &lt;code&gt;Thread.stop()&lt;/code&gt; exists but has been deprecated for a long time, because it can kill a thread mid-update and leave shared data in a half-modified, inconsistent state. Interruption is the sanctioned pattern precisely because it lets the thread finish its current unit of work and clean up before exiting, rather than being killed mid-sentence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Daemon threads&lt;/strong&gt;, set with &lt;code&gt;t.setDaemon(true)&lt;/code&gt; before calling &lt;code&gt;start()&lt;/code&gt;, are background threads the JVM won't wait for. A running JVM stays alive as long as at least one non-daemon thread is alive; the moment every remaining thread is a daemon, the JVM exits immediately, without letting those daemon threads finish whatever they were doing. Ordinary threads default to non-daemon, and this is a real, practical difference: use daemon threads for background housekeeping you're fine with being cut off abruptly, never for work that must complete. &lt;code&gt;yield()&lt;/code&gt; is the least useful of the bunch and worth knowing mainly so you don't overestimate it: it's a hint to the scheduler that the current thread is willing to let other threads of the same priority run, but the scheduler is completely free to ignore it, and most real code never needs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Race conditions: the counter that doesn't count
&lt;/h2&gt;

&lt;p&gt;Here's a small class with a counter, and a separate file that hammers it with ten threads, each incrementing it a thousand times:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Counter.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Counter&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;increment&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;++;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="nf"&gt;getCount&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.ArrayList&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.List&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Counter&lt;/span&gt; &lt;span class="n"&gt;counter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Counter&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="nc"&gt;List&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;threads&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ArrayList&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;();&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;counter&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;increment&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;});&lt;/span&gt;
            &lt;span class="n"&gt;threads&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;add&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
            &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;threads&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;join&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;

        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Final count: "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;counter&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getCount&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
        &lt;span class="c1"&gt;// Expected: 10000&lt;/span&gt;
        &lt;span class="c1"&gt;// Actual: something smaller, and different almost every run&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frlqytunn7fvk1lu1vkqq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frlqytunn7fvk1lu1vkqq.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You would expect 10,000. You reliably get something smaller, and a different smaller number nearly every time you run it. This is the entire concept of a race condition, and the reason is that &lt;code&gt;count++&lt;/code&gt; isn't one operation, it's three: read the current value, add one to it, write the new value back. It's worth being able to narrate the exact interleaving that breaks it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Thread A reads &lt;code&gt;count&lt;/code&gt;, gets 5.&lt;/li&gt;
&lt;li&gt;Thread A is paused by the OS scheduler before it writes anything back.&lt;/li&gt;
&lt;li&gt;Thread B reads &lt;code&gt;count&lt;/code&gt;, also gets 5 (A's increment hasn't been written yet).&lt;/li&gt;
&lt;li&gt;Thread B computes 6 and writes it.&lt;/li&gt;
&lt;li&gt;Thread A resumes, computes 6 from the value it read &lt;em&gt;earlier&lt;/em&gt;, and writes 6 too.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two increments happened. The counter only went up by one. Nothing crashed, nothing threw an exception, the code just silently produced the wrong answer, which is exactly what makes race conditions dangerous in real systems: they don't announce themselves. The next few sections fix this same &lt;code&gt;Counter&lt;/code&gt; class in three different ways.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;synchronized&lt;/code&gt;, monitors, and reentrancy
&lt;/h2&gt;

&lt;p&gt;Java's oldest fix for the counter problem is the &lt;code&gt;synchronized&lt;/code&gt; keyword. Every object in Java has an associated monitor (an intrinsic lock), and &lt;code&gt;synchronized&lt;/code&gt; makes a thread acquire that lock before entering a block of code, and release it on the way out, automatically, even if an exception is thrown inside.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Counter.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Counter&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;increment&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;++;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="nf"&gt;getCount&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the exact same &lt;code&gt;Main&lt;/code&gt; class from before, ten threads, a thousand increments each, against this version, and you get 10,000, every time. Only one thread can be inside a &lt;code&gt;synchronized&lt;/code&gt; method (or block) on the same object at once; every other thread that wants in has to wait for the lock to be free.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvzplk54mejge6183ejrc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvzplk54mejge6183ejrc.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can synchronize a whole method, as above, which locks on &lt;code&gt;this&lt;/code&gt; (the object instance), or synchronize a specific block, which lets you name exactly which object to lock on and keep the locked region as small as possible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Counter.java (block form instead of method form)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Counter&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;increment&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;++;&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="nf"&gt;getCount&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Smaller synchronized blocks are generally better, since anything inside the lock is a stretch of code where every other thread wanting that lock is stuck waiting, and keeping that stretch short keeps contention low.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Object-level vs class-level locking.&lt;/strong&gt; &lt;code&gt;synchronized&lt;/code&gt; on an instance method locks on &lt;code&gt;this&lt;/code&gt;, so it only blocks other threads calling synchronized instance methods on &lt;em&gt;that same object&lt;/em&gt;. Two different &lt;code&gt;Counter&lt;/code&gt; instances don't block each other at all. If you instead want to lock across every instance of a class, for a static counter shared globally, you synchronize on the class object itself, or mark a static method &lt;code&gt;synchronized&lt;/code&gt;, which does the same thing implicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// GlobalCounter.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;GlobalCounter&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="c1"&gt;// synchronized on GlobalCounter.class, shared by every call, from any instance&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;increment&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;++;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="nf"&gt;getCount&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Reentrancy&lt;/strong&gt; is the property that lets a thread that already holds a lock re-acquire the same lock without deadlocking itself, which matters constantly in practice because one synchronized method commonly calls another synchronized method on the same object.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// ReentrantExample.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ReentrantExample&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;outer&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"in outer(), about to call inner()"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;inner&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// the current thread already holds the lock, and re-enters safely&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;inner&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;++;&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"in inner(), count is now "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;synchronized&lt;/code&gt; locks weren't reentrant, that call to &lt;code&gt;inner()&lt;/code&gt; from inside &lt;code&gt;outer()&lt;/code&gt; would deadlock the thread against itself, waiting forever for a lock it's already holding. Java's intrinsic locks are reentrant by design, and that's a deliberate design choice, not an accident, one that &lt;code&gt;ReentrantLock&lt;/code&gt; (covered next) takes its name directly from.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;wait()&lt;/code&gt;, &lt;code&gt;notify()&lt;/code&gt;, and coordinating threads
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;synchronized&lt;/code&gt; stops two threads from corrupting shared data at the same time. It doesn't help when one thread needs to &lt;em&gt;wait&lt;/em&gt; for another thread to do something first, that's a different problem, and Java's answer is &lt;code&gt;wait()&lt;/code&gt; and &lt;code&gt;notify()&lt;/code&gt;/&lt;code&gt;notifyAll()&lt;/code&gt;, defined on every &lt;code&gt;Object&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The rules are strict and worth memorizing exactly, because getting any one of them wrong produces a program that hangs instead of one that crashes, which is far harder to debug. &lt;code&gt;wait()&lt;/code&gt;, &lt;code&gt;notify()&lt;/code&gt;, and &lt;code&gt;notifyAll()&lt;/code&gt; can only be called from inside a &lt;code&gt;synchronized&lt;/code&gt; block on the object you're calling them on, or Java throws &lt;code&gt;IllegalMonitorStateException&lt;/code&gt;. Calling &lt;code&gt;wait()&lt;/code&gt; releases the lock and puts the thread to sleep until another thread calls &lt;code&gt;notify()&lt;/code&gt; or &lt;code&gt;notifyAll()&lt;/code&gt; on the same object. &lt;code&gt;notify()&lt;/code&gt; wakes up one waiting thread, chosen arbitrarily; &lt;code&gt;notifyAll()&lt;/code&gt; wakes every thread waiting on that object, and they all then compete for the lock again.&lt;/p&gt;

&lt;p&gt;There's one rule that's very easy to get wrong: &lt;strong&gt;always call &lt;code&gt;wait()&lt;/code&gt; in a loop that rechecks the condition, never in a plain &lt;code&gt;if&lt;/code&gt;.&lt;/strong&gt; Java allows spurious wakeups, where a thread can wake from &lt;code&gt;wait()&lt;/code&gt; even though nobody called &lt;code&gt;notify()&lt;/code&gt;. If your code assumes an &lt;code&gt;if&lt;/code&gt; check before &lt;code&gt;wait()&lt;/code&gt; is still true the moment it wakes up, you've introduced a bug that shows up rarely, unpredictably, and is miserable to reproduce.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// SafeWaitPattern.java — the shape every wait() call should follow&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SafeWaitPattern&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;boolean&lt;/span&gt; &lt;span class="n"&gt;conditionIsTrue&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;waitUntilReady&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;conditionIsTrue&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;   &lt;span class="c1"&gt;// loop, never if&lt;/span&gt;
                &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;wait&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="c1"&gt;// safe to proceed here, condition is genuinely true&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;A worked example: print odd and even alternately using two threads.&lt;/strong&gt; This is a genuinely useful example to build once by hand, since it forces you to actually use &lt;code&gt;wait()&lt;/code&gt;/&lt;code&gt;notify()&lt;/code&gt; to coordinate two threads, rather than just define what they do.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// OddEvenPrinter.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;OddEvenPrinter&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;number&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;max&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;boolean&lt;/span&gt; &lt;span class="n"&gt;oddTurn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;OddEvenPrinter&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;max&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;max&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;max&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;printOdd&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;number&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;max&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;oddTurn&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                    &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Odd: "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;number&lt;/span&gt;&lt;span class="o"&gt;++);&lt;/span&gt;
            &lt;span class="n"&gt;oddTurn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="n"&gt;notify&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;printEven&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;number&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;max&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;oddTurn&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                    &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Even: "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;number&lt;/span&gt;&lt;span class="o"&gt;++);&lt;/span&gt;
            &lt;span class="n"&gt;oddTurn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="n"&gt;notify&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;OddEvenPrinter&lt;/span&gt; &lt;span class="n"&gt;printer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OddEvenPrinter&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;oddThread&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nl"&gt;printer:&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="n"&gt;printOdd&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"odd-thread"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;evenThread&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nl"&gt;printer:&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="n"&gt;printEven&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"even-thread"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="n"&gt;oddThread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;evenThread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Walk through what's actually happening: both threads call into &lt;code&gt;synchronized&lt;/code&gt; methods on the same &lt;code&gt;printer&lt;/code&gt; object, so only one runs at a time. The odd-thread checks a shared boolean flag; if it's not its turn, it calls &lt;code&gt;wait()&lt;/code&gt;, releasing the lock and going to sleep. The even-thread does the same in reverse. Whichever thread does get to print flips the flag and calls &lt;code&gt;notify()&lt;/code&gt; to wake the other one up. Neither thread ever busy-waits, burning CPU checking a condition in a loop, they genuinely sleep until told to wake up, which is the entire point of &lt;code&gt;wait()&lt;/code&gt;/&lt;code&gt;notify()&lt;/code&gt; over a naive spin-loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Java Memory Model and &lt;code&gt;volatile&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;This is one of the places where understanding what's actually happening at the hardware level pays off, because the core distinction to hold onto is that &lt;strong&gt;visibility and atomicity are two separate problems&lt;/strong&gt;, and &lt;code&gt;synchronized&lt;/code&gt; happens to solve both at once, but &lt;code&gt;volatile&lt;/code&gt; only solves one of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Visibility&lt;/strong&gt; is about whether a change one thread makes to a variable is actually seen by another thread, at all, ever. Modern CPUs have per-core caches, and a thread running on one core can update a variable in its own cache without that update being immediately, or ever, flushed to main memory where another thread on another core would see it. Without any coordination, one thread can update a &lt;code&gt;boolean running = true&lt;/code&gt; to &lt;code&gt;false&lt;/code&gt;, and another thread checking &lt;code&gt;running&lt;/code&gt; in a tight loop can keep reading a stale cached &lt;code&gt;true&lt;/code&gt; forever, spinning in a loop that should have exited.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9pn1aiccm69s3z42wbbb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9pn1aiccm69s3z42wbbb.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;volatile&lt;/code&gt; fixes exactly this, and only this. A &lt;code&gt;volatile&lt;/code&gt; field is never cached privately; every read goes to main memory and every write goes straight back to it, so every thread always sees the latest value. It also establishes a happens-before relationship: everything a thread wrote &lt;em&gt;before&lt;/em&gt; writing a volatile variable becomes visible to any thread that reads that same volatile variable afterward, not just the volatile field itself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Worker.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Worker&lt;/span&gt; &lt;span class="kd"&gt;implements&lt;/span&gt; &lt;span class="nc"&gt;Runnable&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;volatile&lt;/span&gt; &lt;span class="kt"&gt;boolean&lt;/span&gt; &lt;span class="n"&gt;running&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;running&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="nd"&gt;@Override&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;cycles&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;running&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;cycles&lt;/span&gt;&lt;span class="o"&gt;++;&lt;/span&gt; &lt;span class="c1"&gt;// stand-in for real work&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"stopped after "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;cycles&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" cycles"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Worker&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Worker&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sleep&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;stop&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// without volatile, this update might never be seen by t&lt;/span&gt;
        &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;join&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But here's the part worth being careful about: &lt;code&gt;volatile&lt;/code&gt; gives you visibility, it does not give you atomicity. Go back to &lt;code&gt;count++&lt;/code&gt; and make &lt;code&gt;count&lt;/code&gt; volatile instead of synchronizing it: the race condition is completely unchanged, because &lt;code&gt;count++&lt;/code&gt; is still three separate operations (read, add, write), and &lt;code&gt;volatile&lt;/code&gt; doesn't make those three operations happen as one atomic unit. Two threads can still both read the same value before either writes, and one increment still vanishes. &lt;code&gt;volatile&lt;/code&gt; is the right tool for a single flag one thread sets and others read, like the &lt;code&gt;running&lt;/code&gt; boolean above. It is the wrong tool for anything involving a read-modify-write sequence, like a counter, that needs &lt;code&gt;synchronized&lt;/code&gt;, an explicit lock, or an atomic variable instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Happens-before&lt;/strong&gt;, briefly, since it's the deeper idea underneath all of this: it's the JVM's guarantee about which memory effects are visible in which order across threads. A &lt;code&gt;synchronized&lt;/code&gt; block establishes happens-before between the thread releasing a lock and the next thread acquiring that same lock, meaning everything the first thread did before releasing is guaranteed visible to the second thread after it acquires. &lt;code&gt;volatile&lt;/code&gt; writes and reads establish the same guarantee for that specific field. &lt;code&gt;Thread.start()&lt;/code&gt; happens-before anything the started thread does, and everything a thread does happens-before another thread's &lt;code&gt;join()&lt;/code&gt; on it returns. Without one of these explicit happens-before relationships in place, the JVM and the CPU are both free to reorder instructions and cache values in ways that can genuinely surprise you, which is the deeper reason unsynchronized shared-memory code is unsafe even when it "looks" correct on paper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Explicit locks: &lt;code&gt;ReentrantLock&lt;/code&gt;, &lt;code&gt;ReadWriteLock&lt;/code&gt;, &lt;code&gt;StampedLock&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;synchronized&lt;/code&gt; is easy to use correctly but inflexible. &lt;code&gt;java.util.concurrent.locks&lt;/code&gt; gives you explicit lock objects with more control, at the cost of having to remember to release them yourself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Counter.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.locks.ReentrantLock&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Counter&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;ReentrantLock&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ReentrantLock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;increment&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;++;&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;unlock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// must be in finally, or an exception leaves the lock held forever&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="nf"&gt;getCount&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;unlock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;finally&lt;/code&gt; block is not optional. With &lt;code&gt;synchronized&lt;/code&gt;, the JVM releases the lock automatically even if an exception is thrown inside. With &lt;code&gt;ReentrantLock&lt;/code&gt;, nothing releases it for you, so a forgotten &lt;code&gt;unlock()&lt;/code&gt; on an exception path leaves the lock held forever, and every other thread waiting for it waits forever too.&lt;/p&gt;

&lt;p&gt;Here's what &lt;code&gt;ReentrantLock&lt;/code&gt; gets you that &lt;code&gt;synchronized&lt;/code&gt; doesn't, laid out directly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;synchronized&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;ReentrantLock&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Acquire/release&lt;/td&gt;
&lt;td&gt;Automatic (block-scoped)&lt;/td&gt;
&lt;td&gt;Manual (&lt;code&gt;lock()&lt;/code&gt; / &lt;code&gt;unlock()&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Try without blocking&lt;/td&gt;
&lt;td&gt;Not possible&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;tryLock()&lt;/code&gt;, returns immediately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timeout on waiting&lt;/td&gt;
&lt;td&gt;Not possible&lt;/td&gt;
&lt;td&gt;&lt;code&gt;tryLock(time, unit)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interruptible while waiting&lt;/td&gt;
&lt;td&gt;Not possible&lt;/td&gt;
&lt;td&gt;&lt;code&gt;lockInterruptibly()&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fairness (FIFO ordering)&lt;/td&gt;
&lt;td&gt;Not possible&lt;/td&gt;
&lt;td&gt;&lt;code&gt;new ReentrantLock(true)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multiple wait conditions&lt;/td&gt;
&lt;td&gt;One implicit condition (&lt;code&gt;wait&lt;/code&gt;/&lt;code&gt;notify&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Several, via &lt;code&gt;newCondition()&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;tryLock()&lt;/code&gt; is how you'd avoid one specific flavour of deadlock rather than just waiting forever:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// TryLockExample.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.locks.ReentrantLock&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.TimeUnit&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;TryLockExample&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;ReentrantLock&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ReentrantLock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;attemptWork&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;tryLock&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;TimeUnit&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;SECONDS&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;doProtectedWork&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;unlock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"couldn't get the lock in time, doing something else instead"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;doProtectedWork&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// the actual work that needs the lock&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fairness is worth a one-line caveat: a fair lock grants access in the order threads requested it, which sounds strictly better, but it comes with a real throughput cost, so it's an opt-in, not the default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;ReadWriteLock&lt;/code&gt;&lt;/strong&gt; is the answer to a common, specific inefficiency: data that's read constantly and written rarely. A plain lock makes every reader wait for every other reader, even though two threads only &lt;em&gt;reading&lt;/em&gt; the same data was never actually unsafe. &lt;code&gt;ReadWriteLock&lt;/code&gt; splits this into a read lock, which any number of threads can hold at once as long as nobody holds the write lock, and a write lock, which is exclusive against everyone, readers included.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// SharedData.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.locks.ReadWriteLock&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.locks.ReentrantReadWriteLock&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SharedData&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;ReadWriteLock&lt;/span&gt; &lt;span class="n"&gt;rwLock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ReentrantReadWriteLock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"initial value"&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;rwLock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;readLock&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;rwLock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;readLock&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;unlock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;rwLock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;writeLock&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;rwLock&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;writeLock&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;unlock&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a read-heavy cache with occasional writes, this can meaningfully beat a plain &lt;code&gt;ReentrantLock&lt;/code&gt;, since readers stop blocking each other entirely. &lt;strong&gt;&lt;code&gt;StampedLock&lt;/code&gt;&lt;/strong&gt; takes this further with an optimistic read mode that doesn't block writers at all, a reader just validates afterward that no write happened in the meantime and retries if one did. It's a genuinely more advanced tool, worth being able to name and describe in one sentence, but you'd typically only reach for it once you've confirmed a &lt;code&gt;ReadWriteLock&lt;/code&gt; isn't fast enough for what you're building.&lt;/p&gt;

&lt;h2&gt;
  
  
  Atomic variables and compare-and-swap
&lt;/h2&gt;

&lt;p&gt;There's a third way to fix the counter, and it's usually the best one for a simple case like this: skip locking entirely and use a class built for lock-free atomic updates.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Counter.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.atomic.AtomicInteger&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Counter&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;AtomicInteger&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;AtomicInteger&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;increment&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;incrementAndGet&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="nf"&gt;getCount&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the same ten-threads &lt;code&gt;Main&lt;/code&gt; class from the race condition section against this version and, again, you get exactly 10,000, with no &lt;code&gt;synchronized&lt;/code&gt; keyword anywhere. &lt;code&gt;AtomicInteger&lt;/code&gt;, &lt;code&gt;AtomicLong&lt;/code&gt;, and &lt;code&gt;AtomicReference&lt;/code&gt; work using a CPU-level instruction called compare-and-swap (CAS): "update this memory location to a new value, but only if it still holds the value I expect it to; tell me whether that succeeded."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0tjj3h5an1m9kammg5op.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0tjj3h5an1m9kammg5op.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;incrementAndGet()&lt;/code&gt; is implemented, roughly, as a loop: read the current value, compute the new value, attempt a CAS from old to new, and if the CAS fails because another thread changed the value in between, loop around and try again with the fresh value. Here's that same loop using the lower-level &lt;code&gt;compareAndSet()&lt;/code&gt; method directly, which shows what's happening underneath, though you'd never actually write it this way in real code, &lt;code&gt;incrementAndGet()&lt;/code&gt; already does it for you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// ManualCasIncrement.java — for illustration, showing what incrementAndGet() does internally&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.atomic.AtomicInteger&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ManualCasIncrement&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;incrementManually&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;AtomicInteger&lt;/span&gt; &lt;span class="n"&gt;counter&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;counter&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;updated&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;counter&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;compareAndSet&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;updated&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="c1"&gt;// succeeded: nothing else changed the value between get() and here&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="c1"&gt;// else: another thread updated it first, loop around and retry&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No thread ever blocks waiting for a lock, they just occasionally retry, which tends to perform noticeably better than locking under high contention on simple operations like this, since there's no thread ever sitting idle waiting to be woken up.&lt;/p&gt;

&lt;p&gt;The tradeoff is scope: atomics are excellent for a single variable, a counter, a flag, a reference that gets swapped. They don't help when you need to update &lt;em&gt;several&lt;/em&gt; related variables together as one atomic unit, like moving money between two account balances, that's still a job for a lock.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deadlock, livelock, and starvation: a bank transfer gone wrong
&lt;/h2&gt;

&lt;p&gt;A deadlock happens when two or more threads are each waiting for a lock the other one holds, and neither will ever let go. Here's the version that shows up in real code, not a toy two-lock example: two bank accounts, two threads transferring money in opposite directions at the same time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Account.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;double&lt;/span&gt; &lt;span class="n"&gt;balance&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;Account&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;double&lt;/span&gt; &lt;span class="n"&gt;balance&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;balance&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="nf"&gt;getId&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt; &lt;span class="nf"&gt;getLock&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;withdraw&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;double&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;-=&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;deposit&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;double&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// BankTransfer.java — the broken version&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;BankTransfer&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;transfer&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;double&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getLock&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getName&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" locked account "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getId&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
            &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getLock&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getName&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" locked account "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getId&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
                &lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;withdraw&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;deposit&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="n"&gt;accountA&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Account&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="n"&gt;accountB&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Account&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
            &lt;span class="nc"&gt;BankTransfer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;transfer&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;accountA&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;accountB&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="o"&gt;),&lt;/span&gt; &lt;span class="s"&gt;"Thread-1"&lt;/span&gt;
        &lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
            &lt;span class="nc"&gt;BankTransfer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;transfer&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;accountB&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;accountA&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="o"&gt;),&lt;/span&gt; &lt;span class="s"&gt;"Thread-2"&lt;/span&gt;
        &lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="n"&gt;t1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;t2&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="c1"&gt;// this can hang forever, depending on timing&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhjh6t3vkzmqsrftqjmpg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhjh6t3vkzmqsrftqjmpg.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Walk through why this locks up: Thread 1 acquires &lt;code&gt;accountA&lt;/code&gt;'s lock first, then tries to acquire &lt;code&gt;accountB&lt;/code&gt;'s lock. At almost the same moment, Thread 2 acquires &lt;code&gt;accountB&lt;/code&gt;'s lock first, then tries to acquire &lt;code&gt;accountA&lt;/code&gt;'s lock. If the timing lines up so each thread grabs its first lock before either reaches for its second, Thread 1 is now stuck waiting for a lock Thread 2 holds, and Thread 2 is stuck waiting for a lock Thread 1 holds. Neither thread will ever release what it's holding, because releasing only happens after the &lt;code&gt;synchronized&lt;/code&gt; block completes, and neither block can complete. The program doesn't crash, it just stops, silently, forever, which is exactly what makes deadlocks nasty in production: nothing looks wrong until it does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix is lock ordering:&lt;/strong&gt; always acquire locks in the same global order, no matter which direction the transfer is going.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// BankTransfer.java — the fixed version&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;BankTransfer&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;transfer&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;double&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="n"&gt;first&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getId&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getId&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="n"&gt;from&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="nc"&gt;Account&lt;/span&gt; &lt;span class="n"&gt;second&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getId&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getId&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

        &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;first&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getLock&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;second&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getLock&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;withdraw&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;deposit&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now every thread, regardless of which direction it's transferring, locks the lower-ID account first. Thread 1 and Thread 2 can never both be holding one lock while wanting the other, because they always reach for the same first lock, so one of them simply waits at the very first &lt;code&gt;synchronized&lt;/code&gt; block instead of deadlocking deep inside two nested ones. The general pattern worth taking away: assign a fixed, consistent order to every lock in the system, and always acquire in that order. The other standard tool is a timeout: &lt;code&gt;tryLock(time, unit)&lt;/code&gt; on &lt;code&gt;ReentrantLock&lt;/code&gt; lets a thread give up and back off instead of waiting forever, which avoids the deadlock at the cost of needing retry logic.&lt;/p&gt;

&lt;p&gt;Two related terms are worth being able to define crisply, since they're easy to confuse with deadlock and with each other. &lt;strong&gt;Livelock&lt;/strong&gt; is when threads aren't blocked, they're actively running, but they keep responding to each other in a way that prevents any of them from making real progress, like two people repeatedly stepping aside for each other in a hallway and never actually passing. &lt;strong&gt;Starvation&lt;/strong&gt; is when one specific thread never gets access to a resource because other threads keep getting priority over it, it's not stuck waiting on a cycle, it's just perpetually unlucky (or deprioritized). Both come up less often than deadlock, but being able to name the difference correctly reflects real understanding of the distinction rather than a memorized list of terms.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Executor framework and thread pools
&lt;/h2&gt;

&lt;p&gt;Creating a &lt;code&gt;new Thread()&lt;/code&gt; for every unit of work doesn't scale. Each thread carries real OS overhead, and a burst of ten thousand incoming requests each spinning up its own thread will exhaust memory and grind the machine to a halt on context switching long before it exhausts CPU. The Executor framework separates &lt;em&gt;submitting&lt;/em&gt; work from &lt;em&gt;how it actually gets run&lt;/em&gt;, so you write task-submission code once and swap the execution strategy underneath it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.ExecutorService&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.Executors&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;ExecutorService&lt;/span&gt; &lt;span class="n"&gt;executor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Executors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;newFixedThreadPool&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;taskNumber&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;submit&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"running task "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;taskNumber&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" on "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getName&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
            &lt;span class="o"&gt;});&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;shutdown&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// stop accepting new tasks, let submitted ones finish&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Executors&lt;/code&gt; offers several convenience factory methods, and it's worth knowing them, along with why most production codebases have moved away from several of them in favour of building a &lt;code&gt;ThreadPoolExecutor&lt;/code&gt; directly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Factory method&lt;/th&gt;
&lt;th&gt;What it creates&lt;/th&gt;
&lt;th&gt;The catch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;newFixedThreadPool(n)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A pool of exactly n threads, an unbounded task queue&lt;/td&gt;
&lt;td&gt;The unbounded queue can grow without limit if tasks arrive faster than they're processed, hiding a backlog until memory runs out&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;newCachedThreadPool()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Creates threads as needed, reuses idle ones, no upper bound&lt;/td&gt;
&lt;td&gt;No upper bound means a burst of load can create an unbounded number of threads, exactly the problem pooling was meant to prevent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;newSingleThreadExecutor()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;One thread, tasks run strictly in order&lt;/td&gt;
&lt;td&gt;Fine for its specific purpose, just not a general-purpose pool&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;newScheduledThreadPool(n)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Like fixed, plus delayed and repeating tasks&lt;/td&gt;
&lt;td&gt;Same unbounded-queue concern as the fixed pool&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The common thread through the catches: several of these factories quietly use an unbounded queue or an unbounded thread count, which trades one failure mode (an &lt;code&gt;OutOfMemoryError&lt;/code&gt; from too many raw threads) for a subtler one (silent, unbounded queue growth, or the same &lt;code&gt;OutOfMemoryError&lt;/code&gt; from an unbounded cached pool under sustained load). This is exactly why it's worth reaching for &lt;code&gt;ThreadPoolExecutor&lt;/code&gt; directly once you actually care about these bounds, where every one of them is explicit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.*&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;ExecutorService&lt;/span&gt; &lt;span class="n"&gt;executor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ThreadPoolExecutor&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
            &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;                                          &lt;span class="c1"&gt;// core pool size: threads kept alive even when idle&lt;/span&gt;
            &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;                                          &lt;span class="c1"&gt;// maximum pool size: the ceiling under load&lt;/span&gt;
            &lt;span class="mi"&gt;60L&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;TimeUnit&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;SECONDS&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;                      &lt;span class="c1"&gt;// how long extra threads beyond core sit idle before dying&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;LinkedBlockingQueue&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="o"&gt;),&lt;/span&gt;              &lt;span class="c1"&gt;// bounded queue: caps how much backlog can build up&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ThreadPoolExecutor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;CallerRunsPolicy&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;    &lt;span class="c1"&gt;// what to do when the queue is also full&lt;/span&gt;
        &lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;taskNumber&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;submit&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
                &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"task "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;taskNumber&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" on "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getName&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
            &lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;shutdown&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frz8afel6rio7gaarfray.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frz8afel6rio7gaarfray.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Walk through what each parameter is actually deciding. &lt;code&gt;corePoolSize&lt;/code&gt; threads stay alive even with nothing to do. Once the queue fills up, the pool grows beyond core, up to &lt;code&gt;maximumPoolSize&lt;/code&gt;, to absorb the extra load. &lt;code&gt;keepAliveTime&lt;/code&gt; controls how long those extra threads beyond core linger idle before being shut down and reclaimed. The queue itself is bounded on purpose, so backlog has a hard ceiling instead of growing invisibly. And the rejection policy decides what happens once both the pool and the queue are completely full and a new task still arrives, options include throwing an exception (&lt;code&gt;AbortPolicy&lt;/code&gt;, the default), silently dropping the task (&lt;code&gt;DiscardPolicy&lt;/code&gt;), or, as used above, running the task on the calling thread itself as a form of backpressure (&lt;code&gt;CallerRunsPolicy&lt;/code&gt;), which has the effect of slowing down whoever's submitting work until the pool catches up.&lt;/p&gt;

&lt;p&gt;Sizing a pool sensibly depends on what the work actually is. For CPU-bound work, a pool size around the number of available CPU cores is typically a reasonable starting point, since more threads than cores just means more context switching with no extra throughput. For I/O-bound work, where threads spend most of their time waiting on a network call or a disk read rather than actually using the CPU, a pool considerably larger than the core count usually performs better, since idle-waiting threads aren't competing for CPU anyway. This exact tension, threads mostly sitting idle waiting on I/O, is precisely the problem virtual threads are built to solve.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;Callable&lt;/code&gt;, &lt;code&gt;Future&lt;/code&gt;, and &lt;code&gt;CompletableFuture&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;Runnable&lt;/code&gt; doesn't return anything and can't throw a checked exception. &lt;code&gt;Callable&amp;lt;V&amp;gt;&lt;/code&gt; fixes both: it returns a value and is allowed to throw.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.*&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;Exception&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;ExecutorService&lt;/span&gt; &lt;span class="n"&gt;executor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Executors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;newFixedThreadPool&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="nc"&gt;Callable&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Integer&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;task&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sleep&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="o"&gt;};&lt;/span&gt;

        &lt;span class="nc"&gt;Future&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Integer&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;future&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;submit&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"doing other work while the task runs..."&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="nc"&gt;Integer&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;future&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// blocks until the result is ready, or throws&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"result: "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;shutdown&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;future.get()&lt;/code&gt; blocks the calling thread until the result is available, which is useful but limited: you can't easily chain what happens next, and you can't combine the results of several futures without writing a fair amount of blocking, coordinating code by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;CompletableFuture&lt;/code&gt;&lt;/strong&gt; solves that by letting you describe a pipeline of what should happen when a result becomes available, without ever explicitly blocking:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.CompletableFuture&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;CompletableFuture&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Integer&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;future&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;CompletableFuture&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;supplyAsync&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nl"&gt;Main:&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="n"&gt;fetchUserId&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;thenApply&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;thenApply&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doubled&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;doubled&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="n"&gt;future&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;thenAccept&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Final: "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;

        &lt;span class="c1"&gt;// give the async pipeline a moment to finish before the program exits&lt;/span&gt;
        &lt;span class="n"&gt;future&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;join&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="nf"&gt;fetchUserId&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// stand-in for a slow call, like hitting a database&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;thenApply&lt;/code&gt; transforms a result once it's ready and hands you a new &lt;code&gt;CompletableFuture&lt;/code&gt; wrapping the transformed value, &lt;code&gt;thenAccept&lt;/code&gt; consumes the final result without producing a new value, and &lt;code&gt;thenCompose&lt;/code&gt; is what you reach for when the next step is itself another asynchronous operation (chaining two &lt;code&gt;CompletableFuture&lt;/code&gt;s together rather than nesting one inside a plain function). &lt;code&gt;CompletableFuture.allOf(f1, f2, f3)&lt;/code&gt; waits for several independent futures to all complete, which is the standard way to fan out several concurrent calls, say, three separate API calls, and combine their results once every one of them is done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Concurrent collections, and building a blocking queue by hand
&lt;/h2&gt;

&lt;p&gt;Wrapping a plain &lt;code&gt;HashMap&lt;/code&gt; or &lt;code&gt;ArrayList&lt;/code&gt; in &lt;code&gt;synchronized&lt;/code&gt; on every call works, but it serializes every single access, one thread at a time, even for operations that could safely happen concurrently. &lt;code&gt;java.util.concurrent&lt;/code&gt; gives you collections designed for concurrent access from the ground up, and they're almost always the better choice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;ConcurrentHashMap&lt;/code&gt;&lt;/strong&gt; allows concurrent reads and writes without locking the entire map for every operation, internally splitting its locking so unrelated keys don't contend with each other. It's the default choice for a shared map accessed by multiple threads, essentially never &lt;code&gt;Collections.synchronizedMap(new HashMap&amp;lt;&amp;gt;())&lt;/code&gt; in new code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.ConcurrentHashMap&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.Map&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Integer&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;visitCounts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ConcurrentHashMap&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;();&lt;/span&gt;

        &lt;span class="nc"&gt;Runnable&lt;/span&gt; &lt;span class="n"&gt;visitor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="n"&gt;visitCounts&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;merge&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"home-page"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nl"&gt;Integer:&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="n"&gt;sum&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;};&lt;/span&gt;

        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;visitor&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;t2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;visitor&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;t1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;t2&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;t1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;join&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;t2&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;join&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;visitCounts&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"home-page"&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt; &lt;span class="c1"&gt;// reliably 2000&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;&lt;code&gt;CopyOnWriteArrayList&lt;/code&gt;&lt;/strong&gt; takes a different approach entirely: every write creates a brand-new copy of the underlying array, while reads always work against a stable, never-changing snapshot and need no locking at all. That makes it excellent for read-heavy, write-rare situations, like a list of event listeners that's iterated constantly and modified rarely, and a poor choice for anything write-heavy, since copying the entire array on every single write gets expensive fast.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;BlockingQueue&lt;/code&gt;&lt;/strong&gt; implementations (&lt;code&gt;ArrayBlockingQueue&lt;/code&gt;, &lt;code&gt;LinkedBlockingQueue&lt;/code&gt;) add blocking behaviour on top of a queue: &lt;code&gt;put()&lt;/code&gt; blocks the calling thread if the queue is full instead of failing, and &lt;code&gt;take()&lt;/code&gt; blocks if the queue is empty instead of returning null. That blocking behaviour is exactly what makes the producer-consumer pattern simple to write correctly instead of needing manual &lt;code&gt;wait()&lt;/code&gt;/&lt;code&gt;notify()&lt;/code&gt; coordination.&lt;/p&gt;

&lt;p&gt;It's worth actually building a small bounded blocking queue yourself once, using nothing but &lt;code&gt;wait()&lt;/code&gt;/&lt;code&gt;notify()&lt;/code&gt;, because it's a great way to see exactly what &lt;code&gt;BlockingQueue&lt;/code&gt; is doing underneath rather than just knowing it exists:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// SimpleBlockingQueue.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.LinkedList&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.Queue&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SimpleBlockingQueue&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="no"&gt;T&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;Queue&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="no"&gt;T&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;LinkedList&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;();&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;SimpleBlockingQueue&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;capacity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="no"&gt;T&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;size&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// full, wait for a consumer to make room&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;add&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;notifyAll&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// wake any consumer waiting on take()&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;synchronized&lt;/span&gt; &lt;span class="no"&gt;T&lt;/span&gt; &lt;span class="nf"&gt;take&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;isEmpty&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// empty, wait for a producer to add something&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="no"&gt;T&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;poll&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;notifyAll&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// wake any producer waiting on put()&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;SimpleBlockingQueue&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Integer&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;SimpleBlockingQueue&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;producer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;put&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                    &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"produced "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;});&lt;/span&gt;

        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;consumer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;take&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                    &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"consumed "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;});&lt;/span&gt;

        &lt;span class="n"&gt;producer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;consumer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both &lt;code&gt;put()&lt;/code&gt; and &lt;code&gt;take()&lt;/code&gt; wait in a loop, not an &lt;code&gt;if&lt;/code&gt;, for exactly the spurious-wakeup reason covered earlier. &lt;code&gt;notifyAll()&lt;/code&gt; rather than &lt;code&gt;notify()&lt;/code&gt; matters here specifically because both producers and consumers might be waiting on the same object at once, &lt;code&gt;notify()&lt;/code&gt; would only wake one arbitrary thread, which might not even be the kind of thread that can currently make progress; &lt;code&gt;notifyAll()&lt;/code&gt; wakes everyone and lets each one recheck its own condition. This is, in miniature, exactly what &lt;code&gt;ArrayBlockingQueue&lt;/code&gt; does internally, just with more tuning and edge-case handling than is worth reproducing by hand in real code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coordination utilities: &lt;code&gt;CountDownLatch&lt;/code&gt;, &lt;code&gt;CyclicBarrier&lt;/code&gt;, &lt;code&gt;Semaphore&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;These three exist for a specific, recurring shape of problem: several threads need to coordinate their timing with each other, not just their access to shared data.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8qwls1dalcdqgs5lgo5t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8qwls1dalcdqgs5lgo5t.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;CountDownLatch&lt;/code&gt;&lt;/strong&gt; lets one or more threads wait until a fixed number of events have happened. It's initialized with a count, other threads call &lt;code&gt;countDown()&lt;/code&gt; as they finish their part, and any thread that called &lt;code&gt;await()&lt;/code&gt; unblocks the moment the count hits zero. It cannot be reset or reused once it reaches zero, it's a strictly one-time gate.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.CountDownLatch&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;CountDownLatch&lt;/span&gt; &lt;span class="n"&gt;latch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;CountDownLatch&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;workerId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"worker "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;workerId&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" starting work"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="n"&gt;doWork&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"worker "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;workerId&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" finished"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="n"&gt;latch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;countDown&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="o"&gt;}).&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;latch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;await&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// blocks until all 3 threads have called countDown()&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"all workers finished, main can continue"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;doWork&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sleep&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;&lt;code&gt;CyclicBarrier&lt;/code&gt;&lt;/strong&gt; looks similar but solves a different shape of problem: instead of one thread waiting for several others, a &lt;em&gt;group&lt;/em&gt; of threads all wait for each other to reach the same point before any of them proceeds, and once released, the barrier resets and can be used again for the next round.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.CyclicBarrier&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;CyclicBarrier&lt;/span&gt; &lt;span class="n"&gt;barrier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;CyclicBarrier&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;
            &lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"all 3 threads reached the barrier, proceeding together"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
        &lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;workerId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="n"&gt;doWorkPhaseOne&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;workerId&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;barrier&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;await&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Exception&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
                &lt;span class="n"&gt;doWorkPhaseTwo&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;workerId&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// only starts once all 3 threads reached the barrier&lt;/span&gt;
            &lt;span class="o"&gt;}).&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;doWorkPhaseOne&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;workerId&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"worker "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;workerId&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" doing phase one"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;doWorkPhaseTwo&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;workerId&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"worker "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;workerId&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" doing phase two"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The concrete difference worth taking away: &lt;code&gt;CountDownLatch&lt;/code&gt; is for one thread (or several) waiting on a fixed number of other threads to finish, used once; &lt;code&gt;CyclicBarrier&lt;/code&gt; is for a fixed group of threads waiting on each other to all reach the same checkpoint together, and it's reusable round after round, useful for something like a simulation that proceeds in synchronized phases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;Semaphore&lt;/code&gt;&lt;/strong&gt; controls access to a resource with a limited number of permits, rather than the single all-or-nothing lock that &lt;code&gt;synchronized&lt;/code&gt; and &lt;code&gt;ReentrantLock&lt;/code&gt; give you. &lt;code&gt;acquire()&lt;/code&gt; takes a permit, blocking if none are available; &lt;code&gt;release()&lt;/code&gt; gives one back. A semaphore initialized with 3 permits allows exactly three threads through at once, a straightforward way to cap concurrent access to something like a pool of database connections or a rate-limited external API.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.Semaphore&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;Semaphore&lt;/span&gt; &lt;span class="n"&gt;semaphore&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Semaphore&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// at most 3 concurrent&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;callId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;semaphore&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;acquire&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                    &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"call "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;callId&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" got a permit"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                    &lt;span class="n"&gt;callLimitedResource&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;semaphore&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;release&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;}).&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;callLimitedResource&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sleep&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Virtual threads: what changed in Java 21
&lt;/h2&gt;

&lt;p&gt;This is worth understanding properly, because virtual threads became a standard, non-preview feature in Java 21 (they'd been previewed in 19 and 20 first), and they've genuinely changed how high-throughput Java servers get built.&lt;/p&gt;

&lt;p&gt;A traditional web server handles each incoming request on its own platform thread, and each platform thread is a real OS thread, carrying real memory overhead and real scheduling cost. Most of a typical request's time isn't spent computing, it's spent &lt;em&gt;waiting&lt;/em&gt;: on a database query, on a downstream API call, on disk I/O. A platform thread sits there fully allocated, doing nothing, for the entire wait. Scale that to tens of thousands of concurrent requests and you run out of threads, and the machine, long before you run out of actual CPU work to do.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdzd7tbeiva7om8dzp2ge.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdzd7tbeiva7om8dzp2ge.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A virtual thread is a thread implemented by the JVM itself rather than mapped one-to-one to an OS thread. Many thousands of virtual threads run on top of a small pool of ordinary platform threads, called carrier threads. When a virtual thread blocks on I/O, the JVM unmounts it from its carrier thread entirely, freeing that carrier to run a different virtual thread, and remounts the original one onto some carrier once its I/O actually completes. The blocking, from the code's point of view, looks exactly like ordinary blocking code always has, &lt;code&gt;Thread.sleep()&lt;/code&gt;, a blocking database call, a synchronous HTTP client, none of it needs to be rewritten in a reactive or callback style. That's the actual headline: you get the throughput characteristics of async code while writing plain, ordinary, sequential-looking blocking code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.ExecutorService&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.Executors&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ExecutorService&lt;/span&gt; &lt;span class="n"&gt;executor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Executors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;newVirtualThreadPerTaskExecutor&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;10_000&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;taskId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
                &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;submit&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;makeSlowNetworkCall&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// blocks this virtual thread, not a whole OS thread&lt;/span&gt;
                    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
                &lt;span class="o"&gt;});&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="c1"&gt;// executor.close() is called automatically here, waiting for all tasks to finish&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;makeSlowNetworkCall&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sleep&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// stand-in for a real I/O-bound call&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ten thousand platform threads would very plausibly crash a typical server. Ten thousand virtual threads, each cheap to create and cheap to sit idle, generally won't, because the JVM never had to allocate ten thousand OS threads' worth of stack memory and scheduling overhead to begin with.&lt;/p&gt;

&lt;p&gt;The one caveat worth holding onto clearly: virtual threads help &lt;strong&gt;I/O-bound&lt;/strong&gt; workloads, where threads spend most of their time waiting. They do essentially nothing for &lt;strong&gt;CPU-bound&lt;/strong&gt; work, tight loops doing real computation with no blocking at all, because in that case there's no idle waiting time to reclaim in the first place; the bottleneck is genuinely the CPU, and virtual threads don't create more CPU cores. It's also worth knowing that two closely related Loom features, structured concurrency and scoped values, are still preview features even as of Java 25, not yet stable APIs, so it's fine to know what they're for in one sentence each (structured concurrency treats a group of related subtasks as a single unit that succeeds or fails together; scoped values are a safer, immutable alternative to &lt;code&gt;ThreadLocal&lt;/code&gt; for sharing context across threads) without needing to write their exact preview syntax from memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Producer-consumer, done properly
&lt;/h2&gt;

&lt;p&gt;Here's producer-consumer done once, cleanly, with the real &lt;code&gt;BlockingQueue&lt;/code&gt; rather than the hand-rolled &lt;code&gt;SimpleBlockingQueue&lt;/code&gt; from before.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Main.java&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.ArrayBlockingQueue&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.concurrent.BlockingQueue&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Main&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;BlockingQueue&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Integer&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ArrayBlockingQueue&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;producer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;put&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// blocks automatically if the queue is full&lt;/span&gt;
                    &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"produced "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;});&lt;/span&gt;

        &lt;span class="nc"&gt;Thread&lt;/span&gt; &lt;span class="n"&gt;consumer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                    &lt;span class="nc"&gt;Integer&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;take&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// blocks automatically if the queue is empty&lt;/span&gt;
                    &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;out&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;println&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"consumed "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
                &lt;span class="o"&gt;}&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;InterruptedException&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
                &lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;currentThread&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;interrupt&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="o"&gt;}&lt;/span&gt;
        &lt;span class="o"&gt;});&lt;/span&gt;

        &lt;span class="n"&gt;producer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;consumer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice how much of the earlier complexity has simply disappeared. No &lt;code&gt;synchronized&lt;/code&gt;, no manual &lt;code&gt;wait()&lt;/code&gt;/&lt;code&gt;notify()&lt;/code&gt;, no explicit condition-checking loop, &lt;code&gt;BlockingQueue&lt;/code&gt; has absorbed all of it internally, using essentially the same wait-loop-and-notify mechanics behind &lt;code&gt;SimpleBlockingQueue&lt;/code&gt;. That's the real lesson here: the primitives (&lt;code&gt;synchronized&lt;/code&gt;, &lt;code&gt;wait&lt;/code&gt;, &lt;code&gt;notify&lt;/code&gt;) are what you need to &lt;em&gt;understand&lt;/em&gt; deeply, so you can reason about what's happening underneath and debug it when it breaks, but the higher-level tools (&lt;code&gt;BlockingQueue&lt;/code&gt;, &lt;code&gt;ExecutorService&lt;/code&gt;, &lt;code&gt;CompletableFuture&lt;/code&gt;) are almost always what you should actually &lt;em&gt;reach for&lt;/em&gt; in real code, since they've already handled the edge cases you'd otherwise get wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reasoning through a stuck or misbehaving program
&lt;/h2&gt;

&lt;p&gt;There's one more skill worth building: being able to look at a concurrent program that's misbehaving and reason through why, methodically, rather than guessing. Here's a way to think about it that covers most real situations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A program appears to hang.&lt;/strong&gt; The first, fastest diagnostic is a thread dump, triggerable with &lt;code&gt;jstack &amp;lt;pid&amp;gt;&lt;/code&gt; from the command line, or by sending &lt;code&gt;SIGQUIT&lt;/code&gt; to the process, or by capturing one straight from your IDE's debugger. A thread dump lists every thread in the JVM along with its current state and, critically, its full stack trace at that exact instant. Scan specifically for threads in the &lt;code&gt;BLOCKED&lt;/code&gt; state, that tells you they're waiting on a lock, and the dump names exactly which lock and which thread is currently holding it. If thread A is &lt;code&gt;BLOCKED&lt;/code&gt; waiting for a lock thread B holds, and thread B's stack shows it &lt;code&gt;BLOCKED&lt;/code&gt; waiting for a lock thread A holds, you've found your deadlock directly in the dump, no guessing required. Modern JVMs actually detect this specific cycle for you and print "Found one Java-level deadlock" right at the top of the dump, naming the exact threads and locks involved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A program produces wrong results, but doesn't hang.&lt;/strong&gt; This is the race-condition case, and it's harder precisely because nothing crashes and nothing blocks; the program simply computes something subtly incorrect, and often only under real concurrent load, never in a quick single-threaded test. The systematic approach is to look for exactly the pattern from the counter example: any place where a shared, mutable variable is read, modified, and written back without a lock, an atomic type, or a &lt;code&gt;volatile&lt;/code&gt; covering the whole read-modify-write sequence. &lt;code&gt;count++&lt;/code&gt;, &lt;code&gt;list.add()&lt;/code&gt; on a plain &lt;code&gt;ArrayList&lt;/code&gt;, &lt;code&gt;if (map.get(k) == null) map.put(k, v)&lt;/code&gt;, that check-then-act pattern in particular is a very common, very real source of race conditions, since the check and the act aren't atomic together even if each individually looks safe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A program is slow, but not stuck.&lt;/strong&gt; Here a thread dump is still the first move, but you're looking for something different: threads sitting &lt;code&gt;RUNNABLE&lt;/code&gt; for a long stretch doing real work, one very large &lt;code&gt;synchronized&lt;/code&gt; block that most other threads are &lt;code&gt;BLOCKED&lt;/code&gt; waiting on, or, in older code, contention around a single coarse lock that could be split into a &lt;code&gt;ReadWriteLock&lt;/code&gt;, a &lt;code&gt;ConcurrentHashMap&lt;/code&gt;, or several finer-grained locks instead of one lock guarding far more than it needs to.&lt;/p&gt;

&lt;p&gt;The consistent thread through all three: pin down the symptom precisely first, hung completely, wrong output, or just slow, because each symptom points toward a genuinely different class of cause, and starting from that distinction is almost always faster than diving straight into a specific fix and hoping.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.oracle.com/javase/tutorial/essential/concurrency/" rel="noopener noreferrer"&gt;Oracle's official Java Concurrency tutorial&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openjdk.org/jeps/444" rel="noopener noreferrer"&gt;JEP 444: Virtual Threads&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.oracle.com/javase/specs/jls/se21/html/jls-17.html" rel="noopener noreferrer"&gt;Java Language Specification, Chapter 17: Threads and Locks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://jcip.net/" rel="noopener noreferrer"&gt;Brian Goetz et al., &lt;em&gt;Java Concurrency in Practice&lt;/em&gt;&lt;/a&gt;, dated in places, still the deepest treatment of the memory model and locking available&lt;/li&gt;
&lt;li&gt;&lt;a href="https://inside.java/tags/loom/" rel="noopener noreferrer"&gt;Inside Java's ongoing coverage of Project Loom&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>tutorial</category>
      <category>java</category>
      <category>learning</category>
      <category>multithreading</category>
    </item>
    <item>
      <title>Docker for SDE interviews: the guide I wish I'd had</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Fri, 25 Sep 2026 22:57:26 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/docker-for-sde-interviews-the-guide-i-wish-id-had-205h</link>
      <guid>https://dev.to/shogun_the_grt/docker-for-sde-interviews-the-guide-i-wish-id-had-205h</guid>
      <description>&lt;p&gt;&lt;em&gt;What it actually does under the hood, the Dockerfile traps interviewers love, and the debugging playbook that makes you sound like you've shipped something.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqcg65gz4udygambg4hdb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqcg65gz4udygambg4hdb.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Docker questions show up in almost every backend interview now, even ones that aren't about infrastructure. You mention deploying something, and the interviewer asks how you packaged it. You draw a system design diagram, and there's a box labeled "container" that nobody explains. Half the time it's a genuine technical question. The other half it's a filter: does this person actually understand what they've been typing into a terminal, or have they only ever copied a Dockerfile from Stack Overflow?&lt;/p&gt;

&lt;p&gt;This post covers what you actually need: how Docker works under the hood, writing a Dockerfile that isn't naive, images and layers, networking, storage, Compose, the CLI commands worth knowing, security basics, where Docker stops being enough, and a debugging playbook for the scenario questions interviewers like to throw at you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Docker actually is
&lt;/h2&gt;

&lt;p&gt;Here's the confusion to clear up first, because it trips up more candidates than anything else: a container is not a lightweight virtual machine. It's a regular process on the host's Linux kernel, made to look isolated using three kernel features that have existed for years, which Docker packaged into something usable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Namespaces&lt;/strong&gt; give a process its own view of things that are normally global. A PID namespace makes a container's first process think it's PID 1, even though the host sees it as PID 48213. A network namespace gives it its own network interfaces and routing table. There are namespaces for mount points, hostnames, user IDs, and inter-process communication too. Namespaces are about what a process can &lt;em&gt;see&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Control groups (cgroups)&lt;/strong&gt; limit what a process can &lt;em&gt;use&lt;/em&gt;: how much CPU, memory, and I/O bandwidth. &lt;code&gt;docker run --memory=512m&lt;/code&gt; is a cgroup limit. Without one, a single container can eat all the host's RAM and take everything else down with it, which is exactly the kind of thing that happens in production and gets asked about in interviews.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A union filesystem&lt;/strong&gt; stacks read-only image layers with one thin writable layer on top, so containers share the same base files on disk instead of each getting a full copy.&lt;/p&gt;

&lt;p&gt;Put those three together and you get something that starts in milliseconds, because you're not booting a kernel, just starting a process with some walls around it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5lp1ni3c6txa72f91rqn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5lp1ni3c6txa72f91rqn.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That's also the honest answer to "container vs VM," which is asked in nearly every Docker interview. A VM virtualizes hardware: a hypervisor gives each VM its own kernel, so a VM can run a completely different OS than its host, and boots in the order of a minute because it's booting an entire operating system. A container virtualizes at the OS level: it shares the host's kernel, so a Linux container needs a Linux host kernel underneath it (Docker Desktop on Mac and Windows quietly runs a small Linux VM to give you that kernel), and it starts in milliseconds to a couple of seconds because there's no kernel boot involved. The tradeoff is isolation strength: a VM's hypervisor boundary is harder to break out of than a container's kernel-feature boundary, which is part of why nobody runs genuinely hostile, untrusted code in a bare container without extra layers like gVisor or Kata Containers on top.&lt;/p&gt;

&lt;p&gt;This is also why "Docker Desktop" and "Docker Engine" aren't quite the same thing, a distinction worth having straight if you develop on a Mac or Windows laptop and deploy to Linux servers, which describes most people. On Linux, Docker Engine runs natively, talking directly to the host kernel's namespaces and cgroups. On Mac and Windows there is no Linux kernel to talk to, so Docker Desktop quietly runs a small Linux VM in the background and Docker Engine runs inside that VM instead. Functionally you barely notice the difference day to day, but it's the reason volumes on a Mac live inside that hidden VM rather than as a directly browsable folder on your actual filesystem, and it's part of why "works fine on my Mac, does something weird on the server" occasionally isn't actually a lie, there's a real extra layer on your laptop that the server doesn't have.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The architecture.&lt;/strong&gt; Four pieces, and interviewers like to hear you name all four because it explains why Docker commands sometimes fail in confusing ways. The &lt;strong&gt;Docker CLI&lt;/strong&gt; is the &lt;code&gt;docker&lt;/code&gt; command you type. The &lt;strong&gt;Docker daemon&lt;/strong&gt; (&lt;code&gt;dockerd&lt;/code&gt;) is a background service that does the actual work: building images, running containers, managing networks and volumes. The &lt;strong&gt;Docker Engine&lt;/strong&gt; is the daemon plus its REST API and the CLI together, the whole toolset. A &lt;strong&gt;registry&lt;/strong&gt; stores and distributes images; Docker Hub is the default public one. When you run &lt;code&gt;docker build&lt;/code&gt;, the CLI doesn't build anything itself, it sends a request to the daemon over that REST API, and the daemon does the work. That's why "Docker isn't working" is so often really "the daemon isn't running" or "my user doesn't have permission to reach it," and knowing that split is what separates someone who's memorized commands from someone who understands the tool.&lt;/p&gt;

&lt;p&gt;One more distinction worth locking in early, because interviewers ask it directly: an &lt;strong&gt;image&lt;/strong&gt; is a static, read-only blueprint, built once and then reused. A &lt;strong&gt;container&lt;/strong&gt; is a running instance of that image, a live process with its own writable layer on top. One image can spin up many containers at once, each isolated from the others, the way one class definition can produce many objects. A &lt;strong&gt;volume&lt;/strong&gt; is neither of those. It's a mechanism for storing data outside a container's writable layer, so that data survives when the container is removed. We'll come back to volumes properly in the storage section, but keep the three separate in your head: image is the recipe, container is the meal, volume is the fridge you keep leftovers in after the meal's plate gets thrown away.&lt;/p&gt;

&lt;h2&gt;
  
  
  Images, layers, and registries
&lt;/h2&gt;

&lt;p&gt;An image is a stack of read-only layers plus some metadata (which command starts the container, which ports it documents, what its default working directory is). Each instruction in a Dockerfile that changes the filesystem produces one layer, and each layer is content-addressed: Docker hashes its contents, and if two images share a layer with the same hash, they share the actual bytes on disk instead of duplicating them. That's why pulling a new image that shares a base with one you already have is often fast, you're only downloading the layers you don't already have.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F73zgq26ibpaxmzrsn9vq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F73zgq26ibpaxmzrsn9vq.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the whole logic behind the build cache, and it's the single most common "do you actually understand Docker" test. Docker builds a Dockerfile top to bottom, and before running each instruction it checks whether it's seen that exact instruction with the exact same inputs before. If so, it reuses the cached layer instead of redoing the work. The moment one instruction changes, every instruction after it has to rerun, even if nothing else changed, because each layer builds on top of the one before it. That's why you copy &lt;code&gt;requirements.txt&lt;/code&gt; and install dependencies &lt;em&gt;before&lt;/em&gt; copying the rest of your application code: your code changes on every commit, but your dependencies don't, so if you copy code first, every build reinstalls every dependency from scratch. Order layers from least-likely-to-change to most-likely-to-change, and your builds stay fast for months.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="c"&gt;# Slow: any code change reinstalls every dependency&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; . .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt

&lt;span class="c"&gt;# Fast: dependency layer stays cached until requirements.txt itself changes&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; requirements.txt .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; . .&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Storage and registries.&lt;/strong&gt; Images live locally under Docker's data directory until you push them somewhere. A registry is just a storage and distribution service for images: &lt;code&gt;docker pull&lt;/code&gt; fetches from one, &lt;code&gt;docker push&lt;/code&gt; sends to one. Docker Hub is the default and where most official base images live (&lt;code&gt;python&lt;/code&gt;, &lt;code&gt;node&lt;/code&gt;, &lt;code&gt;postgres&lt;/code&gt;, &lt;code&gt;nginx&lt;/code&gt;). In a real company you'll almost always be pushing to a private registry instead, like Amazon ECR, Google Artifact Registry, GitHub Container Registry, or a self-hosted one like Harbor, so images never leave your own infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tagging is where people get careless, and it's a favourite interview trap.&lt;/strong&gt; An image reference is &lt;code&gt;name:tag&lt;/code&gt;, and if you don't specify a tag, Docker assumes &lt;code&gt;latest&lt;/code&gt;. The trap is that &lt;code&gt;latest&lt;/code&gt; doesn't mean "the newest version," it's just a tag like any other that happens to be the default. If you keep pushing new builds tagged &lt;code&gt;latest&lt;/code&gt;, the previous image becomes an untagged, unreferenced layer sitting on disk somewhere, and you've lost your ability to roll back to it by name. In production you want an explicit version or, more commonly now, the git commit SHA as the tag: &lt;code&gt;myapp:a3f9c21&lt;/code&gt;. That gives you full traceability. You always know exactly which commit is running in any environment, and rolling back means simply redeploying the previous SHA.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing a Dockerfile that isn't naive
&lt;/h2&gt;

&lt;p&gt;A Dockerfile is a plain text file of instructions Docker follows, top to bottom, to build an image. Here's a reasonably real one for a small Python service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; python:3.12-slim&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; requirements.txt .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-dir&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; . .&lt;/span&gt;
&lt;span class="k"&gt;EXPOSE&lt;/span&gt;&lt;span class="s"&gt; 8000&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["python", "app.py"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjb0dvx6tb9bmxkfxb26t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjb0dvx6tb9bmxkfxb26t.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Walking through what each line is actually doing: &lt;code&gt;FROM&lt;/code&gt; sets the base image everything else builds on top of, here a slim, minimal Python image rather than the full one, which matters for size (more on that shortly). &lt;code&gt;WORKDIR&lt;/code&gt; sets the working directory inside the image for every instruction after it, and creates the directory if it doesn't exist. &lt;code&gt;COPY&lt;/code&gt; copies files from your build context into the image; &lt;code&gt;RUN&lt;/code&gt; executes a command at build time and bakes its result into a new layer. &lt;code&gt;EXPOSE&lt;/code&gt; is worth being precise about, because it's a common source of confusion in interviews: it does not publish a port. It's purely documentation, a note in the image's metadata saying "this app listens on 8000." The thing that actually makes a port reachable from outside the container is the &lt;code&gt;-p&lt;/code&gt; flag at runtime, &lt;code&gt;docker run -p 8080:8000 myapp&lt;/code&gt;, which maps host port 8080 to container port 8000. You can &lt;code&gt;EXPOSE&lt;/code&gt; a port and never publish it, and you can publish a port you never bothered to &lt;code&gt;EXPOSE&lt;/code&gt;. &lt;code&gt;CMD&lt;/code&gt; sets the default command the container runs when it starts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;COPY&lt;/code&gt; vs &lt;code&gt;ADD&lt;/code&gt;.&lt;/strong&gt; Both copy files into the image, but &lt;code&gt;ADD&lt;/code&gt; quietly does more: it auto-extracts local &lt;code&gt;.tar&lt;/code&gt; archives into the destination, and it can fetch files from a remote URL. That second behaviour is exactly why most teams avoid it: pulling from a URL at build time means your build isn't reproducible, since the same Dockerfile can produce a different image tomorrow if that URL's contents change. Use &lt;code&gt;COPY&lt;/code&gt; by default. Reach for &lt;code&gt;ADD&lt;/code&gt; only on the rare occasion you specifically need local tar extraction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;ENV&lt;/code&gt; vs &lt;code&gt;ARG&lt;/code&gt;.&lt;/strong&gt; Both set variables, but their lifetimes differ. &lt;code&gt;ARG&lt;/code&gt; exists only during the build and isn't present in the final image or in a running container, useful for things like choosing a base image version at build time. &lt;code&gt;ENV&lt;/code&gt; sets an environment variable that persists into the running container, so anything your application reads from the environment at runtime should be &lt;code&gt;ENV&lt;/code&gt;, not &lt;code&gt;ARG&lt;/code&gt;. A subtlety worth knowing: &lt;code&gt;ARG&lt;/code&gt; values do show up in &lt;code&gt;docker history&lt;/code&gt;, so never pass secrets through build args either, they're not hidden, just short-lived.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;CMD&lt;/code&gt; vs &lt;code&gt;ENTRYPOINT&lt;/code&gt; is the classic trap, and it's worth getting exactly right.&lt;/strong&gt; Both define what runs when a container starts, but they behave differently when you pass arguments at runtime. &lt;code&gt;CMD&lt;/code&gt; sets a default command that gets fully replaced if you specify anything after the image name in &lt;code&gt;docker run&lt;/code&gt;. &lt;code&gt;ENTRYPOINT&lt;/code&gt; sets a fixed executable that always runs; anything you pass at runtime gets appended to it as arguments instead of replacing it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="c"&gt;# CMD alone — fully overridable&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["python", "app.py"]&lt;/span&gt;
&lt;span class="c"&gt;# docker run myimage python other.py  →  runs "python other.py", ignoring app.py entirely&lt;/span&gt;

&lt;span class="c"&gt;# ENTRYPOINT + CMD — entrypoint is fixed, CMD is just its default argument&lt;/span&gt;
&lt;span class="k"&gt;ENTRYPOINT&lt;/span&gt;&lt;span class="s"&gt; ["python"]&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["app.py"]&lt;/span&gt;
&lt;span class="c"&gt;# docker run myimage other.py  →  runs "python other.py"&lt;/span&gt;
&lt;span class="c"&gt;# docker run myimage           →  runs "python app.py"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The common, sensible pattern is &lt;code&gt;ENTRYPOINT&lt;/code&gt; for the fixed executable and &lt;code&gt;CMD&lt;/code&gt; for a default argument you're happy to have overridden. If you genuinely need to override the entrypoint itself at runtime, there's an explicit &lt;code&gt;--entrypoint&lt;/code&gt; flag for that; it doesn't happen by accident. For a lot of simple apps &lt;code&gt;CMD&lt;/code&gt; alone is all you need, and &lt;code&gt;ENTRYPOINT&lt;/code&gt; only earns its place when you want to lock in the executable and treat everything else as swappable arguments.&lt;/p&gt;

&lt;p&gt;There's a second, quieter trap sitting right next to this one: the JSON array form above, &lt;code&gt;CMD ["python", "app.py"]&lt;/code&gt;, is called exec form, and it matters for more than syntax. Write it instead as a plain string, &lt;code&gt;CMD python app.py&lt;/code&gt;, and Docker treats it as shell form, silently wrapping it as &lt;code&gt;/bin/sh -c "python app.py"&lt;/code&gt;. That means your app isn't actually PID 1 inside the container, the shell is, with your app running as a child process underneath it. When &lt;code&gt;docker stop&lt;/code&gt; sends &lt;code&gt;SIGTERM&lt;/code&gt;, it goes to PID 1, the shell, which generally has no idea it's supposed to forward that signal on to your app. Your app never hears about the shutdown, so it just keeps running until the grace period expires and Docker escalates to &lt;code&gt;SIGKILL&lt;/code&gt;, meaning every stop takes the full timeout and nothing gets a clean chance to exit. Exec form avoids all of this because your app becomes PID 1 directly and receives signals itself. Default to exec form for both &lt;code&gt;CMD&lt;/code&gt; and &lt;code&gt;ENTRYPOINT&lt;/code&gt; unless you have a specific reason to want the shell in between, like needing shell features such as environment variable expansion in the command itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The build context and &lt;code&gt;.dockerignore&lt;/code&gt;.&lt;/strong&gt; When you run &lt;code&gt;docker build .&lt;/code&gt;, that trailing &lt;code&gt;.&lt;/code&gt; is the build context: the entire directory Docker sends to the daemon before the build even starts. Every file in it gets transferred, whether your Dockerfile uses it or not, which matters for two reasons. A large context slows down every single build, since it all has to ship over before anything happens. And only files inside the context are available to &lt;code&gt;COPY&lt;/code&gt;; anything outside it is invisible to the build. A &lt;code&gt;.dockerignore&lt;/code&gt; file works exactly like &lt;code&gt;.gitignore&lt;/code&gt; and trims the context down, typically excluding &lt;code&gt;.git&lt;/code&gt;, virtual environments, local caches, test directories, and any build artifacts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-stage builds and cutting image size
&lt;/h2&gt;

&lt;p&gt;Multi-stage builds are the answer to a genuinely common interview scenario: "here's a 1.2 GB image, get it under 200 MB." A Dockerfile can have more than one &lt;code&gt;FROM&lt;/code&gt;, and each one starts a fresh stage. You can copy specific files from an earlier stage into a later one, and the final image only contains whatever the last stage explicitly copied in. Compilers, build tools, and test dependencies from earlier stages never make it into the image that actually ships.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="c"&gt;# Stage 1: build, with all the compiler tooling you need&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;python:3.12&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;AS&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;builder&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /build&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; requirements.txt .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-dir&lt;/span&gt; &lt;span class="nt"&gt;--target&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/build/packages &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt

&lt;span class="c"&gt;# Stage 2: the image that actually ships&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; python:3.12-slim&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --from=builder /build/packages /usr/local/lib/python3.12/site-packages&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; . .&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["python", "app.py"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkaova7evz89xamdrve52.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkaova7evz89xamdrve52.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The savings are most dramatic for compiled languages, where a Go or Java build stage can drag in gigabytes of toolchain that a runtime image never touches. For an interpreted language like Python the main win is excluding build-time-only packages like &lt;code&gt;gcc&lt;/code&gt;, needed to compile certain dependencies but useless afterward.&lt;/p&gt;

&lt;p&gt;The other lever, alongside multi-stage builds, is your base image choice. &lt;code&gt;python:3.12&lt;/code&gt; is roughly 900 MB. &lt;code&gt;python:3.12-slim&lt;/code&gt; strips out most tools you don't need at runtime and lands around 130 MB. Alpine-based images go smaller still, but Alpine uses &lt;code&gt;musl&lt;/code&gt; instead of &lt;code&gt;glibc&lt;/code&gt;, which occasionally breaks Python packages with compiled C extensions, so it's not a free upgrade; slim is usually the safer default. Beyond slim there are &lt;strong&gt;distroless&lt;/strong&gt; images, which strip out even the shell and package manager, leaving close to nothing but your app and its runtime, the smallest attack surface but also the hardest to debug, since you can't &lt;code&gt;exec&lt;/code&gt; into a shell that isn't there. Combine a slim or distroless final stage with multi-stage builds and a &lt;code&gt;.dockerignore&lt;/code&gt;, and a 1.2 GB image getting down to under 200 MB is a completely normal outcome, not a special trick.&lt;/p&gt;

&lt;h2&gt;
  
  
  Container lifecycle and the essential CLI
&lt;/h2&gt;

&lt;p&gt;Before the states themselves, one thing worth being clear on, because it's the source of a very common fresher confusion: &lt;code&gt;docker run ubuntu&lt;/code&gt; with nothing else after it starts a container and it exits immediately, often before you've even finished reading the output. This looks broken, but it isn't. A container stays alive for exactly as long as its main process, PID 1, stays alive, and nothing else. The &lt;code&gt;ubuntu&lt;/code&gt; image's default command is a shell with no terminal attached to keep it open, so that shell starts, has nothing to do, and exits, and the moment PID 1 exits, the whole container stops, regardless of anything else that happened to be running. This is genuinely different from a VM, where the "machine" stays up as its own thing independent of whatever process you happen to be running inside it. A container has no concept of "staying up" separate from its main process; the process &lt;em&gt;is&lt;/em&gt; the container's lifetime.&lt;/p&gt;

&lt;p&gt;This is also why you'll rarely see a well-designed container running more than one real service. It's tempting, especially early on, to think "why not run nginx and my app and a cron job all in one container," but Docker only supervises one PID 1, so if you cram three services into one entrypoint script, Docker has no idea if two of them silently died, it only knows whether the wrapper script itself is still alive. The convention, and the thing to say in an interview if it comes up, is one process per container: if you need several services, run several containers and let Compose or an orchestrator manage them as a group, each with its own lifecycle Docker can actually see and act on.&lt;/p&gt;

&lt;p&gt;With that settled, a container moves through a small number of states, and knowing the exact commands and signals for each transition is table stakes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;How you get there&lt;/th&gt;
&lt;th&gt;What's happening&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Created&lt;/td&gt;
&lt;td&gt;&lt;code&gt;docker create&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Filesystem is set up from the image, nothing is running yet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Running&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;docker run&lt;/code&gt; (create + start) or &lt;code&gt;docker start&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;The main process is executing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Paused&lt;/td&gt;
&lt;td&gt;&lt;code&gt;docker pause&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Processes are frozen in place, memory held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stopped&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;docker stop&lt;/code&gt; or &lt;code&gt;docker kill&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Main process has exited, filesystem still on disk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Removed&lt;/td&gt;
&lt;td&gt;&lt;code&gt;docker rm&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Container and its writable layer are gone&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;docker run&lt;/code&gt; and &lt;code&gt;docker start&lt;/code&gt; get mixed up constantly, so it's worth being precise: &lt;code&gt;docker run&lt;/code&gt; always creates a brand-new container from an image and starts it. &lt;code&gt;docker start&lt;/code&gt; restarts a container that already exists but is currently stopped. If you &lt;code&gt;docker run&lt;/code&gt; the same image five times, you get five separate containers; if you &lt;code&gt;docker start&lt;/code&gt; a stopped one, you get the same container back, with the same writable layer and any data it had.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;docker stop&lt;/code&gt; and &lt;code&gt;docker kill&lt;/code&gt; also aren't the same thing, and interviewers like this pairing because it has a real production consequence. &lt;code&gt;docker stop&lt;/code&gt; sends &lt;code&gt;SIGTERM&lt;/code&gt;, gives the process a grace period (10 seconds by default) to shut down cleanly, and only sends &lt;code&gt;SIGKILL&lt;/code&gt; if it hasn't exited by then. &lt;code&gt;docker kill&lt;/code&gt; sends &lt;code&gt;SIGKILL&lt;/code&gt; immediately, no grace period, no chance to flush a buffer or finish a write. For anything talking to a database mid-transaction, that difference matters.&lt;/p&gt;

&lt;p&gt;Containers don't restart on their own unless you tell them to, via &lt;code&gt;--restart&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Policy&lt;/th&gt;
&lt;th&gt;Behaviour&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;no&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Never restart automatically (the default)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;always&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Restart whenever it stops, even after the Docker daemon itself restarts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;on-failure[:N]&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Restart only on a non-zero exit code, optionally capped at N retries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;unless-stopped&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Like &lt;code&gt;always&lt;/code&gt;, but won't restart a container you stopped manually&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;unless-stopped&lt;/code&gt; is usually the sensible production default, since &lt;code&gt;always&lt;/code&gt; will happily restart a container you stopped for maintenance the moment the daemon comes back up, which is rarely what you want.&lt;/p&gt;

&lt;p&gt;Exit codes are worth glancing at when a container has already died: &lt;code&gt;0&lt;/code&gt; is a clean exit, and &lt;code&gt;137&lt;/code&gt; specifically means the process was killed by &lt;code&gt;SIGKILL&lt;/code&gt;, most often because the kernel's OOM killer stepped in after the container hit its memory limit. Seeing &lt;code&gt;137&lt;/code&gt; in &lt;code&gt;docker ps -a&lt;/code&gt; should make you reach straight for &lt;code&gt;docker stats&lt;/code&gt; and your memory limits, not the application logs.&lt;/p&gt;

&lt;p&gt;Here's the CLI reference worth actually knowing, not memorizing flag-by-flag, but knowing what each one is for and why you'd reach for it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Command&lt;/th&gt;
&lt;th&gt;What it's for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker build -t name .&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Build an image from a Dockerfile, tagging it &lt;code&gt;name&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker run -d -p 8080:80 name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Create and start a container, &lt;code&gt;-d&lt;/code&gt; runs it detached in the background instead of tying up your terminal in the foreground&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker run --env-file .env name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Load many environment variables from a file at once, instead of a long chain of &lt;code&gt;-e KEY=value&lt;/code&gt; flags&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;docker ps&lt;/code&gt; / &lt;code&gt;docker ps -a&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;List running containers / all containers including stopped ones&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker exec -it name sh&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Run a new process (usually a shell) inside a running container, for debugging&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker logs --tail 50 -f name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Show recent stdout/stderr, follow live with &lt;code&gt;-f&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker cp file.txt name:/app/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Copy a file between the host and a running container, in either direction, without needing a shell&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;docker stop&lt;/code&gt; / &lt;code&gt;docker kill&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Graceful shutdown (SIGTERM) vs immediate (SIGKILL)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;docker rm&lt;/code&gt; / &lt;code&gt;docker rmi&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Remove a container / remove an image&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker inspect name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Full JSON metadata: IPs, mounts, env vars, health status, everything&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;docker stats&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Live CPU, memory, and I/O usage per container&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;docker pull&lt;/code&gt; / &lt;code&gt;docker push&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Fetch an image from a registry / send one to a registry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;docker image prune&lt;/code&gt; / &lt;code&gt;docker system prune&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Clean up dangling images / clean up all unused containers, networks, and images&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A couple of these are worth a sentence more than the table gives them. &lt;code&gt;-d&lt;/code&gt; matters because without it your terminal is attached to the container's output and blocked until it exits, fine for a quick test, useless for anything you want running in the background while you keep working. &lt;code&gt;--env-file&lt;/code&gt; is the practical way to pass a real application's worth of configuration, database URLs, API keys, feature flags, without a &lt;code&gt;docker run&lt;/code&gt; command that's twenty &lt;code&gt;-e&lt;/code&gt; flags long; Compose has the equivalent &lt;code&gt;env_file:&lt;/code&gt; key for the same reason. &lt;code&gt;docker cp&lt;/code&gt; is the one people forget exists and then do something roundabout instead, like rebuilding an image just to add one file, when copying it straight into a running container takes one command.&lt;/p&gt;

&lt;p&gt;One easy-to-miss distinction: &lt;code&gt;docker exec&lt;/code&gt; starts a brand-new process inside an already-running container, which is what you want for debugging, poking around, or running a one-off script. &lt;code&gt;docker attach&lt;/code&gt; instead connects your terminal directly to the container's main process (PID 1). If you &lt;code&gt;Ctrl+C&lt;/code&gt; out of an attached session instead of detaching properly with &lt;code&gt;Ctrl+P, Ctrl+Q&lt;/code&gt;, you can send a kill signal straight to that main process and stop the container by accident. For everyday debugging, &lt;code&gt;exec&lt;/code&gt; is almost always the right tool.&lt;/p&gt;

&lt;p&gt;A last cleanup note that catches people running CI pipelines: a &lt;strong&gt;dangling image&lt;/strong&gt; is an old, untagged image layer left behind when you rebuild an image with the same name and tag, orphaning the previous version. They pile up quietly and can fill a CI runner's disk within days if nothing ever cleans them up. &lt;code&gt;docker image prune&lt;/code&gt; handles just the dangling ones; &lt;code&gt;docker system prune -a&lt;/code&gt; is far more aggressive and removes every image not currently used by a running container, so use it carefully.&lt;/p&gt;

&lt;h2&gt;
  
  
  Networking
&lt;/h2&gt;

&lt;p&gt;By default, Docker gives you a &lt;strong&gt;bridge&lt;/strong&gt; network, a private virtual network on the host that containers attach to. The default bridge is fairly limited: containers on it can only reach each other by IP address. A &lt;strong&gt;user-defined bridge network&lt;/strong&gt;, which you create yourself with &lt;code&gt;docker network create&lt;/code&gt;, is much more useful, because Docker runs an embedded DNS server on it, so containers can resolve each other by name instead of by IP. This is exactly what makes Docker Compose feel effortless: every service in a Compose file gets attached to the same user-defined network automatically, and a &lt;code&gt;web&lt;/code&gt; service can just connect to &lt;code&gt;postgres://db:5432&lt;/code&gt; using &lt;code&gt;db&lt;/code&gt; as a hostname, no IP addresses anywhere.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fae1nom9fqkgnhr0c2bo4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fae1nom9fqkgnhr0c2bo4.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There are four network drivers worth knowing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bridge&lt;/strong&gt; (the default, described above): single-host, containers reach each other by name on a shared virtual network.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Host&lt;/strong&gt;: the container skips network isolation entirely and shares the host's network stack directly. Fastest possible networking, since there's no virtual bridge in between, but no port mapping and no isolation, since the container binds directly to host ports.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;None&lt;/strong&gt;: no networking at all, useful for fully isolated batch jobs that need zero network access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Overlay&lt;/strong&gt;: extends a virtual network across &lt;em&gt;multiple&lt;/em&gt; Docker hosts, which is what makes Swarm's multi-host clustering possible in the first place. You won't reach for this on a single machine.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The one detail that trips people up in a real debugging scenario: inside a container, &lt;code&gt;localhost&lt;/code&gt; refers to that container itself, not to another container or to the host machine. If service A needs to reach service B, it has to use B's container name (or Compose service name) on a shared user-defined network, &lt;code&gt;localhost&lt;/code&gt; will just fail silently or connect to nothing.&lt;/p&gt;

&lt;p&gt;Port publishing is worth restating cleanly here since it connects back to &lt;code&gt;EXPOSE&lt;/code&gt;: &lt;code&gt;-p 8080:8000&lt;/code&gt; maps host port 8080 to container port 8000, and that's the only thing that actually makes a container reachable from outside Docker's network. Nothing else does it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Storage: volumes, bind mounts, and tmpfs
&lt;/h2&gt;

&lt;p&gt;Anything written to a container's writable layer disappears the moment that container is removed. That's fine for a stateless web server, and a real problem for a database. Docker gives you three ways to persist or share data, and interviewers care about you knowing which one fits which situation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo55jtna8xugysmxn03ic.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo55jtna8xugysmxn03ic.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Where it lives&lt;/th&gt;
&lt;th&gt;Survives container removal?&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Named volume&lt;/td&gt;
&lt;td&gt;Docker-managed location on the host&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Databases, production data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bind mount&lt;/td&gt;
&lt;td&gt;A specific path you choose on the host&lt;/td&gt;
&lt;td&gt;Yes, it's just the host filesystem&lt;/td&gt;
&lt;td&gt;Local development, live code reload&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tmpfs&lt;/td&gt;
&lt;td&gt;RAM only, never touches disk&lt;/td&gt;
&lt;td&gt;No, gone the moment the container stops&lt;/td&gt;
&lt;td&gt;Secrets, sensitive temp data&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A &lt;strong&gt;named volume&lt;/strong&gt; (&lt;code&gt;docker volume create my_data&lt;/code&gt;, then &lt;code&gt;docker run -v my_data:/app/data&lt;/code&gt;) is storage Docker manages for you, outside of any specific container's filesystem, kept at &lt;code&gt;/var/lib/docker/volumes/&lt;/code&gt; on Linux. It's the right default for a database or anything you'd genuinely be upset to lose, and it's independent of any one container's lifecycle. On macOS and Windows, Docker Desktop runs the engine inside a lightweight Linux VM under the hood, so volumes live inside that VM rather than as a directly browsable host path, worth knowing so you're not confused when you go looking for the files and they're not where you expected.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;bind mount&lt;/strong&gt; links a specific path on the host directly into the container: &lt;code&gt;docker run -v $(pwd):/app&lt;/code&gt;. It's what makes local development pleasant, since editing a file on your host shows up inside the running container instantly, no rebuild needed. The tradeoff is that it depends entirely on your host's directory layout and gives the container direct read/write access to a real host path, which is a real security concern if that container is ever compromised, and part of why bind mounts are common in development and far less common in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;tmpfs&lt;/strong&gt; mounts a filesystem that only ever lives in memory, never touching disk, and vanishes completely the instant the container stops. &lt;code&gt;docker run --tmpfs /app/tmp&lt;/code&gt;. Reach for it when you genuinely don't want data persisted anywhere, like a decrypted secret you only need for the lifetime of one process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Docker Compose
&lt;/h2&gt;

&lt;p&gt;A single Dockerfile builds one image. Real applications are rarely one container, so Compose exists to define, network, and start several containers together with one command, all described in a YAML file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;web&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;.&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8000:8000"&lt;/span&gt;
    &lt;span class="na"&gt;depends_on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;condition&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service_healthy&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;DB_HOST=db&lt;/span&gt;

  &lt;span class="na"&gt;db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres:16&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;POSTGRES_PASSWORD=devpass&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;pgdata:/var/lib/postgresql/data&lt;/span&gt;
    &lt;span class="na"&gt;healthcheck&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CMD-SHELL"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pg_isready&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;-U&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;postgres"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5s&lt;/span&gt;
      &lt;span class="na"&gt;timeout&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;3s&lt;/span&gt;
      &lt;span class="na"&gt;retries&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;

&lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pgdata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;docker compose up&lt;/code&gt; reads this file, creates a shared user-defined network automatically, builds or pulls whatever images are needed, and starts every service, in dependency order where you've declared it. &lt;code&gt;docker compose down&lt;/code&gt; tears the whole thing back down. Worth knowing as current: the standalone &lt;code&gt;docker-compose&lt;/code&gt; Python binary (with the hyphen) has been end-of-life since 2021, and what you actually run today is &lt;code&gt;docker compose&lt;/code&gt; (no hyphen), a plugin built into the Docker CLI itself. If you see a tutorial using the hyphenated form, it's outdated; both still work on most installs for backward compatibility, but the space-separated version is the one you'll actually see in current documentation and in any interview that's testing whether you're current.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp6ossi1q0z6gdx9l3lu6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp6ossi1q0z6gdx9l3lu6.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;depends_on: condition: service_healthy&lt;/code&gt; pattern in the example above is worth understanding, not just copying. A plain &lt;code&gt;depends_on: db&lt;/code&gt; only waits for the &lt;code&gt;db&lt;/code&gt; &lt;em&gt;container&lt;/em&gt; to start, not for Postgres inside it to actually be ready to accept connections, which is a very common cause of a container that keeps crash-looping on startup because it tried to connect to a database that technically existed but wasn't listening yet. Pairing a &lt;code&gt;healthcheck&lt;/code&gt; on the dependency with &lt;code&gt;condition: service_healthy&lt;/code&gt; on the dependent service fixes exactly that: &lt;code&gt;web&lt;/code&gt; won't start until Docker has actually confirmed &lt;code&gt;db&lt;/code&gt; is healthy, not just running.&lt;/p&gt;

&lt;p&gt;A few Compose commands worth knowing by name, since they come up constantly in real work and in interviews: &lt;code&gt;docker compose up -d&lt;/code&gt; starts everything in the background, &lt;code&gt;docker compose down&lt;/code&gt; stops and removes containers and the network (add &lt;code&gt;-v&lt;/code&gt; to also drop volumes), &lt;code&gt;docker compose build&lt;/code&gt; rebuilds images without starting anything, and &lt;code&gt;docker compose logs -f service_name&lt;/code&gt; tails logs from just one service instead of the whole stack.&lt;/p&gt;

&lt;p&gt;One more feature worth a mention: &lt;strong&gt;Compose profiles&lt;/strong&gt; let you tag certain services so they only start when you explicitly ask for them, useful for things like debugging tools or a monitoring stack you don't want cluttering a normal &lt;code&gt;docker compose up&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;debug-tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;my-debug-image&lt;/span&gt;
    &lt;span class="na"&gt;profiles&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;debug"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;docker compose up&lt;/code&gt; skips it entirely; &lt;code&gt;docker compose --profile debug up&lt;/code&gt; includes it. Handy for keeping one Compose file instead of maintaining several near-duplicate ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security basics
&lt;/h2&gt;

&lt;p&gt;Security questions have become more common in Docker interviews, not just for DevOps-flavoured roles, and the expectations are fairly specific.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run as a non-root user.&lt;/strong&gt; By default, whatever runs inside a container runs as root, and root inside a container that escapes its isolation is root on the host. It's a small addition to a Dockerfile:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;RUN &lt;/span&gt;adduser &lt;span class="nt"&gt;--system&lt;/span&gt; &lt;span class="nt"&gt;--group&lt;/span&gt; appuser
&lt;span class="k"&gt;USER&lt;/span&gt;&lt;span class="s"&gt; appuser&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Use a read-only filesystem where you can.&lt;/strong&gt; &lt;code&gt;docker run --read-only&lt;/code&gt; makes the container's entire filesystem read-only except for anything you've explicitly mounted as a volume or tmpfs. If your app has no legitimate reason to write to its own filesystem, this closes off an entire class of attack where a compromised process tries to write and execute something malicious inside the container.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep secrets out of environment variables and image layers.&lt;/strong&gt; Environment variables are convenient but not secure: they show up in &lt;code&gt;docker inspect&lt;/code&gt; output and often end up in logs. And anything baked into an image layer at build time, including via &lt;code&gt;ARG&lt;/code&gt;, is recoverable from &lt;code&gt;docker history&lt;/code&gt; even if you delete it in a later layer, since earlier layers are still there underneath. For a Swarm setup, Docker Secrets encrypts secrets at rest and mounts them as files under &lt;code&gt;/run/secrets/&lt;/code&gt; only inside containers that explicitly request them. Outside Swarm, most teams reach for a dedicated secrets manager: HashiCorp Vault, AWS Secrets Manager, or similar.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scan your images before they ship.&lt;/strong&gt; Tools like Trivy and Snyk check your base image and dependencies against known CVE databases, and wiring a scan step into CI catches a vulnerable base image before it ever reaches production instead of after.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never run &lt;code&gt;latest&lt;/code&gt; in production&lt;/strong&gt;, for the same rollback reason covered in the tagging section earlier, and it's worth repeating here because it's as much a security practice as a hygiene one: you want to know, with certainty, exactly which image is running where.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Freo0b2clbtnl9owrlbo3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Freo0b2clbtnl9owrlbo3.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One more thing worth a mention if the role leans toward platform or security: &lt;strong&gt;rootless Docker&lt;/strong&gt; runs the daemon itself as a non-root user on the host, so even if something escapes a container, it only has the privileges of an unprivileged host user, not host root. It comes with some tradeoffs (a handful of networking and storage features aren't available), but it's increasingly the recommended default for security-conscious setups, especially shared CI runners.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Docker stops being enough
&lt;/h2&gt;

&lt;p&gt;A single Docker host works fine until you need more than one, and that's where orchestration comes in. &lt;strong&gt;Docker Swarm&lt;/strong&gt; is Docker's own built-in orchestrator: turn a group of Docker hosts into a cluster, and Swarm handles basic load balancing, service placement, and restarting failed containers. It's genuinely simple to set up, since it's already part of the Docker CLI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Kubernetes&lt;/strong&gt; is a separate, far more capable orchestration platform: sophisticated auto-scaling, fine-grained scheduling, self-healing, rolling updates, and support for deployments at a scale Swarm was never really built for. The CNCF's 2025 annual survey put Kubernetes adoption in production at 82% of container users, up from 66% just two years earlier, and that gap keeps widening rather than closing. For most SDE interviews, especially junior and mid-level ones, you're not expected to know Kubernetes deeply unless the job description specifically calls for it. What you are expected to know is the shape of the tradeoff: Swarm is simpler and genuinely fine for smaller deployments, Kubernetes is what nearly everyone reaches for once things get big enough to need serious orchestration, and being able to say that clearly, with the reasoning behind it, covers this topic completely for the vast majority of interviews.&lt;/p&gt;

&lt;h2&gt;
  
  
  The debugging playbook
&lt;/h2&gt;

&lt;p&gt;This is where interviews increasingly go: not "define a Dockerfile instruction," but "here's a container doing something wrong, walk me through how you'd figure out why." The specific scenario changes, but a consistent method covers nearly all of them: &lt;strong&gt;logs first, then network, then a shell, then metrics&lt;/strong&gt;, in that order, because each step is cheaper and faster than the next, and most problems resolve before you reach the expensive ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A container is running but not responding to requests.&lt;/strong&gt; Start with &lt;code&gt;docker logs --tail 50 -f name&lt;/code&gt;. If the application never actually started, or crashed and restarted, the logs usually say so immediately. If the app looks fine in the logs, check whether the port is actually published: &lt;code&gt;docker inspect name&lt;/code&gt; and look at the port bindings, since a container that's healthy internally but never had &lt;code&gt;-p&lt;/code&gt; set is invisible from outside no matter how correctly it's running. If it still looks fine, get a shell inside with &lt;code&gt;docker exec -it name sh&lt;/code&gt; and check from there directly. If you've configured a &lt;code&gt;HEALTHCHECK&lt;/code&gt;, &lt;code&gt;docker inspect --format='{{.State.Health.Status}}' name&lt;/code&gt; gives you Docker's own verdict without any guesswork.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One container can't reach another.&lt;/strong&gt; This is a networking problem almost every time, not an application bug, and it's worth saying that out loud in an interview because it shows you're not about to go digging through application code for a routing issue. First, confirm both containers are actually on the same network with &lt;code&gt;docker network inspect network_name&lt;/code&gt;, since containers on different networks simply can't see each other by default. Second, and this catches almost everyone at least once: confirm you're connecting using the &lt;em&gt;container name&lt;/em&gt; (or Compose service name) as the hostname, not &lt;code&gt;localhost&lt;/code&gt;, since &lt;code&gt;localhost&lt;/code&gt; inside a container always means that container, never a different one. Third, confirm the target is actually listening on the port you expect from inside that container, using &lt;code&gt;ss -tlnp&lt;/code&gt; (the modern replacement for &lt;code&gt;netstat&lt;/code&gt;, which most slim images don't even include anymore). If the image is minimal enough that even &lt;code&gt;ss&lt;/code&gt; isn't there, check from the host side instead with &lt;code&gt;docker port container_name&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;COPY&lt;/code&gt; fails during a build even though the file is right there in the repo.&lt;/strong&gt; Check &lt;code&gt;.dockerignore&lt;/code&gt; first, it's the single most common cause by a wide margin: the file is present in your repo but explicitly excluded from the build context. Second, remember the build context is whatever directory you pointed &lt;code&gt;docker build&lt;/code&gt; at, not necessarily wherever the Dockerfile itself lives, so if you're running the build from a parent directory with &lt;code&gt;--file path/to/Dockerfile&lt;/code&gt;, your context is that parent directory, and any path in &lt;code&gt;COPY&lt;/code&gt; is relative to it, not to the Dockerfile. Third, check the exact filename: Linux is case-sensitive, so &lt;code&gt;Data.csv&lt;/code&gt; and &lt;code&gt;data.csv&lt;/code&gt; are genuinely different files, a mismatch that's invisible on a case-insensitive filesystem like macOS's default but breaks immediately inside a Linux container.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A container keeps restarting in production.&lt;/strong&gt; &lt;code&gt;docker ps -a&lt;/code&gt; first, to see the exit code of whatever's crash-looping, then &lt;code&gt;docker logs --tail 100 name&lt;/code&gt;. Three causes cover almost every real case. Exit code &lt;code&gt;137&lt;/code&gt; means the OOM killer stepped in because the container hit its memory limit, confirm with &lt;code&gt;docker stats&lt;/code&gt; and raise the limit if the workload genuinely needs it. A dependency not being ready, like a database the app tries to connect to before it's actually accepting connections, is fixed with the &lt;code&gt;depends_on: condition: service_healthy&lt;/code&gt; pattern from the Compose section. And a genuine application crash, a bad input, an unhandled null, a schema mismatch, will usually be sitting plainly in the logs once you look, and needs an actual code fix rather than an infrastructure one. The useful distinction to name out loud in an interview: infrastructure failures (memory, timing, dependencies) and application failures (bugs, bad data) get diagnosed differently and fixed in completely different places, and being able to tell which kind you're looking at quickly is most of what "debugging skill" means here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Using Docker in a CI/CD pipeline&lt;/strong&gt; deserves its own mention, since it's a near-guaranteed follow-up once a Docker conversation gets this far. The pattern that comes up over and over: build an image once, tag it with the commit SHA, run your test suite &lt;em&gt;inside that exact image&lt;/em&gt; rather than in some separate CI environment, and only push to your registry once tests pass.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# GitHub Actions, roughly&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Build image&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker build -t myapp:${{ github.sha }} .&lt;/span&gt;

&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run tests inside the built image&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker run --rm myapp:${{ github.sha }} pytest tests/&lt;/span&gt;

&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Push to registry&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;docker tag myapp:${{ github.sha }} ghcr.io/org/myapp:${{ github.sha }}&lt;/span&gt;
    &lt;span class="s"&gt;docker push ghcr.io/org/myapp:${{ github.sha }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The point of testing inside the image you're about to ship, rather than in a separate CI runner environment, is that you're validating the actual artifact that goes to production, not something that merely resembles it. Tagging with the commit SHA the whole way through means you can trace, at any point later, exactly which commit is running in any given environment, and a "what changed between this deploy and the last one" question has a one-line answer instead of an investigation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.docker.com/build/cache/" rel="noopener noreferrer"&gt;Docker's own documentation on build cache and Dockerfile best practices&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.docker.com/compose/" rel="noopener noreferrer"&gt;Docker Compose specification and reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.cncf.io/announcements/2026/01/20/kubernetes-established-as-the-de-facto-operating-system-for-ai-as-production-use-hits-82-in-2025-cncf-annual-cloud-native-survey/" rel="noopener noreferrer"&gt;CNCF 2025 Annual Survey, on Kubernetes adoption&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.docker.com/engine/security/rootless/" rel="noopener noreferrer"&gt;Docker's guidance on rootless mode&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/aquasecurity/trivy" rel="noopener noreferrer"&gt;Trivy, an open-source image and filesystem vulnerability scanner&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>docker</category>
      <category>tutorial</category>
      <category>devops</category>
      <category>software</category>
    </item>
    <item>
      <title>Redis for SDE interviews: the guide I wish I'd had</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Mon, 21 Sep 2026 00:12:26 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/redis-for-sde-interviews-the-guide-i-wish-id-had-132j</link>
      <guid>https://dev.to/shogun_the_grt/redis-for-sde-interviews-the-guide-i-wish-id-had-132j</guid>
      <description>&lt;p&gt;&lt;em&gt;How it works, what to say when an interviewer asks about it, and where it falls over.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fym2heun34m5kz1v7abdk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fym2heun34m5kz1v7abdk.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Redis comes up in a strange number of interviews. Sometimes it's a direct question, like "what happens when a key expires?" More often it hides inside a system design round. You draw a cache box, the interviewer points at it, and asks what happens when it goes down.&lt;/p&gt;

&lt;p&gt;The trouble with learning it is that most material is either a five-minute "Redis is an in-memory key-value store" intro or a reference manual. Neither is what you need. You need the middle: enough depth to explain why things work the way they do, and enough awareness to notice when an interviewer is walking you toward a trap.&lt;/p&gt;

&lt;p&gt;This post is my attempt at that middle. It goes through how Redis works, the data types, expiry and eviction, caching patterns, persistence, replication and clustering, transactions and Lua, Pub/Sub, and the system design problems where Redis actually earns its place (locks, rate limiters, leaderboards, queues). It ends with when not to use Redis, which gets asked more often than you'd think. The details match Redis 7 and 8.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Redis is and how it works
&lt;/h2&gt;

&lt;p&gt;Redis stands for Remote Dictionary Server. It's an in-memory data structure store written in C. The words "data structure" matter more than "in-memory." A plain key-value store maps a key to a blob. In Redis the key is a string and the value is a real data structure: a list, a set, a sorted set, a hash. When you create a sorted set, it simply lives as the value under whatever key you picked. Nearly everything else about Redis follows from that.&lt;/p&gt;

&lt;p&gt;People use it as a cache, a session store, a rate limiter, a leaderboard, a lightweight message broker, and sometimes as the main database for data they can afford to lose or rebuild.&lt;/p&gt;

&lt;p&gt;Here's what happens when a client sends a command:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The client opens a TCP connection to the server (port 6379 by default) and sends the command using a simple protocol called RESP. It's close enough to plain text that you can type commands into &lt;code&gt;redis-cli&lt;/code&gt; and see exactly what goes over the wire.&lt;/li&gt;
&lt;li&gt;The server runs an event loop built on &lt;code&gt;epoll&lt;/code&gt; (Linux) or &lt;code&gt;kqueue&lt;/code&gt; (BSD and macOS). One thread watches thousands of connections at once without needing a thread per client.&lt;/li&gt;
&lt;li&gt;When a command is ready, the main thread parses it, runs it against the in-memory data, and writes the reply.&lt;/li&gt;
&lt;li&gt;The next command starts only after that one finishes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last point is the heart of Redis. Commands run one at a time, on a single thread, in the order they arrive.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fytk5mdal0jy2f449o345.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fytk5mdal0jy2f449o345.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A quick note on threads, because interviewers like to poke at "Redis is single-threaded." It's true for command execution. Since Redis 6 you can enable extra I/O threads that handle reading and writing network data. Background threads handle things like fsync for the append-only file and freeing large objects. Snapshots are written by a forked child process. But the commands themselves still run one after another on the main thread. To use more CPU cores, you run more Redis instances, which is what Redis Cluster does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it's fast&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Data lives in RAM, so there are no disk seeks on the read path. Running commands on one thread means no locks and no context switching inside the server, and it also means every single command is atomic for free. The event loop lets one thread serve many clients. And the operations themselves are cheap, mostly O(1) or O(log N).&lt;/p&gt;

&lt;p&gt;A single node handles on the order of 100k operations per second. The command takes microseconds to run, and over a network you'll see sub-millisecond replies. That speed makes some habits that are terrible against SQL survivable here. Sending 100 small queries in a loop would hurt a database badly. Against Redis, each one is cheap, and you can batch them with pipelining (more on that later) to pay for one network round trip instead of 100.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What that design costs you&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Everything has to fit in memory, and memory is the most expensive place to store data. One slow command blocks every other client, which is why &lt;code&gt;KEYS *&lt;/code&gt; on a big database is a well-known way to take down production. Durability is weaker than in a relational database (the persistence section covers this). And there are no joins or ad-hoc queries. You decide how you'll read the data first, and design your keys around that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two comparisons you'll get asked&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Redis vs Memcached: Memcached is a simpler cache. It's multi-threaded and stores plain strings. Redis adds rich data types, persistence, replication, Lua scripting and Pub/Sub. If all you need is a plain cache, either works. Once you want a sorted set or an atomic counter, Redis wins.&lt;/p&gt;

&lt;p&gt;Redis vs a SQL database: they're rarely alternatives. Redis trades query flexibility and strong durability for speed and simplicity, so it usually sits in front of or next to a database instead of replacing one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data types and what's under them
&lt;/h2&gt;

&lt;p&gt;Interviewers use data types as a shortcut for "does this person understand Redis or have they only used GET and SET?" You should know each type, its main commands, its cost, and a real use for it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmgvuy5wuf7wks9fhzmjm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmgvuy5wuf7wks9fhzmjm.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Strings are the base type. A string is binary-safe and can hold up to 512 MB, though you'll almost never go near that. &lt;code&gt;SET&lt;/code&gt; and &lt;code&gt;GET&lt;/code&gt; are the obvious commands. &lt;code&gt;SET&lt;/code&gt; also takes options: &lt;code&gt;EX&lt;/code&gt; or &lt;code&gt;PX&lt;/code&gt; for a TTL, &lt;code&gt;NX&lt;/code&gt; to set only if the key doesn't exist, &lt;code&gt;XX&lt;/code&gt; to set only if it does. &lt;code&gt;INCR&lt;/code&gt; and &lt;code&gt;DECR&lt;/code&gt; treat the string as an integer and change it atomically, which is why counters are trivially safe in Redis. Cache values are usually strings, often JSON.&lt;/p&gt;

&lt;p&gt;Hashes are a map of fields to values under one key, so they fit objects nicely. &lt;code&gt;HSET user:42 name Asha city Pune&lt;/code&gt; creates one, &lt;code&gt;HGET&lt;/code&gt; reads a field, &lt;code&gt;HINCRBY&lt;/code&gt; bumps a numeric field. The classic interview question is hash versus a JSON string for an object. With a JSON string you read and rewrite the whole thing to change one field. With a hash you can update one field directly, and it's usually more memory-friendly for small objects. A JSON string is simpler if you always read and write the whole object.&lt;/p&gt;

&lt;p&gt;Lists are ordered sequences, implemented as a linked list of compact blocks (called a quicklist). Pushing and popping at either end is O(1), while reaching into the middle is O(N). &lt;code&gt;LPUSH&lt;/code&gt;, &lt;code&gt;RPUSH&lt;/code&gt;, &lt;code&gt;LPOP&lt;/code&gt;, &lt;code&gt;RPOP&lt;/code&gt; do what you'd expect, and &lt;code&gt;BRPOP&lt;/code&gt; blocks until an item arrives, which is how people built simple job queues before Streams existed. &lt;code&gt;LRANGE&lt;/code&gt; reads a slice.&lt;/p&gt;

&lt;p&gt;Sets hold unique, unordered strings. &lt;code&gt;SADD&lt;/code&gt;, &lt;code&gt;SISMEMBER&lt;/code&gt;, &lt;code&gt;SCARD&lt;/code&gt; and &lt;code&gt;SMEMBERS&lt;/code&gt; are the basics, all cheap except &lt;code&gt;SMEMBERS&lt;/code&gt; on a huge set. The interesting commands are &lt;code&gt;SINTER&lt;/code&gt;, &lt;code&gt;SUNION&lt;/code&gt; and &lt;code&gt;SDIFF&lt;/code&gt;, which give you things like mutual friends or users who share tags. Sets suit exact unique tracking, like "which users have already seen this notification."&lt;/p&gt;

&lt;p&gt;Sorted sets are the type that comes up most in system design. Each member is unique and has a floating-point score, and the set stays ordered by score. Inside, Redis combines a skip list with a hash table. The skip list gives ordered traversal and rank lookups in O(log N), and the hash table gives O(1) lookup of a member's score. &lt;code&gt;ZADD&lt;/code&gt; adds or updates a member, &lt;code&gt;ZINCRBY&lt;/code&gt; bumps a score, &lt;code&gt;ZRANGE&lt;/code&gt; (with &lt;code&gt;REV&lt;/code&gt; for descending order) reads by rank, &lt;code&gt;ZRANK&lt;/code&gt; and &lt;code&gt;ZREVRANK&lt;/code&gt; give a member's position, and &lt;code&gt;ZRANGEBYSCORE&lt;/code&gt;-style queries read by score. Leaderboards, rate limiters, priority queues and "delayed job" schedulers are all sorted sets in disguise.&lt;/p&gt;

&lt;p&gt;Streams are an append-only log. Every entry gets an ID (a millisecond timestamp plus a sequence number) and holds field-value pairs. &lt;code&gt;XADD&lt;/code&gt; appends, &lt;code&gt;XREAD&lt;/code&gt; reads, and consumer groups (&lt;code&gt;XREADGROUP&lt;/code&gt;, &lt;code&gt;XACK&lt;/code&gt;) let several workers share a stream with acknowledgements. We'll come back to these in the queue section.&lt;/p&gt;

&lt;p&gt;Bitmaps aren't a separate type. They're bit operations (&lt;code&gt;SETBIT&lt;/code&gt;, &lt;code&gt;GETBIT&lt;/code&gt;, &lt;code&gt;BITCOUNT&lt;/code&gt;, &lt;code&gt;BITOP&lt;/code&gt;) on a string. If you use a user ID as the bit offset, one bit per user per day costs about 12.5 MB for 100 million users. That's how you track daily active users cheaply.&lt;/p&gt;

&lt;p&gt;HyperLogLog counts unique items approximately. &lt;code&gt;PFADD&lt;/code&gt; adds an item, &lt;code&gt;PFCOUNT&lt;/code&gt; returns the estimated number of distinct items. It uses about 12 KB per key no matter how many items you add, with a standard error around 0.8%. The catch is that you can't list what's inside or remove things. If an interviewer asks "how do you count unique visitors across billions of events," this is the answer, along with an honest note about the error.&lt;/p&gt;

&lt;p&gt;Geo commands (&lt;code&gt;GEOADD&lt;/code&gt;, &lt;code&gt;GEOSEARCH&lt;/code&gt;) store coordinates and find points within a radius or box. Under the hood it's a sorted set where the score is a geohash of the coordinates. The search grabs candidates from grid-aligned boxes first, then filters to the exact radius.&lt;/p&gt;

&lt;p&gt;Redis 8 also ships probabilistic structures like Bloom filters, plus JSON and time series support, in the core. On older versions these came from separate Redis Stack modules. A Bloom filter says either "definitely not in the set" or "probably in the set," never the other way round, and it will matter when we get to cache penetration.&lt;/p&gt;

&lt;p&gt;One habit that's worth building early: name keys like &lt;code&gt;user:42:profile&lt;/code&gt; or &lt;code&gt;order:9001:items&lt;/code&gt;. Colons are just a convention, but every team uses them, and your key design is how you'll shard the data later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Expiration and eviction
&lt;/h2&gt;

&lt;p&gt;People mix these two up constantly, so keep them apart in your head. Expiration is about staleness: a key has a time to live and disappears when it runs out. Eviction is about memory: when Redis is full, it throws away keys to make room. They're separate mechanisms with separate settings.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fscmz0az0tm1lk00ynb3k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fscmz0az0tm1lk00ynb3k.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expiration.&lt;/strong&gt; You set a TTL with &lt;code&gt;EXPIRE key 60&lt;/code&gt;, or directly in &lt;code&gt;SET key value EX 60&lt;/code&gt;. &lt;code&gt;TTL&lt;/code&gt; tells you how many seconds remain, and &lt;code&gt;PERSIST&lt;/code&gt; removes the TTL. One gotcha: a plain &lt;code&gt;SET&lt;/code&gt; on an existing key wipes its TTL unless you pass &lt;code&gt;KEEPTTL&lt;/code&gt;. That has caused real bugs.&lt;/p&gt;

&lt;p&gt;Redis removes expired keys in two ways. When a client touches an expired key, Redis notices and deletes it on the spot (lazy expiration). Separately, a background task samples keys that have TTLs several times a second and removes the expired ones. So an expired key can sit in memory for a short while, but you'll never read it. Redis guarantees that you won't see a value after its TTL has passed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Eviction.&lt;/strong&gt; By default Redis has no memory limit on a 64-bit machine. You set one with &lt;code&gt;maxmemory 2gb&lt;/code&gt;, and you choose what happens at the limit with &lt;code&gt;maxmemory-policy&lt;/code&gt;. The default policy is &lt;code&gt;noeviction&lt;/code&gt;, which means Redis refuses new writes once it's full. That surprises people who assumed a cache would just make room. For a cache, you pick something else:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;allkeys-lru&lt;/code&gt; evicts the least recently used keys from the whole keyspace. It's a good default when some keys are much more popular than others, which is true of most workloads.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;allkeys-lfu&lt;/code&gt; evicts the least frequently used keys. It's better when a key that was popular for a long time shouldn't be dropped just because nobody touched it in the last minute.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;allkeys-random&lt;/code&gt; evicts random keys. It suits workloads that scan through everything evenly.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;volatile-&lt;/code&gt; versions (&lt;code&gt;volatile-lru&lt;/code&gt;, &lt;code&gt;volatile-lfu&lt;/code&gt;, &lt;code&gt;volatile-random&lt;/code&gt;) only consider keys that have a TTL. &lt;code&gt;volatile-ttl&lt;/code&gt; evicts keys with the least time left.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If no keys have a TTL, the &lt;code&gt;volatile-&lt;/code&gt; policies behave exactly like &lt;code&gt;noeviction&lt;/code&gt;. So if you pick one and forget to set TTLs, your cache will start rejecting writes. Newer versions also add "least recently modified" policies, but the ones above cover what interviews ask.&lt;/p&gt;

&lt;p&gt;Redis's LRU is an approximation. It doesn't track the exact order of every key, which would cost memory. Instead it samples a handful of keys (the &lt;code&gt;maxmemory-samples&lt;/code&gt; setting, 5 by default) and evicts the best candidate among them. LFU works similarly, using a small counter per key that decays over time.&lt;/p&gt;

&lt;p&gt;A practical detail: eviction happens when a command that would add data arrives while memory is over the limit. And Redis needs headroom beyond &lt;code&gt;maxmemory&lt;/code&gt;, for replication buffers, fragmentation, and the extra pages copied while a child process is writing a snapshot. Setting &lt;code&gt;maxmemory&lt;/code&gt; to 100% of the machine's RAM is a common way to get the process killed by the operating system.&lt;/p&gt;

&lt;p&gt;If you're asked "cache is full, what happens?" the answer is: it depends on &lt;code&gt;maxmemory-policy&lt;/code&gt;, and the default rejects writes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caching patterns and how they fail
&lt;/h2&gt;

&lt;p&gt;Caching is the most common reason Redis is in a diagram, and the part interviewers dig into.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The four patterns.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Cache-aside is the default. The application checks Redis first. On a miss it reads the database, writes the result to Redis (with a TTL), and returns it. The app talks to both systems, and the cache only ever holds data someone asked for.&lt;/p&gt;

&lt;p&gt;Read-through looks the same from the outside, but the cache layer loads the data itself on a miss and the app only talks to the cache. Redis has no built-in loader, so read-through means a library or a small service that wraps Redis and your database.&lt;/p&gt;

&lt;p&gt;Write-through sends every write to the cache and the database together. Reads stay fresh, writes cost a bit more.&lt;/p&gt;

&lt;p&gt;Write-behind (also called write-back) writes to the cache and flushes to the database later, asynchronously. Writes are fast, but if Redis dies before the flush, those writes are gone, and you have to handle ordering and retries.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F67zuusfzgc9fdp7ud1l3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F67zuusfzgc9fdp7ud1l3.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Invalidation.&lt;/strong&gt; The three ways to keep cached data from going stale are a TTL, deleting the key when the underlying data changes, and write-through. Most systems use a TTL as a safety net and also delete on write.&lt;/p&gt;

&lt;p&gt;On writes, deleting the cached key is usually safer than updating it. Two writers racing to update the cache can leave the older value winning. Even deleting has a race: a reader misses, reads the old row from the database, and gets delayed. Meanwhile a writer updates the database and deletes the key. Then the slow reader finally writes the old value into the cache. Now the cache is stale until the TTL runs out. It's rare, but it's why you always want a TTL, and why some teams delete the key a second time after a short delay.&lt;/p&gt;

&lt;p&gt;The cache is also eventually consistent by nature. Reads from a replica can lag behind the primary too. If your product can't tolerate any staleness, say so in the interview and explain how you'd narrow it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The three failure modes.&lt;/strong&gt; These names are used inconsistently across blog posts. One popular article uses "penetration" for the hot-key case, while most sources use it for missing keys. Define the term in your own words before you answer, and you'll be fine.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn1bz10rsum349n34lzne.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn1bz10rsum349n34lzne.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cache penetration is when lots of requests ask for keys that don't exist anywhere, so they miss the cache every time and hit the database every time. Someone scanning random IDs does this, deliberately or not. There are two standard fixes. You can cache the "not found" answer for a short time (a null value with a small TTL). Or you can put a Bloom filter in front holding every valid key, and reject requests that the filter says are definitely absent. The filter can have false positives, which just means an occasional wasted lookup.&lt;/p&gt;

&lt;p&gt;Cache stampede (also called the thundering herd, or "breakdown") is when one very popular key expires and, in the split second before it's refilled, hundreds or thousands of requests all miss and all hit the database to rebuild the same value. Three fixes come up. The first is a mutex: the first request to miss grabs a lock with &lt;code&gt;SET lock:key token NX PX 5000&lt;/code&gt; and rebuilds the value, while the others wait a moment and retry, or serve a slightly stale copy. The second is probabilistic early expiration (sometimes called X-Fetch), where each request has a small, growing chance of refreshing the value shortly before it actually expires, so one request quietly refreshes it while everyone else keeps reading the old one. The third is never letting the hottest keys expire at all, and refreshing them from a background job.&lt;/p&gt;

&lt;p&gt;Cache avalanche is many keys expiring at once, or the cache itself going down. If you filled the cache at 9:00 with a 60 minute TTL on everything, then at 10:00 the database gets the entire load at once. The fix is jitter: add a random amount to each TTL so expiry spreads out. For the "cache is down" case, you need a plan too: rate limit or shed load in front of the database, and maybe replicate Redis so one failure doesn't empty the cache.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hot keys.&lt;/strong&gt; Sometimes traffic doesn't spread evenly. Imagine a cluster of 100 nodes caching product data, and one product goes viral. The single node holding that key takes as much traffic as the other 99 combined. Adding nodes doesn't help, because the key lives on exactly one of them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flsdxzbtgk1ptkme6j5ut.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flsdxzbtgk1ptkme6j5ut.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There are three usual remedies, each with a cost. A small in-process cache on each app server absorbs most reads for the hottest keys, at the price of data that can be stale for as long as that local cache's TTL. Key copies store the same value under several names (&lt;code&gt;product:123:1&lt;/code&gt; through &lt;code&gt;product:123:10&lt;/code&gt;), which hash to different nodes, and readers pick one at random. Writes then have to update every copy. Read replicas add read capacity, but only if your clients are set up to read from replicas, and they do nothing for a key that's hot for writes.&lt;/p&gt;

&lt;p&gt;Spotting a possible hot key in your own design, before the interviewer asks, is the kind of thing that stands out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Persistence
&lt;/h2&gt;

&lt;p&gt;Redis lives in memory, but it can write to disk so it survives a restart. There are two mechanisms, and you can use them together or turn both off.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnw9bjhk5r64lj0ut1puz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnw9bjhk5r64lj0ut1puz.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RDB snapshots.&lt;/strong&gt; Redis writes the whole dataset to a compact file (&lt;code&gt;dump.rdb&lt;/code&gt;) at intervals. A config line like &lt;code&gt;save 60 1000&lt;/code&gt; means "snapshot if at least 1000 keys changed in the last 60 seconds." You can also trigger it with &lt;code&gt;BGSAVE&lt;/code&gt;. To do this without pausing, Redis forks: the child process writes the file while the parent keeps serving clients, and copy-on-write means memory is only duplicated for pages that change during the snapshot.&lt;/p&gt;

&lt;p&gt;RDB files are small, great for backups, and load fast on restart. The downsides are that you lose everything since the last snapshot on a crash (often minutes), and &lt;code&gt;fork()&lt;/code&gt; on a very large dataset can pause Redis for a noticeable moment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AOF (append-only file).&lt;/strong&gt; Redis logs every write command, and on restart it replays the log to rebuild the data. AOF is off by default (&lt;code&gt;appendonly yes&lt;/code&gt; turns it on). How much you can lose depends on the &lt;code&gt;appendfsync&lt;/code&gt; setting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;always&lt;/code&gt; flushes to disk on every write batch. It's the safest and by far the slowest.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;everysec&lt;/code&gt; flushes once a second, and it's the default. You can lose up to about a second of writes.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;no&lt;/code&gt; leaves flushing to the operating system, which is fast and least safe.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The log grows forever if you leave it, so Redis periodically rewrites it in the background into the shortest list of commands that produces the current data. Since Redis 7 the AOF is split into a base file plus incremental files, tracked by a manifest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which should you use?&lt;/strong&gt; Redis's own guidance is to use both if you want durability comparable to a traditional database. If you can live with a few minutes of loss, RDB alone is fine. For a pure cache, you can use neither. If both are on and Redis restarts, it loads the AOF, because that's the most complete record.&lt;/p&gt;

&lt;p&gt;The honest answer to "is Redis durable?" is "only as durable as you configure it, and never quite like a database that flushes every commit." With the default &lt;code&gt;everysec&lt;/code&gt;, an acknowledged write can still be lost. And replication adds another window on top, which brings us to the next section.&lt;/p&gt;

&lt;h2&gt;
  
  
  Replication, high availability and scaling
&lt;/h2&gt;

&lt;p&gt;One Redis node has two limits: if it dies you lose availability, and it can only hold as much as one machine's memory. Replication, Sentinel and Cluster each deal with a piece of that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Replication.&lt;/strong&gt; You point a replica at a primary (&lt;code&gt;REPLICAOF host port&lt;/code&gt;). The replica does a full sync first: the primary produces an RDB snapshot and sends it over, followed by the writes that piled up during the transfer. After that the primary streams every write command to the replica. If the connection drops briefly, the replica reconnects and asks to continue from its last offset. If that offset is still in the primary's replication backlog (a fixed-size buffer), only the missing part is sent (a partial resync). Otherwise it's another full sync.&lt;/p&gt;

&lt;p&gt;The thing to remember is that replication is asynchronous. The primary acknowledges your write before the replica has seen it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fomo2c83qg0hd2xgx4iwc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fomo2c83qg0hd2xgx4iwc.png" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So if the primary dies right after acknowledging a write, and a replica that hasn't received it gets promoted, that write is gone. This is the core reason Redis isn't a system of record, and it shows up again in the distributed lock section. The &lt;code&gt;WAIT&lt;/code&gt; command makes a client block until N replicas confirm a write, which shrinks the window but doesn't eliminate it.&lt;/p&gt;

&lt;p&gt;Replicas are read-only by default. You can spread reads across them, but every replica read can be a little behind the primary. That's fine for a product page and wrong for something like "did my payment go through."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sentinel.&lt;/strong&gt; Replication alone doesn't promote anyone when a primary dies. Sentinel is a set of separate processes that watch the primary and replicas and handle failover. You run at least three, on different machines. When a Sentinel can't reach the primary, it marks it as "subjectively down." Once enough Sentinels agree (the quorum you configured), it's "objectively down." The Sentinels then elect a leader, and that leader picks the best replica (by priority and how much data it has) and promotes it, and reconfigures the others. Clients ask Sentinel "who is the current primary?" instead of hardcoding an address.&lt;/p&gt;

&lt;p&gt;Sentinel gives you high availability, and nothing more. All your data still lives on one primary, so it doesn't add capacity. It also can't bring back writes lost to asynchronous replication.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cluster.&lt;/strong&gt; When the data no longer fits on one machine, or one machine can't keep up with the traffic, you shard it, and Redis Cluster is the built-in way.&lt;/p&gt;

&lt;p&gt;The keyspace is divided into 16,384 hash slots. A key's slot is &lt;code&gt;CRC16(key) mod 16384&lt;/code&gt;. Each primary owns a range of slots, and each primary usually has one or more replicas for failover. Nodes exchange information with each other using a gossip protocol, so every node knows the full slot-to-node map.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxb1ui8vbe5nz5j69ui8x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxb1ui8vbe5nz5j69ui8x.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Clients cache the slot map and send each command straight to the right node. If a slot has moved, the node answers with a &lt;code&gt;MOVED&lt;/code&gt; error that includes the correct address, and the client refreshes its map. During a resharding operation you might see &lt;code&gt;ASK&lt;/code&gt;, which is a one-time redirect for a slot that's midway through moving. Nodes don't forward requests for you, and there's no router that splits a request across nodes and merges the results.&lt;/p&gt;

&lt;p&gt;That last point is what makes Cluster feel restrictive. Commands that touch several keys (&lt;code&gt;MGET&lt;/code&gt;, &lt;code&gt;SINTER&lt;/code&gt;, &lt;code&gt;MULTI&lt;/code&gt; transactions, Lua scripts) only work when every key is in the same slot. To force that, you use hash tags: only the part of the key inside curly braces is hashed. &lt;code&gt;{user:42}:posts&lt;/code&gt; and &lt;code&gt;{user:42}:likes&lt;/code&gt; always land in the same slot. The takeaway is that with Cluster, how you name your keys is how you scale.&lt;/p&gt;

&lt;p&gt;A few more things worth knowing. Failover in Cluster works without Sentinel: replicas get promoted when a majority of primaries agree a primary has failed. You need at least three primaries to have a majority. Cluster only supports database 0. And during a network partition, a primary stuck on the minority side can keep accepting writes for a short time before it stops, and those writes can be lost.&lt;/p&gt;

&lt;p&gt;Interviewers sometimes ask how hash slots differ from consistent hashing. With consistent hashing, nodes and keys sit on a ring and a key belongs to the next node clockwise, so adding a node automatically takes over part of a neighbour's range. Redis uses a fixed number of slots and an explicit table of who owns which. Moving data means moving specific slots, which gives operators direct control, at the cost of that table being something the cluster has to keep consistent.&lt;/p&gt;

&lt;p&gt;To choose between them: if everything fits on one machine and you want failover, primary plus replicas with Sentinel is enough (or a managed service that does the same). If you need more memory or write throughput than one node offers, you need Cluster.&lt;/p&gt;

&lt;h2&gt;
  
  
  Atomicity: transactions, Lua and pipelining
&lt;/h2&gt;

&lt;p&gt;Because commands run one at a time, every single command is atomic. &lt;code&gt;INCR&lt;/code&gt; can't lose an update the way &lt;code&gt;GET&lt;/code&gt; followed by &lt;code&gt;SET&lt;/code&gt; can, because two clients doing GET then SET can both read 5 and both write 6. That's why &lt;code&gt;INCR&lt;/code&gt; exists.&lt;/p&gt;

&lt;p&gt;The trouble starts when you need several steps to happen together. Redis has three tools that get confused with each other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Transactions with MULTI and EXEC.&lt;/strong&gt; &lt;code&gt;MULTI&lt;/code&gt; starts queuing commands, and &lt;code&gt;EXEC&lt;/code&gt; runs the queue. Nothing from any other client runs in the middle of it. But there's no rollback. If one command fails at runtime (say, &lt;code&gt;INCR&lt;/code&gt; on a key holding text), the others still run. If a command is malformed and fails while being queued, the whole transaction is refused at &lt;code&gt;EXEC&lt;/code&gt;. That's very different from a SQL transaction, and interviewers like to check that you know it.&lt;/p&gt;

&lt;p&gt;For optimistic concurrency, use &lt;code&gt;WATCH key&lt;/code&gt; before &lt;code&gt;MULTI&lt;/code&gt;. If the watched key changes before &lt;code&gt;EXEC&lt;/code&gt;, the transaction aborts, and you retry. It's the Redis version of compare-and-swap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lua scripts.&lt;/strong&gt; &lt;code&gt;EVAL&lt;/code&gt; sends a script to the server, and the whole script runs as a single command. Nothing else interleaves, and you can read a value, make a decision, and write, all atomically. That's the thing MULTI can't do, because inside MULTI you can't use an earlier result to decide a later command. Lua is how you build a safe lock release or a rate limiter. Two cautions: a long script blocks everybody, and in Cluster you must pass every key the script touches as a key argument so Redis can check they share a slot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pipelining.&lt;/strong&gt; A pipeline just sends many commands without waiting for each reply, then reads all the replies. It saves network round trips and can speed things up dramatically. It is not atomic. Other clients' commands can slip in between yours.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzxnemmb7wcp5ngcuzd9p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzxnemmb7wcp5ngcuzd9p.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Pub/Sub
&lt;/h2&gt;

&lt;p&gt;Pub/Sub is a messaging pattern where publishers send messages to a named channel and every client subscribed to that channel at that moment receives them. Publishers and subscribers don't know each other exist.&lt;/p&gt;

&lt;p&gt;The commands are short: &lt;code&gt;SUBSCRIBE news&lt;/code&gt; listens on a channel, &lt;code&gt;PSUBSCRIBE news:*&lt;/code&gt; listens on a pattern, and &lt;code&gt;PUBLISH news "hello"&lt;/code&gt; sends a message and returns how many subscribers received it. A channel needs no setup. It exists as long as someone is subscribed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2cfc93fy7801vqtg4u3v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2cfc93fy7801vqtg4u3v.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The one property you must state clearly is that delivery is at-most-once and nothing is stored. If a subscriber is offline, or its connection blips, it never sees the message. There are no acknowledgements, no replay and no history. A subscriber that can't keep up gets disconnected once its output buffer passes a limit, so it doesn't eat the server's memory. And a client that's subscribed can't run normal commands on that connection under the older protocol, so apps use a separate connection for their other work.&lt;/p&gt;

&lt;p&gt;There's a common outdated claim about scaling. In classic cluster Pub/Sub, every message was broadcast to every node, so adding nodes didn't add capacity. Since Redis 7 there's sharded Pub/Sub (&lt;code&gt;SPUBLISH&lt;/code&gt; and &lt;code&gt;SSUBSCRIBE&lt;/code&gt;), where a channel is hashed to a slot like a key, and only that shard handles it. Capacity grows with the cluster. Connection cost is per node, not per channel, so millions of channels doesn't mean millions of connections.&lt;/p&gt;

&lt;p&gt;How does it compare with the alternatives? Streams keep messages, let consumers catch up after downtime, and support acknowledgements. They give at-least-once delivery. Kafka does the same at much larger scale with long retention and replay for many independent consumers. Pub/Sub is the right tool when you only care about whoever is listening right now and losing a message is acceptable.&lt;/p&gt;

&lt;p&gt;In system design, that usually means something like chat across many app servers. Each server subscribes to the channels its connected users care about. When someone sends a message, it's published once, every server with an interested user receives it, and pushes it down the WebSocket. Live dashboards and presence indicators work the same way. If the requirement says offline users must receive it later, Pub/Sub alone isn't enough. Write the message to a database or a stream too.&lt;/p&gt;

&lt;p&gt;A related trap: keyspace notifications (getting an event when a key expires, for example) are delivered over Pub/Sub, so they inherit the same lack of guarantees. Don't build correctness-critical logic on them.&lt;/p&gt;

&lt;h2&gt;
  
  
  System design problems where Redis shows up
&lt;/h2&gt;

&lt;p&gt;This is where everything above gets used. For each one, say which data structure you'd pick and why, then say what could go wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A distributed lock.&lt;/strong&gt; Say several app servers must not do the same thing at once, like two people booking the same concert seat. Redis works as a lock manager because it's one shared place every server can reach, and each command is atomic. The lock is just a key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SET lock:seat:343 &amp;lt;random-token&amp;gt; NX PX 30000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;NX&lt;/code&gt; means "only set it if it doesn't exist," so exactly one caller gets &lt;code&gt;OK&lt;/code&gt;. &lt;code&gt;PX 30000&lt;/code&gt; gives the lock a 30 second lease, so a crashed process can't hold it forever. The random token identifies who owns it.&lt;/p&gt;

&lt;p&gt;Releasing is where people slip. If you just &lt;code&gt;DEL&lt;/code&gt; the key, you might delete someone else's lock, because yours may have expired while you were still working and another client took it. So you delete only if the token still matches, and the check and delete have to be atomic, which means a Lua script:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight lua"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"GET"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;ARGV&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"DEL"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3xs0t4gybevwi1rwufnj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3xs0t4gybevwi1rwufnj.png" alt=" " width="800" height="413"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now the part interviewers actually care about, which is why this isn't a real guarantee. If Client A stalls (a long garbage collection pause, a slow network) past the lease, the lock expires, Client B takes it, and then Client A wakes up and carries on, believing it still holds the lock. Both write. A second problem is replication: if the primary grants the lock and dies before the replica hears about it, the promoted replica happily grants the same lock again.&lt;/p&gt;

&lt;p&gt;The Redlock algorithm tries to solve the second problem by acquiring the lock on a majority of independent Redis nodes. It's controversial. Martin Kleppmann wrote a well-known critique showing that it still can't survive the paused-client problem. The standard defence for that is a fencing token: each lock grant comes with an increasing number, and the storage layer rejects any write carrying an older number than one it has already seen. Redis doesn't give you that out of the box.&lt;/p&gt;

&lt;p&gt;So a Redis lock is good for efficiency (avoiding duplicate work most of the time) and shaky for correctness. If a stale lock holder would corrupt data, enforce the rule where the data lives. A &lt;code&gt;SELECT ... FOR UPDATE&lt;/code&gt; row lock or a conditional update like &lt;code&gt;UPDATE ... WHERE version = 7&lt;/code&gt; in your database often removes the need for a distributed lock entirely. Saying this out loud in an interview scores well.&lt;/p&gt;

&lt;p&gt;A cousin of the lock, for things like ticket booking, is a temporary hold: &lt;code&gt;SET seat:343:A12 user42 NX EX 600&lt;/code&gt; reserves a seat for ten minutes, and the TTL releases it if the user walks away.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A rate limiter.&lt;/strong&gt; The simplest version is a fixed window. For each user and time window you keep a counter: run &lt;code&gt;INCR&lt;/code&gt; on a key like &lt;code&gt;rl:user42:1700000000&lt;/code&gt;, and if the result goes over the limit, reject with a 429 and a &lt;code&gt;Retry-After&lt;/code&gt; header. One subtlety trips people up. You want to set the expiry only when &lt;code&gt;INCR&lt;/code&gt; returns 1, meaning the first request in the window. If you call &lt;code&gt;EXPIRE&lt;/code&gt; on every request, steady traffic keeps pushing the expiry forward and the window never resets. And if the process crashes between &lt;code&gt;INCR&lt;/code&gt; and &lt;code&gt;EXPIRE&lt;/code&gt;, you've got a counter that never expires. So do both steps in one Lua script:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight lua"&gt;&lt;code&gt;&lt;span class="kd"&gt;local&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"INCR"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt;
  &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"EXPIRE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;ARGV&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The weakness of a fixed window is bursts at the edges. A user can send the full limit at the very end of one window and again at the start of the next, doubling the rate for a moment.&lt;/p&gt;

&lt;p&gt;A sliding window fixes that with a sorted set per user, where each request is a member and its timestamp is the score. On each request you drop entries older than the window, count what's left, and add the new one if the count is under the limit. Again, do it in one script so it's atomic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight lua"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- KEYS[1] = key, ARGV[1] = now (ms), ARGV[2] = window (ms)&lt;/span&gt;
&lt;span class="c1"&gt;-- ARGV[3] = limit, ARGV[4] = a unique id for this request&lt;/span&gt;
&lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"ZREMRANGEBYSCORE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ARGV&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;ARGV&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"ZCARD"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nb"&gt;tonumber&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ARGV&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt;
  &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"ZADD"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;ARGV&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;ARGV&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
  &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"PEXPIRE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;ARGV&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use a unique member for each request. Two requests in the same millisecond would otherwise overwrite each other. The cost is memory: the set holds up to &lt;code&gt;limit&lt;/code&gt; entries per user. If that's too much, mention the token bucket or a sliding window counter as cheaper approximations, and be ready to explain the trade-off between accuracy and memory.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr7ebpc1hofkrela5a9ck.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr7ebpc1hofkrela5a9ck.png" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A leaderboard.&lt;/strong&gt; This is a sorted set, and it's about the easiest system design win there is. &lt;code&gt;ZADD board 1500 alice&lt;/code&gt; sets a score (and re-adding the same member just updates it), &lt;code&gt;ZINCRBY board 10 alice&lt;/code&gt; adds points, &lt;code&gt;ZREVRANGE board 0 9 WITHSCORES&lt;/code&gt; returns the top ten, and &lt;code&gt;ZREVRANK board alice&lt;/code&gt; gives a player's rank. Every operation is O(log N) plus the size of the slice you read. To keep only the top N, trim with &lt;code&gt;ZREMRANGEBYRANK board 0 -101&lt;/code&gt;, which removes everything below the top 100. Ties are broken by member name, so if ties matter, say how you'd handle them. If someone pushes you to a huge scale, note that one sorted set is one key on one node, so you'd split it by region or by score range and merge for global queries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A job queue.&lt;/strong&gt; Lists with &lt;code&gt;LPUSH&lt;/code&gt; and &lt;code&gt;BRPOP&lt;/code&gt; work, but if a worker pops a job and then dies, the job is gone. Streams with consumer groups fix that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;XADD jobs * type email to a@b.com
XGROUP CREATE jobs workers 0 MKSTREAM
XREADGROUP GROUP workers worker-1 COUNT 10 BLOCK 5000 STREAMS jobs &amp;gt;
XACK jobs workers 1700000000000-0
XAUTOCLAIM jobs workers worker-2 60000 0-0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A producer appends with &lt;code&gt;XADD&lt;/code&gt;. Workers in a group read with &lt;code&gt;XREADGROUP&lt;/code&gt;, and Redis tracks each delivered but unacknowledged entry as pending for that worker. When the worker finishes, it calls &lt;code&gt;XACK&lt;/code&gt;. If a worker dies, its pending entries sit idle, and another worker claims them with &lt;code&gt;XAUTOCLAIM&lt;/code&gt; after they've been idle long enough (60 seconds in the example).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9dzkvfnmsrk42hvbo7az.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9dzkvfnmsrk42hvbo7az.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The catch: this gives you at-least-once processing. Redis can't tell a slow worker from a dead one, so a job can occasionally run twice. Make your jobs idempotent. Also, stream entries are only as durable as your persistence settings, so with defaults a crash can lose recent ones. Streams fit modest queues where you already run Redis, like background jobs, notification fan-out and work distribution. Kafka is the better answer when you need long retention, replay for many independent consumers, or throughput where losing a message is unacceptable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A session store.&lt;/strong&gt; Store each session as a hash (or a JSON string) under &lt;code&gt;session:&amp;lt;id&amp;gt;&lt;/code&gt; with a TTL, and refresh the TTL on each request so active users stay logged in. Logging out is a &lt;code&gt;DEL&lt;/code&gt;. It's fast and the TTL cleans up abandoned sessions. The question to raise yourself is what happens if Redis restarts and forgets everyone. If that's unacceptable, turn on AOF, or keep sessions in a database and use Redis as a cache in front of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Counting things.&lt;/strong&gt; &lt;code&gt;INCR&lt;/code&gt; for exact counters. HyperLogLog (&lt;code&gt;PFADD&lt;/code&gt;, &lt;code&gt;PFCOUNT&lt;/code&gt;) for approximate unique counts in tiny memory. Bitmaps for per-user yes/no flags across a huge user base, like daily active users, and &lt;code&gt;BITOP AND&lt;/code&gt; across days to find users who came back. Geo commands for "nearby" queries.&lt;/p&gt;

&lt;h2&gt;
  
  
  When not to use Redis
&lt;/h2&gt;

&lt;p&gt;Interviewers ask this to see whether you understand the limits, so have an answer ready.&lt;/p&gt;

&lt;p&gt;Don't make it the only copy of data you can't lose. Between asynchronous replication and the persistence loss windows, acknowledged writes can vanish. Don't use it when your working set can't fit in memory at a cost that makes sense. Don't expect query flexibility: there are no joins, no secondary indexes out of the box, and in a cluster, multi-key operations only work within one slot. And don't use it as a replayable event log with long retention for many consumers, which is what Kafka is built for.&lt;/p&gt;

&lt;h2&gt;
  
  
  A few operational things worth knowing
&lt;/h2&gt;

&lt;p&gt;You won't be asked to be a Redis administrator, but a handful of commands and habits make you sound like someone who has run it.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;INFO&lt;/code&gt; prints server stats in sections, such as memory, persistence, replication and stats. The &lt;code&gt;keyspace_hits&lt;/code&gt; and &lt;code&gt;keyspace_misses&lt;/code&gt; counters let you work out your cache hit ratio. &lt;code&gt;SLOWLOG GET&lt;/code&gt; shows the slowest recent commands. &lt;code&gt;MONITOR&lt;/code&gt; streams every command the server receives, which is handy for debugging and expensive in production, so avoid leaving it running. Use &lt;code&gt;SCAN&lt;/code&gt; to walk the keyspace in small steps instead of &lt;code&gt;KEYS&lt;/code&gt;, and &lt;code&gt;UNLINK&lt;/code&gt; to delete big keys without blocking (&lt;code&gt;DEL&lt;/code&gt; frees memory on the main thread). &lt;code&gt;redis-cli --bigkeys&lt;/code&gt; and &lt;code&gt;MEMORY USAGE key&lt;/code&gt; help you find oversized keys.&lt;/p&gt;

&lt;p&gt;When latency spikes, the usual suspects are a slow command on a big collection, a large key being deleted, a fork for a snapshot on a big dataset, or the machine swapping to disk. On security, don't expose Redis to the public internet, set authentication and ACLs, and use TLS if traffic crosses networks you don't control.&lt;/p&gt;

&lt;p&gt;One more piece of context, in case it comes up. In 2024 Redis moved away from its permissive BSD license, and the Linux Foundation backed a fork called Valkey, based on Redis 7.2.4 and still BSD-licensed. Redis 8 later added AGPLv3 as a license option. The two projects speak the same protocol, so most clients work with either, but their features are starting to diverge. Managed services like AWS ElastiCache offer Valkey too.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I'd practise this
&lt;/h2&gt;

&lt;p&gt;Reading only gets you partway. Run Redis in Docker (&lt;code&gt;docker run -p 6379:6379 redis&lt;/code&gt;) and type the commands into &lt;code&gt;redis-cli&lt;/code&gt; until they feel obvious. Then build the three things interviews keep asking for: a lock with a safe release script, a sliding window rate limiter, and a leaderboard. Once those work, try to break them: kill the process mid-script, let a TTL expire at a bad moment, and see what happens.&lt;/p&gt;

&lt;p&gt;Then try answering these out loud, without notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why is Redis fast, and what's the cost of that design?&lt;/li&gt;
&lt;li&gt;What's the difference between expiration and eviction, and what does the default policy do when memory is full?&lt;/li&gt;
&lt;li&gt;Walk me through cache-aside. What are the stale-data races?&lt;/li&gt;
&lt;li&gt;What's the difference between cache penetration, stampede and avalanche, and how do you fix each?&lt;/li&gt;
&lt;li&gt;How do RDB and AOF differ, and what would you use for a cache versus for sessions?&lt;/li&gt;
&lt;li&gt;A primary fails right after acknowledging a write. What can happen?&lt;/li&gt;
&lt;li&gt;When would you use Sentinel, and when Cluster? How does a key find its node?&lt;/li&gt;
&lt;li&gt;What's the difference between a pipeline, MULTI/EXEC and a Lua script?&lt;/li&gt;
&lt;li&gt;Why doesn't Pub/Sub work for messages that must not be lost, and what would you use instead?&lt;/li&gt;
&lt;li&gt;Design a rate limiter. Now design a distributed lock, and tell me why it isn't fully safe.&lt;/li&gt;
&lt;li&gt;Give me three reasons not to use Redis as your main database.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you can answer those in your own words, with an example each, you know Redis well enough for most SDE interviews.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://redis.io/docs/latest/operate/oss_and_stack/management/persistence/" rel="noopener noreferrer"&gt;Redis persistence docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://redis.io/docs/latest/develop/reference/eviction/" rel="noopener noreferrer"&gt;Redis key eviction docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://redis.io/docs/latest/develop/clients/patterns/distributed-locks/" rel="noopener noreferrer"&gt;Redis docs on distributed locks and Redlock&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://martin.kleppmann.com/2016/02/08/how-to-do-distributed-locking.html" rel="noopener noreferrer"&gt;Martin Kleppmann, "How to do distributed locking"&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.hellointerview.com/learn/courses/system-design/lesson/scaling-reads/redis" rel="noopener noreferrer"&gt;Hello Interview's Redis deep dive for system design&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>redis</category>
      <category>software</category>
      <category>systemdesign</category>
      <category>programming</category>
    </item>
    <item>
      <title>My AI Support Bot Was Slow. Then SigNoz Told Me My Alert Would Fire in 158 Years.</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Tue, 14 Jul 2026 16:18:20 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/my-ai-support-bot-was-slow-then-signoz-told-me-my-alert-would-fire-in-158-years-1f5f</link>
      <guid>https://dev.to/shogun_the_grt/my-ai-support-bot-was-slow-then-signoz-told-me-my-alert-would-fire-in-158-years-1f5f</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Author's Note&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This project was built as part of the &lt;strong&gt;Agents of SigNoz&lt;/strong&gt; hackathon organized by &lt;strong&gt;WeMakeDevs&lt;/strong&gt; and &lt;strong&gt;SigNoz&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Source code:&lt;/strong&gt; &lt;a href="https://github.com/SarthakRawat-1/signoz-assistant" rel="noopener noreferrer"&gt;https://github.com/SarthakRawat-1/signoz-assistant&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;How I traced a real latency problem, a real auth failure, and one genuinely baffling unit-conversion bug in an AI support assistant, using OpenTelemetry and self-hosted SigNoz.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I built a small AI support assistant for an imaginary ecommerce company, the kind that answers "where's my order" and "will the rain delay my delivery." It worked, then I sent it 25 requests back to back and some took over three seconds. No errors, no crashes, just slow, some of the time. I had no idea which part, so I did what I should have done from the start: instrumented it and watched.&lt;/p&gt;

&lt;p&gt;Two things came out of that afternoon. One was a real, boring, entirely explainable latency bug. The other was SigNoz telling me an alert would trigger in roughly 158 years, which is not a sentence I expected to type while debugging a chatbot.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnzitgizatyz6eag41zwx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnzitgizatyz6eag41zwx.png" alt="Architecture diagram: Client/Seed Script → FastAPI App → SQLite DB, Gemini API, Open-Meteo API, and SigNoz OTel Collector" width="800" height="479"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The app is small on purpose: a FastAPI endpoint (&lt;code&gt;POST /support/chat&lt;/code&gt;) that takes a message and a &lt;code&gt;conversation_id&lt;/code&gt;, stores history in SQLite, optionally calls a weather API if the message looks shipping-related, and generates a reply. A small frontend on top shows the trace ID and classified intent for whatever you just sent, mostly so I didn't have to keep tabbing over to SigNoz to find which trace belonged to which message.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv9vqd63h2t93mf1xn3ym.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv9vqd63h2t93mf1xn3ym.png" alt="The chat interface, with a live panel showing the trace ID and classified intent for the last message" width="800" height="428"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One honest disclosure: I didn't have a Gemini API key wired in for most of this session, so the LLM calls below are running against a stub client, not live Gemini. My code logs this plainly: &lt;code&gt;No GEMINI_API_KEY set. Falling back to Stub LLM Client&lt;/code&gt;. It doesn't affect the story, since the point was never to benchmark Gemini, it was to see whether SigNoz would correctly point at the real bottleneck regardless of which piece was mocked. It did.&lt;/p&gt;

&lt;p&gt;The weather lookup is where the interesting instrumentation lives, a manual span wrapping both outbound calls with business-relevant attributes attached along the way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_weather&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;city&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;tracer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start_as_current_span&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;weather.lookup_operation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;weather.city&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;city&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;httpx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AsyncClient&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;geo_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;GEOCODING_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;city&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
            &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;weather.geocoding.http.status_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;geo_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

            &lt;span class="n"&gt;lat&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lon&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;_extract_coords&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;geo_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
            &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attributes&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;weather.latitude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;lat&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;weather.longitude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;lon&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;

            &lt;span class="n"&gt;forecast_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;FORECAST_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latitude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;lat&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;longitude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;lon&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
            &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;weather.forecast.http.status_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;forecast_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="c1"&gt;# ...response is parsed into a simplified condition below, trimmed here
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two &lt;code&gt;client.get()&lt;/code&gt; calls, one geocoding, one forecast, are the two GET spans stacked under &lt;code&gt;weather.lookup_operation&lt;/code&gt; in the waterfall below. Attaching &lt;code&gt;customer.intent&lt;/code&gt; and &lt;code&gt;conversation.id&lt;/code&gt; to spans like this let me filter traces by business behavior, "show me every shipping_delay conversation," instead of only by HTTP endpoint, which mattered more once several conversations were running at once.&lt;/p&gt;

&lt;p&gt;No agent framework, no tool-calling loop, no vector database. Just enough moving parts to produce interesting telemetry.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting visibility
&lt;/h2&gt;

&lt;p&gt;I self-hosted SigNoz using Foundry, their newer install method that replaced the old &lt;code&gt;docker-compose&lt;/code&gt;-based install a few months back: install the CLI, point it at a small YAML file, and it stands up ClickHouse, the OTel collector, and the UI itself. I added &lt;code&gt;HTTPXClientInstrumentor&lt;/code&gt; on top of FastAPI's auto-instrumentation, plus manual spans around the LLM call and the weather lookup.&lt;/p&gt;

&lt;p&gt;None of it went smoothly first try. Docker Desktop was running, but its CLI binaries weren't on my shell's PATH, so &lt;code&gt;foundryctl cast&lt;/code&gt; failed outright with &lt;code&gt;tools are not available, please install them and try again: docker, docker-compose&lt;/code&gt;, a two-minute fix once I found it. Then logs didn't show up in SigNoz at all, because auto-instrumented logging assumes your app and collector share stdout, which isn't true when your app runs on the host and SigNoz runs in Docker. Configuring an OTLP log exporter to ship logs over gRPC fixed it, and correlation with traces worked immediately after.&lt;/p&gt;

&lt;h2&gt;
  
  
  Investigation #1: the actual slow part
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmu5s5ruiymbf6wamtsf5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmu5s5ruiymbf6wamtsf5.png" alt="Trace waterfall showing weather.lookup_operation taking 3.06s against a 405ms LLM call" width="800" height="413"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once data was flowing, I sorted traces by duration and opened the slowest one. Without the waterfall view, I'd probably have blamed Gemini first, since it's the component that looks expensive on paper. The trace showed I was about to optimize the wrong thing entirely: &lt;code&gt;weather.lookup_operation&lt;/code&gt; took 3.06 seconds. The instrumented LLM span, right next to it, took 405 milliseconds. This wasn't a fluke. I checked several of the slow traces, and the same pattern held every time.&lt;/p&gt;

&lt;p&gt;I built a small dashboard to track this at the aggregate level too: p95 latency on the endpoint, error rate, and average LLM token count.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8dp1q984t7v7iv2h3lev.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8dp1q984t7v7iv2h3lev.png" alt="Custom dashboard with three panels: P95 Latency for POST /support/chat, Request Error Rate, and Avg LLM Token Count" width="800" height="430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across 25 seeded requests, the endpoint's overall p50 was a perfectly reasonable 484ms, but p99 jumped to 3.37 seconds. &lt;code&gt;weather.lookup_operation&lt;/code&gt; alone had a p50 of 2.4 seconds. Four requests triggered it; all four were the slow ones.&lt;/p&gt;

&lt;p&gt;Neither call is artificially delayed; there's no &lt;code&gt;sleep()&lt;/code&gt; anywhere in that path. Two things make it slow: the calls are sequential, since you need coordinates before requesting a forecast, and each request creates a new &lt;code&gt;AsyncClient&lt;/code&gt; instead of reusing one, so every lookup pays the cost of a fresh connection twice. Mundane, fixable, and something I'd never have found by reading logs alone; the waterfall made both GET calls and their relative cost obvious at a glance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Investigation #2: an intentional failure
&lt;/h2&gt;

&lt;p&gt;The second investigation, unlike the first, was deliberate. I temporarily set an invalid Gemini API key to see what a real auth failure would look like end to end.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3na42thiqgl8x7jw08kg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3na42thiqgl8x7jw08kg.png" alt="Log entry showing " width="800" height="421"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The result showed up as an ERROR-level log, &lt;code&gt;Request failed with exception&lt;/code&gt;, with a &lt;code&gt;trace_id&lt;/code&gt; attached directly to the log line. Clicking that ID took me straight to the exact trace where the failure happened, no manual correlation required. This is the kind of thing that's genuinely hard to appreciate until you've done the alternative: grepping through unstructured logs trying to match timestamps to a request you think might be related.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug I didn't expect: nanoseconds pretending to be years
&lt;/h2&gt;

&lt;p&gt;This is the one I'd actually lead with if I were telling this story out loud. I set up an alert: notify me if p95 latency on &lt;code&gt;/support/chat&lt;/code&gt; goes above 4 seconds. Simple enough. Except when I opened the alert's chart to sanity-check it, the y-axis read &lt;strong&gt;158.44 years&lt;/strong&gt;, with the critical threshold drawn at &lt;strong&gt;126.76 years&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz7dph835gfqnuict1p31.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz7dph835gfqnuict1p31.png" alt="Alert chart with a y-axis labeled in years: 158.44 years at the top, a critical threshold line at 126.76 years" width="800" height="411"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SigNoz's alert query builder operates on &lt;code&gt;duration_nano&lt;/code&gt;, nanoseconds, not milliseconds. I'd set my threshold to &lt;code&gt;4000&lt;/code&gt;, assuming milliseconds. It meant 4000 nanoseconds, roughly the time it takes an electron to get mildly annoyed, so every request blew past it instantly, and the chart rescaled itself into whatever unit made the numbers legible: years. The fix was &lt;code&gt;4000000000&lt;/code&gt; (4 billion nanoseconds, i.e., 4 seconds). The lesson wasn't "SigNoz is confusing," it was that observability tools expose raw units because precision matters at the storage layer, and it's on you to know exactly what you're alerting on.&lt;/p&gt;

&lt;p&gt;With the units corrected, the same alert did exactly what I'd originally wanted. It stayed quiet during normal traffic and fired only once a request genuinely crossed 4 seconds.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6zjptxp9zd78mth26xpe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6zjptxp9zd78mth26xpe.png" alt="Alert Rules page showing " width="800" height="238"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I couldn't find this mentioned anywhere in the getting-started docs, and it took me longer than I'd like to admit to realize the alert itself was working correctly and my units weren't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually changed for me
&lt;/h2&gt;

&lt;p&gt;Before this, I thought of observability mostly as dashboards: numbers on a screen telling you something's wrong. What I didn't expect is how much of the actual debugging happens in traces, not metrics. A metric can tell you p99 latency is bad; it can't tell you why. The trace waterfall explained the "why," span by span, in a way I could act on immediately instead of forming a hypothesis and checking it against logs one line at a time.&lt;/p&gt;

&lt;p&gt;The other shift is smaller but sticks with me: I used to think of alerts as something you configure once you already understand your system. This one taught me the alert itself needs debugging too, not just the app it's watching.&lt;/p&gt;

&lt;p&gt;One thing I appreciated only after the fact: none of this instrumentation is actually SigNoz-specific. The spans and attributes are emitted through OpenTelemetry, a vendor-neutral standard; SigNoz just happens to be the backend consuming them here. That's why it felt like a real part of the application rather than something bolted on for one dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Before this, "my app is slow" would have meant staring at logs and guessing. After wiring up OpenTelemetry and SigNoz, every request told its own story: which span cost what, which log line belonged to which trace, and, once, which threshold was off by a factor of a million. Next, I'd share one long-lived &lt;code&gt;httpx.AsyncClient&lt;/code&gt; instead of opening a new one per request, run the geocoding and forecast calls concurrently, and wire in a real Gemini key just to see what surprises a live model call brings. I didn't get a dramatic engineering war story out of this. I got a small, boring, entirely real latency bug, an honest auth failure, and one legitimately funny unit-conversion mistake. That's a fair trade for an afternoon.&lt;/p&gt;

</description>
      <category>signoz</category>
      <category>observability</category>
      <category>opentelemetry</category>
      <category>python</category>
    </item>
    <item>
      <title>Disaggregated Prefill/Decode: The Architecture Quietly Rewiring How LLMs Get Served</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Wed, 08 Jul 2026 16:59:09 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/disaggregated-prefilldecode-the-architecture-quietly-rewiring-how-llms-get-served-4fcl</link>
      <guid>https://dev.to/shogun_the_grt/disaggregated-prefilldecode-the-architecture-quietly-rewiring-how-llms-get-served-4fcl</guid>
      <description>&lt;p&gt;If you have used any large language model product in the last year, you have run into a piece of infrastructure most people never think about: the serving system that turns a prompt into a stream of tokens. For a long time, that system worked the same way regardless of scale. One GPU, or one small cluster of GPUs, would take your prompt, process it, and then generate the response token by token, all on the same hardware, in the same process.&lt;/p&gt;

&lt;p&gt;That approach is quietly breaking down. Context windows have stretched past 100,000 tokens, request volumes have grown by orders of magnitude, and the cost of running inference has become a board level conversation at most AI companies. The fix that has emerged, and is now running in production at some of the largest LLM providers in the world, is called disaggregated prefill/decode serving. It sounds like a narrow infrastructure detail, but it is one of the more interesting distributed systems problems in production today, and it has not been written about nearly as much as caching or load balancing.&lt;/p&gt;

&lt;p&gt;This post explains what the problem actually is, how the architecture works, how real systems like Mooncake and NVIDIA Dynamo implement it, and when it is genuinely worth adopting versus when it is just extra operational burden.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbyp42zkqtprth6b7tbl7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbyp42zkqtprth6b7tbl7.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why prefill and decode do not want to live together
&lt;/h2&gt;

&lt;p&gt;Every LLM inference request has two distinct phases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prefill&lt;/strong&gt; is when the model reads your entire prompt and builds up its internal representation of it, layer by layer, for every token you sent in. This step is compute bound. It uses the GPU's raw math throughput heavily, and it happens once per request, regardless of how long the eventual response is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decode&lt;/strong&gt; is what happens after that: the model generates the response one token at a time, each new token depending on everything that came before it. This step is memory bandwidth bound. The GPU spends most of its time moving the key value cache (the KV cache, which stores the model's running memory of the conversation) in and out of memory rather than doing heavy computation.&lt;/p&gt;

&lt;p&gt;These two phases have almost opposite hardware profiles. Prefill wants raw FLOPs. Decode wants memory bandwidth and low latency access to the KV cache. When both run on the same GPU in the same batch, they compete for the same resources, and one interferes with the other. A long prompt arriving for prefill can stall the token generation of requests that are already mid decode, which shows up to the end user as a sudden latency spike.&lt;/p&gt;

&lt;p&gt;Academic work on this problem, including the DistServe and Splitwise papers, and production systems like Mooncake, converge on the same conclusion: separate the two phases onto different hardware, and let each one be scheduled and scaled independently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core architecture
&lt;/h2&gt;

&lt;p&gt;At a high level, disaggregated serving introduces a new component that does not exist in a monolithic setup: a layer that moves the KV cache from the machines doing prefill to the machines doing decode.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqyr1l08229btukfkr9uf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqyr1l08229btukfkr9uf.png" alt=" " width="800" height="96"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The router (sometimes called a scheduler, sometimes a conductor, depending on the system) decides which prefill instance and which decode instance should handle a given request. It is trying to balance several things at once: current load on each pool, whether part of the KV cache for this request already exists somewhere in the cluster from a previous turn, and the latency target for this particular request.&lt;/p&gt;

&lt;p&gt;Once a prefill instance finishes processing the prompt, it does not throw away its work. It streams the resulting KV cache, often layer by layer as it is produced rather than waiting for the whole thing to finish, over to the assigned decode instance. The decode instance loads that cache and begins generating tokens, which get streamed back to the client as they are produced.&lt;/p&gt;

&lt;p&gt;Here is what a single request actually experiences as it moves through the system.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7flffvk9uuzsychlo9mc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7flffvk9uuzsychlo9mc.png" alt=" " width="799" height="555"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The hard part: moving the KV cache fast enough
&lt;/h2&gt;

&lt;p&gt;The entire architecture only works if the KV cache transfer is fast and does not become its own bottleneck. If moving the cache takes longer than the time saved by separating the phases, disaggregation is pointless.&lt;/p&gt;

&lt;p&gt;This is why the transfer layer has become its own area of serious engineering effort. In 2026, most production systems rely on RDMA (remote direct memory access) over InfiniBand or RoCE networking to move KV cache data directly between GPU memory pools with very low latency, falling back to plain TCP when that hardware is not available. NVIDIA's NIXL and Moonshot AI's Transfer Engine (the foundation of Mooncake) are the two most widely used implementations of this idea, and both vLLM and SGLang have built native support for them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpe9tms40r5zz7d5j4qnl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpe9tms40r5zz7d5j4qnl.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The amount of data being moved is not trivial either. For a rough sense of scale, the size of the KV cache for a single request scales with the number of layers in the model, the number of attention heads, the size of each head, and the length of the sequence. For a long conversation or a large document being summarized, this can be gigabytes of data that needs to move between machines before the first output token is even generated, which is why the transfer engine, not just the model itself, has become a serious point of optimization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Case study: Mooncake, the platform behind Kimi
&lt;/h2&gt;

&lt;p&gt;The clearest public example of this architecture running at real scale is Mooncake, the serving platform Moonshot AI built for its Kimi chatbot. Mooncake was one of the first systems to fully commit to a KV cache centric design rather than treating the cache as a side effect of inference.&lt;/p&gt;

&lt;p&gt;At the center of Mooncake sits a global scheduler called the Conductor. For every incoming request, the Conductor decides which prefill and decode instance pair should handle it, taking into account how much of the relevant KV cache already exists somewhere in the cluster and can be reused rather than recomputed from scratch. Mooncake also does something unusual: it uses otherwise idle CPU memory, DRAM, and even SSD storage across the GPU cluster to build a much larger, distributed KV cache pool than would fit in GPU memory alone.&lt;/p&gt;

&lt;p&gt;The results reported by the team are substantial. In their published research, Mooncake increased effective request capacity by roughly 60 to nearly 500 percent over baseline approaches in various tested scenarios, while still meeting latency service level objectives, with the biggest gains showing up in long context workloads where the cost of recomputing versus reusing cache matters most. As of their most recent public updates, the system is running across thousands of nodes and processing over 100 billion tokens per day in production, not a lab benchmark.&lt;/p&gt;

&lt;p&gt;One detail that is easy to miss: Mooncake does not assume every request should be processed no matter what. Under heavy load, it uses a prediction based early rejection policy, essentially deciding as early as possible that a request cannot realistically meet its latency target and rejecting it immediately, rather than letting it consume GPU resources partway through and then failing anyway. That is a meaningfully different design philosophy from most traditional web backends, where the instinct is usually to queue and retry rather than reject early.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6044x3si5wv3x92vkoe7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6044x3si5wv3x92vkoe7.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The rest of the ecosystem is converging on the same idea
&lt;/h2&gt;

&lt;p&gt;Mooncake is not an isolated experiment. By 2026, prefill/decode disaggregation has effectively become the industry default rather than a niche optimization.&lt;/p&gt;

&lt;p&gt;vLLM, one of the most widely used open source inference engines, supports disaggregated serving natively, using NIXL as its standard transfer mechanism alongside NVIDIA Dynamo. SGLang offers a similar setup through its own router and disaggregation mode, and its RadixAttention design gives it a particular advantage for workloads with heavy prompt reuse. LMDeploy has added support through a comparable mechanism. NVIDIA's own Dynamo platform builds a routing and orchestration layer on top of this pattern for large scale deployments.&lt;/p&gt;

&lt;p&gt;Perhaps the clearest sign that this has moved from research to infrastructure standard is llm-d, a project from Red Hat, IBM, and Google that brings Kubernetes native disaggregated inference on top of vLLM. It was donated to the CNCF in early 2026, and its default configuration reportedly delivers a meaningful performance improvement with no manual tuning required, with later releases showing significantly lower per token latency for large models on modern GPU hardware.&lt;/p&gt;

&lt;p&gt;Production adoption backs this up. Reports throughout 2026 point to Meta, LinkedIn, Mistral, and Hugging Face running vLLM with disaggregated serving in production, and DeepSeek and Google's Gemini team have both discussed similar architectures at scale. Even at the hardware level, the idea is expanding. AWS partnered with Cerebras this year to bring disaggregated inference to Amazon Bedrock, using one type of chip for prompt processing and an entirely different type of chip for token generation, running side by side in the same cloud region.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foqn8sccb9rpcyb8gki5m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foqn8sccb9rpcyb8gki5m.png" alt=" " width="800" height="246"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Pushing disaggregation further: splitting more than just prefill and decode
&lt;/h2&gt;

&lt;p&gt;Once you accept the idea that different parts of inference have different resource profiles, it is natural to ask whether the split should go deeper than just prefill versus decode. Some of the newest research does exactly that.&lt;/p&gt;

&lt;p&gt;ByteDance's MegaScale Infer system takes disaggregation a step further by splitting the attention computation from the feedforward network (FFN) computation within the decode phase itself, since these two sub components also have different bandwidth and compute characteristics. Their reported deployment runs across close to 10,000 GPUs, using one GPU type optimized for the bandwidth heavy attention work and a different GPU type for the compute heavy expert layers.&lt;/p&gt;

&lt;p&gt;There is also work extending this pattern to multimodal models, where an additional pool of instances handles encoding images, audio, or video before the text pipeline even begins, effectively creating a three stage pipeline of encode, prefill, and decode rather than just two.&lt;/p&gt;

&lt;p&gt;The direction is clear: the industry is moving toward treating inference as a pipeline of specialized stages that can each be placed on the hardware best suited to them, rather than a single monolithic job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scheduling: the part nobody talks about enough
&lt;/h2&gt;

&lt;p&gt;A disaggregated architecture is only as good as the scheduler deciding how to use it. This is where most of the ongoing research effort is actually concentrated, more so than the transfer mechanics themselves.&lt;/p&gt;

&lt;p&gt;The key metric these systems optimize for is not raw throughput but goodput, meaning the number of requests that are fully completed within their latency service level objective. A request that gets 80 percent of the way through and then misses its deadline has still consumed GPU time and produced nothing useful, so it counts against the system rather than for it. Recent scheduling research also focuses on request imbalance, meaning what happens when a burst of long prompts arrives and threatens to starve the decode pool of attention, since prefill and decode instances scale independently and a mismatch between the two creates its own kind of bottleneck.&lt;/p&gt;

&lt;p&gt;This is a genuinely different scheduling problem from most systems most engineers have built before. It is closer to job scheduling in a data center than request routing in a typical web service, because the "jobs" have two dependent phases running on different machines with a real time data transfer between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  When disaggregation is not worth it
&lt;/h2&gt;

&lt;p&gt;None of this is free. Splitting prefill and decode introduces real operational cost: two separate clusters to monitor, a KV cache transfer layer that can fail independently of either compute pool, cache miss handling, and a more complex deployment story overall. Teams that adopt this pattern before they actually need it often spend a meaningful amount of extra engineering time in the first several months for little practical benefit.&lt;/p&gt;

&lt;p&gt;A reasonable way to think about whether disaggregation is worth adopting:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp6r69g4kfolim669lv9x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp6r69g4kfolim669lv9x.png" alt=" " width="775" height="1430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The signal worth paying attention to is a correlation between long prompt arrivals and latency spikes in ongoing token generation. If that pattern is not showing up in your metrics, the complexity of disaggregation is probably not paying for itself yet, and the better investment is tuning batch sizes and KV cache management on a single unified cluster.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing thoughts
&lt;/h2&gt;

&lt;p&gt;Disaggregated prefill/decode serving is a good example of a broader trend in system design right now: the systems that matter most in 2026 are the ones built specifically for how large language models actually behave, rather than general purpose infrastructure patterns borrowed from earlier eras of web architecture. Prefill and decode are fundamentally different computational problems, and treating them that way, with dedicated hardware pools and a fast cache transfer layer between them, has produced real, measurable gains in production at some of the largest AI companies in the world.&lt;/p&gt;

&lt;p&gt;It is also a useful case study in judgment. The right answer is not "always disaggregate." It is understanding your own workload well enough to know whether the added complexity buys you something real, which is the same question that sits underneath almost every good system design decision.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>systemdesign</category>
      <category>software</category>
    </item>
    <item>
      <title>Why AI Assistants Have an Amnesia Problem. And How We're Fixing It</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Sat, 06 Jun 2026 08:43:32 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/why-ai-assistants-have-an-amnesia-problem-and-how-were-fixing-it-3m6</link>
      <guid>https://dev.to/shogun_the_grt/why-ai-assistants-have-an-amnesia-problem-and-how-were-fixing-it-3m6</guid>
      <description>&lt;p&gt;&lt;em&gt;A deep dive into HGVM: the memory architecture that makes AI agents remember like humans, forget like humans, and learn like humans.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;There is a moment every frequent AI user has experienced. You spent twenty minutes in a conversation explaining your project, your constraints, your preferences, your stack. The assistant was helpful, precise, contextually aware. Then you closed the tab.&lt;/p&gt;

&lt;p&gt;The next day, you open a new conversation. The assistant greets you like a stranger.&lt;/p&gt;

&lt;p&gt;You explain everything again.&lt;/p&gt;

&lt;p&gt;This is not a minor inconvenience. It is a fundamental architectural failure, and it is not the only one. On the other side of the spectrum are systems that remember &lt;em&gt;everything&lt;/em&gt; forever: every offhand comment, every temporary preference, every piece of stale context from six months ago. These systems don't forget, but they accumulate noise until their memory becomes a liability. They surface outdated information as confidently as current facts. They become sycophantic, anchored to what you said before rather than what is true now.&lt;/p&gt;

&lt;p&gt;Neither model is how memory should work. Neither model is how &lt;em&gt;human&lt;/em&gt; memory works.&lt;/p&gt;

&lt;p&gt;We built HGVM, Hierarchical Graph-Vector Memory, to fix this. This blog explains what it is, how it works, and why we think it represents a meaningfully different approach to AI agent memory.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Human Memory Analogy
&lt;/h2&gt;

&lt;p&gt;Human memory is not a database. It does not store everything with equal fidelity and retrieve it on demand. It is a dynamic, adaptive system that continuously organizes, compresses, reinforces, and discards information based on relevance, recency, and repeated exposure.&lt;/p&gt;

&lt;p&gt;When you learn something new, your brain does not simply file it. It connects it to existing structures. It consolidates related fragments into coherent abstractions during sleep. It strengthens memories that are repeatedly accessed and allows others to fade. And crucially, it reflects, it notices patterns across experiences and crystallizes them into durable knowledge that shapes future reasoning.&lt;/p&gt;

&lt;p&gt;HGVM is built on this intuition. Memory should be a &lt;em&gt;managed resource&lt;/em&gt;, not a growing pile. It should be structured hierarchically, retrieved selectively, compressed into abstractions over time, and actively pruned when it is no longer relevant. Four behaviors, working together.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Four Pillars of HGVM
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Hierarchical Structure
&lt;/h3&gt;

&lt;p&gt;Most existing memory systems store memories as a flat collection of text chunks with embeddings. When you query, you run a vector similarity search across all of them and return the top results. This works adequately at small scale, but it fails in two important ways as the memory grows.&lt;/p&gt;

&lt;p&gt;First, it causes &lt;strong&gt;context leakage&lt;/strong&gt;, semantically similar but contextually irrelevant memories surface because the retrieval system has no way to distinguish between "similar in meaning" and "relevant to this situation." A question about your Python project might surface memories about a Python script you wrote for a completely different purpose two months ago.&lt;/p&gt;

&lt;p&gt;Second, it is expensive. Vector search across all memories grows linearly. The more you remember, the slower and noisier retrieval becomes.&lt;/p&gt;

&lt;p&gt;HGVM organizes memory into a strict four-layer hierarchy:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domain → Category → Topic → Episode&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A Domain is the broadest grouping: Work, Health, Personal, Finance. A Category is a sub-domain: ProjectX, Morning Runs, Family. A Topic is a coherent theme within a category: Backend API Design, Sprint Planning, Dad's Hospital Visit. An Episode is a single atomic memory item attached to a topic.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmbw5t2jnfh23f03dqqla.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmbw5t2jnfh23f03dqqla.png" alt="Core Architecture Diagram" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When a query arrives, the system routes at the Domain and Category level first, a cheap semantic match against summaries, and then runs vector search only within the relevant sub-graph. The candidate set is small. Retrieval is precise. Context leakage is dramatically reduced because irrelevant branches are never even considered.&lt;/p&gt;

&lt;p&gt;This is the structural foundation everything else builds on.&lt;/p&gt;




&lt;h3&gt;
  
  
  2. Active Forgetting
&lt;/h3&gt;

&lt;p&gt;Forgetting is not a failure mode. It is a feature.&lt;/p&gt;

&lt;p&gt;The reason most AI memory systems remember everything forever is that forgetting feels dangerous, what if you delete something important? The result is a system that accumulates noise indefinitely, where the signal-to-noise ratio of the memory store degrades over time and retrieval quality degrades with it.&lt;/p&gt;

&lt;p&gt;HGVM takes a different position: forgetting should be &lt;strong&gt;tiered by memory type&lt;/strong&gt;, not uniform. Not all memories deserve the same treatment.&lt;/p&gt;

&lt;p&gt;We define five memory classes, each with its own decay behavior:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Memory Class&lt;/th&gt;
&lt;th&gt;What It Stores&lt;/th&gt;
&lt;th&gt;Decay Rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;permanent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Explicit identity facts ("I am vegetarian", "never forget this")&lt;/td&gt;
&lt;td&gt;Never decays&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;semantic&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Consolidated summaries and learned patterns&lt;/td&gt;
&lt;td&gt;Very slow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;preference&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Stable user preferences&lt;/td&gt;
&lt;td&gt;Very slow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;observation&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Raw observations and task outcomes&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;task_state&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Active work state and in-progress details&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxb34x0oaj93h0ych0m0r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxb34x0oaj93h0ych0m0r.png" alt="Memory Class Spectrum" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A permanent memory :- something the user explicitly told the system to remember, never decays. A preference :- "I prefer TypeScript", fades very slowly, remaining useful for months. A task_state memory :- "currently blocked on the auth middleware", fades quickly, becoming stale within days because task context changes fast.&lt;/p&gt;

&lt;p&gt;The decay formula itself is borrowed from cognitive science. We use an exponential decay function where the strength of each memory degrades over time since last access:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;new_strength = current_strength × exp(−λ × days_since_access)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each memory class has its own λ value. When strength falls below a minimum threshold, the memory is &lt;strong&gt;soft-deleted&lt;/strong&gt; :- marked as invalid but preserved in the graph for auditability. Nothing is ever hard-deleted during forgetting.&lt;/p&gt;

&lt;p&gt;Retrieval is reinforced too: every time a memory is returned in a query, its strength is boosted. Memories that are repeatedly useful stay strong. Memories that are never relevant fade gracefully.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7sxk3i23uj0asf1xqwnd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7sxk3i23uj0asf1xqwnd.png" alt="Forgetting Curve Comparison" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the forgetting curve most existing systems ignore. HGVM embraces it as a design principle.&lt;/p&gt;




&lt;h3&gt;
  
  
  3. Consolidation
&lt;/h3&gt;

&lt;p&gt;A topic with fifty raw episode nodes is not useful. Fifty individually embedded fragments retrieved piecemeal create noisy, redundant context. At some point, accumulated raw memories need to be compressed into something denser and more useful.&lt;/p&gt;

&lt;p&gt;This is consolidation.&lt;/p&gt;

&lt;p&gt;HGVM runs a scheduled consolidation pipeline that works as follows: when a topic accumulates enough raw episodes, the system clusters them by semantic similarity, extracts atomic facts from each cluster, generates a concise summary from those facts, and then runs a &lt;strong&gt;batched verification pass&lt;/strong&gt; to confirm the summary is faithful to the source material.&lt;/p&gt;

&lt;p&gt;The verification step is non-negotiable. We check for three failure conditions: missing facts (something in the source that did not make it into the summary), altered facts (something that changed in meaning), and contradicted facts (something that directly conflicts with the source). If any of these conditions are non-empty, the consolidation is rejected and logged. No summary episode is created until verification passes completely.&lt;/p&gt;

&lt;p&gt;When verification passes, the system creates a new &lt;code&gt;semantic&lt;/code&gt; episode with subtype &lt;code&gt;consolidation_summary&lt;/code&gt;, linked to the source episodes via &lt;code&gt;SUMMARIZES&lt;/code&gt; relationships. Critically, &lt;strong&gt;the source episodes are never deleted&lt;/strong&gt;. Consolidation adds compressed knowledge — it does not destroy provenance.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhy7xf43tg2ckxwwwjuz2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhy7xf43tg2ckxwwwjuz2.png" alt="Retrieval Pipeline Flow" width="800" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The retrieval system then naturally prefers the summary for most queries, it is denser, more representative, and semantically richer, while the raw episodes remain available for queries that need granular detail.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F90cgsat0d5ttoitbmua6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F90cgsat0d5ttoitbmua6.png" alt="Consolidation Before/After" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This mirrors how human memory consolidates during sleep: the raw experiences remain somewhere in the system, but what you access day-to-day is the compressed, organized version.&lt;/p&gt;




&lt;h3&gt;
  
  
  4. Reflection
&lt;/h3&gt;

&lt;p&gt;Consolidation compresses. Reflection &lt;em&gt;learns&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Reflection is the most novel component of HGVM. While consolidation creates summaries within a single topic, reflection operates across topics and sessions, looking for higher-order patterns in behavior, preference, and reasoning.&lt;/p&gt;

&lt;p&gt;Imagine a user who across ten sessions has: rejected Java twice, chosen Python for three consecutive projects, consistently used FastAPI over Django, and complained about verbose enterprise frameworks. No single episode says "this user values developer productivity." But the pattern across episodes does.&lt;/p&gt;

&lt;p&gt;Reflection detects this. It fetches recent valid episodes from the non-reflective memory classes :- &lt;code&gt;permanent&lt;/code&gt;, &lt;code&gt;preference&lt;/code&gt;, &lt;code&gt;task_state&lt;/code&gt;, and &lt;code&gt;observation&lt;/code&gt;, groups them into coherent evidence bundles, and generates candidate higher-order abstractions. These are persisted as &lt;code&gt;semantic&lt;/code&gt; episodes with subtype &lt;code&gt;reflection_pattern&lt;/code&gt;, linked to their supporting evidence.&lt;/p&gt;

&lt;p&gt;The result: the next time the user asks for a project scaffold recommendation, the system does not just recall that they used FastAPI before. It retrieves the reflection, "this user prioritizes developer productivity over ecosystem conservatism", and reasons from that durable insight.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Famwbbs3n9jn4wdl77xtw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Famwbbs3n9jn4wdl77xtw.png" alt="Reflection Explained" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There is one hard constraint we enforce at the architecture level: &lt;strong&gt;reflection memories can never be used as evidence for future reflections&lt;/strong&gt;. This prevents a feedback loop where the system reflects on its own abstractions, compounding distortions over time. Reflection operates only on primary evidence. This constraint is enforced at the database query level, not just the prompt level.&lt;/p&gt;




&lt;h2&gt;
  
  
  How HGVM Compares to What Exists
&lt;/h2&gt;

&lt;p&gt;The agent memory landscape has become more sophisticated in the last two years. It is worth being honest about where existing systems land and where HGVM goes further.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MemGPT / Letta&lt;/strong&gt; introduced the idea of paging between active context and external memory using function calls, and Letta added sleep-time consolidation agents. HGVM builds directly on these ideas but adds hierarchical routing, tiered forgetting, and the reflection pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mem0&lt;/strong&gt; offers a clean multi-signal retrieval system (semantic + keyword + entity) with good API design. It does not offer hierarchical memory organization, tiered decay, or reflection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zep / Graphiti&lt;/strong&gt; provide excellent graph-based memory with temporal reasoning and the bitemporal model (facts carry both a valid time and a transaction time). HGVM adopts the bitemporal approach for contradiction handling and uses Graphiti's &lt;code&gt;valid_at&lt;/code&gt; / &lt;code&gt;invalid_at&lt;/code&gt; pattern for soft deletion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT Memory&lt;/strong&gt; is a static injection system, remembered facts are prepended to context. This causes cross-domain leakage (facts from unrelated contexts surface inappropriately) and sycophancy (the system defers to stored preferences even when they are stale or wrong). PersistBench has documented both failure modes.&lt;/p&gt;

&lt;p&gt;HGVM is not a critique of any of these systems. Each solved a real problem. HGVM synthesizes the best ideas :- graph structure, temporal modeling, consolidation, multi-signal retrieval and adds forgetting as a first-class feature and reflection as a novel capability.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Multi-Agent Layer
&lt;/h2&gt;

&lt;p&gt;HGVM is not just a memory system. It is a memory system designed for an agent society.&lt;/p&gt;

&lt;p&gt;Four specialized LLM agents share the same memory graph and collaborate on tasks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Planner&lt;/strong&gt; :- memory-first orchestration. Queries memory before planning anything. Owns task-state memory. Decides what gets persisted after each turn.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Executor&lt;/strong&gt; :- task execution. Writes observation memories. Lower trust than Critic for factual claims.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Critic&lt;/strong&gt; :- factual review and contradiction correction. Highest trust for permanent and preference memories. Can invalidate incorrect stored facts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reflection&lt;/strong&gt; :- scheduled background agent. Generates semantic patterns from non-reflective evidence. Does not participate in the synchronous chat loop.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F57f3gk4jd0dt1suqi74s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F57f3gk4jd0dt1suqi74s.png" alt="Agent Society" width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A critical design decision: &lt;strong&gt;no agent writes directly to the memory graph&lt;/strong&gt;. All persistent memory writes go through a single &lt;code&gt;MemoryManager&lt;/code&gt; service. This enforces a trust hierarchy; when Planner and Critic disagree about a stored fact, the Critic's version wins for permanent and preference memories. When write conflicts occur simultaneously, both versions are preserved with provenance tags until a resolution signal arrives.&lt;/p&gt;

&lt;p&gt;This means the memory is not just accurate, it is &lt;em&gt;accountable&lt;/em&gt;. Every episode knows which agent created it, when it was created, and whether it was ever superseded.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the System Looks Like
&lt;/h2&gt;

&lt;p&gt;The HGVM dashboard is built to make memory formation visible in real time. The Memory Graph page shows the full hierarchy as an interactive graph on a dark background: Domain hexagons in teal, Category rectangles, Topic circles with summary previews, and Episode nodes color-coded by strength, bright cyan for strong memories, yellow for medium, orange for fading.&lt;/p&gt;

&lt;p&gt;New episodes appear with a fade-in animation. Strength color transitions animate as decay runs. When consolidation fires on a topic, a brief pulse effect signals the compression happening. Reflection outputs appear as double-ring nodes, visually distinct from raw and summary episodes.&lt;/p&gt;

&lt;p&gt;A live Memory Activity Feed alongside the chat interface shows every memory operation in real time, green for ADD, blue for QUERY, red for INVALIDATE, purple for CONSOLIDATE, teal for REFLECT, grey for FORGET. You can watch the system's memory evolve turn by turn.&lt;/p&gt;

&lt;p&gt;An Analytics page tracks memory growth by class over time, token savings from compression, consolidation verification health, reflection utilization rates, retrieval configuration, and topic reuse effectiveness.&lt;/p&gt;

&lt;p&gt;The goal is to make memory legible, not just to researchers, but to anyone using the system.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Hard Problems We Have Not Solved
&lt;/h2&gt;

&lt;p&gt;We want to be clear about what remains genuinely difficult, because intellectual honesty matters more than clean marketing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decay parameter tuning.&lt;/strong&gt; The λ values for each memory class are educated starting points informed by cognitive science and ablation experiments. The right values for a specific deployment domain — medical, coding, creative writing, likely differ. Universal optimal values do not exist yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Machine unlearning.&lt;/strong&gt; GDPR's right to be forgotten is straightforward for the graph: a cascade deletion removes all a user's data atomically. It is not straightforward for the LLM itself. If an LLM has been exposed to a user's memories in-context during inference, traces of that exposure may persist in ways that are not cleanly deletable. Machine unlearning for in-context exposure is still an open research problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Semantic conflict resolution at scale.&lt;/strong&gt; When two agents write contradictory facts at the same timestamp, the trust hierarchy handles it. But there are edge cases :- domain-dependent conflicts, partial contradictions, nuanced updates that partially overlap with stored facts — where the right resolution depends on context that no static rule captures. We use a conservative approach (preserve both, resolve later) but this is a heuristic, not a solution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reflection quality at low evidence density.&lt;/strong&gt; When a user is new and has generated few episodes, reflection has little to work with. Thin-evidence reflections risk being generic or incorrect. We suppress reflections with insufficient supporting evidence, but the threshold is calibrated manually rather than learned.&lt;/p&gt;

&lt;p&gt;These are real limitations. We name them because they are the honest frontier of this problem space, and because solving them is where the interesting research goes next.&lt;/p&gt;




&lt;h2&gt;
  
  
  What We Are Building Toward
&lt;/h2&gt;

&lt;p&gt;HGVM v2.0 is being built for the Global AI Hackathon Series with Qwen Cloud with a tight deadline, but the ideas behind it are not hackathton ideas. The question of how AI agents should manage long-term memory is one of the most practically important open questions in applied AI right now.&lt;/p&gt;

&lt;p&gt;Every production AI assistant deployment faces this problem. Every enterprise deploying agents across thousands of users faces this problem. Every personal assistant that fails to remember what you told it last week fails because of this problem.&lt;/p&gt;

&lt;p&gt;We think the right direction involves all four pillars working together: hierarchical structure for precision, tiered forgetting for noise control, consolidation for compression, and reflection for higher-order learning. Not as research curiosities but as production-grade engineering.&lt;/p&gt;

&lt;p&gt;The full architecture, schema, pipeline specifications, and implementation plan are documented internally and will be shared progressively as the system matures. The dashboard will be publicly demoed when it is stable.&lt;/p&gt;

&lt;p&gt;If you are working on agent memory, long-term personalization, or multi-agent coordination, we would genuinely like to hear what you think. The hard problems above are not going to be solved alone.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;HGVM is being built as part of a focused 30-day engineering sprint. Architecture is frozen. Implementation is active.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If this resonates with work you are doing, reach out.&lt;/em&gt;&lt;/p&gt;




</description>
      <category>agents</category>
      <category>agentsmemory</category>
      <category>llm</category>
      <category>ai</category>
    </item>
    <item>
      <title>Why AI Assistants Have an Amnesia Problem. And How We're Fixing It</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Sat, 06 Jun 2026 08:30:00 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/why-ai-assistants-have-an-amnesia-problem-and-how-were-fixing-it-3g1e</link>
      <guid>https://dev.to/shogun_the_grt/why-ai-assistants-have-an-amnesia-problem-and-how-were-fixing-it-3g1e</guid>
      <description>&lt;p&gt;&lt;em&gt;A deep dive into HGVM: the memory architecture that makes AI agents remember like humans, forget like humans, and learn like humans.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;There is a moment every frequent AI user has experienced. You spent twenty minutes in a conversation explaining your project, your constraints, your preferences, your stack. The assistant was helpful, precise, contextually aware. Then you closed the tab.&lt;/p&gt;

&lt;p&gt;The next day, you open a new conversation. The assistant greets you like a stranger.&lt;/p&gt;

&lt;p&gt;You explain everything again.&lt;/p&gt;

&lt;p&gt;This is not a minor inconvenience. It is a fundamental architectural failure, and it is not the only one. On the other side of the spectrum are systems that remember &lt;em&gt;everything&lt;/em&gt; forever: every offhand comment, every temporary preference, every piece of stale context from six months ago. These systems don't forget, but they accumulate noise until their memory becomes a liability. They surface outdated information as confidently as current facts. They become sycophantic, anchored to what you said before rather than what is true now.&lt;/p&gt;

&lt;p&gt;Neither model is how memory should work. Neither model is how &lt;em&gt;human&lt;/em&gt; memory works.&lt;/p&gt;

&lt;p&gt;We built HGVM, Hierarchical Graph-Vector Memory, to fix this. This blog explains what it is, how it works, and why we think it represents a meaningfully different approach to AI agent memory.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Human Memory Analogy
&lt;/h2&gt;

&lt;p&gt;Human memory is not a database. It does not store everything with equal fidelity and retrieve it on demand. It is a dynamic, adaptive system that continuously organizes, compresses, reinforces, and discards information based on relevance, recency, and repeated exposure.&lt;/p&gt;

&lt;p&gt;When you learn something new, your brain does not simply file it. It connects it to existing structures. It consolidates related fragments into coherent abstractions during sleep. It strengthens memories that are repeatedly accessed and allows others to fade. And crucially, it reflects, it notices patterns across experiences and crystallizes them into durable knowledge that shapes future reasoning.&lt;/p&gt;

&lt;p&gt;HGVM is built on this intuition. Memory should be a &lt;em&gt;managed resource&lt;/em&gt;, not a growing pile. It should be structured hierarchically, retrieved selectively, compressed into abstractions over time, and actively pruned when it is no longer relevant. Four behaviors, working together.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Four Pillars of HGVM
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Hierarchical Structure
&lt;/h3&gt;

&lt;p&gt;Most existing memory systems store memories as a flat collection of text chunks with embeddings. When you query, you run a vector similarity search across all of them and return the top results. This works adequately at small scale, but it fails in two important ways as the memory grows.&lt;/p&gt;

&lt;p&gt;First, it causes &lt;strong&gt;context leakage&lt;/strong&gt;, semantically similar but contextually irrelevant memories surface because the retrieval system has no way to distinguish between "similar in meaning" and "relevant to this situation." A question about your Python project might surface memories about a Python script you wrote for a completely different purpose two months ago.&lt;/p&gt;

&lt;p&gt;Second, it is expensive. Vector search across all memories grows linearly. The more you remember, the slower and noisier retrieval becomes.&lt;/p&gt;

&lt;p&gt;HGVM organizes memory into a strict four-layer hierarchy:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domain → Category → Topic → Episode&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A Domain is the broadest grouping: Work, Health, Personal, Finance. A Category is a sub-domain: ProjectX, Morning Runs, Family. A Topic is a coherent theme within a category: Backend API Design, Sprint Planning, Dad's Hospital Visit. An Episode is a single atomic memory item attached to a topic.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmbw5t2jnfh23f03dqqla.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmbw5t2jnfh23f03dqqla.png" alt="Core Architecture Diagram" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When a query arrives, the system routes at the Domain and Category level first, a cheap semantic match against summaries, and then runs vector search only within the relevant sub-graph. The candidate set is small. Retrieval is precise. Context leakage is dramatically reduced because irrelevant branches are never even considered.&lt;/p&gt;

&lt;p&gt;This is the structural foundation everything else builds on.&lt;/p&gt;




&lt;h3&gt;
  
  
  2. Active Forgetting
&lt;/h3&gt;

&lt;p&gt;Forgetting is not a failure mode. It is a feature.&lt;/p&gt;

&lt;p&gt;The reason most AI memory systems remember everything forever is that forgetting feels dangerous, what if you delete something important? The result is a system that accumulates noise indefinitely, where the signal-to-noise ratio of the memory store degrades over time and retrieval quality degrades with it.&lt;/p&gt;

&lt;p&gt;HGVM takes a different position: forgetting should be &lt;strong&gt;tiered by memory type&lt;/strong&gt;, not uniform. Not all memories deserve the same treatment.&lt;/p&gt;

&lt;p&gt;We define five memory classes, each with its own decay behavior:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Memory Class&lt;/th&gt;
&lt;th&gt;What It Stores&lt;/th&gt;
&lt;th&gt;Decay Rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;permanent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Explicit identity facts ("I am vegetarian", "never forget this")&lt;/td&gt;
&lt;td&gt;Never decays&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;semantic&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Consolidated summaries and learned patterns&lt;/td&gt;
&lt;td&gt;Very slow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;preference&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Stable user preferences&lt;/td&gt;
&lt;td&gt;Very slow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;observation&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Raw observations and task outcomes&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;task_state&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Active work state and in-progress details&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxb34x0oaj93h0ych0m0r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxb34x0oaj93h0ych0m0r.png" alt="Memory Class Spectrum" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A permanent memory :- something the user explicitly told the system to remember, never decays. A preference :- "I prefer TypeScript", fades very slowly, remaining useful for months. A task_state memory :- "currently blocked on the auth middleware", fades quickly, becoming stale within days because task context changes fast.&lt;/p&gt;

&lt;p&gt;The decay formula itself is borrowed from cognitive science. We use an exponential decay function where the strength of each memory degrades over time since last access:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;new_strength = current_strength × exp(−λ × days_since_access)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each memory class has its own λ value. When strength falls below a minimum threshold, the memory is &lt;strong&gt;soft-deleted&lt;/strong&gt; :- marked as invalid but preserved in the graph for auditability. Nothing is ever hard-deleted during forgetting.&lt;/p&gt;

&lt;p&gt;Retrieval is reinforced too: every time a memory is returned in a query, its strength is boosted. Memories that are repeatedly useful stay strong. Memories that are never relevant fade gracefully.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7sxk3i23uj0asf1xqwnd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7sxk3i23uj0asf1xqwnd.png" alt="Forgetting Curve Comparison" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the forgetting curve most existing systems ignore. HGVM embraces it as a design principle.&lt;/p&gt;




&lt;h3&gt;
  
  
  3. Consolidation
&lt;/h3&gt;

&lt;p&gt;A topic with fifty raw episode nodes is not useful. Fifty individually embedded fragments retrieved piecemeal create noisy, redundant context. At some point, accumulated raw memories need to be compressed into something denser and more useful.&lt;/p&gt;

&lt;p&gt;This is consolidation.&lt;/p&gt;

&lt;p&gt;HGVM runs a scheduled consolidation pipeline that works as follows: when a topic accumulates enough raw episodes, the system clusters them by semantic similarity, extracts atomic facts from each cluster, generates a concise summary from those facts, and then runs a &lt;strong&gt;batched verification pass&lt;/strong&gt; to confirm the summary is faithful to the source material.&lt;/p&gt;

&lt;p&gt;The verification step is non-negotiable. We check for three failure conditions: missing facts (something in the source that did not make it into the summary), altered facts (something that changed in meaning), and contradicted facts (something that directly conflicts with the source). If any of these conditions are non-empty, the consolidation is rejected and logged. No summary episode is created until verification passes completely.&lt;/p&gt;

&lt;p&gt;When verification passes, the system creates a new &lt;code&gt;semantic&lt;/code&gt; episode with subtype &lt;code&gt;consolidation_summary&lt;/code&gt;, linked to the source episodes via &lt;code&gt;SUMMARIZES&lt;/code&gt; relationships. Critically, &lt;strong&gt;the source episodes are never deleted&lt;/strong&gt;. Consolidation adds compressed knowledge — it does not destroy provenance.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhy7xf43tg2ckxwwwjuz2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhy7xf43tg2ckxwwwjuz2.png" alt="Retrieval Pipeline Flow" width="800" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The retrieval system then naturally prefers the summary for most queries, it is denser, more representative, and semantically richer, while the raw episodes remain available for queries that need granular detail.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F90cgsat0d5ttoitbmua6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F90cgsat0d5ttoitbmua6.png" alt="Consolidation Before/After" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This mirrors how human memory consolidates during sleep: the raw experiences remain somewhere in the system, but what you access day-to-day is the compressed, organized version.&lt;/p&gt;




&lt;h3&gt;
  
  
  4. Reflection
&lt;/h3&gt;

&lt;p&gt;Consolidation compresses. Reflection &lt;em&gt;learns&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Reflection is the most novel component of HGVM. While consolidation creates summaries within a single topic, reflection operates across topics and sessions, looking for higher-order patterns in behavior, preference, and reasoning.&lt;/p&gt;

&lt;p&gt;Imagine a user who across ten sessions has: rejected Java twice, chosen Python for three consecutive projects, consistently used FastAPI over Django, and complained about verbose enterprise frameworks. No single episode says "this user values developer productivity." But the pattern across episodes does.&lt;/p&gt;

&lt;p&gt;Reflection detects this. It fetches recent valid episodes from the non-reflective memory classes :- &lt;code&gt;permanent&lt;/code&gt;, &lt;code&gt;preference&lt;/code&gt;, &lt;code&gt;task_state&lt;/code&gt;, and &lt;code&gt;observation&lt;/code&gt;, groups them into coherent evidence bundles, and generates candidate higher-order abstractions. These are persisted as &lt;code&gt;semantic&lt;/code&gt; episodes with subtype &lt;code&gt;reflection_pattern&lt;/code&gt;, linked to their supporting evidence.&lt;/p&gt;

&lt;p&gt;The result: the next time the user asks for a project scaffold recommendation, the system does not just recall that they used FastAPI before. It retrieves the reflection, "this user prioritizes developer productivity over ecosystem conservatism", and reasons from that durable insight.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Famwbbs3n9jn4wdl77xtw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Famwbbs3n9jn4wdl77xtw.png" alt="Reflection Explained" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There is one hard constraint we enforce at the architecture level: &lt;strong&gt;reflection memories can never be used as evidence for future reflections&lt;/strong&gt;. This prevents a feedback loop where the system reflects on its own abstractions, compounding distortions over time. Reflection operates only on primary evidence. This constraint is enforced at the database query level, not just the prompt level.&lt;/p&gt;




&lt;h2&gt;
  
  
  How HGVM Compares to What Exists
&lt;/h2&gt;

&lt;p&gt;The agent memory landscape has become more sophisticated in the last two years. It is worth being honest about where existing systems land and where HGVM goes further.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MemGPT / Letta&lt;/strong&gt; introduced the idea of paging between active context and external memory using function calls, and Letta added sleep-time consolidation agents. HGVM builds directly on these ideas but adds hierarchical routing, tiered forgetting, and the reflection pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mem0&lt;/strong&gt; offers a clean multi-signal retrieval system (semantic + keyword + entity) with good API design. It does not offer hierarchical memory organization, tiered decay, or reflection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zep / Graphiti&lt;/strong&gt; provide excellent graph-based memory with temporal reasoning and the bitemporal model (facts carry both a valid time and a transaction time). HGVM adopts the bitemporal approach for contradiction handling and uses Graphiti's &lt;code&gt;valid_at&lt;/code&gt; / &lt;code&gt;invalid_at&lt;/code&gt; pattern for soft deletion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT Memory&lt;/strong&gt; is a static injection system, remembered facts are prepended to context. This causes cross-domain leakage (facts from unrelated contexts surface inappropriately) and sycophancy (the system defers to stored preferences even when they are stale or wrong). PersistBench has documented both failure modes.&lt;/p&gt;

&lt;p&gt;HGVM is not a critique of any of these systems. Each solved a real problem. HGVM synthesizes the best ideas :- graph structure, temporal modeling, consolidation, multi-signal retrieval and adds forgetting as a first-class feature and reflection as a novel capability.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Multi-Agent Layer
&lt;/h2&gt;

&lt;p&gt;HGVM is not just a memory system. It is a memory system designed for an agent society.&lt;/p&gt;

&lt;p&gt;Four specialized LLM agents share the same memory graph and collaborate on tasks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Planner&lt;/strong&gt; :- memory-first orchestration. Queries memory before planning anything. Owns task-state memory. Decides what gets persisted after each turn.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Executor&lt;/strong&gt; :- task execution. Writes observation memories. Lower trust than Critic for factual claims.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Critic&lt;/strong&gt; :- factual review and contradiction correction. Highest trust for permanent and preference memories. Can invalidate incorrect stored facts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reflection&lt;/strong&gt; :- scheduled background agent. Generates semantic patterns from non-reflective evidence. Does not participate in the synchronous chat loop.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F57f3gk4jd0dt1suqi74s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F57f3gk4jd0dt1suqi74s.png" alt="Agent Society" width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A critical design decision: &lt;strong&gt;no agent writes directly to the memory graph&lt;/strong&gt;. All persistent memory writes go through a single &lt;code&gt;MemoryManager&lt;/code&gt; service. This enforces a trust hierarchy; when Planner and Critic disagree about a stored fact, the Critic's version wins for permanent and preference memories. When write conflicts occur simultaneously, both versions are preserved with provenance tags until a resolution signal arrives.&lt;/p&gt;

&lt;p&gt;This means the memory is not just accurate, it is &lt;em&gt;accountable&lt;/em&gt;. Every episode knows which agent created it, when it was created, and whether it was ever superseded.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the System Looks Like
&lt;/h2&gt;

&lt;p&gt;The HGVM dashboard is built to make memory formation visible in real time. The Memory Graph page shows the full hierarchy as an interactive graph on a dark background: Domain hexagons in teal, Category rectangles, Topic circles with summary previews, and Episode nodes color-coded by strength, bright cyan for strong memories, yellow for medium, orange for fading.&lt;/p&gt;

&lt;p&gt;New episodes appear with a fade-in animation. Strength color transitions animate as decay runs. When consolidation fires on a topic, a brief pulse effect signals the compression happening. Reflection outputs appear as double-ring nodes, visually distinct from raw and summary episodes.&lt;/p&gt;

&lt;p&gt;A live Memory Activity Feed alongside the chat interface shows every memory operation in real time, green for ADD, blue for QUERY, red for INVALIDATE, purple for CONSOLIDATE, teal for REFLECT, grey for FORGET. You can watch the system's memory evolve turn by turn.&lt;/p&gt;

&lt;p&gt;An Analytics page tracks memory growth by class over time, token savings from compression, consolidation verification health, reflection utilization rates, retrieval configuration, and topic reuse effectiveness.&lt;/p&gt;

&lt;p&gt;The goal is to make memory legible, not just to researchers, but to anyone using the system.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Hard Problems We Have Not Solved
&lt;/h2&gt;

&lt;p&gt;We want to be clear about what remains genuinely difficult, because intellectual honesty matters more than clean marketing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decay parameter tuning.&lt;/strong&gt; The λ values for each memory class are educated starting points informed by cognitive science and ablation experiments. The right values for a specific deployment domain — medical, coding, creative writing, likely differ. Universal optimal values do not exist yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Machine unlearning.&lt;/strong&gt; GDPR's right to be forgotten is straightforward for the graph: a cascade deletion removes all a user's data atomically. It is not straightforward for the LLM itself. If an LLM has been exposed to a user's memories in-context during inference, traces of that exposure may persist in ways that are not cleanly deletable. Machine unlearning for in-context exposure is still an open research problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Semantic conflict resolution at scale.&lt;/strong&gt; When two agents write contradictory facts at the same timestamp, the trust hierarchy handles it. But there are edge cases :- domain-dependent conflicts, partial contradictions, nuanced updates that partially overlap with stored facts — where the right resolution depends on context that no static rule captures. We use a conservative approach (preserve both, resolve later) but this is a heuristic, not a solution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reflection quality at low evidence density.&lt;/strong&gt; When a user is new and has generated few episodes, reflection has little to work with. Thin-evidence reflections risk being generic or incorrect. We suppress reflections with insufficient supporting evidence, but the threshold is calibrated manually rather than learned.&lt;/p&gt;

&lt;p&gt;These are real limitations. We name them because they are the honest frontier of this problem space, and because solving them is where the interesting research goes next.&lt;/p&gt;




&lt;h2&gt;
  
  
  What We Are Building Toward
&lt;/h2&gt;

&lt;p&gt;HGVM v2.0 is being built for the Global AI Hackathon Series with Qwen Cloud with a tight deadline, but the ideas behind it are not hackathton ideas. The question of how AI agents should manage long-term memory is one of the most practically important open questions in applied AI right now.&lt;/p&gt;

&lt;p&gt;Every production AI assistant deployment faces this problem. Every enterprise deploying agents across thousands of users faces this problem. Every personal assistant that fails to remember what you told it last week fails because of this problem.&lt;/p&gt;

&lt;p&gt;We think the right direction involves all four pillars working together: hierarchical structure for precision, tiered forgetting for noise control, consolidation for compression, and reflection for higher-order learning. Not as research curiosities but as production-grade engineering.&lt;/p&gt;

&lt;p&gt;The full architecture, schema, pipeline specifications, and implementation plan are documented internally and will be shared progressively as the system matures. The dashboard will be publicly demoed when it is stable.&lt;/p&gt;

&lt;p&gt;If you are working on agent memory, long-term personalization, or multi-agent coordination, we would genuinely like to hear what you think. The hard problems above are not going to be solved alone.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;HGVM is being built as part of a focused 30-day engineering sprint. Architecture is frozen. Implementation is active.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If this resonates with work you are doing, reach out.&lt;/em&gt;&lt;/p&gt;




</description>
      <category>agents</category>
      <category>agentsmemory</category>
      <category>llm</category>
    </item>
    <item>
      <title>From Scripts to Systems: What Learning Kestra Taught Me About Workflow Orchestration</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Fri, 08 May 2026 18:32:13 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/from-scripts-to-systems-what-learning-kestra-taught-me-about-workflow-orchestration-5575</link>
      <guid>https://dev.to/shogun_the_grt/from-scripts-to-systems-what-learning-kestra-taught-me-about-workflow-orchestration-5575</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Before starting the Kestra Fundamentals course, I used to think workflow orchestration was mostly about scheduling scripts.&lt;/p&gt;

&lt;p&gt;Run a cron job.&lt;br&gt;
Execute a Python script.&lt;br&gt;
Send an email.&lt;br&gt;
Done.&lt;/p&gt;

&lt;p&gt;But the deeper I went into workflow orchestration, the more I realized modern systems are far more complex than that.&lt;/p&gt;

&lt;p&gt;Today, applications depend on APIs, databases, cloud services, analytics pipelines, notifications, and event-driven systems all working together in the correct order.&lt;/p&gt;

&lt;p&gt;The challenge is no longer just writing code.&lt;/p&gt;

&lt;p&gt;The real challenge is coordinating systems reliably.&lt;/p&gt;

&lt;p&gt;That’s where workflow orchestration comes in.&lt;/p&gt;

&lt;p&gt;Through the Kestra Fundamentals course by WeMakeDevs, I learned how orchestration platforms help engineers build reliable, automated, observable workflows instead of disconnected scripts.&lt;/p&gt;

&lt;p&gt;And honestly, this course changed the way I think about backend systems.&lt;/p&gt;


&lt;h2&gt;
  
  
  What Workflow Orchestration Actually Means
&lt;/h2&gt;

&lt;p&gt;The easiest analogy for workflow orchestration is an orchestra.&lt;/p&gt;

&lt;p&gt;Different musicians play different instruments.&lt;br&gt;
Some enter early.&lt;br&gt;
Some wait.&lt;br&gt;
Some depend on others.&lt;/p&gt;

&lt;p&gt;Without coordination, everything becomes noise.&lt;/p&gt;

&lt;p&gt;The same thing happens in software systems.&lt;/p&gt;

&lt;p&gt;You might have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;APIs fetching data&lt;/li&gt;
&lt;li&gt;scripts processing it&lt;/li&gt;
&lt;li&gt;databases storing it&lt;/li&gt;
&lt;li&gt;analytics pipelines transforming it&lt;/li&gt;
&lt;li&gt;notifications sending results&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All these systems need to work together in the right order.&lt;/p&gt;

&lt;p&gt;Workflow orchestration is the layer that coordinates all of this.&lt;/p&gt;

&lt;p&gt;It helps with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;sequencing tasks&lt;/li&gt;
&lt;li&gt;handling dependencies&lt;/li&gt;
&lt;li&gt;retries and failures&lt;/li&gt;
&lt;li&gt;automation&lt;/li&gt;
&lt;li&gt;scheduling&lt;/li&gt;
&lt;li&gt;monitoring&lt;/li&gt;
&lt;li&gt;observability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of writing isolated scripts and hoping they work, orchestration platforms let you design reliable systems.&lt;/p&gt;


&lt;h2&gt;
  
  
  Why Kestra Felt Different
&lt;/h2&gt;

&lt;p&gt;One thing I immediately liked about Kestra was how structured everything felt.&lt;/p&gt;

&lt;p&gt;Workflows are defined declaratively using YAML.&lt;/p&gt;

&lt;p&gt;Instead of manually stitching together scripts, you define workflows in a clean and readable format.&lt;/p&gt;

&lt;p&gt;Kestra also provides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;execution tracking&lt;/li&gt;
&lt;li&gt;visual workflow monitoring&lt;/li&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;scheduling&lt;/li&gt;
&lt;li&gt;triggers&lt;/li&gt;
&lt;li&gt;plugins&lt;/li&gt;
&lt;li&gt;reusable blueprints&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What stood out to me most was that Kestra didn’t feel like just another automation tool.&lt;/p&gt;

&lt;p&gt;It felt like a system orchestration platform.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frhwfenjo06iumxg5m21n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frhwfenjo06iumxg5m21n.png" alt=" " width="800" height="413"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Core Concepts That Finally Clicked
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Flows
&lt;/h3&gt;

&lt;p&gt;A Flow is the main orchestration unit in Kestra.&lt;/p&gt;

&lt;p&gt;This is where the overall workflow is defined:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tasks&lt;/li&gt;
&lt;li&gt;triggers&lt;/li&gt;
&lt;li&gt;inputs&lt;/li&gt;
&lt;li&gt;outputs&lt;/li&gt;
&lt;li&gt;orchestration logic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At first, I thought flows were just pipelines.&lt;/p&gt;

&lt;p&gt;But eventually, they started feeling more like system blueprints.&lt;/p&gt;


&lt;h3&gt;
  
  
  Tasks
&lt;/h3&gt;

&lt;p&gt;Tasks are the individual units of work.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;calling an API&lt;/li&gt;
&lt;li&gt;running a script&lt;/li&gt;
&lt;li&gt;querying a database&lt;/li&gt;
&lt;li&gt;sending notifications&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One thing I liked about Kestra’s design is how composable tasks feel.&lt;/p&gt;

&lt;p&gt;Each task focuses on one responsibility.&lt;/p&gt;

&lt;p&gt;That makes workflows much easier to reason about.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbtmkmlkfttyse57bfgrl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbtmkmlkfttyse57bfgrl.png" alt=" " width="800" height="413"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h3&gt;
  
  
  Inputs and Outputs
&lt;/h3&gt;

&lt;p&gt;This was one of the concepts that made orchestration really click for me.&lt;/p&gt;

&lt;p&gt;Inputs allow workflows to receive data dynamically.&lt;br&gt;
Outputs allow tasks to pass data to other tasks.&lt;/p&gt;

&lt;p&gt;So instead of disconnected scripts, you get connected systems.&lt;/p&gt;

&lt;p&gt;Example:&lt;/p&gt;

&lt;p&gt;Task A fetches API data → Task B processes it → Task C stores it.&lt;/p&gt;

&lt;p&gt;The clean data flow model makes workflows much easier to scale and debug.&lt;/p&gt;


&lt;h3&gt;
  
  
  Triggers
&lt;/h3&gt;

&lt;p&gt;Triggers are what make workflows truly automated.&lt;/p&gt;

&lt;p&gt;Instead of manually executing workflows, Kestra can trigger them based on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;schedules (cron)&lt;/li&gt;
&lt;li&gt;API events&lt;/li&gt;
&lt;li&gt;file arrivals&lt;/li&gt;
&lt;li&gt;workflow completions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This shifts workflows from manual execution to event-driven systems.&lt;/p&gt;


&lt;h3&gt;
  
  
  Expressions
&lt;/h3&gt;

&lt;p&gt;Expressions were another powerful concept.&lt;/p&gt;

&lt;p&gt;Kestra uses templating syntax to dynamically reference values and outputs.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight twig"&gt;&lt;code&gt;&lt;span class="cp"&gt;{{&lt;/span&gt; &lt;span class="nv"&gt;outputs.task_id.value&lt;/span&gt; &lt;span class="cp"&gt;}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This allows workflows to become dynamic and programmable instead of static pipelines.&lt;/p&gt;




&lt;h2&gt;
  
  
  Flowable Tasks Changed How I Think About Orchestration
&lt;/h2&gt;

&lt;p&gt;This section was probably one of the biggest mindset shifts for me.&lt;/p&gt;

&lt;p&gt;Flowable tasks are not just tasks that execute work.&lt;/p&gt;

&lt;p&gt;They control orchestration logic itself.&lt;/p&gt;

&lt;p&gt;This includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;loops&lt;/li&gt;
&lt;li&gt;conditionals&lt;/li&gt;
&lt;li&gt;parallel execution&lt;/li&gt;
&lt;li&gt;branching&lt;/li&gt;
&lt;li&gt;subflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Subflows especially felt similar to reusable functions in programming.&lt;/p&gt;

&lt;p&gt;Define once.&lt;br&gt;
Reuse everywhere.&lt;/p&gt;

&lt;p&gt;That realization made orchestration feel much closer to software architecture than simple automation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Execution, Observability, and Reliability
&lt;/h2&gt;

&lt;p&gt;One thing many beginners underestimate is observability.&lt;/p&gt;

&lt;p&gt;Running workflows is only part of the problem.&lt;/p&gt;

&lt;p&gt;Understanding:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what failed&lt;/li&gt;
&lt;li&gt;why it failed&lt;/li&gt;
&lt;li&gt;where it failed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;is equally important.&lt;/p&gt;

&lt;p&gt;Kestra’s execution tracking and logs make workflows much easier to debug.&lt;/p&gt;

&lt;p&gt;Instead of guessing what happened, you can inspect executions step-by-step.&lt;/p&gt;

&lt;p&gt;Retries and execution monitoring also make workflows significantly more reliable.&lt;/p&gt;




&lt;h2&gt;
  
  
  Secrets, Plugins, and Blueprints
&lt;/h2&gt;

&lt;p&gt;Another thing I appreciated was how practical the platform felt.&lt;/p&gt;

&lt;h3&gt;
  
  
  Secrets
&lt;/h3&gt;

&lt;p&gt;Secrets allow sensitive values like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API keys&lt;/li&gt;
&lt;li&gt;credentials&lt;/li&gt;
&lt;li&gt;tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;to stay secure instead of being hardcoded.&lt;/p&gt;




&lt;h3&gt;
  
  
  Plugins
&lt;/h3&gt;

&lt;p&gt;Kestra has a huge plugin ecosystem.&lt;/p&gt;

&lt;p&gt;This means workflows can integrate with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;APIs&lt;/li&gt;
&lt;li&gt;cloud services&lt;/li&gt;
&lt;li&gt;databases&lt;/li&gt;
&lt;li&gt;scripting environments&lt;/li&gt;
&lt;li&gt;data systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;without constantly reinventing integrations.&lt;/p&gt;




&lt;h3&gt;
  
  
  Blueprints
&lt;/h3&gt;

&lt;p&gt;Blueprints are reusable workflow templates.&lt;/p&gt;

&lt;p&gt;This is extremely useful when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;learning orchestration&lt;/li&gt;
&lt;li&gt;prototyping workflows quickly&lt;/li&gt;
&lt;li&gt;exploring integrations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of starting from scratch, you can learn from working examples.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsq1vvajfp0ex3e9slzl5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsq1vvajfp0ex3e9slzl5.png" alt=" " width="800" height="394"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Project I Built
&lt;/h2&gt;

&lt;p&gt;One of the biggest things I wanted after completing the course was to actually apply what I learned instead of stopping at theory.&lt;/p&gt;

&lt;p&gt;So I built a small workflow project using Kestra, an automated daily sales report pipeline.&lt;/p&gt;

&lt;p&gt;The goal of the workflow was simple:&lt;/p&gt;

&lt;p&gt;Automatically fetch sales data, process it, generate a report, and send it through email without any manual intervention.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the workflow does
&lt;/h3&gt;

&lt;p&gt;The workflow follows this sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A cron trigger starts the workflow automatically on schedule.&lt;/li&gt;
&lt;li&gt;Sales data is fetched from a PostgreSQL database.&lt;/li&gt;
&lt;li&gt;The raw data is processed and transformed using Python scripts.&lt;/li&gt;
&lt;li&gt;A summary report is generated.&lt;/li&gt;
&lt;li&gt;The report is exported as a CSV file.&lt;/li&gt;
&lt;li&gt;The final report is sent via email.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This project helped me understand how orchestration systems coordinate multiple dependent tasks reliably.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fw1c78s7km0wb8gp98fu0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fw1c78s7km0wb8gp98fu0.png" alt=" " width="800" height="412"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  What I Learned While Building It
&lt;/h3&gt;

&lt;p&gt;Building this workflow made several orchestration concepts finally click for me.&lt;/p&gt;

&lt;h4&gt;
  
  
  Dependencies and execution order
&lt;/h4&gt;

&lt;p&gt;Each task depended on outputs from previous tasks.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;report generation depends on transformed data&lt;/li&gt;
&lt;li&gt;email sending depends on successful CSV generation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This made workflow dependencies feel much more real than isolated scripts.&lt;/p&gt;




&lt;h4&gt;
  
  
  Inputs and Outputs
&lt;/h4&gt;

&lt;p&gt;Passing data between tasks using outputs was one of the most interesting parts.&lt;/p&gt;

&lt;p&gt;Instead of manually handling intermediate files everywhere, the workflow itself became the coordination layer.&lt;/p&gt;




&lt;h4&gt;
  
  
  Automation Design
&lt;/h4&gt;

&lt;p&gt;The trigger system showed how workflows can become completely automated.&lt;/p&gt;

&lt;p&gt;Once configured, the entire pipeline could run on schedule without manual execution.&lt;/p&gt;




&lt;h4&gt;
  
  
  Observability and Reliability
&lt;/h4&gt;

&lt;p&gt;Execution logs and task tracking made debugging much easier.&lt;/p&gt;

&lt;p&gt;Instead of guessing where failures happened, I could inspect each task execution step-by-step.&lt;/p&gt;

&lt;p&gt;That visibility is something normal scripts usually lack.&lt;/p&gt;




&lt;h4&gt;
  
  
  System Thinking
&lt;/h4&gt;

&lt;p&gt;The biggest realization for me was this:&lt;/p&gt;

&lt;p&gt;I wasn’t just writing scripts anymore.&lt;/p&gt;

&lt;p&gt;I was designing a system where:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tasks coordinate together&lt;/li&gt;
&lt;li&gt;data flows between steps&lt;/li&gt;
&lt;li&gt;failures can be monitored&lt;/li&gt;
&lt;li&gt;workflows execute automatically&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  That shift in mindset was probably the most valuable part of the entire course.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;The biggest shift for me after learning Kestra was realizing this:&lt;/p&gt;

&lt;p&gt;Modern engineering isn’t just about writing code.&lt;/p&gt;

&lt;p&gt;It’s about designing systems that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;coordinate work reliably&lt;/li&gt;
&lt;li&gt;automate execution&lt;/li&gt;
&lt;li&gt;handle dependencies&lt;/li&gt;
&lt;li&gt;recover from failures&lt;/li&gt;
&lt;li&gt;provide visibility into what’s happening&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Workflow orchestration sits at the center of all of this.&lt;/p&gt;

&lt;p&gt;And Kestra made these ideas surprisingly approachable.&lt;/p&gt;

&lt;p&gt;I originally started this course thinking orchestration was just advanced scheduling.&lt;/p&gt;

&lt;p&gt;I finished it understanding why orchestration is such an important part of modern software systems.&lt;/p&gt;

</description>
      <category>kestraacademy</category>
      <category>workfloworchestration</category>
      <category>kestra</category>
    </item>
    <item>
      <title>Stop Building Reactive AI — Build AI That Acts Before You Ask</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Sat, 25 Apr 2026 20:41:22 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/stop-building-reactive-ai-build-ai-that-acts-before-you-ask-4goj</link>
      <guid>https://dev.to/shogun_the_grt/stop-building-reactive-ai-build-ai-that-acts-before-you-ask-4goj</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/openclaw-2026-04-16"&gt;OpenClaw Writing Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem With Most AI
&lt;/h2&gt;

&lt;p&gt;Most AI systems today follow the same pattern:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You ask → it responds&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Chatbots, assistants, copilots — they all wait.&lt;/p&gt;

&lt;p&gt;But real-world problems don't work like that.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You don't notice when a task is slowly being ignored&lt;/li&gt;
&lt;li&gt;You don't realize work is at risk until it's too late&lt;/li&gt;
&lt;li&gt;You don't ask for help because you don't know there's a problem yet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI that waits for prompts is fundamentally limited.&lt;/p&gt;

&lt;p&gt;It reacts to awareness — but awareness is often the problem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsorfehbqx4i6br0v4jm0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsorfehbqx4i6br0v4jm0.png" alt="Reactive AI Model" width="800" height="82"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Reactive AI waits for input before acting&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Enter OpenClaw
&lt;/h2&gt;

&lt;p&gt;When I first heard about OpenClaw, I thought of it as just another AI assistant — something that sits on your machine and answers questions faster.&lt;/p&gt;

&lt;p&gt;I was wrong.&lt;/p&gt;

&lt;p&gt;OpenClaw describes itself as a personal AI assistant that runs on your own devices and answers you on the channels you already use — Telegram, Slack, Discord, WhatsApp, and dozens more. But the real insight isn't in the channels. It's in the &lt;strong&gt;Gateway&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The Gateway is OpenClaw's local control plane: a persistent process that's always running, always watching, and always ready to act. That changes everything.&lt;/p&gt;

&lt;p&gt;Because now the question isn't just &lt;em&gt;"what can my AI respond to?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The question becomes: &lt;em&gt;"what should my AI notice?"&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Shift: From Reactive to Proactive AI
&lt;/h2&gt;

&lt;p&gt;After building a workflow monitoring agent on top of OpenClaw, I realized something:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The real value of AI isn't answering questions.&lt;br&gt;
It's noticing problems before you do.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of the classic model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User → Prompt → AI Response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OpenClaw lets you build something fundamentally different:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Observe → Detect → Decide → Intervene
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI is no longer just a tool you use — it becomes something that &lt;strong&gt;works alongside you&lt;/strong&gt;, even when you're not thinking about it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fi9adfn8fxd7811ix1v9i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fi9adfn8fxd7811ix1v9i.png" alt="Proactive AI Model" width="455" height="1030"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Proactive AI continuously observes, detects, and intervenes&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  How OpenClaw Makes This Possible
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Gateway: Always On, Never Waiting
&lt;/h3&gt;

&lt;p&gt;Most AI tools live in a tab you have to open. OpenClaw's Gateway runs as a system service (via &lt;code&gt;launchd&lt;/code&gt; on macOS, &lt;code&gt;systemd&lt;/code&gt; on Linux). It's always on — like a background process for your life.&lt;/p&gt;

&lt;p&gt;This is the prerequisite for proactive AI. You can't observe the world if you only wake up when someone talks to you.&lt;/p&gt;




&lt;h3&gt;
  
  
  Skills: Teaching Your Agent What to Watch
&lt;/h3&gt;

&lt;p&gt;OpenClaw's extensible &lt;strong&gt;Skills&lt;/strong&gt; system is what lets you encode custom logic into your agent. A skill is just a folder with a &lt;code&gt;SKILL.md&lt;/code&gt; file — plain, hackable, yours.&lt;/p&gt;

&lt;p&gt;For my workflow monitoring agent, I wrote a skill that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Checks open tasks across my project board&lt;/li&gt;
&lt;li&gt;Measures how many days have passed since the last commit&lt;/li&gt;
&lt;li&gt;Counts files that are modified but uncommitted&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this required complex infrastructure. Just a skill, some shell commands, and OpenClaw's tool-calling loop.&lt;/p&gt;




&lt;h3&gt;
  
  
  Cron-Style Loops: The Heartbeat of Proactive AI
&lt;/h3&gt;

&lt;p&gt;OpenClaw supports scheduled and recurring agent runs. This is the heartbeat of the proactive pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;check → analyze → act → repeat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every 20 minutes, my agent wakes up, runs the skill, evaluates current state, and decides whether to intervene. Most of the time, it does nothing. But when it detects something worth flagging — a task ignored for 3 days, an uncommitted spike of files — it reaches out.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fr8vt3ig761kt0a3boyni.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fr8vt3ig761kt0a3boyni.png" alt="Proactive Loop Example" width="713" height="127"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Simple loop: check → analyze → act → repeat&lt;/em&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Channels: Reaching You Where You Already Are
&lt;/h3&gt;

&lt;p&gt;Because OpenClaw routes through messaging platforms, the intervention doesn't feel like a notification from yet another app.&lt;/p&gt;

&lt;p&gt;It feels like a message from someone paying attention.&lt;/p&gt;

&lt;p&gt;When my agent pinged me on Telegram — &lt;em&gt;"This task has been sitting for 3 days. You'll likely miss the deadline — prioritize it today."&lt;/em&gt; — I actually acted on it. Because it arrived in the same place my real messages do, in plain language, at the right moment.&lt;/p&gt;

&lt;p&gt;That's the difference between logging a risk and actually changing behavior.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Changes When AI Stops Waiting
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. It Works Without Being Asked
&lt;/h3&gt;

&lt;p&gt;You don't need to remember to use it.&lt;/p&gt;

&lt;p&gt;It's already watching:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;patterns&lt;/li&gt;
&lt;li&gt;delays&lt;/li&gt;
&lt;li&gt;risks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The system becomes ambient — always present, never intrusive.&lt;/p&gt;




&lt;h3&gt;
  
  
  2. It Solves Problems You Didn't See
&lt;/h3&gt;

&lt;p&gt;Reactive AI can only solve known problems.&lt;/p&gt;

&lt;p&gt;Proactive AI surfaces &lt;strong&gt;unknown problems&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;forgotten tasks&lt;/li&gt;
&lt;li&gt;risky workflows&lt;/li&gt;
&lt;li&gt;accumulating technical debt&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where most real-world value comes from.&lt;/p&gt;




&lt;h3&gt;
  
  
  3. It Feels Less Like a Tool, More Like a System
&lt;/h3&gt;

&lt;p&gt;A chatbot feels like software.&lt;/p&gt;

&lt;p&gt;A proactive OpenClaw agent feels like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a safety net&lt;/li&gt;
&lt;li&gt;a second layer of awareness&lt;/li&gt;
&lt;li&gt;something that actively supports your workflow&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That shift is subtle — but powerful.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Learned Building This
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Simplicity Beats Complexity
&lt;/h3&gt;

&lt;p&gt;The system didn't need to be perfect.&lt;/p&gt;

&lt;p&gt;Even simple signals — like "days since last commit" — were enough to detect meaningful risk. OpenClaw's skill system made it easy to start small and iterate.&lt;/p&gt;




&lt;h3&gt;
  
  
  Context Matters More Than Rules
&lt;/h3&gt;

&lt;p&gt;Hardcoded thresholds failed quickly.&lt;/p&gt;

&lt;p&gt;What mattered was context. Five uncommitted files might be fine — or catastrophic, depending on the deadline. OpenClaw's LLM-powered evaluation layer is what made the difference: instead of a rule engine, I had something that could &lt;em&gt;interpret&lt;/em&gt;, not just calculate.&lt;/p&gt;




&lt;h3&gt;
  
  
  Small Interventions Create Big Impact
&lt;/h3&gt;

&lt;p&gt;The agent didn't need to automate everything.&lt;/p&gt;

&lt;p&gt;It just needed to say the right thing at the right time, through the right channel.&lt;/p&gt;

&lt;p&gt;That was enough to change behavior.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to Build This Yourself with OpenClaw
&lt;/h2&gt;

&lt;p&gt;If you want to try this pattern, here's where to start:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Get OpenClaw Running
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://openclaw.ai/install.sh | bash
openclaw onboard &lt;span class="nt"&gt;--install-daemon&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The onboarding wizard walks you through model provider setup, API keys, and Gateway configuration in about 2 minutes.&lt;/p&gt;




&lt;h3&gt;
  
  
  2. Connect a Channel
&lt;/h3&gt;

&lt;p&gt;Telegram is the easiest first channel. Once it's connected, your agent can reach you wherever you are — no new app needed.&lt;/p&gt;




&lt;h3&gt;
  
  
  3. Write a Simple Skill
&lt;/h3&gt;

&lt;p&gt;Create a folder at &lt;code&gt;~/.openclaw/workspace/skills/workflow-monitor/&lt;/code&gt; with a &lt;code&gt;SKILL.md&lt;/code&gt; file. Define what your agent should observe: stale tasks, missed commits, overdue tickets — whatever your workflow surfaces.&lt;/p&gt;




&lt;h3&gt;
  
  
  4. Set Up a Loop
&lt;/h3&gt;

&lt;p&gt;Schedule your skill to run every 15–30 minutes. This is your agent's heartbeat. Most runs will be quiet. The ones that aren't are the ones that matter.&lt;/p&gt;




&lt;h3&gt;
  
  
  5. Generate Interventions, Not Just Data
&lt;/h3&gt;

&lt;p&gt;The difference between useful and useless proactive AI is this:&lt;/p&gt;

&lt;p&gt;Don't just log:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Task is overdue"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead, have your agent reason about it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"This task has been ignored for 3 days. You'll likely miss the deadline — prioritize it today."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;OpenClaw's LLM layer handles this naturally. You don't need to write the reasoning — you just need to give it the signal.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where This Goes Next
&lt;/h2&gt;

&lt;p&gt;This pattern isn't limited to developer workflows.&lt;/p&gt;

&lt;p&gt;Anywhere you have data + delay + risk, you can apply proactive AI with OpenClaw:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sales pipelines&lt;/strong&gt; → catching stalled deals before they die&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Support systems&lt;/strong&gt; → flagging SLA risks before they breach&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Personal productivity&lt;/strong&gt; → identifying burnout patterns before they compound&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpenClaw's open ecosystem — with ClawHub, community skills, and multi-platform channels — means you're not building from scratch. You're building on top of something already watching.&lt;/p&gt;




&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;We've spent years improving how AI responds.&lt;/p&gt;

&lt;p&gt;But maybe that's the wrong direction.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The next step isn't better answers.&lt;br&gt;
It's AI that knows when to act.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;OpenClaw gave me the infrastructure to stop asking &lt;em&gt;"what can my AI respond to?"&lt;/em&gt; and start asking &lt;em&gt;"what should my AI notice?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That question changes everything.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>openclawchallenge</category>
      <category>openclaw</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Let's build a Production-Grade Bloom Filter in Python</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Sat, 21 Mar 2026 19:04:40 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/lets-build-a-production-grade-bloom-filter-in-python-j34</link>
      <guid>https://dev.to/shogun_the_grt/lets-build-a-production-grade-bloom-filter-in-python-j34</guid>
      <description>&lt;p&gt;Ever wondered how databases can tell you "this username is definitely not taken" in milliseconds without scanning millions of records? Or how caching systems avoid expensive database lookups for keys that don't exist? The secret is a probabilistic data structure called a &lt;strong&gt;Bloom Filter&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Let's build one from scratch :- with production features like persistence, serialization, and monitoring.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's a Bloom Filter?
&lt;/h2&gt;

&lt;p&gt;A Bloom filter is a space-efficient probabilistic data structure that tells you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"Definitely not in the set"&lt;/strong&gt; (100% certain)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Probably in the set"&lt;/strong&gt; (with a configurable false positive rate)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's like a bouncer who sometimes lets the wrong person in but never turns away someone who should be there.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Trade-off
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Traditional Set&lt;/th&gt;
&lt;th&gt;Bloom Filter&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Space&lt;/td&gt;
&lt;td&gt;O(n) per element&lt;/td&gt;
&lt;td&gt;~2-10 bytes per element&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time&lt;/td&gt;
&lt;td&gt;O(1) average&lt;/td&gt;
&lt;td&gt;O(k) where k ~ 5-10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;False Positives&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Configurable (0.1% - 5%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deletions&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;td&gt;Not supported&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For 10 million items, a hash set might use 500MB+ of memory. A Bloom filter with 1% false positive rate? Just ~12MB.&lt;/p&gt;

&lt;h2&gt;
  
  
  How It Works
&lt;/h2&gt;

&lt;p&gt;Imagine a massive array of bits, all initially 0. When you add an element:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run it through &lt;code&gt;k&lt;/code&gt; different hash functions&lt;/li&gt;
&lt;li&gt;Set the bits at positions &lt;code&gt;h1(item)&lt;/code&gt;, &lt;code&gt;h2(item)&lt;/code&gt;, ..., &lt;code&gt;hk(item)&lt;/code&gt; to 1&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;To check if an item exists:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run it through the same &lt;code&gt;k&lt;/code&gt; hash functions&lt;/li&gt;
&lt;li&gt;If ANY of those bits are 0 → &lt;strong&gt;definitely not in set&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;If ALL bits are 1 → &lt;strong&gt;probably in set&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Adding "apple":
Hash1("apple") = 42 → set bit 42
Hash2("apple") = 157 → set bit 157
Hash3("apple") = 891 → set bit 891

Checking "banana":
Hash1("banana") = 42 → bit 42 is 1
Hash2("banana") = 203 → bit 203 is 0 → definitely not present!
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Math Behind It
&lt;/h2&gt;

&lt;p&gt;Given:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;n&lt;/code&gt; = expected number of elements&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;p&lt;/code&gt; = desired false positive rate&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We can calculate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Optimal bits&lt;/strong&gt;: &lt;code&gt;m = -n * ln(p) / (ln(2))²&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optimal hash functions&lt;/strong&gt;: &lt;code&gt;k = (m/n) * ln(2)&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For 10,000 elements at 1% false positives:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;m&lt;/code&gt; = 95,851 bits (~12KB)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;k&lt;/code&gt; = 7 hash functions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each query checks 7 bits. No matter if you have 10 or 10 million items.&lt;/p&gt;

&lt;h2&gt;
  
  
  Let's Build It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Storage Backend
&lt;/h3&gt;

&lt;p&gt;First, we need a place to store bits. Let's make it extensible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;abc&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ABC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;abstractmethod&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;StorageBackend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ABC&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nd"&gt;@abstractmethod&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;pass&lt;/span&gt;

    &lt;span class="nd"&gt;@abstractmethod&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;set_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;pass&lt;/span&gt;

    &lt;span class="nd"&gt;@abstractmethod&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;clear&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;pass&lt;/span&gt;

    &lt;span class="nd"&gt;@abstractmethod&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;pass&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;InMemoryStorage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;StorageBackend&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size_bits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bit_array&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;bytearray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size_bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;RLock&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;byte_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
            &lt;span class="n"&gt;bit_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;
            &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bit_array&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;byte_index&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;bit_index&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;set_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;byte_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
            &lt;span class="n"&gt;bit_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bit_array&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;byte_index&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;|=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;bit_index&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bit_array&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Text Analysis
&lt;/h3&gt;

&lt;p&gt;We need to normalize items before adding them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Set&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;mmh3&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;HashFunction&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;MURMUR3&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;murmur3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;SHA1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sha1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;false_positive_rate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fpr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;false_positive_rate&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_calculate_size&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;false_positive_rate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_hashes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_calculate_num_hashes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;InMemoryStorage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_calculate_size&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ceil&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_calculate_num_hashes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ceil&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_get_hash_values&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Generate k hash positions using double hashing&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
        &lt;span class="n"&gt;item_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="c1"&gt;# Get two independent hash values
&lt;/span&gt;        &lt;span class="n"&gt;hash1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mmh3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;hash64&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item_bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;seed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;signed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="n"&gt;hash2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mmh3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;hash64&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item_bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;seed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;signed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

        &lt;span class="n"&gt;positions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_hashes&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="c1"&gt;# Combine using double hashing formula
&lt;/span&gt;            &lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hash1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;hash2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;
            &lt;span class="n"&gt;positions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;positions&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Add an item to the filter&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_get_hash_values&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Check if item might be in the filter&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_get_hash_values&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;  &lt;span class="c1"&gt;# Definitely not present
&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;  &lt;span class="c1"&gt;# Probably present
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Adding Statistics
&lt;/h3&gt;

&lt;p&gt;Production systems need monitoring:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;BloomFilterStats&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;total_insertions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="n"&gt;total_queries&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="n"&gt;false_positives&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="n"&gt;memory_usage_bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="n"&gt;last_reset_time&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_false_positive_rate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_queries&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;false_positives&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_queries&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# ... previous code ...
&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fpr&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;enable_stats&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# ... existing init ...
&lt;/span&gt;        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilterStats&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;enable_stats&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_get_hash_values&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_insertions&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;memory_usage_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_get_hash_values&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
                &lt;span class="k"&gt;break&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_queries&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
            &lt;span class="c1"&gt;# Track false positives (would need external verification)
&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_stats&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_dict&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;estimate_current_capacity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Estimate how many items are currently stored&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
        &lt;span class="n"&gt;bits_set&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;bin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;1&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
        &lt;span class="n"&gt;ratio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bits_set&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ratio&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="c1"&gt;# n = -m * ln(1 - ratio) / k
&lt;/span&gt;        &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;ratio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_hashes&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  File Persistence
&lt;/h3&gt;

&lt;p&gt;Make it survive restarts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;FileStorage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;StorageBackend&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;filepath&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size_bits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;filepath&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;filepath&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;RLock&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_load_or_create&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_load_or_create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;filepath&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;filepath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bit_array&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;bytearray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bit_array&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;bytearray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size_bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_persist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_persist&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;filepath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;wb&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bit_array&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;set_bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;byte_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
            &lt;span class="n"&gt;bit_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bit_array&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;byte_index&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;|=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;bit_index&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_persist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# Save after each change
&lt;/span&gt;
&lt;span class="c1"&gt;# Usage
&lt;/span&gt;&lt;span class="n"&gt;bf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;false_positive_rate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;storage_backend&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;storage_path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bloom_filter.bin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Serialization
&lt;/h3&gt;

&lt;p&gt;Share filters across services:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SerializationFormat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;JSON&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;PICKLE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pickle&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;BASE64&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;base64&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# ... existing code ...
&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;serialize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;format&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SerializationFormat&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;SerializationFormat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;capacity&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;false_positive_rate&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fpr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;size&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;num_hashes&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_hashes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;bit_array&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;b64encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ascii&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nb"&gt;format&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;SerializationFormat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="nb"&gt;format&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;SerializationFormat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PICKLE&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;b64encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pickle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ascii&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="nd"&gt;@classmethod&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;deserialize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cls&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;serialized&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;format&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SerializationFormat&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;SerializationFormat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;BloomFilter&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nb"&gt;format&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;SerializationFormat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;serialized&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pickle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;b64decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;serialized&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

        &lt;span class="n"&gt;bf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cls&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;capacity&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;false_positive_rate&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;b64decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;bit_array&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;bf&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Set Operations
&lt;/h3&gt;

&lt;p&gt;Merge multiple filters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;merge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;BloomFilter&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Merge another filter into this one (OR operation)&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_hashes&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_hashes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Filters must have same parameters&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;self_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;other_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;merged&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;bytearray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self_bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;other_bytes&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;merged&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;intersection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;BloomFilter&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;BloomFilter&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Create new filter from intersection (AND operation)&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_hashes&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_hashes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Filters must have same parameters&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;new_bf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fpr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;self_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;other_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_bytes&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;intersected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;bytearray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self_bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;other_bytes&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;new_bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;intersected&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;new_bf&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Complete Working Example
&lt;/h2&gt;

&lt;p&gt;Here's the full code in action:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Example 1: Basic Usage
&lt;/span&gt;&lt;span class="n"&gt;bf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;false_positive_rate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Add some items
&lt;/span&gt;&lt;span class="n"&gt;items&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apple&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;banana&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cherry&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Test membership
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apple in filter?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apple&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;      &lt;span class="c1"&gt;# True
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;grape in filter?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;grape&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;      &lt;span class="c1"&gt;# False
&lt;/span&gt;
&lt;span class="c1"&gt;# Example 2: With Statistics
&lt;/span&gt;&lt;span class="n"&gt;bf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;false_positive_rate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;enable_stats&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;stats&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_stats&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Insertions: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;total_insertions&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Memory: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;memory_usage_bytes&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; KB&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Estimated capacity: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;estimated_capacity&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Example 3: File Persistence
&lt;/span&gt;&lt;span class="n"&gt;bf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;false_positive_rate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;storage_backend&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;storage_path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;my_filter.bin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;persistent_data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# Filter survives restarts!
&lt;/span&gt;
&lt;span class="c1"&gt;# Example 4: Serialization
&lt;/span&gt;&lt;span class="n"&gt;serialized&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;serialize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SerializationFormat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;restored&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;deserialize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;serialized&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SerializationFormat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Gallery
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fr4wefq6luwtsoiejypxq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fr4wefq6luwtsoiejypxq.png" alt=" " width="800" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance Numbers
&lt;/h2&gt;

&lt;p&gt;Running on a modest laptop with 10 million items:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Operation&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;th&gt;Space&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Insert 1 item&lt;/td&gt;
&lt;td&gt;~0.5 µs&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Check 1 item&lt;/td&gt;
&lt;td&gt;~0.4 µs&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Build 10M items&lt;/td&gt;
&lt;td&gt;8 seconds&lt;/td&gt;
&lt;td&gt;12 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory per item&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;td&gt;1.2 bytes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The false positive rate? Exactly 1% as configured.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Applications
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. &lt;strong&gt;Cache Filtering&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Avoid cache lookups for keys that don't exist
&lt;/span&gt;&lt;span class="n"&gt;cache_filter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fpr&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;cache_filter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;  &lt;span class="c1"&gt;# Definitely not in cache
&lt;/span&gt;
    &lt;span class="c1"&gt;# Might be in cache - do the expensive lookup
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. &lt;strong&gt;Database Query Optimization&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Check if username exists without hitting the DB
&lt;/span&gt;&lt;span class="n"&gt;username_filter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fpr&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.001&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;username_available&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;username&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;username_filter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;username&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# Probably taken - check database
&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;user_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;username&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;  &lt;span class="c1"&gt;# Definitely available
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. &lt;strong&gt;Web Crawler Deduplication&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Remember visited URLs without storing them all
&lt;/span&gt;&lt;span class="n"&gt;visited&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1_000_000_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fpr&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.001&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;should_crawl&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;visited&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;  &lt;span class="c1"&gt;# Probably visited
&lt;/span&gt;    &lt;span class="n"&gt;visited&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. &lt;strong&gt;Distributed Systems&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Sync filters across services
&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;DistributedBloomFilter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fpr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;redis_client&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;redis&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;redis_client&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bloom_filter&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;filter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BloomFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fpr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="c1"&gt;# Periodically sync to Redis
&lt;/span&gt;        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;serialize&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# Try local first
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="c1"&gt;# Fall back to Redis if needed
&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_check_redis&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Limitations and Gotchas
&lt;/h2&gt;

&lt;h3&gt;
  
  
  No Deletions
&lt;/h3&gt;

&lt;p&gt;Bloom filters don't support deletion. Removing an item would require clearing bits that might belong to other items. For deletion support, check out &lt;strong&gt;Counting Bloom Filters&lt;/strong&gt; (store counters instead of bits).&lt;/p&gt;

&lt;h3&gt;
  
  
  False Positives Increase with Saturation
&lt;/h3&gt;

&lt;p&gt;As you add more items than capacity, the false positive rate increases:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csvs"&gt;&lt;code&gt;&lt;span class="nv"&gt;Capacity:&lt;/span&gt; &lt;span class="mf"&gt;1000&lt;/span&gt; &lt;span class="k"&gt;items&lt;/span&gt;
&lt;span class="k"&gt;Items&lt;/span&gt; &lt;span class="k"&gt;Added&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="k"&gt;Actual&lt;/span&gt; &lt;span class="k"&gt;FPR&lt;/span&gt;
&lt;span class="mf"&gt;1000&lt;/span&gt;        &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="mf"&gt;0.99&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;
&lt;span class="mf"&gt;2000&lt;/span&gt;        &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="mf"&gt;3.2&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;
&lt;span class="mf"&gt;5000&lt;/span&gt;        &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="mf"&gt;15.8&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Hash Function Quality Matters
&lt;/h3&gt;

&lt;p&gt;Use high-quality, independent hash functions. Our double hashing with Murmur3 works well, but avoid Python's built-in &lt;code&gt;hash()&lt;/code&gt; it's not designed for this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to Go From Here
&lt;/h2&gt;

&lt;p&gt;This is just the beginning. Real production systems add:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Counting Bloom Filters&lt;/strong&gt;: Support deletions with counters&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scalable Bloom Filters&lt;/strong&gt;: Grow dynamically as you add items&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cuckoo Filters&lt;/strong&gt;: Support deletions with better space efficiency&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Partitioned Filters&lt;/strong&gt;: Split across multiple machines&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bit-sliced Bloom Filters&lt;/strong&gt;: SIMD-optimized for speed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The inverted index may be the heart of search engines, but the Bloom filter is the secret sauce that makes them fast. It's the reason databases can say "no" in microseconds, and caching systems can avoid millions of useless lookups.&lt;/p&gt;

</description>
      <category>python</category>
      <category>tutorial</category>
      <category>algorithms</category>
      <category>datastructures</category>
    </item>
    <item>
      <title>Building TaskPilot: An AI Agent That Sees Your Screen and Takes Control</title>
      <dc:creator>Sarthak Rawat</dc:creator>
      <pubDate>Mon, 16 Mar 2026 21:07:44 +0000</pubDate>
      <link>https://dev.to/shogun_the_grt/building-taskpilot-an-ai-agent-that-sees-your-screen-and-takes-control-4i2e</link>
      <guid>https://dev.to/shogun_the_grt/building-taskpilot-an-ai-agent-that-sees-your-screen-and-takes-control-4i2e</guid>
      <description>&lt;p&gt;I created this blog for detailing about my project in the Gemini Live Agent Challenge hackathon. #GeminiLiveAgentChallenge&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem: Automation That Breaks the Moment the UI Changes
&lt;/h2&gt;

&lt;p&gt;Every developer has been there. You write a Selenium script, it works perfectly, and then the website updates its CSS class names and the whole thing falls apart. You set up an RPA workflow, it runs fine for a week, and then someone moves a button and it starts clicking the wrong thing.&lt;/p&gt;

&lt;p&gt;Traditional automation is brittle because it's blind. It relies on DOM selectors, API hooks, and hardcoded coordinates. It doesn't actually &lt;em&gt;see&lt;/em&gt; the screen. It just pokes at it.&lt;/p&gt;

&lt;p&gt;But humans don't automate that way. When you ask a colleague to "find the cheapest flight to New York and book it," they open a browser, look at the screen, read what's there, and make decisions based on what they see. They don't need an API. They don't need a DOM inspector. They just need eyes.&lt;/p&gt;

&lt;p&gt;That's the gap TaskPilot fills. It's an AI agent that observes your screen the way a human would, understands what it sees using Gemini's multimodal vision, and executes actions based on natural language intent. No selectors. No APIs. No brittle scripts. Just vision, reasoning, and action.&lt;/p&gt;

&lt;p&gt;Built for the &lt;strong&gt;Gemini Live Agent Challenge&lt;/strong&gt; hackathon. #GeminiLiveAgentChallenge&lt;/p&gt;




&lt;h2&gt;
  
  
  The Architecture: Two Agents, One Interface
&lt;/h2&gt;

&lt;p&gt;TaskPilot has two distinct execution environments that share a single Electron frontend:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Desktop Mode&lt;/strong&gt; — A TypeScript/Node.js agent (&lt;code&gt;clawd-cursor&lt;/code&gt;) that runs locally and controls your actual OS. It can open apps, type text, click buttons, switch windows, and execute multi-app workflows across your entire desktop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Browser Mode&lt;/strong&gt; — A Python WebSocket server (&lt;code&gt;computer-use-preview&lt;/code&gt;) deployed on Google Cloud Run that spins up a Playwright browser, runs a Gemini Computer Use vision loop, and streams screenshots and reasoning back to the frontend in real time.&lt;/p&gt;

&lt;p&gt;The Electron frontend connects to whichever mode the user selects. Desktop mode talks to a local REST API on &lt;code&gt;127.0.0.1:3847&lt;/code&gt;. Browser mode connects to a Cloud Run WebSocket endpoint. The UI is identical either way — live screenshots, a reasoning panel, an action timeline, and a voice input button.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5dq0p0y5wkdrfsu0zce5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5dq0p0y5wkdrfsu0zce5.png" alt=" " width="800" height="893"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The 5-Layer Pipeline: Why We Don't Always Need Vision
&lt;/h2&gt;

&lt;p&gt;The most important architectural decision in TaskPilot is the layered execution pipeline in &lt;code&gt;clawd-cursor&lt;/code&gt;. The core insight: taking a screenshot and sending it to a vision LLM is expensive and slow. Most tasks don't need it.&lt;/p&gt;

&lt;p&gt;So we built five layers, each cheaper and faster than the last. The agent tries them in order and only escalates when the current layer can't handle the task.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 0 — Browser (Playwright CDP)&lt;/strong&gt;&lt;br&gt;
For any task that involves a browser, we go straight to Chrome DevTools Protocol. No screenshots. No LLM. We read the DOM directly, find elements, and interact with them programmatically. This handles a huge chunk of web tasks instantly and for free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1 — Action Router (Regex + Keyboard Shortcuts)&lt;/strong&gt;&lt;br&gt;
A pattern-matching layer that recognizes common intents and maps them to direct actions. "Scroll down" becomes a keyboard shortcut. "Copy" becomes Ctrl+C. "Open Notepad" becomes a shell command. No LLM involved. This layer handles the majority of simple desktop tasks in under a second.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// From action-router.ts — pattern matching before any LLM call&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;SHORTCUT_PATTERNS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sr"&gt;/scroll&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+down/i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;keyboard&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;Key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;PageDown&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sr"&gt;/copy/i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;keyboard&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pressKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;Key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;LeftControl&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;Key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;C&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sr"&gt;/open&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;(\w&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;/i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;match&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;shell&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`start &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;match&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Layer 1.5 — Smart Interaction (1 LLM Call)&lt;/strong&gt;&lt;br&gt;
When pattern matching isn't enough, we make a single cheap text LLM call to plan the steps, then execute them via CDP or the accessibility tree. One call, no screenshots.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2 — A11y Reasoner (Accessibility Tree + Cheap LLM)&lt;/strong&gt;&lt;br&gt;
We read the OS accessibility tree — the structured representation of every UI element on screen — and feed it to a cheap text model. The model reasons about which element to interact with based on its label, role, and position. Still no screenshots.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3 — Computer Use (Vision LLM)&lt;/strong&gt;&lt;br&gt;
Only when all else fails do we take a screenshot and send it to Gemini or Anthropic Computer Use. This is the most powerful layer but also the most expensive. By the time we reach it, we've already handled 80%+ of tasks in the layers above.&lt;/p&gt;

&lt;p&gt;The performance difference is dramatic:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Without Pipeline&lt;/th&gt;
&lt;th&gt;With Pipeline&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Calculator (255×38)&lt;/td&gt;
&lt;td&gt;43s (18 LLM calls)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2.6s&lt;/strong&gt; (0 LLM calls)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Notepad (type hello)&lt;/td&gt;
&lt;td&gt;73s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.0s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;File Explorer&lt;/td&gt;
&lt;td&gt;53s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.9s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gmail compose&lt;/td&gt;
&lt;td&gt;162s&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;21.7s&lt;/strong&gt; (1 LLM call)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;h2&gt;
  
  
  The Browser Agent: Gemini Computer Use in a Loop
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwfhd2f1jyl3si1i2ixfo.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwfhd2f1jyl3si1i2ixfo.jpeg" alt=" " width="800" height="455"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Python &lt;code&gt;computer-use-preview&lt;/code&gt; service is where Gemini's multimodal capabilities really shine. It runs a tight vision loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Capture a screenshot of the Playwright browser&lt;/li&gt;
&lt;li&gt;Send it to Gemini along with the task and conversation history&lt;/li&gt;
&lt;li&gt;Gemini returns a function call (click, type, navigate, scroll, etc.)&lt;/li&gt;
&lt;li&gt;Execute the action in Playwright&lt;/li&gt;
&lt;li&gt;Capture the next screenshot and repeat
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# From agent.py — the core Gemini Computer Use loop
&lt;/span&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;agent_loop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;aio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_model_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;contents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_contents&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;GenerateContentConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="n"&gt;system_instruction&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;SYSTEM_PROMPT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="c1"&gt;# Extract function calls from response
&lt;/span&gt;        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;function_call&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_execute_action&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;function_call&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_contents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;  &lt;span class="c1"&gt;# append result to history
&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;finish_reason&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;FinishReason&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;STOP&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;final_reasoning&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Every screenshot is streamed back to the Electron frontend over WebSocket as a base64-encoded JPEG. The user watches the agent think and act in real time — they see the same screen the agent sees, with a cursor indicator showing exactly where it's about to click.&lt;/p&gt;

&lt;p&gt;We also implemented screenshot pruning: only the last 3 screenshots are kept in the Gemini context window. Older ones are replaced with text summaries. This keeps token costs manageable for long-running tasks without losing context.&lt;/p&gt;


&lt;h2&gt;
  
  
  The WebSocket Server: Bridging Frontend to Agent
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9li1jxkzkh0j3mjypmqp.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9li1jxkzkh0j3mjypmqp.jpeg" alt=" " width="800" height="452"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Python server (&lt;code&gt;server.py&lt;/code&gt;) is the glue between the Electron frontend and the browser agent. It manages session lifecycle, handles both browser and desktop modes, and routes voice input.&lt;/p&gt;

&lt;p&gt;Each connection gets an &lt;code&gt;AgentSession&lt;/code&gt; with a dedicated worker thread. Commands from the frontend go into a queue; the worker processes them sequentially. This keeps the WebSocket handler non-blocking while the agent does its work.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AgentSession&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_ws&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ws&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_loop&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_cmd_queue&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Queue&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_worker_thread&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_worker_loop&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daemon&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_worker_thread&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_worker_loop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_closed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;cmd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_cmd_queue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_run_agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mode&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For desktop mode, the server proxies the task to the local &lt;code&gt;clawd-cursor&lt;/code&gt; REST API via &lt;code&gt;ClawdBridge&lt;/code&gt;. The frontend doesn't need to know which backend is handling the task — it just sends a message and receives a stream of screenshots and reasoning updates.&lt;/p&gt;




&lt;h2&gt;
  
  
  Voice Input: Talking to Your Agent
&lt;/h2&gt;

&lt;p&gt;One of the more satisfying features to build was voice input. The user clicks the microphone button in the Electron UI, speaks their task, and the audio is sent to the Python server for transcription via Google Cloud Speech-to-Text.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# From voice_input.py
&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;VoiceTranscriber&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;audio_bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;speech&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;SpeechClient&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;speech&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;RecognitionAudio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;audio_bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;speech&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;RecognitionConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;speech&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RecognitionConfig&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AudioEncoding&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;WEBM_OPUS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;sample_rate_hertz&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;48000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;language_code&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en-US&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recognize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;alternatives&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;transcript&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The transcribed text drops straight into the task input field. It's a small thing but it makes the agent feel much more natural to use — especially for longer, more complex instructions where typing is tedious.&lt;/p&gt;




&lt;h2&gt;
  
  
  Safety: The Agent Needs to Know When to Ask
&lt;/h2&gt;

&lt;p&gt;Giving an AI agent control of your desktop is powerful. It's also potentially dangerous. We built a safety layer that classifies every action before executing it.&lt;/p&gt;

&lt;p&gt;Actions fall into three tiers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Auto&lt;/strong&gt; — Navigation, scrolling, reading. Execute immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Preview&lt;/strong&gt; — Typing text, opening files. Show the user what's about to happen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confirm&lt;/strong&gt; — Sending emails, deleting files, form submissions. Pause and require explicit approval.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The web dashboard (served by the Express server in &lt;code&gt;clawd-cursor&lt;/code&gt;) shows pending confirmations in real time. The user can approve or reject any action before it executes. There's also a kill switch that immediately halts the agent.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// From safety.ts — tier classification&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SafetyLayer&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;InputAction&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;SafetyTier&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;BLOCKED_PATTERNS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;action&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;description&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;SafetyTier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Block&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;CONFIRM_PATTERNS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;action&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;description&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;SafetyTier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Confirm&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;PREVIEW_PATTERNS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;action&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;description&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;SafetyTier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Preview&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;SafetyTier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Auto&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Provider-Agnostic by Design
&lt;/h2&gt;

&lt;p&gt;We didn't want to lock TaskPilot into a single AI provider. The &lt;code&gt;doctor&lt;/code&gt; command runs an interactive setup wizard that scans your environment, detects available providers, tests each one, and recommends the optimal pipeline configuration.&lt;/p&gt;

&lt;p&gt;The provider is auto-detected from the API key format:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;sk-ant-*&lt;/code&gt; → Anthropic&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;sk-*&lt;/code&gt; → OpenAI&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;AIza*&lt;/code&gt; → Gemini&lt;/li&gt;
&lt;li&gt;Local endpoint → Ollama&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can also mix providers: use Ollama for cheap text tasks (free, runs locally) and Gemini or Anthropic for vision tasks (best quality). The pipeline config is saved to &lt;code&gt;.clawd-config.json&lt;/code&gt; and loaded on startup.&lt;/p&gt;

&lt;p&gt;For the hackathon submission, Gemini is the primary provider — &lt;code&gt;@google/genai&lt;/code&gt; for the TypeScript desktop agent and &lt;code&gt;google-genai&lt;/code&gt; for the Python browser agent. Vertex AI mode is supported for cloud deployments.&lt;/p&gt;




&lt;h2&gt;
  
  
  Challenges We Faced
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The "blind clicking" problem.&lt;/strong&gt; Early versions of the browser agent would sometimes click coordinates that looked right in the screenshot but were slightly off in the actual browser due to scaling and DPI differences. We fixed this by normalizing coordinates relative to the Playwright viewport size and adding a cursor indicator in the frontend so users can see exactly where the agent is clicking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context window management.&lt;/strong&gt; Long browser sessions accumulate a lot of screenshots. Sending all of them to Gemini on every iteration would be prohibitively expensive and slow. The screenshot pruning strategy — keeping only the last 3 screenshots and summarizing older ones as text — was the right balance between context retention and cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thread safety in the Python server.&lt;/strong&gt; The WebSocket handler runs in an async event loop, but the Playwright browser and Gemini client calls are blocking. Getting the threading model right — async WebSocket handler, queue-based worker thread, &lt;code&gt;run_coroutine_threadsafe&lt;/code&gt; for sending messages back — took several iterations to get stable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The 5-layer pipeline ordering.&lt;/strong&gt; Deciding which layer handles which task isn't always obvious. We went through many iterations of the routing logic before settling on the current approach: browser CDP first, then pattern matching, then accessibility tree, then vision. The key insight was that the accessibility tree is almost always faster and cheaper than a screenshot, and it's surprisingly capable for most UI tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cross-platform desktop control.&lt;/strong&gt; Windows uses PowerShell for accessibility queries. macOS uses JXA (JavaScript for Automation) and System Events. Linux has neither. We ended up with platform-specific script directories (&lt;code&gt;scripts/mac/&lt;/code&gt;) and a runtime check that routes to the right implementation. Linux falls back to browser-only mode.&lt;/p&gt;




&lt;h2&gt;
  
  
  What We Learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Vision is the fallback, not the foundation.&lt;/strong&gt; The instinct when building a visual agent is to route everything through the vision model. That's wrong. Vision is expensive and slow. Build the cheap layers first — pattern matching, accessibility trees, keyboard shortcuts — and use vision only when they fail. Your users will notice the difference.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The accessibility tree is underrated.&lt;/strong&gt; Most developers don't think about the OS accessibility tree as an automation primitive. But it's a structured, real-time representation of every UI element on screen, with labels, roles, and positions. For a huge range of tasks, it's more reliable than a screenshot and orders of magnitude cheaper.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Streaming makes agents feel alive.&lt;/strong&gt; Sending screenshots and reasoning updates to the frontend in real time — rather than waiting for the task to complete — fundamentally changes how the agent feels to use. Users can see the agent thinking. They can intervene if it's going wrong. It transforms a black box into a collaborator.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Safety gates build trust.&lt;/strong&gt; The confirmation flow for risky actions isn't just a safety feature — it's a trust-building mechanism. When users see the agent pause and ask "I'm about to send this email, confirm?" they feel in control. That feeling of control is what makes people comfortable giving an AI agent access to their desktop.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;TaskPilot is one of three projects our team built for the Gemini Live Agent Challenge. We also built:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A real-time visual companion (Live Agents category) — an agent you can talk to naturally that sees through your camera, grounded in Cloud Vision, Document AI, and Natural Language API&lt;/li&gt;
&lt;li&gt;A 3D interactive history explorer (Creative Storyteller category) — a globe you can explore, clicking locations to get rich interleaved historical narratives with Imagen-generated imagery and Cloud TTS narration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each project targets a different category, but they all share the same philosophy: use the right tool for the right job, ground AI reasoning in structured data, and build experiences that feel genuinely useful rather than impressive demos.&lt;/p&gt;




&lt;h2&gt;
  
  
  Try It Yourself
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/tarinagarwal/task-pilot" rel="noopener noreferrer"&gt;https://github.com/tarinagarwal/task-pilot&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Demo Video:&lt;/strong&gt; &lt;a href="https://vimeo.com/1174159668?share=copy&amp;amp;fl=sv&amp;amp;fe=ci" rel="noopener noreferrer"&gt;https://vimeo.com/1174159668?share=copy&amp;amp;fl=sv&amp;amp;fe=ci&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Built with Gemini (&lt;code&gt;@google/genai&lt;/code&gt;, &lt;code&gt;google-genai&lt;/code&gt;), Playwright, Google Cloud Run, Artifact Registry, Cloud Speech-to-Text, Cloud Firestore, Terraform, Electron, TypeScript, Python, and Express.&lt;/p&gt;

&lt;p&gt;Created for the Gemini Live Agent Challenge. #GeminiLiveAgentChallenge&lt;/p&gt;

</description>
      <category>geminiliveagentchallenge</category>
      <category>electron</category>
      <category>productivity</category>
      <category>openclaw</category>
    </item>
  </channel>
</rss>
