<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: lukman lukman</title>
    <description>The latest articles on DEV Community by lukman lukman (@lukman-ss).</description>
    <link>https://dev.to/lukman-ss</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3474743%2Fcbcf077b-0089-4c5f-8236-7f53d2acd53c.jpg</url>
      <title>DEV Community: lukman lukman</title>
      <link>https://dev.to/lukman-ss</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lukman-ss"/>
    <language>en</language>
    <item>
      <title>Database Isolation: Why Two Committed Transfers Can Create Money</title>
      <dc:creator>lukman lukman</dc:creator>
      <pubDate>Fri, 02 Oct 2026 07:16:37 +0000</pubDate>
      <link>https://dev.to/lukman-ss/database-isolation-why-two-committed-transfers-can-create-money-49f9</link>
      <guid>https://dev.to/lukman-ss/database-isolation-why-two-committed-transfers-can-create-money-49f9</guid>
      <description>&lt;p&gt;Two wallet transfers can both commit successfully while creating money. In Lab 08, the problem is a specific query pattern: read a balance, validate it in Go, calculate a replacement value, and write that value later. A transaction boundary does not make the original read remain current.&lt;/p&gt;

&lt;p&gt;The useful question is whether overlapping transactions preserve the business invariant. This lab approaches that question through a three-account wallet, PostgreSQL isolation experiments, and concurrent tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with what must remain true
&lt;/h2&gt;

&lt;p&gt;Alice, Bob, and Charlie each begin with 1,000,000. A transfer should change the distribution of those balances while keeping their sum at 3,000,000. Every balance must also remain non-negative.&lt;/p&gt;

&lt;p&gt;The schema implements the row-level invariant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;CHECK&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That constraint does not enforce conservation of money across accounts. A broken transfer can leave every individual balance positive while increasing their sum. Correctness therefore needs both checks.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;TransferNaive&lt;/code&gt; uses &lt;code&gt;sql.LevelReadCommitted&lt;/code&gt;. It begins a transaction, reads the sender, checks sufficient funds, reads the receiver, updates both accounts, inserts an audit record, and commits. These operations share an atomic boundary. Their decisions can still depend on stale application values.&lt;/p&gt;

&lt;h2&gt;
  
  
  Follow the overlap instead of the happy path
&lt;/h2&gt;

&lt;p&gt;Transfer A sends 800,000 from Alice to Bob. Transfer B sends 800,000 from Alice to Charlie. Both transactions read Alice at 1,000,000 before either writes. Each concludes that the transfer is allowed.&lt;/p&gt;

&lt;p&gt;Each then calculates Alice's replacement balance as 200,000. The naive sender update uses that calculated value rather than deriving the new balance from the current row:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="c"&gt;// 4. Overwrite sender balance with calculated value (Lost Update vulnerability)&lt;/span&gt;
    &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ExecContext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"UPDATE isolation_accounts SET balance = $1 WHERE id = $2"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fromBalance&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fromID&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Errorf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"update from balance: %w"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Excerpt from &lt;code&gt;transfer.go&lt;/code&gt;; surrounding validation, receiver update, audit, and commit are omitted.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The second write can wait for the first writer and still overwrite Alice with the same 200,000. Bob and Charlie receive separate credits of 800,000. The expected final state in &lt;code&gt;TestNaiveTransfer_LostUpdate&lt;/code&gt; is Alice 200,000, Bob 1,800,000, Charlie 1,800,000: total 3,800,000.&lt;/p&gt;

&lt;p&gt;The test orchestrates this overlap with channels. Both transactions signal that their reads have completed; the coordinator then releases their writes. It reproduces the read-calculate-write pattern explicitly, rather than calling &lt;code&gt;TransferNaive&lt;/code&gt; directly. That distinction matters when describing what the test covers.&lt;/p&gt;

&lt;p&gt;These numbers are assertions in the repository, not measurements from a test run for this article. No lab tests were executed during preparation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why the constraint does not reject it
&lt;/h3&gt;

&lt;p&gt;Alice still has 200,000. There is no negative balance to reject. The missing debit is a lost update, and the violated invariant is the cross-account sum. A successful commit and a satisfied CHECK constraint are insufficient evidence that the transfer was correct.&lt;/p&gt;

&lt;p&gt;This does not mean every READ COMMITTED update loses data. The vulnerable implementation reads, calculates in application memory, and later writes an absolute value. Its query pattern is part of the explanation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Separate the SQL standard from PostgreSQL behavior
&lt;/h2&gt;

&lt;p&gt;The README describes the standard anomaly ladder: READ UNCOMMITTED permits dirty reads; READ COMMITTED prevents dirty reads but permits non-repeatable and phantom reads; standard REPEATABLE READ prevents dirty and non-repeatable reads but may permit phantoms; SERIALIZABLE prevents those anomalies.&lt;/p&gt;

&lt;p&gt;PostgreSQL provides stronger behavior in two important places. READ UNCOMMITTED behaves as READ COMMITTED and does not permit dirty reads. PostgreSQL REPEATABLE READ uses a stable snapshot and prevents the classic phantom read demonstrated by the lab.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PostgreSQL level&lt;/th&gt;
&lt;th&gt;Visibility in this lab&lt;/th&gt;
&lt;th&gt;Remaining concern&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;READ UNCOMMITTED&lt;/td&gt;
&lt;td&gt;Same behavior as READ COMMITTED&lt;/td&gt;
&lt;td&gt;Does not grant a stable transaction snapshot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;READ COMMITTED&lt;/td&gt;
&lt;td&gt;New snapshot per statement&lt;/td&gt;
&lt;td&gt;Repeated reads may change; naive writes may lose updates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;REPEATABLE READ&lt;/td&gt;
&lt;td&gt;Snapshot acquired at the first read&lt;/td&gt;
&lt;td&gt;Concurrent same-row update can fail with 40001&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SERIALIZABLE&lt;/td&gt;
&lt;td&gt;SSI with serial execution equivalence&lt;/td&gt;
&lt;td&gt;Conflicts may abort a transaction; retry must be handled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The table describes the engine used by the repository. It should not be treated as a portable implementation table for other databases.&lt;/p&gt;

&lt;h2&gt;
  
  
  Read Committed: two statements, two views
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;TestReadCommitted_NonRepeatableRead&lt;/code&gt; coordinates a reader and a writer. The reader first sees Alice at 1,000,000. The writer adds 500,000 and commits. The reader's second SELECT, still inside the same transaction, is expected to see 1,500,000.&lt;/p&gt;

&lt;p&gt;The change is committed, so this is not a dirty read. It is a non-repeatable read caused by statement snapshots.&lt;/p&gt;

&lt;p&gt;The phantom experiment uses a predicate instead of a single account. Two invoices initially have status PAID. A writer inserts a third PAID invoice and commits between the reader's queries:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;isolation_invoices&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'PAID'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under READ COMMITTED, the test expects counts of two and three. Under PostgreSQL REPEATABLE READ, it expects two and two. The predicate is unchanged; the visible set of matching rows differs only in the first strategy.&lt;/p&gt;

&lt;h2&gt;
  
  
  A stable snapshot is not exclusive access
&lt;/h2&gt;

&lt;p&gt;PostgreSQL MVCC lets ordinary readers observe versions without taking a FOR UPDATE row lock. &lt;code&gt;TestRepeatableRead_SnapshotIsolation&lt;/code&gt; expects both reads to return the original 1,000,000 even after another transaction commits an increase.&lt;/p&gt;

&lt;p&gt;The concurrent-update experiment is different. Two REPEATABLE READ transactions read the same account, then try to deduct 800,000. Its assertions require exactly one success and one serialization failure, SQLSTATE 40001.&lt;/p&gt;

&lt;p&gt;The application must therefore handle an aborted transaction. Choosing REPEATABLE READ does not make the second transfer silently succeed correctly. It changes the failure behavior.&lt;/p&gt;

&lt;p&gt;The README also discusses write skew as a possible serialization anomaly under snapshot isolation. The supplied experiments demonstrate same-row update conflicts; they do not reproduce a separate write-skew scenario. Keep the conceptual warning separate from the experiment's actual evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Row locking moves validation behind coordination
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;TransferWithLock&lt;/code&gt; stays at READ COMMITTED but obtains both account locks before validating the sender. Here is the locking excerpt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;firstID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;secondID&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;DeterministicLockOrder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fromID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;toID&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c"&gt;// Lock both accounts deterministically&lt;/span&gt;
    &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="n"&gt;b1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b2&lt;/span&gt; &lt;span class="kt"&gt;int64&lt;/span&gt;
    &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;QueryRowContext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"SELECT balance FROM isolation_accounts WHERE id = $1 FOR UPDATE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;firstID&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Scan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;b1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Errorf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"lock account %d: %w"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;firstID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;QueryRowContext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"SELECT balance FROM isolation_accounts WHERE id = $1 FOR UPDATE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;secondID&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Scan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;b2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Errorf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"lock account %d: %w"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;secondID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Excerpt from &lt;code&gt;transfer.go&lt;/code&gt;; the rest of the transfer is omitted.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A competing FOR UPDATE or UPDATE on those rows waits until the holder ends its transaction. An ordinary SELECT continues to use MVCC. The lock is coordination for the write path, not a promise that every reader is blocked.&lt;/p&gt;

&lt;p&gt;After locking, the implementation identifies which locked balance belongs to the sender and checks sufficient funds. Debit and credit use arithmetic SQL updates, and the audit record is written before commit. The second competing transfer validates after obtaining access to the row, rather than trusting an earlier unlocked balance.&lt;/p&gt;

&lt;h3&gt;
  
  
  Acquire locks in a shared order
&lt;/h3&gt;

&lt;p&gt;If A-to-B locks A first while B-to-A locks B first, the two transactions can create circular wait. The lab uses a helper independent of transfer direction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;DeterministicLockOrder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id2&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;firstID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;secondID&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;id1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;id2&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;id1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id2&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;id2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id1&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both directions acquire the smaller account ID, then the larger. The bidirectional test runs two workers with 50 transfers each. It checks errors and total balance. This order addresses the two-account lock pattern used here; it does not guarantee that arbitrary additional locking elsewhere can never deadlock.&lt;/p&gt;

&lt;p&gt;Lock waits are a real cost. The README's phrase about deterministic latency should not be read as a guarantee from this implementation. Contention and transaction duration still affect waiting, and the repository provides no latency benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Serializable means serial equivalence, not a global queue
&lt;/h2&gt;

&lt;p&gt;PostgreSQL SERIALIZABLE uses Serializable Snapshot Isolation. Concurrent transactions can run together; committed results must be equivalent to an allowed serial execution. Conflicts can require an abort.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;TestSerializable_ConcurrentUpdate_SerializationFailure&lt;/code&gt; coordinates two same-row updates and expects one success plus one 40001. It is a conflict experiment, not a throughput comparison or proof of every possible invariant.&lt;/p&gt;

&lt;p&gt;The recovery path needs a new transaction with fresh reads and validation. Retrying only the failed statement would not repeat the decision that caused it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retry the whole operation, with a limit
&lt;/h2&gt;

&lt;p&gt;The wrapper passes a complete transfer as the operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;PostgresWalletRepo&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;TransferSerializableWithRetryPolicy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fromID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;toID&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="kt"&gt;int64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;maxAttempts&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;policy&lt;/span&gt; &lt;span class="n"&gt;RetryPolicy&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;repo&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;RetryTransaction&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;maxAttempts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;policy&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;repo&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TransferSerializable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fromID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;toID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Complete wrapper from &lt;code&gt;transfer.go&lt;/code&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Each attempt calls &lt;code&gt;TransferSerializable&lt;/code&gt; again, including BeginTx, read, validation, debit, credit, audit, and commit. Failed attempts defer rollback. The retry helper accepts only serialization failure 40001 or deadlock 40P01, through the PostgreSQL error or corresponding sentinel.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;IsRetryableTxError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="kt"&gt;bool&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="no"&gt;false&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Is&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ErrSerializationFailure&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Is&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ErrDeadlockDetected&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="no"&gt;true&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="n"&gt;pqErr&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;pq&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Error&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;As&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;pqErr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;pqErr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s"&gt;"40001"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;pqErr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s"&gt;"40P01"&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="no"&gt;false&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Complete filter from &lt;code&gt;transfer.go&lt;/code&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;WrapTxError&lt;/code&gt; preserves both sentinel matching with &lt;code&gt;errors.Is&lt;/code&gt; and access to &lt;code&gt;*pq.Error&lt;/code&gt; with &lt;code&gt;errors.As&lt;/code&gt;. Constraint violations such as 23505 and 23503, generic errors, and insufficient funds are not automatically retried.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;maxAttempts&lt;/code&gt; includes the initial attempt. Non-positive values are rejected. The default base delay is 10 milliseconds and the maximum is one second. Before another attempt, the helper computes an exponential upper bound, caps it, and multiplies it by a random value for full jitter.&lt;/p&gt;

&lt;p&gt;Context cancellation is checked before attempts and can interrupt the default sleep. When retryable errors exhaust the limit, &lt;code&gt;ErrMaxRetryExceeded&lt;/code&gt; wraps the last error. &lt;code&gt;retry_test.go&lt;/code&gt; covers filtering, error preservation, success after a conflict, attempt exhaustion, cancellation, invalid limits, and delay behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the invariant under contention
&lt;/h2&gt;

&lt;p&gt;The stress fixture differs from the introductory wallet: three accounts begin at 50,000 each, giving total 150,000. One hundred goroutines attempt transfers of 1,000 using &lt;code&gt;TransferWithLock&lt;/code&gt; after a shared start gate is released.&lt;/p&gt;

&lt;p&gt;The test checks transfer errors, non-negative final balances, and the conserved sum of 150,000. A start gate creates competing work; it does not force every goroutine to reach every query at an identical moment. The naive test uses a more specific barrier after reads to reproduce its anomaly.&lt;/p&gt;

&lt;p&gt;To run the supplied experiments, start the repository's PostgreSQL infrastructure and execute:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;labs/08-database-isolation-level
go &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; ./...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database helper skips integration tests when PostgreSQL is unavailable unless &lt;code&gt;REQUIRE_POSTGRES=1&lt;/code&gt; makes absence a failure. A test command that reports success alongside skipped experiments is not evidence that database behavior was verified.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose from the invariant outward
&lt;/h2&gt;

&lt;p&gt;For this wallet's known rows, READ COMMITTED with FOR UPDATE and a shared lock order coordinates the critical read-check-write path. REPEATABLE READ serves a stable multi-query view but requires conflict handling for updates. SERIALIZABLE addresses serial equivalence at the cost of aborts, retry work, and dependency tracking.&lt;/p&gt;

&lt;p&gt;The README mentions an atomic conditional update as an alternative, but the lab does not implement that strategy. It also offers no measured throughput or latency ranking. The practical conclusion is to specify the invariant, identify the relevant reads and writes, choose coordination or conflict detection, and verify the resulting state under overlap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Repository and source map
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/08-database-isolation-level" rel="noopener noreferrer"&gt;Lab 08&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/README.md" rel="noopener noreferrer"&gt;README.md&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/schema.sql" rel="noopener noreferrer"&gt;schema.sql&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/isolation.go" rel="noopener noreferrer"&gt;isolation.go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/transfer.go" rel="noopener noreferrer"&gt;transfer.go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/read_committed_test.go" rel="noopener noreferrer"&gt;read_committed_test.go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/repeatable_read_test.go" rel="noopener noreferrer"&gt;repeatable_read_test.go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/phantom_read_test.go" rel="noopener noreferrer"&gt;phantom_read_test.go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/transfer_test.go" rel="noopener noreferrer"&gt;transfer_test.go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/serializable_test.go" rel="noopener noreferrer"&gt;serializable_test.go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/blob/main/labs/08-database-isolation-level/retry_test.go" rel="noopener noreferrer"&gt;retry_test.go&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>go</category>
      <category>postgres</category>
      <category>database</category>
      <category>backend</category>
    </item>
    <item>
      <title>Observability: Making an Invoice Request Explain Itself</title>
      <dc:creator>lukman lukman</dc:creator>
      <pubDate>Wed, 30 Sep 2026 23:55:00 +0000</pubDate>
      <link>https://dev.to/lukman-ss/observability-making-an-invoice-request-explain-itself-1o4i</link>
      <guid>https://dev.to/lukman-ss/observability-making-an-invoice-request-explain-itself-1o4i</guid>
      <description>&lt;p&gt;A successful response says the invoice workflow finished. It does not explain where the request spent its time. Lab 07 explores that gap through a Go service that loads an invoice, reserves inventory, calculates commission, generates a PDF, and sends a notification.&lt;/p&gt;

&lt;p&gt;The useful question is not just “Is the endpoint working?” It is “What happened inside this request, and which dependency explains its latency or failure?”&lt;/p&gt;

&lt;h2&gt;
  
  
  One workflow, two levels of visibility
&lt;/h2&gt;

&lt;p&gt;The unsafe service executes the same five dependency operations but exposes only a coarse completion duration or a generic failure log. Its HTTP response can still contain a request ID. What is missing is the service-level breakdown needed to investigate the request.&lt;/p&gt;

&lt;p&gt;The safe service adds structured logs, Prometheus metrics, and OpenTelemetry spans around the workflow. This changes the evidence available during investigation; it does not make the business operations themselves faster.&lt;/p&gt;

&lt;p&gt;The demo dependencies simulate work using configurable delays and errors. A delay waits on a timer while also listening for context cancellation. These are controlled diagnostic scenarios, not measurements of a real database, inventory service, or PDF engine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Logs, metrics, and traces answer different questions
&lt;/h2&gt;

&lt;p&gt;Logs describe discrete events. The safe service emits a structured event when a dependency completes, with its component, operation, duration, and outcome. On failure, the event also carries an error.&lt;/p&gt;

&lt;p&gt;Metrics summarize many requests. Request counters describe traffic, duration histograms describe latency, error counters describe failures, and an in-flight gauge describes active requests. The gauge is a concurrency signal; it is not a complete measurement of CPU capacity or resource saturation.&lt;/p&gt;

&lt;p&gt;Traces explain a particular execution. A server span contains an invoice-processing span, which contains one span for each dependency operation. That hierarchy connects the HTTP request to the work performed inside it.&lt;/p&gt;

&lt;p&gt;A latency histogram can reveal a shift across requests. A trace can identify the slow operation in one request. A structured log can provide the event details at that operation. None replaces the other two.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the slow step visible
&lt;/h2&gt;

&lt;p&gt;The demo's normal configuration assigns 20 ms to invoice loading, 15 ms to inventory, 10 ms to commission, 30 ms to PDF generation, and 20 ms to notification. The configured delays sum to approximately 95 ms before runtime and instrumentation overhead.&lt;/p&gt;

&lt;p&gt;In the slow-PDF scenario, PDF generation receives a 4,800 ms delay. Keeping the other fake dependencies unchanged gives a configured total of approximately 4,865 ms. This is scenario arithmetic, not a benchmark result. An HTTP notification client also changes the execution path when enabled.&lt;/p&gt;

&lt;p&gt;The point of the scenario is the contrast between “the request took several seconds” and “pdf.generate consumed most of the request.” A total-duration log provides the first statement. Dependency spans provide the second.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST /invoices/{id}/process
  invoice.process
    database.load_invoice
    inventory.reserve
    commission.calculate
    pdf.generate
    notification.send
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the hierarchy for the five-dependency workflow. A real HTTP notification call adds transport and downstream spans when those boundaries are instrumented.&lt;/p&gt;

&lt;h2&gt;
  
  
  Carry context through every dependency
&lt;/h2&gt;

&lt;p&gt;The safe service starts a span for each dependency and passes that span's context into the dependency call. The context carries the causal relationship and cancellation signal together.&lt;/p&gt;

&lt;p&gt;On a dependency failure, the implementation records the error on the dependency span and the invoice-processing span, marks them as failed, observes an error outcome, and returns a wrapped error identifying the component and operation. Processing stops at that failed step.&lt;/p&gt;

&lt;p&gt;The HTTP handler then maps a processing error to a response: a deadline becomes 504, cancellation becomes 499, and other processing failures become 502. The response uses the generic message “invoice processing failed” and includes the request ID. Detailed failure information stays in telemetry.&lt;/p&gt;

&lt;p&gt;These are mappings implemented by this lab, rather than a universal status-code policy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Correlate events without turning IDs into metric labels
&lt;/h2&gt;

&lt;p&gt;ContextLogger attaches request_id and, when an active span is valid, trace_id and span_id. The request ID helps connect the HTTP response to logs. The trace ID groups spans in one execution. The span ID identifies the active operation within that trace.&lt;/p&gt;

&lt;p&gt;The correlation fields are useful in logs, but they are deliberately absent from metric labels. Request metrics use method, a route template, and status class. Dependency metrics use component, operation, and outcome.&lt;/p&gt;

&lt;p&gt;The route label is &lt;code&gt;/invoices/{id}/process&lt;/code&gt;, rather than a separate label value for each invoice URL. The registry tests also check that request IDs and invoice IDs do not leak into label values.&lt;/p&gt;

&lt;p&gt;That distinction preserves aggregate metrics while leaving request-level investigation to logs and traces. The fixture tests check selected sensitive keys in emitted logs; they do not implement a general-purpose redaction system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Read the metric names as an investigation map
&lt;/h2&gt;

&lt;p&gt;The collector exposes six metric families:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;What it records&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;lab07_http_requests_total&lt;/td&gt;
&lt;td&gt;Requests by method, route, and status class&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lab07_http_request_duration_seconds&lt;/td&gt;
&lt;td&gt;Request duration observations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lab07_http_request_errors_total&lt;/td&gt;
&lt;td&gt;Processing errors at the HTTP boundary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lab07_http_in_flight_requests&lt;/td&gt;
&lt;td&gt;Currently active requests for the route&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lab07_dependency_duration_seconds&lt;/td&gt;
&lt;td&gt;Dependency duration by component, operation, and outcome&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lab07_dependency_errors_total&lt;/td&gt;
&lt;td&gt;Dependency failures&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Histograms store duration observations in seconds. Structured dependency logs report duration_ms. Keep those units explicit when comparing the two.&lt;/p&gt;

&lt;p&gt;The missing-ID branch records a 4xx request but does not increment the processing-error counter. The name “request errors” therefore needs to be read alongside the implementation's accounting rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cross an HTTP boundary without losing the trace
&lt;/h2&gt;

&lt;p&gt;The handler extracts incoming W3C trace context before starting its server span. The demo configures TraceContext and Baggage propagation globally.&lt;/p&gt;

&lt;p&gt;For outgoing notifications, HTTPNotificationClient creates a request with the current context and uses an OpenTelemetry HTTP transport. The transport propagates trace context, while the client explicitly copies X-Request-ID from the context into the downstream request.&lt;/p&gt;

&lt;p&gt;The default notification client has a five-second timeout. A supplied client is copied before its transport is wrapped, and an already instrumented transport is not wrapped again. Request cancellation still reaches the outgoing HTTP call through its context.&lt;/p&gt;

&lt;p&gt;The local &lt;code&gt;/notifications&lt;/code&gt; endpoint lets the demo exercise this HTTP boundary. It runs in the same demo process; this is not evidence of a deployed multi-service production system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Exercise specific failure modes
&lt;/h2&gt;

&lt;p&gt;The demo supports normal, slow-pdf, slow-database, slow-external, and pdf-error scenarios. They isolate different dependency delays or a PDF failure so the resulting telemetry can be inspected.&lt;/p&gt;

&lt;p&gt;For example, with the demo and its configured telemetry services running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST   &lt;span class="s1"&gt;'http://localhost:8087/invoices/INV-007/process?scenario=slow-pdf'&lt;/span&gt;   &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'X-Request-ID: lab07-slow-pdf'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The corresponding unsafe route is &lt;code&gt;/unsafe/invoices/INV-007/process&lt;/code&gt;. Comparing the routes makes the difference in diagnostic detail visible.&lt;/p&gt;

&lt;p&gt;Use a stable request ID to find the relevant logs, inspect the request's dependency spans, and compare request-level duration observations with dependency duration observations. The useful outcome is an explanation of the configured scenario, not merely confirmation that an endpoint returned a response.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat telemetry behavior as a contract
&lt;/h2&gt;

&lt;p&gt;The tests inspect span hierarchy, shared trace IDs, incoming trace-context extraction, outgoing propagation, request-ID preservation, structured log fields, and metric accounting.&lt;/p&gt;

&lt;p&gt;The fake-dependency hierarchy test expects seven ended spans: one HTTP server span, one invoice-processing span, and five dependency spans. The HTTP notification path has additional instrumentation and should not be forced into that seven-span expectation.&lt;/p&gt;

&lt;p&gt;The slow-PDF test uses a shorter, 10 ms delay and checks the recorded PDF span. Concurrent-request tests check mixed success and error responses, matching counters, and an in-flight gauge that returns to zero.&lt;/p&gt;

&lt;p&gt;There is also a test for response-encoding failures. Such failures are logged with correlation fields, but the implementation has already recorded success accounting before encoding the successful response. That limitation matters when interpreting metrics.&lt;/p&gt;

&lt;p&gt;These observations come from reading the source and its assertions. The Go tests were not executed in this content-generation environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model
&lt;/h2&gt;

&lt;p&gt;Observability is useful when the evidence explains a request: what happened, where its time went, and which operation failed. This lab makes that idea concrete by connecting dependency events, aggregate measurements, and a causal trace through one invoice workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A system should explain what happened, where time was spent, and why it was slow or failed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Repository: &lt;a href="https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/07-observability" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/07-observability&lt;/a&gt;&lt;/p&gt;

</description>
      <category>go</category>
      <category>observability</category>
      <category>backend</category>
      <category>opentelemetry</category>
    </item>
    <item>
      <title>API Versioning with Separate Invoice Contracts</title>
      <dc:creator>lukman lukman</dc:creator>
      <pubDate>Wed, 30 Sep 2026 07:23:01 +0000</pubDate>
      <link>https://dev.to/lukman-ss/api-versioning-with-separate-invoice-contracts-42cg</link>
      <guid>https://dev.to/lukman-ss/api-versioning-with-separate-invoice-contracts-42cg</guid>
      <description>&lt;p&gt;An invoice endpoint returns customer information as a string. A new representation replaces that string with an object containing the customer's ID, name, and phone number. The response remains valid JSON and the server still returns HTTP 200. The legacy consumer expects a string and cannot decode the new object.&lt;/p&gt;

&lt;p&gt;Lab 06 of Software Engineering Lab demonstrates this change using Go HTTP handlers and an explicit consumer model. It then keeps the old representation on V1 and introduces the nested representation on V2. A third example adds a field without changing the fields that the legacy model already understands.&lt;/p&gt;

&lt;p&gt;The tests described here are source-defined scenarios and assertions. They were inspected, not executed for this article. The legacy consumer is simulated by a Go struct; this is not an Android device test.&lt;/p&gt;

&lt;h2&gt;
  
  
  The original invoice contract
&lt;/h2&gt;

&lt;p&gt;The documented invoice response contains four fields:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1001&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"customer"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Budi"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;500000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"PAID"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;domain.go&lt;/code&gt; makes the consumer's expectation concrete:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;LegacyInvoice&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;ID&lt;/span&gt;       &lt;span class="kt"&gt;int&lt;/span&gt;    &lt;span class="s"&gt;`json:"id"`&lt;/span&gt;
    &lt;span class="n"&gt;Customer&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="s"&gt;`json:"customer"`&lt;/span&gt;
    &lt;span class="n"&gt;Total&lt;/span&gt;    &lt;span class="kt"&gt;int64&lt;/span&gt;  &lt;span class="s"&gt;`json:"total"`&lt;/span&gt;
    &lt;span class="n"&gt;Status&lt;/span&gt;   &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="s"&gt;`json:"status"`&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The helper &lt;code&gt;ParseLegacyInvoice&lt;/code&gt; calls &lt;code&gt;json.Unmarshal&lt;/code&gt; into that struct. This turns compatibility into an operation that the tests can exercise. The question is whether the existing model can still consume the body and obtain the expected values.&lt;/p&gt;

&lt;p&gt;The fixture uses invoice ID &lt;code&gt;1001&lt;/code&gt;, customer name &lt;code&gt;Budi&lt;/code&gt;, total &lt;code&gt;500000&lt;/code&gt;, and status &lt;code&gt;PAID&lt;/code&gt;. These are demonstration values, not a production invoice dataset.&lt;/p&gt;

&lt;h2&gt;
  
  
  The unsafe handler changes an existing field's type
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;UnsafeHandler&lt;/code&gt; returns the following customer representation while retaining the invoice's other fields:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1001&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"customer"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Budi"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"phone"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"08123"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;500000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"PAID"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The new information is useful, but it occupies a field whose published representation was a string. &lt;code&gt;LegacyInvoice.Customer&lt;/code&gt; cannot receive this object through the supplied decoder.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;TestBreakingChange_LegacyClientFails&lt;/code&gt; first requires HTTP 200, then calls &lt;code&gt;ParseLegacyInvoice&lt;/code&gt; on the response body. It expects decoding to fail. A passing test in this example would confirm the deliberate incompatibility; it would not certify the handler as backward compatible.&lt;/p&gt;

&lt;p&gt;The unsafe endpoint is synthetic. It creates its payload directly after validating the path, rather than retrieving the invoice through the repository used by V1 and V2. It demonstrates the shape change, not a complete persistence-backed invoice service.&lt;/p&gt;

&lt;h2&gt;
  
  
  One domain model, two response models
&lt;/h2&gt;

&lt;p&gt;The internal &lt;code&gt;Invoice&lt;/code&gt; model already contains a nested &lt;code&gt;Customer&lt;/code&gt;. The public V1 response does not have to copy that shape.&lt;/p&gt;

&lt;p&gt;In &lt;code&gt;safe_server.go&lt;/code&gt;, the response types establish separate contracts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;InvoiceV1Response&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;ID&lt;/span&gt;       &lt;span class="kt"&gt;int&lt;/span&gt;    &lt;span class="s"&gt;`json:"id"`&lt;/span&gt;
    &lt;span class="n"&gt;Customer&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="s"&gt;`json:"customer"`&lt;/span&gt;
    &lt;span class="n"&gt;Total&lt;/span&gt;    &lt;span class="kt"&gt;int64&lt;/span&gt;  &lt;span class="s"&gt;`json:"total"`&lt;/span&gt;
    &lt;span class="n"&gt;Status&lt;/span&gt;   &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="s"&gt;`json:"status"`&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;CustomerV2Response&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;ID&lt;/span&gt;    &lt;span class="kt"&gt;int&lt;/span&gt;    &lt;span class="s"&gt;`json:"id"`&lt;/span&gt;
    &lt;span class="n"&gt;Name&lt;/span&gt;  &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="s"&gt;`json:"name"`&lt;/span&gt;
    &lt;span class="n"&gt;Phone&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="s"&gt;`json:"phone"`&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;InvoiceV2Response&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;ID&lt;/span&gt;       &lt;span class="kt"&gt;int&lt;/span&gt;                &lt;span class="s"&gt;`json:"id"`&lt;/span&gt;
    &lt;span class="n"&gt;Customer&lt;/span&gt; &lt;span class="n"&gt;CustomerV2Response&lt;/span&gt; &lt;span class="s"&gt;`json:"customer"`&lt;/span&gt;
    &lt;span class="n"&gt;Total&lt;/span&gt;    &lt;span class="kt"&gt;int64&lt;/span&gt;              &lt;span class="s"&gt;`json:"total"`&lt;/span&gt;
    &lt;span class="n"&gt;Status&lt;/span&gt;   &lt;span class="kt"&gt;string&lt;/span&gt;             &lt;span class="s"&gt;`json:"status"`&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mapper for V1 selects only the customer name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;mapToV1&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;domain&lt;/span&gt; &lt;span class="n"&gt;Invoice&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;InvoiceV1Response&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;InvoiceV1Response&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;ID&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;       &lt;span class="n"&gt;domain&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Customer&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;domain&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Customer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Total&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;    &lt;span class="n"&gt;domain&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Total&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Status&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;   &lt;span class="n"&gt;domain&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;mapToV2&lt;/code&gt; maps the same domain invoice into the nested customer DTO, preserving ID, name, and phone. Both handlers use the same repository interface. The repository fixture supplies the same invoice to each mapper.&lt;/p&gt;

&lt;p&gt;This is the boundary the lab implements: the domain model holds the invoice information, while a version-specific mapper decides what the consumer receives. There is no database schema or database migration in this lab.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route each contract explicitly
&lt;/h2&gt;

&lt;p&gt;The supported endpoints are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET /api/v1/invoices/1001
GET /api/v2/invoices/1001
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;V1Handler&lt;/code&gt; maps its result through &lt;code&gt;mapToV1&lt;/code&gt;; &lt;code&gt;V2Handler&lt;/code&gt; maps through &lt;code&gt;mapToV2&lt;/code&gt;. Registering V2 does not replace V1's handler or DTO.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;TestVersionedRoutes_CanRunSideBySide&lt;/code&gt; registers both handlers on the same mux and requests both paths. It decodes each response into its version's type and checks the customer values.&lt;/p&gt;

&lt;p&gt;Another regression test registers V2 and then checks V1's wire contract again. That test targets a specific risk: a new route exists, but the old response must still retain the fields the legacy consumer understands.&lt;/p&gt;

&lt;p&gt;The lab uses URL versioning. Header-based selection appears in the README as a design alternative; there is no header-version negotiation implementation here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the consumer and inspect the response
&lt;/h2&gt;

&lt;p&gt;The safe tests use two complementary views of the body. One decodes V1 into &lt;code&gt;LegacyInvoice&lt;/code&gt; and checks its invoice fields. Another reads the response into a map of &lt;code&gt;json.RawMessage&lt;/code&gt; and checks field presence and values.&lt;/p&gt;

&lt;p&gt;For the central breaking change, the important assertion is explicit: V1's &lt;code&gt;customer&lt;/code&gt; is decoded into a string and compared with &lt;code&gt;Budi&lt;/code&gt;. V2's &lt;code&gt;customer&lt;/code&gt; is decoded into an object, then its &lt;code&gt;id&lt;/code&gt;, &lt;code&gt;name&lt;/code&gt;, and &lt;code&gt;phone&lt;/code&gt; are checked separately.&lt;/p&gt;

&lt;p&gt;Typed response tests also check the invoice ID, total, and status. This gives the tests a concrete fixture-based contract rather than merely requiring that the body be JSON.&lt;/p&gt;

&lt;p&gt;These assertions cover the fields and requests they exercise. They are not an exhaustive schema validator for all payloads, nullability combinations, or consumer implementations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Adding a field is a different change
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;AdditiveHandler&lt;/code&gt; keeps &lt;code&gt;customer&lt;/code&gt; as a string and adds currency:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1001&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"customer"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Budi"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;500000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"PAID"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"currency"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"IDR"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;TestAdditiveField_LegacyClientStillWorks&lt;/code&gt; verifies that the body contains &lt;code&gt;currency&lt;/code&gt; with value &lt;code&gt;IDR&lt;/code&gt;. It then decodes the response into the old &lt;code&gt;LegacyInvoice&lt;/code&gt;, which has no currency field, and checks the customer and total.&lt;/p&gt;

&lt;p&gt;The supplied decoder tolerates the extra field. The README qualifies the broader design rule: additive response fields are usually compatible when consumers tolerate unknown fields. Consumers using strict validation may behave differently.&lt;/p&gt;

&lt;p&gt;The demonstrated outcome belongs to this Go decoding path. The lab does not include a strict-schema consumer test, so it does not prove compatibility with every possible reader.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compatibility includes request requirements
&lt;/h2&gt;

&lt;p&gt;The source also protects an expectation outside the JSON body. &lt;code&gt;TestV1Contract_DoesNotIntroduceRequiredTenantHeader&lt;/code&gt; sends a V1 request without &lt;code&gt;X-Tenant-ID&lt;/code&gt; and expects HTTP 200.&lt;/p&gt;

&lt;p&gt;That test records that this V1 handler does not introduce the header as a requirement. It does not establish a broader authentication policy; the lab contains no tenant authentication middleware.&lt;/p&gt;

&lt;p&gt;The README lists other contract changes, such as renamed fields, required inputs, date formats, and changed error representations. The implemented string-to-object case is the main demonstration. The listed possibilities should not be mistaken for separately implemented experiments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Routing behavior is tested separately
&lt;/h2&gt;

&lt;p&gt;The helper &lt;code&gt;parseInvoiceID&lt;/code&gt; validates the prefix, requires an ID, rejects extra path segments, parses a numeric value, and requires it to be positive. The versioned handlers then look up that ID.&lt;/p&gt;

&lt;p&gt;For the versioned fixture, the expected cases include:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Request&lt;/th&gt;
&lt;th&gt;Expected response&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /api/v1/invoices/1001&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;200, invoice JSON&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /api/v1/invoices/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;400, missing ID&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /api/v1/invoices/abc&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;400, invalid numeric ID&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /api/v1/invoices/0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;400, nonpositive ID&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /api/v1/invoices/1001/extra&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;400, extra segment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /api/v1/invoices/9999&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;404, invoice not found&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;POST /api/v1/invoices/1001&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;405, &lt;code&gt;Allow: GET&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /api/v3/invoices/1001&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;404 from the router&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The fixture repository contains only invoice &lt;code&gt;1001&lt;/code&gt;. Thus &lt;code&gt;9999&lt;/code&gt; is a valid ID format but has no corresponding invoice in the fixture.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;newVersionedMux&lt;/code&gt; registers both the exact collection path and the subtree path for each version. The request without a trailing slash therefore reaches the handler and produces the lab's 400 response, rather than relying on an implicit redirect.&lt;/p&gt;

&lt;p&gt;Handler-generated responses use &lt;code&gt;application/json&lt;/code&gt;, including method errors. An unmatched V3 path is handled by the mux's default 404 behavior; the README documents that response as plain text. The common JSON writer does not wrap every router response.&lt;/p&gt;

&lt;h2&gt;
  
  
  Migration remains a documented policy
&lt;/h2&gt;

&lt;p&gt;The README describes releasing V2, retaining V1 for older consumers, monitoring adoption, communicating deprecation, and removing V1 after sunset criteria are met.&lt;/p&gt;

&lt;p&gt;It includes example timeframes and traffic thresholds. They are illustrative policy values, not measured adoption data or an implemented shutdown schedule. This directory contains no traffic-monitoring pipeline or automated sunset mechanism.&lt;/p&gt;

&lt;p&gt;The code demonstrates the prerequisite: both representations can be served through separate routes, with tests that continue checking V1. Deciding when V1 can be removed requires the consumer information described in the README.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run the lab from its module
&lt;/h2&gt;

&lt;p&gt;The directory contains its own &lt;code&gt;go.mod&lt;/code&gt;, declaring Go &lt;code&gt;1.25.0&lt;/code&gt;. Its README lists these test commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;labs/06-api-versioning
go &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; ./...
go &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="nt"&gt;-race&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; ./...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tests use &lt;code&gt;httptest&lt;/code&gt; and a mock repository rather than an external database. No test execution results are claimed here.&lt;/p&gt;

&lt;p&gt;The lab's mental model is precise: HTTP 200 describes a server response, while compatibility must be evaluated against the consumer's contract. Separate DTOs and regression tests keep that contract visible when a new representation is introduced.&lt;/p&gt;

&lt;h2&gt;
  
  
  Source
&lt;/h2&gt;

&lt;p&gt;Based solely on the README, Go handlers, models, and tests in Software Engineering Lab, Lab 06. Test outcomes are described from source assertions, not a new run.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/06-api-versioning" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/06-api-versioning&lt;/a&gt;&lt;/p&gt;

</description>
      <category>go</category>
      <category>api</category>
      <category>testing</category>
      <category>backend</category>
    </item>
    <item>
      <title>Race Condition on Concurrent Stock Updates</title>
      <dc:creator>lukman lukman</dc:creator>
      <pubDate>Tue, 29 Sep 2026 05:58:05 +0000</pubDate>
      <link>https://dev.to/lukman-ss/race-condition-on-concurrent-stock-updates-398h</link>
      <guid>https://dev.to/lukman-ss/race-condition-on-concurrent-stock-updates-398h</guid>
      <description>&lt;p&gt;A stock update can be implemented with a sequence that looks reasonable when only one request is running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read stock
check availability
decrease stock
save the new value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Lab 05 focuses on what happens when multiple operations execute that sequence against the same state.&lt;/p&gt;

&lt;p&gt;The important part of the experiment is not whether each request contains valid business logic in isolation. The test is whether the invariant still holds when those requests overlap.&lt;/p&gt;

&lt;p&gt;The lab covers this from several angles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;an unsafe read/check/write flow;&lt;/li&gt;
&lt;li&gt;deterministic lost-update tests;&lt;/li&gt;
&lt;li&gt;an atomic conditional SQL update;&lt;/li&gt;
&lt;li&gt;PostgreSQL row locking;&lt;/li&gt;
&lt;li&gt;concurrent inventory attempts;&lt;/li&gt;
&lt;li&gt;concurrent booking attempts protected by a database uniqueness constraint;&lt;/li&gt;
&lt;li&gt;a separate Go memory data-race example.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The main experiment is the application/database race.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Smallest Inventory Case
&lt;/h2&gt;

&lt;p&gt;Start with one unit of stock:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;stock = 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two operations attempt to consume it.&lt;/p&gt;

&lt;p&gt;The expected business result is straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;successful operations = 1
remaining stock       = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The unsafe implementation separates the decision into multiple steps.&lt;/p&gt;

&lt;p&gt;Conceptually, the execution is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;READ
  ↓
CHECK
  ↓
CALCULATE
  ↓
WRITE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With one caller, the sequence behaves as expected.&lt;/p&gt;

&lt;p&gt;With two callers, those steps can interleave.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two Operations Read the Same State
&lt;/h2&gt;

&lt;p&gt;Consider this execution:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request A                    Request B
---------                    ---------
READ stock = 1

                             READ stock = 1

CHECK stock &amp;gt; 0              CHECK stock &amp;gt; 0

newStock = 0                 newStock = 0

WRITE stock = 0

                             WRITE stock = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both operations observed the same original value.&lt;/p&gt;

&lt;p&gt;Both checks passed.&lt;/p&gt;

&lt;p&gt;Both calculated the same replacement value.&lt;/p&gt;

&lt;p&gt;The final database value is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;stock = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Looking only at the final stock value does not expose the full problem.&lt;/p&gt;

&lt;p&gt;Both operations may have been treated as successful even though only one unit existed.&lt;/p&gt;

&lt;p&gt;The broken invariant is therefore not simply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;stock must never be negative
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test also needs to account for the number of successful decrements.&lt;/p&gt;

&lt;p&gt;For one available item:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;successful decrements &amp;lt;= 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Lost Update
&lt;/h2&gt;

&lt;p&gt;The unsafe flow calculates a new value outside the final write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read stock
calculate stock - 1
write calculated value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That creates a window where another operation can read the same state.&lt;/p&gt;

&lt;p&gt;The important interleaving is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A reads 1
B reads 1

A calculates 0
B calculates 0

A writes 0
B writes 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second write is not based on the state produced by the first write.&lt;/p&gt;

&lt;p&gt;It is based on the earlier value both callers observed.&lt;/p&gt;

&lt;p&gt;That is the lost-update behavior exercised by the lab.&lt;/p&gt;

&lt;p&gt;The repository uses deterministic concurrency coordination for these tests rather than depending only on timing or arbitrary sleeps. Barriers/start gates are used to force the relevant operations to overlap at the point required by the experiment.&lt;/p&gt;

&lt;p&gt;That matters because a concurrency test that passes or fails only depending on scheduler luck is a weak proof.&lt;/p&gt;

&lt;h2&gt;
  
  
  Moving the Condition Into the Update
&lt;/h2&gt;

&lt;p&gt;One implementation in the lab removes the separate read/check/write sequence.&lt;/p&gt;

&lt;p&gt;The SQL statement is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;inventory_products&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="n"&gt;RETURNING&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There are two important details in this statement.&lt;/p&gt;

&lt;p&gt;First:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application does not read a value, calculate a replacement, then send that replacement back.&lt;/p&gt;

&lt;p&gt;Second:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The availability condition is part of the statement that mutates the row.&lt;/p&gt;

&lt;p&gt;The operation becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;attempt decrement
only when stock &amp;gt; 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read
decide in application
write later
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This removes the application-level race window demonstrated by the unsafe implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Invariant Is Checked at the Write Boundary
&lt;/h2&gt;

&lt;p&gt;The atomic version changes where the decision happens.&lt;/p&gt;

&lt;p&gt;Unsafe model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application:
    read stock
    check stock
    calculate new stock

Database:
    write supplied value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Atomic model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database:
    decrement stock
    only if stock &amp;gt; 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference is visible under contention.&lt;/p&gt;

&lt;p&gt;The repository includes a concurrent inventory scenario with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;initial stock = 100
concurrent attempts = 500
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The useful assertion is not merely that all concurrent calls return.&lt;/p&gt;

&lt;p&gt;The final state must still be consistent with the initial inventory.&lt;/p&gt;

&lt;p&gt;The number of successful decrements cannot exceed the amount that was available.&lt;/p&gt;

&lt;h2&gt;
  
  
  PostgreSQL Row Locking
&lt;/h2&gt;

&lt;p&gt;The lab also contains a PostgreSQL row-lock implementation and tests.&lt;/p&gt;

&lt;p&gt;The relevant read uses &lt;code&gt;FOR UPDATE&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;inventory_products&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="k"&gt;UPDATE&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The row lock changes the execution shape.&lt;/p&gt;

&lt;p&gt;Instead of two transactions independently reading the same stock and later trying to update it, the critical read/check/update flow is coordinated around the same row.&lt;/p&gt;

&lt;p&gt;A simplified execution is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Transaction A
-------------
BEGIN

SELECT ... FOR UPDATE

check stock

update stock

COMMIT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A concurrent transaction targeting the same locked row does not proceed through the same critical section independently while the first transaction still owns the lock.&lt;/p&gt;

&lt;p&gt;The repository demonstrates this as another way to protect the same business invariant.&lt;/p&gt;

&lt;p&gt;The lab does not use that result to claim that every concurrency problem should be solved with row locking. It is one implementation demonstrated against the stock-update problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Atomic Update and Row Lock Solve the Problem Differently
&lt;/h2&gt;

&lt;p&gt;Both implementations target the same unsafe boundary, but the shape is different.&lt;/p&gt;

&lt;p&gt;With the atomic update:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="n"&gt;RETURNING&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the condition and mutation fit into one statement.&lt;/p&gt;

&lt;p&gt;With row locking:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;lock row
↓
read state
↓
make decision
↓
update
↓
commit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the operation retains a multi-step decision while coordinating access to the row.&lt;/p&gt;

&lt;p&gt;The lab therefore gives two concrete implementations rather than reducing the topic to "put everything in a transaction."&lt;/p&gt;

&lt;p&gt;A transaction establishes a transaction boundary.&lt;/p&gt;

&lt;p&gt;The concurrency behavior still depends on how the shared row is accessed inside that boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Booking Uses a Different Invariant
&lt;/h2&gt;

&lt;p&gt;The repository also applies the concurrency problem to booking.&lt;/p&gt;

&lt;p&gt;The invariant is no longer a numeric stock value.&lt;/p&gt;

&lt;p&gt;A booking is identified by a combination including:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;branch_id
service_date
slot_time
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The schema protects that combination with a unique constraint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UNIQUE(branch_id, service_date, slot_time)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important concurrency case is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;many requests
↓
same branch
same date
same slot
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An application can check whether the slot exists before inserting.&lt;/p&gt;

&lt;p&gt;That check is useful for normal control flow, but it is not enough to enforce the invariant under concurrency.&lt;/p&gt;

&lt;p&gt;Two requests can both observe:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;slot is available
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;before either insert is visible to the other operation.&lt;/p&gt;

&lt;p&gt;The database constraint is the final storage boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Concurrent Booking Test
&lt;/h2&gt;

&lt;p&gt;The booking test exercises a single slot with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;500 concurrent attempts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only the valid database state should remain.&lt;/p&gt;

&lt;p&gt;PostgreSQL uniqueness conflicts are mapped through SQLSTATE:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;23505
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The conflict is treated as an expected concurrency outcome rather than as proof that the database is broken.&lt;/p&gt;

&lt;p&gt;The database is enforcing the invariant that the application needs.&lt;/p&gt;

&lt;p&gt;This is the same design principle as the inventory test, applied to a different state shape.&lt;/p&gt;

&lt;p&gt;Inventory invariant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;successful decrements &amp;lt;= available stock
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Booking invariant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one row for one unique branch/date/slot combination
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mechanism is chosen around the invariant being protected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Application Pre-Checks Do Not Replace Constraints
&lt;/h2&gt;

&lt;p&gt;Consider the booking sequence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SELECT booking
WHERE branch/date/slot = requested slot

if not found:
    INSERT booking
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under concurrent execution:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request A                       Request B
---------                       ---------
SELECT → not found

                                SELECT → not found

INSERT

                                INSERT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without storage enforcement, both requests reached the same application decision.&lt;/p&gt;

&lt;p&gt;The unique constraint prevents the database from accepting both final rows for the same slot.&lt;/p&gt;

&lt;p&gt;This is another example of why code that is correct sequentially can still be incomplete as concurrent business logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory Data Race vs Business Race
&lt;/h2&gt;

&lt;p&gt;Lab 05 also keeps a Go memory data-race example.&lt;/p&gt;

&lt;p&gt;That experiment is useful, but it is not the same failure as the inventory and booking cases.&lt;/p&gt;

&lt;p&gt;A memory data race is concerned with unsynchronized memory access between goroutines.&lt;/p&gt;

&lt;p&gt;The inventory and booking experiments focus on business invariants over shared application/database state.&lt;/p&gt;

&lt;p&gt;Those two categories can overlap in some systems, but they should not be treated as synonyms.&lt;/p&gt;

&lt;p&gt;The repository keeps the intentional memory race separated so the application/database concurrency tests can be reasoned about independently.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Tests Need to Prove
&lt;/h2&gt;

&lt;p&gt;Concurrency tests need stronger assertions than:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;no panic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;all goroutines completed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For inventory, useful assertions concern the invariant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;successful operations
remaining stock
initial stock
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For booking, the relevant assertion concerns how many rows survive for the unique slot.&lt;/p&gt;

&lt;p&gt;The final repository revision also made the PostgreSQL lost-update reproduction deterministic rather than relying on an arbitrary timing window.&lt;/p&gt;

&lt;p&gt;That is an important testing property.&lt;/p&gt;

&lt;p&gt;The test should deliberately create the interleaving being investigated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Commands
&lt;/h2&gt;

&lt;p&gt;Lab 05 is a nested Go module.&lt;/p&gt;

&lt;p&gt;The repository exposes dedicated Make targets including:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make lab-05-test
make lab-05-test-race
make lab-05-vet
make lab-05-fmt
make lab-05-integration
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The nested module boundary matters when running tests from the repository root: a plain root-level &lt;code&gt;go test ./...&lt;/code&gt; does not implicitly replace the dedicated Lab 05 targets.&lt;/p&gt;

&lt;p&gt;The README was updated to reflect that layout.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Lab Demonstrates
&lt;/h2&gt;

&lt;p&gt;The unsafe inventory example starts from code that is locally understandable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read
check
modify
write
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Concurrency changes the correctness requirement.&lt;/p&gt;

&lt;p&gt;Another operation can run between those steps.&lt;/p&gt;

&lt;p&gt;The fixes in the repository move the invariant to a boundary that can protect it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;atomic database update
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;row lock around the critical database flow
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For booking, the invariant is enforced through:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;database uniqueness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The broader result from the experiment is specific:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sequential correctness
does not prove
concurrent correctness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The invariant has to survive the interleavings that the system actually permits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Source
&lt;/h2&gt;

&lt;p&gt;This article is based on an implementation from my Software Engineering Lab:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/05-race-condition" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/05-race-condition&lt;/a&gt;&lt;/p&gt;

</description>
      <category>concurrency</category>
      <category>database</category>
      <category>backend</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Caching: Why Faster Reads Create Consistency Problems</title>
      <dc:creator>lukman lukman</dc:creator>
      <pubDate>Tue, 22 Sep 2026 07:48:22 +0000</pubDate>
      <link>https://dev.to/lukman-ss/caching-why-faster-reads-create-consistency-problems-kp9</link>
      <guid>https://dev.to/lukman-ss/caching-why-faster-reads-create-consistency-problems-kp9</guid>
      <description>&lt;p&gt;A dashboard request can look harmless.&lt;/p&gt;

&lt;p&gt;The problem appears when hundreds of users ask for the same expensive data at the same time.&lt;/p&gt;

&lt;p&gt;In Software Engineering Lab 04, the scenario is a workshop dashboard. Without caching, every request runs six database queries with joins and aggregations.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;500 concurrent users
        ↓
Dashboard request
        ↓
6 queries + join/aggregation per request
        ↓
3000 total DB queries
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first instinct is easy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;put Redis in front of the database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That reduces repeated work.&lt;/p&gt;

&lt;p&gt;But it also creates a new set of problems.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;How stale may the cached value be?
When should it be invalidated?
What happens when the key expires under load?
What happens if Redis is unavailable?
Can the cache return data from the wrong tenant?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the part I wanted to explore in Lab 04.&lt;/p&gt;

&lt;p&gt;The main mental model is not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;slow query → add cache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;reduce repeated work
        ↓
accept a consistency boundary
        ↓
design expiration, invalidation, failure, and concurrency behavior
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  PostgreSQL is still the source of truth
&lt;/h2&gt;

&lt;p&gt;The lab uses two storage layers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Technology&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary&lt;/td&gt;
&lt;td&gt;PostgreSQL&lt;/td&gt;
&lt;td&gt;Durable, persistent, authoritative storage that can rebuild the cache&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache&lt;/td&gt;
&lt;td&gt;Redis&lt;/td&gt;
&lt;td&gt;Derived data, TTL-bound, rebuilt on demand&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The important decision is ownership of correctness.&lt;/p&gt;

&lt;p&gt;In this lab, PostgreSQL remains authoritative for business data. Redis stores derived data that can disappear, expire, or be rebuilt.&lt;/p&gt;

&lt;p&gt;That means cache correctness has to work even when the cache is empty.&lt;/p&gt;

&lt;p&gt;The cache is an optimization layer, not the only copy of the business state.&lt;/p&gt;

&lt;h2&gt;
  
  
  The read path: Cache Aside
&lt;/h2&gt;

&lt;p&gt;Lab 04 uses Cache Aside for reads.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET cache
   ↓
hit? ───── yes ───→ return cached value
   ↓ no
query PostgreSQL
   ↓
populate Redis
   ↓
return value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A cache miss is not an application failure.&lt;/p&gt;

&lt;p&gt;It means the application has to rebuild the value from the authoritative source.&lt;/p&gt;

&lt;p&gt;The application explicitly knows about both storage layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Redis
PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That control is useful when Redis is unavailable because the application can still attempt to read from PostgreSQL, as long as the database and fallback capacity can handle the traffic.&lt;/p&gt;

&lt;p&gt;But Cache Aside also means the application now owns cache freshness.&lt;/p&gt;

&lt;p&gt;That is where the interesting failures start.&lt;/p&gt;

&lt;h2&gt;
  
  
  TTL is really a staleness decision
&lt;/h2&gt;

&lt;p&gt;The useful question is not simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does this data change?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The lab asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How long can stale data be accepted?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Examples recorded in the lab:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Data&lt;/th&gt;
&lt;th&gt;Max staleness&lt;/th&gt;
&lt;th&gt;Reasonable?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Dashboard statistics&lt;/td&gt;
&lt;td&gt;30s–2min&lt;/td&gt;
&lt;td&gt;Yes, operational metrics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stock display&lt;/td&gt;
&lt;td&gt;1–5s&lt;/td&gt;
&lt;td&gt;Yes, UI/UX only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wallet balance&lt;/td&gt;
&lt;td&gt;0s&lt;/td&gt;
&lt;td&gt;No, audit risk&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The difference matters.&lt;/p&gt;

&lt;p&gt;A dashboard metric can tolerate a freshness window that would be unacceptable for a transactional balance.&lt;/p&gt;

&lt;p&gt;So TTL is not just an expiry configuration. In this implementation, it represents an accepted freshness window and also acts as a recovery path when stale cache survives longer than expected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Invalidating before commit can reintroduce stale data
&lt;/h2&gt;

&lt;p&gt;Consider this write flow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DELETE cache
↓
update database
↓
COMMIT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It looks reasonable. Remove the old cache first, then write the new database value.&lt;/p&gt;

&lt;p&gt;The race appears when a reader enters between those operations.&lt;/p&gt;

&lt;p&gt;Lab 04 describes this timeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1 Writer: DELETE cache
T2 Reader: cache MISS → reads old DB value
T3 Reader: SET old value into cache
T4 Writer: DB COMMIT new value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Final state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database = new value
Cache    = old value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is stale cache, not data loss. The authoritative business data in PostgreSQL is still correct.&lt;/p&gt;

&lt;p&gt;A safer order is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB COMMIT
↓
DELETE cache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now a reader that misses after the delete can rebuild from the committed database value.&lt;/p&gt;

&lt;p&gt;But even this is not strong consistency.&lt;/p&gt;

&lt;p&gt;Another interleaving still exists:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1 Reader: cache MISS
T2 Reader: reads old DB value
T3 Writer: DB COMMIT new value
T4 Writer: DELETE cache
T5 Reader: SET old DB result into cache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Final state again:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database = new value
Cache    = old value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the conclusion is narrower:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;COMMIT → DELETE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;is safer than:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DELETE → COMMIT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but Cache Aside is still an eventually consistent optimization in this lab.&lt;/p&gt;

&lt;p&gt;TTL remains useful as a safety net for the residual stale window.&lt;/p&gt;

&lt;h2&gt;
  
  
  Updating Redis after a database write is not atomic either
&lt;/h2&gt;

&lt;p&gt;The lab also uses an application-managed update-on-write flow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB update
↓
DB COMMIT succeeds
↓
best-effort Redis SET
↓
return success
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database returns the authoritative value, then the application tries to update the cache.&lt;/p&gt;

&lt;p&gt;The problem is the boundary between PostgreSQL and Redis.&lt;/p&gt;

&lt;p&gt;These are separate systems, so this sequence is not atomic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB COMMIT
↓
Redis SET
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A process can fail between them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB COMMIT succeeds
↓
process crashes
↓
Redis SET never happens
↓
old cache remains
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Concurrent writers create another race:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Writer A commits value A
Writer B commits value B
Writer B SET cache = B
Writer A performs a late SET cache = A
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Final state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database = B
Cache    = A
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This was an important distinction for me:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cache Aside
→ read strategy

Invalidate-on-write / update-on-write
→ write strategy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They solve different parts of the cache lifecycle.&lt;/p&gt;

&lt;p&gt;Neither turns PostgreSQL and Redis into one atomic system.&lt;/p&gt;

&lt;h2&gt;
  
  
  One miss is normal. One thousand simultaneous misses are not
&lt;/h2&gt;

&lt;p&gt;Now consider a popular key reaching expiration.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cache expires
      ↓
1000 concurrent requests arrive
      ↓
1000 cache misses
      ↓
1000 parallel DB queries
      ↓
database overload / crash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the cache stampede scenario used in Lab 04.&lt;/p&gt;

&lt;p&gt;The cache successfully removes repeated database work while the entry exists, but expiration can suddenly send that work back to the database at the same time.&lt;/p&gt;

&lt;p&gt;For duplicate rebuilds inside one process, the lab uses &lt;code&gt;golang.org/x/sync/singleflight&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The flow includes a second cache check:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;initial cache GET
      ↓
miss
      ↓
singleflight.Do
      ↓
check cache again
      ↓
query DB
      ↓
populate cache
      ↓
share result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second check matters because another caller may already have populated the cache between the first miss and the point where this caller becomes the rebuild leader.&lt;/p&gt;

&lt;h2&gt;
  
  
  Singleflight stops at the process boundary
&lt;/h2&gt;

&lt;p&gt;Singleflight coordinates callers inside one process.&lt;/p&gt;

&lt;p&gt;A multi-instance deployment needs a different coordination boundary.&lt;/p&gt;

&lt;p&gt;Lab 04 implements a distributed-lock primitive:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;WithLock() = try-once lock primitive
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The conceptual cache regeneration flow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cache GET
   ↓ miss
acquire distributed lock
   ↓
check cache again
   ↓
query DB
   ↓
populate cache
   ↓
safe release
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lock requirements in the lab are explicit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;unique token or owner;&lt;/li&gt;
&lt;li&gt;TTL to avoid a permanent deadlock;&lt;/li&gt;
&lt;li&gt;atomic compare-and-delete when releasing;&lt;/li&gt;
&lt;li&gt;one holder must not remove another holder's lock.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But even this has a boundary.&lt;/p&gt;

&lt;p&gt;If regeneration takes longer than the lock TTL:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Instance A acquires lock
↓
lock expires
↓
Instance B acquires a new lock
↓
Instance A is still rebuilding
↓
duplicate rebuild can run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lock in Lab 04 reduces duplicate cache regeneration. It is not used as a correctness primitive for business transactions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Expiration can also be spread out
&lt;/h2&gt;

&lt;p&gt;The lab adds TTL jitter so many keys do not expire at nearly the same moment.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;60s + random 0–15s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The implementation produces values in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[base, base + maxJitter)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The TTL never goes below &lt;code&gt;base&lt;/code&gt;, and the upper bound is exclusive.&lt;/p&gt;

&lt;p&gt;The lab also discusses background refresh: refresh the cached value before expiry while clients continue receiving the existing cache value.&lt;/p&gt;

&lt;p&gt;Both techniques target expiration behavior rather than changing the source of truth.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cache keys are part of the data boundary
&lt;/h2&gt;

&lt;p&gt;The canonical key format in the lab is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{app}:{tenant}:{branch}:{resource}:{dimension}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cmms:tenant:42:branch:7:dashboard:2026-09-01
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rule is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every input that changes the result belongs in the cache key.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For the dashboard example, that includes tenant, branch, and business date.&lt;/p&gt;

&lt;p&gt;This is also a security boundary in a multi-tenant system.&lt;/p&gt;

&lt;p&gt;A sensitive cached value must include tenant scope. Missing isolation can expose data under the wrong tenant context.&lt;/p&gt;

&lt;p&gt;Key design also affects reuse.&lt;/p&gt;

&lt;p&gt;If 10,000 users read the same branch dashboard, this shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cache:tenant:42:user:{user_id}:dashboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;creates high-cardinality entries for data that is actually shared.&lt;/p&gt;

&lt;p&gt;The lab contrasts it with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cache:tenant:42:branch:{branch_id}:dashboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One branch-scoped value can be reused by those readers.&lt;/p&gt;

&lt;p&gt;The trade-off is not "specific keys are bad." The point is that key dimensions should match the result boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Redis failure changes where the traffic goes
&lt;/h2&gt;

&lt;p&gt;Cache Aside allows database fallback when Redis fails, but only while the authoritative dependency and fallback capacity remain available.&lt;/p&gt;

&lt;p&gt;The traffic pattern becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;traffic previously absorbed by Redis
      ↓
directly reaches PostgreSQL
      ↓
load spike / cache-failure amplification
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So graceful degradation does not mean the outage has no effect.&lt;/p&gt;

&lt;p&gt;It means the main function may continue with degraded performance while PostgreSQL can still handle the fallback traffic.&lt;/p&gt;

&lt;p&gt;The lab also separates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cache_miss
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cache_error
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;because an expected miss and an unavailable cache backend are different operational events.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hit ratio alone does not tell me whether the cache is worth it
&lt;/h2&gt;

&lt;p&gt;Lab 04 evaluates cache value together with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;avoided query cost
cache latency
memory cost
invalidation complexity
failure amplification
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The README gives this comparison:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;30% hit ratio for a 100ms operation
may still be valuable

99% hit ratio for a 0.1ms operation
may not justify the complexity
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The useful signal is not the percentage alone.&lt;/p&gt;

&lt;p&gt;The cost of the work being avoided matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cache comes after understanding the database work
&lt;/h2&gt;

&lt;p&gt;The diagnostic order in the lab is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;measure endpoint
↓
check N+1 queries
↓
inspect execution plan
↓
add / optimize indexes
↓
reduce selected columns
↓
optimize joins / subqueries
↓
evaluate caching if the workload needs it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prevents cache from becoming a way to hide an unexplained query problem.&lt;/p&gt;

&lt;p&gt;For a very cheap query and some workloads, the extra cache hop and operational complexity may not provide meaningful end-to-end benefit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model I take from this lab
&lt;/h2&gt;

&lt;p&gt;The naive model is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;query is expensive
→ add Redis
→ problem solved
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Lab 04 forces a longer chain of questions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What is the source of truth?
↓
How stale may this value be?
↓
What dimensions belong in the key?
↓
How is the value invalidated after writes?
↓
What happens during concurrent misses?
↓
What happens when Redis fails?
↓
Is the avoided work worth the added complexity?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the main lesson from the experiment.&lt;/p&gt;

&lt;p&gt;Caching can reduce repeated work in a read-heavy workload.&lt;/p&gt;

&lt;p&gt;But the moment a second storage layer is introduced, freshness, invalidation, concurrency, failure behavior, and isolation become part of the design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running the lab
&lt;/h2&gt;

&lt;p&gt;From the repository root:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker compose up &lt;span class="nt"&gt;-d&lt;/span&gt; redis

make lab-04-test
make lab-04-test-race
make lab-04-vet
make lab-04-demo
make lab-04-integration
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The demo scenarios can also be run directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;labs/04-caching

go run ./cmd/demo &lt;span class="nt"&gt;-scenario&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;without-cache
go run ./cmd/demo &lt;span class="nt"&gt;-scenario&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;cache-aside
go run ./cmd/demo &lt;span class="nt"&gt;-scenario&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;stampede-unprotected
go run ./cmd/demo &lt;span class="nt"&gt;-scenario&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;stampede-protected
go run ./cmd/demo &lt;span class="nt"&gt;-scenario&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;write-through
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Source Code
&lt;/h2&gt;

&lt;p&gt;Software Engineering Lab 04 — Caching&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/04-caching" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/04-caching&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Repository:&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/lukman-ss/software-engineering-lab" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Author: Lukman (&lt;code&gt;lukman-ss&lt;/code&gt;)&lt;/p&gt;

</description>
      <category>redis</category>
      <category>backend</category>
      <category>go</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Database Transactions Are a Boundary, Not a Safety Blanket</title>
      <dc:creator>lukman lukman</dc:creator>
      <pubDate>Mon, 21 Sep 2026 02:21:39 +0000</pubDate>
      <link>https://dev.to/lukman-ss/database-transactions-are-a-boundary-not-a-safety-blanket-1d9m</link>
      <guid>https://dev.to/lukman-ss/database-transactions-are-a-boundary-not-a-safety-blanket-1d9m</guid>
      <description>&lt;p&gt;The uncomfortable part of a failed payment flow is not always the error.&lt;/p&gt;

&lt;p&gt;Sometimes the function returns an error and the database still keeps half of the work.&lt;/p&gt;

&lt;p&gt;That is the failure Lab 03 demonstrates. A payment is inserted. The order is marked as paid. Then the flow fails before the wallet transaction is inserted.&lt;/p&gt;

&lt;p&gt;The result is not a clean failure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payment persisted = 1
order.status = paid
wallet_transactions = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the useful lesson: database transactions are not just a way to call &lt;code&gt;ROLLBACK&lt;/code&gt;. They are a boundary. Anything inside the boundary can commit or roll back together. Anything outside it cannot be magically undone by the database.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Scenario
&lt;/h2&gt;

&lt;p&gt;The lab starts with a local payment operation.&lt;/p&gt;

&lt;p&gt;Initial state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;orders.id = 101
orders.status = pending

invoices.order_id = 101
invoices.status = unpaid
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The service needs to do three local writes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;payments&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'completed'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'paid'&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;wallet_transactions&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;type&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'credit'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The injected failure happens after the first two statements, but before the wallet transaction insert.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;PaymentServiceUnsafe&lt;/code&gt; executes statements directly with &lt;code&gt;db.ExecContext&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The flow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;INSERT payment
  ↓
UPDATE order to paid
  ↓
injected failure
  ↓
wallet transaction is not inserted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no shared transaction boundary around the related writes. So when the failure happens, the database does not roll the earlier statements back.&lt;/p&gt;

&lt;p&gt;The test result is the whole point:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payment persisted = 1
order.status = paid
wallet_transactions = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application failed. The database did not return to the original state.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Went Wrong
&lt;/h2&gt;

&lt;p&gt;The bug is not that SQL failed.&lt;/p&gt;

&lt;p&gt;The bug is that the service treated a multi-step business operation as separate database statements.&lt;/p&gt;

&lt;p&gt;From the business side, the payment row, the paid order status, and the wallet transaction belong to the same local invariant. From the unsafe database flow, they are just separate statements executed in order.&lt;/p&gt;

&lt;p&gt;Without a transaction, the database has no reason to undo statement one and statement two just because the application never reaches statement three.&lt;/p&gt;

&lt;h2&gt;
  
  
  Using a Local Database Transaction
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;PaymentServiceSafe&lt;/code&gt; changes the boundary.&lt;/p&gt;

&lt;p&gt;It starts a transaction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;tx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;BeginTx&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It keeps rollback as the default if commit does not happen:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;committed&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="no"&gt;false&lt;/span&gt;
&lt;span class="k"&gt;defer&lt;/span&gt; &lt;span class="k"&gt;func&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;committed&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Rollback&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the related writes are executed through &lt;code&gt;tx.ExecContext&lt;/code&gt; instead of &lt;code&gt;db.ExecContext&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If all local writes succeed, the transaction commits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;tx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Commit&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Successful Flow
&lt;/h2&gt;

&lt;p&gt;The successful local flow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BEGIN TRANSACTION
  ↓
INSERT payment
  ↓
UPDATE order status
  ↓
INSERT wallet transaction
  ↓
COMMIT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At commit time, the local database state becomes complete together.&lt;/p&gt;

&lt;p&gt;The lab uses the same idea again in the outbox example. &lt;code&gt;InvoiceServiceOutbox&lt;/code&gt; updates the invoice and inserts an outbox event in one local transaction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BEGIN
  ↓
UPDATE invoices SET status = 'paid'
  ↓
INSERT INTO outbox_events (... status = 'pending' ...)
  ↓
COMMIT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After commit, both the business state and the event intent exist locally.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure Flow
&lt;/h2&gt;

&lt;p&gt;The failure path is where the difference becomes visible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BEGIN TRANSACTION
  ↓
INSERT payment
  ↓
UPDATE order status
  ↓
injected failure
  ↓
ROLLBACK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lab verifies the final state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payments = 0
order.status = pending
wallet_transactions = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a clean local failure. The service still returns an error, but the database does not keep half of the payment flow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before vs After
&lt;/h2&gt;

&lt;p&gt;Unsafe failure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payments = 1
order.status = paid
wallet_transactions = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Safe local transaction failure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payments = 0
order.status = pending
wallet_transactions = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is what local atomicity buys you.&lt;/p&gt;

&lt;p&gt;But the lab does not stop there, because this is only the easy boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rollback Does Not Undo WhatsApp
&lt;/h2&gt;

&lt;p&gt;The next example uses &lt;code&gt;DistributedOrderService&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The service opens a database transaction, inserts a payment, updates an invoice, sends a WhatsApp notification, then hits a simulated ERP integration error.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BEGIN TRANSACTION
  ↓
INSERT payment
  ↓
UPDATE invoice
  ↓
Send WhatsApp notification
  ↓
simulated ERP integration error
  ↓
ROLLBACK database transaction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The final state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;WhatsApp sent count = 1
payments = 0
paid invoices = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database rollback worked. The WhatsApp message was still sent.&lt;/p&gt;

&lt;p&gt;That is not a contradiction. WhatsApp is outside the database transaction boundary. The database can roll back its own rows. It cannot recall a message that has already been sent through another system.&lt;/p&gt;

&lt;p&gt;The lab uses the same boundary idea for email, SMS, ERP APIs, payment gateway APIs, and message broker publishes.&lt;/p&gt;

&lt;h2&gt;
  
  
  HTTP Inside a Transaction
&lt;/h2&gt;

&lt;p&gt;The lab also shows a blocking external call while a database transaction is open:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BEGIN TRANSACTION
  ↓
UPDATE invoice SET status = 'paid'
  ↓
HTTP call blocks
  ↓
transaction stays open
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test verifies that the transaction remains open during the blocking external call, then closes after commit.&lt;/p&gt;

&lt;p&gt;The issue here is not only failure. The lifetime of the database transaction is now tied to an external resource.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Dual-Write Gap
&lt;/h2&gt;

&lt;p&gt;Another failure appears when database commit and event publish are separate operations.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UPDATE invoice
  ↓
COMMIT succeeds
  ↓
process crashes
  ↓
Publish event never happens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test shows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;invoice.status = paid
published events = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The invoice was paid, but no event was published.&lt;/p&gt;

&lt;p&gt;Reversing the order does not make the operation atomic. Publishing before commit can produce an event for a database state that never commits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Transactional Outbox
&lt;/h2&gt;

&lt;p&gt;The lab uses transactional outbox to keep the business state and event intent in one local transaction.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Local transaction
  ↓
Record business state + event intent
  ↓
COMMIT
  ↓
Dispatcher publishes pending events
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;InvoiceServiceOutbox&lt;/code&gt; inserts an &lt;code&gt;outbox_events&lt;/code&gt; row with status &lt;code&gt;pending&lt;/code&gt;. The dispatcher later reads pending events, publishes them to the broker, and marks them as &lt;code&gt;published&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The important part is the local atomic step: invoice state and event intent are saved together.&lt;/p&gt;

&lt;h2&gt;
  
  
  Idempotent Consumer
&lt;/h2&gt;

&lt;p&gt;The lab then shows the consumer side.&lt;/p&gt;

&lt;p&gt;The commission worker stores a processed-event marker and the business state in the same transaction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;processed_events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;consumer_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;processed_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;CONFLICT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;consumer_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;DO&lt;/span&gt; &lt;span class="k"&gt;NOTHING&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the same consumer receives the same event again, the insert affects zero rows and the event is skipped.&lt;/p&gt;

&lt;p&gt;The test result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;same consumer: processed once
same consumer duplicate: skipped
same event, different consumer: processed independently
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This matters because the outbox side and the consumer side are connected. Recording an event intent is not enough. The receiver also has to handle repeated delivery safely.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Lab Demonstrates
&lt;/h2&gt;

&lt;p&gt;A local transaction is the right tool for local state that must commit or roll back together.&lt;/p&gt;

&lt;p&gt;It does not roll back external side effects.&lt;/p&gt;

&lt;p&gt;It does not make database commit and broker publish atomic when those are executed as separate operations.&lt;/p&gt;

&lt;p&gt;Outbox solves the local database/event-intent part by storing both in one transaction.&lt;/p&gt;

&lt;p&gt;Idempotent consumer logic handles repeated processing on the consumer side by storing a dedup marker with the business update.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;First define the transaction boundary.&lt;/li&gt;
&lt;li&gt;Put related local writes inside the same transaction.&lt;/li&gt;
&lt;li&gt;Do not treat external calls as rollbackable database work.&lt;/li&gt;
&lt;li&gt;Do not keep a transaction open longer than necessary while waiting on external systems.&lt;/li&gt;
&lt;li&gt;Use outbox when database state and event publishing must be coordinated.&lt;/li&gt;
&lt;li&gt;Make consumers idempotent when the same event can be processed more than once.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Source Code
&lt;/h2&gt;

&lt;p&gt;Repository:&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/lukman-ss/software-engineering-lab" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Lab:&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/03-database-transaction" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/03-database-transaction&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Author:&lt;br&gt;&lt;br&gt;
Lukman (&lt;code&gt;lukman-ss&lt;/code&gt;)&lt;/p&gt;

</description>
      <category>database</category>
      <category>backend</category>
      <category>go</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Database Index: Why Queries Become Slow as Data Grows</title>
      <dc:creator>lukman lukman</dc:creator>
      <pubDate>Thu, 17 Sep 2026 07:36:43 +0000</pubDate>
      <link>https://dev.to/lukman-ss/database-index-why-queries-become-slow-as-data-grows-4oae</link>
      <guid>https://dev.to/lukman-ss/database-index-why-queries-become-slow-as-data-grows-4oae</guid>
      <description>&lt;p&gt;This query looks ordinary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;branch_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'FINISHED'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;BETWEEN&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-01'&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-31'&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With a small dataset, a query like this may appear completely fine.&lt;/p&gt;

&lt;p&gt;Lab 02 does not rely on assumptions. The query is tested using &lt;code&gt;EXPLAIN (ANALYZE, BUFFERS)&lt;/code&gt; against the same dataset.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dataset
&lt;/h2&gt;

&lt;p&gt;Table: &lt;code&gt;service&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Total rows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;500,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Status distribution:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Fraction&lt;/th&gt;
&lt;th&gt;Rows&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;FINISHED&lt;/td&gt;
&lt;td&gt;70.00%&lt;/td&gt;
&lt;td&gt;350,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CANCELLED&lt;/td&gt;
&lt;td&gt;20.00%&lt;/td&gt;
&lt;td&gt;100,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IN_PROGRESS&lt;/td&gt;
&lt;td&gt;5.00%&lt;/td&gt;
&lt;td&gt;25,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WAITING&lt;/td&gt;
&lt;td&gt;4.90%&lt;/td&gt;
&lt;td&gt;24,500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PENDING_REFUND&lt;/td&gt;
&lt;td&gt;0.10%&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Branch 2 is the busiest branch, containing 25.00% of the data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Baseline
&lt;/h2&gt;

&lt;p&gt;The baseline only has constraint-backed indexes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;service_pkey
service_invoice_no_key
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no secondary index on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;branch_id
status
service_date
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Baseline query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;EXPLAIN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;ANALYZE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;BUFFERS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;branch_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'FINISHED'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;BETWEEN&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-01'&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-31'&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What should be inspected:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;plan node;&lt;/li&gt;
&lt;li&gt;actual rows;&lt;/li&gt;
&lt;li&gt;rows removed by filter;&lt;/li&gt;
&lt;li&gt;explicit &lt;code&gt;Sort&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;shared read / shared hit;&lt;/li&gt;
&lt;li&gt;planning time;&lt;/li&gt;
&lt;li&gt;execution time.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Add the Candidate Index
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;idx_service_branch_status_date&lt;/span&gt;
&lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;branch_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This index follows the query shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;branch_id equality
→ status equality
→ service_date range/order
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is a candidate for this specific query, not a universal index.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run the Same Query Again
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;EXPLAIN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;ANALYZE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;BUFFERS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;branch_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'FINISHED'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;BETWEEN&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-01'&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-31'&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;What to inspect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Before index&lt;/td&gt;
&lt;td&gt;Seq Scan, Sort, buffers, rows, time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;After index&lt;/td&gt;
&lt;td&gt;Index usage, Index Cond, Sort presence, buffers, rows, time&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The lab does not store fixed execution-time numbers. Those values should come from the local run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cardinality Is Not Match Fraction
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;status&lt;/code&gt; has 5 distinct values. That is its cardinality.&lt;/p&gt;

&lt;p&gt;Match fraction is the proportion of rows that match a predicate.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Predicate&lt;/th&gt;
&lt;th&gt;Match fraction&lt;/th&gt;
&lt;th&gt;Expected / likely plan&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;status = 'FINISHED'&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;70.0%&lt;/td&gt;
&lt;td&gt;Seq Scan likely cheaper&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;status = 'PENDING_REFUND'&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.1%&lt;/td&gt;
&lt;td&gt;index-based plan likely cheaper&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The lesson is that low cardinality does not automatically make an index useless.&lt;/p&gt;

&lt;h2&gt;
  
  
  Column Order Matters
&lt;/h2&gt;

&lt;p&gt;The lab compares:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;idx_service_a_branch_status_date&lt;/span&gt;
    &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;branch_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;idx_service_b_date_branch_status&lt;/span&gt;
    &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;service_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;branch_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;idx_service_c_status_date_branch&lt;/span&gt;
    &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;branch_id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the main query, the first index becomes a strong candidate because the leading equality predicates can narrow the scan range more effectively.&lt;/p&gt;

&lt;h2&gt;
  
  
  ORDER BY + LIMIT
&lt;/h2&gt;

&lt;p&gt;Dashboard query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;branch_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'FINISHED'&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Index:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;idx_service_branch_status_date_desc&lt;/span&gt;
    &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;branch_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service_date&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What to inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;whether &lt;code&gt;Sort&lt;/code&gt; is absent;&lt;/li&gt;
&lt;li&gt;whether the plan uses &lt;code&gt;Index Scan&lt;/code&gt; or &lt;code&gt;Index Only Scan&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;how many rows are examined before &lt;code&gt;LIMIT&lt;/code&gt; is satisfied.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Index Is Not Free
&lt;/h2&gt;

&lt;p&gt;The lab also tests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;inserting 1,000 rows without secondary indexes;&lt;/li&gt;
&lt;li&gt;inserting 1,000 rows with 1 composite index;&lt;/li&gt;
&lt;li&gt;inserting 1,000 rows with 4 secondary indexes;&lt;/li&gt;
&lt;li&gt;updating indexed vs non-indexed columns;&lt;/li&gt;
&lt;li&gt;storage cost using &lt;code&gt;pg_relation_size&lt;/code&gt;, &lt;code&gt;pg_indexes_size&lt;/code&gt;, and &lt;code&gt;pg_total_relation_size&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An index can improve specific reads, but it also adds write and storage cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Do not optimize based on guesses.&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;EXPLAIN (ANALYZE, BUFFERS)&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A composite index should follow the query shape.&lt;/li&gt;
&lt;li&gt;Column order matters.&lt;/li&gt;
&lt;li&gt;Cardinality is different from match fraction.&lt;/li&gt;
&lt;li&gt;A Seq Scan can be the correct plan.&lt;/li&gt;
&lt;li&gt;Indexes have write and storage costs.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Hands-on Lab
&lt;/h2&gt;

&lt;p&gt;Software Engineering Lab #02 — Database Index&lt;/p&gt;

&lt;p&gt;Author: Lukman&lt;/p&gt;

&lt;p&gt;GitHub: lukman-ss&lt;/p&gt;

&lt;p&gt;Repository:&lt;br&gt;
&lt;a href="https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/02-database-index" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/02-database-index&lt;/a&gt;&lt;/p&gt;

</description>
      <category>database</category>
      <category>postgres</category>
      <category>sql</category>
      <category>backend</category>
    </item>
    <item>
      <title>Making POST Requests Safe to Retry with Idempotency Keys</title>
      <dc:creator>lukman lukman</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:05:50 +0000</pubDate>
      <link>https://dev.to/lukman-ss/making-post-requests-safe-to-retry-with-idempotency-keys-3ip0</link>
      <guid>https://dev.to/lukman-ss/making-post-requests-safe-to-retry-with-idempotency-keys-3ip0</guid>
      <description>&lt;p&gt;&lt;em&gt;A practical guide to duplicate execution, request fingerprints, concurrency, and safe retries.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A request timeout does not mean the operation failed.&lt;/p&gt;

&lt;p&gt;Sometimes the server has already completed the operation, but the response never reaches the client.&lt;/p&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /orders/order-123/pay
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"amount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;500000&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The timeline may look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client -&amp;gt; POST /pay
Server -&amp;gt; payment succeeds
Server -&amp;gt; response lost
Client -&amp;gt; timeout
Client -&amp;gt; retry
Server -&amp;gt; payment succeeds again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The problem is not the duplicate HTTP request itself. The problem is &lt;strong&gt;duplicate execution&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The target invariant
&lt;/h2&gt;

&lt;p&gt;For a mutation API, the goal should be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 logical operation
=
1 effective side effect
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;even if the request is delivered multiple times.&lt;/p&gt;

&lt;h2&gt;
  
  
  Introduce an Idempotency Key
&lt;/h2&gt;

&lt;p&gt;The client generates an identifier for one logical operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;Idempotency-Key: 550e8400-e29b-41d4-a716-446655440000
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key must remain the same across retries.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;same logical operation -&amp;gt; same key
new logical operation  -&amp;gt; new key
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The basic flow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
  ↓
Check Idempotency-Key
  ↓
New key?
  ↓
Process operation
  ↓
Store result
  ↓
Return response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On retry:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Same key
  ↓
Existing completed result
  ↓
Replay response
  ↓
Do not execute side effect again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why check-then-act fails
&lt;/h2&gt;

&lt;p&gt;This is not enough:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;if key does not exist:
    process payment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two concurrent requests may both observe the key as missing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A -&amp;gt; key not found
B -&amp;gt; key not found

A -&amp;gt; process
B -&amp;gt; process
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a race condition. The uniqueness decision needs to be atomic.&lt;/p&gt;

&lt;p&gt;For a relational database, that normally means enforcing the invariant in storage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;UNIQUE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;idx_idempotency_key&lt;/span&gt;
&lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;payments&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;idempotency_key&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Application checks are useful. The unique constraint is the final guard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same key, different payload
&lt;/h2&gt;

&lt;p&gt;An idempotency key must represent exactly one logical operation.&lt;/p&gt;

&lt;p&gt;This should be accepted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;key: abc
amount: 500000

retry

key: abc
amount: 500000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This should not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;key: abc
amount: 500000

retry

key: abc
amount: 800000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lab uses a request fingerprint based on SHA-256:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;hashRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="kt"&gt;byte&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;sum&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;sha256&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Sum256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;hex&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;EncodeToString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;same key + same fingerprint
-&amp;gt; replay safely

same key + different fingerprint
-&amp;gt; 409 Conflict
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  PROCESSING vs COMPLETED
&lt;/h2&gt;

&lt;p&gt;The safe flow uses two important states:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PROCESSING
COMPLETED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When a second request arrives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PROCESSING
-&amp;gt; return 409
-&amp;gt; do not execute payment again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the operation is already finished:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;COMPLETED
-&amp;gt; load stored response
-&amp;gt; replay it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Idempotency is not a transaction
&lt;/h2&gt;

&lt;p&gt;A database transaction provides atomicity inside one execution.&lt;/p&gt;

&lt;p&gt;Idempotency protects a logical operation across multiple execution attempts.&lt;/p&gt;

&lt;p&gt;These can both succeed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request A
BEGIN
create payment
COMMIT

Request B
BEGIN
create payment
COMMIT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database is consistent. The business result is not. The customer was charged twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  External side effects have a different boundary
&lt;/h2&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Call payment provider
↓
Provider SUCCESS
↓
Local commit fails
↓
ROLLBACK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The local rollback does not undo the external charge.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;local rollback
!=
external rollback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the provider supports provider-side idempotency, the backend should reuse a stable key for that external request too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab implementation
&lt;/h2&gt;

&lt;p&gt;The lab intentionally keeps the infrastructure small.&lt;/p&gt;

&lt;p&gt;Unsafe:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;ProcessPayment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt; &lt;span class="n"&gt;PaymentRequest&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PaymentResult&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;gateway&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Charge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Safe:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;validate key
↓
calculate fingerprint
↓
reserve operation
↓
PROCESSING
↓
execute payment
↓
store result
↓
COMPLETED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The storage implementation uses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;map + sync.RWMutex
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This simulates atomic uniqueness without introducing a real database into the lab.&lt;/p&gt;

&lt;p&gt;Run both versions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;go &lt;span class="nb"&gt;test&lt;/span&gt; ./labs/01-idempotency/unsafe/... &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="nt"&gt;-count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1

go &lt;span class="nb"&gt;test&lt;/span&gt; ./labs/01-idempotency/safe/... &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="nt"&gt;-count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Final mental model
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Retry is normal.
Duplicate execution is the problem.
Idempotency makes retry safe.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Source code:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/01-idempotency" rel="noopener noreferrer"&gt;https://github.com/lukman-ss/software-engineering-lab/tree/main/labs/01-idempotency&lt;/a&gt;&lt;/p&gt;

</description>
      <category>softwaredevelopment</category>
      <category>softwareengineering</category>
      <category>distributedsystems</category>
      <category>go</category>
    </item>
  </channel>
</rss>
