Your feature passed every test. Then two users clicked at the same moment. Why concurrency bugs keep coming back, and what developers and QA should each do about it.
The Feature Worked. The System Failed.
Here is a story I have seen in many forms. (This is a hypothetical scenario built from common real-world patterns.)
A team ships a "last item in stock" feature. It passes functional testing, API testing, automation, regression, and staging. QA signs off.
On the first busy evening, two customers press "Buy" within a few milliseconds of each other. Both get a success message. The warehouse has one item and two orders.
The feature worked. The system failed under concurrency.
Next sprint, a coupon limit ships with the same bug: one coupon, used twice. Then a wallet balance. Then a booking slot. Different features, same root problem.
A race condition is not a bug you fix once. It returns whenever a new feature touches shared state.
Let's take it in order: when concurrent work happens, how it becomes a race condition, what database and design decisions create the risk, and what developers and QA should each do.
Part 1: When Does Concurrent Work Happen?
Concurrent work means more than one thing is happening at the same time on the same system. This is normal. It is not a bug. Every real product is concurrent all day.
A simple picture: two people stand at two different ATMs and withdraw money from the same joint account at the same moment. The bank must handle both correctly. That is concurrency, and the shared account is the "shared state."
Real situations in software:
- Two users, one resource. Two customers buy the last item. Two agents book the same seat.
- One user, two requests. A user double-clicks "Pay," or a mobile app retries after a slow network. The server receives the same action twice.
- Two admins, one record. Two people open the same order and both change its status.
- User and background job. A customer cancels an order while a scheduled job is shipping it.
- Two workers, one event. A queue delivers the same message and two workers process it.
- Many servers, one database. Production runs several copies of your app, all writing to the same rows.
- A slow job that overlaps itself. A nightly job runs longer than expected and the next run starts before it ends.
None of these needs a huge traffic spike. Two requests are enough. Concurrency is not only a performance topic. It is a correctness topic.
Concurrent work is normal. It becomes a problem only when it touches shared state without protection.
Part 2: How Concurrent Work Becomes a Race Condition
A race condition happens when the result depends on which request finishes first, or on how the steps of two requests interleave.
Three ingredients must be present:
- Shared state: something both requests read or change (stock, balance, status).
- Overlapping timing: both requests are in progress at the same time.
- A gap between checking and changing: the code reads a value, decides, and writes later, with no protection in between.
Remove any one ingredient and the race disappears.
Example 1: Overselling the last item
Stock is 1. The code does this:
Step 1 → Read the stock (returns 1)
Step 2 → Check: is stock greater than 0? Yes, continue
Step 3 → Update stock to 0
Step 4 → Create the order
Request A and Request B both finish Step 1 before either reaches Step 3. Both see stock = 1. Both pass the check. Both create an order.
sequenceDiagram
participant A as Request A
participant DB as Database (stock = 1)
participant B as Request B
A->>DB: Read stock
DB-->>A: stock = 1
B->>DB: Read stock
DB-->>B: stock = 1
Note over A,B: Both pass the "stock > 0" check
A->>DB: Update stock = 0
A->>DB: Create order A
B->>DB: Update stock = 0
B->>DB: Create order B
Note over DB: Result: stock = 0, but 2 orders exist for 1 item
The time between the read (Step 1) and the write (Step 3) is called the race window. It may last only a few milliseconds, but production traffic will find it.
Example 2: The lost update
A wallet has a balance of 100. Two deposits of 50 arrive together.
Request A reads 100 → Request B reads 100 → A writes 150 → B writes 150
The customer should have 200. They have 150. No error was shown. A lost update is dangerous because it fails silently.
Example 3: Invalid state
An order is PAID. A customer cancels while the warehouse marks it SHIPPED. Both updates succeed. The order ends as CANCELLED, but the package is already on a truck. Each update was valid alone. The combination is not.
Example 4: Duplicate processing
A payment request times out. The client retries. But the first request was still running and finished. The customer is charged twice.
Why is it so hard to catch?
- It depends on exact timing, so it is intermittent.
- Sequential tests never create the overlap.
- It often passes in staging, where traffic is low.
- It leaves quiet damage (wrong numbers, duplicates, bad state), not crashes.
Why does it keep returning?
Each new feature has its own code path and its own read-check-write. The fix for the stock feature protects only the stock feature. The coupon feature starts with no protection, because nobody asked the concurrency question for it.
Every feature that changes shared state should trigger one question: what happens if two actors do this at the same time?
Part 3: How It Arises: Database and System Design Concerns
A race condition is created in two places: the database layer and the system design.
Database concerns
1. Read-modify-write outside a safe boundary. The most common cause. The read and the write are separate steps with a gap between them.
2. Transaction boundary in the wrong place. If the read is outside the transaction and the write is inside, the gap is still open. A transaction that is too wide also hurts: it holds locks longer and causes waiting and deadlocks.
3. Isolation level. Isolation controls what one running transaction can see of another. Different levels allow different anomalies. A test that passes on one setting can fail on another. Check that your test database uses the same isolation setting as production.
4. Missing constraints. If a rule matters ("one coupon per user," "one active booking per slot"), the database should enforce it with a unique constraint. Application code can have bugs or forgotten code paths. A constraint cannot be bypassed.
5. Locking choices. Row-level locks block other writers on the same row. They protect data but make requests wait. Two transactions can also lock each other, causing a deadlock, where the database kills one of them.
6. Optimistic vs pessimistic locking.
- Pessimistic: lock first, work after. Safe, but slower under contention.
- Optimistic: no lock, but a version number. If the version changed, the update is rejected. Fast, but the app must handle the conflict error.
System design concerns
A request travels through several layers, and each one carries its own concurrency risk:
flowchart LR
U[User] --> F[Frontend]
F --> G[API Gateway]
G --> AI["App instances<br/>(multiple copies)"]
AI --> BL[Business logic]
BL --> DB[(Database)]
BL --> Q["Queue & workers"]
BL --> EXT[External service]
1. Multiple app instances. A lock in one server's memory does nothing for the other servers. Protection must live in shared places, like the database or a shared lock service.
2. Queues that deliver at least once. Many message systems can deliver the same message more than once. Consumers must be safe to run twice.
3. Retries and timeouts. When a call times out, the caller does not know if the work happened. A blind retry can duplicate it.
4. Missing idempotency. Idempotent means doing something twice has the same effect as doing it once. Without an idempotency key or a natural unique rule, the system cannot recognize a repeated request.
5. Cache and database out of sync. If the cache says 5 in stock and the database says 0, a decision from the cache is wrong. That is a race between two copies of the truth.
6. Async flows and external services. A payment provider or warehouse system adds delay, and delay widens race windows.
7. Unclear source of truth. If two services can both change the same data, races between services are very hard to prevent.
Part 4: Developer Concerns
Developers own the implementation. Here is what I would expect, and what I would gladly review.
1. Write down the rule first. "Stock can never go below zero." "One coupon per user." If the rule is not clear, nobody can protect it.
2. Make the check and the change one step. For the stock example, do it in one atomic database statement:
UPDATE products
SET stock = stock - 1
WHERE id = 101 AND stock > 0;
Then check how many rows were updated. One row means success. Zero rows means someone else got it, so return a clear "out of stock" response.
3. Choose the right control for the problem. There is no universal answer. Common options:
- Atomic update: best for counters and stock.
- Row lock (pessimistic): useful when several steps must happen together.
BEGIN;
SELECT * FROM products WHERE id = 101 FOR UPDATE;
-- check stock, then update
UPDATE products SET stock = stock - 1 WHERE id = 101;
COMMIT;
- Optimistic version check: useful for records people edit.
UPDATE orders
SET status = 'SHIPPED', version = version + 1
WHERE id = 55 AND version = 7;
If zero rows changed, someone else edited first.
- Unique constraint plus idempotency key: best for "do this only once." A repeat request fails on the unique index, and you return the first result.
- Serialize through a queue: all changes for one item go through one ordered path.
The right choice depends on your database, transaction boundary, business rule, isolation level, traffic, and how strict consistency must be. One SQL pattern does not solve every case.
4. Do not rely on in-memory locks in a system with multiple instances.
5. Handle conflict errors kindly. A losing request should get a clear response, not a 500 error. Decide when it is safe to retry.
6. Keep transactions short. Long transactions hold locks and cause waiting and deadlocks.
7. Make consumers and retries safe. Any handler that a queue or client may call twice must be idempotent.
8. Add request IDs to logs. When something breaks at 2 a.m., correlation IDs are how you reconstruct what happened.
9. Tell QA where the risk is. A short note in the pull request, like "this updates stock inside a transaction using X," helps QA target tests.
Part 5: QA Concerns
This is where I spend most of my time. The SDET does not have to write the locking code. The SDET has to notice the risk, design the tests, and check the truth in the data.
Step 1: Spot the risk before testing
The most valuable sentence an SDET can say in grooming:
"This feature changes shared state, so concurrency is part of its risk profile."
Use this checklist for every feature.
Shared state
- Do multiple users touch the same resource?
- Is stock, balance, or status involved?
Parallel requests
- Can two requests arrive at nearly the same time?
- Can a user double-click, or a client retry?
Async processing
- Are queues or workers involved?
- Can a job run twice?
Database
- Is there read → modify → write logic?
- Which constraint protects the rule?
External dependencies
- What if a payment call times out? Can the client retry?
Distributed system
- Are there multiple instances? Is a cache involved?
Several "yes" or "not sure" answers mean the feature needs concurrency tests.
Step 2: Test in four levels
Most teams stop at level 2.
- Level 1, Functional: Request A → Success
- Level 2, Sequential: Request A completes → then Request B completes
- Level 3, Concurrent: Request A and Request B hit the API at the same moment
- Level 4, High concurrency: 10 users → 50 requests → 100 requests → 500 requests
Measure more than the status code: correctness, duplicate records, lost updates, wrong state, transaction failures, deadlocks, timeouts, retry behavior, error rate, response time, and throughput.
A concurrency test is not passed because the API returned HTTP 200. Two 200 responses for one item is a failure. The final database state and business outcome are the real result.
Step 3: A practical test design
Scenario: only one user can buy the final item.
- Initial state: Stock = 1
-
Concurrent requests: User A and User B both send
POST /orders - Expected: one request succeeds, one fails gracefully (for example, "out of stock")
- Database: stock = 0, exactly 1 successful order
There must be no negative stock, no duplicate reservation, no inconsistent order status, and no double payment.
How to run it: write a small script (JavaScript or Python) that fires 20 requests at the same instant using a shared start time, and prints every status code. You do not need a full load-testing platform for a first reproduction.
Then verify the truth in the database:
- Query the stock for the product. It must be 0, never negative.
- Count the successful orders for the product. It must be exactly 1.
- Group orders by user and look for any user with more than one.
Check five things:
- API response: one success, the rest handled cleanly, no 500 errors
- Database state: stock is 0, never negative
- Queue or event state: one event published, not two
- Logs: no deadlocks or unexpected exceptions
- Business outcome: one order, one charge, one shipment
Step 4: Reproduce it reliably
Race conditions are intermittent, so make them appear on purpose:
- Start all requests at the same instant
- Use more parallel workers
- Repeat the test many times
- Add a small artificial delay between the read and the write in a test build, to widen the race window
- Hit the same row from many requests
- Simulate retries and network delay
- Reset test data to a known state before each run
- Tag every request with a unique ID for log tracing
Flaky test or race condition?
Run once → Pass. Run 100 times → Failure. This does not automatically mean the test is flaky.
- A flaky test is unreliable because of the test itself: bad waits, shared test data, an unstable environment.
- A race condition is real product behavior that depends on timing. The test is reliable and is telling you the truth.
To tell them apart, fix the test setup (isolated data, deterministic waits) and rerun. If it still fails under concurrency, it is the product. Never label a concurrency failure "flaky, ignore." That is how these bugs reach production.
Step 5: Use tools by purpose
To generate concurrent requests
- A small custom script is often best for the first reproduction
- k6 or JMeter when you need many virtual users and metrics
- Postman/Newman or Playwright API tests when your suite already exists
To investigate the database
- Lock and transaction views for waiting queries
- Query, slow query, and deadlock logs
- Direct queries comparing final state to expected state
To investigate the application
- Request and correlation IDs
- Timestamps in logs
- Distributed tracing, if available
- Queue dashboards and worker logs
In CI/CD
A small concurrency check can run before release. Even 20 concurrent requests that verify final database state can catch a broken feature on every deployment.
Step 6: Follow a feature-level workflow
flowchart TD
A[New feature] --> B[Identify shared state]
B --> C[Identify concurrent actors]
C --> D[Identify race window]
D --> E[Review system design]
E --> F[Review database & transaction strategy]
F --> G[Create concurrent test]
G --> H[Run controlled reproduction]
H --> I[Increase concurrency]
I --> J[Verify API, database, and business state]
J --> K[Analyze logs and tracing]
K --> L[Regression]
L --> M[Release decision]
The first four steps happen before any test is written. Concurrency thinking is a design-time activity. If you start at "create concurrent test," you are already late.
Step 7: Ask these in feature grooming
- What shared state does this feature modify?
- Can two actors modify it at the same time?
- What prevents duplicate processing?
- Is this operation idempotent?
- Where does the transaction start and end?
- What happens if the request times out after the server has processed it?
- Can the queue deliver the same message twice?
- Which database constraint protects this business rule?
- What happens when two workers pick up the same entity?
- How can QA reproduce the expected concurrency behavior?
If nobody can answer question 10, you have found a testability problem. Better to find it now than after release.
Step 8: Release readiness
Before release, a feature that changes shared state should have answers to these:
- What happens when two requests arrive at once?
- What happens when the same request is retried?
- What happens when a worker processes an event twice?
- What happens when an update conflicts?
- What happens when an external service times out and the client retries?
- Is the operation idempotent where it needs to be?
- Can invalid state or duplicate records exist?
- Is the final database state correct?
- Are failures visible in logs or alerts?
Release readiness is not only "does the happy path work?" It is "does the system keep its business rules true when concurrency goes wrong?"
Be honest about where this effort is not worth it. A single-user internal tool, data written by only one process, or a read-only feature does not need the same depth. Match the effort to the risk.
Who Owns What
Nobody owns this alone. It is built across the chain:
- Product: states the business rule clearly
- Architect: decides the source of truth and how requests are ordered
- Developer: implements the control and documents the transaction boundary
- Database engineer: reviews constraints, indexes, isolation, and locking
- QA / SDET: spots the risk, designs the concurrent tests, verifies final data
- DevOps / SRE: makes failures visible with logs, metrics, alerts, and correct retry and timeout settings
Closing
Every new feature is a new opportunity for a concurrency bug.
That is why QA should ask two questions, not one. The first is the one we always ask: "Does this feature work?" The second is the one that catches these bugs: "What happens when multiple actors try to change the same thing at the same time?"
Good SDET work is not only proving that the system works. It is discovering the conditions under which the system stops being correct, and making those conditions visible before production does.
What was the last concurrency bug your team shipped, and which check would have caught it earlier? I'd like to compare notes in the comments.
Top comments (0)