Designing a Railway Booking System Where Double-Booking Is Not an Acceptable Failure
Most railway booking projects begin with a relatively simple idea:
Find a train → choose a seat → make a payment → confirm the ticket.
The difficult part begins when hundreds or thousands of users attempt to book the same limited inventory at almost exactly the same time.
If two users see the same seat as available and both successfully reserve it, the system has violated one of the most fundamental invariants of a reservation system:
One inventory item cannot be allocated to two passengers for overlapping journeys.
That problem became the central focus of this project.
I built a railway reservation engine from scratch in Go and PostgreSQL, focusing less on recreating the IRCTC interface and more on the backend engineering problems behind a high-contention reservation system.
The project is inspired by railway reservation systems such as IRCTC/PRS, but it is not intended to reproduce their internal implementation.
The Problem
Imagine a train with 100 seats.
At 10:00:00.000, seat A1 is available.
Two users send requests:
User A → Book A1
User B → Book A1
Both requests reach the backend nearly simultaneously.
A naïve implementation might perform:
SELECT *
FROM seats
WHERE id = 'A1'
AND status = 'AVAILABLE';
followed by:
UPDATE seats
SET status = 'BOOKED'
WHERE id = 'A1';
The problem is that both transactions can potentially observe the seat before either transaction has committed its update.
The result can become:
User A → A1
User B → A1
The application has now sold the same inventory twice.
This isn't merely a theoretical problem.
Reservation systems are fundamentally concurrency-control systems.
The UI, REST API and payment integration are secondary to one question:
What happens when multiple valid requests compete for the same inventory simultaneously?
Project Goal
The goal of this project was to build a reservation engine with a strong correctness invariant:
The same inventory cannot be allocated to overlapping journeys.
Instead of relying on a single application-level check, I designed multiple layers of protection:
HTTP Request
│
▼
Reservation API
│
▼
Idempotency Layer
│
▼
Allocation Engine
│
┌──────────┴──────────┐
▼ ▼
Candidate Discovery Lock Ordering
│ │
└──────────┬──────────┘
▼
PostgreSQL Transaction
│
▼
Re-validation + Allocation
│
▼
Database Constraint
│
▼
Inventory State
The important design principle is that the database remains the final authority.
Technology Stack
Backend
- Go 1.22+
net/http- PostgreSQL 15+
- SQL migrations
- PostgreSQL transactions
- PostgreSQL row locking
- GiST exclusion constraints
Concurrency & Reliability
SELECT ... FOR UPDATE- Canonical lock ordering
- Bounded retries
- Jittered retries
- Idempotency keys
- Transactional outbox
FOR UPDATE SKIP LOCKED
Domain Components
- Inventory
- Allocation
- Reservations
- Payments
- Waitlist
- RAC
- Quotas
- Authentication
- RBAC
- Event publishing
The repository currently contains 15 database migrations and separate packages for these major domains.
The Most Important Design Decision: Segment-Based Inventory
One subtle problem with railway inventory is that a seat isn't necessarily unavailable for an entire train journey.
Consider a route:
Delhi → Agra → Gwalior → Bhopal → Nagpur
Suppose passenger A books:
Delhi → Gwalior
The same physical seat could potentially be allocated to passenger B for:
Gwalior → Nagpur
Therefore, inventory isn't simply:
seat = occupied / free
It is actually:
seat + journey segment + time/range
This led to representing allocations using segment ranges.
Conceptually:
Seat A1
Delhi ───── Agra ───── Gwalior ───── Bhopal ───── Nagpur
[──────────────]
Passenger A
[──────────────────────]
Passenger B
These allocations don't overlap.
But this would be invalid:
Seat A1
Delhi ───── Agra ───── Gwalior ───── Bhopal ───── Nagpur
[──────────────────────]
Passenger A
[────────────────]
Passenger B
The database therefore needs to understand range overlap, not simply a boolean seat status.
PostgreSQL as the Final Safety Layer
The most important piece of the system is a PostgreSQL exclusion constraint.
The inventory allocation table contains a constraint equivalent to:
EXCLUDE USING gist (
inventory_id WITH =,
segment_range WITH &&
)
WHERE (status IN ('HELD', 'CONFIRMED'));
This creates a database-level invariant:
Same inventory
+
Overlapping segment
+
Active reservation
↓
REJECT
In other words, even if an application bug somehow causes two concurrent transactions to attempt the same allocation, PostgreSQL remains the final line of defense.
The project explicitly treats this constraint as the ultimate correctness mechanism rather than relying exclusively on application-level checks.
This is one of the most important lessons from the project:
Concurrency correctness should not depend on application code alone.
Why Locking Alone Isn't Enough
A common solution is:
SELECT ...
FOR UPDATE;
Row-level locking is useful, but it isn't the entire solution.
The system uses several mechanisms together.
1. Candidate discovery
First, the allocation engine identifies potentially available inventory.
2. Canonical lock ordering
Multiple inventory records are locked in a deterministic order.
For example:
Seat 1
Seat 2
Seat 3
Seat 4
rather than allowing transactions to acquire locks in arbitrary order.
This helps reduce deadlock scenarios.
3. Re-validation
After acquiring locks, availability is checked again.
This is important because the world may have changed between:
candidate discovery
and:
lock acquisition
4. Allocation
Only after the candidate has survived those checks does the system create the allocation.
5. Database constraint
Finally, PostgreSQL independently enforces the no-overlap invariant.
The allocation flow therefore becomes:
DISCOVER
↓
LOCK
↓
RE-VALIDATE
↓
RANK
↓
ALLOCATE
↓
DATABASE CONSTRAINT
The implementation follows this general flow inside the allocation engine.
The Reservation State Machine
A reservation isn't simply:
BOOKED
There are several states involved.
The system models the lifecycle as:
CREATED
│
▼
HELD
│
▼
PAYMENT_PENDING
│
▼
CONFIRMED
with failure paths:
HELD ─────────────→ EXPIRED
│
└───────────────→ CANCELLED
PAYMENT_PENDING ──→ FAILED
This distinction is critical.
For example, imagine a user selects a seat but doesn't complete payment.
If the system immediately treats that seat as permanently booked, inventory becomes unavailable indefinitely.
Instead, the reservation can enter a temporary hold state.
A background worker periodically finds expired holds and releases the inventory.
Temporary Holds
The system therefore separates:
Available
Held
Confirmed
Expired
Cancelled
This gives the booking flow a much more realistic lifecycle.
For example:
User starts booking
│
▼
Seat HELD
│
├──── Payment succeeds ────→ CONFIRMED
│
└──── Timeout ─────────────→ EXPIRED
The ExpireWorker periodically sweeps for expired reservations.
This is implemented as part of the reservation package.
Idempotency: Handling Retries Safely
Another subtle problem appears when clients retry requests.
Suppose a user clicks:
Confirm Booking
and the request reaches the server successfully.
But the network connection dies before the client receives the response.
The client doesn't know whether the booking succeeded.
So it retries:
POST /reservations
Without idempotency, the backend could create a second reservation.
The system therefore supports durable idempotency keys.
Conceptually:
Idempotency-Key: 8f2c...
The server records the operation associated with that key.
If the same request arrives again:
same key
↓
existing operation
↓
return previous result
rather than executing the booking again.
This makes retries safe.
Payment Is Deliberately Decoupled From Inventory Locks
Another important architectural decision was separating payment processing from the inventory transaction.
A dangerous architecture would look like:
BEGIN TRANSACTION
Lock seat
Call payment provider
↓
Wait...
Payment response
↓
Commit
COMMIT
This is problematic because external payment systems can be slow or unavailable.
Holding database locks while waiting for an external service increases contention.
Instead, the system separates:
Inventory transaction
│
▼
Reservation / Payment Intent
│
▼
External payment provider
│
▼
Webhook
│
▼
Confirm reservation
The payment package therefore handles payment intents and idempotent webhook processing without unnecessarily coupling external payment latency to inventory locks.
Transactional Outbox
Once a reservation is confirmed, the system may need to publish events:
ReservationConfirmed
PaymentCompleted
ReservationCancelled
A classic distributed-systems problem appears here.
Suppose the application does:
UPDATE reservation
INSERT event into message broker
and the database update succeeds but the broker operation fails.
Now the reservation exists but the event was never published.
The system uses the transactional outbox pattern.
Instead of directly publishing an event:
DB Transaction
├── Update reservation
└── Insert outbox event
Both operations happen in the same database transaction.
A separate worker then publishes events from the outbox.
The project uses PostgreSQL's SKIP LOCKED mechanism for worker-side claiming.
Conceptually:
PostgreSQL
│
┌─────────┴─────────┐
│ │
reservation outbox
│ │
└─────────┬─────────┘
│
▼
Outbox Worker
│
▼
Event System
This makes event publication much more resilient to process crashes.
Waitlist
Railway inventory isn't always simply:
AVAILABLE
or
SOLD
When inventory is exhausted, passengers may enter a waitlist.
The project implements a priority-ordered waitlist.
Conceptually:
Train full
Passenger A → WL 1
Passenger B → WL 2
Passenger C → WL 3
When inventory becomes available:
Cancellation
│
▼
Available inventory
│
▼
Promotion Worker
│
▼
Waitlist candidate
│
▼
Allocation Engine
An important architectural choice is that waitlist promotion does not implement a completely separate seat-allocation algorithm.
It reuses the existing allocation engine.
That means the same concurrency guarantees apply to:
Normal booking
Waitlist promotion
RAC promotion
This avoids having multiple implementations of the most critical piece of business logic.
RAC
The project also models Reservation Against Cancellation (RAC) as a separate domain.
Instead of treating RAC as simply another waitlist status, it has its own queue and promotion flow.
The architecture follows the same general principle:
Cancellation
│
▼
RAC Queue
│
▼
Promotion Worker
│
▼
Allocation Engine
Again, the allocation engine remains the central authority for assigning inventory.
Quotas
Railway reservations can also have inventory divided into different quotas.
The project models quota-scoped inventory assignment and eventual release back to general inventory.
Conceptually:
Inventory
│
┌────────────┼────────────┐
▼ ▼ ▼
General Quota A Quota B
│ │ │
└────────────┴────────────┘
│
Release Policy
│
▼
General Inventory
This introduces another interesting consistency problem:
Inventory isn't merely "available or unavailable"; it can belong to different allocation pools with different release rules.
API Architecture
The backend exposes REST endpoints for:
- Train availability
- PNR lookup
- Reservation creation
- Reservation confirmation
- Reservation cancellation
- Payment webhooks
- Administrative operations
- RBAC/role management
The public read path is separated from authenticated booking operations, while administrative operations require authorization and permission checks.
A simplified API structure looks like:
/v1
├── /trains
│ └── /{id}/availability
│
├── /pnr
│ └── /{pnr}
│
├── /reservations
│ ├── POST
│ ├── GET
│ ├── /confirm
│ └── /cancel
│
├── /payments
│ └── /webhook
│
└── /admin
└── RBAC
Authentication and RBAC
The project intentionally keeps authentication relatively small.
Requests can carry bearer tokens, while administrative actions additionally go through role/permission checks.
An important design choice is that authorization isn't treated purely as middleware.
Policy checks also happen inside handlers.
This makes authorization closer to the actual business operation rather than relying entirely on route-level protection.
Background Workers
Some operations should not happen synchronously inside a user's HTTP request.
The server therefore wires several background workers:
Expiry Worker
│
└── Expire temporary holds
Outbox Worker
│
└── Publish pending events
Waitlist Worker
│
└── Promote eligible passengers
RAC Worker
│
└── Promote RAC passengers
This allows the API layer to remain focused on request/response operations while asynchronous domain processing happens independently.
Testing the Actual Concurrency Problem
One of the most important parts of the project is that concurrency correctness isn't tested purely with mocks.
The repository contains a dedicated concurrency test suite that runs against a real PostgreSQL instance.
This is important because the behavior being tested depends on PostgreSQL itself:
Row locks
Exclusion constraints
Transactions
SKIP LOCKED
Concurrent transactions
Mocking these mechanisms would not actually prove that PostgreSQL behaves correctly under contention.
The repository documents 11 concurrency/integration scenarios, including a test that exposed an actual bug during development.
The Testing Philosophy
The project deliberately separates two kinds of testing.
Unit tests
Used for deterministic domain behavior:
go test ./internal/...
These don't require a database.
Concurrency integration tests
Used for database behavior:
go test ./tests/concurrency/... -v
These require a real PostgreSQL instance.
This distinction matters.
A unit test can prove:
allocateSeat()
returns the expected result.
It cannot prove that:
100 concurrent transactions
cannot violate a PostgreSQL locking invariant.
For that, the actual database needs to participate in the test.
Architecture at a Glance
The complete system can be thought of as:
CLIENT
│
▼
REST API
│
┌─────────┴─────────┐
│ │
Auth/RBAC Idempotency
│ │
└─────────┬─────────┘
▼
Reservation
Coordinator
│
┌───────────┼───────────┐
▼ ▼ ▼
Inventory Payment Rules
│ │ │
└─────┬─────┴───────────┘
▼
PostgreSQL
│
┌────────────┼────────────┐
▼ ▼ ▼
Outbox Waitlist RAC
Worker Worker Worker
│ │ │
└────────────┴────────────┘
│
▼
Event Processing
The database is at the center because correctness depends heavily on transactional guarantees.
Project Structure
The repository is intentionally divided by domain rather than putting all business logic into one large service.
internal/
├── inventory/
├── allocation/
├── rules/
├── reservation/
├── idempotency/
├── payment/
├── outbox/
├── waitlist/
├── rac/
├── quota/
├── auth/
├── rbac/
├── api/
└── platform/
The repository also keeps:
docs/
├── architecture.md
├── concurrency.md
├── status.md
└── adr/
alongside:
tests/concurrency/
This separation makes the repository itself part of the design documentation rather than simply being a collection of source files.
Architecture Decision Records
The project contains individual ADRs documenting major architectural trade-offs.
This was intentional.
In real systems, architecture isn't simply:
"I chose PostgreSQL."
There is usually a question behind every decision:
Why PostgreSQL?
Why database-level exclusion?
Why row locking?
Why canonical lock ordering?
Why idempotency?
Why an outbox?
Why SKIP LOCKED?
Why a state machine?
Why separate payment from inventory?
Writing these decisions down makes the system easier to reason about and makes trade-offs explicit.
What I Learned
The biggest lesson from building this system was that high-concurrency backend engineering is fundamentally about invariants.
It is tempting to think about a booking system as:
API → Database → Response
But the more useful model is:
Invariant
↓
Concurrency model
↓
Transaction boundaries
↓
Database constraints
↓
Recovery mechanisms
↓
Tests
For example:
Invariant
A seat cannot be allocated twice for overlapping journey segments.
Concurrency mechanism
Locks + canonical ordering + transactions.
Database guarantee
PostgreSQL exclusion constraint.
Recovery
Bounded/jittered retries.
Verification
Real PostgreSQL concurrency tests.
This approach changes how I think about backend architecture.
What I Would Build Next
The current implementation intentionally focuses on correctness rather than claiming production-scale performance.
There are several areas still left for future work:
- OpenTelemetry tracing
- Prometheus metrics
- Docker / Docker Compose setup
- k6 load-testing scenarios
- Formal throughput benchmarks
These are explicitly tracked as unfinished areas in the repository.
That distinction is important.
The current concurrency suite demonstrates correctness under contention, but it does not claim a particular production throughput or capacity figure.
The next stage would therefore be to measure the system rather than simply assume it scales.
Why This Project Exists
I didn't want to build another CRUD-heavy railway booking application where the main engineering challenge was:
POST /book
Instead, I wanted to investigate a question that appears simple but becomes difficult under concurrent load:
How do you safely allocate scarce, segment-based inventory when many independent requests arrive simultaneously?
That question naturally led into:
- database transactions
- row-level locking
- deadlock avoidance
- exclusion constraints
- idempotency
- state machines
- payment boundaries
- transactional outbox
- background workers
- waitlists
- RAC
- quota management
- concurrency testing
The resulting system is therefore less about building an "IRCTC clone" and more about building a concurrency-correct reservation engine.
Final Takeaway
A railway booking system looks like a straightforward application until you introduce concurrency.
Then the problem becomes:
MANY USERS
│
▼
LIMITED INVENTORY
│
▼
CONCURRENT REQUESTS
│
▼
TRANSACTION CONTROL
│
▼
DATABASE INVARIANTS
│
▼
CORRECT RESULT
The most important architectural decision in this project is that correctness is not entrusted to a single if seat.available check.
Instead, correctness is defended through multiple layers:
Idempotency
+
Canonical locking
+
Re-validation
+
Transactions
+
Database exclusion constraints
+
Retry handling
+
Concurrency integration tests
The result is a backend reservation engine designed around a very specific principle:
When inventory is scarce and requests are concurrent, correctness has to be designed into the system — not assumed from the application code.
Research & Architectural References
Architecture disclaimer: This project is an independent engineering implementation inspired by publicly available information about Indian Railways' Passenger Reservation System (PRS), CONCERT and CRIS. It does not claim to reproduce IRCTC/CRIS's proprietary production architecture, implementation, infrastructure, algorithms, or internal systems.
The architecture and engineering problems explored in this project were informed by publicly available government, academic, and institutional material describing the evolution and architecture of Indian Railways' computerized reservation systems.
1. CRIS — CONCERT / Passenger Reservation System
The Centre for Railway Information Systems (CRIS) describes CONCERT as a mission-critical online transaction-processing application built around a distributed database model and a three-tier client-server architecture.
This provided the primary architectural context for treating railway reservation as a distributed, high-concurrency transaction-processing problem rather than as a conventional CRUD application.
Relevant concepts:
- Distributed reservation infrastructure
- Online Transaction Processing (OLTP)
- Three-tier architecture
- Distributed databases
- Geographically distributed reservation infrastructure
- Mission-critical transaction processing
Reference:
CRIS — Passenger Reservation System / CONCERT
2. CAG of India — Computerised Passenger Reservation System
The Comptroller and Auditor General of India (CAG) published an audit report covering the Computerised Passenger Reservation System of Indian Railways.
The report provides historical information about the evolution of the reservation system, including the earlier IMPRESS system, its deployment across multiple PRS locations, and the subsequent evolution toward CONCERT.
It is particularly useful for understanding the historical distributed architecture of Indian Railways' reservation infrastructure.
Relevant concepts:
- Evolution from IMPRESS to CONCERT
- Distributed PRS locations
- Reservation database architecture
- Networked reservation terminals
- Transaction processing
- Capacity and operational constraints
Reference:
CAG of India — Computerised Passenger Reservation System of Indian Railways
3. Computerized Passenger Reservation System for Indian Railways — System Architecture
Seema Agarwal's research paper, "Computerized Passenger Reservation System for Indian Railways — Its Development and System Architecture," examines the evolution and architecture of India's computerized railway reservation system.
The paper discusses the transition from earlier reservation platforms toward CONCERT and describes the architectural characteristics of the system, including its three-tier client-server model.
This research was particularly useful for understanding how a railway reservation system evolved from individual reservation systems into a connected, distributed reservation environment.
Relevant concepts:
- PRS evolution
- IMPRESS
- CONCERT
- Three-tier architecture
- Distributed reservation systems
- Centralized vs. distributed processing
References:
ResearchGate — Computerized Passenger Reservation System for Indian Railways
4. IIM Ahmedabad — Passenger Reservation System of Indian Railways
The Indian Institute of Management Ahmedabad published a case study titled:
"Management of Large IT Projects: The Passenger Reservation System of Indian Railways."
The study examines the PRS as a large-scale information technology project and provides historical architectural context around its distributed database infrastructure and geographically distributed reservation terminals.
This was useful for understanding the reservation system not merely as a ticket-booking application, but as a large-scale distributed information system with significant operational and transaction-processing requirements.
Relevant concepts:
- Large-scale IT systems
- Distributed databases
- Geographically distributed infrastructure
- Transaction processing
- System evolution
- Operational scalability
Reference:
IIM Ahmedabad — Management of Large IT Projects: The Passenger Reservation System of Indian Railways
5. Modernization of Passenger Reservation System
The paper "Modernization of Passenger Reservation System: Indian Railways' Dilemma", published in the Journal of Information Technology, examines the challenges involved in modernizing a large, established railway reservation system.
The paper is useful for understanding the architectural tension between:
- Existing reliable infrastructure
- New functional requirements
- Changing user expectations
- Legacy system constraints
- Modernization of large-scale information systems
These considerations influenced the decision to build this project as a modern independent reservation engine rather than attempting to replicate the historical architecture directly.
Reference:
DOI:
10.1057/palgrave.jit.2000112
6. CRIS — New PRS in RDBMS Platform and NGeT
A historical CRIS presentation discussing the modernization of the Passenger Reservation System provides additional technical context around:
- CONCERT
- RDBMS migration
- Distributed architecture
- Transaction routing
- Reliability
- Next Generation e-Ticketing (NGeT)
- Reservation system modernization
This material is particularly useful for understanding the transition from older reservation infrastructure toward newer database and Internet-based architectures.
Note: This is historical material and should not be interpreted as a specification of the current production IRCTC/CRIS architecture.
Reference:
CRIS — New PRS in RDBMS Platform and NGeT
What These References Contributed to This Project
The purpose of researching these systems was not to reproduce IRCTC's internal implementation.
Instead, the research helped identify the fundamental engineering characteristics of railway reservation systems:
text
Railway Reservation
│
▼
Distributed System
│
┌─────────────┼─────────────┐
▼ ▼ ▼
High Volume Limited Distributed
Transactions Inventory Infrastructure
│ │ │
└─────────────┼─────────────┘
▼
Concurrency Control
│
▼
Transaction Integrity
│
▼
Inventory Correctness
Top comments (0)