The Amazon Great Indian Festival and Flipkart Big Billion Days have started.
Millions of users are visiting these platforms, searching for products, checking offers, adding items to carts, and placing orders—sometimes at almost the same moment.
But behind every search, product page, cart update, and payment, there are thousands of backend requests being processed continuously.
So the interesting question is:
How can an application handle millions of users and huge traffic spikes without bringing the system down?
One possible approach is to build a highly scalable architecture using Amazon EKS (Elastic Kubernetes Service) along with load balancing, Kubernetes autoscaling, caching, databases, and other supporting services.
The important part isn't simply running an application inside EKS.
The real challenge is making the infrastructure react to changing traffic.
When traffic increases, the application needs more capacity.
When traffic decreases, that extra capacity should no longer be required.
This is where Kubernetes autoscaling becomes interesting.
Note: Amazon and Flipkart's actual production architectures are proprietary. This article describes a conceptual architecture for understanding how an EKS-based platform can handle large-scale e-commerce traffic.
The problem: traffic doesn't arrive evenly
On a normal day, an e-commerce application might receive a relatively predictable amount of traffic.
For example:
10,000 requests/minute
The infrastructure can be sized around this workload.
But when a major sale begins, the traffic pattern can change very quickly.
It might look conceptually like:
Normal
|
v
10K requests/min
|
v
50K
|
v
200K
|
v
500K+
The exact traffic numbers are hypothetical, but the engineering problem is real:
Traffic can increase much faster than normal capacity requirements.
Running peak capacity all day would be expensive.
Running too little capacity during the spike could result in:
- High latency
- Failed requests
- Timeouts
- Poor user experience
- Overloaded backend services
So the infrastructure needs to be elastic.
A simplified EKS architecture
A simplified architecture could look something like this:
Users
|
v
Route 53
|
v
CloudFront
|
v
Application Load Balancer
|
v
EKS Cluster
|
+-----------+-----------+
| | |
v v v
Product Cart Search
Pods Pods Pods
| | |
+-----------+-----------+
|
+--------+--------+
| |
v v
Cache Database
The idea is simple:
Don't make one server handle everything.
Instead, distribute the workload across multiple application instances.
Step 1: A user sends a request
Imagine someone searches for a smartphone during the sale.
The request can travel through several layers:
User
↓
Route 53
↓
CloudFront
↓
Application Load Balancer
↓
EKS
↓
Application Pod
The load balancer helps distribute incoming requests across healthy application targets.
Instead of sending every request to one application instance, traffic can be spread across multiple Pods.
For example:
Load Balancer
|
+---------+---------+
| | |
v v v
Pod 1 Pod 2 Pod 3
If traffic increases, additional Pods can be added.
Step 2: Kubernetes manages the Pods
Inside EKS, the application runs as containers managed by Kubernetes.
Suppose the application starts with:
Pod 1
Pod 2
Pod 3
Pod 4
Pod 5
This might be enough during normal traffic.
But then a flash sale begins.
More users arrive.
Instead of manually creating Pods one by one, Kubernetes can use autoscaling policies to adjust the number of replicas.
Before sale:
5 Pods
During higher traffic:
5 Pods
↓
10 Pods
↓
20 Pods
↓
40 Pods
The actual number depends on the application's resource requirements and scaling configuration.
Step 3: Horizontal Pod Autoscaler
This is where Horizontal Pod Autoscaler (HPA) comes into the picture.
HPA can monitor metrics such as CPU and memory utilization. It can also be configured to use custom or application-level metrics.
For example:
Minimum replicas: 5
Maximum replicas: 100
Target CPU: 60%
During normal traffic:
5 Pods
As the workload increases:
CPU utilization increases
↓
HPA evaluates metrics
↓
Desired replica count increases
↓
New Pods are created
The important concept here is horizontal scaling.
Instead of making one server bigger:
Small Server
↓
Huge Server
we can run more application instances:
Pod
Pod
Pod
Pod
Pod
Pod
This is especially useful for stateless services that can handle requests independently.
But what if there isn't enough compute?
There is an important limitation.
Suppose HPA decides that the application needs 50 more Pods.
But the existing EKS nodes are already full.
Node 1 → Full
Node 2 → Full
Node 3 → Full
Kubernetes knows that additional Pods are needed, but there isn't enough capacity to schedule them.
This is where node autoscaling becomes important.
Step 4: Node autoscaling with Karpenter
An EKS cluster can use Karpenter to help provision compute capacity when Pods cannot be scheduled because the existing nodes don't have enough resources.
The flow can look like:
Traffic increases
↓
HPA requests more Pods
↓
Pods cannot be scheduled
↓
Insufficient node capacity
↓
Karpenter provisions suitable compute
↓
New node becomes available
↓
Pending Pods are scheduled
Now we have two different scaling levels:
Application level
↓
HPA
↓
More Pods
Infrastructure level
↓
Karpenter
↓
More compute capacity
This is an important distinction.
HPA scales the application.
Node autoscaling provides the infrastructure needed to run that application.
Step 5: What happens during a traffic spike?
Let's put the pieces together.
Imagine traffic suddenly increases:
Massive Traffic
|
v
Application Load
Balancer
|
+------------+------------+
| | |
v v v
Pod 1 Pod 2 Pod 3
| | |
+------------+------------+
|
v
HPA
|
More Pods?
|
Yes
|
v
Node capacity?
|
+-----+-----+
| |
Yes No
| |
| v
| Karpenter
| |
| v
| More compute
| |
+-----------+
|
v
More Pods
The result is an architecture that can adapt to changing demand.
Step 6: Why caching is important
There is another problem.
Even if EKS can run hundreds of Pods, what happens if all those Pods send requests directly to the database?
You could end up with:
100 Pods
|
v
Database
|
X
Bottleneck
This is why caching can be extremely important.
Frequently requested information can potentially be served from a cache instead of repeatedly querying the database.
For example:
User
↓
Application
↓
Cache
↓
Response
Services such as Amazon ElastiCache can be used for distributed caching.
For an e-commerce application, caching can be useful for data such as product information, sessions, or other information where the application's consistency requirements allow it.
The key idea is:
Scaling the application layer doesn't automatically mean the database can handle unlimited traffic.
Every layer needs to be considered.
Step 7: Different services can scale independently
A large e-commerce application usually isn't one giant application.
It may contain different services:
EKS
|
+----------+----------+
| | |
v v v
Product Cart Search
Service Service Service
| | |
v v v
Cache Database Search
During a sale, search traffic might increase dramatically.
Cart traffic might behave differently.
Payment traffic might have another pattern.
Therefore, each service can have its own resource requirements and scaling policies.
For example:
Product Service
10 → 50 Pods
Search Service
20 → 100 Pods
Cart Service
5 → 30 Pods
These numbers are only examples.
The important idea is that not every service needs to scale by the same amount.
Step 8: Scaling down after the sale
This is the part people often forget.
Autoscaling isn't only about adding capacity.
It's also about removing capacity when it is no longer needed.
Suppose the sale ends.
Traffic starts dropping:
500K requests/min
↓
200K
↓
50K
↓
10K
HPA can reduce the number of application replicas according to its configuration.
For example:
100 Pods
↓
70 Pods
↓
40 Pods
↓
15 Pods
Once there is unused node capacity, node autoscaling can also reduce infrastructure where appropriate.
So the system can move from:
Peak:
30 Nodes
100 Pods
toward something closer to:
Normal:
5 Nodes
15 Pods
The exact behavior depends on the autoscaling configuration, workload constraints, Pod disruption rules, and other factors.
Autoscaling isn't instant
There is one important thing to understand.
Autoscaling is not magic.
If traffic suddenly changes from:
10K → 500K requests/minute
the infrastructure doesn't instantly create everything required.
There is a reaction time.
Metrics need to be collected.
HPA needs to evaluate them.
New Pods need to start.
Container images may need to be pulled.
If additional nodes are required, those nodes need to become available.
Because of this, predictable large-scale events often require capacity planning and pre-scaling, rather than relying entirely on reactive autoscaling.
Pre-scaling before a major sale
For a known event, engineers can increase the baseline capacity before traffic arrives.
For example:
Normal day
5 Pods
3 Nodes
Before the sale:
30 Pods
10 Nodes
During the peak:
100 Pods
30 Nodes
Autoscaling can then handle additional fluctuations.
This is a much safer approach for predictable traffic spikes because the system doesn't have to build all of its capacity after the traffic has already arrived.
The database is still a critical bottleneck
One of the biggest mistakes when thinking about Kubernetes scaling is:
"If I add more Pods, I can handle unlimited traffic."
Not necessarily.
Imagine:
200 Pods
|
v
One database
|
X
Overloaded
The application layer may scale perfectly while the database becomes the bottleneck.
Large-scale systems therefore need to think about:
- Read replicas
- Connection pooling
- Caching
- Database scaling
- Index optimization
- Partitioning
- Asynchronous processing
- Queues
- Rate limiting
- CDN caching
- Backpressure
- Graceful degradation
EKS is an important part of the architecture, but it isn't the entire architecture.
What happens behind a simple "Buy Now"?
When you click Buy Now, the request might trigger multiple backend operations.
Conceptually:
User
|
v
Load Balancer
|
v
Cart Service
|
+------> Inventory
|
+------> Pricing
|
+------> Order Service
|
+------> Payment
|
v
Order Confirmation
Each of these components can have different scaling and reliability requirements.
The user sees:
Order Confirmed ✓
But behind that simple message, multiple distributed services may have communicated with each other.
The bigger lesson
When millions of users arrive during a major shopping sale, the solution isn't simply:
"Add more servers."
The real engineering challenge is:
How do we automatically provide the right amount of capacity at the right time?
Too little capacity:
High latency
Timeouts
Errors
Failed requests
Too much capacity:
Unused resources
Higher cloud costs
The goal is elasticity.
Traffic increases
↓
More Pods
↓
More compute when required
↓
Traffic is distributed
↓
Traffic decreases
↓
Pods scale down
↓
Unused capacity can be removed
That's the interesting part of EKS.
It's not just Kubernetes running in AWS.
It's the combination of container orchestration, load balancing, autoscaling, compute capacity, caching, databases, observability, and application architecture working together.
Final architecture
Putting everything together:
USERS
|
v
Route 53
|
v
CloudFront
|
v
Application Load Balancer
|
v
+-------------+
| EKS Cluster |
+-------------+
|
+-------------+-------------+
| | |
v v v
Product Search Cart
Pods Pods Pods
| | |
+-------------+-------------+
|
HPA
|
More Pods required?
|
v
Node capacity?
/ \
Yes No
| |
| Karpenter
| |
| v
| More compute
| |
+-----------+
|
v
Cache / Database
The next time you see a massive sale and thousands of products being sold within seconds, remember that there is a huge amount of engineering behind that experience.
Traffic comes in.
Pods scale out.
Compute capacity grows.
Requests are distributed.
Caches reduce database pressure.
And when traffic goes down, the infrastructure can scale back.
That's the real power of designing applications for elasticity instead of fixed capacity.

Top comments (0)