DEV Community

Cover image for The Amazon Great Indian Festival and Flipkart Big Billion Days Started: How AWS EKS Handles Millions of Users
Rahul R
Rahul R

Posted on

The Amazon Great Indian Festival and Flipkart Big Billion Days Started: How AWS EKS Handles Millions of Users

The Amazon Great Indian Festival and Flipkart Big Billion Days have started.

Millions of users are visiting these platforms, searching for products, checking offers, adding items to carts, and placing orders—sometimes at almost the same moment.

But behind every search, product page, cart update, and payment, there are thousands of backend requests being processed continuously.

So the interesting question is:

How can an application handle millions of users and huge traffic spikes without bringing the system down?

One possible approach is to build a highly scalable architecture using Amazon EKS (Elastic Kubernetes Service) along with load balancing, Kubernetes autoscaling, caching, databases, and other supporting services.

The important part isn't simply running an application inside EKS.

The real challenge is making the infrastructure react to changing traffic.

When traffic increases, the application needs more capacity.

When traffic decreases, that extra capacity should no longer be required.

This is where Kubernetes autoscaling becomes interesting.

Note: Amazon and Flipkart's actual production architectures are proprietary. This article describes a conceptual architecture for understanding how an EKS-based platform can handle large-scale e-commerce traffic.


The problem: traffic doesn't arrive evenly

On a normal day, an e-commerce application might receive a relatively predictable amount of traffic.

For example:

10,000 requests/minute
Enter fullscreen mode Exit fullscreen mode

The infrastructure can be sized around this workload.

But when a major sale begins, the traffic pattern can change very quickly.

It might look conceptually like:

Normal
  |
  v
10K requests/min
  |
  v
50K
  |
  v
200K
  |
  v
500K+
Enter fullscreen mode Exit fullscreen mode

The exact traffic numbers are hypothetical, but the engineering problem is real:

Traffic can increase much faster than normal capacity requirements.

Running peak capacity all day would be expensive.

Running too little capacity during the spike could result in:

  • High latency
  • Failed requests
  • Timeouts
  • Poor user experience
  • Overloaded backend services

So the infrastructure needs to be elastic.


A simplified EKS architecture

A simplified architecture could look something like this:

                    Users
                      |
                      v
                Route 53
                      |
                      v
                  CloudFront
                      |
                      v
           Application Load Balancer
                      |
                      v
                EKS Cluster
                      |
          +-----------+-----------+
          |           |           |
          v           v           v
      Product       Cart        Search
       Pods         Pods         Pods
          |           |           |
          +-----------+-----------+
                      |
             +--------+--------+
             |                 |
             v                 v
           Cache           Database
Enter fullscreen mode Exit fullscreen mode

The idea is simple:

Don't make one server handle everything.

Instead, distribute the workload across multiple application instances.


Step 1: A user sends a request

Imagine someone searches for a smartphone during the sale.

The request can travel through several layers:

User
 ↓
Route 53
 ↓
CloudFront
 ↓
Application Load Balancer
 ↓
EKS
 ↓
Application Pod
Enter fullscreen mode Exit fullscreen mode

The load balancer helps distribute incoming requests across healthy application targets.

Instead of sending every request to one application instance, traffic can be spread across multiple Pods.

For example:

             Load Balancer
                  |
        +---------+---------+
        |         |         |
        v         v         v
      Pod 1     Pod 2     Pod 3
Enter fullscreen mode Exit fullscreen mode

If traffic increases, additional Pods can be added.


Step 2: Kubernetes manages the Pods

Inside EKS, the application runs as containers managed by Kubernetes.

Suppose the application starts with:

Pod 1
Pod 2
Pod 3
Pod 4
Pod 5
Enter fullscreen mode Exit fullscreen mode

This might be enough during normal traffic.

But then a flash sale begins.

More users arrive.

Instead of manually creating Pods one by one, Kubernetes can use autoscaling policies to adjust the number of replicas.

Before sale:

5 Pods
Enter fullscreen mode Exit fullscreen mode

During higher traffic:

5 Pods
 ↓
10 Pods
 ↓
20 Pods
 ↓
40 Pods
Enter fullscreen mode Exit fullscreen mode

The actual number depends on the application's resource requirements and scaling configuration.


Step 3: Horizontal Pod Autoscaler

This is where Horizontal Pod Autoscaler (HPA) comes into the picture.

HPA can monitor metrics such as CPU and memory utilization. It can also be configured to use custom or application-level metrics.

For example:

Minimum replicas: 5
Maximum replicas: 100

Target CPU: 60%
Enter fullscreen mode Exit fullscreen mode

During normal traffic:

5 Pods
Enter fullscreen mode Exit fullscreen mode

As the workload increases:

CPU utilization increases
        ↓
HPA evaluates metrics
        ↓
Desired replica count increases
        ↓
New Pods are created
Enter fullscreen mode Exit fullscreen mode

The important concept here is horizontal scaling.

Instead of making one server bigger:

Small Server
     ↓
Huge Server
Enter fullscreen mode Exit fullscreen mode

we can run more application instances:

Pod
Pod
Pod
Pod
Pod
Pod
Enter fullscreen mode Exit fullscreen mode

This is especially useful for stateless services that can handle requests independently.


But what if there isn't enough compute?

There is an important limitation.

Suppose HPA decides that the application needs 50 more Pods.

But the existing EKS nodes are already full.

Node 1 → Full
Node 2 → Full
Node 3 → Full
Enter fullscreen mode Exit fullscreen mode

Kubernetes knows that additional Pods are needed, but there isn't enough capacity to schedule them.

This is where node autoscaling becomes important.


Step 4: Node autoscaling with Karpenter

An EKS cluster can use Karpenter to help provision compute capacity when Pods cannot be scheduled because the existing nodes don't have enough resources.

The flow can look like:

Traffic increases
       ↓
HPA requests more Pods
       ↓
Pods cannot be scheduled
       ↓
Insufficient node capacity
       ↓
Karpenter provisions suitable compute
       ↓
New node becomes available
       ↓
Pending Pods are scheduled
Enter fullscreen mode Exit fullscreen mode

Now we have two different scaling levels:

Application level
        ↓
HPA
        ↓
More Pods

Infrastructure level
        ↓
Karpenter
        ↓
More compute capacity
Enter fullscreen mode Exit fullscreen mode

This is an important distinction.

HPA scales the application.

Node autoscaling provides the infrastructure needed to run that application.


Step 5: What happens during a traffic spike?

Let's put the pieces together.

Imagine traffic suddenly increases:

                    Massive Traffic
                          |
                          v
                 Application Load
                    Balancer
                          |
             +------------+------------+
             |            |            |
             v            v            v
           Pod 1        Pod 2        Pod 3
             |            |            |
             +------------+------------+
                          |
                          v
                         HPA
                          |
                    More Pods?
                          |
                         Yes
                          |
                          v
                  Node capacity?
                          |
                    +-----+-----+
                    |           |
                   Yes          No
                    |           |
                    |           v
                    |       Karpenter
                    |           |
                    |           v
                    |     More compute
                    |           |
                    +-----------+
                          |
                          v
                    More Pods
Enter fullscreen mode Exit fullscreen mode

The result is an architecture that can adapt to changing demand.


Step 6: Why caching is important

There is another problem.

Even if EKS can run hundreds of Pods, what happens if all those Pods send requests directly to the database?

You could end up with:

100 Pods
   |
   v
Database
   |
   X
Bottleneck
Enter fullscreen mode Exit fullscreen mode

This is why caching can be extremely important.

Frequently requested information can potentially be served from a cache instead of repeatedly querying the database.

For example:

User
 ↓
Application
 ↓
Cache
 ↓
Response
Enter fullscreen mode Exit fullscreen mode

Services such as Amazon ElastiCache can be used for distributed caching.

For an e-commerce application, caching can be useful for data such as product information, sessions, or other information where the application's consistency requirements allow it.

The key idea is:

Scaling the application layer doesn't automatically mean the database can handle unlimited traffic.

Every layer needs to be considered.


Step 7: Different services can scale independently

A large e-commerce application usually isn't one giant application.

It may contain different services:

                 EKS
                  |
       +----------+----------+
       |          |          |
       v          v          v
    Product     Cart       Search
    Service     Service    Service
       |          |          |
       v          v          v
     Cache      Database   Search
Enter fullscreen mode Exit fullscreen mode

During a sale, search traffic might increase dramatically.

Cart traffic might behave differently.

Payment traffic might have another pattern.

Therefore, each service can have its own resource requirements and scaling policies.

For example:

Product Service
10 → 50 Pods

Search Service
20 → 100 Pods

Cart Service
5 → 30 Pods
Enter fullscreen mode Exit fullscreen mode

These numbers are only examples.

The important idea is that not every service needs to scale by the same amount.


Step 8: Scaling down after the sale

This is the part people often forget.

Autoscaling isn't only about adding capacity.

It's also about removing capacity when it is no longer needed.

Suppose the sale ends.

Traffic starts dropping:

500K requests/min
       ↓
200K
       ↓
50K
       ↓
10K
Enter fullscreen mode Exit fullscreen mode

HPA can reduce the number of application replicas according to its configuration.

For example:

100 Pods
   ↓
70 Pods
   ↓
40 Pods
   ↓
15 Pods
Enter fullscreen mode Exit fullscreen mode

Once there is unused node capacity, node autoscaling can also reduce infrastructure where appropriate.

So the system can move from:

Peak:

30 Nodes
100 Pods
Enter fullscreen mode Exit fullscreen mode

toward something closer to:

Normal:

5 Nodes
15 Pods
Enter fullscreen mode Exit fullscreen mode

The exact behavior depends on the autoscaling configuration, workload constraints, Pod disruption rules, and other factors.


Autoscaling isn't instant

There is one important thing to understand.

Autoscaling is not magic.

If traffic suddenly changes from:

10K → 500K requests/minute
Enter fullscreen mode Exit fullscreen mode

the infrastructure doesn't instantly create everything required.

There is a reaction time.

Metrics need to be collected.

HPA needs to evaluate them.

New Pods need to start.

Container images may need to be pulled.

If additional nodes are required, those nodes need to become available.

Because of this, predictable large-scale events often require capacity planning and pre-scaling, rather than relying entirely on reactive autoscaling.


Pre-scaling before a major sale

For a known event, engineers can increase the baseline capacity before traffic arrives.

For example:

Normal day

5 Pods
3 Nodes
Enter fullscreen mode Exit fullscreen mode

Before the sale:

30 Pods
10 Nodes
Enter fullscreen mode Exit fullscreen mode

During the peak:

100 Pods
30 Nodes
Enter fullscreen mode Exit fullscreen mode

Autoscaling can then handle additional fluctuations.

This is a much safer approach for predictable traffic spikes because the system doesn't have to build all of its capacity after the traffic has already arrived.


The database is still a critical bottleneck

One of the biggest mistakes when thinking about Kubernetes scaling is:

"If I add more Pods, I can handle unlimited traffic."

Not necessarily.

Imagine:

200 Pods
   |
   v
One database
   |
   X
Overloaded
Enter fullscreen mode Exit fullscreen mode

The application layer may scale perfectly while the database becomes the bottleneck.

Large-scale systems therefore need to think about:

  • Read replicas
  • Connection pooling
  • Caching
  • Database scaling
  • Index optimization
  • Partitioning
  • Asynchronous processing
  • Queues
  • Rate limiting
  • CDN caching
  • Backpressure
  • Graceful degradation

EKS is an important part of the architecture, but it isn't the entire architecture.


What happens behind a simple "Buy Now"?

When you click Buy Now, the request might trigger multiple backend operations.

Conceptually:

User
  |
  v
Load Balancer
  |
  v
Cart Service
  |
  +------> Inventory
  |
  +------> Pricing
  |
  +------> Order Service
  |
  +------> Payment
  |
  v
Order Confirmation
Enter fullscreen mode Exit fullscreen mode

Each of these components can have different scaling and reliability requirements.

The user sees:

Order Confirmed ✓
Enter fullscreen mode Exit fullscreen mode

But behind that simple message, multiple distributed services may have communicated with each other.


The bigger lesson

When millions of users arrive during a major shopping sale, the solution isn't simply:

"Add more servers."

The real engineering challenge is:

How do we automatically provide the right amount of capacity at the right time?

Too little capacity:

High latency
Timeouts
Errors
Failed requests
Enter fullscreen mode Exit fullscreen mode

Too much capacity:

Unused resources
Higher cloud costs
Enter fullscreen mode Exit fullscreen mode

The goal is elasticity.

Traffic increases
       ↓
More Pods
       ↓
More compute when required
       ↓
Traffic is distributed
       ↓
Traffic decreases
       ↓
Pods scale down
       ↓
Unused capacity can be removed
Enter fullscreen mode Exit fullscreen mode

That's the interesting part of EKS.

It's not just Kubernetes running in AWS.

It's the combination of container orchestration, load balancing, autoscaling, compute capacity, caching, databases, observability, and application architecture working together.


Final architecture

Putting everything together:

                         USERS
                           |
                           v
                      Route 53
                           |
                           v
                       CloudFront
                           |
                           v
                Application Load Balancer
                           |
                           v
                    +-------------+
                    | EKS Cluster |
                    +-------------+
                           |
             +-------------+-------------+
             |             |             |
             v             v             v
          Product        Search        Cart
           Pods           Pods          Pods
             |             |             |
             +-------------+-------------+
                           |
                          HPA
                           |
                  More Pods required?
                           |
                           v
                     Node capacity?
                       /        \
                     Yes         No
                      |           |
                      |       Karpenter
                      |           |
                      |           v
                      |     More compute
                      |           |
                      +-----------+
                           |
                           v
                    Cache / Database
Enter fullscreen mode Exit fullscreen mode

The next time you see a massive sale and thousands of products being sold within seconds, remember that there is a huge amount of engineering behind that experience.

Traffic comes in.
Pods scale out.
Compute capacity grows.
Requests are distributed.
Caches reduce database pressure.
And when traffic goes down, the infrastructure can scale back.

That's the real power of designing applications for elasticity instead of fixed capacity.

Top comments (0)