Building a Web App: A Beginner's Guide to Cloud Architecture (Part 5 of 5)
In Part 1, our main server died, but our load balancer instantly shifted traffic to a backup (High Availability). In Part 2, 50,000 users rushed our registration page, and our infrastructure automatically cloned itself to handle the massive spike (Scalability). In Part 3, malicious actors tried to bypass our website, but our private networks and firewalls blocked them (Security). In Part 4, international users complained about lag, so we moved our static files to global edge networks to make the app lightning fast (Speed).
We now have a highly available, endlessly scalable, secure, and blazing-fast web application.
But there is one final bottleneck to overcome: Synchronous Processing.
Imagine our 50,000 users are now asked to upload a 2GB video portfolio to complete their registration. When a user clicks "Upload," our web server accepts the file and begins processing it. Because it is working synchronously, that specific web server is now completely locked up. For the next 10 minutes, it cannot handle any other requests.
The user is left staring at a spinning loading wheel. If other users try to reach that same server, they are put on hold. Eventually, the browser times out, the connection drops, and the server crashes from memory exhaustion.
To fix this, we need to separate the servers taking the requests from the servers doing the heavy compute. In cloud architecture, this is called Decoupling.
1. The Restaurant Analogy
Think of your web server as a waiter in a restaurant.
A bad restaurant has the waiter take your order, walk into the kitchen, cook the meal themselves, and bring it back to you. While they are cooking, they can't help any other tables.
A good restaurant uses a ticket system. The waiter takes your order, slaps a ticket onto a rail, and immediately goes back to serving other tables. The chefs in the kitchen look at the rail, pull the tickets one by one, and do the heavy lifting in the background.
We need to build that ticket rail and hire some chefs.
2. The Ticket Rail: Message Queues
In the cloud, the "ticket rail" is called a Message Queue. It is a simple, highly durable service that holds onto text messages (tickets) until a server is ready to read them.
- AWS provides Amazon SQS (Simple Queue Service
- Azure provides Azure Service Bus.
How the new flow works?
- The user uploads a video.
- The Web Server saves the raw video file to S3 (Object Storage).
- The Web Server drops a tiny text message into the SQS Queue saying: "Hey, there is a new video at this S3 link that needs processing. User ID is 45."
- The Web Server immediately replies to the user: "Thanks! We are processing your video and will email you when it's ready." The user is happy, the web server is instantly freed up to handle the next user, and the app remains lightning fast. But who actually processes the video?
3. The Chefs: Background Workers
We don't want our front-line Web Servers doing heavy compute. Instead, we spin up a completely separate, hidden group of resources called Workers.
While you could use a second group of standard web servers (like EC2 or Azure VMs) for this, modern cloud teams usually use Containers or Serverless Functions because they are incredibly lightweight and can scale up in seconds rather than minutes.
- In AWS: You might use AWS Fargate (Serverless Containers) or AWS Lambda.
- In Azure: You would use Azure Container Apps or Azure Functions.
These Worker containers have one job: they constantly ask the Message Queue, "Do you have any new tickets?"
When a worker sees the message about the video, it picks it up, downloads the raw video from S3, spends 10 minutes processing it, saves the final version, and updates the Database to mark the job as "Complete."
Because the Workers are completely decoupled, you can scale them based on the exact length of the queue. If there are 0 messages, you scale the containers down to zero so you pay nothing. If 1,000 users upload videos at once, the queue fills up, and the cloud provider instantly spins up 50 container apps to churn through the backlog.
Architecture Diagrams
Here is the final, fully decoupled architecture showing the Web Servers writing to the Queue, and the isolated Worker Servers reading from it.
Further Reading & Official Documentation
If you want to dive deeper into the concepts covered in this article, check out the official documentation from AWS and Microsoft:
AWS Queues & Background Workers
- Amazon SQS: What is Amazon Simple Queue Service?
- AWS Fargate: What is AWS Fargate for Serverless Containers?
- Asynchronous Integration: Asynchronous messaging patterns in AWS
Azure Queues & Background Workers
- Azure Service Bus: What is Azure Service Bus?
- Azure Container Apps: Azure Container Apps Overview
- Asynchronous Messaging: Asynchronous messaging options in Azure
That brings us to the end of our five-part journey into cloud architecture. We took a single, fragile web server and evolved it into a highly available, endlessly scalable, secure, fast, and fully decoupled enterprise application.
If you found this series helpful and want to keep leveling up your engineering skills, be sure to follow for more such content!


Top comments (0)