DEV Community

Cover image for Cloud computing concepts: scaling, serverless, HA and VPCs explained

Cloud computing concepts: scaling, serverless, HA and VPCs explained

Cloud computing is easier to learn as a handful of architecture ideas than as a catalogue of three hundred product names. Every production backend on AWS, Google Cloud or Azure is built from the same eleven cloud computing concepts: scaling, load balancing, autoscaling, serverless, event-driven design, container orchestration, the storage hierarchy, high availability, durability, infrastructure as code and private networking. This article walks through all eleven with the numbers that matter and the trade-offs the marketing pages leave out, then puts them into one diagram.

TL;DR

  • Scale out, not up. A bigger machine needs no code changes but hits a ceiling; many small stateless machines behind a load balancer survive the loss of any one of them.
  • Serverless still has servers. AWS Lambda runs your function in a short-lived microVM, bills nothing when idle, and stops every invocation at 15 minutes.
  • Queues beat call chains. Publishing an event and letting workers catch up means a slow email provider can't fail a checkout.
  • Availability is not durability. Amazon S3's SLA is 99.9 % uptime, about 43 minutes a month; its durability is designed at eleven nines. One is "can I reach it", the other is "is it still there".
  • Write the infrastructure down. Terraform and friends turn console clicks into reviewed pull requests, and a VPC with private subnets keeps your database off the internet.

Cloud computing concepts 1–3: scaling, load balancing and autoscaling

Vertical vs horizontal scaling

Vertical scaling ("scaling up") means giving the one machine more: from four cores to thirty-two, from 32 GB of RAM to 128. Nothing in the code changes. But no single machine has ten thousand cores, the biggest instances cost disproportionately more, and one machine is still a single point of failure.

Horizontal scaling ("scaling out") keeps each server small and runs many copies side by side. Lose one of ten and you lose a tenth of your capacity while the service stays up. The price is one rule: the application tier must be stateless. No sessions, uploads or state on a server's local disk. They live in a database, a cache or object storage, so any server can answer any request.

Load balancing: Layer 4 vs Layer 7

Once you have many servers, something has to pick one for each request. A load balancer is a reverse proxy between the internet and your fleet, and it comes in two kinds:

Layer 4 (e.g. AWS NLB) Layer 7 (e.g. AWS ALB)
Sees TCP/UDP packets: IP and port HTTP: paths, headers, cookies, methods
Good for raw throughput, very low latency routing /api to one service, /static to another

The part people forget is health checks. The balancer probes an endpoint such as /healthz on every target; on an AWS Application Load Balancer the interval is 5 to 300 seconds, 30 by default. After a configured number of consecutive failures the target is taken out of service, and put back after enough successes. That is what turns "a server crashed" into "a server was quietly removed".

On algorithms: round robin suits short, uniform requests, least connections suits long-lived ones such as WebSockets, and sticky sessions are a sign that state has leaked into the app tier.

Autoscaling and the flapping trap

Provisioning for your peak means paying for it all night. An Auto Scaling Group watches a metric (average CPU, request latency, or the depth of a queue) and changes the number of instances between a minimum and a maximum. A typical rule: if average CPU stays above 70 % for three minutes, add instances and register them with the load balancer.

Scaling in matters as much as scaling out, because idle instances are the bill. Scale in or out too eagerly and the group flaps: instances are created and destroyed in a loop. The fix is a cooldown period, a wait after each scaling action before the next one, for example 300 seconds.

Cloud architecture concepts 4–6: serverless, event-driven architecture, containers

How does serverless work?

"Serverless" means you do not own, patch or pay for servers while no code runs. With Function-as-a-Service (AWS Lambda, Google Cloud Functions) you write a handler; an HTTP request, a file upload or a database change triggers it. Under the hood, Lambda runs on Firecracker microVMs, which the project describes as having "a < 125 ms startup time and a < 5 MiB memory footprint". Billing is per millisecond of execution and memory. No traffic for three months means a compute bill of zero.

The trade-offs are the reason serverless is not the answer to everything:

  • Cold starts. A fresh runtime has to initialise before your code runs, typically adding somewhere between 100 ms and a couple of seconds.
  • A hard time limit. A Lambda invocation can run for at most 900 seconds (15 minutes).
  • No local state between invocations.

Serverless fits event pipelines and spiky, sporadic APIs. It fits persistent WebSocket servers and multi-hour jobs badly.

Event-driven architecture: why queues decouple services

The synchronous version of a checkout looks like this: checkout calls payment, payment calls inventory, inventory calls fraud, fraud calls email. If the email provider takes ten seconds to answer, the customer's checkout times out.

In an event-driven design the checkout does not call anyone. It publishes one event and returns:

checkout ──► OrderPlaced ──► event bus / topic (EventBridge, SNS, Pub/Sub)
                                   ├──► queue ──► payment worker
                                   ├──► queue ──► inventory worker
                                   └──► queue ──► email worker   (down for an hour? messages wait)
Enter fullscreen mode Exit fullscreen mode

Two primitives do the work. A pub/sub topic (Amazon SNS, Google Cloud Pub/Sub, EventBridge as an event bus) fans one event out to many subscribers. A message queue (SQS, RabbitMQ) buffers messages for one consumer, so producer and consumer can run at different speeds. If the email worker is down for an hour, its queue fills and drains later, and no order is lost. The cost is that you now reason about eventual consistency and duplicate delivery instead of a single stack trace.

Container orchestration: what Kubernetes actually does

Docker solved packaging: code, runtime, libraries and config in one immutable image. Running five hundred containers across fifty machines is a different problem (placement, networking, rolling deploys, restarts), and that is what Kubernetes and AWS ECS are for.

Kubernetes' control plane has an API server, a scheduler, a controller manager and etcd, a distributed key-value store holding the cluster's state. You declare the state you want ("ten replicas of the auth service, 2 GB of memory each"). The scheduler places pods on nodes with room, and controllers keep comparing what is running with what you asked for. A crashed container is restarted; a dead node's pods are rescheduled onto healthy nodes.

Cloud storage types: object, block, database, cache

Cloud storage is four different things, picked by access pattern:

Type Examples Access Use it for
Object storage Amazon S3, Google Cloud Storage, Azure Blob HTTP PUT/GET on whole objects uploads, video, logs, backups
Block storage Amazon EBS, Persistent Disk a virtual disk attached to one VM, formatted as ext4/NTFS database engines, anything needing random writes
Managed databases RDS, Cloud SQL, Aurora; DynamoDB SQL with ACID transactions; or key-value/document at scale your system of record
In-memory cache Redis, Memcached RAM, sub-millisecond hot reads, sessions, rate limits

Object storage costs cents per gigabyte-month and has no practical capacity limit, but it is not a filesystem: you can't edit bytes in the middle of an object. Block storage is the opposite, fast random writes, one instance at a time. A cache in front of the database is often the cheapest scaling you will ever do.

High availability vs durability

High availability and the nines

Availability is the share of time your system answers. It is quoted in nines, and each nine cuts allowed downtime by ten:

Availability Downtime per year
99 % (two nines) about 3.65 days
99.9 % (three nines) about 8.76 hours
99.99 % (four nines) about 52.6 minutes
99.999 % (five nines) about 5.26 minutes

You buy nines by removing single points of failure across fault domains. AWS describes Availability Zones as "multiple, isolated locations within each Region"; each is one or more data centres with its own power and cooling, miles from the others. High availability means running active instances in at least two zones, with the database replicated between them, so losing a whole building is a failover instead of an outage.

Availability vs durability: the interview trap

These two words get used interchangeably and measure different things. Availability asks: can I read or write my data right now? Durability asks: will my data still exist, uncorrupted, years from now?

Amazon S3 is the standard example. Its service level agreement pays credits when monthly uptime drops below 99.9 %, which allows roughly 43 minutes a month of errors. Its durability, per the S3 FAQ, is designed at 99.999999999 %, eleven nines, with objects stored "across multiple devices spanning a minimum of three Availability Zones". Do the arithmetic: at eleven nines, ten million objects lose on average one object every ten thousand years. An outage makes S3 unreachable; it doesn't delete your data. Design for both, separately: replicas across zones for availability, backups and versioning for durability.

Infrastructure as code and VPC networking

Infrastructure as code vs ClickOps

Creating servers, subnets and firewall rules by clicking through a web console ("ClickOps") leaves no audit trail, no rollback and, sooner or later, a staging environment that no longer matches production. Infrastructure as code describes the desired end state in files kept in Git: Terraform or OpenTofu, AWS CloudFormation or CDK, Pulumi. Every change to a port or a replica goes through a pull request. terraform plan shows what would change before anything does.

What is a VPC? Public and private subnets

A Virtual Private Cloud is a private, software-defined network. You give it an address range such as 10.0.0.0/16 (65,536 addresses) and split it into subnets:

  • A public subnet has a route to an Internet Gateway. It holds only what must face the internet: load balancers and NAT gateways.
  • A private subnet has no inbound route from the internet. Application servers and databases live here. When they need to reach out, for package updates for example, the traffic goes through a NAT gateway in the public subnet.

On top of that, security groups are stateful firewalls around each instance, and network ACLs filter at the subnet level. Least privilege in practice: the database accepts PostgreSQL traffic only from the application servers' security group. An illustrative Terraform example of that one rule:

# illustrative example: Postgres reachable only from the app tier
resource "aws_security_group_rule" "db_from_app" {
  type                     = "ingress"
  from_port                = 5432
  to_port                  = 5432
  protocol                 = "tcp"
  security_group_id        = aws_security_group.db.id
  source_security_group_id = aws_security_group.app.id
}
Enter fullscreen mode Exit fullscreen mode

That doesn't make a breach impossible, whatever my video's enthusiasm suggested; a compromised app server still reaches the database. It does shrink the ways in to the ones you meant to build.

How the cloud computing concepts fit together

Put together, the eleven make the blueprint most production backends share:

DNS ──► load balancer (public subnet, 2+ AZs)
            │
            ▼
   autoscaling group of stateless app servers (private subnets, 2+ AZs)
            │                 │                         │
            ▼                 ▼                         ▼
   cache (Redis) ──► managed database (replicated)   event bus ──► queues ──► workers / functions
                                                                               │
                                                                    object storage (S3)
   everything above: defined in Git (IaC), reached only through security groups
Enter fullscreen mode Exit fullscreen mode

Verdict: SHIP IT

I stamped the masterclass SHIP IT because the fundamentals are the durable part of cloud computing. Learn the eleven patterns, keep your state out of your servers, and build systems where a failure costs you a component instead of the whole product. Which of these concepts gave you the most trouble when you started? I'd like to know in the comments.

FAQ

What is the difference between vertical and horizontal scaling?
Vertical scaling makes one machine bigger. Horizontal scaling adds machines behind a load balancer and needs stateless servers.

What is the difference between availability and durability?
Availability is whether you can access your data right now; durability is whether it survives long-term. Amazon S3's SLA covers 99.9 % availability while its durability is designed at eleven nines.

Does serverless mean there are no servers?
No. The provider runs your code on its servers, AWS Lambda on Firecracker microVMs, and you pay only while it runs.

Sources


This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.

Top comments (0)