Elastic Beanstalk's Cluster Mode is one day old, announced on September 17, 2026, and the incident below has not happened to anyone I know. I wrote the postmortem anyway, because everything in it is in the documentation: environments that use the same subnet set land on the same EKS cluster, the service detects drift when someone touches the cluster from outside, and while the drift lasts every update to every environment on that cluster fails. After 16 years operating financial platforms, I learned the question is not 'does this cut cost?', it's 'how many services stop accepting hotfixes when one person makes a mistake at 5 PM on a Friday?'.
What happened (in the rehearsal)
The scenario: a payments team with nine services in separate Standard environments migrates all of them to Cluster Mode on the same three private subnets. The first create-environment builds the beanstalk-cluster-{uuid} stack in about ten minutes; the other eight join that cluster without creating anything. The EC2 bill drops because EKS Auto Mode bin-packs the containers, and the team celebrates.
The trigger: an engineer with eks:* on the admin role spots a node pool that looks idle and edits it directly in EKS, as they would on any cluster in the company. Beanstalk compares the cluster against the service-managed configuration it expects, records ERROR Cluster drift detected in the environment events and pauses maintenance for that cluster.
The consequence: from that moment the service places no new environment on the cluster, applies no add-on update and, the part that hurts, rejects every update-environment for all nine environments. Replicas keep running; traffic keeps arriving through the ALB. What dies is the ability to change. Two hours later a bug in the reconciliation service needs a hotfix and the deploy fails with a message nobody on call had seen before.
The retro that follows is not looking for someone to blame. It is looking for the boundary that was missing.
Timeline
T-14d: Consolidation: Nine environments created with the same subnet set and the same
cluster-role,node-roleandobservability-role. One cluster, one control plane, a Kubernetes version Beanstalk picked that stays fixed for the cluster's life.T-0 5:05 PM: Manual edit: Node pool changed in the EKS console. No alarm fires: the cluster is healthy from Kubernetes' point of view.
T+0h20: Drift detected: The next
update-environmentmakes Beanstalk re-evaluate the cluster and emitCluster drift detected ... Service will skip cluster maintenance. The event sits in the environment's event list, nobody subscribes to that stream.T+2h10: Hotfix blocked: Reconciliation service deploy fails. On-call tries another environment on the same cluster: same failure. Nine services with no path to change.
T+3h40: Recovery: The drift message names what changed; the team reverts the node pool edit. The next environment operation re-evaluates the cluster, Beanstalk resumes management and the failed deploy is retried. Recovery is per cluster, not per environment.
Root cause: The root cause was not the node pool edit, it was treating the subnet set as a networking decision when, in Cluster Mode, it is the blast-radius decision. The documentation is explicit: same subnet set, same cluster; a cluster-wide failure affects every environment on it; subnets and roles cannot be changed after the environment is created. The team consolidated nine services with different SLOs into a single failure boundary and left humans with write permission on an EKS cluster the service considers its own.
Subnets pick the cluster, the cluster sets the blast radius
Two subnet sets produce two service-managed clusters. Drift freezes only the cluster where the manual edit happened; the second keeps accepting deploys.
🟧 AWS: Elastic Beanstalk (control plane do serviço)
- Beanstalk agrupa por subnet set (compute)
- Detector de drift config esperada vs real (security)
🟧 AWS: Cluster A (subnets s-1,s-2,s-3): DRIFTED
- EKS Auto Mode beanstalk-cluster-a (compute)
- eb-payments-api 3 replicas (compute)
- eb-reconciliation hotfix bloqueado (compute)
- Node pool editado à mão (network)
🟧 AWS: Cluster B (subnets s-4,s-5,s-6), saudável
- EKS Auto Mode beanstalk-cluster-b (compute)
- eb-ledger node-pool dedicado (compute)
🟧 AWS: Observabilidade compartilhada
- CloudWatch 4 log groups sem retenção (data)
Flows
- gha -> eb: update-environment
- eb -> eksa: same subnet set → cluster A
- eb -> eksb: other subnet set → cluster B
- eng -> np: edits directly in EKS
- np -> drift: config ≠expected
- drift -> pay: update fails
- drift -> rec: update fails
- eb -> ledger: deploy continues
- pay -> cw: OTel sidecar
- ledger -> cw: OTel sidecar
Remediation: design the boundary before the first environment
Subnets per failure domain, not per VPC: the cheapest day-zero decision is the most expensive to fix later, changing an existing environment's subnets is rejected, and the way out is a new environment plus a blue/green CNAME swap. I would allocate one subnet set per SLO and compliance regime: one cluster for what is in PCI DSS scope and carries card data, another for what can go two hours without a hotfix, another for dev. Separating production from development is the case the documentation itself cites as common.
No human write access to the service's EKS: drift only exists because someone could edit. An SCP or a permission policy with Deny on eks:Update*, eks:Delete* and eks:Create* conditioned on the aws:ResourceTag of the beanstalk-cluster-* stacks forces every change through update-environment. Read access can stay open; the deletion-verification runbook needs it.
node-pool for what cannot compete for CPU: without it, replicas from different environments share the same node. Giving the ledger service a node-pool value no other environment uses removes the noisy neighbor, but changing that value restarts every replica, decide at creation.
One application-role per environment: network between environments is blocked by default, but outbound is not restricted and a shared credential cancels the separation. A per-environment role, scoped to what it touches, is the minimum.
Three boundaries, three blast radii
| Criterion | How you get it | What it separates | What it does not separate | When to use |
|---|---|---|---|---|
| Shared cluster (default) | Same subnet set | Own eb-<env> namespace; inter-environment traffic blocked |
Control plane, nodes, Kubernetes version, drift, egress | Services of one application, same team, same SLO |
| Dedicated nodes |
node-pool with a unique value |
Node CPU and memory; noisy neighbor | Control plane and drift remain shared | Latency-sensitive service inside a trusted cluster |
| Separate clusters | Different subnet sets | Everything: control plane, nodes, network, drift | Nothing, but each cluster adds US$ 0.10/h | Distinct customers, code you do not control, regulatory scope |
The defaults that only show up at 2 AM
The drift retro was hiding a second incident waiting its turn. The defaults in the aws:elasticbeanstalk:eks:* namespaces are reasonable for a demo and dangerous for financial production.
Probes off: readiness, liveness and startup ship with enabled=false. Without readiness, a new replica gets ALB traffic before its connection pool opens; without liveness, a hung process stays in service until someone notices. Enable all three, with period-seconds below the default 30 and failure-threshold of 3: 90 seconds of worst-case detection.
cpu at 250m with no cpu-limit: the request is small and the ceiling does not exist. A replica with a CPU leak competes for the node with whoever sits next to it. Set cpu-limit and memory-limit (the memory default is 1Gi) per service, measured.
CPU scaling at 80%: with no explicit trigger the service scales between min-replica=1 and max-replica=10 on CPU utilization. For a payments queue, scaler-type=metrics-api reading the backlog from your own endpoint is the metric that matters; min-replica=1 in production is a single point of failure with a nice name.
Secrets without application-role: the secrets option and scaler-auth-secret mount through the pod identity, which only exists when application-role is set. Without it the mount fails and replicas never start, silently, on the first deploy.
Log groups without retention: the four /aws/elasticbeanstalk/* groups are shared per account and region and are created with no expiration; every deploy creates a new stream. Set retention on day one.
The math that justifies, or doesn't, the shared cluster
Cluster Mode's argument is per-application cost falling as the application count rises. The math has three fixed terms that Standard did not have.
Control plane: US$ 0.10 per cluster per hour, roughly US$ 73 per month per subnet set. Every failure boundary I recommended above costs that. Once the Kubernetes standard support window ends, extended support rises to US$ 0.60/h, and since the version is fixed for the cluster's life, that clock starts at creation.
EKS Auto Mode: a per-instance management fee on top of EC2. In the pricing page's example, an m5a.xlarge costs US$ 0.172/h of EC2 plus US$ 0.02064/h of Auto Mode: 12% on top of compute.
Custom metrics: Standard publishes to AWS/ElasticBeanstalk, free. Cluster Mode publishes to ElasticBeanstalk/Infrastructure, /System and /Application, custom namespaces billed per metric, and the two container metrics are emitted per replica. Twenty replicas become dozens of paid metrics before the application emits its first one.
The launch post itself gives the yardstick: under roughly US$ 500 per month of workload, EKS overhead outweighs the savings and Standard remains the better fit. My reading: the mode pays off when you have six to eight services in the same domain, each underusing a dedicated instance today, and a platform willing to pay US$ 73 per extra failure boundary. With two services, it is a migration that pays more to change less.
Well-Architected lenses
-
security: Inter-environment network blocked by default; open it only with
ingress-groupsor an allowlist on the receiver. Egress is not restricted, handle it at the subnet. A shared cluster is soft multi-tenancy; PCI DSS scope or a distinct customer calls for its own subnets. Per-environmentapplication-roleandwafv2-acl-arnon the ALB. -
reliability: Blast radius is decided by the subnet set and is immutable per environment. Enable all three probes, set
min-replica≥ 2 andmax-unavailable=0on rolling updates, and subscribe to environment events to alert onCluster drift detected, the event exists, nobody reads it by default.
Anti-patterns
- One VPC, one subnet set, every service: turns the whole account into a single failure domain that can only be undone by recreating environment by environment.
-
Operating the cluster with
kubectlbecause 'it's EKS': it is EKS, but it belongs to the service; the first edit freezes deploys for every neighbor. -
Passing classic namespaces (
aws:autoscaling:*) in a migrated.ebextensions: Cluster Mode returnsInvalidParameterValueExceptioninstead of ignoring them, the pipeline breaks on the first deploy. -
GetSecretValuewithoutDescribeSecreton the custom observability role: the environment starts, then every scheduled credential refresh fails: Datadog stops receiving data with no deploy error. - Consolidating to save money with two services: the control plane and Auto Mode fee cost more than the dedicated instance you stopped paying for.
Curator's note: I would adopt Cluster Mode for the internal services of a single team, the ones paying for a dedicated instance today to serve 200 requests per minute, and for nothing that carries card data in the first six months. Before the first environment, I would draw the subnet map the way I draw an account map: each set is a cluster, each cluster is what stops together. I would deny EKS write access to every human, alarm on the drift event and set log group retention in the same commit. The hard-won lesson behind this: a platform that hides Kubernetes does not hide the blast radius, it only renames the knob that defines it.
Verdict
Use Cluster Mode when: the services belong to the same team and the same SLO, add up to more than US$ 500 per month of underused compute, and the platform accepts paying US$ 73 per month for each additional failure boundary. Stay on Standard when: it is one or two services, the application keeps state on local disk, or the regulatory scope demands infrastructure separation you are not going to pay for in separate clusters. Either way, the decision with no undo is the subnet set, treat it as an architecture decision, with an ADR, not as a form field.
Rating: adopt-with-boundaries
References
- AWS What's New: Elastic Beanstalk introduces Cluster Mode (17 Sep 2026)
- AWS News Blog: AWS Elastic Beanstalk introduces Cluster Mode
- Elastic Beanstalk Developer Guide: Beanstalk Cluster architecture (grouping, drift, deletion)
- Elastic Beanstalk Developer Guide: Multi-tenancy for Beanstalk Cluster environments
- Elastic Beanstalk Developer Guide: Configuration options for Beanstalk Cluster environments
- Elastic Beanstalk Developer Guide: Monitoring Beanstalk Cluster environments
- Amazon EKS pricing, control plane and Auto Mode
- Amazon EKS Best Practices: Tenant isolation (soft vs hard multi-tenancy)
Originally published at fernando.moretes.com. By Fernando F. Azevedo: Senior Solutions Architect.
Top comments (0)