DEV Community

Fernando Azevedo
Fernando Azevedo

Posted on Originally published at fernando.moretes.com

Beanstalk Cluster Mode: the postmortem I wrote before the incident

Elastic Beanstalk's Cluster Mode is one day old, announced on September 17, 2026, and the incident below has not happened to anyone I know. I wrote the postmortem anyway, because everything in it is in the documentation: environments that use the same subnet set land on the same EKS cluster, the service detects drift when someone touches the cluster from outside, and while the drift lasts every update to every environment on that cluster fails. After 16 years operating financial platforms, I learned the question is not 'does this cut cost?', it's 'how many services stop accepting hotfixes when one person makes a mistake at 5 PM on a Friday?'.

What happened (in the rehearsal)

The scenario: a payments team with nine services in separate Standard environments migrates all of them to Cluster Mode on the same three private subnets. The first create-environment builds the beanstalk-cluster-{uuid} stack in about ten minutes; the other eight join that cluster without creating anything. The EC2 bill drops because EKS Auto Mode bin-packs the containers, and the team celebrates.

The trigger: an engineer with eks:* on the admin role spots a node pool that looks idle and edits it directly in EKS, as they would on any cluster in the company. Beanstalk compares the cluster against the service-managed configuration it expects, records ERROR Cluster drift detected in the environment events and pauses maintenance for that cluster.

The consequence: from that moment the service places no new environment on the cluster, applies no add-on update and, the part that hurts, rejects every update-environment for all nine environments. Replicas keep running; traffic keeps arriving through the ALB. What dies is the ability to change. Two hours later a bug in the reconciliation service needs a hotfix and the deploy fails with a message nobody on call had seen before.

The retro that follows is not looking for someone to blame. It is looking for the boundary that was missing.

Timeline

  1. T-14d: Consolidation: Nine environments created with the same subnet set and the same cluster-role, node-role and observability-role. One cluster, one control plane, a Kubernetes version Beanstalk picked that stays fixed for the cluster's life.

  2. T-0 5:05 PM: Manual edit: Node pool changed in the EKS console. No alarm fires: the cluster is healthy from Kubernetes' point of view.

  3. T+0h20: Drift detected: The next update-environment makes Beanstalk re-evaluate the cluster and emit Cluster drift detected ... Service will skip cluster maintenance. The event sits in the environment's event list, nobody subscribes to that stream.

  4. T+2h10: Hotfix blocked: Reconciliation service deploy fails. On-call tries another environment on the same cluster: same failure. Nine services with no path to change.

  5. T+3h40: Recovery: The drift message names what changed; the team reverts the node pool edit. The next environment operation re-evaluates the cluster, Beanstalk resumes management and the failed deploy is retried. Recovery is per cluster, not per environment.

Root cause: The root cause was not the node pool edit, it was treating the subnet set as a networking decision when, in Cluster Mode, it is the blast-radius decision. The documentation is explicit: same subnet set, same cluster; a cluster-wide failure affects every environment on it; subnets and roles cannot be changed after the environment is created. The team consolidated nine services with different SLOs into a single failure boundary and left humans with write permission on an EKS cluster the service considers its own.

Subnets pick the cluster, the cluster sets the blast radius

Two subnet sets produce two service-managed clusters. Drift freezes only the cluster where the manual edit happened; the second keeps accepting deploys.

🟧 AWS: Elastic Beanstalk (control plane do serviço)

  • Beanstalk agrupa por subnet set (compute)
  • Detector de drift config esperada vs real (security)

🟧 AWS: Cluster A (subnets s-1,s-2,s-3): DRIFTED

  • EKS Auto Mode beanstalk-cluster-a (compute)
  • eb-payments-api 3 replicas (compute)
  • eb-reconciliation hotfix bloqueado (compute)
  • Node pool editado à mão (network)

🟧 AWS: Cluster B (subnets s-4,s-5,s-6), saudável

  • EKS Auto Mode beanstalk-cluster-b (compute)
  • eb-ledger node-pool dedicado (compute)

🟧 AWS: Observabilidade compartilhada

  • CloudWatch 4 log groups sem retenção (data)

Flows

  • gha -> eb: update-environment
  • eb -> eksa: same subnet set → cluster A
  • eb -> eksb: other subnet set → cluster B
  • eng -> np: edits directly in EKS
  • np -> drift: config ≠ expected
  • drift -> pay: update fails
  • drift -> rec: update fails
  • eb -> ledger: deploy continues
  • pay -> cw: OTel sidecar
  • ledger -> cw: OTel sidecar

Remediation: design the boundary before the first environment

Subnets per failure domain, not per VPC: the cheapest day-zero decision is the most expensive to fix later, changing an existing environment's subnets is rejected, and the way out is a new environment plus a blue/green CNAME swap. I would allocate one subnet set per SLO and compliance regime: one cluster for what is in PCI DSS scope and carries card data, another for what can go two hours without a hotfix, another for dev. Separating production from development is the case the documentation itself cites as common.

No human write access to the service's EKS: drift only exists because someone could edit. An SCP or a permission policy with Deny on eks:Update*, eks:Delete* and eks:Create* conditioned on the aws:ResourceTag of the beanstalk-cluster-* stacks forces every change through update-environment. Read access can stay open; the deletion-verification runbook needs it.

node-pool for what cannot compete for CPU: without it, replicas from different environments share the same node. Giving the ledger service a node-pool value no other environment uses removes the noisy neighbor, but changing that value restarts every replica, decide at creation.

One application-role per environment: network between environments is blocked by default, but outbound is not restricted and a shared credential cancels the separation. A per-environment role, scoped to what it touches, is the minimum.

Three boundaries, three blast radii

Criterion How you get it What it separates What it does not separate When to use
Shared cluster (default) Same subnet set Own eb-<env> namespace; inter-environment traffic blocked Control plane, nodes, Kubernetes version, drift, egress Services of one application, same team, same SLO
Dedicated nodes node-pool with a unique value Node CPU and memory; noisy neighbor Control plane and drift remain shared Latency-sensitive service inside a trusted cluster
Separate clusters Different subnet sets Everything: control plane, nodes, network, drift Nothing, but each cluster adds US$ 0.10/h Distinct customers, code you do not control, regulatory scope

The defaults that only show up at 2 AM

The drift retro was hiding a second incident waiting its turn. The defaults in the aws:elasticbeanstalk:eks:* namespaces are reasonable for a demo and dangerous for financial production.

Probes off: readiness, liveness and startup ship with enabled=false. Without readiness, a new replica gets ALB traffic before its connection pool opens; without liveness, a hung process stays in service until someone notices. Enable all three, with period-seconds below the default 30 and failure-threshold of 3: 90 seconds of worst-case detection.

cpu at 250m with no cpu-limit: the request is small and the ceiling does not exist. A replica with a CPU leak competes for the node with whoever sits next to it. Set cpu-limit and memory-limit (the memory default is 1Gi) per service, measured.

CPU scaling at 80%: with no explicit trigger the service scales between min-replica=1 and max-replica=10 on CPU utilization. For a payments queue, scaler-type=metrics-api reading the backlog from your own endpoint is the metric that matters; min-replica=1 in production is a single point of failure with a nice name.

Secrets without application-role: the secrets option and scaler-auth-secret mount through the pod identity, which only exists when application-role is set. Without it the mount fails and replicas never start, silently, on the first deploy.

Log groups without retention: the four /aws/elasticbeanstalk/* groups are shared per account and region and are created with no expiration; every deploy creates a new stream. Set retention on day one.

The math that justifies, or doesn't, the shared cluster

Cluster Mode's argument is per-application cost falling as the application count rises. The math has three fixed terms that Standard did not have.

Control plane: US$ 0.10 per cluster per hour, roughly US$ 73 per month per subnet set. Every failure boundary I recommended above costs that. Once the Kubernetes standard support window ends, extended support rises to US$ 0.60/h, and since the version is fixed for the cluster's life, that clock starts at creation.

EKS Auto Mode: a per-instance management fee on top of EC2. In the pricing page's example, an m5a.xlarge costs US$ 0.172/h of EC2 plus US$ 0.02064/h of Auto Mode: 12% on top of compute.

Custom metrics: Standard publishes to AWS/ElasticBeanstalk, free. Cluster Mode publishes to ElasticBeanstalk/Infrastructure, /System and /Application, custom namespaces billed per metric, and the two container metrics are emitted per replica. Twenty replicas become dozens of paid metrics before the application emits its first one.

The launch post itself gives the yardstick: under roughly US$ 500 per month of workload, EKS overhead outweighs the savings and Standard remains the better fit. My reading: the mode pays off when you have six to eight services in the same domain, each underusing a dedicated instance today, and a platform willing to pay US$ 73 per extra failure boundary. With two services, it is a migration that pays more to change less.

Well-Architected lenses

  • security: Inter-environment network blocked by default; open it only with ingress-groups or an allowlist on the receiver. Egress is not restricted, handle it at the subnet. A shared cluster is soft multi-tenancy; PCI DSS scope or a distinct customer calls for its own subnets. Per-environment application-role and wafv2-acl-arn on the ALB.
  • reliability: Blast radius is decided by the subnet set and is immutable per environment. Enable all three probes, set min-replica ≥ 2 and max-unavailable=0 on rolling updates, and subscribe to environment events to alert on Cluster drift detected, the event exists, nobody reads it by default.

Anti-patterns

  • One VPC, one subnet set, every service: turns the whole account into a single failure domain that can only be undone by recreating environment by environment.
  • Operating the cluster with kubectl because 'it's EKS': it is EKS, but it belongs to the service; the first edit freezes deploys for every neighbor.
  • Passing classic namespaces (aws:autoscaling:*) in a migrated .ebextensions: Cluster Mode returns InvalidParameterValueException instead of ignoring them, the pipeline breaks on the first deploy.
  • GetSecretValue without DescribeSecret on the custom observability role: the environment starts, then every scheduled credential refresh fails: Datadog stops receiving data with no deploy error.
  • Consolidating to save money with two services: the control plane and Auto Mode fee cost more than the dedicated instance you stopped paying for.

Curator's note: I would adopt Cluster Mode for the internal services of a single team, the ones paying for a dedicated instance today to serve 200 requests per minute, and for nothing that carries card data in the first six months. Before the first environment, I would draw the subnet map the way I draw an account map: each set is a cluster, each cluster is what stops together. I would deny EKS write access to every human, alarm on the drift event and set log group retention in the same commit. The hard-won lesson behind this: a platform that hides Kubernetes does not hide the blast radius, it only renames the knob that defines it.

Verdict

Use Cluster Mode when: the services belong to the same team and the same SLO, add up to more than US$ 500 per month of underused compute, and the platform accepts paying US$ 73 per month for each additional failure boundary. Stay on Standard when: it is one or two services, the application keeps state on local disk, or the regulatory scope demands infrastructure separation you are not going to pay for in separate clusters. Either way, the decision with no undo is the subnet set, treat it as an architecture decision, with an ADR, not as a form field.

Rating: adopt-with-boundaries

References


Originally published at fernando.moretes.com. By Fernando F. Azevedo: Senior Solutions Architect.

Top comments (0)