Moving a fleet of on-premises Kubernetes clusters to Amazon EKS is rarely blocked by the cloud transition itself — it's blocked by the version gap between where those clusters have been sitting and where EKS lives today. This post walks through that gap, how the AWS DevOps Agent closes it before deployment, and what changes operationally once your workloads are running on EKS.
The modernization challenge: the Kubernetes version gap
On-premises clusters tend to stagnate — often frozen at v1.23 or v1.25 to avoid destabilizing aging applications. Amazon EKS, by contrast, follows a forward-leaning lifecycle that typically targets v1.30 and newer. That delta isn't cosmetic; it's a real shift in the Kubernetes API surface and resource schema, and it's where most migrations stall.
The friction shows up hardest at API deprecation boundaries. Manifests that deploy cleanly on-premises fail outright on EKS when they reference removed resource kinds — extensions/v1beta1, apps/v1beta1, and policy/v1beta1 are the classic blockers. policy/v1beta1, for example, was removed in v1.25, so any manifest still using it will fail on a v1.30 target without prior refactoring.
There's a second dimension beyond API versions: storage and networking drift. Moving off local CSI drivers and custom ingress controllers onto the Amazon EBS CSI driver, the AWS Load Balancer Controller, and the VPC CNI is a structural re-architecture, not a lift-and-shift. At scale, manually auditing every YAML manifest against these two axes stops being viable — human review misses the nuanced OpenAPI schema changes between versions, and trial-and-error deployment against production EKS clusters is an expensive way to find them.
Pre-migration validation: the AWS DevOps Agent CLI compatibility engine
This calls for a shift-left validation model. The AWS DevOps Agent acts as a context-aware, pre-flight compatibility engine — it analyzes on-premises manifests against the OpenAPI specification of the target EKS version before a single resource is provisioned.
1. Export the source cluster state
# Capture every namespaced resource from the on-prem cluster
kubectl get all --all-namespaces -o yaml > onprem-cluster-state.yaml
2. Define the migration task
agent-config.json
{
"migrationTask": {
"sourceEnvironment": {
"platform": "on-premises",
"kubernetesVersion": "1.24"
},
"targetEnvironment": {
"platform": "aws-eks",
"kubernetesVersion": "1.30",
"region": "us-east-1"
},
"inputManifests": "./manifests/",
"outputDirectory": "./eks-migration-report/"
}
}
3. Run the analysis
aws-devops-agent analyze --config agent-config.json
The agent returns a Migration Report with a compatibility score, a breaking-changes log, and remediation manifests that are ready to deploy — no separate refactoring pass required.
Remediation in practice
Take a legacy PodDisruptionBudget written against policy/v1beta1, which is non-functional starting in EKS 1.26:
Legacy Manifest
apiVersion: policy/v1beta1
kind: PodDisruptionBudget
metadata:
name: payment-service-pdb
spec:
minAvailable: 2
selector:
matchLabels:
app: payment-service
Agent-remediated Manifest
# =====================================================================
# AWS DEVOPS AGENT REMEDIATION REPORT
# Resource: PodDisruptionBudget/payment-service-pdb
# Issue: API version policy/v1beta1 removed in Kubernetes 1.26+
# Action: Upgraded apiVersion to policy/v1
# =====================================================================
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: payment-service-pdb
spec:
minAvailable: 2
selector:
matchLabels:
app: payment-service
This removes the manual burden of chasing schema changes resource-by-resource, and guarantees manifests are compliant before they ever reach the EKS control plane.
Operations: the invisible throttling problem
Migrations frequently surface architectural issues that older clusters simply never enforced. Chief among them: API Priority and Fairness (APF), the mechanism that protects the EKS control plane from overload — and a system that's historically hard for SRE teams to reason about
The APF transparency problem. When the API server is under pressure, APF returns HTTP 429 responses. Because client-go retries these transparently, the offending workload often shows no errors in its own logs — only rising latency. Your dashboards can stay green while the control plane is quietly degrading. WATCH calls are especially dangerous here: unlike short-lived GET/LIST calls, they hold a concurrency "seat" for the life of the connection, so a handful of misbehaving watchers can exhaust available seats fast.
In a fault-injection exercise, a misbehaving Python async controller was scaled to 50 replicas, generating 1,600–2,000 requests per second against the API server. The effects were immediate:
Correlating CloudWatch audit logs, EKS metrics, and CloudTrail, the AWS DevOps Agent isolated the root cause to the controller service account, entirely inside the workload-low APF priority level, with traffic shifting from simple LIST calls to a heavy mix of WATCH and CREATE/DELETE mutations — the signature of a client that stopped caching and started polling.
Automation: the Operator and volatile event capture
EKS state is volatile by design — when a pod is rescheduled or deleted, the logs and events needed for root-cause analysis often disappear with it. Where a static tool only gives you a snapshot after the fact, the EKS DevOps Agent Operator captures a "movie" of the failure by collecting evidence at the moment a state transition happens.
Unlike a Bedrock Agent, which needs manual tool wiring and pipeline configuration, the Operator runs as an autonomous detection-and-collection engine. It uses AWS Systems Manager Run Command to reach node-level data that's invisible to kubectl — kernel dmesg output, IPAMD networking internals — though this depth is limited to managed nodes; on Fargate, collection stays at the Kubernetes manifest and log level.
An OOMKilled failure, narrated
- Detection — the Operator's informer catches the Running → OOMKilled transition within milliseconds.
- Collection — SSM (on managed nodes) is triggered to pull manifests, container logs, and kernel logs immediately.
- Storage — evidence lands in S3 and CloudWatch under a 14-day lifecycle policy to keep costs down.
- Trigger — an HMAC-SHA256-signed webhook notifies the DevOps Agent to start the investigation.
In one such investigation of a web-python container, the Agent traced the failure to an unbounded processed_records list inside the _cache_worker thread. Modeling the leak at 20Mi/min against a 200Mi container limit let it mathematically explain why the container died at the 10-minute mark, every time — turning on-call from reactive log spelunking into an architectural finding delivered up front.
Production results and operational guidelines
The ROI shows up as measurable MTTR reduction in production environments:
| Organization | Scenario | Result |
|---|---|---|
| Customer A | 38,000 OneAgents across 500+ accounts, integrated with Dynatrace | Single pane of glass; eliminated troubleshooting "black boxes" |
| Customer B | Lambda configuration disruptions | MTTR: 2 hrs → 28 min (77% reduction) |
| Customer C | API integration regressions during peak events | MTTR: 75% reduction, down to 20–30 min |
Architectural guidelines to carry into your own environment
- Access management — configure EKS Access Entries with the Standard type and attach AmazonAIOpsAssistantPolicy so the Agent has the introspection rights it needs.
- Secure identity — use EKS Pod Identity for role-to-role communication between the Operator and the Agent, keeping least privilege intact.
- Cost governance — apply 14-day lifecycle policies to the CloudWatch log groups and S3 buckets used for diagnostic data; value drops off fast after the point of impact.
- Pipeline gating — run the Agent CLI as a mandatory shift-left gate in CI/CD so incompatible manifests never reach production.
Conclusion
Closing the Kubernetes version gap is a one-time migration problem; keeping an EKS control plane healthy afterward is an ongoing one. Pairing pre-migration validation with Operator-driven incident capture turns both into a single, intelligent operational layer — one that catches breaking API changes before they ship and explains production incidents in the time it used to take to notice them.

Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.