The result first
I tested how much of my cloud engineering work I could delegate to AI in two experiments.
- Level 2: Implement and operate a small environment after a human defines the architecture and safety boundaries.
- Level 3: Turn customer requirements into a cloud design that satisfies explicit acceptance criteria.
The live Level 2 experiments completed deployment, selected failure tests, recovery, and cleanup across AWS, Azure, and Google Cloud. The latest Level 3 trial, D5, scored 77/100 in an independent AI review: FAIL against the acceptance criteria, HOLD for adoption.
The distinction matters throughout this article. Level 2 has live execution records. Level 3 evaluates design documents. Neither experiment measures the percentage of a profession that can be replaced.
“Level” is my own classification, not an industry standard or an official mapping to cloud certifications. I did not run the introductory Level 1 stage.
Level 2: Delegate a defined infrastructure lifecycle
What the human specified
The common target was a Layer 7 load balancer serving HTTP through two private VMs. The VMs had no public IP addresses and no direct SSH access from the internet.
| Component | AWS | Azure | Google Cloud |
|---|---|---|---|
| Public entry point | Regional ALB | Application Gateway | Global external Application Load Balancer |
| Backend | Two EC2 instances in separate AZs | Two Linux VMs in separate zones | Two VMs in a regional managed instance group |
| Outbound path | S3 Gateway Endpoint | NAT Gateway | Cloud Router and Cloud NAT |
| Remote state | S3 | Blob Storage | Cloud Storage |
These were implementations of common requirements, not identical architectures. Google Cloud used a managed instance group, while AWS and Azure used individual VMs. A global load balancer did not make the backend multi-region.
The human retained responsibility for requirements, allowed changes, identity and exposure boundaries, and approvals. AI handled Terraform, GitHub Actions, short-lived federated authentication, remote state, monitoring, selected failure tests, recovery, and cleanup.
What counted as completion
I looked beyond a successful terraform apply:
- HTTP 200 through the load balancer and healthy backends
- Successful short-lived CI authentication and remote-state initialization
- No changes after migration, repair, and recovery
- Observed alert-state changes for selected failures
- A delete-only plan, empty state, and checks for remaining managed resources
The execution records report completion across all three clouds. Cleanup removed 139 Terraform-managed items: AWS 48, Azure 53, and Google Cloud 38. These totals include the application infrastructure and bootstrap resources for state and authentication.
That is not a count of VMs or individually billable resources. “Zero remaining” refers to the managed active resources checked in this experiment. It does not mean an empty account, physical erasure of every retained copy, or a guarantee of no later charges. See the Level 2 final report for the evaluation method and counts.
Three failures made the boundaries concrete
AWS: the instance was private, but its startup dependencies were unreachable.
Insufficient outbound connectivity prevented Nginx installation, and the ALB returned 502. AI diagnosed the problem and added narrowly scoped HTTPS access to the S3 prefix list. The lesson was to review startup dependencies alongside inbound isolation. This S3-based path depended on the selected OS package source; it was not general internet access.
Azure: a successful apply removed tags from 12 resources.
Local and CI inputs differed. The CI apply completed successfully while removing additional tags. The inputs were aligned, and a tag-only repair followed human approval. The exit status alone did not establish that the change matched the intended result.
Google Cloud: missing backend configuration made 18 existing resources look new.
The backend declaration was not tracked in Git. CI could not use the existing remote state and planned 18 additions. The process stopped before apply and repaired the declaration. Unexpected creation, deletion, or replacement needs to be a stop condition, not just another line in the output.
These incidents and their recovery are documented in the failure analysis.
What Level 2 did not establish
Monitoring checks covered selected signals and alert states. External notification receivers were not configured, so delivery to an operator was not demonstrated. The failure scenarios also differed between providers.
The experiments ran sequentially: AWS, then Azure, then Google Cloud, carrying lessons forward. Human-intervention counts used different units. That rules out a clean provider ranking or a defensible intervention-reduction percentage.
There was also no standardized human-only baseline for time, cost, or quality. The final report identifies code and CI controls that need checking before reuse; successful historical execution is not proof that the current checkout reproduces unchanged.
The supported conclusion is narrower: an experienced engineer could delegate a substantial implementation and operations lifecycle in small, disposable environments while retaining design, review, and approval. Production readiness, long-term SLOs, and unattended operation remain outside that result.
Level 3: Can AI produce an acceptable design from requirements?
The second experiment started with a member-facing application: registration, login, profile changes, list/detail views, favorites, and administrative updates.
Its main constraints were:
- Initially 10,000 members and one million monthly page views
- Dynamic traffic of 10 requests/second normally and 100 requests/second for a 15-minute peak, with a 9:1 read/write ratio; static requests counted separately; 500 concurrent users
- Primary API p95 at or below 500 ms and server-side errors below 1%
- 99.9% monthly availability, including planned downtime and underlying platform failures
- RTO of 30 minutes from service-impact onset, including detection and five stable minutes after recovery; RPO of five minutes
- Production data and backups stored in Japan; departing members' data removed from active systems within 30 days, backups retained for no more than 35 days, and deletion reapplied after restoration
- Five developers and one infrastructure engineer, supporting weekday daytime operations with no guaranteed immediate response at night
- JPY 100,000/month including tax for production and minimal development/test environments, including mandatory networking, monitoring, logs, and backups; labor and other exclusions listed separately
The RTO/RPO targets covered process, runtime, database, and single-AZ failures. Region-wide failure was excluded. These were design requirements, not observed service metrics. The customer answers and requirements document preserve the input.
Budget compliance could not depend on free tiers, temporary credits, or long-term commitment discounts.
The tension between a 30-minute recovery target and weekday daytime staffing is a useful example: selecting managed services does not by itself establish who can act, which actions are authorized, or whether unattended recovery works.
Acceptance required more than a total score
The September work was self-assessed: A moved from 61 to 65, and B scored 66. October introduced a reviewer separate from the designer. D4 and D5 used anonymously reviewed, frozen design-and-evidence packages.
A passing design had to meet every condition:
- Weighted total of at least 80/100
- Every one of 11 categories rated at least 3/5
- Six mandatory categories rated at least 4/5
- No unresolved P0 items
- All six adoption gates marked PASS
The mandatory categories were requirements, architecture, security/IAM, availability, cost, and backup/disaster recovery. No human scoring or adoption approval was obtained. These were independent AI document reviews, not production acceptance tests.
The latest result was 77, not the highest result of 78
| Trial | Requested designer configuration | Independent score |
|---|---|---|
| D1 | GPT-6 Luna / Low | 53 |
| D2 | GPT-6 Luna / Medium | 72 |
| D3 | GPT-6 Luna / High | 69 |
| D4 | GPT-6.1 Sol / Medium | 78 |
| D5 | GPT-6.1 Sol / High | 77 |
These are requested settings. Actual runtime models and applied reasoning settings are UNKNOWN, including for the reviewers.
The series also carries earlier findings forward. D5 revised D4, added controls and evidence, and used a fresh reviewer. The scores cannot isolate a model or reasoning-effort effect. September self-scores are a different evaluation method and are not part of this series.
The latest authoritative trial is D5, dated October 2, 2026. It failed the acceptance criteria and remained on hold for adoption. The methods and results and D5 report explain the chronology and comparison limits.
Requirements changed separately on October 2 remain unscored and were not D5 input. The score of 77 does not apply to those changed requirements.
Why adoption stayed on hold
Cost was 3/5, below its mandatory minimum of 4. The design did not establish an all-inclusive budget fit with headroom, adequately supported effort for custom controls and rehearsals, or sufficiently concrete cheaper alternatives. Keeping unknown prices explicit was credited; unknowns did not become zero or evidence of affordability.
Deletion gate G3 remained HOLD. Deletion-trigger events, retention of linking information, independent custody of evidence, and responsibility for external services or recipient copies were not approved.
Operations gate G4 also remained HOLD. Named primary and backup owners, approved effort, absence coverage, and authority for unattended recovery were unresolved.
Added controls brought an evidence burden. D5 added a SQL-backed AdmissionCoordinator to principal request paths. Evidence for operation counts, contention, latency, and recovery ordering was insufficient. Complexity fell from 4 to 3, accounting for the one-point total decrease from D4. Complexity 3 meets the general category floor; it is distinct from Cost failing a mandatory minimum.
Unresolved P0 items concerned owner decisions, source access, and prerequisites for live validation, rather than identified P0 desk-design defects. The D5 report and gate record keep these distinctions explicit.
What I would carry into real work
Level 2 demonstrated more than initial code generation. AI investigated provider constraints, authentication mismatches, partial applies, and state inconsistencies, then made scoped repairs and completed cleanup.
Level 3 showed increasingly specific design documents, but no accepted design. Missing prices, owner decisions, and empirical evidence remained missing. A score near 80 does not establish that acceptance is close when mandatory conditions remain unresolved.
My practical takeaway is to specify the verification and decision process alongside the delegated task:
- Define completion beyond command success: service behavior, state consistency, recovery, and cleanup
- Stop unexpected changes before apply
- Separate proposed designs, execution records, measured behavior, and approved decisions
- Name the people and permissions behind recovery, operation, and deletion
A useful next experiment would hold inputs, evidence, and evaluation conditions constant while verifying actual model settings. Live performance, recovery, and cost validation would be a separate step. Those are research proposals, not work already completed.
For now, the evidence supports broad delegation of predefined implementation under human control. It does not establish replacement of cloud engineering as a whole, or that a stronger model alone would satisfy the remaining adoption conditions.
Sources and AI disclosure
Source review date: October 6, 2026. References are pinned to Level 2 commit c5d25b6 and Level 3 commit e1ad7ab. The existing Japanese and English slides were also cross-checked. Preparing this article did not rerun cloud operations or produce a new assessment.
AI assisted the structure, drafting, and English wording of this article. The experiment conditions, results, and limitations are based on the linked records.
Top comments (0)