DEV Community

Herbert
Herbert

Posted on

You Migrated to the Cloud, But Why Are You Still Thinking Like a Data Center?

You Migrated to the Cloud, But Why Are You Still Thinking Like a Data Center?

A Real Dilemma

A traditional industry IT manager has been frustrated lately. Three years ago, the company responded to the "cloud migration" initiative and moved their ERP system from an on-premises data center to a public cloud. The migration went smoothly — the software vendor directly installed Oracle 19 database on ECS instances, application servers were also on ECS, and the network topology was identical to the data center era.

But recently, problems emerged:

  • Database disk IOPS alerts frequently, month-end batch processing time doubled
  • The vendor suggested "adding IOPS" — the price was shocking
  • During business peak hours, CPU is insufficient, requiring VM shutdown for resizing, 15 minutes downtime each time
  • Three to four shutdowns per month

The IT manager asked: "We're already on the cloud, why is it still so painful?"

Because you only moved the data center to the cloud, but didn't truly "use" the cloud.


I. Three Symptoms of Fake Cloud Migration

Symptom 1: Giving Software Vendors SSH Access, Treating Cloud as Physical Machines

Many enterprise cloud migration approaches:

  • Self-install Oracle database on ECS
  • Mount cloud disks as data volumes
  • Grant software vendors SSH root access
  • When encountering performance issues, add CPU, memory, or optimize SQL

This is the root cause: SSH access locks in the "treat cloud as physical machine" mindset.

What do software vendors do after getting SSH access?

  • Manually install Nginx on ECS for load balancing
  • Set up open-source monitoring tools (Prometheus/Grafana) themselves
  • Write shell scripts for health checks and failover
  • SSH login to troubleshoot logs when issues occur

Result: Software vendors will never learn and use cloud-managed services (load balancers, monitoring, managed databases) because they only have SSH access and cannot access the cloud console.

Symptom 2: Continue Using Oracle, Running Even Slower on Cloud Disks

A more critical issue: Oracle's architecture is inherently unsuitable for cloud environments.

Oracle was designed for local disks on physical servers:

  • Synchronous write mechanism: Every transaction COMMIT must wait for redo log to be written to disk before returning
  • Multi-block read operations: Full table scans read multiple consecutive data blocks at once
  • Direct path writes: Bulk operations write directly to disk, bypassing buffers

These designs work fast on local NVMe SSDs (10-50 microseconds latency), but face fundamental obstacles on cloud disks:

  • Network latency: Each I/O crosses the network, 1-5 milliseconds latency (100-500x slower than local)
  • IOPS ceiling: Cloud ultra-high I/O disks max at 50,000 IOPS (requiring ~1TB capacity), far below local SSDs' hundreds of thousands
  • Throughput bottleneck: Limited by network bandwidth

Cost Comparison:

Hidden costs of self-managed Oracle:

  • Oracle licensing: Charged per vCPU, Standard Edition 2 ~$17,000/core (~$120,000 RMB)
  • DBA labor: 2 people × $2,000/month = $4,000/month
  • Backup HA: Additional ECS + storage + manual maintenance
  • Disaster recovery: Self-managed may take hours, managed RDS only takes 30 seconds

True TCO: Self-managed Oracle can be 2-3x more expensive than managed PostgreSQL/cloud-native databases.

Symptom 3: Single-Point Deployment, No Load Balancing

Many enterprises still use single-point architecture after cloud migration:

  • All traffic hits one ECS instance
  • During business peaks, CPU hits 100%, users experience timeouts
  • Can only shut down at midnight to manually upgrade to larger specs

This completely wastes the cloud's distributed architecture capabilities.

The correct cloud approach: Multiple instances + load balancing

  • Deploy 3-5 medium-spec ECS instances (like 4-core 8GB or 8-core 16GB)
  • Add a cloud load balancer in front
  • Traffic automatically distributes to multiple servers
  • Single instance failure automatically removed, business uninterrupted

II. Root Cause: SSH Access vs IAM Delegation

The Death Loop of Traditional Approach

Enterprise IT Department
  ↓ Only grants SSH
Software Vendor
  ↓ Can only use ECS as physical machines
  ↓ Install Nginx/MySQL/monitoring on ECS
Continue data center mindset
  ↓ Cannot use cloud-managed services
  ↓ When hitting bottlenecks, only change configs/optimize code
Performance and cost issues continue to worsen
Enter fullscreen mode Exit fullscreen mode

The Right Approach: Console Delegation (IAM Delegation)

Core transformation: Don't give SSH, give Console permissions (via delegation)

How cloud IAM delegation works:

  1. Enterprise Account A creates delegation, authorizing Software Vendor Account B to access specific resources

  2. Fine-grained permission control:

    • ✅ Can operate load balancers (configure balancing, health checks)
    • ✅ Can use managed databases (view monitoring, backup recovery)
    • ✅ Can view cloud monitoring (CPU/memory/network metrics)
    • ✅ Can access log services
    • ✅ Can use database diagnostics
    • ❌ But CANNOT SSH into ECS
  3. Resource isolation: Vendors can only see authorized resources, not other projects

  4. Operation audit: All Console operations recorded in cloud audit service, traceable

  5. Revocable anytime: Click to revoke when project ends, no password changes needed

The Essential Difference in Accountability

SSH authorization vs cloud account delegation is not just a technical difference, but a fundamental difference in accountability:

Dimension SSH Authorization Mode Cloud Account Delegation Mode
Authorization Target Authorize individuals (employee Zhang, Li) Authorize company (XX Software Corp, legal entity)
Accountability Unclear, keys may remain after employee leaves Clear and explicit, company account ops = company fully responsible
Audit Trail Only see IP and time, cannot confirm operator Audit logs precise to the second, complete operator/time/action
Staff Changes Departed employees' SSH keys hard to revoke Company account permissions centrally managed, staff departures don't affect
When Issues Occur Vendor claims "departed employee did it," client cannot hold accountable Vendor cannot deflect, company account operations have complete audit evidence
Professionalism Temporary worker mode, person-to-person B2B professional mode, company-to-company

SSH authorization to individuals means no accountability when issues occur. Vendors can claim "a departed employee did it," and enterprises have no recourse.

Account delegation is the B2B professional authorization model: Company to company, clear accountability, complete audit evidence, no deflection possible.

Required Transformation for Software Vendors

Dimension Traditional Mode (SSH) Cloud-Native Mode (Console Delegation)
Access Method SSH to ECS Cloud Console or API
Deployment Method Manual software installation Configure managed services
Monitoring Method Self-built Prometheus/Grafana Use cloud monitoring
Log Management Local log files on ECS Cloud log service
Load Balancing Install Nginx on ECS Configure cloud load balancer
Database Install MySQL/Oracle on ECS Use managed RDS or cloud-native DB
Skill Requirements Linux ops + open-source software Cloud service config + IaC

This requires software vendors to also transform: from "can install software" to "can use cloud services".


III. The Way Forward: Three Steps

Step 1 (3-6 months): Oracle → Managed PostgreSQL/Cloud-Native Database

Why must we replace Oracle?

Technical hard constraints:

  • Oracle optimized for local disks, cloud disks have 100-500x higher latency
  • Licensing explodes per vCPU (~$17k/core)
  • High operational costs (requires dedicated DBAs)

Recommended migration path:

Business Scenario Target Database Reason
OLTP transactional business Managed PostgreSQL Good compatibility, strong performance, mature ecosystem
High-concurrency web apps Managed MySQL Good connection pooling, lowest cost
Domestic compliance requirements Cloud-native database Fully autonomous, performance matches Oracle

Real benefits (based on industry cases):

  • Performance: Cloud-native databases on cloud are 1-2x faster than self-managed Oracle (no network latency drag)
  • Cost: Save 60-70% (no Oracle licensing + managed saves DBA labor)
  • Availability: Managed DB multi-AZ auto-failover in 30 seconds, self-managed Oracle may take hours

Step 2 (1-2 months): Configure Cloud Load Balancer

Architecture Evolution Comparison:

❌ Wrong: Users → 1 ECS (16-core 32GB) → Total outage on failure, resize requires shutdown

✅ Right: Users → Cloud Load Balancer → 4-5 ECS (4-core 8GB) → Single failure auto-removed, scale anytime

Cost and Reliability Comparison (monthly):

Single-machine approach:

  • 1x 16-core 32GB ≈ Higher cost
  • Failure = total business outage
  • Scaling requires downtime

Load-balanced approach:

  • 4x 4-core 8GB + Load Balancer ≈ Potentially lower cost
  • Any single failure doesn't affect business
  • Scale up/down anytime

Implementation Steps:

  1. Create application image (including code and dependencies)
  2. Launch 3-5 ECS instances, quickly deploy using image
  3. Create cloud load balancer, configure listeners and backend server groups
  4. Configure health checks (HTTP GET /health or TCP port probe)
  5. Test failover: Manually stop one ECS, verify traffic automatically switches away

Step 3 (Long-term): Push Vendors to Transform

Collaboration model with software vendors must change:

Old model:

  • Enterprise: Here's SSH access, you handle it
  • Vendor: OK, I'll install software on ECS

New model:

  • Enterprise: Create IAM delegation, authorize you to use load balancer, managed DB, monitoring
  • Vendor: OK, I'll configure load balancing, connect to managed DB, set up monitoring alerts
  • Enterprise: I can see all operations in audit logs
  • Vendor: Revoke delegation with one click when project ends

Contract terms should also adjust:

  • Explicitly require vendors to use cloud-managed services (load balancer/managed DB), no self-building on ECS
  • Deliverables include IaC scripts (Terraform/CloudFormation), not manual deployment docs
  • Training requirements: Vendor teams must pass cloud certifications

Conclusion

Cloud migration is not moving apps to cloud data centers, but redesigning systems using cloud methods.

Many enterprises spent money migrating to the cloud but didn't enjoy cloud benefits:

  • Give vendors SSH access → Vendors continue treating ECS as physical machines
  • Oracle continues running on cloud disks → 100-500x higher latency, worse performance
  • Single-point deployment without load balancing → Total outage on failure

It's not that cloud doesn't work, but that you're still using data center methods with cloud.

Cloud value is not in the IaaS layer (virtual machines), but in the PaaS layer (managed services):

  • Replace database: PostgreSQL/cloud-native DB more suitable for cloud than Oracle, better performance and lower cost
  • Load balancing: Multiple instances cheaper than single machine, eliminates single point of failure
  • Cloudify permissions: IAM delegation more secure than SSH, operations auditable

Start taking action today:

  1. Evaluate current architecture: Are vendors still using SSH? Still installing open-source software on ECS?
  2. Learn cloud IAM delegation mechanisms
  3. Renegotiate collaboration model with vendors: require using cloud-managed services
  4. Migrate Oracle to managed PostgreSQL/cloud-native database
  5. Configure cloud load balancer
  6. Establish Console-based collaboration model, not SSH

Cloud is a tool, cloud-native is a mindset. Only when the mindset is in place can the tool deliver value.


Additional Note: Troubleshooting Doesn't Need SSH

Some enterprises worry: "If I don't give SSH, how do vendors troubleshoot issues?"

The answer: In cloud-native architecture, troubleshooting doesn't need SSH.

  • Application logs → Cloud log service unified collection, Console search
  • System monitoring → Cloud monitoring service, visualized CPU/memory/network
  • Fault diagnosis → Database diagnostics, slow query analysis
  • Performance optimization → APM application performance monitoring, call chain tracing

Through Console and APIs, vendors get diagnostic capabilities more powerful than SSH, and all operations are auditable and revocable.

SSH is last generation's operations method, Console delegation is the cloud era's collaboration model.


This article is based on verified technical parameters from cloud vendor documentation, suitable for enterprise IT decision-makers, CTOs, and architects who are migrating to cloud or have migrated but with poor results. Migrating Oracle to cloud-native databases is the trend, technically fully feasible — the earlier you act, the more proactive you are.

Top comments (0)