DEV Community

Usman Khan
Usman Khan

Posted on Originally published at ctousman.com

Cloud Infrastructure Cost Optimization: FinOps Engineering, Graviton Migrations, and Auto-Scaling

Cloud cost runaway is the silent margin killer in scaling B2B SaaS platforms. Unoptimized auto-scaling groups, over-provisioned database instances, and idle compute nodes rapidly erode gross margins.

Here is an actionable FinOps engineering framework for migrating to ARM64, tuning Kubernetes pod bin-packing, and enforcing compute ceiling controls.


1. The Financial Engineering Framework: Unit Economics Over Flat Bills

Engineers often evaluate cloud costs by looking at total monthly spend. High-agency CTOs instead manage infrastructure spend through Cost Per Active Tenant (CPAT) or Cost Per API Transaction.

If infrastructure spend grows in line with revenue, gross margins stay healthy. If cloud costs outpace transaction growth, architectural inefficiency is taking hold.

  • Identify idle waste: EC2 instances or EKS nodes running under 15% average CPU/memory utilization.
  • Right-size compute: Shift workloads from memory-optimized (r6i) to compute or balanced types (c7g / m7g) based on profiling data.
  • Eliminate unused storage: Unattached EBS volumes, S3 buckets without lifecycle transition rules, and orphaned snapshot chains.

2. Architectural Leverage: x86 to Graviton (ARM64) Migration

Moving backend services from Intel/AMD x86 to AWS Graviton (ARM64) typically yields around a 20% cost reduction and up to 40% better price-performance. For containerized services (Node.js, Go, Python, Java), multi-architecture Docker builds make deployment straightforward.

Multi-Arch Dockerfile for Graviton

# Build stage runs natively on the CI machine's architecture
FROM --platform=$BUILDPLATFORM node:20-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci                      # dev dependencies needed for the build
COPY . .
RUN npm run build

# Runtime stage is built per target architecture (amd64 / arm64)
FROM node:20-alpine AS runner
WORKDIR /app
ENV NODE_ENV=production
COPY package*.json ./
RUN npm ci --omit=dev           # native modules compiled for the target arch
COPY --from=builder /app/dist ./dist

EXPOSE 3000
USER node
CMD ["node", "dist/main.js"]
Enter fullscreen mode Exit fullscreen mode

Build and push multi-platform images directly from your CI/CD pipeline:

docker buildx build \
  --platform linux/amd64,linux/arm64 \
  -t registry.internal.net/saas-api:v2.4.0 \
  --push .
Enter fullscreen mode Exit fullscreen mode

3. Kubernetes Pod Bin-Packing and Karpenter Auto-Scaling

The traditional Kubernetes Cluster Autoscaler (CAS) reacts slowly and often leaves nodes underutilized because it works with static node groups. Modern platforms use Karpenter for consolidated, right-sized node provisioning on demand.

Optimizing Pod Resource Requests

If a pod requests 2 vCPUs and 4 GB RAM but only uses 200m CPU at peak, the scheduler still reserves the full 2 vCPUs, which forces unnecessary node scale-outs.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: billing-pipeline-worker
  namespace: production
spec:
  replicas: 12
  selector:
    matchLabels:
      app: billing-pipeline-worker
  template:
    metadata:
      labels:
        app: billing-pipeline-worker
    spec:
      containers:
        - name: worker
          image: registry.internal.net/billing-worker:v1.8.2
          resources:
            requests:
              cpu: "250m"       # Minimal baseline for tight bin-packing
              memory: "512Mi"
            limits:
              cpu: "1000m"      # Burst capacity during high throughput
              memory: "1024Mi"
Enter fullscreen mode Exit fullscreen mode

4. Spot Instance Orchestration for Fault-Tolerant Queue Workers

Stateless background workers, asynchronous LLM processors, and batch ETL jobs should run almost entirely on Spot Instances, which offer up to 90% off On-Demand pricing.

Handling Spot Interruptions Gracefully

  • Diversify Spot pools: Never restrict a Spot group to a single instance type. Configure Karpenter or your ASGs to draw from at least 4–8 instance types (e.g. c6g.large, c7g.large, m6g.large, m7g.large).
  • React to interruption signals: Handle both the EC2 Rebalance Recommendation (an early warning) and the 2-minute Spot Interruption Notice via EventBridge, so you can drain active tasks before the node is terminated.
  • Make tasks idempotent: Every job in BullMQ, SQS, or RabbitMQ should be safe to interrupt mid-execution and retry on another worker without duplicating business side effects.

Originally published at ctousman.com.

About the author: I'm Usman Khan, a Fractional CTO & Systems Architect. I advise high-growth SaaS platforms on backend performance, distributed systems, and cloud cost efficiency. Need an architectural audit of your infrastructure? Book a 30-min strategy call.

Top comments (0)