DEV Community

Dinesh Kumar Sarangapani
Dinesh Kumar Sarangapani

Posted on Originally published at dineshkumars.dev

Scaling 10,000+ AI Batch Jobs with EKS Auto Mode and Karpenter

Building a Retrieval-Augmented Generation (RAG) platform looks deceptively simple in proof-of-concept tutorials. You read a PDF, split it into chunks, call an embedding API, and store the vectors in a database.

In an enterprise, scale changes the equation completely.

When project teams upload hundreds of thousands of technical documents, contracts, and manuals, your platform experiences massive, unpredictable traffic spikes:

  • During peak business hours, users queue 20,000 documents for indexing.
  • Overnight and on weekends, the queue sits empty.

If you provision enough on-demand cloud servers to process the peak 20,000-document backlog immediately, you pay thousands of dollars for idle compute during off-peak periods.

If you downsize your cluster to save money, documents sit in queue backlogs for hours, frustrating users who expect instant search availability.

Here is the idea we used to solve this on Amazon EKS, how it performed in production, and what to watch out for.


The Idea: Event-Driven Queue Scaling with Spot Compute

To achieve both high throughput and extreme cost efficiency, we combined three architectural patterns:

  1. Decoupled Asynchronous Ingestion: Document uploads are decoupled from processing. When files arrive, lightweight ingest endpoints place processing jobs onto high-throughput cloud queues.
  2. Backlog-Driven Pod Autoscaling (KEDA): Rather than scaling based on lagging CPU or memory metrics, we scale worker pods based directly on queue depth. If thousands of jobs land in the queue, autoscalers immediately request hundreds of worker pods within seconds, and scale completely down to zero when the queue clears.
  3. Just-In-Time Node Provisioning (Karpenter Spot): When hundreds of worker pods become pending, an intelligent node provisioner communicates directly with the cloud fleet API. It dynamically selects from a diverse pool of discounted Spot instance types and architectures, launching right-sized nodes in under 45 seconds.

How It Worked Well

  1. 70% to 90% Cost Reduction: Spot instances trade availability guarantees for massive discounts. Because batch document parsing is naturally fault-tolerant, running batch workers on Spot instances reduced our compute bill by up to 90% compared to on-demand instances.
  2. Instant Queue Draining: When thousands of documents are uploaded simultaneously, the autoscaler skips gradual ramp-up intervals, jumping from zero to hundreds of concurrent pods. The node provisioner launches dozens of varied EC2 instances in parallel, draining massive backlogs in minutes instead of hours.
  3. True Scale-to-Zero Efficiency: When queues empty, pods scale to zero, and the node provisioner automatically consolidates and terminates the underlying worker nodes within seconds. We pay zero compute costs when no documents are being processed.
  4. Architectural Flexibility Across Chipsets: By allowing the node provisioner to schedule across both standard x86 and ARM-based Graviton chipsets, the system automatically selects the cheapest available capacity in each cloud region.

What to Watch Out For

  1. Handling Spot Interruptions Gracefully: Spot instances can be reclaimed by cloud providers with a two-minute warning. If a node is reclaimed while a worker is halfway through parsing a 100-page document, the job must not be lost. Configure queue visibility timeouts and dead-letter queues so that interrupted jobs become visible for other workers to retry automatically.
  2. Avoid Throttling Downstream Embedding APIs: When scaling from 2 pods to 500 pods in 60 seconds, your workers will make hundreds of concurrent calls to your embedding model endpoints. If you do not configure rate-limiting and connection pooling, you will overwhelm your upstream model gateway with HTTP 429 errors.
  3. Spot Pool Diversification: If you restrict your node provisioner to only one or two specific instance types (e.g. only c6i.4xlarge), your cluster will experience provisioning delays when cloud capacity pools run low. Always configure your provisioners to draw from dozens of instance families across multiple availability zones.
  4. Node Startup Optimization: A pod cannot start until its container image is pulled. For large AI worker containers (which often bundle OCR libraries and Python dependencies), pull times can take several minutes. Optimize your container images by stripping unnecessary packages, using lightweight base images, and taking advantage of node caching or fast-launch disk snapshots.

Top comments (0)