DEV Community

Dinesh_gowtham
Dinesh_gowtham

Posted on

Claude‑Generated Aurora Serverless v2 Scaling Policies: A Step‑By‑Step Guide with the RDS SDK

LLMs are no longer just for writing code snippets—they can design cloud infrastructure too. Imagine describing your traffic pattern in plain English and getting a ready‑to‑apply Aurora Serverless scaling policy. This post shows exactly how that works.

Why Aurora Serverless v2 Needs Dynamic Scaling

Aurora Serverless v2 (ASv2) is Amazon’s on‑demand relational database that automatically adds or removes Aurora Capacity Units (ACUs). An ACU is a bundle of CPU, memory, and networking that the database can use. Think of ACUs like seats on a bus: when more passengers (queries) show up, you add seats; when the bus empties, you remove seats to save money.

Without a policy that matches your real‑world traffic, two bad things happen:

  1. Under‑provisioning – the database cannot keep up, causing high latency or throttling.
  2. Over‑provisioning – you pay for capacity you never use. Remember, ASv2 has a minimum of 0.5 ACU which still costs roughly $50 / month.

A dynamic scaling policy tells the database when to grow (increase max ACU) and when to shrink (decrease min ACU). The policy is expressed as a JSON block that the RDS API understands.

In plain English: A good scaling policy is like a thermostat for your database—it heats up (adds capacity) when it gets cold (high load) and cools down (removes capacity) when things calm.

Prompting Claude to Design a Scaling Policy

Claude is Anthropic’s conversational model. By feeding it a short description of recent metrics, we can ask it to suggest a safe minACU and maxACU. The prompt should contain:

  • The average CPU utilization over the last 5 minutes.
  • The peak number of connections seen in the same window.
  • Any cost constraints you care about (e.g., “keep monthly cost under $200”).

Here’s a minimal prompt we can send via HTTP:

{
  "model": "claude-3-5-sonnet-20241022",
  "max_tokens": 500,
  "messages": [
    {
      "role": "user",
      "content": "We run an Aurora Serverless v2 cluster named prod-db. In the last 5 minutes CPU was 68 % on average and we saw a peak of 220 concurrent connections. Our budget allows up to $150 per month for this cluster. Suggest a minACU and maxACU that keep latency low but stay within budget. Return only a JSON object with keys minACU and maxACU."
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Claude replies with something like:

{
  "minACU": 2,
  "maxACU": 8
}
Enter fullscreen mode Exit fullscreen mode

Why this works: The model has read the Aurora documentation and learned typical ACU ranges for given workloads. By asking for only the JSON, we make parsing easy and avoid extra text that could confuse our script.

Tip: Always ask Claude to return only the data you need. Adding “Return only a JSON object …” removes the chance of receiving explanatory sentences that break your parser.

Gotcha: Quota Mismatch

Claude may suggest a maxACU that exceeds the ACU quota for your AWS account, or it might accidentally flip minACU > maxACU. The RDS API will then throw an InvalidParameterCombinationException. We’ll guard against that later.

Translating Claude’s Output into RDS SDK Calls

Now we move from a plain JSON response to an actual AWS API call. We’ll use the @aws-sdk/client-rds package, the official JavaScript/TypeScript client for Amazon RDS.

First, install the SDK (run once):

npm install @aws-sdk/client-rds
Enter fullscreen mode Exit fullscreen mode

Next, write a small function that:

  1. Validates the numbers (ensures minACU ≤ maxACU and both are within your account limits).
  2. Creates a new DBClusterParameterGroup that holds the scaling configuration.
  3. Calls modifyDBCluster to attach the new parameter group.
// scaling.ts
import {
  RDSClient,
  CreateDBClusterParameterGroupCommand,
  ModifyDBClusterCommand,
  DBClusterParameterGroup,
} from "@aws-sdk/client-rds";

// Load credentials from the environment (AWS_ACCESS_KEY_ID, etc.)
const rds = new RDSClient({ region: "us-east-1" });

/**
 * Ensures the ACU values are sensible.
 * Returns a tuple [validMin, validMax] or throws an error.
 */
function validateAcus(min: number, max: number, quota: number): [number, number] {
  // Aurora Serverless v2 minimum is 0.5 ACU
  if (min < 0.5) throw new Error("minACU cannot be below 0.5");
  // Ensure min ≤ max
  if (min > max) throw new Error(`minACU (${min}) cannot be greater than maxACU (${max})`);
  // Make sure we don’t exceed the account quota
  if (max > quota) throw new Error(`maxACU (${max}) exceeds your ACU quota of ${quota}`);
  return [min, max];
}

/**
 * Builds a parameter group name that is unique per deployment.
 */
function buildGroupName(base: string): string {
  const timestamp = Date.now().toString(36);
  return `${base}-${timestamp}`;
}

/**
 * Main function that receives Claude’s JSON and applies the scaling policy.
 */
export async function applyScalingPolicy(
  clusterIdentifier: string,
  raw: { minACU: number; maxACU: number },
  accountAcusQuota: number // fetch from Service Quotas API or hard‑code for demo
) {
  // 1️⃣ Validate the numbers
  const [minACU, maxACU] = validateAcus(raw.minACU, raw.maxACU, accountAcusQuota);

  // 2️⃣ Create a new parameter group that holds the scaling settings
  const groupName = buildGroupName("aurora-scaling");
  const createCmd = new CreateDBClusterParameterGroupCommand({
    DBClusterParameterGroupName: groupName,
    DBParameterGroupFamily: "aurora-mysql8.0", // adjust to your engine
    Description: `Auto‑generated scaling group for ${clusterIdentifier}`,
    // Parameter values are key/value pairs; the scaling keys are documented by AWS
    Parameters: [
      { ParameterName: "aurora_serverless_max_capacity", ParameterValue: maxACU.toString() },
      { ParameterName: "aurora_serverless_min_capacity", ParameterValue: minACU.toString() },
    ],
  });

  await rds.send(createCmd);
  console.log(`Created parameter group ${groupName}`);

  // 3️⃣ Attach the new group to the DB cluster
  const modifyCmd = new ModifyDBClusterCommand({
    DBClusterIdentifier: clusterIdentifier,
    DBClusterParameterGroupName: groupName,
    ApplyImmediately: true, // For demo; in production you may prefer a maintenance window
  });

  await rds.send(modifyCmd);
  console.log(`Applied scaling policy to ${clusterIdentifier}`);
}
Enter fullscreen mode Exit fullscreen mode

Key points:

  • aurora_serverless_min_capacity and aurora_serverless_max_capacity are the exact parameter names that control ASv2 scaling.
  • ApplyImmediately: true forces the change now; otherwise you must wait for the next maintenance window.
  • The function logs each step so you can see what happened in CloudWatch logs.

Takeaway: Turning Claude’s plain numbers into a working policy requires three safety nets – validation, a unique parameter group, and an explicit modifyDBCluster call.

Testing the Policy with a Safe Dry‑Run

Before we touch a production cluster, we should simulate the whole flow using a dry‑run flag. The AWS SDK does not provide a built‑in dry‑run for RDS, but we can mimic it by:

  1. Skipping the actual send calls.
  2. Printing the exact request objects that would be sent.
  3. Running the validation logic to catch obvious errors.

Add a simple wrapper around applyScalingPolicy:

// dryRun.ts
import { applyScalingPolicy } from "./scaling";

/**
 * Executes the scaling workflow without making real API calls.
 */
export async function dryRun(
  clusterId: string,
  claudeJson: { minACU: number; maxACU: number },
  quota: number
) {
  try {
    // Run validation – this will throw if numbers are bad
    const [min, max] = (await (async () => {
      // Re‑use the internal validation (exposed for demo)
      const { validateAcus } = await import("./scaling");
      return validateAcus(claudeJson.minACU, claudeJson.maxACU, quota);
    })()) as [number, number];

    console.log("✅ Validation passed");
    console.log(`🧪 Would create parameter group with minACU=${min}, maxACU=${max}`);

    // Show the request bodies that would be sent
    const groupName = `dryrun-${Date.now()}`;
    console.log("🧾 CreateDBClusterParameterGroupCommand payload:");
    console.log({
      DBClusterParameterGroupName: groupName,
      DBParameterGroupFamily: "aurora-mysql8.0",
      Description: `Dry‑run group for ${clusterId}`,
      Parameters: [
        { ParameterName: "aurora_serverless_min_capacity", ParameterValue: min.toString() },
        { ParameterName: "aurora_serverless_max_capacity", ParameterValue: max.toString() },
      ],
    });

    console.log("🧾 ModifyDBClusterCommand payload:");
    console.log({
      DBClusterIdentifier: clusterId,
      DBClusterParameterGroupName: groupName,
      ApplyImmediately: true,
    });

    console.log("🔍 Dry‑run complete – no changes made.");
  } catch (e) {
    console.error("❌ Dry‑run failed:", e);
  }
}
Enter fullscreen mode Exit fullscreen mode

Run it with Node 22 (which ships with the native fetch API for the Claude call) and you’ll see a clear picture of what would happen, without paying for a new parameter group or risking an invalid configuration.

Helpful tip: Keep the dry‑run script in version control. Treat it like a unit test for your infrastructure‑as‑code pipeline.

Real‑world Gotcha: Snapshot Restore Time

If you ever need to roll back a bad scaling change, restoring a snapshot can take 5–20 minutes. That delay is a common surprise during incidents. A dry‑run helps you catch mistakes before they land on a live cluster, saving you from a lengthy rollback.

Deploying the Policy via Node.js and RDS Proxy

In production we usually run the scaling script inside an AWS Lambda function (Node.js 22 runtime). To keep connection latency low, we place an RDS Proxy in front of the Aurora cluster. The proxy pools connections, reducing the “too many connections” error that many Lambda‑driven services see.

Setting Up the Proxy (quick recap)

  1. In the RDS console, create a Proxy targeting your Aurora Serverless v2 cluster.
  2. Grant the Lambda execution role the rds-db:connect permission for the proxy.
  3. Update the Lambda environment variable DB_PROXY_ENDPOINT with the proxy endpoint.

Analogy: Think of the proxy as a receptionist who hands out a limited number of phones to callers. Instead of each caller dialing the main line (which would overload it), they use the receptionist’s pool of phones.

Full Lambda Handler

Below is a complete Lambda handler that:

  • Calls Claude with the latest metrics (simulated here).
  • Parses the JSON response.
  • Performs a dry‑run first (optional, can be toggled with an env var).
  • Applies the policy using the RDS SDK.
// handler.ts
import { APIGatewayProxyEvent, APIGatewayProxyResult } from "aws-lambda";
import { RDSClient, CreateDBClusterParameterGroupCommand, ModifyDBClusterCommand } from "@aws-sdk/client-rds";
import fetch from "node-fetch"; // Node 22 includes fetch, but we import for type safety

// Environment variables (set in Lambda config)
const {
  CLAUDE_API_KEY,
  CLUSTER_ID,
  ACU_QUOTA = "16", // default quota if you haven’t fetched it
  DRY_RUN = "true", // toggle dry‑run with "false"
} = process.env;

const rds = new RDSClient({ region: "us-east-1" });

/**
 * Helper to ask Claude for a scaling recommendation.
 */
async function askClaude(prompt: string): Promise<{ minACU: number; maxACU: number }> {
  const response = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "x-api-key": CLAUDE_API_KEY!,
      "Content-Type": "application/json",
      "anthropic-version": "2023-06-01",
    },
    body: JSON.stringify({
      model: "claude-3-5-sonnet-20241022",
      max_tokens: 300,
      messages: [{ role: "user", content: prompt }],
    }),
  });

  if (!response.ok) {
    const txt = await response.text();
    throw new Error(`Claude API error: ${response.status} ${txt}`);
  }

  const data = await response.json();
  // Claude returns a JSON string inside `content`. Extract it.
  const raw = JSON.parse(data.content[0].text);
  return { minACU: Number(raw.minACU), maxACU: Number(raw.maxACU) };
}

/**
 * Core Lambda handler.
 */
export const handler = async (event: APIGatewayProxyEvent): Promise<APIGatewayProxyResult> => {
  try {
    // 1️⃣ Gather recent metrics – in a real system you would read CloudWatch.
    const cpuAvg = 68; // placeholder
    const peakConns = 220; // placeholder

    const prompt = `We run an Aurora Serverless v2 cluster named ${CLUSTER_ID}. In the last 5 minutes CPU was ${cpuAvg}% on average and we saw a peak of ${peakConns} concurrent connections. Our budget allows up to $150 per month for this cluster. Suggest a minACU and maxACU that keep latency low but stay within budget. Return only a JSON object with keys minACU and maxACU.`;

    // 2️⃣ Ask Claude
    const recommendation = await askClaude(prompt);
    console.log("Claude recommendation:", recommendation);

    // 3️⃣ Validate numbers (reuse logic from scaling.ts)
    const { validateAcus } = await import("./scaling");
    const [minACU, maxACU] = validateAcus(
      recommendation.minACU,
      recommendation.maxACU,
      Number(ACU_QUOTA)
    );

    // 4️⃣ Dry‑run? If DRY_RUN is true, just log the would‑be calls.
    if (DRY_RUN?.toLowerCase() === "true") {
      console.log("🔎 Dry‑run mode – not touching AWS.");
      console.log(`Would create parameter group with minACU=${minACU}, maxACU=${maxACU}`);
      return {
        statusCode: 200,
        body: JSON.stringify({ message: "Dry‑run successful", minACU, maxACU }),
      };
    }

    // 5️⃣ Create parameter group
    const groupName = `auto-scaling-${Date.now()}`;
    await rds.send(
      new CreateDBClusterParameterGroupCommand({
        DBClusterParameterGroupName: groupName,
        DBParameterGroupFamily: "aurora-mysql8.0",
        Description: `Auto‑generated scaling group for ${CLUSTER_ID}`,
        Parameters: [
          { ParameterName: "aurora_serverless_min_capacity", ParameterValue: minACU.toString() },
          { ParameterName: "aurora_serverless_max_capacity", ParameterValue: maxACU.toString() },
        ],
      })
    );
    console.log(`Created parameter group ${groupName}`);

    // 6️⃣ Attach it to the cluster (via the proxy endpoint)
    await rds.send(
      new ModifyDBClusterCommand({
        DBClusterIdentifier: CLUSTER_ID,
        DBClusterParameterGroupName: groupName,
        ApplyImmediately: true,
      })
    );
    console.log(`Applied new scaling policy to ${CLUSTER_ID}`);

    return {
      statusCode: 200,
      body: JSON.stringify({ message: "Scaling policy applied", minACU, maxACU }),
    };
  } catch (err) {
    console.error("Error in scaling handler:", err);
    return {
      statusCode: 500,
      body: JSON.stringify({ error: (err as Error).message }),
    };
  }
};
Enter fullscreen mode Exit fullscreen mode

Why we use RDS Proxy:

When the Lambda runs, each invocation could open a fresh TCP connection. Without a proxy, a burst of invocations could exceed the max connections limit of the Aurora cluster, leading to the Too many connections error. The proxy keeps a small pool (often 5‑10 connections) and re‑uses them, keeping latency low (the added overhead is typically 1‑3 ms, which is acceptable for most p99 latency budgets).

Key tip: Keep the ApplyImmediately flag true only for non‑critical clusters. Production workloads that cannot tolerate a brief pause should schedule the change in a maintenance window instead.

The Takeaway

  • Dynamic scaling prevents both performance bottlenecks and unnecessary spend in Aurora Serverless v2.
  • Claude can translate plain‑English traffic descriptions into concrete minACU/maxACU values, but you must still validate the numbers against your account quota.
  • The @aws-sdk/client-rds package lets you create a parameter group and attach it to a DB cluster with just a few lines of TypeScript.
  • Dry‑run scripts act as a safety net, printing the exact API payloads before any real changes hit your database.
  • RDS Proxy removes the “too many connections” problem for Lambda‑driven workloads, adding only a few milliseconds of latency.
  • Always test (dry‑run → staging → production) and remember that restoring a snapshot can take minutes, so catching errors early saves time and money.

By following these steps, you can let a language model do the heavy mental lifting while you keep the final guardrails firmly in your code. Happy scaling!


Transparency notice

This article was written with the help of an AI system — Groq (GPT OSS 120B).

Published: 2026-10-08 · Primary focus: RDS

All code blocks are intended to be correct and runnable, but please verify them
against the official docs for the tools mentioned before using in production.

Find an error? Drop a comment — corrections are always welcome.

Top comments (0)