DEV Community

Dinesh_gowtham
Dinesh_gowtham

Posted on

How to Build a RAG Pipeline with DynamoDB’s Native Vector Search, Explained Simply

Most engineers reach for a dedicated vector database the moment they need similarity search, adding latency and operational overhead. DynamoDB just announced native vector indexes, letting you store embeddings alongside your business data. In this post you’ll see how to turn that new feature into a fully‑functional Retrieval‑Augmented Generation (RAG) flow with Claude, using only DynamoDB and plain Node.js.

In plain English – DynamoDB now does the math to find “similar looking” vectors, so you no longer need a separate service just to ask “which pieces of text look like this query?”


Why DynamoDB’s Vector Search Changes the RAG Landscape

The “why” first

A Retrieval‑Augmented Generation (RAG) system works in two steps:

  1. Retrieve the most relevant text chunks from a knowledge store.
  2. Generate a response that weaves those chunks together, usually with a large language model (LLM) like Claude.

Historically the retrieve step required a vector store—a specialized database that can compare high‑dimensional vectors (numeric representations of text) quickly. Adding a second store meant:

  • extra network hops (higher latency)
  • two places to back up (more operational work)
  • consistency headaches (the DB and the vector store can drift apart)

DynamoDB’s new vector index removes the separate store. Your embeddings live right next to the rest of the item, and DynamoDB’s built‑in KNN (k‑nearest‑neighbors) engine does the similarity math for you.

Key takeaway – Cutting the stack in half also cuts the chances of data mismatches and reduces the time it takes to answer a user query.

Quick vocab recap

  • Embedding – a list of numbers that captures the meaning of a piece of text; think of it as a fingerprint.
  • Cosine similarity – a way to measure how close two fingerprints are, ranging from -1 (opposite) to 1 (identical).
  • KNN (k‑nearest‑neighbors) – an algorithm that returns the k most similar items to a query vector.
  • GSI (Global Secondary Index) – a secondary view of a DynamoDB table that can be queried with different keys or, now, with a vector index.

Setting Up a Vector‑Enabled Table in DynamoDB

The “why” first

Before we can store or search embeddings, DynamoDB needs a table that knows where the vector lives and how to compare it. This is done by creating a GSI with a vector attribute type and declaring that the index should use cosine similarity.

Step‑by‑step code

// src/createTable.ts
import {
  DynamoDBClient,
  CreateTableCommand,
} from "@aws-sdk/client-dynamodb";

// The low‑level client is used for the CreateTable call because
// the lib‑dynamodb wrapper does not expose table‑creation APIs.
const client = new DynamoDBClient({ region: "us-east-1" });

async function createVectorTable() {
  // Table name – keep it short and unique.
  const TableName = "RagDocs";

  // Primary key (partition key) – we use a simple UUID string.
  const KeySchema = [{ AttributeName: "docId", KeyType: "HASH" }];

  // The attribute that will hold the embedding as a binary blob.
  const AttributeDefinitions = [
    { AttributeName: "docId", AttributeType: "S" }, // S = String
    { AttributeName: "embedding", AttributeType: "B" }, // B = Binary
  ];

  // Global Secondary Index (GSI) that enables vector search.
  const GlobalSecondaryIndexes = [
    {
      IndexName: "EmbeddingKNN",
      // No projection needed for this demo; we pull the whole item later.
      Projection: { ProjectionType: "ALL" },
      // The GSI's key schema uses a placeholder attribute; the
      // vector is supplied at query time, not stored in the index key.
      KeySchema: [{ AttributeName: "docId", KeyType: "HASH" }],
      // Vector configuration tells DynamoDB to treat `embedding` as a
      // 1536‑dimensional Float32 vector and to compare using cosine similarity.
      // The `VectorConfig` field is only available on the GSI definition.
      VectorConfig: {
        VectorDimensions: 1536,
        VectorDataType: "FLOAT32", // 32‑bit floating point numbers
        MetricType: "COSINE", // similarity metric
        VectorField: "embedding", // attribute that holds the binary vector
      },
      // Provisioned read capacity – adjust for your traffic.
      ProvisionedThroughput: { ReadCapacityUnits: 5, WriteCapacityUnits: 5 },
    },
  ];

  const command = new CreateTableCommand({
    TableName,
    KeySchema,
    AttributeDefinitions,
    BillingMode: "PROVISIONED", // explicit to control cost
    ProvisionedThroughput: { ReadCapacityUnits: 5, WriteCapacityUnits: 5 },
    GlobalSecondaryIndexes,
    // Optional TTL (time‑to‑live) for automatic cleanup of old chunks.
    TimeToLiveSpecification: {
      AttributeName: "expiresAt", // epoch seconds
      Enabled: true,
    },
  });

  try {
    const response = await client.send(command);
    console.log("Table created:", response.TableDescription?.TableName);
  } catch (err) {
    console.error("Failed to create table:", err);
  }
}

createVectorTable();
Enter fullscreen mode Exit fullscreen mode

Explanation of tricky bits

  • VectorConfig lives inside the GSI definition; DynamoDB uses it at query time, not at write time.
  • The embedding attribute must be declared as Binary (B) because the service expects a base64‑encoded Float32Array.
  • TTL does not delete instantly; items can stick around up to 48 hours after the expiration timestamp.

Tip – Even though the GSI has a dummy hash key (docId), DynamoDB still requires it. Think of it as a “placeholder seat” that lets the engine focus on the vector field.


Storing and Querying Embeddings with the AWS SDK

The “why” first

Now that the table exists, we need to persist text chunks together with their embeddings. The embeddings are produced elsewhere (e.g., OpenAI’s embedding endpoint) and are 1536‑dimensional Float32 vectors. DynamoDB expects those vectors as binary data, so we must encode them as base64 before sending.

When we later need the most relevant chunks for a user question, we will:

  1. Convert the question into an embedding (same model, same dimensions).
  2. Run a KNN query against the EmbeddingKNN GSI, asking for the top 3 nearest neighbors.

Code – inserting a document

// src/putDocument.ts
import {
  DynamoDBDocumentClient,
  PutCommand,
} from "@aws-sdk/lib-dynamodb";
import { DynamoDBClient } from "@aws-sdk/client-dynamodb";

// Low‑level client needed only for the Document wrapper.
const ddbClient = new DynamoDBClient({ region: "us-east-1" });
const ddbDoc = DynamoDBDocumentClient.from(ddbClient);

// Helper: turn a Float32Array into a base64 string.
function encodeEmbedding(vec: Float32Array): string {
  // Convert the raw bytes to a Uint8Array, then to a base64 string.
  const bytes = new Uint8Array(vec.buffer);
  return Buffer.from(bytes).toString("base64");
}

// Example: a single text chunk and its embedding.
async function putChunk(
  docId: string,
  chunkText: string,
  embedding: Float32Array,
  ttlSeconds: number // optional expiration
) {
  const command = new PutCommand({
    TableName: "RagDocs",
    Item: {
      docId, // partition key
      chunk: chunkText,
      embedding: encodeEmbedding(embedding), // stored as Binary (base64)
      // DynamoDB treats a base64 string as Binary automatically.
      // Adding a TTL helps keep the table size in check.
      expiresAt: Math.floor(Date.now() / 1000) + ttlSeconds,
    },
  });

  try {
    await ddbDoc.send(command);
    console.log(`Chunk ${docId} stored`);
  } catch (err) {
    console.error("Write error:", err);
  }
}

/* -------------------------------------------------------------
   Imagine we already fetched an embedding from OpenAI:
   const embedding = await getOpenAIEmbedding(chunkText);
   ------------------------------------------------------------- */
Enter fullscreen mode Exit fullscreen mode

Code – querying the top‑3 similar chunks

// src/queryKnn.ts
import {
  DynamoDBDocumentClient,
  QueryCommand,
} from "@aws-sdk/lib-dynamodb";
import { DynamoDBClient } from "@aws-sdk/client-dynamodb";

const ddbClient = new DynamoDBClient({ region: "us-east-1" });
const ddbDoc = DynamoDBDocumentClient.from(ddbClient);

// Same encoder used for writes; DynamoDB expects the query vector
// in the same binary format.
function encodeEmbedding(vec: Float32Array): string {
  const bytes = new Uint8Array(vec.buffer);
  return Buffer.from(bytes).toString("base64");
}

/**
 * Returns the `k` most similar chunks for a given query embedding.
 */
async function knnSearch(
  queryEmbedding: Float32Array,
  k: number = 3
) {
  const command = new QueryCommand({
    TableName: "RagDocs",
    IndexName: "EmbeddingKNN", // the vector‑enabled GSI
    // The special `VectorSearch` field tells DynamoDB to perform a
    // KNN lookup using the supplied binary vector.
    // `KNN` is the number of neighbors to return.
    // `Metric` defaults to the one defined in the GSI (COSINE).
    VectorSearch: {
      KNN: k,
      VectorField: "embedding",
      QueryVector: encodeEmbedding(queryEmbedding),
    },
    // ConsistentRead is not allowed on GSI; we accept eventual consistency.
    // This matches the reality of DynamoDB’s GSI behavior.
  });

  try {
    const result = await ddbDoc.send(command);
    // Each item contains the original chunk text plus the embedding.
    return result.Items?.map((item) => ({
      docId: item.docId,
      chunk: item.chunk,
    })) ?? [];
  } catch (err) {
    console.error("KNN query error:", err);
    return [];
  }
}
Enter fullscreen mode Exit fullscreen mode

Gotchas highlighted

  • Binary encoding – If you forget to base64‑encode the Float32Array, DynamoDB will reject the write with a ValidationException: Attribute type mismatch.
  • GSI attribute type – The attribute must be declared as Binary (B). Declaring it as a string leads to the same validation error.
  • Eventual consistency – Reads from a GSI are eventually consistent, so a newly written chunk may not appear instantly in a KNN query. In high‑throughput systems, add a small retry loop or write‑through cache.

In plain English – Think of the embedding as a secret code. DynamoDB only understands the code when it’s wrapped in a base64 envelope; missing the envelope makes the service shout “I don’t know this type!”.


Calling Claude with Retrieved Context via Node.js fetch

The “why” first

The RAG pattern finishes by feeding the retrieved chunks to an LLM so it can craft a response. Claude’s HTTP API is straightforward: send a JSON payload with a messages array and receive a completion field. We’ll use Node 22’s built‑in fetch (no extra packages needed).

Minimal fetch wrapper

// src/callClaude.ts
/**
 * Sends a user question together with retrieved context chunks to Claude.
 * Returns Claude’s answer as plain text.
 */
async function askClaude(
  question: string,
  contextChunks: { chunk: string }[],
  apiKey: string
): Promise<string> {
  // Build a single system prompt that concatenates the chunks.
  const systemPrompt = `You are an assistant that answers questions using only the following information:\n\n${contextChunks
    .map((c) => `- ${c.chunk}`)
    .join("\n")}\n\nIf the answer cannot be derived from this data, say "I don't know."`;

  const payload = {
    model: "claude-3-5-sonnet-20241007", // example model name
    messages: [
      { role: "system", content: systemPrompt },
      { role: "user", content: question },
    ],
    max_tokens: 500,
  };

  const response = await fetch(
    "https://api.anthropic.com/v1/messages",
    {
      method: "POST",
      headers: {
        "Content-Type": "application/json",
        // Anthropic expects the key in the `x-api-key` header.
        "x-api-key": apiKey,
      },
      body: JSON.stringify(payload),
    }
  );

  if (!response.ok) {
    const errorBody = await response.text();
    throw new Error(`Claude API error ${response.status}: ${errorBody}`);
  }

  const data = await response.json();
  // The completion text lives in `content[0].text` for most responses.
  return data.content?.[0]?.text ?? "No response";
}
Enter fullscreen mode Exit fullscreen mode

Tip – Keep the system prompt short; Claude’s token limit includes both the prompt and the answer, so overly long context can truncate the model’s output.


Putting It All Together: A Minimal End‑to‑End RAG Service

The “why” first

Having separate snippets is useful for learning, but production code needs a single orchestrator that:

  1. Receives a user query (e.g., via an HTTP endpoint).
  2. Generates an embedding for the query.
  3. Retrieves the top‑3 matching chunks from DynamoDB.
  4. Calls Claude with those chunks and returns the answer.

Below is a compact Node.js script that ties everything together. It uses the same SDKs and fetch we already introduced, plus the OpenAI embedding endpoint (you can swap it for any model that returns a 1536‑dimensional vector).

Full flow script

// src/ragService.ts
import { DynamoDBDocumentClient } from "@aws-sdk/lib-dynamodb";
import { DynamoDBClient } from "@aws-sdk/client-dynamodb";
import { knnSearch } from "./queryKnn";
import { askClaude } from "./callClaude";
import fetch from "node-fetch"; // Node 22 includes global fetch, but keep for type safety

// Re‑use the document client we created earlier.
const ddbClient = new DynamoDBClient({ region: "us-east-1" });
const ddbDoc = DynamoDBDocumentClient.from(ddbClient);

// Replace with your own keys.
const OPENAI_API_KEY = process.env.OPENAI_API_KEY!;
const CLAUDE_API_KEY = process.env.CLAUDE_API_KEY!;

/**
 * Calls OpenAI's embedding endpoint to turn text into a Float32Array.
 */
async function embed(text: string): Promise<Float32Array> {
  const response = await fetch(
    "https://api.openai.com/v1/embeddings",
    {
      method: "POST",
      headers: {
        "Content-Type": "application/json",
        Authorization: `Bearer ${OPENAI_API_KEY}`,
      },
      body: JSON.stringify({
        model: "text-embedding-3-large", // returns 1536‑dim vectors
        input: text,
      }),
    }
  );

  if (!response.ok) {
    const err = await response.text();
    throw new Error(`OpenAI embed error ${response.status}: ${err}`);
  }

  const data = await response.json();
  // OpenAI returns an array of numbers; convert to Float32Array.
  const numbers: number[] = data.data[0].embedding;
  return new Float32Array(numbers);
}

/**
 * Main entry point – receives a user question and returns Claude's answer.
 */
export async function answerQuestion(question: string): Promise<string> {
  // 1️⃣ Turn the question into a vector.
  const queryVec = await embed(question);

  // 2️⃣ Find the 3 most similar stored chunks.
  const matches = await knnSearch(queryVec, 3);

  if (matches.length === 0) {
    return "I couldn't find any relevant information.";
  }

  // 3️⃣ Pass the question + context to Claude.
  const answer = await askClaude(question, matches, CLAUDE_API_KEY);
  return answer;
}

/* -------------------------------------------------------------
   Example usage as a simple HTTP server (Node's built‑in http).
   ------------------------------------------------------------- */
import { createServer } from "http";

const server = createServer(async (req, res) => {
  if (req.method !== "POST" || req.url !== "/ask") {
    res.writeHead(404);
    res.end("Not found");
    return;
  }

  try {
    const body = await new Promise<string>((resolve, reject) => {
      let data = "";
      req.on("data", (chunk) => (data += chunk));
      req.on("end", () => resolve(data));
      req.on("error", reject);
    });
    const { question } = JSON.parse(body);
    const answer = await answerQuestion(question);
    res.writeHead(200, { "Content-Type": "application/json" });
    res.end(JSON.stringify({ answer }));
  } catch (e) {
    console.error(e);
    res.writeHead(500);
    res.end("Server error");
  }
});

server.listen(3000, () => console.log("RAG service listening on :3000"));
Enter fullscreen mode Exit fullscreen mode

What this script does

  • Embedding step – Calls OpenAI once per request; the cost is low compared to a full vector store query.
  • KNN query – Uses the DynamoDB GSI we set up earlier; the request costs 1 read capacity unit per item examined (adjust WCU accordingly).
  • Claude call – Sends a compact system prompt containing the three retrieved chunks.

Operational gotchas

  • Hot partitions – If many queries target the same docId range, DynamoDB can throttle. Mitigate by adding a random prefix to the partition key when you write chunks (sharding).
  • TransactWriteItems limit – When bulk‑loading many chunks, keep each transaction ≤ 100 items; otherwise you’ll hit a TransactionCanceledException.
  • Pricing awareness – At $0.25 per write capacity unit (WCU) in us‑east‑1, a table that writes 10 k items per day can cost a few dollars; monitor usage with CloudWatch.
  • TTL delay – Items scheduled for deletion may still appear in KNN results for up to 48 hours, so you might need to filter them out in application code if strict freshness is required.

In plain English – The whole pipeline is now a single DynamoDB table, one HTTP call to OpenAI, one KNN query, and one call to Claude. No extra services, no extra latency.


The Takeaway

  • DynamoDB’s native vector index lets you store embeddings next to your business data, removing the need for a separate vector database.
  • Create the table with a vector‑enabled GSI; remember to declare the embedding attribute as Binary and to base64‑encode the Float32Array.
  • Write and read embeddings using the @aws-sdk/lib-dynamodb package; the same client handles regular attributes and vectors.
  • A KNN query returns the most similar chunks; keep in mind the eventual‑consistency behavior of GSIs.
  • Use Node 22’s built‑in fetch to call Claude (or any LLM) after you have the context, building a short system prompt that forces the model to stay on‑topic.
  • Operational realities—hot partitions, TTL lag, and capacity pricing—still apply, so plan sharding and monitoring from day 1.

With these pieces in place, you have a lean, production‑ready Retrieval‑Augmented Generation service that lives entirely inside DynamoDB and plain Node.js. Happy building!


Transparency notice

This article was written with the help of an AI system — Groq (GPT OSS 120B).

Published: 2026-10-07 · Primary focus: DynamoDB

All code blocks are intended to be correct and runnable, but please verify them
against the official docs for the tools mentioned before using in production.

Find an error? Drop a comment — corrections are always welcome.

Top comments (0)