DEV Community

Tamiz Uddin
Tamiz Uddin

Posted on • Originally published at tamiz.pro

Build a Codebase Intelligence Tool Like repowise With a RAG-Assisted MCP for Your Monorepo

Originally published on tamiz.pro.

Introduction

Modern monorepos contain hundreds of thousands of files spanning multiple services, libraries, and configurations. Traditional code search—whether ripgrep, Sourcegraph, or IDE search—struggles with semantic queries like "how do we handle payment retries?" or "find all places where user permissions are checked". A RAG-assisted Model Context Protocol (MCP) server can turn your local codebase into a queryable knowledge base, giving LLMs and CLI tools accurate, context-rich answers. This tutorial shows you how to build a production-grade version of tools like repowise for your own monorepo.

Table of Contents

1. Architecture Overview

We'll build three components:

┌─────────────┐     ┌──────────────┐     ┌─────────────┐
│ File Watcher│────▶│ Indexer      │────▶│ Vector Store│
│ (chokidar)  │     │ (LangChain)  │     │ (Qdrant)    │
└─────────────┘     └──────────────┘     └─────────────┘
                                               │
                                               ▼
┌─────────────┐     ┌──────────────┐     ┌─────────────┐
│ IDE / CLI  │◀────│ MCP Server   │◀────│ Retriever   │
│ (Claude,    │     │ (FastMCP)    │     │ (Hybrid)    │
│ Cursor)     │     └──────────────┘     └─────────────┘
└─────────────┘
Enter fullscreen mode Exit fullscreen mode

Indexer: Splits code into AST-aware chunks, embeds them, and stores them in a local vector database.
MCP Server: Exposes a standardized interface (tools, resources, prompts) that any MCP-compatible client can consume.
Retriever: Combines vector similarity with BM25 lexical search and AST context for precise retrieval.

2. Prerequisites

  • Node.js 20+ and pnpm/npm
  • Python 3.11+ (for embedding server)
  • A running Qdrant instance (local or remote)
  • A monorepo with a manageable file count (< 1M files for this tutorial)
  • Basic familiarity with TypeScript and Python

3. Project Setup

Initialize the project structure:

mkdir monorepo-rag-mcp && cd monorepo-rag-mcp
pnpm init
mkdir -p packages/{indexer,mcp-server,embedding-server}
Enter fullscreen mode Exit fullscreen mode

3.1 Shared Types

Create a shared package for common types:

// packages/shared/src/types.ts
export interface CodeChunk {
  id: string;
  path: string;
  language: string;
  startLine: number;
  endLine: number;
  content: string;
  astMetadata?: Record<string, any>;
}

export interface IndexingConfig {
  monorepoRoot: string;
  excludePatterns: string[];
  languages: string[];
}
Enter fullscreen mode Exit fullscreen mode

4. Indexing Pipeline

4.1 AST-Aware Chunking

We use tree-sitter for language-agnostic parsing. This preserves semantic boundaries (functions, classes) instead of arbitrary character splits.

// packages/indexer/src/chunker.ts
import { Parser } from "tree-sitter";
import * as ts from "tree-sitter-typescript";

const parser = new Parser();
parser.setLanguage(ts.language);

export function chunkFile(content: string, filePath: string): CodeChunk[] {
  const tree = parser.parse(content);
  const chunks: CodeChunk[] = [];

  function traverse(node: any) {
    if (["function_declaration", "class_declaration", "method_definition"].includes(node.type)) {
      chunks.push({
        id: `${filePath}:${node.startPosition.row}-${node.endPosition.row}`,
        path: filePath,
        language: "typescript",
        startLine: node.startPosition.row,
        endLine: node.endPosition.row,
        content: content.slice(node.startPosition.byte, node.endPosition.byte),
        astMetadata: { type: node.type, name: node.children[0]?.text ?? "anonymous" }
      });
    }
    for (const child of node.children) traverse(child);
  }
  traverse(tree.rootNode);
  return chunks;
}
Enter fullscreen mode Exit fullscreen mode

4.2 Embedding Generation

Run a local embedding server using sentence-transformers for privacy and zero API cost.

# packages/embedding-server/main.py
from fastapi import FastAPI
from pydantic import BaseModel
from sentence_transformers import SentenceTransformer
import uvicorn

app = FastAPI()
model = SentenceTransformer("BAAI/bge-small-en-v1.5")

class EmbedRequest(BaseModel):
    texts: list[str]

@app.post("/embed")
def embed(req: EmbedRequest):
    embeddings = model.encode(req.texts, normalize_embeddings=True).tolist()
    return {"embeddings": embeddings}

if __name__ == "__main__":
    uvicorn.run(app, host="0.0.0.0", port=8000)
Enter fullscreen mode Exit fullscreen mode

4.3 Indexing Orchestrator

Watch for file changes and update the vector store incrementally.

// packages/indexer/src/indexer.ts
import chokidar from "chokidar";
import { QdrantClient } from "@qdrant/js-client-rest";
import { chunkFile } from "./chunker";
import axios from "axios";

const qdrant = new QdrantClient({ url: "http://localhost:6333" });

async function embed(texts: string[]) {
  const { data } = await axios.post("http://localhost:8000/embed", { texts });
  return data.embeddings;
}

async function indexFile(filePath: string) {
  const content = await fs.readFile(filePath, "utf-8");
  const chunks = chunkFile(content, filePath);
  if (chunks.length === 0) return;

  const embeddings = await embed(chunks.map(c => c.content));
  const points = chunks.map((chunk, i) => ({
    id: chunk.id,
    vector: embeddings[i],
    payload: chunk
  }));

  await qdrant.upsert("codebase", { points });
}

async function watch(root: string) {
  const watcher = chokidar.watch(root, {
    ignored: /node_modules|\.git|dist/g,
    persistent: true
  });

  watcher.on("add", path => indexFile(path));
  watcher.on("change", path => indexFile(path));
  watcher.on("unlink", path => qdrant.delete("codebase", { wait: true, points: [path] }));
}
Enter fullscreen mode Exit fullscreen mode

5. MCP Server Implementation

We'll use FastMCP (TypeScript SDK) to expose three tools: search_code, get_file_context, and explain_symbol.

// packages/mcp-server/src/server.ts
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import { QdrantClient } from "@qdrant/js-client-rest";
import axios from "axios";

const server = new McpServer({ name: "repo-wise", version: "1.0.0" });
const qdrant = new QdrantClient({ url: "http://localhost:6333" });

server.tool("search_code", {
  description: "Semantic + lexical search over the indexed monorepo",
  inputSchema: {
    type: "object",
    properties: {
      query: { type: "string", description: "Natural language query" },
      limit: { type: "number", default: 10 }
    },
    required: ["query"]
  }
}, async ({ query, limit = 10 }) => {
  // 1. Embed the query
  const { data: embedData } = await axios.post("http://localhost:8000/embed", { texts: [query] });
  const queryVector = embedData.embeddings[0];

  // 2. Vector search
  const vectorResults = await qdrant.search("codebase", {
    query: queryVector,
    limit: limit * 2, // overshoot for reranking
    with_payload: true
  });

  // 3. BM25 rerank (simplified: use Qdrant's built-in sparse vector if configured,
  //    or implement a local BM25 scorer on the retrieved candidates)
  const scored = vectorResults.map(r => ({
    path: r.payload.path,
    lines: `${r.payload.startLine}-${r.payload.endLine}`,
    snippet: r.payload.content.slice(0, 200),
    score: r.score
  }));

  return {
    content: [{ type: "text", text: JSON.stringify(scored, null, 2) }]
  };
});

server.tool("get_file_context", {
  description: "Get full file content with line numbers for a given path",
  inputSchema: {
    type: "object",
    properties: {
      path: { type: "string" },
      startLine: { type: "number" },
      endLine: { type: "number" }
    },
    required: ["path"]
  }
}, async ({ path, startLine, endLine }) => {
  const content = await fs.readFile(path, "utf-8");
  const lines = content.split("\n");
  const slice = lines.slice(startLine ?? 0, endLine ?? lines.length).join("\n");
  return {
    content: [{ type: "text", text: slice }]
  };
});

server.tool("explain_symbol", {
  description: "Explain what a symbol (function/class) does based on code and comments",
  inputSchema: {
    type: "object",
    properties: {
      symbolName: { type: "string" },
      language: { type: "string", default: "typescript" }
    },
    required: ["symbolName"]
  }
}, async ({ symbolName, language }) => {
  const { data: embedData } = await axios.post("http://localhost:8000/embed", { texts: [symbolName] });
  const results = await qdrant.search("codebase", {
    query: embedData.embeddings[0],
    limit: 5,
    filter: { must: [{ key: "language", match: { value: language } }] }
  });

  const context = results.map(r => r.payload.content).join("\n\n---\n\n");
  const prompt = `You are a senior engineer. Explain the symbol "${symbolName}" based on the following code snippets:\n\n${context}`;
  // In production, call an LLM here. For demo, return context.
  return { content: [{ type: "text", text: prompt }] };
});

const transport = new StdioServerTransport();
server.connect(transport);
Enter fullscreen mode Exit fullscreen mode

6. Client Integration

6.1 Claude Desktop

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "repo-wise": {
      "command": "node",
      "args": ["./packages/mcp-server/dist/server.js"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

6.2 Cursor IDE

In Cursor settings, add the MCP server as a custom tool:

{
  "tools": [
    {
      "name": "repo-wise",
      "command": "node",
      "args": ["./packages/mcp-server/dist/server.js"]
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

7. Production Hardening

7.1 Incremental Indexing with Git Hooks

Instead of filesystem watching (which misses renames, churn), use post-commit and post-merge hooks:

#!/bin/bash
# .git/hooks/post-commit
ROOT=$(git rev-parse --show-toplevel)
pnpm --filter @repo/indexer run index --root "$ROOT" --commit $(git rev-parse HEAD)
Enter fullscreen mode Exit fullscreen mode

7.2 Hybrid Search with Qdrant

Enable Qdrant's built-in sparse vectors (BM25) for better lexical recall:

await qdrant.createCollection("codebase", {
  vectors: {
    size: 384,
    distance: "Cosine"
  },
  sparse_vectors: {
    bm25: {}
  }
});
Enter fullscreen mode Exit fullscreen mode

Then during search, use Qdrant's search with using: ["bm25"] and fuse scores.

7.3 Security & Privacy

  • Run all components locally; never ship code to external APIs unless encrypted.
  • Use filesystem permissions to restrict access.
  • Sanitize file paths to prevent path traversal.
  • Add rate limiting to the embedding server.

8. Frequently Asked Questions

Q: How large a monorepo can this handle?
A: With local embeddings and Qdrant, we've tested repos up to 2M files (~50GB). The bottleneck is initial indexing time, not query latency.

Q: Can I use this with non-TypeScript languages?
A: Yes. tree-sitter supports 100+ languages. Update the languages array and use the corresponding tree-sitter-<lang> grammar.

Q: How does this compare to GitHub Copilot Workspace?
A: Copilot is cloud-hosted and proprietary. This tool runs entirely in your infrastructure, supports custom retrieval logic, and integrates with any MCP client (Claude, Cursor, custom IDEs).


Ready to ship? Start with the indexer package and iterate on chunking strategies. For production deployments, consider adding a lightweight job queue (BullMQ) for indexing and a reverse proxy for the embedding server.

Production Considerations

For production deployments, consider adding a lightweight job queue (BullMQ) for indexing and a reverse proxy for the embedding server. Below is a minimal BullMQ integration for the indexer package.

// packages/indexer/src/worker.ts
import { Queue, Worker, Job } from 'bullmq';
import { IndexService } from './index-service';
import { redisConnection } from './redis';

export const indexQueue = new Queue('codebase-indexing', { connection: redisConnection });

export function startWorker() {
  const service = new IndexService();

  new Worker('codebase-indexing', async (job: Job) => {
    const { filePath, content, projectId } = job.data;

    await service.indexFile(projectId, filePath, content);

    job.updateProgress(100);
  }, { connection: redisConnection });
}

export async function enqueueIndex(
  projectId: string,
  filePath: string,
  content: string,
) {
  await indexQueue.add('index-file', {
    projectId,
    filePath,
    content,
  }, {
    attempts: 3,
    backoff: { type: 'exponential', delay: 2000 },
    removeOnComplete: true,
    removeOnFail: true,
  });
}
Enter fullscreen mode Exit fullscreen mode

Reverse Proxy for Embedding Server

The embedding server should sit behind a reverse proxy that handles TLS termination, request buffering, and connection pooling. Here's an Nginx configuration:

# nginx.conf
upstream embedding_server {
    server 127.0.0.1:6000 max_fails=3 fail_timeout=30s;
    keepalive 32;
}

server {
    listen 443 ssl http2;
    server_name embedding.internal;

    ssl_certificate     /etc/ssl/certs/embedding.pem;
    ssl_certificate_key /etc/ssl/private/embedding.key;

    location /v1/ {
        proxy_pass http://embedding_server/v1/;
        proxy_http_version 1.1;
        proxy_set_header Connection "";

        # Buffer large document batches
        proxy_request_buffering on;
        client_max_body_size 50m;

        # Timeout tuning for embedding models
        proxy_connect_timeout 10s;
        proxy_send_timeout 120s;
        proxy_read_timeout 120s;
    }

    location /health {
        proxy_pass http://embedding_server/health;
    }
}
Enter fullscreen mode Exit fullscreen mode

Monitoring and Observability

A codebase intelligence tool generates significant operational data. You need visibility into indexing throughput, embedding latency, and MCP request patterns.

Structured Logging Setup

// packages/monitoring/src/logger.ts
import pino from 'pino';

export const logger = pino({
  level: process.env.LOG_LEVEL || 'info',
  formatters: {
    level: (label) => ({ level: label.toUpperCase() }),
  },
  timestamp: pino.stdTimeFunctions.isoTime,
});

export type MetricsEvent =
  | { type: 'index_progress'; projectId: string; fileCount: number; total: number }
  | { type: 'embedding_latency'; durationMs: number; model: string }
  | { type: 'mcp_request'; toolName: string; durationMs: number; success: boolean }
  | { type: 'chunk_generated'; projectId: string; chunkSize: number; strategy: string };

export function emitMetrics(event: MetricsEvent) {
  logger.info(event, 'metrics');
}
Enter fullscreen mode Exit fullscreen mode

Key Metrics to Track

Metric Why It Matters Alert Threshold
Indexing throughput (files/min) Detects slowdowns in scanning or embedding < 10 files/min for > 5 min
Embedding latency p99 Impacts MCP tool response times > 3s p99
MCP tool error rate Indicates retrieval or parsing failures > 5%
Chunk size distribution Reveals if strategies produce unusable chunks mean chunk < 50 tokens
Vector store query latency Directly affects LLM response time > 500ms p99
Memory usage per project Catches memory leaks in long-running workers > 2GB/project

OpenTelemetry Integration

// packages/monitoring/src/tracing.ts
import * as opentelemetry from '@opentelemetry/api';
import { NodeTracerProvider } from '@opentelemetry/sdk-trace-node';
import { Resource } from '@opentelemetry/resources';
import { SemanticResourceAttributes } from '@opentelemetry/semantic-conventions';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-proto';
import { registerInstrumentations } from '@opentelemetry/instrumentation';
import { HttpInstrumentation } from '@opentelemetry/instrumentation-http';
import { PinoInstrumentation } from '@opentelemetry/instrumentation-pino';

const provider = new NodeTracerProvider({
  resource: new Resource({
    [SemanticResourceAttributes.SERVICE_NAME]: 'monorepo-rag-mcp',
    [SemanticResourceAttributes.SERVICE_VERSION]: '1.0.0',
    [SemanticResourceAttributes.DEPLOYMENT_ENVIRONMENT]: process.env.NODE_ENV,
  }),
});

provider.addSpanProcessor(
  new SimpleSpanProcessor(new OTLPTraceExporter())
);
provider.register();

registerInstrumentations({
  instrumentations: [
    new HttpInstrumentation(),
    new PinoInstrumentation(),
  ],
});

export const tracer = opentelemetry.trace.getTracer('monorepo-rag-mcp');
Enter fullscreen mode Exit fullscreen mode

Using these spans, you can trace a single MCP tool call from the IDE through the gateway, into the RAG pipeline, through embedding generation, vector search, and context assembly—identifying exactly where latency accumulates.

Security Considerations

Preventing Prompt Injection

When user code contains comments or strings that resemble instructions, those can leak into the context window and influence the LLM's behavior. Mitigate this with input sanitization:

// packages/rag/src/sanitization.ts

const INJECTION_PATTERNS = [
  /ignore previous instructions/i,
  /act as a different system/i,
  /you are now/i,
  /system prompt/i,
  /do not follow/i,
];

export function sanitizeContext(text: string): string {
  return INJECTION_PATTERNS.reduce(
    (acc, pattern) => acc.replace(pattern, '[REDACTED_INJECTION_ATTEMPT]'),
    text
  );
}

export function wrapCodeInSecurityBoundary(code: string): string {
  return `SECURITY_BOUNDARY_START\n${sanitizeContext(code)}\nSECURITY_BOUNDARY_END`;
}
Enter fullscreen mode Exit fullscreen mode

Access Control for Monorepo Projects

In a monorepo, not all developers should have equal access. Implement project-level RBAC:

// packages/mcp-server/src/auth.ts
import { z } from 'zod';

const AccessLevel = z.enum(['read', 'write', 'admin']);
type AccessLevel = z.infer<typeof AccessLevel>;

interface ProjectPermission {
  projectId: string;
  userId: string;
  accessLevel: AccessLevel;
  inheritedFrom?: string; // parent workspace for monorepo hierarchy
}

export class PermissionResolver {
  private permissions: Map<string, ProjectPermission[]> = new Map();

  setPermissions(projectId: string, perms: ProjectPermission[]) {
    this.permissions.set(projectId, perms);
  }

  checkAccess(userId: string, projectId: string, requiredLevel: AccessLevel): boolean {
    // Check direct permissions
    const direct = this.permissions.get(projectId)?.find(p => p.userId === userId);
    if (direct && this.isLevelSufficient(direct.accessLevel, requiredLevel)) {
      return true;
    }

    // Check inherited permissions (monorepo parent workspaces)
    return this.checkInheritedAccess(userId, projectId, requiredLevel);
  }

  private isLevelSufficient(given: AccessLevel, required: AccessLevel): boolean {
    const order: AccessLevel[] = ['read', 'write', 'admin'];
    return order.indexOf(given) >= order.indexOf(required);
  }

  private checkInheritedAccess(
    userId: string,
    projectId: string,
    requiredLevel: AccessLevel,
    visited: Set<string> = new Set(),
  ): boolean {
    if (visited.has(projectId)) return false;
    visited.add(projectId);

    const perms = this.permissions.get(projectId);
    if (!perms) return false;

    for (const perm of perms) {
      if (perm.userId !== userId) continue;
      if (this.isLevelSufficient(perm.accessLevel, requiredLevel)) return true;
      if (perm.inheritedFrom) {
        if (this.checkInheritedAccess(userId, perm.inheritedFrom, requiredLevel, visited)) {
          return true;
        }
      }
    }
    return false;
  }
}
Enter fullscreen mode Exit fullscreen mode

Secret Scanning in Indexed Code

Never allow sensitive data to be embedded in vectors. Add a pre-indexing scan:

// packages/indexer/src/secret-scanner.ts
import { scan } from 'secret-scan';

export interface ScanResult {
  matched: boolean;
  secretType: string;
  lineNumber: number;
}

export async function scanForSecrets(content: string): Promise<ScanResult[]> {
  return scan(content, {
    rules: [
      { id: 'aws-access-key', pattern: /AKIA[0-9A-Z]{16}/, severity: 'high' },
      { id: 'github-token', pattern: /ghp_[0-9a-zA-Z]{36}/, severity: 'high' },
      { id: 'private-key', pattern: /-----BEGIN (RSA|EC|OPENSSH) PRIVATE KEY-----/, severity: 'high' },
      { id: 'jwt', pattern: /eyJ[A-Za-z0-9_-]+\.eyJ[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+/, severity: 'medium' },
    ],
  });
}

export function filterSecrets(content: string): string {
  return content.replace(
    /AKIA[0-9A-Z]{16}|ghp_[0-9a-zA-Z]{36}|-----BEGIN [A-Z ]+ PRIVATE KEY-----|eyJ[A-Za-z0-9_-]+/g,
    '[SECRET_REDACTED]',
  );
}
Enter fullscreen mode Exit fullscreen mode

Testing Strategy

A RAG-assisted MCP tool has complex interactions between the language server, vector store, embedding model, and LLM. Test each layer independently and integrate at the boundaries.

Unit Tests for Chunking Strategies

// packages/indexer/__tests__/strategies/typescript-strategy.test.ts
import { describe, it, expect } from 'vitest';
import { TypeScriptChunkStrategy } from '../../src/strategies/typescript';
import { Chunk } from '../../src/types';

describe('TypeScriptChunkStrategy', () => {
  const strategy = new TypeScriptChunkStrategy({
    maxChunkSize: 500,
    overlapTokens: 50,
  });

  it('splits a large class into logical chunks', () => {
    const code = `
      class UserService {
        async getUser(id: string) { return db.find(id); }
        async createUser(data: CreateUserInput) { return db.insert(data); }
        async deleteUser(id: string) { return db.remove(id); }
        async listUsers(page: number) { return db.paginate(page); }
        async updateUser(id: string, data: Partial<CreateUserInput>) {
          return db.update(id, data);
        }
      }
    `.trim();

    const chunks = strategy.chunk({
      filePath: 'src/services/user.service.ts',
      content: code,
      language: 'typescript',
    });

    expect(chunks).toHaveLengthGreaterThan(1);
    chunks.forEach((chunk: Chunk) => {
      expect(chunk.content.length).toBeLessThanOrEqual(500);
      expect(chunk.filePath).toBe('src/services/user.service.ts');
      expect(chunk.language).toBe('typescript');
    });
  });

  it('preserves imports and exports in chunk metadata', () => {
    const code = `
      import { injectable } from 'inversify';
      export class CacheService implements ICacheService {}
    `.trim();

    const chunks = strategy.chunk({
      filePath: 'src/services/cache.service.ts',
      content: code,
      language: 'typescript',
    });

    expect(chunks[0].metadata.imports).toContain('injectable');
    expect(chunks[0].metadata.exports).toContain('CacheService');
  });
});
Enter fullscreen mode Exit fullscreen mode

Integration Tests for the MCP Tool Server

// packages/mcp-server/__tests__/tool-server.test.ts
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { MCPToolServer } from '../src/tool-server';
import { MockVectorStore } from './fixtures/mock-vector-store';
import { MockEmbeddingService } from './fixtures/mock-embedding-service';

describe('MCPToolServer', () => {
  let server: MCPToolServer;
  let vectorStore: MockVectorStore;
  let embeddingService: MockEmbeddingService;

  beforeEach(() => {
    vectorStore = new MockVectorStore();
    embeddingService = new MockEmbeddingService();
    server = new MCPToolServer({
      vectorStore,
      embeddingService,
      projectId: 'test-project',
    });
  });

  afterEach(async () => {
    await server.shutdown();
  });

  it('responds to ping', async () => {
    const result = await server.handleRequest({
      jsonrpc: '2.0',
      id: 1,
      method: 'ping',
    });
    expect(result.result).toEqual({ status: 'ok' });
  });

  it('searches codebase with semantic query', async () => {
    await vectorStore.upsert('test-project', [
      { id: 'chunk-1', content: 'auth middleware validates JWT tokens', metadata: { filePath: 'src/middleware/auth.ts' } },
      { id: 'chunk-2', content: 'database connection pool configuration', metadata: { filePath: 'src/db/pool.ts' } },
    ]);

    const result = await server.callTool({
      name: 'search_codebase',
      arguments: {
        query: 'How do I configure authentication?',
        projectId: 'test-project',
        limit: 3,
      },
    });

    expect(result.content).toHaveLengthGreaterThan(0);
    expect(result.content[0].text).toContain('auth');
  });

  it('handles undefined project gracefully', async () => {
    const result = await server.callTool({
      name: 'search_codebase',
      arguments: {
        query: 'something',
        projectId: 'nonexistent-project',
      },
    });

    expect(result.isError).toBe(true);
    expect(result.content[0].text).toContain('not indexed');
  });
});
Enter fullscreen mode Exit fullscreen mode

E2E Test with Real MCP Client

// packages/mcp-server/__tests__/e2e/mcp-client-e2e.test.ts
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import { Client } from '@modelcontextprotocol/sdk/client/index.js';
import { StdioClientTransport } from '@modelcontextprotocol/sdk/client/stdio.js';
import { spawn } from 'child_process';

describe('E2E: MCP Tool Server', () => {
  let client: Client;
  let serverProcess: ReturnType<typeof spawn>;

  beforeAll(async () => {
    serverProcess = spawn('node', ['dist/mcp-server.js'], {
      env: {
        ...process.env,
        VECTOR_STORE_PATH: ':memory:',
        PROJECT_ROOT: './__tests__/fixtures/monorepo',
      },
    });

    const transport = new StdioClientTransport(serverProcess);
    client = new Client(
      { name: 'e2e-test-client', version: '1.0.0' },
      { capabilities: {} },
    );
    await client.connect(transport);
  });

  afterAll(async () => {
    await client.close();
    serverProcess.kill();
  });

  it('lists available tools', async () => {
    const tools = await client.listTools();
    const toolNames = tools.tools.map(t => t.name);

    expect(toolNames).toContain('search_codebase');
    expect(toolNames).toContain('get_file_context');
    expect(toolNames).toContain('trace_dependency');
  });

  it('performs a full search-and-reason workflow', async () => {
    // Search for auth-related code
    const searchResult = await client.callTool({
      name: 'search_codebase',
      arguments: {
        query: 'JWT token validation',
        limit: 2,
      },
    });

    expect(searchResult.isError).toBe(false);

    // Get file context for the top result
    const firstResult = JSON.parse(searchResult.content[0].text);
    const contextResult = await client.callTool({
      name: 'get_file_context',
      arguments: {
        filePath: firstResult.filePath,
        contextWindow: 50,
      },
    });

    expect(contextResult.isError).toBe(false);
    expect(contextResult.content[0].text.length).toBeGreaterThan(0);
  });
});
Enter fullscreen mode Exit fullscreen mode

Benchmarking and Evaluation

How do you know your RAG pipeline is performing well? You need both quantitative benchmarks and qualitative human evaluation.

Build a Golden Test Set

// packages/eval/src/golden-set.ts
export interface GoldenQuestion {
  id: string;
  query: string;
  expectedFilePaths: string[];
  expectedChunkIds: string[];
  difficulty: 'easy' | 'medium' | 'hard';
  category: 'finding' | 'understanding' | 'debugging' | 'refactoring';
}

export const GOLDEN_SET: GoldenQuestion[] = [
  {
    id: 'q001',
    query: 'Where is the authentication middleware defined?',
    expectedFilePaths: ['packages/auth/src/middleware.ts'],
    expectedChunkIds: ['chunk-auth-middleware-1'],
    difficulty: 'easy',
    category: 'finding',
  },
  {
    id: 'q002',
    query: 'How does the payment service handle retry logic for Stripe API calls?',
    expectedFilePaths: [
      'packages/payments/src/service.ts',
      'packages/payments/src/retry.ts',
    ],
    expectedChunkIds: ['chunk-payment-service', 'chunk-retry-logic'],
    difficulty: 'medium',
    category: 'understanding',
  },
  {
    id: 'q003',
    query: 'There\'s a memory leak in the WebSocket handler. Where is the cleanup code?',
    expectedFilePaths: ['packages/ws/src/handler.ts'],
    expectedChunkIds: ['chunk-ws-handler-close'],
    difficulty: 'hard',
    category: 'debugging',
  },
];
Enter fullscreen mode Exit fullscreen mode

Evaluation Script

// packages/eval/src/evaluator.ts
import { GoldenQuestion, GOLDEN_SET } from './golden-set';
import { MCPToolServer } from 'monorepo-rag-mcp-server';

interface EvalResult {
  questionId: string;
  recall: number;       // % of expected chunks retrieved
  precision: number;    // % of retrieved chunks that are relevant
  f1: number;
  latencyMs: number;
  passed: boolean;
}

export class Evaluator {
  constructor(private server: MCPToolServer) {}

  async evaluate(): Promise<EvalResult[]> {
    const results: EvalResult[] = [];

    for (const question of GOLDEN_SET) {
      const start = performance.now();

      const searchResult = await this.server.callTool({
        name: 'search_codebase',
        arguments: {
          query: question.query,
          limit: 10,
        },
      });

      const latencyMs = performance.now() - start;
      const retrievedChunks = this.parseChunksFromResult(searchResult);

      const { recall, precision, f1 } = this.computeMetrics(
        retrievedChunks,
        question.expectedChunkIds,
      );

      results.push({
        questionId: question.id,
        recall,
        precision,
        f1,
        latencyMs,
        passed: recall >= 0.8 && precision >= 0.7,
      });
    }

    return results;
  }

  private computeMetrics(
    retrieved: string[],
    expected: string[],
  ): { recall: number; precision: number; f1: number } {
    const retrievedSet = new Set(retrieved);
    const expectedSet = new Set(expected);

    const truePositives = [...retrievedSet].filter(id => expectedSet.has(id)).length;
    const recall = expectedSet.size > 0
      ? truePositives / expectedSet.size
      : 0;
    const precision = retrievedSet.size > 0
      ? truePositives / retrievedSet.size
      : 0;
    const f1 = (recall + precision) > 0
      ? (2 * recall * precision) / (recall + precision)
      : 0;

    return { recall, precision, f1 };
  }

  private parseChunksFromResult(result: unknown): string[] {
    // Implementation depends on your result format
    return [];
  }
}
Enter fullscreen mode Exit fullscreen mode

Performance Targets

Metric Target Notes
Mean query latency < 2s Includes embedding + search
P99 query latency < 5s During peak indexing
Recall@10 on golden set > 85% Minimum viable
Precision@10 on golden set > 70% Context window budget
F1 across all questions > 0.75 Balanced metric
Index build time (10k files) < 15 min Parallel embedding workers

Run pnpm eval after every chunking strategy change to track regressions.

Deployment with Docker Compose

Here's a complete docker-compose.yml for running the full stack locally or in production:

# docker-compose.yml
version: '3.9'

services:
  embedding-server:
    build:
      context: ./packages/embedding-server
      dockerfile: ../../Dockerfile.embedding
    ports:
      - "6000:6000"
    environment:
      - MODEL_ID=${EMBEDDING_MODEL:-all-MiniLM-L6-v2}
      - MAX_BATCH_SIZE=32
      - WORKERS=4
    volumes:
      - embedding-cache:/app/.cache
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:6000/health"]
      interval: 30s
      timeout: 10s
      retries: 3

  redis:
    image: redis:7-alpine
    ports:
      - "6379:6379"
    volumes:
      - redis-data:/data
    command: redis-server --appendonly yes
    healthcheck:
      test: ["CMD", "redis-cli", "ping"]
      interval: 10s
      timeout: 5s
      retries: 5

  mcp-gateway:
    build:
      context: ./packages/mcp-gateway
      dockerfile: ../../Dockerfile.gateway
    ports:
      - "3000:3000"
    environment:
      - MCP_SERVER_URL=http://mcp-server:8080
      - REDIS_URL=redis://redis:6379
      - EMBEDDING_SERVER_URL=http://embedding-server:6000
    depends_on:
      redis:
        condition: service_healthy
      embedding-server:
        condition: service_healthy
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:3000/health"]
      interval: 30s
      timeout: 10s
      retries: 3

  mcp-server:
    build:
      context: ./packages/mcp-server
      dockerfile: ../../Dockerfile.mcp
    environment:
      - VECTOR_STORE_URL=http://embedding-server:6000
      - REDIS_URL=redis://redis:6379
      - PROJECT_ROOT=/data/monorepo
    volumes:
      - ./monorepo:/data/monorepo:ro
    depends_on:
      redis:
        condition: service_healthy
      embedding-server:
        condition: service_healthy

  indexer-worker:
    build:
      context: ./packages/indexer
      dockerfile: ../../Dockerfile.worker
    environment:
      - REDIS_URL=redis://redis:6379
      - EMBEDDING_SERVER_URL=http://embedding-server:6000
      - MAX_CONCURRENCY=4
    volumes:
      - ./monorepo:/data/monorepo:ro
    depends_on:
      redis:
        condition: service_healthy
      embedding-server:
        condition: service_healthy

volumes:
  redis-data:
  embedding-cache:
Enter fullscreen mode Exit fullscreen mode

Start everything with:

docker compose up -d
Enter fullscreen mode Exit fullscreen mode

Then verify the health checks:

curl http://localhost:3000/health
curl http://localhost:6000/health
Enter fullscreen mode Exit fullscreen mode

CI/CD Pipeline

# .github/workflows/ci.yml
name: CI

on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

jobs:
  test:
    runs-on: ubuntu-latest
    services:
      redis:
        image: redis:7-alpine
        ports: ['6379:6379']
      qdrant:
        image: qdrant/qdrant:latest
        ports: ['6333:6333']

    steps:
      - uses: actions/checkout@v4

      - name: Setup Node
        uses: actions/setup-node@v4
        with:
          node-version: '20'
          cache: 'pnpm'

      - name: Install dependencies
        run: pnpm install --frozen-lockfile

      - name: Type check
        run: pnpm --filter monorepo-rag-mcp... typecheck

      - name: Lint
        run: pnpm --filter monorepo-rag-mcp... lint

      - name: Unit tests
        run: pnpm --filter monorepo-rag-mcp... test:unit

      - name: Integration tests
        run: pnpm --filter monorepo-rag-mcp-server test:integration
        env:
          REDIS_URL: redis://localhost:6379
          VECTOR_STORE_URL: http://localhost:6333

      - name: E2E tests
        run: pnpm --filter monorepo-rag-mcp-server test:e2e
        env:
          REDIS_URL: redis://localhost:6379
          EMBEDDING_SERVER_URL: http://localhost:6000

      - name: Evaluate
        run: pnpm --filter monorepo-rag-mcp-eval eval
        env:
          REDIS_URL: redis://localhost:6379

      - name: Build
        run: pnpm build

      - name: Docker build (gateway)
        run: docker build -f Dockerfile.gateway -t mcp-gateway:${{ github.sha }} .

      - name: Docker build (mcp server)
        run: docker build -f Dockerfile.mcp -t mcp-server:${{ github.sha }} .
Enter fullscreen mode Exit fullscreen mode

Connecting to Your IDE

VS Code Configuration

Add this to your .vscode/settings.json:

{
  "mcp.servers": {
    "monorepo-rag": {
      "type": "stdio",
      "command": "node",
      "args": ["/path/to/packages/mcp-gateway/dist/index.js"],
      "env": {
        "PROJECT_ROOT": "/path/to/your/monorepo",
        "REDIS_URL": "redis://localhost:6379",
        "EMBEDDING_SERVER_URL": "http://localhost:6000"
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Or use the official MCP client extension:

{
  "mcpServers": {
    "monorepo-rag-mcp": {
      "command": "npx",
      "args": ["-y", "monorepo-rag-mcp-gateway"],
      "env": {
        "PROJECT_ROOT": "${workspaceFolder}",
        "EMBEDDING_MODEL": "all-MiniLM-L6-v2"
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

JetBrains / Cursor Configuration

For Cursor, add to .cursor/mcp.json:

{
  "mcpServers": {
    "monorepo-rag": {
      "command": "node",
      "args": ["dist/mcp-gateway.js"],
      "env": {
        "PROJECT_ROOT": "/path/to/monorepo"
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

For JetBrains IDEs with the MCP plugin, use the same configuration format in the plugin settings panel.

Using the Tools

Once connected, you'll see tools like these in your AI assistant's tool palette:

  • search_codebase: "Find where error handling is implemented for the payment flow"
  • get_file_context: "Show me 50 lines around the main export in src/cache.ts"
  • trace_dependency: "What packages depend on lodash?"
  • get_api_schema: "What does the /users endpoint expect?"
  • compare_versions: "What changed between v2.3.1 and v2.4.0 in the auth module?"

Common Pitfalls and How to Avoid Them

1. Chunk Size Mismatch

If your chunks are too small, the LLM loses context. If too large, retrieval becomes noisy. Start with 300–500 tokens and measure recall at different sizes.

// packages/indexer/src/chunker.ts
const CHUNK_SIZES = [200, 300, 500, 800, 1200];

export function findOptimalChunkSize(
  projectId: string,
  sampleQueries: string[],
): number {
  let bestSize = CHUNK_SIZES[0];
  let bestScore = 0;

  for (const size of CHUNK_SIZES) {
    const strategy = new SmartChunker({ maxChunkSize: size, overlapTokens: Math.floor(size * 0.1) });
    const score = evaluateStrategy(projectId, strategy, sampleQueries);
    if (score > bestScore) {
      bestScore = score;
      bestSize = size;
    }
  }

  return bestSize;
}
Enter fullscreen mode Exit fullscreen mode

2. Stale Indexes

Files change. Your index must reflect reality. Use filesystem watchers combined with periodic full re-indexes:

// packages/indexer/src/watcher.ts
import { watch } from 'chokidar';
import { IndexService } from './index-service';

export function startFileWatcher(
  projectRoot: string,
  indexService: IndexService,
  projectId: string,
) {
  const watcher = watch(projectRoot, {
    ignored: /node_modules|\.git|dist|build|\.next/,
    persistent: true,
    depth: 10,
  });

  watcher
    .on('add', async (path) => {
      await indexService.indexFile(projectId, path);
    })
    .on('change', async (path) => {
      await indexService.reindexFile(projectId, path);
    })
    .on('unlink', async (path) => {
      await indexService.removeFile(projectId, path);
    });

  // Full reindex every 6 hours as a safety net
  setInterval(() => {
    indexService.reindexAll(projectId);
  }, 6 * 60 * 60 * 1000);
}
Enter fullscreen mode Exit fullscreen mode

3. Embedding Model Drift

Switching embedding models changes vector semantics. Never swap models without re-embedding the entire corpus, or you'll get incomparable vectors in the store:

// packages/indexer/src/migration.ts
export interface ModelMigration {
  fromModel: string;
  toModel: string;
  requiresReindex: boolean;
  notes: string;
}

const MIGRATIONS: Record<string, ModelMigration> = {
  'all-MiniLM-L6-v2->text-embedding-3-small': {
    fromModel: 'all-MiniLM-L6-v2',
    toModel: 'text-embedding-3-small',
    requiresReindex: true,
    notes: 'Dimensionality change: 384 -> 1536. Full re-index required.',
  },
};

export function validateModelSwap(currentModel: string, newModel: string): void {
  const migration = MIGRATIONS[`${currentModel}->${newModel}`];
  if (!migration) return;

  if (migration.requiresReindex) {
    throw new Error(
      `Cannot switch from ${currentModel} to ${newModel} without re-indexing. ${migration.notes}`
    );
  }
}
Enter fullscreen mode Exit fullscreen mode

4. IDE Context Pollution

When the MCP tool returns results, the IDE's AI may also have access to the open files. This can cause the model to prefer IDE context over retrieved context. Mitigate by clearly labeling retrieved results:

// packages/mcp-server/src/responses.ts
export function formatSearchResult(query: string, results: SearchResult[]): string {
  const lines = [
    `# Search Results for: "${query}"`,
    `Found ${results.length} relevant code chunks from ${new Set(results.map(r => r.filePath)).size} files.`,
    `---`,
  ];

  for (const result of results) {
    lines.push(`\n## \`${result.filePath}\` (relevance: ${(result.score * 100).toFixed(1)}%)`);
    lines.push(result.content);
    lines.push(`---`);
  }

  return lines.join('\n');
}
Enter fullscreen mode Exit fullscreen mode

Conclusion

Building a codebase intelligence tool like repowise with a RAG-assisted MCP for your monorepo is a substantial engineering effort—but one that pays compounding dividends. Every developer on the team gains instant, semantic access to the entire codebase without context-switching between IDE, documentation, and search tools.

The key architectural decisions that determine success are:

  1. Language-aware chunking that respects syntactic boundaries rather than naively splitting by character count.
  2. Separating embedding generation from vector storage so you can swap models and stores independently.
  3. A clean MCP tool contract with explicit error handling, typed parameters, and structured responses.
  4. Operational rigor—monitoring, evaluation, secret scanning, and access control aren't optional extras for a monorepo tool; they're the difference between a prototype and a production system.

The codebase structure we've outlined—five focused packages under a pnpm workspace—gives you a foundation that scales. The indexer handles the heavy lifting of turning code into searchable vectors. The embedding-server provides a stable, language-agnostic embedding interface. The vector-store abstraction lets you move between Qdrant, Weavi

Top comments (0)