📝 Update (Feb 2026): Most AI coding tools now have built-in sub-agents — but each tool only delegates to its own model. If you want to mix backends (e.g. Claude for reviews, Codex for generation, Gemini for large-context analysis) within a single workflow, check out sub-agents-skills. It's a lighter, model-agnostic approach — no MCP server needed.
👉 Full implementation available at shinpr/sub-agents-mcp
I wanted to try Cursor and other emerging AI coding tools. But I kept hitting the same walls without sub-agents — context pollution, inconsistent outputs, and the dreaded mid-task context exhaustion.
Claude Code has this feature called Sub-agents — specialized AI assistants with separate contexts that handle specific tasks. It solves context exhaustion and dramatically improves accuracy. But other AI coding tools don't have it.
So I built an MCP server that makes it happen.
What Sub agents Actually Do
Sub-agents are specialized AI assistants in Claude Code that handle specific tasks with efficient problem-solving and context management.
Why They Matter
Isolated Contexts
Each sub-agent gets its own context window. No more context pollution. No more running out of tokens mid-task.
Task-Specific Precision
A code reviewer needs different context than an implementer. Give each agent exactly what it needs.
Reproducible Results
Define once in Markdown, get consistent quality every time.
Building an MCP Bridge
I built an MCP server that brings sub-agents to any tool that supports Model Context Protocol.
https://github.com/shinpr/sub-agents-mcp
About MCP
Model Context Protocol (MCP) is a standardized protocol that enables host applications (like Cursor and Claude Desktop) to communicate with servers (data sources and tools).
By implementing sub-agents as an MCP server, I made this Claude Code-exclusive feature available to any MCP-compatible tool.
How It Works
Just tell your AI to use a sub-agent:
"Use the document-reviewer agent to check docs/PRD/spec.md"
Your specialized agent takes over with its own fresh context.
Example Run in Cursor
As you can see, the document-reviewer agent produces a structured report rather than freeform text. The output includes a summary of strengths, key issues with severity levels, and actionable suggestions for improvement. This makes it easy to spot gaps and apply consistent review standards across different documents.
Quick Setup
Step 1: Configure Your Tool
Add this to your tool's MCP config (e.g., ~/.cursor/mcp.json):
{
"mcpServers": {
"sub-agents": {
"command": "npx",
"args": ["-y", "sub-agents-mcp"],
"env": {
"AGENTS_DIR": "/Users/username/projects/my-app/.cursor/agents",
"AGENT_TYPE": "cursor"
}
}
}
}
Note: Use absolute paths for AGENTS_DIR. I recommend creating an agents folder in your project.
Step 2: Create Your First Agent
Create a Markdown file in your agents directory:
code-reviewer.md:
# Code Reviewer
You are an AI assistant specialized in code review.
Please review with the following perspectives:
- Finding bugs and potential issues
- Suggesting performance and readability improvements
- Checking compliance with best practices
Your code-reviewer sub-agent is ready to use.
Creating Effective Sub-agents
Claude Code's official documentation explains the concept:
Custom sub agents in Claude Code are specialized AI assistants that can be invoked to handle specific types of tasks. They enable more efficient problem-solving by providing task-specific configurations with customized system prompts, tools and a separate context window.
sub-agents-mcp follows this pattern. It interprets Markdown files in the AGENTS_DIR directory as agent definitions and passes the entire file content as system context to Cursor CLI or Claude Code.
.claude
┗ agents
┣ code-reviewer.md # Works as "code-reviewer" agent
┗ document-reviewer.md # Works as "document-reviewer" agent
Key tips for creating sub-agent definitions:
- Define sub-agents with single responsibility
- Provide necessary information while excluding unnecessary details
- Break down tasks to fit within one context window
For example, I recommend separating generation tasks from review tasks. Implementation gathers lots of context that becomes unnecessary during review. Trying to do both in one task usually exhausts the context before review, resulting in poor quality.
You can find definition samples in my boilerplate repository.
Example: Document Reviewer
Here's a real agent definition I use:
document-reviewer.md:
You are a technical document reviewer.
## Primary Objective
Evaluate document completeness and consistency. Output structured review with actionable improvements.
## Input
- **target**: Absolute path to document file
## Review Process
### Phase 1: Document Analysis
1. Load document from target path
2. Extract all technical claims and requirements
3. Identify document type and expected sections
### Phase 2: Validation Checks
Execute ALL checks in order:
**Consistency Check**
- Find contradictions between sections
- Identify ambiguous statements
- Flag: If found, mark as CRITICAL severity
**Completeness Check**
- Verify mandatory sections exist
- Check technical details depth
- Flag: Missing sections = CRITICAL, insufficient detail = IMPORTANT
**Clarity Check**
- Assess technical terminology usage
- Evaluate logical flow
- Flag: Unclear sections = RECOMMENDED
## Output Requirements
### Mandatory Structure
[METADATA]
document: <filename>
review_date: <ISO-8601>
reviewer: document-reviewer
[SCORES]
consistency: <0-100>
completeness: <0-100>
clarity: <0-100>
[ISSUES]
<For each issue>
id: <ISSUE-001 format>
severity: <critical|important|recommended>
category: <consistency|completeness|clarity>
location: <section/line reference>
description: <what is wrong>
suggestion: <how to fix>
[VERDICT]
decision: <APPROVED|REJECTED|CONDITIONAL>
reason: <one sentence explanation>
### Severity Rules
- CRITICAL: Blocks approval. Contradictions, missing mandatory sections
- IMPORTANT: Should fix. Incomplete technical details, unclear requirements
- RECOMMENDED: Nice to fix. Style, minor clarity improvements
### Decision Logic
- APPROVED: No CRITICAL issues AND <3 IMPORTANT issues
- REJECTED: Any CRITICAL issue exists
- CONDITIONAL: No CRITICAL but ≥3 IMPORTANT issues
## Constraints
- Never skip mandatory sections in output
- Always provide specific line/section references
- Each suggestion must be actionable (not "improve clarity" but "replace X with Y")
- Use exact severity definitions above
Implementation Challenges
CLI Authentication and Timeouts
Cursor CLI can take a long time to respond depending on task complexity. The default timeout is 5 minutes. For complex tasks, extend it in your MCP config:
"EXECUTION_TIMEOUT_MS": "600000" # 10 minutes (maximum)
When I ran the document-reviewer on a 14,000-character document, Cursor CLI sometimes took over 10 minutes. For quick testing, use smaller files or simplify your agent definitions.
You can specify AGENT_TYPE as either cursor (for Cursor CLI) or claude (for Claude Code).
To install Cursor CLI:
curl https://cursor.com/install -fsS | bash
Cursor CLI requires authentication before use. Sessions expire periodically, so if the MCP stops responding, try logging in again:
% cursor-agent login
Technical Implementation Details
MCP Server Structure
export class McpServer {
private setupHandlers(): void {
// Implementing the run_agent tool
this.server.setRequestHandler(
CallToolRequestSchema,
async (request): Promise<CallToolResult> => {
if (request.params.name === 'run_agent') {
const result = await this.runAgentTool.execute(request.params.arguments)
return result as CallToolResult
}
throw new ValidationError(`Unknown tool: ${request.params.name}`)
}
)
// Publishing resources (exposing agent definitions as MCP resources)
this.server.setRequestHandler(
ListResourcesRequestSchema,
async (): Promise<ListResourcesResult> => {
const resources = await this.agentResources.listResources()
return { resources }
}
)
}
}
Building This MCP with Agentic Coding
When implementing this MCP server, I used my Agentic Coding boilerplate, which comes with several sub-agents that helped ensure code quality throughout development:
-
Adaptive Rule Selection: The
rule-advisorsub-agent analyzed each task and selected only the necessary coding rules, keeping contexts lean -
Staged Quality Assurance: The
quality-fixersub-agent automatically ran type checks and tests, fixing errors before they accumulated -
Pre-implementation Approval: The
TodoWritepattern required my approval before code changes, preventing the AI from going off-track
These boilerplate sub-agents made it possible to build this MCP server reliably — they're examples of the very patterns this MCP now enables for everyone.
run_agent Tool Specification
The run_agent tool accepts:
-
agent: Agent name to execute (required) -
prompt: Instructions for the agent (required) -
cwd: Working directory (optional) -
extra_args: Additional command-line arguments (optional)
sub-agents-mcp passes the entire agent definition file content as system context to Cursor CLI or Claude Code.
What I Learned
I initially thought it would be simple — just pass prompts to CLI tools and return responses. But I needed streaming interactions, and response formats differ between LLMs, requiring careful abstraction.
Since I was setting up my Agentic Coding environment as a hobby project, it took longer than expected with significant refactoring mid-project. Still, being able to quickly build exactly what I need is satisfying.
Through this development, I realized that effective AI tool usage requires proper context management and task decomposition.
Future Plans
While it's already usable, I'm looking forward to JSON output support once Cursor CLI adds it. Currently receiving streaming text, structured data would open up more possibilities.
Got ideas for new agents or improvements? Issues and PRs are welcome.
Repository
shinpr
/
sub-agents-mcp
Define task-specific AI sub-agents in Markdown for any MCP-compatible tool.
Sub-Agents MCP Server
Run reusable coding agents from any MCP-compatible client.
Write a reviewer, test writer, or investigator in Markdown, then ask your assistant to use it. The MCP server runs that agent with the coding CLI you choose and returns the result to the same conversation.
What You Can Do
- Delegate code review, test writing, investigation, and documentation to focused agents
- Reuse the same agent definitions across MCP clients with one shared backend and model configuration
- Continue the same agent across multiple calls for longer work
Quick Start
You need Node.js 22 or later, an MCP-compatible client, and one supported coding CLI installed and signed in. This example uses Codex.
1. Create an Agent
Create an agents folder anywhere on your machine, then add code-reviewer.md:
# Code Reviewer
Review code for bugs and maintainability issues.
## Task
- Find concrete problems in the requested changes
- Explain why…Agent Skills Version
Prefer a lighter setup? There's also a Skills-based approach:
shinpr
/
sub-agents-skills
Cross-LLM sub-agent orchestration as an Agent Skills. Route tasks to Codex, Claude Code, Grok, GLM, Kimi, Cursor, Gemini, OpenCode, or Command Code from any compatible tool.
Sub-Agents Skills
English | 简体中文 | Русский | Deutsch | Español
Run task-specific agents on different AI coding backends from a single parent tool.
Write an agent once in Markdown, then choose the backend that runs it. Send implementation, review, investigation, and verification work to different coding tools without duplicating the agent definition.
The skill itself follows the Agent Skills standard; the agents it runs are Markdown files under .agents/.
Quick Start
Requirements: Python 3.9+ and at least one supported backend installed.
1. Install the Skill
Codex (plugin):
codex plugin marketplace add shinpr/sub-agents-skills
Then open the plugin picker, install Runner, and restart Codex:
/plugins
After restart, invoke the skill as $runner:sub-agents.
Claude Code (plugin):
/plugin marketplace add shinpr/sub-agents-skills
/plugin install runner@sub-agents-skills
/reload-plugins
Grok Build (plugin):
grok plugin marketplace add shinpr/sub-agents-skills
grok plugin install runner --trust
Google Antigravity (plugin):
agy plugin install https://github.com/shinpr/sub-agents-skills/tree/main/plugins/runner
Other clients (Cursor CLI, VS…
Just drop the skill files into your project—no MCP server configuration needed. The tradeoff: error handling depends on your AI client, and there's no session persistence between agent calls.



Top comments (2)
The context isolation point is key. I run multiple subagents in parallel for research tasks and the biggest win isn't speed - it's that each agent stays focused on its specific job without getting confused by the other agents' context. The markdown-based agent definitions are a nice touch for team sharing. Curious how you handle the case where a sub-agent needs to pass structured data back to the parent - do you just return it as text in the completion or is there a typed interface?
Not sure which layer you meant, so I'll cover both. The CLI itself (cursor-agent -p) returns JSON/JSONL as the transport format. For the agent's actual output content, it's plain text - no typed interface. If you want some structure, the practical approach is defining the expected JSON format in the agent's markdown definition and having the parent treat it as 'this is probably what comes back.'
Stricter guarantees are on the table as a potential enhancement, though there are inherent limits since you're dealing with LLM output either way. That said, if you find yourself chasing tighter typing, it's often a sign the agent's scope has grown too large - narrowing the task tends to make the schema simple enough that it stops being a concern.