<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kavin Arvind Ragavan</title>
    <description>The latest articles on DEV Community by Kavin Arvind Ragavan (@kavin_arvind_8a1adbd39efd).</description>
    <link>https://dev.to/kavin_arvind_8a1adbd39efd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3120530%2F1e4509d6-3c14-465b-966a-2b2a3d15bac9.jpg</url>
      <title>DEV Community: Kavin Arvind Ragavan</title>
      <link>https://dev.to/kavin_arvind_8a1adbd39efd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kavin_arvind_8a1adbd39efd"/>
    <language>en</language>
    <item>
      <title>Build Your First GitHub Copilot Custom Agent: A Practical LoadRunner Example</title>
      <dc:creator>Kavin Arvind Ragavan</dc:creator>
      <pubDate>Thu, 24 Sep 2026 04:51:35 +0000</pubDate>
      <link>https://dev.to/kavin_arvind_8a1adbd39efd/build-your-first-github-copilot-custom-agent-a-practical-loadrunner-example-4bo0</link>
      <guid>https://dev.to/kavin_arvind_8a1adbd39efd/build-your-first-github-copilot-custom-agent-a-practical-loadrunner-example-4bo0</guid>
      <description>&lt;h2&gt;
  
  
  Creating Custom Agents in GitHub Copilot
&lt;/h2&gt;

&lt;p&gt;A good Copilot conversation can solve a one-off problem. A custom agent turns that conversation into a repeatable engineering capability: a named persona with a defined job, a controlled tool surface, and domain knowledge that can be versioned with the repository.&lt;/p&gt;

&lt;p&gt;This post is a practical guide to building one. It uses a LoadRunner Agent as the running example, but the design applies equally well to agents for test automation, incident response, documentation, code review, or platform operations.&lt;/p&gt;

&lt;p&gt;The goal is not to create a longer system prompt. The goal is to create a small, inspectable workflow that your team can select from Copilot Chat and trust to produce the same kind of result every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What a custom agent is
&lt;/h2&gt;

&lt;p&gt;A GitHub Copilot custom agent is a scoped, named persona with its own instructions, tools, and optional skills. It is defined by a Markdown file with YAML frontmatter and is discoverable from the Chat agent dropdown.&lt;/p&gt;

&lt;p&gt;A custom agent can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Give Copilot a stable role, such as &lt;code&gt;LoadRunner Agent&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Define the tools that the agent is allowed to call.&lt;/li&gt;
&lt;li&gt;Encode a workflow such as &lt;code&gt;EXTRACT -&amp;gt; GENERATE -&amp;gt; REVIEW&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Load focused domain knowledge only when a task requires it.&lt;/li&gt;
&lt;li&gt;Produce consistent output for everyone working in the repository.&lt;/li&gt;
&lt;li&gt;Keep the configuration versioned alongside the code it operates on.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important distinction is between ad-hoc prompting and an engineered agent. A prompt says what you want right now. An agent defines how a class of requests should be handled, which files may be touched, what checks must run, and what a complete result looks like.&lt;/p&gt;

&lt;p&gt;The usual location is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/agents/&amp;lt;name&amp;gt;.agent.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent then appears as a selectable mode in Copilot Chat. Its instructions are not hidden in one engineer's chat history; they are reviewable, changeable, and shareable in source control.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The four customization primitives
&lt;/h2&gt;

&lt;p&gt;GitHub Copilot customization has four closely related primitives. They solve different problems, so choosing the right one matters.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Primitive&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Typical file&lt;/th&gt;
&lt;th&gt;How it is activated&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;A named persona with tools and a system prompt&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.agent.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Selected from the Chat agent dropdown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skill&lt;/td&gt;
&lt;td&gt;Focused domain knowledge or a procedural playbook&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.SKILL.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Loaded on demand by an agent or user&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction&lt;/td&gt;
&lt;td&gt;Always-on rules scoped to matching files&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.instructions.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Applied automatically through &lt;code&gt;applyTo&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt&lt;/td&gt;
&lt;td&gt;A reusable, parameterized slash command&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.prompt.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Invoked with &lt;code&gt;/prompt-name&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There are also repository-wide files that help discovery and consistency:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;copilot-instructions.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Rules every chat should see in the workspace&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.github/&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AGENTS.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Repository-level agent index or discovery guide&lt;/td&gt;
&lt;td&gt;Repository root&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A useful rule of thumb is simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use an &lt;strong&gt;agent&lt;/strong&gt; when the user needs a named role and controlled capabilities.&lt;/li&gt;
&lt;li&gt;Use a &lt;strong&gt;skill&lt;/strong&gt; when the agent needs deep knowledge for one phase or capability.&lt;/li&gt;
&lt;li&gt;Use an &lt;strong&gt;instruction&lt;/strong&gt; when a rule should apply automatically to matching files.&lt;/li&gt;
&lt;li&gt;Use a &lt;strong&gt;prompt&lt;/strong&gt; when people should invoke a repeatable request on demand.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not put every rule in the agent body. A large prompt becomes difficult to review, expensive to load, and easy to contradict. Keep the agent responsible for orchestration and move detailed procedures into skills.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. How the pieces fit together
&lt;/h2&gt;

&lt;p&gt;The relationship between these files is easier to understand as a flow:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    U[Engineer request] --&amp;gt; A[Custom agent]
    A --&amp;gt; I[System instructions]
    A --&amp;gt; S[Load matching skills]
    A --&amp;gt; T[Call allowed tools]
    I --&amp;gt; T
    S --&amp;gt; T
    T --&amp;gt; F[Files, commands, and search]
    F --&amp;gt; R[Reviewed output]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The agent is the coordinator. Skills provide specialized knowledge. Instructions provide rules that are always in force for a matching file pattern. Tools provide the ability to inspect, modify, or execute.&lt;/p&gt;

&lt;p&gt;That separation creates useful boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The agent decides the phase and sequence.&lt;/li&gt;
&lt;li&gt;A skill explains how to perform one phase correctly.&lt;/li&gt;
&lt;li&gt;Instructions protect repository conventions.&lt;/li&gt;
&lt;li&gt;Tool permissions limit what the agent can actually do.&lt;/li&gt;
&lt;li&gt;Human review remains the final control for generated code and side effects.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Anatomy of an &lt;code&gt;.agent.md&lt;/code&gt; file
&lt;/h2&gt;

&lt;p&gt;An agent file has two parts: YAML frontmatter at the top and a Markdown system prompt below it.&lt;/p&gt;

&lt;p&gt;A minimal example looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;LoadRunner Agent&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Generates, reviews, and summarizes LoadRunner Web Vuser scripts from API project files.&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;execute/runInTerminal&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;read/readFile&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;edit/createFile&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;edit/editFiles&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;search/codebase&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;search/fileSearch&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

You are an expert LoadRunner script engineer.

Operate in three phases:
EXTRACT -&amp;gt; GENERATE -&amp;gt; REVIEW.

Do not skip a phase. Keep generated files inside the configured output directory.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The four important fields are:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The label shown in the Chat agent dropdown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;description&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Helps users and Copilot understand when this agent is relevant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tools&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The allow-list of tool or toolset IDs available to the agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Body&lt;/td&gt;
&lt;td&gt;The system prompt containing phases, rules, boundaries, and completion criteria&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The description deserves more care than it usually gets. It should say what the agent produces, what inputs it understands, and when it should be selected. A vague description makes the agent harder to discover and easier to misuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Tool IDs: use the namespaced form
&lt;/h2&gt;

&lt;p&gt;The most disruptive failure in the walkthrough was also one of the easiest to miss: an agent can appear to load while silently losing the tools it needs because the IDs are wrong.&lt;/p&gt;

&lt;p&gt;Use namespaced IDs in the form:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;category/toolName
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;execute/runInTerminal&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;read/readFile&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;edit/createFile&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;edit/editFiles&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;search/codebase&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;search/fileSearch&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The category communicates the capability boundary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;execute/&lt;/code&gt; for commands and terminal execution.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;read/&lt;/code&gt; for reading files and content.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;edit/&lt;/code&gt; for creating or modifying files.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;search/&lt;/code&gt; for codebase, file, or text search.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These forms are not interchangeable with informal names such as &lt;code&gt;runCommands&lt;/code&gt;, &lt;code&gt;editFiles&lt;/code&gt;, &lt;code&gt;create_file&lt;/code&gt;, or &lt;code&gt;read_file&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Incorrect: raw tool names or toolset names&lt;/span&gt;
 &lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;runCommands&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;editFiles&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;create_file&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;read_file&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Correct: namespaced tool IDs&lt;/span&gt;
 &lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;execute/runInTerminal&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;read/readFile&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;edit/createFile&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;edit/editFiles&lt;/span&gt;
   &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;search/codebase&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An unknown ID may be dropped without an obvious YAML error. The resulting agent still appears in Chat, but it cannot write files, run the extractor, or search the repository as intended.&lt;/p&gt;

&lt;p&gt;After changing the YAML, open the Chat tools picker and verify the exact tool IDs available to the agent. Then reload the agent. Configuration changes do not necessarily hot-reload into an already active chat session.&lt;/p&gt;

&lt;p&gt;Tool IDs are a version-sensitive integration detail, not a universal API contract. Treat the names in this post as examples for the environment in which the agent was authored. The Chat tools picker is the source of truth for the exact IDs available in your VS Code installation.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Skills: package domain knowledge on demand
&lt;/h2&gt;

&lt;p&gt;A skill is a focused, self-contained playbook for one capability. Skills keep the main agent prompt short while allowing the workflow to carry detailed rules.&lt;/p&gt;

&lt;p&gt;A typical location is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/agents/skills/&amp;lt;name&amp;gt;.SKILL.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the LoadRunner Agent, the skills can map directly to the workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/agents/skills/
    lr-extract.SKILL.md
    lr-generate.SKILL.md
    lr-review.SKILL.md
    lr-parameterize.SKILL.md
    lr-boilerplate.SKILL.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each file should have one job:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Skill&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;lr-extract&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Detect SoapUI or Postman input and produce normalized request files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;lr-generate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Convert one normalized request into a LoadRunner C script&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;lr-review&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Run the quality checklist and summarize findings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;lr-parameterize&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Build data files from extracted values&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;lr-boilerplate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Generate functional overview documentation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A skill should tell the agent what to read, what to produce, which rules to apply, and how to recognize success. It should not assume that the agent remembers a procedure from an earlier conversation.&lt;/p&gt;

&lt;p&gt;The agent can chain several skills when the task crosses phases. That is more maintainable than putting extraction rules, C coding conventions, and review criteria into one giant system prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. The LoadRunner Agent design
&lt;/h2&gt;

&lt;p&gt;The example agent converts SoapUI or Postman API project files into LoadRunner Web Vuser scripts. Its workflow has three explicit phases:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Input files in input/] --&amp;gt; B{Detect format}
    B --&amp;gt;|SoapUI XML| C[Run SoapUI extractor]
    B --&amp;gt;|Postman JSON| D[Run Postman extractor]
    C --&amp;gt; E[One normalized .txt per request]
    D --&amp;gt; E
    E --&amp;gt; F[Read one request]
    F --&amp;gt; G[Generate one .c script]
    G --&amp;gt; H{More requests?}
    H --&amp;gt;|Yes| F
    H --&amp;gt;|No| I[Run review checklist]
    I --&amp;gt; J[Pass/fail table and recommendations]&lt;/code&gt;&lt;/pre&gt;



&lt;h3&gt;
  
  
  Phase 1: EXTRACT
&lt;/h3&gt;

&lt;p&gt;The agent auto-detects input files in the configured input directory, selects the appropriate extractor, and produces one normalized text payload per API request.&lt;/p&gt;

&lt;p&gt;This phase should answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which source format was found?&lt;/li&gt;
&lt;li&gt;Which extractor will run?&lt;/li&gt;
&lt;li&gt;How many requests were discovered?&lt;/li&gt;
&lt;li&gt;Where were the normalized payloads written?&lt;/li&gt;
&lt;li&gt;Did extraction produce warnings or incomplete requests?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The extraction result is the contract for the generation phase. If the contract is ambiguous, generation should stop and report the problem instead of guessing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Phase 2: GENERATE
&lt;/h3&gt;

&lt;p&gt;The agent reads each normalized request one at a time and writes one C script per request. The one-at-a-time rule is deliberate. It avoids loading all payloads into context, reduces cross-request confusion, and makes it clear which input produced which output.&lt;/p&gt;

&lt;p&gt;The generation skill can encode rules such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use stable transaction names.&lt;/li&gt;
&lt;li&gt;Add Dynatrace correlation headers where required.&lt;/li&gt;
&lt;li&gt;Register expected response content with &lt;code&gt;web_reg_find&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Add think time at realistic business boundaries.&lt;/li&gt;
&lt;li&gt;Handle request and response errors explicitly.&lt;/li&gt;
&lt;li&gt;Preserve the method, URL, headers, and body from the extracted request.&lt;/li&gt;
&lt;li&gt;Keep generated output in the repository's locked folder convention.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent should write the output immediately after processing each input:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read request-001.txt
write request-001.c
read request-002.txt
write request-002.c
read request-003.txt
write request-003.c
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That loop is safer than reading every input file first and attempting to generate all scripts at the end.&lt;/p&gt;

&lt;h3&gt;
  
  
  Phase 3: REVIEW
&lt;/h3&gt;

&lt;p&gt;The review phase runs a fixed checklist against the generated scripts. It should return evidence, not just a statement that the files look correct.&lt;/p&gt;

&lt;p&gt;A useful result contains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A pass/fail status for each checklist item.&lt;/li&gt;
&lt;li&gt;The files examined.&lt;/li&gt;
&lt;li&gt;The exact issue for every failed item.&lt;/li&gt;
&lt;li&gt;The top three issues that should be fixed first.&lt;/li&gt;
&lt;li&gt;The top five recommendations for the next iteration.&lt;/li&gt;
&lt;li&gt;Any assumptions that still require a human decision.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The review phase is where an agent becomes more than a file generator. It gives the team a consistent quality gate after generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. A practical agent body
&lt;/h2&gt;

&lt;p&gt;The system prompt should be explicit about phases and boundaries. For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;You are an expert LoadRunner script engineer.

Your job is to convert SoapUI XML or Postman JSON project files into reviewed
LoadRunner Web Vuser scripts.

Always operate in this order:
&lt;span class="p"&gt;
1.&lt;/span&gt; EXTRACT
&lt;span class="p"&gt;   -&lt;/span&gt; Detect supported input files under input/soap/ or input/postman/.
&lt;span class="p"&gt;   -&lt;/span&gt; Run the matching extractor.
&lt;span class="p"&gt;   -&lt;/span&gt; Confirm the normalized request files that were created.
&lt;span class="p"&gt;
2.&lt;/span&gt; GENERATE
&lt;span class="p"&gt;   -&lt;/span&gt; Read exactly one normalized request file at a time.
&lt;span class="p"&gt;   -&lt;/span&gt; Create exactly one .c file for that request.
&lt;span class="p"&gt;   -&lt;/span&gt; Apply every rule in lr-generate.SKILL.md.
&lt;span class="p"&gt;   -&lt;/span&gt; Continue until every normalized request has an output file.
&lt;span class="p"&gt;
3.&lt;/span&gt; REVIEW
&lt;span class="p"&gt;   -&lt;/span&gt; Run every item in lr-review.SKILL.md.
&lt;span class="p"&gt;   -&lt;/span&gt; Produce a pass/fail table.
&lt;span class="p"&gt;   -&lt;/span&gt; Report the top three issues and top five recommendations.

Do not invent missing request data. Do not skip extraction or review.
Do not batch-read large collections of input files.
Do not create helper scripts to replace this workflow.
Keep all outputs inside the repository's documented directories.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The details belong in skills, but the agent body must still define the non-negotiable orchestration rules. In particular, say what the agent must not do. Otherwise a general-purpose agent may decide that generating a helper script is a convenient shortcut, even when the agent itself is supposed to perform the work.&lt;/p&gt;

&lt;h3&gt;
  
  
  A minimal starter set
&lt;/h3&gt;

&lt;p&gt;The smallest useful implementation can be three files. The agent coordinates the workflow, the skill contains the generation rules, and the workspace instructions lock down paths and naming.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;.github/agents/loadrunner.agent.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;LoadRunner Agent&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Converts API project files into reviewed LoadRunner Web Vuser scripts.&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;execute/runInTerminal&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;read/readFile&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;edit/createFile&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;edit/editFiles&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;search/fileSearch&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

Convert supported files under input/ into LoadRunner scripts under output/scripts/.

Always run these phases in order:
&lt;span class="p"&gt;1.&lt;/span&gt; EXTRACT: identify the input format and create normalized request files.
&lt;span class="p"&gt;2.&lt;/span&gt; GENERATE: read one normalized request and write one .c file at a time.
&lt;span class="p"&gt;3.&lt;/span&gt; REVIEW: run the checklist in lr-generate.SKILL.md and report failures.

Do not invent missing request data, batch-read the entire input directory, or create
helper scripts. Keep all generated files under output/.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.github/agents/skills/lr-generate.SKILL.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# LoadRunner generation rules&lt;/span&gt;

For each normalized request:
&lt;span class="p"&gt;
-&lt;/span&gt; Preserve the HTTP method, URL, headers, and request body.
&lt;span class="p"&gt;-&lt;/span&gt; Use a stable transaction name derived from the request name.
&lt;span class="p"&gt;-&lt;/span&gt; Register an expected response with web_reg_find when a reliable check exists.
&lt;span class="p"&gt;-&lt;/span&gt; Add required Dynatrace headers.
&lt;span class="p"&gt;-&lt;/span&gt; Add realistic think time only at business-flow boundaries.
&lt;span class="p"&gt;-&lt;/span&gt; Report missing correlation candidates instead of guessing replacements.
&lt;span class="p"&gt;-&lt;/span&gt; Write exactly one reviewed candidate script for the input request.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.github/copilot-instructions.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="p"&gt;-&lt;/span&gt; Read source inputs from input/.
&lt;span class="p"&gt;-&lt;/span&gt; Write generated scripts only to output/scripts/.
&lt;span class="p"&gt;-&lt;/span&gt; Write review summaries only to output/reports/.
&lt;span class="p"&gt;-&lt;/span&gt; Use one output file per normalized request.
&lt;span class="p"&gt;-&lt;/span&gt; Never store credentials, tokens, or production data in generated files.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This starter set is intentionally small. Add extraction, parameterization, and review skills as the workflow gains real requirements, rather than making the first agent responsible for every possible LoadRunner task.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Why one file in and one file out matters
&lt;/h2&gt;

&lt;p&gt;Large input collections create two predictable problems:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Context pressure. Reading 29 payloads at once forces summarization, and the agent may need to reread information it already had.&lt;/li&gt;
&lt;li&gt;Output ambiguity. When many files are generated together, it becomes harder to detect which input was skipped or which output contains an accidental mix of requests.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A strict loop provides a small checkpoint after every item:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;for each normalized request:
    read one request
    generate one output
    verify the output exists
    continue
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a workflow rule, not a performance optimization. It makes failures local, recoverable, and visible. It also lets a user stop after a particular request without discarding all completed work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security and approval boundaries
&lt;/h2&gt;

&lt;p&gt;An agent's tool list is an allow-list, but it is not a complete security boundary. A tool that can run a terminal command or write a file can still have meaningful side effects. Design the agent with the same caution you would apply to a CI job or an automation service account.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Grant only the tools required for the workflow. A read-and-review agent should not receive file-write or terminal permissions.&lt;/li&gt;
&lt;li&gt;Keep credentials in environment variables, secret stores, or the platform's approved authentication mechanism. Never place tokens in &lt;code&gt;.agent.md&lt;/code&gt;, skills, prompts, examples, or generated scripts.&lt;/li&gt;
&lt;li&gt;Treat SoapUI, Postman, and other imported project files as untrusted input. Do not execute commands or URLs found in a payload without an explicit rule allowing that behavior.&lt;/li&gt;
&lt;li&gt;Keep generated files inside a known workspace directory and review path-handling rules before enabling terminal execution.&lt;/li&gt;
&lt;li&gt;Require human approval for commands that start load tests, modify shared infrastructure, send messages, delete data, or access production systems.&lt;/li&gt;
&lt;li&gt;Use test accounts and sanitized request data when demonstrating the workflow or committing sample inputs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent should also report what it did: files read, commands run, files created, warnings, and assumptions. That audit trail makes a generated result easier to review and easier to investigate when something goes wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Common pitfalls and fixes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Wrong tool IDs
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; The agent loads but cannot create files or run commands.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cause:&lt;/strong&gt; The &lt;code&gt;tools&lt;/code&gt; list contains raw tool names or unsupported toolset names.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix:&lt;/strong&gt; Use namespaced IDs such as &lt;code&gt;edit/createFile&lt;/code&gt; and &lt;code&gt;execute/runInTerminal&lt;/code&gt;. Verify them in the Chat tools picker after every YAML change.&lt;/p&gt;

&lt;h3&gt;
  
  
  The agent writes a helper script
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; Instead of processing the inputs, the agent creates a Python or shell utility intended to do the work later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cause:&lt;/strong&gt; The prompt did not clearly state that the agent is the program and that helper-script generation is out of scope.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix:&lt;/strong&gt; Add an explicit prohibition to both the agent instructions and the relevant skill. State the permitted tools and the required output directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Batch reads exhaust the context
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; The agent summarizes the input collection, loses details, or repeatedly rereads files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cause:&lt;/strong&gt; Large payloads were loaded in one operation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix:&lt;/strong&gt; Mandate the one-file-in, one-file-out loop and write each result before moving to the next input.&lt;/p&gt;

&lt;h3&gt;
  
  
  A tool is disabled in the session
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; The YAML is correct, but the agent still cannot call a required tool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cause:&lt;/strong&gt; The tool was unchecked in the Chat tool picker, or the chat session was created before the configuration changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix:&lt;/strong&gt; Enable the tool, reload the agent, and start a fresh session when necessary.&lt;/p&gt;

&lt;h3&gt;
  
  
  Output structure drifts
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; Different runs create different folder layouts or naming conventions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cause:&lt;/strong&gt; The required paths were implied instead of declared.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix:&lt;/strong&gt; Lock folder and naming conventions in &lt;code&gt;.github/copilot-instructions.md&lt;/code&gt; and repeat the output contract in the agent or skill that writes the files.&lt;/p&gt;

&lt;h3&gt;
  
  
  YAML changes appear to do nothing
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; The old description or tool set remains active after editing the agent file.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cause:&lt;/strong&gt; The active agent instance has not reloaded its configuration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix:&lt;/strong&gt; Reload the agent and verify the current tool list before debugging the workflow itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. A repository layout that scales
&lt;/h2&gt;

&lt;p&gt;A small project can begin with this structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/
    agents/
        loadrunner.agent.md
        skills/
            lr-extract.SKILL.md
            lr-generate.SKILL.md
            lr-review.SKILL.md
            lr-parameterize.SKILL.md
            lr-boilerplate.SKILL.md
    instructions/
        loadrunner.instructions.md
    prompts/
        review-loadrunner.prompt.md
    copilot-instructions.md
AGENTS.md
input/
    soap/
    postman/
output/
    scripts/
    reports/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The names are conventions, not requirements. What matters is that the locations and naming rules are stable enough for the agent, the team, and code review to agree on where things belong.&lt;/p&gt;

&lt;p&gt;A useful division of responsibility is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;.agent.md&lt;/code&gt;: identity, phases, tool allow-list, and completion behavior.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;.SKILL.md&lt;/code&gt;: detailed rules for extraction, generation, parameterization, and review.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;.instructions.md&lt;/code&gt;: file-pattern-specific coding conventions.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;copilot-instructions.md&lt;/code&gt;: workspace-wide rules and path conventions.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;.prompt.md&lt;/code&gt;: shortcuts for recurring user requests.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;AGENTS.md&lt;/code&gt;: a human-readable index of available agents and their intended jobs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  12. Design checklist
&lt;/h2&gt;

&lt;p&gt;Before sharing a custom agent, verify the following:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The agent has one clear job and a useful description.&lt;/li&gt;
&lt;li&gt;The file is stored under &lt;code&gt;.github/agents/&lt;/code&gt; and is committed with the repository.&lt;/li&gt;
&lt;li&gt;The tool list uses exact namespaced IDs.&lt;/li&gt;
&lt;li&gt;The Chat tools picker shows every required tool as enabled.&lt;/li&gt;
&lt;li&gt;The system prompt names its phases explicitly.&lt;/li&gt;
&lt;li&gt;Detailed domain rules live in focused skills.&lt;/li&gt;
&lt;li&gt;Inputs and outputs have fixed paths and naming conventions.&lt;/li&gt;
&lt;li&gt;Large input sets are processed one item at a time.&lt;/li&gt;
&lt;li&gt;The agent is forbidden from creating helper scripts when direct execution is required.&lt;/li&gt;
&lt;li&gt;The workflow reports warnings and assumptions instead of silently guessing.&lt;/li&gt;
&lt;li&gt;Generated files receive an explicit review pass.&lt;/li&gt;
&lt;li&gt;YAML changes are tested in a newly reloaded agent session.&lt;/li&gt;
&lt;li&gt;Side-effecting commands have an approval or review boundary.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Smoke-test the agent before sharing it
&lt;/h3&gt;

&lt;p&gt;Run a deliberately small test after creating or changing the agent:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Reload the agent and confirm the expected tools are enabled in the Chat tools picker.&lt;/li&gt;
&lt;li&gt;Provide one sanitized request file, not the full project collection.&lt;/li&gt;
&lt;li&gt;Confirm that extraction produces the expected normalized request.&lt;/li&gt;
&lt;li&gt;Confirm that generation writes exactly one output file in the documented directory.&lt;/li&gt;
&lt;li&gt;Inspect the generated script for method, URL, headers, body, checks, and credentials.&lt;/li&gt;
&lt;li&gt;Run the review phase and verify that its report names the file it examined.&lt;/li&gt;
&lt;li&gt;Repeat with a malformed or incomplete input and confirm that the agent reports the problem instead of inventing values.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This test catches the most expensive configuration failures early: missing tools, wrong paths, skipped phases, accidental batching, and silent guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  13. The larger lesson
&lt;/h2&gt;

&lt;p&gt;Custom agents are small software systems. Their Markdown files may look simple, but the same engineering concerns still apply: interface contracts, permissions, state, failure handling, observability, and tests.&lt;/p&gt;

&lt;p&gt;The agent is the orchestration layer. Skills are the domain modules. Instructions are the policy layer. Tools are the execution boundary. A reliable result comes from designing those layers together, then checking that the selected tools and repository paths match the design.&lt;/p&gt;

&lt;p&gt;For the LoadRunner example, the winning pattern is straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXTRACT -&amp;gt; GENERATE -&amp;gt; REVIEW
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Make each phase explicit, keep each skill focused, process large inputs incrementally, and make tool permissions visible. The result is not merely a clever prompt. It is a repeatable engineering workflow that a team can inspect, improve, and run from GitHub Copilot.&lt;/p&gt;

&lt;h2&gt;
  
  
  14. Quick reference
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Need&lt;/th&gt;
&lt;th&gt;Put it in&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Named specialist with tools&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.agent.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;loadrunner.agent.md&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Detailed domain procedure&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.SKILL.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;lr-review.SKILL.md&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rule for matching files&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.instructions.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;loadrunner.instructions.md&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reusable slash command&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.prompt.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;review-loadrunner.prompt.md&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workspace-wide rule&lt;/td&gt;
&lt;td&gt;&lt;code&gt;copilot-instructions.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Output paths and naming&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent discovery index&lt;/td&gt;
&lt;td&gt;&lt;code&gt;AGENTS.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Available agents and responsibilities&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Common tool ID families:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool family&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Execute&lt;/td&gt;
&lt;td&gt;&lt;code&gt;execute/runInTerminal&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read&lt;/td&gt;
&lt;td&gt;&lt;code&gt;read/readFile&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Create&lt;/td&gt;
&lt;td&gt;&lt;code&gt;edit/createFile&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Modify&lt;/td&gt;
&lt;td&gt;&lt;code&gt;edit/editFiles&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic search&lt;/td&gt;
&lt;td&gt;&lt;code&gt;search/codebase&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;File search&lt;/td&gt;
&lt;td&gt;&lt;code&gt;search/fileSearch&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text search&lt;/td&gt;
&lt;td&gt;&lt;code&gt;search/textSearch&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Start with one narrow agent, one clear workflow, and one review checklist. Add skills as the workflow grows; do not make the first agent responsible for everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/copilot/customization/custom-agents" rel="noopener noreferrer"&gt;VS Code custom agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/copilot/customization/prompt-files" rel="noopener noreferrer"&gt;VS Code prompt files&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/agent-customization/custom-instructions" rel="noopener noreferrer"&gt;VS Code custom instructions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/copilot/customizing-copilot/adding-repository-custom-instructions-for-github-copilot" rel="noopener noreferrer"&gt;GitHub Copilot custom instructions&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Documentation and available tool IDs change over time. Check the current VS Code documentation and the Chat tools picker before copying a configuration into a production repository.&lt;/p&gt;

</description>
      <category>githubcopilot</category>
      <category>ai</category>
      <category>loadtesting</category>
      <category>agentskills</category>
    </item>
    <item>
      <title>Creating a Custom Copilot Agent for Chaos Engineering Using LitmusChaos and MCP</title>
      <dc:creator>Kavin Arvind Ragavan</dc:creator>
      <pubDate>Wed, 23 Sep 2026 19:28:31 +0000</pubDate>
      <link>https://dev.to/kavin_arvind_8a1adbd39efd/creating-a-custom-copilot-agent-for-chaos-engineering-using-litmuschaos-and-mcp-2g6n</link>
      <guid>https://dev.to/kavin_arvind_8a1adbd39efd/creating-a-custom-copilot-agent-for-chaos-engineering-using-litmuschaos-and-mcp-2g6n</guid>
      <description>&lt;p&gt;Chaos engineering is often introduced as a command-line or dashboard workflow: install a platform, register a Kubernetes cluster, create an experiment, and run it. That workflow is useful, but it leaves a gap between an SRE's or platform engineer's intent and safe execution.&lt;/p&gt;

&lt;p&gt;The engineer should be able to work with the agent through a sequence of small, reviewable prompts: check readiness, inspect the target, review the experiment, approve execution, and verify recovery. This keeps discovery and approval separate from the destructive action.&lt;/p&gt;

&lt;p&gt;That is the idea behind this project. I created a custom GitHub Copilot agent whose job is to operate LitmusChaos safely. The agent uses a Go-based Model Context Protocol (MCP) server for LitmusChaos API operations and uses Kubernetes read-only checks to verify local targets and recovery.&lt;/p&gt;

&lt;p&gt;This post walks through the agent design, skill workflow, MCP integration, safety model, local Litmus setup, troubleshooting, and lessons learned.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Scope note:&lt;/strong&gt; This article is about the LitmusChaos MCP integration and the custom Copilot agent. The &lt;code&gt;demo-nginx&lt;/code&gt; workload is only an optional local Kubernetes validation target; it is separate from the MCP server and the agent.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Target Audience
&lt;/h2&gt;

&lt;p&gt;This project is intended for engineers who already work with Kubernetes workloads and need a safer, more conversational way to run resilience tests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Site reliability engineers (SREs)&lt;/strong&gt; validating service recovery and steady-state behavior.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Platform engineers&lt;/strong&gt; providing controlled chaos capabilities to development teams.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chaos engineers&lt;/strong&gt; designing and running LitmusChaos experiments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Performance and resilience engineers&lt;/strong&gt; testing failure behavior under controlled conditions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Developers&lt;/strong&gt; who need a disposable local workflow before promoting an experiment to a shared environment.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Throughout this article, the custom agent supports the SRE, platform engineer, or chaos engineer responsible for reviewing and approving a chaos action. It helps make the workflow repeatable, but it does not replace engineering judgment or approval.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Agent Is the Product
&lt;/h2&gt;

&lt;p&gt;The implementation contains five pieces, but the custom agent is the control layer that ties them together:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Docker Desktop Kubernetes&lt;/strong&gt; as the local Kubernetes cluster.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LitmusChaos ChaosCenter&lt;/strong&gt; as the control plane.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Litmus chaos infrastructure&lt;/strong&gt; deployed into the local cluster.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Go MCP server&lt;/strong&gt; that translates MCP tool calls into LitmusChaos GraphQL requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A custom Copilot agent&lt;/strong&gt; that applies preflight checks, safety rules, confirmation gates, execution, and recovery verification.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The agent is designed to work with an explicitly identified workload and environment. It is instructed to reject control-plane namespaces, production environments, and unknown targets unless the responsible engineer provides an approved safety boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Create a Custom Agent?
&lt;/h2&gt;

&lt;p&gt;An MCP server exposes capabilities. It does not automatically define a safe operating procedure.&lt;/p&gt;

&lt;p&gt;Without an agent, an engineer might ask for a chaos run and jump directly to an execution tool. With an agent, the request becomes a controlled workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;natural-language request
  |
  v
find experiment and inspect target
  |
  v
check infrastructure and Kubernetes workload
  |
  v
summarize impact and ask for confirmation
  |
  v
run through MCP and poll status
  |
  v
verify Kubernetes recovery and report evidence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent does not replace LitmusChaos. It provides an opinionated operating model around LitmusChaos.&lt;/p&gt;

&lt;h2&gt;
  
  
  Custom Agent Structure
&lt;/h2&gt;

&lt;p&gt;The workspace agent is defined in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/agents/litmuschaos.agent.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Its frontmatter restricts the agent to the Litmus MCP server plus read-only or observational terminal work:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;LitmusChaos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;MCP&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;operations:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;inspect&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;experiments,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;validate&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;local&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Kubernetes&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;chaos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;targets,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;run&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;or&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;stop&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;disposable&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;chaos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;experiments,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;investigate&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;queued&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;or&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;failed&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;runs,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;report&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;recovery.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Requires&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;confirmation&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;before&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;destructive&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;actions."&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LitmusChaos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Operator"&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;litmuschaos/*&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;execute&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;read&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="na"&gt;user-invocable&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent body defines the behavior that matters more than the persona:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Inspect before executing.&lt;/li&gt;
&lt;li&gt;Verify the infrastructure is active and confirmed.&lt;/li&gt;
&lt;li&gt;Verify the namespace and workload selector.&lt;/li&gt;
&lt;li&gt;Ask for explicit confirmation before run or stop operations.&lt;/li&gt;
&lt;li&gt;Never expose the access token.&lt;/li&gt;
&lt;li&gt;Verify the target recovers after the experiment.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Agent Skill
&lt;/h2&gt;

&lt;p&gt;The repeatable run workflow lives in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/skills/litmus-experiment-runner/SKILL.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The skill is deliberately procedural. It tells the agent to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Find the experiment by name when necessary.&lt;/li&gt;
&lt;li&gt;Load the full experiment definition.&lt;/li&gt;
&lt;li&gt;Check infrastructure state.&lt;/li&gt;
&lt;li&gt;Validate the target namespace and selector.&lt;/li&gt;
&lt;li&gt;Present a confirmation summary.&lt;/li&gt;
&lt;li&gt;Run the experiment through MCP.&lt;/li&gt;
&lt;li&gt;Poll the run until it reaches a terminal state.&lt;/li&gt;
&lt;li&gt;Verify Kubernetes recovery.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This separation is useful: the &lt;code&gt;.agent.md&lt;/code&gt; file defines the agent's role and boundaries, while the skill defines a reusable operational workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tutorial: Create the Agent and Skill
&lt;/h2&gt;

&lt;p&gt;This is the smallest useful implementation. The agent defines the operating policy; the skill defines the repeatable run procedure.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Register the MCP Server
&lt;/h3&gt;

&lt;p&gt;Create &lt;code&gt;.vscode/mcp.json&lt;/code&gt; in the workspace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"servers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"litmuschaos"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stdio"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"${workspaceFolder}/litmus-mcp-server/bin/litmuschaos-mcp-server.exe"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"envFile"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"${workspaceFolder}/litmus-mcp-server/.env"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the token in the ignored &lt;code&gt;.env&lt;/code&gt; file. Do not embed it in the agent, skill, README, screenshots, or chat prompts.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Create the Custom Agent
&lt;/h3&gt;

&lt;p&gt;Create &lt;code&gt;.github/agents/litmuschaos.agent.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Operate&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;LitmusChaos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;through&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;MCP&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;target&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;validation&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;confirmation&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;gates."&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LitmusChaos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Operator"&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;litmuschaos/*&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;execute&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;read&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="na"&gt;user-invocable&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

You are the LitmusChaos Operator.
&lt;span class="p"&gt;
-&lt;/span&gt; Inspect experiments and infrastructure before execution.
&lt;span class="p"&gt;-&lt;/span&gt; Verify the target namespace and workload.
&lt;span class="p"&gt;-&lt;/span&gt; Require explicit confirmation before run or stop operations.
&lt;span class="p"&gt;-&lt;/span&gt; Never expose credentials.
&lt;span class="p"&gt;-&lt;/span&gt; Verify Kubernetes recovery after a local experiment.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The description is important because it is the agent's discovery surface in VS Code. The body establishes behavior and safety boundaries.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Add the Operational Skill
&lt;/h3&gt;

&lt;p&gt;Create &lt;code&gt;.github/skills/litmus-experiment-runner/SKILL.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;litmus-experiment-runner&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Run&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;safe&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;LitmusChaos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;experiments&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;through&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;MCP&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;preflight,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;confirmation,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;polling,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;recovery&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;checks.'&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gh"&gt;# Litmus Experiment Runner&lt;/span&gt;
&lt;span class="p"&gt;
1.&lt;/span&gt; Find the experiment by name or ID.
&lt;span class="p"&gt;2.&lt;/span&gt; Inspect the experiment definition and infrastructure.
&lt;span class="p"&gt;3.&lt;/span&gt; Verify the namespace and workload selector.
&lt;span class="p"&gt;4.&lt;/span&gt; Summarize impact and ask for confirmation.
&lt;span class="p"&gt;5.&lt;/span&gt; Run through MCP only after confirmation.
&lt;span class="p"&gt;6.&lt;/span&gt; Poll the run until it reaches a terminal state.
&lt;span class="p"&gt;7.&lt;/span&gt; Verify Kubernetes recovery.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A skill is a good fit here because the sequence is reusable and procedural. Keep policy in the agent and task-specific procedure in the skill.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Select the Agent in Copilot Chat
&lt;/h3&gt;

&lt;p&gt;Reload the VS Code window if needed, select &lt;strong&gt;LitmusChaos Operator&lt;/strong&gt; in the agent picker, and use Agent mode.&lt;/p&gt;

&lt;p&gt;Start with a read-only request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Check whether Litmus is ready for a chaos test, then list my chaos experiments.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For an execution request against the optional local validation target, the agent should ask for confirmation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run the local `pod-delete-1` experiment against `demo-nginx` in namespace `chaos-test`.
Inspect the experiment and infrastructure first, then ask for confirmation.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After confirmation, the agent can call &lt;code&gt;run_chaos_experiment&lt;/code&gt;, poll &lt;code&gt;list_experiment_runs&lt;/code&gt;, and use Kubernetes read-only commands to verify the local test Deployment recovered. The MCP server and custom agent remain the subject of this article; &lt;code&gt;demo-nginx&lt;/code&gt; is only a sample workload used to exercise them.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Test the MCP Server Directly
&lt;/h3&gt;

&lt;p&gt;The MCP transport can be tested independently of the agent with a read-only JSON-RPC request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"tools/call"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"params"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"list_chaos_experiments"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"arguments"&lt;/span&gt;&lt;span class="p"&gt;:{}}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This distinction helps debugging:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If MCP discovery fails, inspect &lt;code&gt;.vscode/mcp.json&lt;/code&gt; and server startup logs.&lt;/li&gt;
&lt;li&gt;If MCP works but the run fails, inspect the experiment, infrastructure, workflow, and Kubernetes events.&lt;/li&gt;
&lt;li&gt;If the agent skips confirmation, tighten the agent instructions and skill procedure.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;

&lt;p&gt;The request path looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Copilot Chat
    |
    | MCP over stdio
    v
LitmusChaos MCP server (Go)
    |
    | HTTP GraphQL request
    v
ChaosCenter GraphQL server
    |
    | WebSocket/subscriber communication
    v
Litmus chaos infrastructure in Kubernetes
    |
    v
ChaosEngine and experiment runner
    |
    v
Target workload in the approved application namespace
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The MCP server is a local stdio process. VS Code starts it from &lt;code&gt;.vscode/mcp.json&lt;/code&gt;, and the process reads credentials from a local &lt;code&gt;.env&lt;/code&gt; file.&lt;/p&gt;

&lt;p&gt;A simplified MCP configuration looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"servers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"litmuschaos"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stdio"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"${workspaceFolder}/litmus-mcp-server/bin/litmuschaos-mcp-server.exe"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"envFile"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"${workspaceFolder}/litmus-mcp-server/.env"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The MCP server exposes these 16 enabled tools:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_chaos_experiments&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Discover experiments with optional filters and pagination.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_chaos_experiment&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect an experiment, including its manifest and infrastructure.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;run_chaos_experiment&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Start an existing experiment immediately.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;stop_chaos_experiment&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Stop an active experiment or specific run.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_experiment_runs&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Review experiment run history and filter by status.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_experiment_run_details&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect a run, execution state, and optional logs.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_chaos_infrastructures&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Find registered infrastructures and filter by status.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_infrastructure_details&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect infrastructure details and optionally its manifest.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;register_chaos_infrastructure&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Create an infrastructure registration request.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_environments&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;List environments used to organize infrastructures.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;create_environment&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Create a &lt;code&gt;PROD&lt;/code&gt; or &lt;code&gt;NON_PROD&lt;/code&gt; environment.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_resilience_probes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Discover configured resilience probes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;create_resilience_probe&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Create supported HTTP, command, Kubernetes, or Prometheus probes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_chaos_hubs&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Discover available ChaosHubs.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_chaos_faults&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Browse faults exposed by a ChaosHub.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_experiment_statistics&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Review experiment and resiliency-score statistics.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The server currently exposes 16 enabled tools. Experiment creation is intentionally disabled: create and configure experiments in ChaosCenter first, then use MCP to list, inspect, run, stop, and review them.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the MCP Server Does
&lt;/h3&gt;

&lt;p&gt;The Go server is an API adapter. It accepts MCP tool calls over standard input and output, adds the configured project ID and bearer token to the request, calls the ChaosCenter GraphQL endpoint, and returns structured results to Copilot. It does not inspect Kubernetes workloads, decide whether a target is safe, or approve a destructive action.&lt;/p&gt;

&lt;p&gt;Those responsibilities belong to the custom agent and its skill. Before a run, the agent reads the experiment manifest returned by &lt;code&gt;get_chaos_experiment&lt;/code&gt;, checks that the manifest target matches the requested workload, verifies the infrastructure, and asks for explicit confirmation. After starting a run, it treats the start response as an acknowledgement rather than proof of success, polls the run, and investigates the workflow when the run remains active beyond its expected duration.&lt;/p&gt;

&lt;p&gt;This separation also makes failures easier to diagnose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MCP connectivity failure:&lt;/strong&gt; the local GraphQL endpoint or port-forward is unavailable, so the server cannot reach ChaosCenter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API failure:&lt;/strong&gt; ChaosCenter rejects the GraphQL request or returns an error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workflow failure:&lt;/strong&gt; ChaosCenter accepts the request, but Kubernetes or Argo rejects or cannot execute the generated workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Target verification failure:&lt;/strong&gt; the experiment manifest does not match the requested workload, or the local Kubernetes check cannot verify it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent should report which layer failed and should not retry a destructive action automatically.&lt;/p&gt;

&lt;h3&gt;
  
  
  Experiment Coverage
&lt;/h3&gt;

&lt;p&gt;MCP is not a second experiment catalog. It controls experiments already created in the configured ChaosCenter project, so the available experiments depend on the ChaosHub faults and probes installed there. Depending on that project, an engineer may use MCP with pod or container disruption, CPU or memory stress, network or storage faults, node-level faults, and probe-backed experiments that check HTTP, Kubernetes, command, or Prometheus signals.&lt;/p&gt;

&lt;p&gt;The reliable workflow is to discover the actual catalog first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;List my chaos experiments.
Show the available faults in ChaosHub &amp;lt;hub-id&amp;gt;.
Show details for experiment &amp;lt;experiment-id&amp;gt;, including its manifest and recent runs.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are capability categories, not a guarantee that every Litmus installation contains every fault. The ChaosHub and the experiment definition determine which faults and parameters are available.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;p&gt;You need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Windows with Docker Desktop&lt;/li&gt;
&lt;li&gt;Docker Desktop Kubernetes enabled&lt;/li&gt;
&lt;li&gt;&lt;code&gt;kubectl&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Helm 3 or later&lt;/li&gt;
&lt;li&gt;Go 1.21 or later to build the MCP server&lt;/li&gt;
&lt;li&gt;A LitmusChaos project and API token&lt;/li&gt;
&lt;li&gt;VS Code with Copilot Chat and MCP support&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A full ChaosCenter installation is resource-heavy for a small laptop. My host had 8 GB of RAM and an i3 processor, so WSL 2 was configured conservatively:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[wsl2]&lt;/span&gt;
&lt;span class="py"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;4GB&lt;/span&gt;
&lt;span class="py"&gt;processors&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;2&lt;/span&gt;
&lt;span class="py"&gt;swap&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;2GB&lt;/span&gt;
&lt;span class="py"&gt;localhostForwarding&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The WSL configuration lives at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;%UserProfile%\.wslconfig
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After changing it, apply the settings with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;wsl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--shutdown&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Install ChaosCenter
&lt;/h2&gt;

&lt;p&gt;First, select the local Docker Desktop context:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;use-context&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;docker-desktop&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;nodes&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The node should be &lt;code&gt;Ready&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Add the Litmus Helm repository:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;helm&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;repo&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;add&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmuschaos&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;https://litmuschaos.github.io/litmus-helm/&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;helm&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;repo&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;update&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a constrained laptop, use a single MongoDB replica set member rather than the default three-member configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;helm&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;install&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;chaos&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmuschaos/litmus&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--namespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--create-namespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;portal.frontend.service.type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;NodePort&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;mongodb.architecture&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;replicaset&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;mongodb.replicaCount&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nt"&gt;--set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;mongodb.persistence.size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="n"&gt;Gi&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--wait&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--timeout&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;15m&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important detail is &lt;code&gt;architecture=replicaset&lt;/code&gt; with &lt;code&gt;replicaCount=1&lt;/code&gt;. A standalone MongoDB setting looked attractive for a low-resource installation, but Litmus init containers expected the StatefulSet DNS name &lt;code&gt;chaos-mongodb-0.chaos-mongodb-headless&lt;/code&gt;. A one-member replica set preserves that expected service-discovery behavior.&lt;/p&gt;

&lt;p&gt;Verify the installation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;pods&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;svc&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Forward the ChaosCenter frontend:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;port-forward&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;svc/chaos-litmus-frontend-service&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;9091:9091&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://localhost:9091
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sign in with the administrator credentials configured for your local installation, and change the initial password immediately. Do not publish administrator credentials, tokens, or screenshots containing them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Register the Local Chaos Infrastructure
&lt;/h2&gt;

&lt;p&gt;ChaosCenter requires an environment and an infrastructure before it can run experiments.&lt;/p&gt;

&lt;p&gt;Create a non-production environment, for example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Name: local-docker
Type: NON_PROD
Description: Local Docker Desktop Kubernetes test environment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use ChaosCenter's &lt;strong&gt;Enable Chaos&lt;/strong&gt; workflow to generate the infrastructure manifest. Apply the downloaded YAML with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;use-context&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;docker-desktop&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;apply&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-f&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;local-docker-chaos-litmus-chaos-enable.yml&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The manifest installs components including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Chaos operator&lt;/li&gt;
&lt;li&gt;Chaos exporter&lt;/li&gt;
&lt;li&gt;Subscriber&lt;/li&gt;
&lt;li&gt;Event tracker&lt;/li&gt;
&lt;li&gt;Workflow controller&lt;/li&gt;
&lt;li&gt;Litmus CRDs and RBAC&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Check readiness:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;pods&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;logs&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment/subscriber&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--tail&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A useful confirmation in the subscriber log is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AgentID ... has been confirmed
Server connection established, Listening....
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a 2 CPU local node, the generated infrastructure requests may exceed available CPU. For local validation, reduce only the infrastructure deployment requests:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;resources&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment/chaos-operator-ce&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment/chaos-exporter&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment/subscriber&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment/event-tracker&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment/workflow-controller&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--requests&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;cpu&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nx"&gt;memory&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="n"&gt;Mi&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a local scheduling workaround, not a production sizing recommendation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Connect the MCP Backend
&lt;/h2&gt;

&lt;p&gt;From the server directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;cd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus-mcp-server&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;go&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;/...&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;go&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-o&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;\bin\litmuschaos-mcp-server.exe&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The server uses these environment variables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CHAOS_CENTER_ENDPOINT=http://localhost:9002
LITMUS_PROJECT_ID=your-project-uuid
LITMUS_ACCESS_TOKEN=your-token
DEFAULT_INFRA_ID=your-infrastructure-id
DEFAULT_ENVIRONMENT_ID=your-environment-id
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The project ID is the UUID in the ChaosCenter project URL, not the display name. For example, a URL like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/account/&amp;lt;account-id&amp;gt;/project/&amp;lt;project-id&amp;gt;/dashboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;contains the project ID after &lt;code&gt;/project/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Forward the GraphQL server used by the MCP implementation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;port-forward&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;svc/chaos-litmus-server-service&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;9002:9002&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The local MCP endpoint is then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://localhost:9002
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The frontend port &lt;code&gt;9091&lt;/code&gt; is for the browser UI. The MCP server should use the GraphQL port &lt;code&gt;9002&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Set Up a Local Validation Target
&lt;/h2&gt;

&lt;p&gt;Create an isolated test namespace and workload:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;create&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;namespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;chaos-test&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--dry-run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-o&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;yaml&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;apply&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-f&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;create&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;demo-nginx&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;nginx:alpine&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--replicas&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;chaos-test&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;--dry-run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-o&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;yaml&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;apply&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-f&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;rollout&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment/demo-nginx&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;chaos-test&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The expected label is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;app=demo-nginx
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The experiment should target this Deployment, not a pod name. Kubernetes will replace the pod after the fault runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create the Experiment
&lt;/h2&gt;

&lt;p&gt;In ChaosCenter, create a Pod Delete experiment from a ChaosHub template.&lt;/p&gt;

&lt;p&gt;Recommended local values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;App kind: deployment
App namespace: chaos-test
App label: app=demo-nginx
Total chaos duration: 15 seconds
Ramp time: 0
Chaos interval: 5 seconds
Pods affected: 100 percent
Default health check: false
Sequence: parallel
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the first test, either use no probe or create a Kubernetes probe that checks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Group: apps
Version: v1
Resource: deployments
Resource name: demo-nginx
Namespace: chaos-test
Operation: present
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The probe must be attached to the Pod Delete fault, not to a workflow helper step such as &lt;code&gt;run-chaos&lt;/code&gt;. Litmus stores the relationship in an annotation similar to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;probeRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;[{"name":"k8sprobe","mode":"SOT"}]'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The probe name must exactly match the saved probe name. A stale or mismatched probe reference causes errors such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Probe in fault is not attached to a proper reference
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Select and Test the Custom Agent
&lt;/h2&gt;

&lt;p&gt;In Copilot Chat, select the &lt;strong&gt;LitmusChaos Operator&lt;/strong&gt; custom agent and use Agent mode. Start with a read-only request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Check whether Litmus is ready for a chaos test, then list my chaos experiments.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful response includes the experiment name and ID. The agent should also report whether the infrastructure is active and confirmed. The server returned a successful read-only response with zero experiments before the experiment was created, which was a useful first connectivity test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run Through the Custom Agent
&lt;/h2&gt;

&lt;p&gt;The safe conversational workflow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run the pod-delete experiment through LitmusChaos.
First inspect the experiment and infrastructure.
Verify that the target is only demo-nginx in the chaos-test namespace.
Ask for confirmation before running it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent should resolve the display name to an experiment ID and then call MCP. Before execution, it should verify:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Infrastructure is active&lt;/li&gt;
&lt;li&gt;Infrastructure is confirmed&lt;/li&gt;
&lt;li&gt;The target namespace is &lt;code&gt;chaos-test&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The target selector is &lt;code&gt;app=demo-nginx&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The experiment is not pointed at &lt;code&gt;litmus&lt;/code&gt;, &lt;code&gt;kube-system&lt;/code&gt;, or production&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Monitor Kubernetes in another terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;pods&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;chaos-test&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-w&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a successful Pod Delete experiment, the target pod disappears and the Deployment creates a replacement pod.&lt;/p&gt;

&lt;p&gt;The important design decision is that the agent does not silently execute a destructive tool. It turns the request into a preflight report and confirmation gate first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Small Prompts for the Agent
&lt;/h2&gt;

&lt;p&gt;Use these prompts one at a time. Wait for the agent's result before sending the next prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prompt 1: Check Readiness
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Is Litmus ready for a local chaos test?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 2: Check the Target
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Check `Deployment/demo-nginx` in namespace `chaos-test`.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 3: List Experiments
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;List my chaos experiments.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 4: Inspect the Experiment
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Show details for `pod-delete-1`.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 5: Validate the Target
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Verify that `pod-delete-1` targets only `demo-nginx` in `chaos-test`.
Do not run it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 6: Ask for a Preflight
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Prepare the preflight summary for `pod-delete-1`.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 7: Request Execution
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run `pod-delete-1`.
Ask for confirmation first.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 8: Confirm Execution
&lt;/h3&gt;

&lt;p&gt;Only send this after reviewing the preflight summary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Yes, run it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 9: Check the Run
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What is the status of the latest run?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 10: Verify Recovery
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Verify that `demo-nginx` recovered.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 11: Investigate Failure
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Why did the latest run fail or remain queued?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 12: Stop Safely
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Check whether the experiment is still running.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it is still active, review the target and then send:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Stop the active run. Ask for confirmation first.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 13: Inspect Infrastructure
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;List only active chaos infrastructures.
Show details and the installation manifest for infrastructure &amp;lt;infra-id&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 14: Manage Environments
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;List NON_PROD environments.
Create a NON_PROD environment named local-test with description "Disposable Kubernetes test environment".
Ask for confirmation before creating it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 15: Work With Resilience Probes
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;List Kubernetes resilience probes.
Create a Kubernetes probe that checks whether Deployment/demo-nginx exists in namespace chaos-test, with a 5 second timeout and 5 second interval.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 16: Discover ChaosHubs and Faults
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;List all ChaosHubs.
Show pod-related faults available in ChaosHub &amp;lt;hub-id&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 17: Review Statistics
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Show experiment statistics including resiliency score distribution.
Give me a read-only health summary of experiments, infrastructures, environments, and ChaosHubs.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt 18: Register Infrastructure
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Register a namespace-scoped Kubernetes infrastructure named local-test-2 in environment &amp;lt;environment-id&amp;gt;.
Show me the request details and ask for confirmation before creating it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Creating an environment or infrastructure registration request changes ChaosCenter state, so the agent should confirm those actions too. Creating an experiment definition remains a ChaosCenter UI workflow in this implementation.&lt;/p&gt;

&lt;p&gt;This incremental style keeps each decision visible. The agent should never combine discovery, approval, execution, and recovery into one opaque action.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Failed During the Build
&lt;/h2&gt;

&lt;p&gt;The setup produced several useful lessons.&lt;/p&gt;

&lt;h3&gt;
  
  
  Docker Desktop Kubernetes was unstable at 2 GB
&lt;/h3&gt;

&lt;p&gt;The default Docker Desktop resource limit was 2 GB. ChaosCenter starts several services and MongoDB, so the Kubernetes API repeatedly became unavailable with TLS handshake timeouts.&lt;/p&gt;

&lt;p&gt;The practical fix was to configure WSL 2 with 4 GB and 2 CPUs, then use a reduced MongoDB configuration and smaller infrastructure requests.&lt;/p&gt;

&lt;h3&gt;
  
  
  Standalone MongoDB broke expected DNS
&lt;/h3&gt;

&lt;p&gt;A standalone MongoDB deployment removed the StatefulSet hostname expected by Litmus init containers. The compatible low-resource choice was a one-member replica set.&lt;/p&gt;

&lt;h3&gt;
  
  
  Helm retained a pending operation
&lt;/h3&gt;

&lt;p&gt;An interrupted Helm installation left the release in &lt;code&gt;pending-install&lt;/code&gt;. The clean recovery was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;helm&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;uninstall&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;chaos&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;delete&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;namespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;create&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;namespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;litmus&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then install the chart again with the corrected values.&lt;/p&gt;

&lt;h3&gt;
  
  
  Generated Argo labels were invalid
&lt;/h3&gt;

&lt;p&gt;The generated workflow used a label value containing an unresolved template:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;subject&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{workflow.parameters.appNamespace}}_pod-delete"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kubernetes rejected the literal braces because they are not valid label characters. A static value such as this is valid:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;subject&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;chaos-test_pod-delete&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a general lesson: do not assume every workflow field expands templates. Labels are validated before Argo can necessarily substitute workflow parameters.&lt;/p&gt;

&lt;h3&gt;
  
  
  The first MCP run was queued and rejected
&lt;/h3&gt;

&lt;p&gt;The MCP request itself succeeded and returned a success response, but the resulting workflow was rejected by Kubernetes because of the invalid label. That distinction matters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MCP transport: working
ChaosCenter API: reachable
Infrastructure: confirmed
Workflow validation: failed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful MCP request does not guarantee that the generated chaos workflow will execute successfully. Always inspect the run status and Kubernetes events.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep Local Validation Repeatable
&lt;/h2&gt;

&lt;p&gt;The repository includes an optional helper for restarting the local validation environment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;powershell&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-ExecutionPolicy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;Bypass&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-File&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;\scripts\start-litmus-demo.ps1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Starts Docker Desktop if necessary.&lt;/li&gt;
&lt;li&gt;Selects the &lt;code&gt;docker-desktop&lt;/code&gt; Kubernetes context.&lt;/li&gt;
&lt;li&gt;Waits for the local node.&lt;/li&gt;
&lt;li&gt;Waits for Litmus deployments.&lt;/li&gt;
&lt;li&gt;Verifies or recreates the local validation workload.&lt;/li&gt;
&lt;li&gt;Starts the &lt;code&gt;9091&lt;/code&gt; frontend port-forward.&lt;/li&gt;
&lt;li&gt;Starts the &lt;code&gt;9002&lt;/code&gt; GraphQL port-forward.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Stop the port-forwards without deleting Kubernetes resources:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;powershell&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-ExecutionPolicy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;Bypass&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-File&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;\scripts\stop-litmus-demo.ps1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes normal Windows or Docker Desktop restarts cheap. A Docker Desktop &lt;strong&gt;Reset Kubernetes cluster&lt;/strong&gt; is different: it deletes the local Kubernetes resources and requires reinstalling ChaosCenter and the infrastructure manifest.&lt;/p&gt;

&lt;h2&gt;
  
  
  References and Useful Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.litmuschaos.io/" rel="noopener noreferrer"&gt;LitmusChaos documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/litmuschaos" rel="noopener noreferrer"&gt;LitmusChaos GitHub organization&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/litmuschaos/litmus-mcp-server" rel="noopener noreferrer"&gt;LitmusChaos MCP server repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://modelcontextprotocol.io/docs/getting-started/intro" rel="noopener noreferrer"&gt;Model Context Protocol documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/copilot/customization/custom-agents" rel="noopener noreferrer"&gt;VS Code custom agents documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/copilot/customization/agent-skills" rel="noopener noreferrer"&gt;VS Code agent skills documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/copilot/chat/mcp-servers" rel="noopener noreferrer"&gt;VS Code MCP servers documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.litmuschaos.io/docs/getting-started/installation" rel="noopener noreferrer"&gt;LitmusChaos installation guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For the complete workspace implementation, the relevant files are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;.github/agents/litmuschaos.agent.md&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;.github/skills/litmus-experiment-runner/SKILL.md&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;.vscode/mcp.json&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;litmus-mcp-server/MCP_CHAT_GUIDE.md&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Attribution
&lt;/h2&gt;

&lt;p&gt;This project and tutorial integrate and build on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://litmuschaos.io/" rel="noopener noreferrer"&gt;LitmusChaos&lt;/a&gt; and the &lt;a href="https://github.com/litmuschaos" rel="noopener noreferrer"&gt;LitmusChaos project&lt;/a&gt; for chaos orchestration, experiments, ChaosCenter, and GraphQL APIs.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;Model Context Protocol&lt;/a&gt; for the tool-transport contract.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://go.dev/" rel="noopener noreferrer"&gt;Go&lt;/a&gt; for the MCP server implementation.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/" rel="noopener noreferrer"&gt;Kubernetes&lt;/a&gt; and &lt;a href="https://www.docker.com/products/docker-desktop/" rel="noopener noreferrer"&gt;Docker Desktop&lt;/a&gt; for the local execution environment.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://code.visualstudio.com/docs/copilot/customization/custom-agents" rel="noopener noreferrer"&gt;Visual Studio Code Copilot customization&lt;/a&gt; for the custom agent and skill workflow.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The MCP server integration, custom agent instructions, skill workflow, restart scripts, and tutorial content in this repository are custom project work. Please consult each upstream project's repository and license before redistributing code, documentation, or generated experiment assets. LitmusChaos, Kubernetes, Docker, GitHub Copilot, VS Code, and related product names are trademarks or project names of their respective owners.&lt;/p&gt;

&lt;p&gt;The custom agent instructions, skill workflow, restart scripts, and supporting documentation in this tutorial were created for this example. Replace local IDs, endpoints, credentials, and infrastructure details with your own values before publishing or sharing the repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security Notes
&lt;/h2&gt;

&lt;p&gt;Do not publish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;.env&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;API tokens&lt;/li&gt;
&lt;li&gt;Infrastructure access keys&lt;/li&gt;
&lt;li&gt;Downloaded infrastructure manifests containing credentials&lt;/li&gt;
&lt;li&gt;Private project or account identifiers unless intentionally redacted&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;.env&lt;/code&gt; file is ignored by Git, but a Git ignore rule does not protect a token that has already appeared in chat logs, screenshots, terminal transcripts, or public commits. Rotate any token that was exposed during testing.&lt;/p&gt;

&lt;p&gt;For a team or public blog, replace all real values with placeholders:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LITMUS_PROJECT_ID=&amp;lt;project-uuid&amp;gt;
LITMUS_ACCESS_TOKEN=&amp;lt;redacted-token&amp;gt;
DEFAULT_INFRA_ID=&amp;lt;infrastructure-uuid&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Final Takeaways
&lt;/h2&gt;

&lt;p&gt;The custom Copilot agent is the central design artifact. LitmusChaos remains the chaos platform, Kubernetes remains the execution environment, and MCP provides the API bridge. The agent supplies the operational discipline around those systems.&lt;/p&gt;

&lt;p&gt;MCP does not replace Kubernetes or ChaosCenter. It adds a conversational control layer over existing APIs and workflows, while the custom agent turns that capability into a repeatable and safer SRE and resilience-engineering workflow.&lt;/p&gt;

&lt;p&gt;The most reliable pattern is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;preflight -&amp;gt; inspect -&amp;gt; confirm -&amp;gt; execute -&amp;gt; poll -&amp;gt; verify recovery
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The preflight and confirmation steps are not ceremony. They are what prevent a natural-language request from becoming an accidental production outage.&lt;/p&gt;

&lt;p&gt;For a small local machine, resource sizing is part of the experiment design. A constrained target and a restart helper make local validation repeatable without pretending that a laptop-sized cluster is production infrastructure.&lt;/p&gt;

&lt;p&gt;The reusable pattern is broader than LitmusChaos:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;custom agent policy
  +
MCP tools
  +
domain-specific skill
  +
explicit confirmation
  =
conversational operations with guardrails
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>ai</category>
      <category>kubernetes</category>
      <category>chaos</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Building an AI Agent for Performance Engineering: MCP Servers You Can Use Today</title>
      <dc:creator>Kavin Arvind Ragavan</dc:creator>
      <pubDate>Fri, 18 Sep 2026 18:54:40 +0000</pubDate>
      <link>https://dev.to/kavin_arvind_8a1adbd39efd/mcp-servers-for-performance-engineering-2d1i</link>
      <guid>https://dev.to/kavin_arvind_8a1adbd39efd/mcp-servers-for-performance-engineering-2d1i</guid>
      <description>&lt;h2&gt;
  
  
  MCP Servers for Performance, Observability and Resilience Engineering
&lt;/h2&gt;

&lt;p&gt;Many widely used performance, observability, resilience, and browser-testing platforms now have an MCP server sitting in front of them. That means you can stop opening five different consoles to answer one question about a load test, and instead ask for the answer directly: "did checkout regress in the last run," "audit this page on mobile," "run a pod-delete and give me the resiliency score." The server does the API calls; you just describe the outcome.&lt;/p&gt;

&lt;p&gt;This post is the practical map for doing that: what each server actually lets you do as a performance engineer, the real tools behind each capability, the configuration you need to get it running, and a worked usage scenario for every one. Twelve servers, five categories — load and stress testing, observability and evidence, resilience and chaos engineering, front-end/client-side testing, and browser automation — all documented with the parameters you'll pass.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to use this post:&lt;/strong&gt; skim the category you care about, read the "usage in practice" example to see what a real request looks like, then use the tool table as your reference when you wire it into an agent. Within every category, vendor-maintained servers are listed first, community servers next — so the trust tier is visible before you even reach the details.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Catalog snapshot:&lt;/strong&gt; Tool counts, transports, configuration options, and project status were checked against the referenced repositories around September 2026. MCP servers evolve quickly, so treat these counts as a point-in-time snapshot rather than permanent API contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The shape every one of these servers follows
&lt;/h2&gt;

&lt;p&gt;Before the catalog, the pattern worth internalizing: most MCP servers here are thin, typed adapters over an existing platform's API or CLI. Nothing more.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌────────────────────┐        ┌───────────────────────┐        ┌───────────────────────┐
│   MCP Client         │        │   MCP Server            │        │   Underlying Platform  │
│  (Copilot, Claude,    │──────▶│   (this catalog)         │──────▶│  (LoadRunner, Dynatrace,│
│   Cursor, ...)        │◀──────│   typed tools, schemas   │◀──────│   Splunk, JMeter, ...)  │
└────────────────────┘        └───────────────────────┘        └───────────────────────┘
        tools/list                  auth: token / key /              REST, GraphQL, DQL,
        tools/call                  OAuth / CLI child-proc            CLI subprocess, etc.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What varies between servers is the transport (stdio vs. SSE vs. HTTP), the auth model (static keys vs. OAuth vs. browser-based token exchange), and — the part worth reading carefully — what a tool call can actually &lt;em&gt;do&lt;/em&gt;: read-only investigation, or an action with a real-world side effect (spend money, send a message, start a test, delete data).&lt;/p&gt;

&lt;p&gt;The practical flow looks like this: the MCP server exposes capabilities, the agent decides how to use them, and policy controls actions that need approval.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
        A[Engineer request] --&amp;gt; B[Copilot agent]
        B --&amp;gt; C[Discover MCP tools]
        C --&amp;gt; D[Plan tool sequence]
        D --&amp;gt; E{Approval needed?}
        E --&amp;gt;|Yes| F[Human approval]
        E --&amp;gt;|No| G[Call MCP server]
        F --&amp;gt; G
        G --&amp;gt; H[Underlying platform]
        H --&amp;gt; I[Evidence and result]
        I --&amp;gt; B&lt;/code&gt;&lt;/pre&gt;



&lt;h2&gt;
  
  
  2. Trust tiers, at a glance
&lt;/h2&gt;

&lt;p&gt;Not every server here deserves the same confidence. I use two tiers throughout this reference, and within every category below, servers are listed &lt;strong&gt;vendor-maintained first, community next&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;th&gt;Servers in this catalog&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;🟢 Vendor-maintained&lt;/td&gt;
&lt;td&gt;Maintained by the platform vendor or product team&lt;/td&gt;
&lt;td&gt;k6, BlazeMeter, LitmusChaos, Chrome DevTools MCP, Playwright MCP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;🔵 Community, verified&lt;/td&gt;
&lt;td&gt;Third-party, source read and confirmed&lt;/td&gt;
&lt;td&gt;LoadRunner Cloud, Apache JMeter, Artillery, Dynatrace, Splunk, Lighthouse, PageSpeed Insights&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Five of the twelve servers are vendor-maintained projects — more than I expected when I started this catalog. The remaining seven are community-verified projects, with their source and documented tool surfaces reviewed for this catalog.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgcpjjw7e6j0i0a3bp9m6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgcpjjw7e6j0i0a3bp9m6.png" alt="Overview of the five performance-engineering categories and the MCP servers listed in each" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbqstotlm1l3f7v3tyfyv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbqstotlm1l3f7v3tyfyv.png" alt="Where each agent helps across the performance-engineering lifecycle" width="800" height="354"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Load &amp;amp; stress testing
&lt;/h2&gt;

&lt;h3&gt;
  
  
  3.1 k6 MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/grafana/mcp-k6" rel="noopener noreferrer"&gt;&lt;code&gt;grafana/mcp-k6&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; Go · &lt;strong&gt;Status:&lt;/strong&gt; 🟢 vendor-maintained (Grafana Labs) — the maintainers mark it experimental&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; stdio by default, optional Streamable HTTP (&lt;code&gt;-transport=http&lt;/code&gt;) — ships as a Docker image, Homebrew formula, Debian/RPM packages, a native Go binary, or an &lt;code&gt;xk6&lt;/code&gt; subcommand&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; none required for local use beyond having &lt;code&gt;k6&lt;/code&gt; on &lt;code&gt;PATH&lt;/code&gt; (or using the Docker image, which bundles it); HTTP mode adds &lt;code&gt;-addr&lt;/code&gt;, &lt;code&gt;-endpoint&lt;/code&gt;, &lt;code&gt;-stateless&lt;/code&gt;, &lt;code&gt;-preload&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth:&lt;/strong&gt; none built in — the README is explicit that a remote deployment needs a trusted network or a proxy in front of it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;k6 scripts are JavaScript, and JavaScript is forgiving about compiling into something that silently does the wrong thing. Plenty of k6 users have shipped a "test" that ran zero real iterations because of a scenario-config typo, and only found out from a suspiciously fast pass.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;validate_script&lt;/code&gt; catches structural mistakes before you burn a real run on them — a minimal dry run at 1 VU, 1 iteration&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;generate_script&lt;/code&gt; prompt template drafts a starting script from a plain-English description, grounded in the actual k6 docs rather than a stale training snapshot&lt;/li&gt;
&lt;li&gt;The documentation tools (&lt;code&gt;list_sections&lt;/code&gt;, &lt;code&gt;get_documentation&lt;/code&gt;) let the agent look up the current k6 API instead of guessing at option names&lt;/li&gt;
&lt;li&gt;Works the same whether you run it locally via Docker or point your whole team at one shared HTTP instance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; "generate a k6 script that load-tests our login API with authentication, then validate it" pulls the current k6 docs for scenarios and thresholds, drafts a script, and immediately runs &lt;code&gt;validate_script&lt;/code&gt; against it — catching a bad &lt;code&gt;stages&lt;/code&gt; array before you ever spend a real run on it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Parameters&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;validate_script&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;script&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Dry-runs the script (1 VU, 1 iteration) and returns pass/fail plus stdout/stderr&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;run_script&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;script&lt;/code&gt;, &lt;code&gt;vus?&lt;/code&gt;, &lt;code&gt;duration?&lt;/code&gt; (max 5m), &lt;code&gt;iterations?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Executes the test locally and returns metrics and a summary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_sections&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;version?&lt;/code&gt;, &lt;code&gt;category?&lt;/code&gt;, &lt;code&gt;depth?&lt;/code&gt; (default 1, max 5), &lt;code&gt;root_slug?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Browses the k6 docs tree without loading all of it into context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_documentation&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;slug&lt;/code&gt;, &lt;code&gt;version?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Retrieves the full markdown for one docs section&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;info&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(none)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;Reports k6 and server environment information&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;search_terraform&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;query&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Searches the k6 Terraform documentation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The current server exposes six tools. It also provides a &lt;code&gt;generate_script&lt;/code&gt; prompt template (resource URI &lt;code&gt;prompts://k6/generate_script&lt;/code&gt;) that chains research, best practices, and validation into one guided flow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; going from "describe the test in English" to a validated k6 script without leaving the conversation — and it's the one server here where the maintainers are upfront that it's still experimental, so budget for rough edges.&lt;/p&gt;




&lt;h3&gt;
  
  
  3.2 BlazeMeter MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/Blazemeter/bzm-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;Blazemeter/bzm-mcp&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; Python 3.11+ · &lt;strong&gt;Status:&lt;/strong&gt; 🟢 &lt;strong&gt;vendor-maintained&lt;/strong&gt; (published by BlazeMeter/Perforce)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime options:&lt;/strong&gt; pre-built binary, &lt;code&gt;uvx&lt;/code&gt; (from git), or Docker (&lt;code&gt;ghcr.io/blazemeter/bzm-mcp&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;API_KEY_ID&lt;/code&gt; + &lt;code&gt;API_KEY_SECRET&lt;/code&gt; (or a &lt;code&gt;BLAZEMETER_API_KEY&lt;/code&gt; JSON file), &lt;code&gt;SOURCE_WORKING_DIRECTORY&lt;/code&gt; for Docker mounts, optional &lt;code&gt;SSL_CERT_FILE&lt;/code&gt; for corporate CA bundles&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's the one entry on this list you can point at a compliance review without an argument — a vendor-maintained server, from the company that makes the product.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenTelemetry ships out of the box, so you get observability into the &lt;em&gt;agent's&lt;/em&gt; behavior for free, not just the load test's&lt;/li&gt;
&lt;li&gt;Three install paths (binary, &lt;code&gt;uvx&lt;/code&gt;, Docker) mean it fits whatever your team already standardized on&lt;/li&gt;
&lt;li&gt;Cloud-scale execution without you having to run or scale your own load generators&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; a quarterly capacity test that used to mean someone manually clicking through the BlazeMeter console becomes "launch the payments load test in the cloud and tell me how it went" — the agent starts the run, waits, and reports back with the summary, while OpenTelemetry quietly records how long each of those calls actually took.&lt;/p&gt;

&lt;p&gt;This is the one vendor-maintained server in the load-testing category. Its tool surface covers the full cloud workflow — creating and managing load-test workflows, executing them, and retrieving reports — but BlazeMeter documents the exact tool list in their own &lt;a href="https://help.blazemeter.com/docs/guide/integrations-blazemeter-mcp-server.html" rel="noopener noreferrer"&gt;MCP Server guide&lt;/a&gt; rather than enumerating every tool in the README, so treat the specific tool names as vendor-documented rather than independently verified here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Observability bonus:&lt;/strong&gt; it ships OpenTelemetry instrumentation out of the box — every tool call produces a trace (tool name, action, client name/version, session ID) and two metrics (&lt;code&gt;mcp.tool.calls&lt;/code&gt;, &lt;code&gt;mcp.tool.duration&lt;/code&gt;). Telemetry defaults to BlazeMeter's own collector; you can redirect it (&lt;code&gt;OTEL_EXPORTER_OTLP_ENDPOINT&lt;/code&gt;) or disable it (&lt;code&gt;OTEL_SDK_DISABLED=true&lt;/code&gt; or &lt;code&gt;--no-telemetry&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; cloud-scale execution when you're already a BlazeMeter customer and want the vendor's own supported path.&lt;/p&gt;




&lt;h3&gt;
  
  
  3.3 LoadRunner Cloud MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/pbandreddy/loadrunner-cloud-mcp-server" rel="noopener noreferrer"&gt;&lt;code&gt;pbandreddy/loadrunner-cloud-mcp-server&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; JavaScript (ESM) · &lt;strong&gt;Status:&lt;/strong&gt; 🔵 community&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Node.js 18+ (20+ recommended) · &lt;code&gt;@modelcontextprotocol/sdk&lt;/code&gt; 1.9.0 · stdio by default, optional SSE (&lt;code&gt;--sse&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;LRC_BASE_URL&lt;/code&gt;, &lt;code&gt;LRC_TENANT_ID&lt;/code&gt;, &lt;code&gt;LRC_CLIENT_ID&lt;/code&gt;, &lt;code&gt;LRC_CLIENT_SECRET&lt;/code&gt;, optional &lt;code&gt;PORT&lt;/code&gt; (SSE mode)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth:&lt;/strong&gt; client credentials exchanged for a bearer token automatically before every call&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Engineers finish a spike test and then lose the next quarter hour clicking through the LoadRunner Cloud UI to find the run, open the transactions tab, and eyeball whether p95 crossed the line. This server exists to compress that click-through into a question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No more hunting for a run ID by hand — the tool chain resolves project → test → run for you&lt;/li&gt;
&lt;li&gt;Percentile math (p90/p95) comes back built into the response, not something you compute from a raw CSV export&lt;/li&gt;
&lt;li&gt;Read-only by construction, so pointing an agent at it can't accidentally trigger a new test run&lt;/li&gt;
&lt;li&gt;One call (&lt;code&gt;test_runs_getHttpResponses&lt;/code&gt;) gets you straight to the failure evidence instead of a support ticket&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; picture this — your spike test on checkout just finished. Instead of opening a browser, you ask "did checkout regress in the last run?" The agent resolves the project with &lt;code&gt;get_projects&lt;/code&gt;, walks to the latest run through &lt;code&gt;projects_getLoadTestRuns&lt;/code&gt;, then pulls &lt;code&gt;test_runs_getTestRunTransactions&lt;/code&gt; for the percentile table and &lt;code&gt;test_runs_getHttpResponses&lt;/code&gt; if anything looks off. Thirty seconds later you have an answer instead of a browser tab.&lt;/p&gt;

&lt;p&gt;This server is &lt;strong&gt;read-only&lt;/strong&gt; — it's built for investigating existing LoadRunner Cloud projects and runs, not for launching new ones. All nine tools require &lt;code&gt;TENANTID&lt;/code&gt; under the hood; you never pass it yourself.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Parameters&lt;/th&gt;
&lt;th&gt;What it returns&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_projects&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(none)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;All projects in the tenant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;projects_getLoadTests&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;projectId&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Load tests for a project&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;projects_getLoadTestScripts&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;projectId&lt;/code&gt;, &lt;code&gt;loadTestId&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Scripts attached to a load test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;projects_getLoadTestRuns&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;projectId&lt;/code&gt;, &lt;code&gt;loadTestId&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Runs for a load test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_active_test_runs&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;status?&lt;/code&gt;, &lt;code&gt;projectIds?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Currently active runs, filterable by status (&lt;code&gt;RUNNING&lt;/code&gt;, &lt;code&gt;INITIALIZING&lt;/code&gt;, &lt;code&gt;CHECKING_STATUS&lt;/code&gt;, &lt;code&gt;STOPPING&lt;/code&gt;, &lt;code&gt;DELAYED&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;test_runs_getRecentTestRuns&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;projectIds?&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;License usage for runs in the last 30 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;test_runs_getTestRunResults&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;runId&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Overall result/status for a run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;test_runs_getTestRunTransactions&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;runId&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Transaction data — &lt;strong&gt;the call always requests the 90th and 95th percentile&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;test_runs_getHttpResponses&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;runId&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;HTTP response detail for a run&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; "what's the p95 on the latest checkout run" style investigation, without opening the LRC UI.&lt;/p&gt;




&lt;h3&gt;
  
  
  3.4 Apache JMeter MCP ("JMeter Architect")
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/aravindksk7/Jmeter-MCP" rel="noopener noreferrer"&gt;&lt;code&gt;aravindksk7/Jmeter-MCP&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; TypeScript → &lt;code&gt;dist/index.js&lt;/code&gt; · &lt;strong&gt;Status:&lt;/strong&gt; 🔵 community&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Node.js 18+ · Apache JMeter itself only required for the run tool, on &lt;code&gt;PATH&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; invoked as &lt;code&gt;node dist/index.js&lt;/code&gt;; JMeter's &lt;code&gt;bin&lt;/code&gt; directory must be on &lt;code&gt;PATH&lt;/code&gt; for &lt;code&gt;jmeter_run_test&lt;/code&gt; to work&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ask any engineer who's used JMeter's desktop GUI to add a Header Manager, and you'll get a specific kind of sigh. This server routes around the GUI entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No hand-edited XML — the tools assemble a structurally valid &lt;code&gt;.jmx&lt;/code&gt; for you, element by element&lt;/li&gt;
&lt;li&gt;Executes in non-GUI mode from the start, so the same plan you build in chat is the one that runs in CI&lt;/li&gt;
&lt;li&gt;Assertions and listeners are added as explicit tool calls, so nothing gets silently skipped the way it can in a GUI where a checkbox is easy to miss&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; "build a 100-user checkout test with a 200-status assertion and run it" turns into a real sequence: &lt;code&gt;jmeter_init_plan&lt;/code&gt;, then a thread group at 100 users, a sampler for the checkout endpoint, an assertion on the response code, a listener for the aggregate report, and finally &lt;code&gt;jmeter_run_test&lt;/code&gt;. You get a working &lt;code&gt;.jmx&lt;/code&gt; and a completed run without opening the JMeter desktop app once.&lt;/p&gt;

&lt;p&gt;This one doesn't call an existing JMeter installation to &lt;em&gt;build&lt;/em&gt; a plan — it constructs a real &lt;code&gt;.jmx&lt;/code&gt; file, element by element, then hands it to JMeter to execute in non-GUI mode.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Required parameters&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jmeter_init_plan&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;filename&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Creates a fresh, empty &lt;code&gt;.jmx&lt;/code&gt; test plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jmeter_add_thread_group&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;num_threads&lt;/code&gt;, &lt;code&gt;ramp_time&lt;/code&gt;, &lt;code&gt;loops&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Adds virtual users (&lt;code&gt;loops: -1&lt;/code&gt; = infinite)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jmeter_add_sampler&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;domain&lt;/code&gt;, &lt;code&gt;path&lt;/code&gt;, &lt;code&gt;method&lt;/code&gt;, &lt;code&gt;parameters&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Adds an HTTP request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jmeter_add_header&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;headers&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Adds an HTTP Header Manager&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jmeter_add_listener&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;listener_type&lt;/code&gt; (&lt;code&gt;summary&lt;/code&gt; \&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;aggregate&lt;/code&gt; \&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jmeter_add_timer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;delay_ms&lt;/code&gt;, &lt;code&gt;random_delay_ms?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Adds think time between requests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jmeter_add_assertion&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;test_field&lt;/code&gt; (&lt;code&gt;response_data&lt;/code&gt; \&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;response_code&lt;/code&gt; \&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jmeter_run_test&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;output_file?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Executes the plan in non-GUI mode&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; building a correct &lt;code&gt;.jmx&lt;/code&gt; from a sentence instead of hand-editing XML — genuinely useful if you've ever fought JMeter's GUI to add a Header Manager.&lt;/p&gt;




&lt;h3&gt;
  
  
  3.5 Artillery MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/jch1887/artillery-mcp-server" rel="noopener noreferrer"&gt;&lt;code&gt;jch1887/artillery-mcp-server&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; TypeScript · &lt;strong&gt;Status:&lt;/strong&gt; 🔵 community&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Node.js 22.18+ · requires the Artillery CLI on &lt;code&gt;PATH&lt;/code&gt; (or &lt;code&gt;ARTILLERY_BIN&lt;/code&gt;) · stdio&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;ARTILLERY_WORKDIR&lt;/code&gt;, &lt;code&gt;ARTILLERY_BIN&lt;/code&gt;, &lt;code&gt;ARTILLERY_TIMEOUT_MS&lt;/code&gt; (default 1,800,000 ms / 30 min), &lt;code&gt;ARTILLERY_MAX_OUTPUT_MB&lt;/code&gt; (default 10), &lt;code&gt;ARTILLERY_ALLOW_QUICK&lt;/code&gt; (default true), &lt;code&gt;DEBUG&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sandbox:&lt;/strong&gt; every path is resolved inside &lt;code&gt;ARTILLERY_WORKDIR&lt;/code&gt;; the child process environment is an explicit allow-list, not inherited wholesale&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Artillery is one of the fastest load tools to spin up, but its CLI output is a wall of JSON you re-parse by eye after every run, and the YAML configs tend to live wherever the last person who wrote one happened to save them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Saved configs mean "the smoke test for the payments API" lives in one named place instead of six local copies&lt;/li&gt;
&lt;li&gt;Built-in regression thresholds turn "looks about the same to me" into an actual pass/fail&lt;/li&gt;
&lt;li&gt;Sandboxing means you can hand this to a teammate — or an agent — without worrying what paths or env vars it can touch&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;quick_test&lt;/code&gt; skips the YAML entirely when you just need to hit an endpoint a few times right now&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; before every deploy, someone on the team runs the same smoke test by hand. Wired up here, that becomes "run the api-smoke baseline and compare it to last week's." The agent replays the saved config, parses the JSON results into percentiles and error counts, and tells you plainly whether anything regressed — no spreadsheet required.&lt;/p&gt;

&lt;p&gt;Eleven tools, split across running tests and managing saved configurations:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Parameters&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;run_test_from_file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;path&lt;/code&gt;, &lt;code&gt;outputJson?&lt;/code&gt;, &lt;code&gt;reportHtml?&lt;/code&gt;, &lt;code&gt;env?&lt;/code&gt;, &lt;code&gt;cwd?&lt;/code&gt;, &lt;code&gt;validateOnly?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Runs a config file already inside the workdir&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;run_test_inline&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;configText&lt;/code&gt;, &lt;code&gt;outputJson?&lt;/code&gt;, &lt;code&gt;reportHtml?&lt;/code&gt;, &lt;code&gt;env?&lt;/code&gt;, &lt;code&gt;cwd?&lt;/code&gt;, &lt;code&gt;validateOnly?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Config as a YAML/JSON string — Artillery 2.0+ requires &lt;code&gt;flow:&lt;/code&gt; instead of &lt;code&gt;requests:&lt;/code&gt; in scenarios&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;quick_test&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;target&lt;/code&gt;, &lt;code&gt;rate?&lt;/code&gt;, &lt;code&gt;duration?&lt;/code&gt;, &lt;code&gt;count?&lt;/code&gt;, &lt;code&gt;method?&lt;/code&gt;, &lt;code&gt;headers?&lt;/code&gt;, &lt;code&gt;body?&lt;/code&gt;, &lt;code&gt;insecure?&lt;/code&gt;, &lt;code&gt;keepResults?&lt;/code&gt;, &lt;code&gt;outputJson?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;No config file needed; generates a one-request scenario&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;run_saved_config&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;name&lt;/code&gt;, &lt;code&gt;outputJson?&lt;/code&gt;, &lt;code&gt;reportHtml?&lt;/code&gt;, &lt;code&gt;env?&lt;/code&gt;, &lt;code&gt;validateOnly?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Runs a previously saved config by name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;save_config&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;name&lt;/code&gt;, &lt;code&gt;content&lt;/code&gt;, &lt;code&gt;description?&lt;/code&gt;, &lt;code&gt;tags?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Stores under &lt;code&gt;$ARTILLERY_WORKDIR/saved-configs/&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_configs&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;tag?&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Lists saved configs, optionally filtered by tag&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_config&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Retrieves a saved config's content&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;delete_config&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Deletes a saved config&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;parse_results&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;jsonPath&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Summarizes a results file: RPS, latency percentiles (p50/p95/p99), HTTP codes, error counts, vuser stats&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_results&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;limit?&lt;/code&gt; (default 100)&lt;/td&gt;
&lt;td&gt;Lists result files under the workdir, newest first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;list_capabilities&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(none)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;Reports Artillery version, server version, transports, and configured limits&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Note on HTML reports:&lt;/strong&gt; recent Artillery releases removed the &lt;code&gt;report&lt;/code&gt; command; if &lt;code&gt;reportHtml&lt;/code&gt; is set and no file appears, the server falls back to returning the JSON path with a warning rather than failing outright.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; quick load checks and baseline-vs-current regression comparisons, driven entirely from chat.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Observability &amp;amp; evidence
&lt;/h2&gt;

&lt;h3&gt;
  
  
  4.1 Dynatrace MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/dynatrace-oss/dynatrace-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;dynatrace-oss/dynatrace-mcp&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; TypeScript · &lt;strong&gt;License:&lt;/strong&gt; MIT&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Status:&lt;/strong&gt; 🔵 &lt;strong&gt;community, verified&lt;/strong&gt; — source read and tool contract reviewed for this catalog.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Node.js 24+ · npm package &lt;code&gt;@dynatrace-oss/dynatrace-mcp-server&lt;/code&gt; · stdio by default, or &lt;code&gt;--http&lt;/code&gt; for an HTTP/bearer-token mode &lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;DT_ENVIRONMENT&lt;/code&gt; (required — the &lt;em&gt;Platform&lt;/em&gt; URL, &lt;code&gt;…apps.dynatrace.com&lt;/code&gt;, not the classic &lt;code&gt;…live.dynatrace.com&lt;/code&gt;), &lt;code&gt;DT_PLATFORM_TOKEN&lt;/code&gt; or &lt;code&gt;OAUTH_CLIENT_ID&lt;/code&gt;/&lt;code&gt;OAUTH_CLIENT_SECRET&lt;/code&gt; (optional — otherwise browser OAuth + OS keychain), &lt;code&gt;DT_GRAIL_QUERY_BUDGET_GB&lt;/code&gt; (default 1000), &lt;code&gt;DT_SSO_URL&lt;/code&gt; (optional override)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;DQL is powerful and genuinely has a learning curve; the natural-language generate/verify/explain tools exist specifically so you don't have to memorize it to get an answer out of Grail.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;generate_dql_from_natural_language&lt;/code&gt; + &lt;code&gt;verify_dql&lt;/code&gt; means you get a query you can actually read before it runs against your data&lt;/li&gt;
&lt;li&gt;The Grail budget (&lt;code&gt;DT_GRAIL_QUERY_BUDGET_GB&lt;/code&gt;) turns "oops, that scanned way more than I meant" from a bill into a warning&lt;/li&gt;
&lt;li&gt;Davis AI (&lt;code&gt;chat_with_davis_copilot&lt;/code&gt;) is there for the "what does this actually mean" follow-up question a raw query result can't answer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; "during the checkout load test window, latency spiked — what happened?" resolves the service, generates and verifies a DQL query scoped to just that window, and cross-references &lt;code&gt;list_problems&lt;/code&gt; and &lt;code&gt;list_exceptions&lt;/code&gt; for the same period — landing on a concrete, correlated answer instead of a wall of log lines.&lt;/p&gt;

&lt;p&gt;Capabilities are grouped by function (the README doesn't present them as a flat numbered list, but these are every tool named in the source docs):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Group&lt;/th&gt;
&lt;th&gt;Tools&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observability &amp;amp; problems&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;list_problems&lt;/code&gt; · &lt;code&gt;list_vulnerabilities&lt;/code&gt; · &lt;code&gt;list_exceptions&lt;/code&gt; · &lt;code&gt;get_kubernetes_events&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Grail queries&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;execute_dql&lt;/code&gt; · &lt;code&gt;verify_dql&lt;/code&gt; · &lt;code&gt;generate_dql_from_natural_language&lt;/code&gt; · &lt;code&gt;explain_dql_in_natural_language&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Entity discovery&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;find_entity_by_name&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Davis AI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;chat_with_davis_copilot&lt;/code&gt; · &lt;code&gt;list_davis_analyzers&lt;/code&gt; · &lt;code&gt;execute_davis_analyzer&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Automation &amp;amp; sharing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;send_slack_message&lt;/code&gt; · &lt;code&gt;send_email&lt;/code&gt; · &lt;code&gt;send_event&lt;/code&gt; · &lt;code&gt;create_dynatrace_notebook&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Cost matters here:&lt;/strong&gt; &lt;code&gt;execute_dql&lt;/code&gt; scans Grail storage, billed by volume scanned. The server tracks usage against &lt;code&gt;DT_GRAIL_QUERY_BUDGET_GB&lt;/code&gt; per session and warns at 80% of budget. You can audit actual consumption with a DQL query against &lt;code&gt;dt.system.events&lt;/code&gt; filtered to &lt;code&gt;client.client_context&lt;/code&gt; containing &lt;code&gt;"dynatrace-mcp"&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Required OAuth scopes&lt;/strong&gt; vary by tool — at minimum &lt;code&gt;app-engine:apps:run&lt;/code&gt; for nearly everything, plus the specific &lt;code&gt;storage:*:read&lt;/code&gt; scope for whatever Grail data type you're querying (logs, metrics, spans, entities, events, etc.), &lt;code&gt;davis-copilot:*:execute&lt;/code&gt; for the AI features, and &lt;code&gt;email:emails:send&lt;/code&gt; / &lt;code&gt;document:documents:*&lt;/code&gt; for the sharing tools.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; correlating a load-test window with production problems and exceptions via DQL — the worked example in the companion post walks through exactly this.&lt;/p&gt;




&lt;h3&gt;
  
  
  4.2 Splunk MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/livehybrid/splunk-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;livehybrid/splunk-mcp&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; Python&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Status:&lt;/strong&gt; 🔵 &lt;strong&gt;community, verified&lt;/strong&gt; — source read and tool contract reviewed for this catalog.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Python, built on &lt;strong&gt;FastMCP&lt;/strong&gt; · three operating modes: SSE (default), REST API (&lt;code&gt;python splunk_mcp.py api&lt;/code&gt;), or stdio (&lt;code&gt;python splunk_mcp.py stdio&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;SPLUNK_HOST&lt;/code&gt;, &lt;code&gt;SPLUNK_PORT&lt;/code&gt; (default 8089), &lt;code&gt;SPLUNK_TOKEN&lt;/code&gt; (or &lt;code&gt;SPLUNK_USERNAME&lt;/code&gt; + &lt;code&gt;SPLUNK_PASSWORD&lt;/code&gt;), &lt;code&gt;SPLUNK_SCHEME&lt;/code&gt; (default https), &lt;code&gt;VERIFY_SSL&lt;/code&gt; (default true), &lt;code&gt;FASTMCP_LOG_LEVEL&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Understanding the server's shape is exactly what lets you evaluate a Splunk MCP integration with open eyes instead of starting from zero.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Consistent error handling across every tool means a failed search or a permissions issue comes back as something you can actually act on, not a stack trace&lt;/li&gt;
&lt;li&gt;The KV Store tools double as a lightweight state store for automation, a nice trick if you're already scripting around Splunk&lt;/li&gt;
&lt;li&gt;Three transport modes (SSE, REST API, stdio) mean it can fit into however your team's tooling already talks to services&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; during an incident, "search the last 30 minutes for checkout errors and summarize the affected sourcetypes" runs &lt;code&gt;search_splunk&lt;/code&gt; scoped to that window, cross-references &lt;code&gt;indexes_and_sourcetypes&lt;/code&gt; to explain where the noise is coming from, and gives you a triage-ready summary instead of a raw search-results table.&lt;/p&gt;

&lt;p&gt;Thirteen tools across five functional groups:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Group&lt;/th&gt;
&lt;th&gt;Tools&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Meta&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;list_tools&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Health&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;health_check&lt;/code&gt; (lists reachable Splunk apps) · &lt;code&gt;ping&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Users&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;current_user&lt;/code&gt; · &lt;code&gt;list_users&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Indexes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;list_indexes&lt;/code&gt; · &lt;code&gt;get_index_info&lt;/code&gt; (params: &lt;code&gt;index_name&lt;/code&gt;) · &lt;code&gt;indexes_and_sourcetypes&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Search&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;search_splunk&lt;/code&gt; (params: &lt;code&gt;search_query&lt;/code&gt;, &lt;code&gt;earliest_time?&lt;/code&gt;, &lt;code&gt;latest_time?&lt;/code&gt;, &lt;code&gt;max_results?&lt;/code&gt;) · &lt;code&gt;list_saved_searches&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;KV Store&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;list_kvstore_collections&lt;/code&gt; · &lt;code&gt;create_kvstore_collection&lt;/code&gt; (params: &lt;code&gt;collection_name&lt;/code&gt;) · &lt;code&gt;delete_kvstore_collection&lt;/code&gt; (params: &lt;code&gt;collection_name&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Error handling is consistent across the tool set: invalid searches, permission failures, missing resources, and bad input all return a structured error message rather than a bare exception.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; understanding the pattern (search + index introspection + KV store) if you're evaluating whether to build against Splunk's now-official server instead.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Resilience &amp;amp; chaos
&lt;/h2&gt;

&lt;h3&gt;
  
  
  5.1 LitmusChaos MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/litmuschaos/litmus-mcp-server" rel="noopener noreferrer"&gt;&lt;code&gt;litmuschaos/litmus-mcp-server&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; Go · &lt;strong&gt;Status:&lt;/strong&gt; 🟢 &lt;strong&gt;vendor-maintained&lt;/strong&gt; (LitmusChaos project)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; connects to a running LitmusChaos ChaosCenter (3.x) over its GraphQL API&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;CHAOS_CENTER_ENDPOINT&lt;/code&gt;, &lt;code&gt;LITMUS_PROJECT_ID&lt;/code&gt;, &lt;code&gt;LITMUS_ACCESS_TOKEN&lt;/code&gt;, optional &lt;code&gt;DEFAULT_INFRA_ID&lt;/code&gt;, optional &lt;code&gt;DEFAULT_ENVIRONMENT_ID&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Chaos engineering has a well-earned reputation as YAML archaeology — CRDs, manifests, and infrastructure wiring before you even get to break anything on purpose. This collapses all of that into a sentence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sixteen tools cover the supported lifecycle, from discovering faults in a ChaosHub to registering infrastructure to reading back a resiliency score — you're not stitching together &lt;code&gt;kubectl&lt;/code&gt; commands by hand&lt;/li&gt;
&lt;li&gt;Resiliency scoring gives you a number you can actually track release over release, not just a pass/fail&lt;/li&gt;
&lt;li&gt;It plugs into existing ChaosHub content, so you're not authoring fault definitions from scratch every time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; before a release, "run a pod-delete on checkout-service in staging for 30 seconds and tell me the resiliency score" starts the experiment, polls its status, checks the probes attached to it, and reports back a number and a verdict — the entire chaos-engineering loop, in one exchange.&lt;/p&gt;

&lt;p&gt;The LitmusChaos server exposes &lt;strong&gt;16 tools&lt;/strong&gt; grouped into six functional areas:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Group&lt;/th&gt;
&lt;th&gt;Tools (by function)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Chaos experiments&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;list experiments · get experiment details · run an experiment · stop an experiment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Execution monitoring&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;list runs · get run details&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Infrastructure&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;list registered infrastructure · get infrastructure details · register new infrastructure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Environments&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;create an environment · list environments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Resilience probes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;list probes · create supported probes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ChaosHub &amp;amp; statistics&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;list ChaosHubs · get fault details · review experiment statistics&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Target infrastructure is Kubernetes/OpenShift; target applications are typically microservices, databases, APIs, and messaging systems. Repository layout for the server itself: &lt;code&gt;main.go&lt;/code&gt; (server setup), &lt;code&gt;handlers.go&lt;/code&gt; (tool handlers), &lt;code&gt;go.mod&lt;/code&gt;, a &lt;code&gt;Dockerfile&lt;/code&gt;, and a &lt;code&gt;Makefile&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; "run a pod-delete on the checkout service for 30 seconds and tell me the resiliency score" — the whole point of chaos engineering, minus the YAML and CRDs.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Front-end / client-side testing
&lt;/h2&gt;

&lt;h3&gt;
  
  
  6.1 Chrome DevTools MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/ChromeDevTools/chrome-devtools-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;ChromeDevTools/chrome-devtools-mcp&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; TypeScript · &lt;strong&gt;Status:&lt;/strong&gt; 🟢 &lt;strong&gt;vendor-maintained&lt;/strong&gt; (the Google Chrome DevTools team)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Node.js LTS, current stable Chrome · install via &lt;code&gt;npx -y chrome-devtools-mcp@latest&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;--headless&lt;/code&gt;, &lt;code&gt;--isolated&lt;/code&gt;, &lt;code&gt;--slim&lt;/code&gt; (basic-only tool set), &lt;code&gt;--no-performance-crux&lt;/code&gt; (disable CrUX lookups), &lt;code&gt;--no-usage-statistics&lt;/code&gt; or &lt;code&gt;CHROME_DEVTOOLS_MCP_NO_USAGE_STATISTICS&lt;/code&gt; (opt out of telemetry), &lt;code&gt;CHROME_DEVTOOLS_MCP_NO_UPDATE_CHECKS&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth:&lt;/strong&gt; none — it drives a local (or connected) Chrome instance directly via Puppeteer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the biggest tool surface in the whole catalog — 58 tools at the time of writing — because it isn't just a performance tool, it's the entire DevTools panel exposed to an agent: Performance, Network, Elements, Memory profiling, even PWA install/launch and browser extension management.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Real Chrome performance traces (&lt;code&gt;performance_start_trace&lt;/code&gt; / &lt;code&gt;performance_stop_trace&lt;/code&gt; / &lt;code&gt;performance_analyze_insight&lt;/code&gt;) give you actual Core Web Vitals data, not an estimate&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;lighthouse_audit&lt;/code&gt; covers accessibility, SEO, best practices, and "agentic browsing" in the same session as your performance trace — one browser, one context, several kinds of evidence&lt;/li&gt;
&lt;li&gt;The accessibility-tree snapshot (&lt;code&gt;take_snapshot&lt;/code&gt;) lets an agent click, fill, and navigate a real page reliably, without brittle pixel coordinates&lt;/li&gt;
&lt;li&gt;Ships from the same team that builds DevTools itself, so it tracks Chrome's actual internals rather than reverse-engineering them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; "check the performance of the checkout page and tell me what's hurting LCP" navigates to the page, starts a trace with &lt;code&gt;performance_start_trace&lt;/code&gt; (reload enabled), stops it, and calls &lt;code&gt;performance_analyze_insight&lt;/code&gt; for the specific insight — usually something concrete like render-blocking CSS or an oversized hero image, backed by the same trace data you'd get by hand in DevTools.&lt;/p&gt;

&lt;p&gt;The full reference lists all 58 tools; the groups most relevant to performance and front-end testing:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Group&lt;/th&gt;
&lt;th&gt;Tool count&lt;/th&gt;
&lt;th&gt;Examples&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;performance_start_trace&lt;/code&gt;, &lt;code&gt;performance_stop_trace&lt;/code&gt;, &lt;code&gt;performance_analyze_insight&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Network&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;list_network_requests&lt;/code&gt;, &lt;code&gt;get_network_request&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Debugging&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;lighthouse_audit&lt;/code&gt;, &lt;code&gt;take_snapshot&lt;/code&gt;, &lt;code&gt;take_screenshot&lt;/code&gt;, &lt;code&gt;evaluate_script&lt;/code&gt;, &lt;code&gt;get_css_styles&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Navigation automation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;navigate_page&lt;/code&gt;, &lt;code&gt;new_page&lt;/code&gt;, &lt;code&gt;wait_for&lt;/code&gt;, &lt;code&gt;list_pages&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Input automation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;click&lt;/code&gt;, &lt;code&gt;fill&lt;/code&gt;, &lt;code&gt;fill_form&lt;/code&gt;, &lt;code&gt;hover&lt;/code&gt;, &lt;code&gt;press_key&lt;/code&gt;, &lt;code&gt;upload_file&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Emulation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;emulate&lt;/code&gt; (viewport, network throttling, dark mode), &lt;code&gt;resize_page&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Memory&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;heap snapshot capture, comparison, and querying&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Extensions, third-party tools, WebMCP, PWA&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;13 combined&lt;/td&gt;
&lt;td&gt;browser extension and installed-app management&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; the same "audit this page" job as Lighthouse MCP, but backed by real trace data and the option to drive the page first (log in, add items to a cart, then trace the checkout flow) instead of only auditing a cold page load. Use &lt;code&gt;--slim&lt;/code&gt; if you want the input/navigation/debugging basics without the full 58-tool surface.&lt;/p&gt;




&lt;h3&gt;
  
  
  6.2 Lighthouse MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/priyankark/lighthouse-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;priyankark/lighthouse-mcp&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; TypeScript · &lt;strong&gt;Status:&lt;/strong&gt; 🔵 community&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Node.js 22.19+, Chrome/Chromium (sandboxed) · install via MCP Registry, &lt;code&gt;npx lighthouse-mcp&lt;/code&gt;, or global npm&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;AUDIT_ALLOW_LOOPBACK&lt;/code&gt; (default true; set &lt;code&gt;false&lt;/code&gt; for hosted/public-only audits)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Next to Chrome DevTools MCP's 58 tools at the time of writing, it's genuinely refreshing that this one only has two — sometimes you just want a score, not a workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SSRF and DNS-rebinding protections are on by default, so it's safe to expose this to a shared bot or a hosted worker&lt;/li&gt;
&lt;li&gt;Fast enough to run on every pull request without anyone noticing the extra time&lt;/li&gt;
&lt;li&gt;Mobile-first defaults, which matches where most real users actually are&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; a PR touching the checkout page triggers "audit this page on mobile and tell me the top fixes." &lt;code&gt;run_audit&lt;/code&gt; comes back with a score and a category breakdown; if all you need is the headline number, &lt;code&gt;get_performance_score&lt;/code&gt; skips the rest of the audit and answers in a fraction of the time.&lt;/p&gt;

&lt;p&gt;Deliberately minimal — two tools, both wrapping Google Lighthouse directly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Parameters&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;run_audit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;url&lt;/code&gt;, &lt;code&gt;categories?&lt;/code&gt; (&lt;code&gt;performance&lt;/code&gt;, &lt;code&gt;accessibility&lt;/code&gt;, &lt;code&gt;best-practices&lt;/code&gt;, &lt;code&gt;seo&lt;/code&gt; — default all), &lt;code&gt;device?&lt;/code&gt; (&lt;code&gt;mobile&lt;/code&gt; default \&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;desktop&lt;/code&gt;), &lt;code&gt;throttling?&lt;/code&gt; (default true)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_performance_score&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;url&lt;/code&gt;, &lt;code&gt;device?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Performance score only — faster than a full audit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Safety model, worth calling out explicitly:&lt;/strong&gt; Chrome runs sandboxed; loopback/private/link-local/ cloud-metadata destinations are blocked by default (redirects included) to prevent SSRF and DNS rebinding; audits are serialized one-at-a-time with a 120-second timeout; Chrome and the audit proxy are torn down after every run, success or failure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; a fast "what's wrong with this page" check when the worker's network and URL policies are configured appropriately.&lt;/p&gt;




&lt;h3&gt;
  
  
  6.3 PageSpeed Insights MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/ruslanlap/pagespeed-insights-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;ruslanlap/pagespeed-insights-mcp&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; TypeScript · &lt;strong&gt;Status:&lt;/strong&gt; 🔵 community&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Node.js 20.19+ · &lt;code&gt;npx -y pagespeed-insights-mcp&lt;/code&gt;, npm global, or Docker&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;GOOGLE_API_KEY&lt;/code&gt; (with the PageSpeed Insights API enabled in Google Cloud Console)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Lighthouse alone tells only part of the story — it's a lab measurement over a throttled connection, which is pessimistic by design. PageSpeed adds the real-user CrUX data that says what's actually happening in production, so the two together stop you from either overreacting to a lab number or ignoring a real regression.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Lab and field data live behind one conversational interface instead of two separate dashboards&lt;/li&gt;
&lt;li&gt;Built-in baseline comparison means "did the release regress?" is a single tool call, not a manual diff&lt;/li&gt;
&lt;li&gt;Batch analysis triages up to ten pages at once, which matters the moment you're auditing more than a homepage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; after a release, "compare mobile performance of the old and new build and prioritize the fixes" runs &lt;code&gt;pagespeed_analyze_page&lt;/code&gt; against both, pulls &lt;code&gt;pagespeed_get_field_data&lt;/code&gt; for the real-user view, and hands back a ranked list — usually something unglamorous like an unoptimized hero image or a blocking third-party script.&lt;/p&gt;

&lt;p&gt;Version 2 of this server deliberately &lt;strong&gt;replaced 19 endpoint-shaped tools with six workflow tools&lt;/strong&gt; — worth knowing if you find v1 examples online, since the old tool names no longer exist.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Key parameters&lt;/th&gt;
&lt;th&gt;What it covers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pagespeed_analyze_page&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;url&lt;/code&gt;, &lt;code&gt;strategy&lt;/code&gt; (mobile/desktop), &lt;code&gt;report&lt;/code&gt; (&lt;code&gt;full&lt;/code&gt; \&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;summary&lt;/code&gt; \&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pagespeed_diagnose_page&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;url&lt;/code&gt;, &lt;code&gt;focus&lt;/code&gt; (&lt;code&gt;visual&lt;/code&gt; \&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;elements&lt;/code&gt; \&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pagespeed_get_field_data&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;url&lt;/code&gt; or origin, &lt;code&gt;scope&lt;/code&gt; (&lt;code&gt;page&lt;/code&gt; \&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;origin&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pagespeed_compare_pages&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;two URLs, or one URL + &lt;code&gt;mode: baseline&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Page-vs-page or page-vs-saved-baseline comparison&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pagespeed_analyze_batch&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1–10 URLs&lt;/td&gt;
&lt;td&gt;Triage many pages at once, with progress notifications where the client supports it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pagespeed_clear_cache&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(none)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;Clears the in-memory API-response cache (useful right after a deploy)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every data-returning tool accepts &lt;code&gt;responseFormat: markdown&lt;/code&gt; (default) or &lt;code&gt;json&lt;/code&gt;, and results come back as structured MCP &lt;code&gt;structuredContent&lt;/code&gt;, not just prose.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; combining &lt;strong&gt;lab&lt;/strong&gt; results (Lighthouse, throttled and therefore pessimistic) with &lt;strong&gt;field&lt;/strong&gt; results (CrUX, what real visitors experienced) in the same conversation — the README's own example shows GitHub.com scoring 54/100 in the lab while CrUX shows real users seeing a 1.9s FCP, which is a good illustration of why you want both.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Automation
&lt;/h2&gt;

&lt;h3&gt;
  
  
  7.1 Playwright MCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/microsoft/playwright-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;microsoft/playwright-mcp&lt;/code&gt;&lt;/a&gt; · &lt;strong&gt;Language:&lt;/strong&gt; TypeScript · &lt;strong&gt;Status:&lt;/strong&gt; 🟢 &lt;strong&gt;vendor-maintained&lt;/strong&gt; (Microsoft) — source lives in the main &lt;a href="https://github.com/microsoft/playwright" rel="noopener noreferrer"&gt;Playwright monorepo&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; Node.js 18+ · install via &lt;code&gt;npx @playwright/mcp@latest&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config:&lt;/strong&gt; &lt;code&gt;--browser&lt;/code&gt; (chrome/firefox/webkit/msedge), &lt;code&gt;--headless&lt;/code&gt;, &lt;code&gt;--isolated&lt;/code&gt; (fresh profile per session) or persistent profile (default), &lt;code&gt;--device&lt;/code&gt; / &lt;code&gt;--mobile&lt;/code&gt; emulation, &lt;code&gt;--caps vision,pdf,devtools&lt;/code&gt; for optional extra capabilities, &lt;code&gt;--cdp-endpoint&lt;/code&gt; to attach to an already-running browser, &lt;code&gt;--allowed-origins&lt;/code&gt; / &lt;code&gt;--blocked-origins&lt;/code&gt; for network scoping&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth:&lt;/strong&gt; none by default — trust boundaries are enforced through &lt;code&gt;--allowed-origins&lt;/code&gt;/&lt;code&gt;--blocked-origins&lt;/code&gt; and the workspace-root file-access restriction, not credentials&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is also the browser automation server used by the LoadRunner Agent framework described in this series: the &lt;code&gt;lr-record-auto&lt;/code&gt; and &lt;code&gt;lr-record-manual&lt;/code&gt; skills drive Playwright MCP to record a real UI journey, and &lt;code&gt;lr-generate-har&lt;/code&gt; replays it headlessly to produce the HAR that feeds LoadRunner UI Vuser script generation. If you're already scripting browser automation for that pipeline, this is the same tool — you don't need a second one for general-purpose browser automation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Works off the accessibility tree by default, not screenshots — faster, more deterministic, and it doesn't need a vision-capable model&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;--isolated&lt;/code&gt; sessions give you a clean-slate browser for every recording, so cookies and login state from a previous run can't leak into the next one&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;--cdp-endpoint&lt;/code&gt; lets you attach to a browser you already launched — useful for recording against a real, already-authenticated session instead of automating a fresh login every time&lt;/li&gt;
&lt;li&gt;The same server doubles as your UI-automation recorder and your day-to-day "go check this page for me" browser agent — one integration, two jobs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Usage in practice:&lt;/strong&gt; ask it to "log into staging, add an item to the cart, and go to checkout" and it takes a snapshot of the page, clicks and fills using the accessibility tree, and can hand that journey off as the seed for a LoadRunner Web Vuser script — the exact loop this repo's &lt;code&gt;lr-record-auto&lt;/code&gt; skill automates end to end.&lt;/p&gt;

&lt;h3&gt;
  
  
  How a browser recording becomes a performance-test flow
&lt;/h3&gt;

&lt;p&gt;Playwright MCP does not turn one browser session into a finished load test by itself. It records the &lt;em&gt;business journey&lt;/em&gt; — the navigation, clicks, form submissions, waits, and resulting network activity — then a performance-testing tool converts that evidence into a script that can run many virtual users.&lt;/p&gt;

&lt;p&gt;The practical handoff looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Record the journey.&lt;/strong&gt; Ask the agent to drive a realistic flow, such as signing in, searching for a product, adding it to the cart, and checking out. Playwright MCP uses accessibility-tree actions such as &lt;code&gt;browser_navigate&lt;/code&gt;, &lt;code&gt;browser_click&lt;/code&gt;, and &lt;code&gt;browser_type&lt;/code&gt;, so the recording follows user intent rather than screen coordinates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capture the network evidence.&lt;/strong&gt; The recording pipeline observes the requests produced by that journey and saves them as a HAR (HTTP Archive). The HAR contains URLs, methods, headers, cookies, request bodies, responses, and timing details that a script generator can use as its input.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generate the performance script.&lt;/strong&gt; A converter such as this repository's &lt;code&gt;lr-generate-har&lt;/code&gt; skill replays the journey headlessly and produces a HAR for LoadRunner UI Vuser script generation. The result is a repeatable protocol script, not another Playwright automation session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prepare it for load.&lt;/strong&gt; Correlate dynamic values such as tokens and session IDs, parameterize user data, remove one-user-only steps, and add transactions, rendezvous points, checks, pacing, and load-model settings. These are performance-test decisions; they cannot be inferred safely from a single recording.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate at low load, then scale.&lt;/strong&gt; Run one or a few virtual users first, confirm that the generated flow reaches the intended business state, and only then run the broader test while collecting latency, throughput, error, and resource evidence.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That makes Playwright MCP useful as a &lt;strong&gt;journey recorder and script seed&lt;/strong&gt;: it captures the path a real user takes, while LoadRunner or another load engine supplies virtual-user concurrency, scheduling, measurements, and performance analysis. The recording is not automatically production-ready; authentication, test data, dynamic-value correlation, and traffic volume still need review.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Representative tools&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Navigation &amp;amp; interaction&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;browser_navigate&lt;/code&gt;, &lt;code&gt;browser_click&lt;/code&gt;, &lt;code&gt;browser_type&lt;/code&gt;, &lt;code&gt;browser_hover&lt;/code&gt;, &lt;code&gt;browser_press_key&lt;/code&gt;, &lt;code&gt;browser_select_option&lt;/code&gt;, &lt;code&gt;browser_drag&lt;/code&gt;, &lt;code&gt;browser_file_upload&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Structured, accessibility-tree-driven actions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inspection&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;browser_snapshot&lt;/code&gt;, &lt;code&gt;browser_take_screenshot&lt;/code&gt;, &lt;code&gt;browser_console_messages&lt;/code&gt;, &lt;code&gt;browser_network_requests&lt;/code&gt;, &lt;code&gt;browser_evaluate&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Read the page state without guessing from pixels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session &amp;amp; tabs&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;browser_tabs&lt;/code&gt;, &lt;code&gt;browser_wait_for&lt;/code&gt;, &lt;code&gt;browser_resize&lt;/code&gt;, &lt;code&gt;browser_close&lt;/code&gt;, &lt;code&gt;browser_install&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Multi-tab and lifecycle management&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The exact tool list ships with the server and is documented in full in its own reference — treat the names above as the stable, widely-referenced surface rather than an independently re-verified list from source in this pass, the same treatment given to BlazeMeter's tool names earlier.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good for:&lt;/strong&gt; the automation category exists because of this one server — general-purpose browser scripting that happens to be the same tool this repo's own UI-recording pipeline is built on, so wiring it in once covers both "record a user journey for load testing" and "go check something on this page for me."&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Server comparison at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Server&lt;/th&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Language&lt;/th&gt;
&lt;th&gt;Tool count&lt;/th&gt;
&lt;th&gt;Transport&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;k6&lt;/td&gt;
&lt;td&gt;Load &amp;amp; stress&lt;/td&gt;
&lt;td&gt;🟢 vendor-maintained&lt;/td&gt;
&lt;td&gt;Go&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;stdio / HTTP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BlazeMeter&lt;/td&gt;
&lt;td&gt;Load &amp;amp; stress&lt;/td&gt;
&lt;td&gt;🟢 vendor-maintained&lt;/td&gt;
&lt;td&gt;Python&lt;/td&gt;
&lt;td&gt;cloud workflows (vendor-documented)&lt;/td&gt;
&lt;td&gt;stdio / HTTP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LoadRunner Cloud&lt;/td&gt;
&lt;td&gt;Load &amp;amp; stress&lt;/td&gt;
&lt;td&gt;🔵 community&lt;/td&gt;
&lt;td&gt;JavaScript&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;stdio / SSE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apache JMeter&lt;/td&gt;
&lt;td&gt;Load &amp;amp; stress&lt;/td&gt;
&lt;td&gt;🔵 community&lt;/td&gt;
&lt;td&gt;TypeScript&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;stdio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Artillery&lt;/td&gt;
&lt;td&gt;Load &amp;amp; stress&lt;/td&gt;
&lt;td&gt;🔵 community&lt;/td&gt;
&lt;td&gt;TypeScript&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;stdio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dynatrace&lt;/td&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;🔵 community&lt;/td&gt;
&lt;td&gt;TypeScript&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;stdio / HTTP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Splunk&lt;/td&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;🔵 community&lt;/td&gt;
&lt;td&gt;Python&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;stdio / SSE / API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LitmusChaos&lt;/td&gt;
&lt;td&gt;Resilience&lt;/td&gt;
&lt;td&gt;🟢 vendor-maintained&lt;/td&gt;
&lt;td&gt;Go&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;GraphQL-backed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chrome DevTools MCP&lt;/td&gt;
&lt;td&gt;Front-end&lt;/td&gt;
&lt;td&gt;🟢 vendor-maintained&lt;/td&gt;
&lt;td&gt;TypeScript&lt;/td&gt;
&lt;td&gt;58&lt;/td&gt;
&lt;td&gt;stdio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lighthouse&lt;/td&gt;
&lt;td&gt;Front-end&lt;/td&gt;
&lt;td&gt;🔵 community&lt;/td&gt;
&lt;td&gt;TypeScript&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;stdio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PageSpeed Insights&lt;/td&gt;
&lt;td&gt;Front-end&lt;/td&gt;
&lt;td&gt;🔵 community&lt;/td&gt;
&lt;td&gt;TypeScript&lt;/td&gt;
&lt;td&gt;6 (v2)&lt;/td&gt;
&lt;td&gt;stdio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Playwright MCP&lt;/td&gt;
&lt;td&gt;Automation&lt;/td&gt;
&lt;td&gt;🟢 vendor-maintained&lt;/td&gt;
&lt;td&gt;TypeScript&lt;/td&gt;
&lt;td&gt;varies by enabled capabilities&lt;/td&gt;
&lt;td&gt;stdio&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;140+ independently-catalogued tools across ten servers, plus two more (BlazeMeter, Playwright) with vendor-documented tool surfaces — twelve servers, five categories, and that's before anyone writes the agent layer on top of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. What trips people up in practice
&lt;/h2&gt;

&lt;p&gt;These are the specific things that cause a working integration to quietly break, not a generic best-practices list:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Discover the live tool surface.&lt;/strong&gt; MCP clients can discover the tools exposed by a server at connection time. Use that discovery result as the source of truth for the version and configuration you have installed, and keep an explicit allow-list for workflows that depend on specific tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Old tool names stop working silently.&lt;/strong&gt; PageSpeed Insights replaced 19 endpoint-shaped tools with 6 workflow tools in its v2 release. Any example referencing the old names (&lt;code&gt;analyze_page_speed&lt;/code&gt;, &lt;code&gt;get_recommendations&lt;/code&gt;, etc.) will fail against a current install — check the version before you copy a snippet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Community tool surfaces need version checks.&lt;/strong&gt; Dynatrace and Splunk expose useful contracts for investigation and automation, but verify the exact tool names and parameters against the version you install.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost-bearing tools need a budget before you use them, not after.&lt;/strong&gt; &lt;code&gt;execute_dql&lt;/code&gt; (Dynatrace) scans Grail storage by volume. Nothing stops a broad query across 90 days of data the first time someone asks a vague question — set the budget first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not every tool is safe to auto-run.&lt;/strong&gt; Several servers mix read tools with ones that have real side effects — &lt;code&gt;send_email&lt;/code&gt;, &lt;code&gt;send_slack_message&lt;/code&gt;, &lt;code&gt;create_kvstore_collection&lt;/code&gt;, &lt;code&gt;delete_kvstore_collection&lt;/code&gt;. The tool schema itself won't tell you which is which; you have to decide that before wiring one up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SSE and stdio aren't interchangeable.&lt;/strong&gt; LoadRunner Cloud and Splunk both support multiple transports, but with different startup flags and, in Splunk's case, different default behavior per mode. Check the mode-specific section, not just the quickstart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A single vendor-maintained server can have a much larger tool surface than its peers.&lt;/strong&gt; Chrome DevTools MCP alone exposes 58 tools — the largest individual tool surface in this catalog. If your client loads every tool schema into context by default, use &lt;code&gt;--slim&lt;/code&gt; or an explicit allow-list instead of eating that cost on every request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Playwright's persistent profile is exclusive.&lt;/strong&gt; Only one browser instance can use the default persistent profile at a time; running two MCP clients against the same workspace will conflict. Use &lt;code&gt;--isolated&lt;/code&gt; or a distinct &lt;code&gt;--user-data-dir&lt;/code&gt; for parallel sessions.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  10. Production considerations
&lt;/h2&gt;

&lt;p&gt;A short checklist worth running before any of these servers goes anywhere near a shared environment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pin tool names, don't trust discovery blindly.&lt;/strong&gt; MCP's dynamic tool listing means a server operator can rename or remove a tool at any time. If your integration depends on a specific tool existing, maintain an explicit allow-list and fail loudly if it's missing, rather than silently adapting to whatever the server currently exposes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cap cost-bearing queries.&lt;/strong&gt; Dynatrace exposes &lt;code&gt;DT_GRAIL_QUERY_BUDGET_GB&lt;/code&gt; for exactly this reason — use it, and default to short timeframes (12–24h) rather than open-ended windows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gate side-effecting tools behind approval.&lt;/strong&gt; Anything that sends a message, creates a resource, or deletes data should require an explicit human or policy-engine confirmation, the same way you'd gate a write-capable API call in any other system.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify auth scope requirements up front.&lt;/strong&gt; Dynatrace alone has more than a dozen distinct OAuth scopes depending on which tools you use (&lt;code&gt;storage:logs:read&lt;/code&gt;, &lt;code&gt;storage:spans:read&lt;/code&gt;, &lt;code&gt;davis-copilot:*:execute&lt;/code&gt;, and so on) — request only what the tools you actually use require.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Track server status over time.&lt;/strong&gt; Community projects evolve quickly, so revisit this catalog's status column periodically rather than treating it as a one-time check.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't stand up a second browser-automation server.&lt;/strong&gt; If a team already wires up Playwright MCP for UI-script generation (as this repo's LoadRunner Agent framework does), reuse that connection instead of adding Chrome DevTools MCP or a third browser tool for the same job — decide which one owns browser automation and route everything through it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this — the tool names, the parameters, the trust tiers — matters on its own. What matters is what happens once one of these is wired into a Copilot agent with a clear job and a few guardrails: a lengthy investigation can turn into a short conversation, and a tool you'd otherwise have to look up becomes a capability your team just &lt;em&gt;has&lt;/em&gt;. The companion posts in this series walk through exactly that wiring, with a full worked example against Dynatrace — the config, the agent instructions, and a real conversation, not just a table of tool names.&lt;/p&gt;

&lt;p&gt;The MCP server is the capability layer; the agent is the reasoning and orchestration layer. The real engineering value appears when the two are combined with explicit tool ordering, context management, and guardrails. An MCP server exposes what can be done; the agent determines when and in what order to use it, while policy determines what requires approval.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 2 is coming&lt;/strong&gt; — a follow-up post on building the actual GitHub Copilot custom agents that sit in front of these MCP servers: the agent files, the tool-order and guardrail decisions, and worked examples beyond Dynatrace.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Credits &amp;amp; references
&lt;/h2&gt;

&lt;p&gt;Every project below is third-party open source. Thank you to the maintainers — please check each repository's license and current maintenance status before depending on it in production.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;k6 — &lt;a href="https://github.com/grafana/mcp-k6" rel="noopener noreferrer"&gt;https://github.com/grafana/mcp-k6&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;BlazeMeter — &lt;a href="https://github.com/Blazemeter/bzm-mcp" rel="noopener noreferrer"&gt;https://github.com/Blazemeter/bzm-mcp&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;LoadRunner Cloud — &lt;a href="https://github.com/pbandreddy/loadrunner-cloud-mcp-server" rel="noopener noreferrer"&gt;https://github.com/pbandreddy/loadrunner-cloud-mcp-server&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Apache JMeter — &lt;a href="https://github.com/aravindksk7/Jmeter-MCP" rel="noopener noreferrer"&gt;https://github.com/aravindksk7/Jmeter-MCP&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Artillery — &lt;a href="https://github.com/jch1887/artillery-mcp-server" rel="noopener noreferrer"&gt;https://github.com/jch1887/artillery-mcp-server&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Dynatrace — &lt;a href="https://github.com/dynatrace-oss/dynatrace-mcp" rel="noopener noreferrer"&gt;https://github.com/dynatrace-oss/dynatrace-mcp&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Splunk — &lt;a href="https://github.com/livehybrid/splunk-mcp" rel="noopener noreferrer"&gt;https://github.com/livehybrid/splunk-mcp&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;LitmusChaos — &lt;a href="https://github.com/litmuschaos/litmus-mcp-server" rel="noopener noreferrer"&gt;https://github.com/litmuschaos/litmus-mcp-server&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Chrome DevTools MCP — &lt;a href="https://github.com/ChromeDevTools/chrome-devtools-mcp" rel="noopener noreferrer"&gt;https://github.com/ChromeDevTools/chrome-devtools-mcp&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Lighthouse — &lt;a href="https://github.com/priyankark/lighthouse-mcp" rel="noopener noreferrer"&gt;https://github.com/priyankark/lighthouse-mcp&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;PageSpeed Insights — &lt;a href="https://github.com/ruslanlap/pagespeed-insights-mcp" rel="noopener noreferrer"&gt;https://github.com/ruslanlap/pagespeed-insights-mcp&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Playwright MCP — &lt;a href="https://github.com/microsoft/playwright-mcp" rel="noopener noreferrer"&gt;https://github.com/microsoft/playwright-mcp&lt;/a&gt; (source in &lt;a href="https://github.com/microsoft/playwright" rel="noopener noreferrer"&gt;microsoft/playwright&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This post is the reference you come back to when you're deciding which server to wire up next. Part 2 of this blog picks up from here and walks through building the Copilot agents themselves.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>mcp</category>
      <category>performance</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
