<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Muthu Kumar Koodalingam</title>
    <description>The latest articles on DEV Community by Muthu Kumar Koodalingam (@muthu_kumarkoodalingam).</description>
    <link>https://dev.to/muthu_kumarkoodalingam</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3684562%2Fbc53bfd4-846e-4a96-9cc5-67ead73741be.jpg</url>
      <title>DEV Community: Muthu Kumar Koodalingam</title>
      <link>https://dev.to/muthu_kumarkoodalingam</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/muthu_kumarkoodalingam"/>
    <language>en</language>
    <item>
      <title>Your Karate Test Failed in CI. Should AI Fix It?</title>
      <dc:creator>Muthu Kumar Koodalingam</dc:creator>
      <pubDate>Tue, 22 Sep 2026 07:58:49 +0000</pubDate>
      <link>https://dev.to/muthu_kumarkoodalingam/your-karate-test-failed-in-ci-should-ai-fix-it-2bcd</link>
      <guid>https://dev.to/muthu_kumarkoodalingam/your-karate-test-failed-in-ci-should-ai-fix-it-2bcd</guid>
      <description>&lt;h1&gt;
  
  
  Your Karate Test Failed in CI. Should AI Fix It?
&lt;/h1&gt;

&lt;p&gt;A failed API test often creates an awkward handoff.&lt;/p&gt;

&lt;p&gt;CI knows &lt;em&gt;that&lt;/em&gt; a scenario failed. The Karate report may know the failed step, HTTP status, request and response. The source repository knows the scenario that produced it. But by the time someone investigates, that evidence is scattered across logs, artifacts and source files.&lt;/p&gt;

&lt;p&gt;It is tempting to solve this with a simple prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here is the failing Karate feature and the error.
Fix the test.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is also a good way to create a test suite that slowly learns to accept broken behavior.&lt;/p&gt;

&lt;p&gt;The interesting problem is not whether an LLM can edit Gherkin. It can. The problem is deciding &lt;strong&gt;what evidence is strong enough to justify a repair, how narrowly the repair is allowed to operate, and where a human must remain in control.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This article describes the architecture I use in Karate Test Management for evidence-driven CI repair.&lt;/p&gt;

&lt;h2&gt;
  
  
  A failing test is not proof that the test is wrong
&lt;/h2&gt;

&lt;p&gt;Consider this scenario:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get an existing order
  &lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders', orderId
  &lt;span class="nf"&gt;When &lt;/span&gt;method get
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 200
  &lt;span class="nf"&gt;And &lt;/span&gt;match response.status == 'CONFIRMED'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CI reports:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;expected: 200
actual:   404
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A repair model could make the build green immediately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;- Then status 200
&lt;/span&gt;&lt;span class="gi"&gt;+ Then status 404
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is syntactically valid and operationally dangerous.&lt;/p&gt;

&lt;p&gt;The failure might mean:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the API contract changed;&lt;/li&gt;
&lt;li&gt;test data was not created;&lt;/li&gt;
&lt;li&gt;an environment reset removed the order;&lt;/li&gt;
&lt;li&gt;authentication selected the wrong tenant;&lt;/li&gt;
&lt;li&gt;the endpoint is broken;&lt;/li&gt;
&lt;li&gt;the test is genuinely stale.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A repair system that starts by editing the assertion has confused &lt;strong&gt;observed behavior&lt;/strong&gt; with &lt;strong&gt;expected behavior&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So the first rule is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Treat a CI failure as evidence to investigate, not permission to weaken a test.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Normalize CI failures before asking AI anything
&lt;/h2&gt;

&lt;p&gt;Different pipelines expose failures differently. GitHub Actions may provide downloadable test artifacts and logs. Jenkins may expose console output. A custom runner may already have structured JSON.&lt;/p&gt;

&lt;p&gt;The repair layer should not care.&lt;/p&gt;

&lt;p&gt;Normalize the useful evidence into one payload first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"github-actions"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"featurePath"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"src/test/java/orders/orders.feature"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"scenarioName"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Get an existing order"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"scenarioLine"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;18&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"failedStep"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Then status 200"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"errorMessage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"status code was: 404, expected: 200"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"httpRequest"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GET"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.example.test/orders/9821"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"httpResponse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;404&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"body"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;code&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;ORDER_NOT_FOUND&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;}"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"runId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"814521:1"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is much more useful than sending an entire CI log to a model.&lt;/p&gt;

&lt;p&gt;It gives the repair workflow explicit anchors:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CI run
  ↓
failed job / step
  ↓
feature + scenario identity
  ↓
HTTP evidence
  ↓
source scenario
  ↓
repair candidate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It also makes the process inspectable. If a proposed repair is questionable, you can see exactly which evidence produced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Locate the scenario precisely
&lt;/h2&gt;

&lt;p&gt;Scenario identity is an underrated part of automated repair.&lt;/p&gt;

&lt;p&gt;Imagine a repository containing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get user
  ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;in several feature files, or a feature with generated scenarios that have similar names.&lt;/p&gt;

&lt;p&gt;A repair engine should use more than a free-text scenario name where possible. Useful coordinates include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;feature path
scenario name
scenario line
scenario tags
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important safety property is &lt;strong&gt;unique identification&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If the system cannot identify exactly one source scenario, it should refuse the edit rather than guessing.&lt;/p&gt;

&lt;p&gt;This principle generalizes beyond Karate: autonomous code changes become safer when uncertainty causes the system to stop instead of broadening its write scope.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give the model evidence, not authority
&lt;/h2&gt;

&lt;p&gt;Once the source scenario is located, AI can help reason about the failure. But the prompt should define a narrow repair boundary.&lt;/p&gt;

&lt;p&gt;A useful repair context looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;REPAIR CONTEXT
Feature: src/test/java/orders/orders.feature
Scenario: Get an existing order
Failed step: Then status 200
Error: status code was: 404, expected: 200

HTTP EVIDENCE
GET https://api.example.test/orders/9821
Response: 404
Body: {"code":"ORDER_NOT_FOUND"}

CONSTRAINT
Fix only the failed step and its immediate dependencies.
Do not change any other scenario.
Do not invent endpoints or fields not visible in the evidence.

OUTPUT
Return the complete repaired Scenario block only.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice what is deliberately absent: "make the test pass."&lt;/p&gt;

&lt;p&gt;The goal is to produce a &lt;strong&gt;candidate repair&lt;/strong&gt;, not optimize for a green build at any cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the complete scenario is a useful repair unit
&lt;/h2&gt;

&lt;p&gt;Returning a single replacement line sounds safer, but failures often involve a small dependency chain.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="err"&gt;* def order = call read('classpath&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="err"&gt;helpers/create-order.feature')&lt;/span&gt;
&lt;span class="nf"&gt;* &lt;/span&gt;def orderId = order.id
&lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders', orderId
&lt;span class="nf"&gt;When &lt;/span&gt;method get
&lt;span class="nf"&gt;Then &lt;/span&gt;status 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;order.id&lt;/code&gt; changed to &lt;code&gt;order.orderId&lt;/code&gt;, repairing only the failed HTTP assertion cannot solve the root problem.&lt;/p&gt;

&lt;p&gt;On the other hand, returning the entire feature file gives the model unnecessary authority over unrelated tests.&lt;/p&gt;

&lt;p&gt;The scenario is a useful middle ground:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;single line        → often insufficient context
single scenario    → useful bounded repair surface
entire feature     → unnecessarily broad write scope
repository         → far too broad
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The repair system can then replace exactly that scenario in the original feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Diff before apply should be the default
&lt;/h2&gt;

&lt;p&gt;Suppose the model proposes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="p"&gt;Scenario: Get an existing order
&lt;/span&gt;&lt;span class="gd"&gt;- * def orderId = order.id
&lt;/span&gt;&lt;span class="gi"&gt;+ * def orderId = order.orderId
&lt;/span&gt;  Given path 'orders', orderId
  When method get
  Then status 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the moment where automation should slow down.&lt;/p&gt;

&lt;p&gt;Show the original and candidate side by side. Let the engineer inspect whether the change preserves intent.&lt;/p&gt;

&lt;p&gt;In Karate Test Management, CI repair defaults to a reviewable diff. Automatic application is an explicit setting rather than the default behavior. A backup can also be created before applying a repair.&lt;/p&gt;

&lt;p&gt;That distinction matters because there are really two separate capabilities:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI can propose a change
        ≠
AI is authorized to modify the test suite
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keeping those permissions separate is one of the simplest guardrails for AI-assisted engineering.&lt;/p&gt;

&lt;h2&gt;
  
  
  Never let repair erase the oracle
&lt;/h2&gt;

&lt;p&gt;The most dangerous repair is one that makes a test less capable of detecting defects.&lt;/p&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="nf"&gt;Then &lt;/span&gt;status 200
&lt;span class="nf"&gt;And &lt;/span&gt;match response ==
&lt;span class="s"&gt;"""
{
  id: '#number',
  state: 'CONFIRMED'
}
"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A naive repair could turn this into:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="nf"&gt;Then &lt;/span&gt;status 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The build becomes green, but the test has lost most of its value.&lt;/p&gt;

&lt;p&gt;A practical repair review should therefore ask:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Did the candidate change the expected business behavior?&lt;/li&gt;
&lt;li&gt;Did it remove or weaken assertions?&lt;/li&gt;
&lt;li&gt;Did it introduce an endpoint, field or status not supported by evidence?&lt;/li&gt;
&lt;li&gt;Did it modify anything outside the failed scenario?&lt;/li&gt;
&lt;li&gt;Is the failure more likely product behavior, environment behavior or test behavior?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That final question is important. Some failures should produce &lt;strong&gt;no repair at all&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Separate deterministic work from AI work
&lt;/h2&gt;

&lt;p&gt;Most of the CI-repair pipeline does not require an LLM.&lt;/p&gt;

&lt;p&gt;Deterministic code can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;identify failed jobs;&lt;/li&gt;
&lt;li&gt;download artifacts and logs;&lt;/li&gt;
&lt;li&gt;extract &lt;code&gt;.feature&lt;/code&gt; references;&lt;/li&gt;
&lt;li&gt;locate scenario names and source lines;&lt;/li&gt;
&lt;li&gt;detect common status/assertion failures;&lt;/li&gt;
&lt;li&gt;extract HTTP methods, URLs, statuses and response bodies;&lt;/li&gt;
&lt;li&gt;locate the exact scenario in the repository;&lt;/li&gt;
&lt;li&gt;generate a diff;&lt;/li&gt;
&lt;li&gt;enforce write boundaries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI becomes useful after that evidence has been assembled, where the task requires reasoning about how the scenario might need to change.&lt;/p&gt;

&lt;p&gt;A safer architecture therefore looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GitHub Actions / CI
        ↓
Failure extraction          deterministic
        ↓
Structured evidence         deterministic
        ↓
Scenario location           deterministic
        ↓
Repair proposal             AI-assisted
        ↓
Scope validation            deterministic
        ↓
Diff / review               human
        ↓
Apply
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a much stronger pattern than giving an agent a repository and asking it to "fix CI."&lt;/p&gt;

&lt;h2&gt;
  
  
  Pulling evidence from GitHub Actions
&lt;/h2&gt;

&lt;p&gt;For GitHub Actions, a useful workflow is to inspect the failed run and gather the evidence already produced by the test job.&lt;/p&gt;

&lt;p&gt;The current Karate Test Management implementation can pull workflow runs, jobs, artifacts and logs, then search those sources for a feature reference, scenario name, error and HTTP evidence.&lt;/p&gt;

&lt;p&gt;The extraction is intentionally best-effort because CI output varies. A Karate failure might expose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;src/test/java/orders/orders.feature:18
Scenario: Get an existing order
status code was: 404, expected: 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and request logs may contain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&amp;gt; GET https://api.example.test/orders/9821
response status: 404
{"code":"ORDER_NOT_FOUND"}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From those fragments, the system can build the structured payload shown earlier.&lt;/p&gt;

&lt;p&gt;If it cannot infer the feature path reliably, the correct behavior is to stop the automated repair path rather than invent a target.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about Jenkins or GitLab?
&lt;/h2&gt;

&lt;p&gt;The same architecture does not require every CI provider to expose the same API.&lt;/p&gt;

&lt;p&gt;The normalized failure contract can represent several sources:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;FailureSource&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;github-actions&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;jenkins&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gitlab-ci&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;generic&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That means provider-specific adapters can evolve independently while the repair engine continues to consume one evidence shape.&lt;/p&gt;

&lt;p&gt;A Jenkins pipeline, for example, could POST a compact failure payload after a failed Karate run instead of granting a desktop extension broad access to Jenkins itself.&lt;/p&gt;

&lt;p&gt;This separation is useful operationally and from a security perspective.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI repair should be optional
&lt;/h2&gt;

&lt;p&gt;There is another architectural boundary worth keeping: the test suite must remain usable when AI is unavailable.&lt;/p&gt;

&lt;p&gt;Generation, execution, coverage analysis and normal test management should not depend on a model provider. Repair is an enhancement layer.&lt;/p&gt;

&lt;p&gt;If AI is disabled, quota is exhausted, or a provider is unavailable, the CI evidence is still useful:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Failed scenario
Failed step
Request
Response
Source location
Run identity
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An engineer can investigate from that information directly.&lt;/p&gt;

&lt;p&gt;This prevents AI availability from becoming a new dependency in the delivery pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger lesson: automate evidence before decisions
&lt;/h2&gt;

&lt;p&gt;CI repair is one example of a broader pattern I increasingly use for AI-assisted test engineering.&lt;/p&gt;

&lt;p&gt;Do not begin with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;failure → AI → code change
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Build this instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;failure
  ↓
collect evidence
  ↓
normalize evidence
  ↓
identify exact change surface
  ↓
reason about candidate change
  ↓
validate boundaries
  ↓
review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI step becomes smaller, better informed and easier to audit.&lt;/p&gt;

&lt;p&gt;That is important for testing tools because the test suite is itself part of the safety system. An automated repair that silently weakens the oracle can be worse than a failing build.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this fits into Karate Test Management
&lt;/h2&gt;

&lt;p&gt;Karate Test Management treats CI repair as one part of a larger test lifecycle rather than a standalone "self-healing" trick.&lt;/p&gt;

&lt;p&gt;The current project includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;structured CI failure ingestion;&lt;/li&gt;
&lt;li&gt;GitHub Actions run, artifact and log intake;&lt;/li&gt;
&lt;li&gt;feature/scenario location;&lt;/li&gt;
&lt;li&gt;HTTP evidence extraction;&lt;/li&gt;
&lt;li&gt;AI-assisted scenario repair;&lt;/li&gt;
&lt;li&gt;bounded scenario replacement;&lt;/li&gt;
&lt;li&gt;reviewable diffs;&lt;/li&gt;
&lt;li&gt;optional backups and explicit auto-apply;&lt;/li&gt;
&lt;li&gt;the same provider-routing guardrails used by other AI-assisted workflows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The principle behind all of them is the same: &lt;strong&gt;AI can help reason over test evidence, but deterministic controls decide what it is allowed to touch.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you use Karate and want to inspect or experiment with the implementation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/mov2day/KaratePlugin" rel="noopener noreferrer"&gt;https://github.com/mov2day/KaratePlugin&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;VS Code Marketplace: &lt;a href="https://marketplace.visualstudio.com/items?itemName=MuthuKumarKoodalingam.karate-test-generator" rel="noopener noreferrer"&gt;https://marketplace.visualstudio.com/items?itemName=MuthuKumarKoodalingam.karate-test-generator&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result I want from AI-assisted maintenance is not "tests that heal themselves." It is something less magical and much more useful: &lt;strong&gt;failures that arrive with enough evidence to make a small, reviewable repair possible.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://muthukumarkoodalingam.com/blog/karate-ci-failure-repair-evidence/" rel="noopener noreferrer"&gt;muthukumarkoodalingam.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>api</category>
      <category>vscode</category>
    </item>
    <item>
      <title>Your Karate Tests Are Green. But Is Your API Actually Covered?</title>
      <dc:creator>Muthu Kumar Koodalingam</dc:creator>
      <pubDate>Tue, 15 Sep 2026 07:35:34 +0000</pubDate>
      <link>https://dev.to/muthu_kumarkoodalingam/your-karate-tests-are-green-but-is-your-api-actually-covered-2583</link>
      <guid>https://dev.to/muthu_kumarkoodalingam/your-karate-tests-are-green-but-is-your-api-actually-covered-2583</guid>
      <description>&lt;h1&gt;
  
  
  Your Karate Tests Are Green. But Is Your API Actually Covered?
&lt;/h1&gt;

&lt;p&gt;A green Karate run answers one question well:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Did the tests we executed pass?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It does &lt;strong&gt;not&lt;/strong&gt; answer a different question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Did we test the API surface that matters?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That distinction becomes painful in mature API suites. You can have 300 passing scenarios, a beautiful CI report, and still have newly added operations, destructive methods, or entire resource paths that no Karate scenario exercises.&lt;/p&gt;

&lt;p&gt;The useful coverage metric for an API suite is therefore not simply scenario pass rate. It is the relationship between the &lt;strong&gt;contract you expose&lt;/strong&gt; and the &lt;strong&gt;tests you can map back to that contract&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This article shows a practical way to build that map using OpenAPI and an existing Karate suite, including where simple matching fails and what evidence you should retain before generating missing tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pass rate and API coverage are different dimensions
&lt;/h2&gt;

&lt;p&gt;Assume CI reports:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Karate scenarios: 184
Passed:           184
Failed:             0
Pass rate:         100%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That looks excellent.&lt;/p&gt;

&lt;p&gt;Now compare the suite with the current OpenAPI contract:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET    /customers                 covered
POST   /customers                 covered
GET    /customers/{customerId}    covered
PATCH  /customers/{customerId}    missing
DELETE /customers/{customerId}    missing
GET    /customers/{customerId}/orders  missing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The suite is still 100% green. It is simply green over an incomplete slice of the API.&lt;/p&gt;

&lt;p&gt;This is why I treat &lt;strong&gt;execution health&lt;/strong&gt; and &lt;strong&gt;contract coverage&lt;/strong&gt; as separate signals:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Execution health = did known tests pass?
Contract coverage = which API operations have test evidence?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Neither replaces the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with operations, not lines of code
&lt;/h2&gt;

&lt;p&gt;For API testing, an immediately useful unit of coverage is an OpenAPI operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HTTP method + normalized path
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET  /orders/{id}
POST /orders
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is more actionable than application code coverage when the question is whether your externally documented API surface is exercised.&lt;/p&gt;

&lt;p&gt;Given this OpenAPI fragment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;/orders&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;post&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;createOrder&lt;/span&gt;
      &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;201'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Created&lt;/span&gt;

  &lt;span class="s"&gt;/orders/{orderId}&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;getOrder&lt;/span&gt;
      &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;orderId&lt;/span&gt;
          &lt;span class="na"&gt;in&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;path&lt;/span&gt;
          &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
          &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;we can derive two contract operations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST /orders
GET  /orders/{orderId}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next job is to find evidence for those operations in Karate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mapping Karate scenarios to OpenAPI operations
&lt;/h2&gt;

&lt;p&gt;Consider this feature:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kd"&gt;Feature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Order lookup

&lt;span class="kn"&gt;Background&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
  &lt;span class="nf"&gt;* &lt;/span&gt;url baseUrl

&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get an existing order
  &lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders', orderId
  &lt;span class="nf"&gt;When &lt;/span&gt;method get
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 200
  &lt;span class="nf"&gt;And &lt;/span&gt;match response.id == orderId
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A human immediately sees that it probably covers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /orders/{orderId}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A coverage engine has to establish that relationship systematically.&lt;/p&gt;

&lt;p&gt;At minimum it needs to extract:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;method = GET
path   = /orders/&amp;lt;dynamic-value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and normalize the dynamic segment against the contract path.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Karate:  GET /orders/8c12...
OpenAPI: GET /orders/{orderId}
                     ↓
                  match
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You do not need an LLM for this basic comparison. It should be deterministic and explainable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Normalize paths before comparing them
&lt;/h2&gt;

&lt;p&gt;Literal string matching is too weak.&lt;/p&gt;

&lt;p&gt;These can all represent the same contract operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders', orderId
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders/' + orderId
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="nf"&gt;Given &lt;/span&gt;url baseUrl + '/orders/' + orderId
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while OpenAPI represents it as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/orders/{orderId}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A useful mapper normalizes both sides into comparable structures.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Contract segments: [orders, {param}]
Test segments:     [orders, &amp;lt;dynamic&amp;gt;]
Method:            GET == GET
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then coverage can be based on structural compatibility rather than source-code spelling.&lt;/p&gt;

&lt;p&gt;Be conservative when evidence is ambiguous. A false &lt;code&gt;covered&lt;/code&gt; result is more dangerous than an &lt;code&gt;unknown&lt;/code&gt; result because it hides work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coverage should retain evidence
&lt;/h2&gt;

&lt;p&gt;A percentage alone is not enough.&lt;/p&gt;

&lt;p&gt;Suppose a dashboard says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;API coverage: 82%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first useful question is: &lt;em&gt;why 82%?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For every covered operation, retain evidence such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /orders/{orderId}
  -&amp;gt; src/test/java/orders/orders.feature
  -&amp;gt; Scenario: Get an existing order
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For uncovered operations, retain the opposite evidence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DELETE /orders/{orderId}
  -&amp;gt; no mapped Karate scenario
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This turns coverage into something engineers can inspect rather than a vanity metric.&lt;/p&gt;

&lt;p&gt;A practical result model might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"DELETE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/orders/{orderId}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"uncovered"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"evidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;versus:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GET"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/orders/{orderId}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"covered"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"evidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"feature"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"orders/orders.feature"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"scenario"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Get an existing order"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the number is reproducible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not confuse endpoint coverage with assertion quality
&lt;/h2&gt;

&lt;p&gt;Finding a mapped scenario proves that an operation is exercised. It does not prove the scenario is strong.&lt;/p&gt;

&lt;p&gt;This scenario technically reaches the endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get order
  &lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders', orderId
  &lt;span class="nf"&gt;When &lt;/span&gt;method get
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A more meaningful test may include contract-relevant assertions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get order
  &lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders', orderId
  &lt;span class="nf"&gt;When &lt;/span&gt;method get
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 200
  &lt;span class="nf"&gt;And &lt;/span&gt;match response ==
    &lt;span class="s"&gt;"""
    {
      id: '#string',
      status: '#string',
      total: '#number',
      items: '#array'
    }
    """&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I therefore avoid pretending operation coverage is a complete test-quality score.&lt;/p&gt;

&lt;p&gt;Think of quality as several independent questions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Was the operation exercised?
Did the scenario pass?
Are meaningful assertions present?
Are important response/error variants tested?
Is the test stable over repeated runs?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A useful test-management system keeps these signals separate instead of compressing all of them into one opaque score.&lt;/p&gt;

&lt;h2&gt;
  
  
  Positive coverage is only the first layer
&lt;/h2&gt;

&lt;p&gt;OpenAPI can also expose candidate coverage dimensions inside an operation.&lt;/p&gt;

&lt;p&gt;Suppose the contract documents:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;200'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Order returned&lt;/span&gt;
  &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;404'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Order not found&lt;/span&gt;
  &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;401'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Authentication required&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A single happy-path scenario means the operation is exercised, but the documented behavior is not necessarily well covered.&lt;/p&gt;

&lt;p&gt;You can model progressively richer coverage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Level 1: operation coverage
  GET /orders/{id}

Level 2: documented response coverage
  200 covered
  404 missing
  401 missing

Level 3: schema / boundary coverage
  required fields
  enums
  minimum / maximum
  formats
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is to label these dimensions accurately. Do not report Level 1 as though it proves Level 3.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generate from gaps, not from the whole specification
&lt;/h2&gt;

&lt;p&gt;Once the comparison is trustworthy, generation becomes much safer.&lt;/p&gt;

&lt;p&gt;Imagine the analysis produces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;12 uncovered operations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The wrong next step is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;regenerate 86 operations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The better workflow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenAPI contract
       +
Existing Karate suite
       ↓
Operation mapping
       ↓
Coverage evidence
       ↓
Uncovered operations only
       ↓
Generate candidate tests
       ↓
Review + execute
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This matters because mature suites contain engineering intent that does not exist in OpenAPI: authentication helpers, setup flows, domain fixtures, custom assertions, tags and team conventions.&lt;/p&gt;

&lt;p&gt;Targeted generation reduces the chance of replacing that intent with generic generated code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example: turning one coverage gap into a test
&lt;/h2&gt;

&lt;p&gt;Suppose coverage identifies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DELETE /orders/{orderId} — uncovered
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The contract says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;/orders/{orderId}&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;delete&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deleteOrder&lt;/span&gt;
    &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;orderId&lt;/span&gt;
        &lt;span class="na"&gt;in&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;path&lt;/span&gt;
        &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
        &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
    &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;204'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Deleted&lt;/span&gt;
      &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;404'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Not found&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not invent a random order ID and call that complete.&lt;/p&gt;

&lt;p&gt;First inspect how the repository creates valid orders. If an existing reusable feature exists:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="err"&gt;* def created = call read('classpath&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="err"&gt;helpers/create-order.feature')&lt;/span&gt;
&lt;span class="nf"&gt;* &lt;/span&gt;def orderId = created.response.id
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;then a generated candidate can fit the suite:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Delete an existing order
  &lt;span class="err"&gt;* def created = call read('classpath&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="err"&gt;helpers/create-order.feature')&lt;/span&gt;
  &lt;span class="nf"&gt;* &lt;/span&gt;def orderId = created.response.id

  &lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders', orderId
  &lt;span class="nf"&gt;When &lt;/span&gt;method delete
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 204
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A second candidate can address documented negative behavior:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Delete an unknown order
  &lt;span class="nf"&gt;Given &lt;/span&gt;path 'orders', 'does-not-exist'
  &lt;span class="nf"&gt;When &lt;/span&gt;method delete
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 404
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is far more useful than blindly producing one isolated scenario per OpenAPI operation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when the specification changes?
&lt;/h2&gt;

&lt;p&gt;Coverage becomes even more valuable when it is recomputed after contract changes.&lt;/p&gt;

&lt;p&gt;Suppose yesterday's contract contained:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET  /orders/{id}
POST /orders
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and today's contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET   /orders/{id}
POST  /orders
PATCH /orders/{id}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The interesting result is not simply "OpenAPI changed."&lt;/p&gt;

&lt;p&gt;It is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;New operation:
PATCH /orders/{id}

Mapped scenarios:
none

Coverage state:
uncovered
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is actionable change impact.&lt;/p&gt;

&lt;p&gt;Likewise, if a response schema adds a required field, the tool should distinguish that from a completely new operation. The likely test impact is different.&lt;/p&gt;

&lt;p&gt;This is where contract coverage starts becoming a maintenance mechanism rather than a reporting feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where AI helps — and where it should not decide the truth
&lt;/h2&gt;

&lt;p&gt;AI is useful after deterministic evidence exists.&lt;/p&gt;

&lt;p&gt;Given:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PATCH /orders/{id} is uncovered
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;plus the operation schema and nearby repository patterns, an AI assistant can help propose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;realistic domain scenarios&lt;/li&gt;
&lt;li&gt;useful boundaries&lt;/li&gt;
&lt;li&gt;repository-consistent naming&lt;/li&gt;
&lt;li&gt;stronger assertions&lt;/li&gt;
&lt;li&gt;reuse of nearby setup helpers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But I would not ask the model to be the source of truth for whether the operation is covered.&lt;/p&gt;

&lt;p&gt;The safer split is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Deterministic layer
  parse OpenAPI
  index Karate
  map operations
  identify gaps
  retain evidence

AI-assisted layer
  interpret domain intent
  suggest scenarios
  adapt to local style
  propose repairs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That separation makes the result inspectable when the model is wrong or unavailable.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Karate Test Management approaches this
&lt;/h2&gt;

&lt;p&gt;This workflow is now part of &lt;strong&gt;Karate Test Management for VS Code&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The current extension can analyze an OpenAPI specification by itself, or combine the specification with Karate feature files for exact scenario-to-endpoint mapping. Coverage gaps enter the Quality workflow, and missing-test generation can start from those gaps rather than regenerating the entire API suite.&lt;/p&gt;

&lt;p&gt;The same workspace also keeps execution, run history, flakiness, specification-change findings and repair review alongside coverage. That matters because an uncovered endpoint, a failing endpoint and a flaky endpoint are three different engineering problems.&lt;/p&gt;

&lt;p&gt;The implementation deliberately keeps deterministic coverage available without an AI provider. AI enhancement is optional and uses the evidence produced by the underlying workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical coverage review for your own Karate repository
&lt;/h2&gt;

&lt;p&gt;Even without tooling, you can apply the same model manually or in a script:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Parse the current OpenAPI document into &lt;code&gt;(method, normalized path)&lt;/code&gt; operations.&lt;/li&gt;
&lt;li&gt;Index Karate scenarios and extract the HTTP method and resolved/structural path evidence you can establish.&lt;/li&gt;
&lt;li&gt;Map tests to operations conservatively.&lt;/li&gt;
&lt;li&gt;Mark ambiguous mappings as unknown instead of covered.&lt;/li&gt;
&lt;li&gt;Keep the feature and scenario name as evidence for every match.&lt;/li&gt;
&lt;li&gt;Review uncovered operations by risk, not alphabetically.&lt;/li&gt;
&lt;li&gt;Generate or write tests only for justified gaps.&lt;/li&gt;
&lt;li&gt;Re-run the mapping whenever the contract changes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The key output should look less like a single percentage and more like an engineering queue:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HIGH   DELETE /customers/{id}       uncovered
HIGH   POST   /payments/refunds     uncovered
MEDIUM PATCH  /orders/{id}          uncovered
LOW    GET    /catalog/metadata     uncovered
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now API coverage can influence release decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Green is necessary, not sufficient
&lt;/h2&gt;

&lt;p&gt;A passing suite is good evidence about the tests you ran.&lt;/p&gt;

&lt;p&gt;It is not evidence about tests that do not exist.&lt;/p&gt;

&lt;p&gt;For contract-driven APIs, comparing Karate scenarios with OpenAPI gives you a practical way to expose that blind spot. Keep the mapping deterministic, retain evidence, distinguish endpoint coverage from assertion quality, and generate from gaps instead of repeatedly generating the whole suite.&lt;/p&gt;

&lt;p&gt;That is the difference between asking &lt;strong&gt;"are my tests green?"&lt;/strong&gt; and asking the more useful question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What part of my API do those green tests actually protect?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Karate Test Management is open source if you want to try the workflow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/mov2day/KaratePlugin" rel="noopener noreferrer"&gt;https://github.com/mov2day/KaratePlugin&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;VS Code Marketplace: &lt;a href="https://marketplace.visualstudio.com/items?itemName=MuthuKumarKoodalingam.karate-test-generator" rel="noopener noreferrer"&gt;https://marketplace.visualstudio.com/items?itemName=MuthuKumarKoodalingam.karate-test-generator&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>vscode</category>
      <category>testing</category>
      <category>api</category>
    </item>
    <item>
      <title>How to Generate Karate API Tests from OpenAPI — Without Creating a Maintenance Nightmare</title>
      <dc:creator>Muthu Kumar Koodalingam</dc:creator>
      <pubDate>Fri, 11 Sep 2026 16:16:06 +0000</pubDate>
      <link>https://dev.to/muthu_kumarkoodalingam/how-to-generate-karate-api-tests-from-openapi-without-creating-a-maintenance-nightmare-5131</link>
      <guid>https://dev.to/muthu_kumarkoodalingam/how-to-generate-karate-api-tests-from-openapi-without-creating-a-maintenance-nightmare-5131</guid>
      <description>&lt;h1&gt;
  
  
  How to Generate Karate API Tests from OpenAPI — Without Creating a Maintenance Nightmare
&lt;/h1&gt;

&lt;p&gt;Generating API tests from an OpenAPI file sounds like an easy automation win.&lt;/p&gt;

&lt;p&gt;Parse the specification. Create one test per endpoint. Add a few assertions. Commit the generated files.&lt;/p&gt;

&lt;p&gt;That approach works surprisingly well in a demo — and often becomes painful as soon as the API or test suite starts evolving.&lt;/p&gt;

&lt;p&gt;The hard problem is not producing Karate syntax. The hard problem is producing tests that fit the way your project already works and remain useful after the first generation run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tempting approach
&lt;/h2&gt;

&lt;p&gt;Imagine an API contract containing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;/orders&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;post&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;createOrder&lt;/span&gt;
  &lt;span class="s"&gt;/orders/{id}&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;getOrder&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A generator can easily produce something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kd"&gt;Feature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Orders API

&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Create an order
  &lt;span class="nf"&gt;Given &lt;/span&gt;url baseUrl
  &lt;span class="nf"&gt;And &lt;/span&gt;path 'orders'
  &lt;span class="err"&gt;And request { productId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;123, quantity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;1&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;
  &lt;span class="nf"&gt;When &lt;/span&gt;method post
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 201

&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get an order
  &lt;span class="nf"&gt;Given &lt;/span&gt;url baseUrl
  &lt;span class="nf"&gt;And &lt;/span&gt;path 'orders', 123
  &lt;span class="nf"&gt;When &lt;/span&gt;method get
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is valid Karate. It may even pass.&lt;/p&gt;

&lt;p&gt;But several important questions are still unanswered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where did &lt;code&gt;baseUrl&lt;/code&gt; come from?&lt;/li&gt;
&lt;li&gt;Does the project already have authentication helpers?&lt;/li&gt;
&lt;li&gt;Is &lt;code&gt;productId: 123&lt;/code&gt; valid test data?&lt;/li&gt;
&lt;li&gt;What should the response schema contain?&lt;/li&gt;
&lt;li&gt;What happens for an invalid product?&lt;/li&gt;
&lt;li&gt;What happens when &lt;code&gt;quantity&lt;/code&gt; is missing?&lt;/li&gt;
&lt;li&gt;Is the endpoint already covered by another scenario?&lt;/li&gt;
&lt;li&gt;Does the team have naming, tagging or reusable-feature conventions?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A generated test that ignores those questions creates code, but not necessarily useful coverage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat OpenAPI as a contract, not a test suite
&lt;/h2&gt;

&lt;p&gt;OpenAPI tells us a lot about the API surface:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;paths and HTTP methods&lt;/li&gt;
&lt;li&gt;parameters&lt;/li&gt;
&lt;li&gt;request schemas&lt;/li&gt;
&lt;li&gt;response schemas&lt;/li&gt;
&lt;li&gt;required fields&lt;/li&gt;
&lt;li&gt;enums and formats&lt;/li&gt;
&lt;li&gt;documented response codes&lt;/li&gt;
&lt;li&gt;security schemes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That information is excellent input for test generation.&lt;/p&gt;

&lt;p&gt;But the specification usually does not contain all the information needed for an executable test suite. Environment configuration, reusable authentication, realistic data creation and project conventions typically live elsewhere.&lt;/p&gt;

&lt;p&gt;A good generator therefore needs two kinds of context:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Contract context&lt;/strong&gt; — what the API says it accepts and returns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repository context&lt;/strong&gt; — how this project already executes Karate tests.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Ignoring the second is where many generated suites become difficult to maintain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generate deterministic coverage first
&lt;/h2&gt;

&lt;p&gt;I prefer to begin with rules that do not need an LLM.&lt;/p&gt;

&lt;p&gt;For each operation, deterministic generation can derive candidate scenarios from the contract.&lt;/p&gt;

&lt;p&gt;For example, a required field gives us at least two obvious cases:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;quantity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;integer&lt;/span&gt;
  &lt;span class="na"&gt;minimum&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Positive case:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="err"&gt;And request { productId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;123, quantity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;1&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;
&lt;span class="nf"&gt;When &lt;/span&gt;method post
&lt;span class="nf"&gt;Then &lt;/span&gt;status 201
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Boundary or negative candidates can be derived from the same schema:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="err"&gt;And request { productId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;123, quantity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;0&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;
&lt;span class="nf"&gt;When &lt;/span&gt;method post
&lt;span class="nf"&gt;Then &lt;/span&gt;status 400
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="err"&gt;And request { productId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;123&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;
&lt;span class="nf"&gt;When &lt;/span&gt;method post
&lt;span class="nf"&gt;Then &lt;/span&gt;status 400
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This has an important advantage: the reasoning is explainable. We know exactly why each scenario exists.&lt;/p&gt;

&lt;p&gt;AI can still improve the suite later, but it should not be the only mechanism deciding what needs to be tested.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reuse the project before inventing configuration
&lt;/h2&gt;

&lt;p&gt;Before generating a new feature, inspect the repository.&lt;/p&gt;

&lt;p&gt;A Karate project may already contain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;src/test/java/
  karate-config.js
  auth/
    token.feature
  common/
    create-user.feature
  orders/
    orders.feature
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If authentication already comes from &lt;code&gt;karate-config.js&lt;/code&gt;, generated tests should use it.&lt;/p&gt;

&lt;p&gt;If the team creates test users through a reusable feature, the generator should reuse that pattern instead of inventing credentials.&lt;/p&gt;

&lt;p&gt;If scenarios consistently use tags such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="nt"&gt;@orders&lt;/span&gt; &lt;span class="nt"&gt;@smoke&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;new tests should follow the same convention.&lt;/p&gt;

&lt;p&gt;Repository awareness is the difference between generating a standalone example and adding a maintainable test to an existing suite.&lt;/p&gt;

&lt;h2&gt;
  
  
  Detect coverage before generating more tests
&lt;/h2&gt;

&lt;p&gt;Another mistake is assuming every OpenAPI operation needs a new feature file.&lt;/p&gt;

&lt;p&gt;First compare the API contract with the existing Karate scenarios.&lt;/p&gt;

&lt;p&gt;Think of the API surface as a matrix:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST   /orders         covered
GET    /orders/{id}    covered
DELETE /orders/{id}    missing
PATCH  /orders/{id}    missing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now generation becomes targeted.&lt;/p&gt;

&lt;p&gt;Instead of producing another complete test suite, generate tests only for the gaps.&lt;/p&gt;

&lt;p&gt;That makes the workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenAPI
   ↓
Existing Karate suite
   ↓
Coverage mapping
   ↓
Missing operations
   ↓
Generate only what is missing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is much more useful in a real repository than repeatedly regenerating everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use AI as an enhancement layer
&lt;/h2&gt;

&lt;p&gt;There are places where AI is genuinely useful.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;understanding domain intent from descriptions&lt;/li&gt;
&lt;li&gt;suggesting edge cases that are not encoded in the schema&lt;/li&gt;
&lt;li&gt;adapting generated scenarios to an existing style&lt;/li&gt;
&lt;li&gt;explaining a failed test&lt;/li&gt;
&lt;li&gt;proposing a repair after an API change&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But those workflows should sit on top of deterministic evidence.&lt;/p&gt;

&lt;p&gt;A useful architecture looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenAPI specification
        ↓
Deterministic parser
        ↓
Coverage + schema rules
        ↓
Repository conventions
        ↓
Candidate Karate scenarios
        ↓
Optional AI enhancement
        ↓
Validation / review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key word is &lt;strong&gt;optional&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;You should still be able to generate, execute and analyse tests when an AI provider is unavailable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Validate what was generated
&lt;/h2&gt;

&lt;p&gt;Generation is not finished when a &lt;code&gt;.feature&lt;/code&gt; file exists.&lt;/p&gt;

&lt;p&gt;At minimum, generated output should be checked for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;valid Karate syntax&lt;/li&gt;
&lt;li&gt;valid referenced configuration&lt;/li&gt;
&lt;li&gt;duplicate scenarios&lt;/li&gt;
&lt;li&gt;unresolved variables&lt;/li&gt;
&lt;li&gt;incorrect paths or HTTP methods&lt;/li&gt;
&lt;li&gt;missing required request data&lt;/li&gt;
&lt;li&gt;assertions that are too weak&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last item deserves particular attention.&lt;/p&gt;

&lt;p&gt;This test can pass while proving very little:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="nf"&gt;When &lt;/span&gt;method get
&lt;span class="nf"&gt;Then &lt;/span&gt;status 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Where the contract supports it, stronger assertions might verify:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="nf"&gt;Then &lt;/span&gt;status 200
&lt;span class="nf"&gt;And &lt;/span&gt;match response.id == '#number'
&lt;span class="nf"&gt;And &lt;/span&gt;match response.status == '#string'
&lt;span class="nf"&gt;And &lt;/span&gt;match response.items == '#array'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to maximize assertion count. It is to make sure the test can detect the failures it was created to protect against.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is the workflow I built into Karate Test Management
&lt;/h2&gt;

&lt;p&gt;I originally built &lt;strong&gt;Karate Test Management&lt;/strong&gt;, an open-source VS Code extension, around test generation. As the project grew, the interesting problem became everything around generation.&lt;/p&gt;

&lt;p&gt;The current workflow can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;generate Karate tests from OpenAPI&lt;/li&gt;
&lt;li&gt;import Postman, HAR and GraphQL sources&lt;/li&gt;
&lt;li&gt;index an existing Karate test library&lt;/li&gt;
&lt;li&gt;execute features and individual scenarios&lt;/li&gt;
&lt;li&gt;compare OpenAPI operations with existing test coverage&lt;/li&gt;
&lt;li&gt;identify missing tests&lt;/li&gt;
&lt;li&gt;track quality findings&lt;/li&gt;
&lt;li&gt;analyse failures and flakiness&lt;/li&gt;
&lt;li&gt;optionally use AI for generation and repair&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important architectural decision is that deterministic generation, execution and coverage analysis do not depend on AI.&lt;/p&gt;

&lt;p&gt;AI is there to enhance test-engineering work, not to replace the evidence underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A better mental model
&lt;/h2&gt;

&lt;p&gt;Instead of thinking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;OpenAPI → generated test files&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I find this model more useful:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Discover
   ↓
Understand the contract and repository
   ↓
Measure existing coverage
   ↓
Generate missing scenarios
   ↓
Execute
   ↓
Analyse evidence
   ↓
Maintain
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That turns test generation from a one-off code-generation trick into part of a sustainable API testing workflow.&lt;/p&gt;

&lt;p&gt;If you use Karate and want to experiment with this approach, the project is open source:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/mov2day/KaratePlugin" rel="noopener noreferrer"&gt;https://github.com/mov2day/KaratePlugin&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;VS Code Marketplace: &lt;a href="https://marketplace.visualstudio.com/items?itemName=MuthuKumarKoodalingam.karate-test-generator" rel="noopener noreferrer"&gt;https://marketplace.visualstudio.com/items?itemName=MuthuKumarKoodalingam.karate-test-generator&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The next article in this series will look at a related problem: &lt;strong&gt;your Karate tests may all be green while large parts of your API are still untested.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>githubcopilot</category>
      <category>testing</category>
      <category>api</category>
    </item>
    <item>
      <title>Karate Test Generator for VS Code: Automate API Testing with AI</title>
      <dc:creator>Muthu Kumar Koodalingam</dc:creator>
      <pubDate>Mon, 29 Dec 2025 15:02:11 +0000</pubDate>
      <link>https://dev.to/muthu_kumarkoodalingam/karate-test-generator-for-vs-code-automate-api-testing-with-ai-2lfk</link>
      <guid>https://dev.to/muthu_kumarkoodalingam/karate-test-generator-for-vs-code-automate-api-testing-with-ai-2lfk</guid>
      <description>&lt;h1&gt;
  
  
  🚀 Automate Karate API Testing with This VS Code Extension
&lt;/h1&gt;

&lt;p&gt;Automate &lt;strong&gt;Karate API testing&lt;/strong&gt; directly from &lt;strong&gt;OpenAPI specs&lt;/strong&gt;, keep tests auto-synced with API changes, and enhance coverage using &lt;strong&gt;GitHub Copilot AI&lt;/strong&gt; — saving up to &lt;strong&gt;80% maintenance time&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  📚 Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Why Manual API Testing Fails in 2025&lt;/li&gt;
&lt;li&gt;Introducing Karate API Test Generator&lt;/li&gt;
&lt;li&gt;Key Features &amp;amp; Capabilities&lt;/li&gt;
&lt;li&gt;Real-World Results &amp;amp; Case Studies&lt;/li&gt;
&lt;li&gt;Quick Start Guide&lt;/li&gt;
&lt;li&gt;FAQ&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  ❌ Why Manual API Testing Fails in 2025
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Hidden Cost of Manual API Testing
&lt;/h3&gt;

&lt;p&gt;Are you spending &lt;strong&gt;40% of your sprint fixing broken API tests&lt;/strong&gt;?&lt;/p&gt;

&lt;p&gt;You’re not alone. Surveys show QA teams now spend &lt;strong&gt;more time maintaining tests than validating new features&lt;/strong&gt;. Manual Karate DSL test writing creates four major bottlenecks.&lt;/p&gt;

&lt;h3&gt;
  
  
  🔥 The Manual Testing Crisis
&lt;/h3&gt;

&lt;h4&gt;
  
  
  1. Constant API Changes Break Test Suites
&lt;/h4&gt;

&lt;p&gt;Every endpoint update triggers hours of test fixes. With microservices averaging &lt;strong&gt;20+ API changes per sprint&lt;/strong&gt;, maintenance becomes unsustainable.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Repetitive Boilerplate Crushes Productivity
&lt;/h4&gt;

&lt;p&gt;Writing &lt;code&gt;Given url&lt;/code&gt;, &lt;code&gt;When method&lt;/code&gt;, &lt;code&gt;Then status&lt;/code&gt; hundreds of times isn’t just boring — it’s where inconsistencies and bugs hide.&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Invisible Coverage Gaps
&lt;/h4&gt;

&lt;p&gt;Manual testing misses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;429 rate limits&lt;/li&gt;
&lt;li&gt;500 server errors&lt;/li&gt;
&lt;li&gt;Malformed payloads&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These gaps cost real money when they reach production.&lt;/p&gt;

&lt;h4&gt;
  
  
  4. Team Style Inconsistency
&lt;/h4&gt;

&lt;p&gt;10 developers = 10 styles.&lt;br&gt;
Code reviews become &lt;strong&gt;style debates&lt;/strong&gt;, not logic reviews.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;📉 Real cost:&lt;/strong&gt; Teams report spending &lt;strong&gt;60–70% of QA time&lt;/strong&gt; on test maintenance.&lt;/p&gt;


&lt;h2&gt;
  
  
  🧩 Meet Karate API Test Generator
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fuvxtwgxgrldhjd63jtx7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fuvxtwgxgrldhjd63jtx7.png" alt=" " width="800" height="446"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;VS Code extension&lt;/strong&gt; purpose-built to automate Karate API testing.&lt;/p&gt;

&lt;p&gt;👉 Available on the &lt;strong&gt;&lt;a href="https://marketplace.visualstudio.com/items?itemName=MuthuKumarKoodalingam.karate-test-generator" rel="noopener noreferrer"&gt;VS Code Marketplace&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  What Makes It Different?
&lt;/h3&gt;

&lt;p&gt;Unlike generic testing tools, Karate API Test Generator uniquely combines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;✅ &lt;strong&gt;Native Karate DSL support&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;OpenAPI 3.1 compatibility&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;GitHub Copilot AI enhancement&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;Automatic test sync with API changes&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;Style learning for team consistency&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  ⚙️ Complete Feature Breakdown
&lt;/h2&gt;
&lt;h3&gt;
  
  
  🚀 1. Multi-Source Test Generation
&lt;/h3&gt;
&lt;h4&gt;
  
  
  OpenAPI Support
&lt;/h4&gt;

&lt;p&gt;Supports &lt;strong&gt;OpenAPI 2.0, 3.0, and 3.1&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="c"&gt;# Generated from OpenAPI spec in seconds&lt;/span&gt;
&lt;span class="kd"&gt;Feature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; User Management API Tests

&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Create new user - Happy path
  &lt;span class="nf"&gt;Given &lt;/span&gt;url baseUrl + '/api/v1/users'
  &lt;span class="err"&gt;And request { name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;'John Doe', email&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;'john&lt;/span&gt;&lt;span class="nt"&gt;@example.com'&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;
  &lt;span class="nf"&gt;When &lt;/span&gt;method POST
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 201
  &lt;span class="nf"&gt;And &lt;/span&gt;match response.id == '#number'
  &lt;span class="nf"&gt;And &lt;/span&gt;match response.email == 'john@example.com'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Automatically generates:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All HTTP methods&lt;/li&gt;
&lt;li&gt;Path &amp;amp; query parameters&lt;/li&gt;
&lt;li&gt;Request/response validation&lt;/li&gt;
&lt;li&gt;Auth headers (Bearer, API Key, OAuth)&lt;/li&gt;
&lt;li&gt;Examples from OpenAPI docs&lt;/li&gt;
&lt;li&gt;Error scenarios (400–500)&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Confluence Integration
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Parse tables into test data&lt;/li&gt;
&lt;li&gt;Convert flow diagrams to scenarios&lt;/li&gt;
&lt;li&gt;Extract acceptance criteria&lt;/li&gt;
&lt;li&gt;Link tests to requirements&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Combined Generation
&lt;/h4&gt;

&lt;p&gt;Merge &lt;strong&gt;OpenAPI contracts + business docs&lt;/strong&gt; for full coverage of &lt;strong&gt;API + business logic&lt;/strong&gt;.&lt;/p&gt;




&lt;h3&gt;
  
  
  🤖 2. GitHub Copilot AI Enhancement
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Before AI
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get user by ID
  &lt;span class="nf"&gt;Given &lt;/span&gt;url baseUrl + '/users/1'
  &lt;span class="nf"&gt;When &lt;/span&gt;method GET
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  After AI
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight gherkin"&gt;&lt;code&gt;&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get user by ID - Valid request
  &lt;span class="nf"&gt;Given &lt;/span&gt;url baseUrl + '/users/1'
  &lt;span class="nf"&gt;And &lt;/span&gt;header Authorization = 'Bearer ' + authToken
  &lt;span class="nf"&gt;When &lt;/span&gt;method GET
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 200
  &lt;span class="nf"&gt;And &lt;/span&gt;match response == 
    &lt;span class="err"&gt;{ id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;1, name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;'#string', email&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;'&lt;/span&gt;&lt;span class="c"&gt;#regex ^[\\w-\\.]+@([\\w-]+\\.)+[\\w-]{2,4}$' }&lt;/span&gt;

&lt;span class="kn"&gt;Scenario&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; Get user by ID - Invalid ID returns 404
  &lt;span class="nf"&gt;Given &lt;/span&gt;url baseUrl + '/users/999999'
  &lt;span class="nf"&gt;When &lt;/span&gt;method GET
  &lt;span class="nf"&gt;Then &lt;/span&gt;status 404
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;AI automatically adds:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Realistic test data&lt;/li&gt;
&lt;li&gt;Edge cases&lt;/li&gt;
&lt;li&gt;Security scenarios&lt;/li&gt;
&lt;li&gt;Boundary testing&lt;/li&gt;
&lt;li&gt;Performance considerations&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  🔄 3. Automatic Test Maintenance
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Problem:&lt;/strong&gt; New API field → dozens of broken tests&lt;br&gt;
&lt;strong&gt;Solution:&lt;/strong&gt; Auto-sync with OpenAPI specs&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it works:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Link OpenAPI spec to tests&lt;/li&gt;
&lt;li&gt;Extension watches for changes&lt;/li&gt;
&lt;li&gt;Smart notifications:&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;Update with AI&lt;/li&gt;
&lt;li&gt;Ignore&lt;/li&gt;
&lt;li&gt;Review diff&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;📊 Result:&lt;/strong&gt; &lt;strong&gt;80% reduction&lt;/strong&gt; in maintenance time&lt;/p&gt;




&lt;h3&gt;
  
  
  🎨 4. Style Learning Engine
&lt;/h3&gt;

&lt;p&gt;Learns from existing Karate tests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Indentation &amp;amp; quoting style&lt;/li&gt;
&lt;li&gt;Naming conventions&lt;/li&gt;
&lt;li&gt;Background &amp;amp; tag usage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Applies conventions automatically to all new tests — perfect for large teams.&lt;/p&gt;




&lt;h3&gt;
  
  
  🧭 5. Developer-Friendly UI
&lt;/h3&gt;

&lt;p&gt;No CLI required.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inside VS Code:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dashboard &amp;amp; stats&lt;/li&gt;
&lt;li&gt;OpenAPI import &amp;amp; generation&lt;/li&gt;
&lt;li&gt;Confluence integration&lt;/li&gt;
&lt;li&gt;Auto-sync management&lt;/li&gt;
&lt;li&gt;Template editor&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🏆 Real-World Case Studies
&lt;/h2&gt;

&lt;h3&gt;
  
  
  🏃‍♂️ Startup: 3-Person Team
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Results:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Test generation: &lt;strong&gt;4 hours → 15 minutes&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Maintenance: &lt;strong&gt;60% → 10% of sprint&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Coverage: &lt;strong&gt;45% → 87%&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;“We stopped firefighting broken tests and started testing features.”&lt;br&gt;
— Lead QA Engineer&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  🏢 Enterprise: 50+ Engineers
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Results:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;100% test style consistency&lt;/li&gt;
&lt;li&gt;Code review time: &lt;strong&gt;45 → 15 minutes&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Coverage up &lt;strong&gt;34%&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;3× faster onboarding&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  🔗 Contract Testing at Scale
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Results:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;~80% maintenance reduction&lt;/li&gt;
&lt;li&gt;5× larger contract test suite&lt;/li&gt;
&lt;li&gt;Prevented &lt;strong&gt;$50K+&lt;/strong&gt; in incidents&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  ⚡ Quick Start: 5 Minutes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Install Extension
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Open VS Code&lt;/li&gt;
&lt;li&gt;Extensions → Search &lt;strong&gt;Karate API Test Generator&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Install&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. (Recommended) Install GitHub Copilot
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Enables AI enhancement&lt;/li&gt;
&lt;li&gt;Free for students &amp;amp; OSS maintainers&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Generate Tests
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Method 1:&lt;/strong&gt;&lt;br&gt;
Right-click OpenAPI spec → &lt;strong&gt;Generate Karate Tests Now&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Method 2:&lt;/strong&gt;&lt;br&gt;
Open extension UI → Import spec → Generate&lt;/p&gt;

&lt;p&gt;✅ Tests generated in under &lt;strong&gt;30 seconds&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  🧠 Best Practices
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Run style learning first&lt;/li&gt;
&lt;li&gt;Combine specs + docs&lt;/li&gt;
&lt;li&gt;Review AI-enhanced tests&lt;/li&gt;
&lt;li&gt;Enable auto-sync for stable APIs&lt;/li&gt;
&lt;li&gt;Use shared templates&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Comparison: Karate Test Generator vs. Alternatives
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7q6xsqw1nrwzmluha89t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7q6xsqw1nrwzmluha89t.png" alt=" " width="800" height="306"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  🛣️ Roadmap
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GraphQL support&lt;/li&gt;
&lt;li&gt;Run tests inside VS Code&lt;/li&gt;
&lt;li&gt;Test results dashboard&lt;/li&gt;
&lt;li&gt;Jira integration&lt;/li&gt;
&lt;li&gt;Git-based auto-commits&lt;/li&gt;
&lt;li&gt;SOAP + REST hybrid support&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  ❓ FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does it support Karate 1.x &amp;amp; 2.x?&lt;/strong&gt;&lt;br&gt;
Yes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use it without Copilot?&lt;/strong&gt;&lt;br&gt;
Absolutely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will it replace QA engineers?&lt;/strong&gt;&lt;br&gt;
No — it removes boilerplate so QAs focus on value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is my OpenAPI spec sent externally?&lt;/strong&gt;&lt;br&gt;
No. Everything runs locally.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 Pricing &amp;amp; Support
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Karate API Test Generator:&lt;/strong&gt;&lt;br&gt;
🆓 100% Free (VS Code Marketplace)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub Copilot:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;$10/month (Individual)&lt;/li&gt;
&lt;li&gt;Free for students &amp;amp; OSS maintainers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Support:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;📖 Docs&lt;/li&gt;
&lt;li&gt;💬 Discord&lt;/li&gt;
&lt;li&gt;🐛 GitHub Issues&lt;/li&gt;
&lt;li&gt;✉️ Email&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🧨 The Bottom Line
&lt;/h2&gt;

&lt;p&gt;If you’re still writing Karate tests manually in 2025, you’re wasting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;⏱️ &lt;strong&gt;5–10 hours per sprint per dev&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;💵 &lt;strong&gt;$50K+ annually per team&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;📉 Missing &lt;strong&gt;30–40% edge cases&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;🔄 Fixing tests instead of shipping features&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🎯 Ready to Transform Your API Testing?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Get started in 3 steps:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Install Karate API Test Generator&lt;/li&gt;
&lt;li&gt;Open any OpenAPI spec&lt;/li&gt;
&lt;li&gt;Right-click → Generate Tests&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then enjoy reclaiming &lt;strong&gt;10+ hours per sprint&lt;/strong&gt; 🚀&lt;/p&gt;

</description>
      <category>testing</category>
      <category>vscode</category>
      <category>ai</category>
      <category>githubcopilot</category>
    </item>
  </channel>
</rss>
