<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Albina Blazhko</title>
    <description>The latest articles on DEV Community by Albina Blazhko (@albinator).</description>
    <link>https://dev.to/albinator</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4108203%2F892dbeb8-a67b-4738-b836-c1ca5a706d66.jpg</url>
      <title>DEV Community: Albina Blazhko</title>
      <link>https://dev.to/albinator</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/albinator"/>
    <language>en</language>
    <item>
      <title>OpenAPI linters: which one is the best in a real test?</title>
      <dc:creator>Albina Blazhko</dc:creator>
      <pubDate>Thu, 08 Oct 2026 18:44:01 +0000</pubDate>
      <link>https://dev.to/albinator/openapi-linters-which-one-is-the-best-in-a-real-test-1n2j</link>
      <guid>https://dev.to/albinator/openapi-linters-which-one-is-the-best-in-a-real-test-1n2j</guid>
      <description>&lt;p&gt;Every few months somebody posts a comparison of OpenAPI linters. It is usually one number per tool, with no spec name, no versions and no machine, and nobody reruns it. I wanted numbers I could defend. So I took the API descriptions of products most of us use every day, or have at least heard of, and ran the five linters that get compared most often on them: Redocly CLI, Vacuum, Spectral, Scalar CLI, and Speakeasy CLI.&lt;/p&gt;

&lt;p&gt;The setup is simple: latest release of each tool, default rules, one GitHub-hosted runner, one tool at a time, five runs each, median reported. The commands, the runner specs and every raw output are in the public repo.&lt;/p&gt;

&lt;h3&gt;
  
  
  Here come the true numbers
&lt;/h3&gt;

&lt;p&gt;On Stripe's API description, the most popular one in the set, Redocly CLI finished first at 1.26 s, Vacuum second at 2 s and Speakeasy third at 3.3 s. Spectral came last, far behind at 12 s. Scalar CLI crashed.&lt;/p&gt;

&lt;p&gt;The GitHub, AWS EC2 and Cloudflare API descriptions give the same order: Redocly CLI first and Vacuum second. Speakeasy is third on GitHub and Cloudflare and fourth on AWS EC2, after Spectral. Cloudflare is the largest description in the set, with 3,663 operations, three times as many as GitHub, and Redocly CLI lints it in 5 s. I could have stopped there and called Redocly CLI the fastest linter, measured on the most popular spec around.&lt;/p&gt;

&lt;p&gt;Then I looked at the specs where Redocly CLI is not first, and the results are worth a closer look.&lt;/p&gt;

&lt;p&gt;On DigitalOcean, Speakeasy CLI is first. It does not follow the references DigitalOcean puts on its operations, so it never opens those files. It reports the operations as broken and finishes in under a second, because it did almost nothing. When I bundled the DigitalOcean description, Speakeasy finally read every operation, and the gap closed. Redocly CLI and Speakeasy finished almost together. And Speakeasy's number still leaves out the time it takes to bundle the description first. Redocly CLI needs no extra step: it lints the split files as published and resolves the references itself.&lt;/p&gt;

&lt;p&gt;Azure is the boring case. The description is small, and the difference there is Node's startup time, nothing else.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0kjbaznjxraniu1rb0c0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0kjbaznjxraniu1rb0c0.png" alt="Lint time against API size for the five linters" width="800" height="453"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Redocly CLI draws the flattest line and the spec size affects it the least. Vacuum wins on Azure, and Speakeasy wins on DigitalOcean, where it skips the referenced files. Scalar CLI does not finish four of the six specs.&lt;/p&gt;

&lt;h3&gt;
  
  
  The layout that teams actually work in
&lt;/h3&gt;

&lt;p&gt;Look at how these specs are published and you see two shapes. Some come split into many files. Others are bundled into one file just for publishing. Both shapes come from real production flows. Teams split their descriptions for governance: nobody edits a 100,000-line YAML file by hand, a pull request that touches one endpoint is impossible to review, two writers can't work on the file at the same time, and you can't give one team ownership of one part of the description while it all lives in one file. To publish, they bundle it back into a single file, so anyone can read the API without cloning the repository.&lt;/p&gt;

&lt;p&gt;So I did both directions to simulate the real flow: I split the descriptions that come bundled, and I bundled the ones that come split.&lt;/p&gt;

&lt;p&gt;The first experiment was Stripe. I split the bundled spec with &lt;code&gt;redocly split&lt;/code&gt; and ran the linters on the result.&lt;/p&gt;

&lt;p&gt;Here are the raw numbers.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
  &lt;tbody&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Time (median)&lt;/th&gt;
&lt;th&gt;Errors&lt;/th&gt;
&lt;th&gt;Warnings + info&lt;/th&gt;
&lt;/tr&gt;
  &lt;tr&gt;
&lt;td&gt;Redocly CLI&lt;/td&gt;
&lt;td&gt;1.12 s&lt;/td&gt;
&lt;td&gt;641&lt;/td&gt;
&lt;td&gt;1,012&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
&lt;td&gt;Speakeasy CLI&lt;/td&gt;
&lt;td&gt;3.17 s&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;4,345&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
&lt;td&gt;Vacuum&lt;/td&gt;
&lt;td colspan="3"&gt;❌ did not finish in 5 minutes&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
&lt;td&gt;Spectral&lt;/td&gt;
&lt;td colspan="3"&gt;❌ did not finish in 5 minutes&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
&lt;td&gt;Scalar CLI&lt;/td&gt;
&lt;td colspan="3"&gt;❌ did not finish in 5 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two tools finished. Redocly CLI got a bit faster and reported the same 641 errors and 1,012 warnings as before. Speakeasy kept its errors and warnings too, but lost about 1,500 hints, because it doesn't check every referenced schema.&lt;/p&gt;

&lt;p&gt;Vacuum printed its banner, announced 55 rules, and then did nothing: no CPU, no output, until the workflow killed it after five minutes. Spectral did the opposite: it kept a CPU core busy and never finished. Scalar hung like Vacuum, because its lint command runs a bundled copy of Vacuum.&lt;/p&gt;

&lt;p&gt;DigitalOcean publishes its API description as split files. I bundled it into one file with &lt;code&gt;redocly bundle&lt;/code&gt;, and the reports changed a lot. The chart below compares the counts in the two layouts.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkb4w3b35o90pgf4g4ue5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkb4w3b35o90pgf4g4ue5.png" alt="Errors, warnings and info added up on the DigitalOcean API, split files against the bundled file, for the five linters" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Redocly CLI was the one tool whose report did not change between the two layouts. Vacuum's warnings more than doubled, and Scalar CLI's, which runs an older Vacuum, tripled. Spectral had read the operations all along, so the bundle only removed its 1,320 errors about the invalid reference. Speakeasy's count went down, because in the split files it never opened the operation references and reported each one as broken. On the bundle those false errors went away, from more than 1,000 to 2.&lt;/p&gt;

&lt;p&gt;You might ask why I reached for Redocly CLI for splitting and bundling the spec. Because it was the only tool that gave me files I could use.&lt;/p&gt;

&lt;p&gt;Of the five tools, two can split a description: Redocly CLI and Scalar. Scalar turned the Stripe file into 2,051 JSON files that break the OpenAPI specification: a &lt;code&gt;$ref&lt;/code&gt; on every operation, pointers to a components section that isn't there, &lt;code&gt;$global: true&lt;/code&gt; next to every &lt;code&gt;$ref&lt;/code&gt;, and its own &lt;code&gt;__scalar_&lt;/code&gt; property in nearly 2,000 places, which other linters report. It also silently upgraded the description from OpenAPI 3.0.0 to 3.1.1. Scalar CLI itself cannot resolve references in the files it just produced.&lt;/p&gt;

&lt;p&gt;Bundling was the same story. Three can bundle: Redocly CLI, Vacuum and Scalar. Vacuum's default mode writes a large 28 MB file full of duplicated blocks that still has unresolved references, and its composed mode is invalid too and drops &lt;code&gt;responses&lt;/code&gt;, &lt;code&gt;summary&lt;/code&gt;, &lt;code&gt;tags&lt;/code&gt; and &lt;code&gt;parameters&lt;/code&gt; from every operation. Scalar keeps a &lt;code&gt;$ref&lt;/code&gt; on every operation, so Speakeasy still can't read the result. It also moves the content of all 2,852 referenced files under an &lt;code&gt;x-ext&lt;/code&gt; key with hash names and leaves &lt;code&gt;components.schemas&lt;/code&gt; empty, which makes the file hard to read and maintain.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the numbers in the "errors" column actually mean
&lt;/h3&gt;

&lt;p&gt;The other thing these comparisons get wrong is reading the error count as a quality score. Each tool ships its own default ruleset, and the counts measure different things. The raw outputs are linked from &lt;a href="https://github.com/AlbinaBlazhko17/openapi-lint-benchmark#stripe-single-file" rel="noopener noreferrer"&gt;every row in the repo&lt;/a&gt;, and this is what is in them for Stripe.&lt;/p&gt;

&lt;p&gt;Vacuum's 2 errors are circular &lt;code&gt;$ref&lt;/code&gt; chains, which OpenAPI permits. Its 23,000 warnings and infos are almost all style rules, useful if you want prettier docs, but none of them says the document is invalid. Spectral reports 0 errors and 600 warnings, 594 of them operations without tags. Speakeasy's 2 errors are enum values that collide after normalization, which matters if you generate SDKs, and that is what Speakeasy is for.&lt;/p&gt;

&lt;p&gt;Redocly CLI is the only tool that reports a problem written into the &lt;a href="https://spec.openapis.org/oas/v3.0.3#schema-object" rel="noopener noreferrer"&gt;OpenAPI specification&lt;/a&gt;: almost all of its errors are schemas where &lt;code&gt;nullable: true&lt;/code&gt; has no effect, because the schema has no &lt;code&gt;type&lt;/code&gt;. The rest are operations without a &lt;code&gt;summary&lt;/code&gt; and server URLs with a trailing slash. The warnings are about missing 4xx responses, &lt;code&gt;allOf&lt;/code&gt;/&lt;code&gt;oneOf&lt;/code&gt; combinations that can't match anything, and one missing license.&lt;/p&gt;

&lt;p&gt;The DigitalOcean chart in the previous section makes the same point from the other side. Speakeasy reported more than 1,000 errors there, almost all of them because it couldn't resolve a reference written in a way the specification doesn't allow. Point it at the bundled file and the count drops to 2. The number measures what the tool managed to read, not how good the API description is.&lt;/p&gt;

&lt;p&gt;So "2 errors" and "641 errors" don't compare. One tool checks style and mostly skips validity, the other found 622 instances of one real problem. You have to look at the rule behind the number.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where this leaves us
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/AlbinaBlazhko17/openapi-lint-benchmark" rel="noopener noreferrer"&gt;The repo is public&lt;/a&gt;: pinned spec commits, recorded tool versions, a machine nobody owns, the commands in the workflow file, and every raw output committed next to the tables. The workflow reruns everything on demand. If you think the setup is wrong, change it and run it.&lt;/p&gt;

&lt;p&gt;Scalar CLI was the least stable tool in this test. A new release came out while I worked on this article, and it made Scalar CLI slower and less stable: it now crashes on Stripe and times out on three more descriptions.&lt;/p&gt;

&lt;p&gt;The performance numbers are not the whole story. Some linters look fast only because they skipped references or rules, and on a large split description most of them need a bundling step before every run, which raises the overall linting time even more.&lt;/p&gt;

&lt;p&gt;In production you need a tool you can rely on: one that finishes on every layout, keeps your pipeline short, and gives the same valid answer whether the API description is one file or multiple files. In this benchmark that tool is Redocly CLI. That is what the numbers say.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: I'm a Redocly CLI maintainer&lt;/em&gt;&lt;/p&gt;

</description>
      <category>openapi</category>
      <category>tooling</category>
      <category>api</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
