DEV Community

Cover image for OpenAPI linters: which one is the best in a real test?
Albina Blazhko
Albina Blazhko

Posted on

OpenAPI linters: which one is the best in a real test?

Every few months somebody posts a comparison of OpenAPI linters. It is usually one number per tool, with no spec name, no versions and no machine, and nobody reruns it. I wanted numbers I could defend. So I took the API descriptions of products most of us use every day, or have at least heard of, and ran the five linters that get compared most often on them: Redocly CLI, Vacuum, Spectral, Scalar CLI, and Speakeasy CLI.

The setup is simple: latest release of each tool, default rules, one GitHub-hosted runner, one tool at a time, five runs each, median reported. The commands, the runner specs and every raw output are in the public repo.

Here come the true numbers

On Stripe's API description, the most popular one in the set, Redocly CLI finished first at 1.26 s, Vacuum second at 2 s and Speakeasy third at 3.3 s. Spectral came last, far behind at 12 s. Scalar CLI crashed.

The GitHub, AWS EC2 and Cloudflare API descriptions give the same order: Redocly CLI first and Vacuum second. Speakeasy is third on GitHub and Cloudflare and fourth on AWS EC2, after Spectral. Cloudflare is the largest description in the set, with 3,663 operations, three times as many as GitHub, and Redocly CLI lints it in 5 s. I could have stopped there and called Redocly CLI the fastest linter, measured on the most popular spec around.

Then I looked at the specs where Redocly CLI is not first, and the results are worth a closer look.

On DigitalOcean, Speakeasy CLI is first. It does not follow the references DigitalOcean puts on its operations, so it never opens those files. It reports the operations as broken and finishes in under a second, because it did almost nothing. When I bundled the DigitalOcean description, Speakeasy finally read every operation, and the gap closed. Redocly CLI and Speakeasy finished almost together. And Speakeasy's number still leaves out the time it takes to bundle the description first. Redocly CLI needs no extra step: it lints the split files as published and resolves the references itself.

Azure is the boring case. The description is small, and the difference there is Node's startup time, nothing else.

Lint time against API size for the five linters

Redocly CLI draws the flattest line and the spec size affects it the least. Vacuum wins on Azure, and Speakeasy wins on DigitalOcean, where it skips the referenced files. Scalar CLI does not finish four of the six specs.

The layout that teams actually work in

Look at how these specs are published and you see two shapes. Some come split into many files. Others are bundled into one file just for publishing. Both shapes come from real production flows. Teams split their descriptions for governance: nobody edits a 100,000-line YAML file by hand, a pull request that touches one endpoint is impossible to review, two writers can't work on the file at the same time, and you can't give one team ownership of one part of the description while it all lives in one file. To publish, they bundle it back into a single file, so anyone can read the API without cloning the repository.

So I did both directions to simulate the real flow: I split the descriptions that come bundled, and I bundled the ones that come split.

The first experiment was Stripe. I split the bundled spec with redocly split and ran the linters on the result.

Here are the raw numbers.

Tool Time (median) Errors Warnings + info
Redocly CLI 1.12 s 641 1,012
Speakeasy CLI 3.17 s 2 4,345
Vacuum ❌ did not finish in 5 minutes
Spectral ❌ did not finish in 5 minutes
Scalar CLI ❌ did not finish in 5 minutes

Two tools finished. Redocly CLI got a bit faster and reported the same 641 errors and 1,012 warnings as before. Speakeasy kept its errors and warnings too, but lost about 1,500 hints, because it doesn't check every referenced schema.

Vacuum printed its banner, announced 55 rules, and then did nothing: no CPU, no output, until the workflow killed it after five minutes. Spectral did the opposite: it kept a CPU core busy and never finished. Scalar hung like Vacuum, because its lint command runs a bundled copy of Vacuum.

DigitalOcean publishes its API description as split files. I bundled it into one file with redocly bundle, and the reports changed a lot. The chart below compares the counts in the two layouts.

Errors, warnings and info added up on the DigitalOcean API, split files against the bundled file, for the five linters

Redocly CLI was the one tool whose report did not change between the two layouts. Vacuum's warnings more than doubled, and Scalar CLI's, which runs an older Vacuum, tripled. Spectral had read the operations all along, so the bundle only removed its 1,320 errors about the invalid reference. Speakeasy's count went down, because in the split files it never opened the operation references and reported each one as broken. On the bundle those false errors went away, from more than 1,000 to 2.

You might ask why I reached for Redocly CLI for splitting and bundling the spec. Because it was the only tool that gave me files I could use.

Of the five tools, two can split a description: Redocly CLI and Scalar. Scalar turned the Stripe file into 2,051 JSON files that break the OpenAPI specification: a $ref on every operation, pointers to a components section that isn't there, $global: true next to every $ref, and its own __scalar_ property in nearly 2,000 places, which other linters report. It also silently upgraded the description from OpenAPI 3.0.0 to 3.1.1. Scalar CLI itself cannot resolve references in the files it just produced.

Bundling was the same story. Three can bundle: Redocly CLI, Vacuum and Scalar. Vacuum's default mode writes a large 28 MB file full of duplicated blocks that still has unresolved references, and its composed mode is invalid too and drops responses, summary, tags and parameters from every operation. Scalar keeps a $ref on every operation, so Speakeasy still can't read the result. It also moves the content of all 2,852 referenced files under an x-ext key with hash names and leaves components.schemas empty, which makes the file hard to read and maintain.

What the numbers in the "errors" column actually mean

The other thing these comparisons get wrong is reading the error count as a quality score. Each tool ships its own default ruleset, and the counts measure different things. The raw outputs are linked from every row in the repo, and this is what is in them for Stripe.

Vacuum's 2 errors are circular $ref chains, which OpenAPI permits. Its 23,000 warnings and infos are almost all style rules, useful if you want prettier docs, but none of them says the document is invalid. Spectral reports 0 errors and 600 warnings, 594 of them operations without tags. Speakeasy's 2 errors are enum values that collide after normalization, which matters if you generate SDKs, and that is what Speakeasy is for.

Redocly CLI is the only tool that reports a problem written into the OpenAPI specification: almost all of its errors are schemas where nullable: true has no effect, because the schema has no type. The rest are operations without a summary and server URLs with a trailing slash. The warnings are about missing 4xx responses, allOf/oneOf combinations that can't match anything, and one missing license.

The DigitalOcean chart in the previous section makes the same point from the other side. Speakeasy reported more than 1,000 errors there, almost all of them because it couldn't resolve a reference written in a way the specification doesn't allow. Point it at the bundled file and the count drops to 2. The number measures what the tool managed to read, not how good the API description is.

So "2 errors" and "641 errors" don't compare. One tool checks style and mostly skips validity, the other found 622 instances of one real problem. You have to look at the rule behind the number.

Where this leaves us

The repo is public: pinned spec commits, recorded tool versions, a machine nobody owns, the commands in the workflow file, and every raw output committed next to the tables. The workflow reruns everything on demand. If you think the setup is wrong, change it and run it.

Scalar CLI was the least stable tool in this test. A new release came out while I worked on this article, and it made Scalar CLI slower and less stable: it now crashes on Stripe and times out on three more descriptions.

The performance numbers are not the whole story. Some linters look fast only because they skipped references or rules, and on a large split description most of them need a bundling step before every run, which raises the overall linting time even more.

In production you need a tool you can rely on: one that finishes on every layout, keeps your pipeline short, and gives the same valid answer whether the API description is one file or multiple files. In this benchmark that tool is Redocly CLI. That is what the numbers say.

Disclosure: I'm a Redocly CLI maintainer

Top comments (2)

Collapse
 
yuliiakhm profile image
yuliia-khm •

Really interesting comparison! I especially liked the point that faster doesn't always mean better if a linter skips references or important checks. It's great to see real-world benchmarks with reproducible results instead of just performance claims. Curious how these tools would compare with custom rulesets in a production CI/CD pipeline!💻️

Collapse
 
kanoru profile image
Viktor •

Nice benchmark. What stood out to me is that Redocly CLI was the only tool whose report didn’t change between split and bundled layouts. That’s the part people usually don’t test. Did you look into why Spectral and Vacuum never finished on split Stripe? Is it the ref resolution itself, or do they re-parse the referenced files on every lookup?