An AI agent can only reliably call a well-documented API. Most "OpenAPI specs" in the wild are missing operationIds, have undocumented 4xx/5xx responses, don't define a reusable error schema, and lack security schemes — which makes them impossible to call programmatically. After auditing a few hundred APIs for the 153-route URL Metadata API, I needed a single endpoint that scores a spec for AI-consumability.
What I built: /api/openapi-quality
A Flask endpoint that accepts either GET ?url=<openapi.json URL> or a POST application/json body with a spec object or a spec_text string (YAML or JSON), then scores the spec across 12 quality dimensions. Each dimension catches a specific failure mode that makes APIs hard for AI agents to consume.
The 12 dimensions
(1) OpenAPI version compliance — top-level openapi: 3.0.x or 3.1.x field present. Catches specs that are really Swagger 2.0 mislabeled, or AsyncAPI/Postman files passed in error.
(2) Info object completeness — 5 sub-checks: info.title, info.version, info.description, info.contact (name/url/email), info.license (name/url). Each missing field loses 3 of 15 points.
(3) OperationId coverage — every operation (get/post/put/delete/patch/options/head/trace) must have a unique operationId. Coverage = operations with ID / total operations. This is THE field an LLM uses to call the API. 95%+ coverage = full points.
(4) Duplicate operationIds — operationId must be unique. Specs with duplicates (e.g. multiple GET /items with the same getItems ID) break LLM tool-calling.
(5) Response documentation — every operation should document a 2xx success, a 4xx client error, and a 5xx server error. Counts documented / expected responses.
(6) Error schema consistency — does components.schemas define an Error/Problem/ErrorResponse schema, AND is it referenced by at least one 4xx/5xx response? This catches specs that use inline ad-hoc error objects everywhere instead of a reusable contract.
(7) Security schemes — components.securitySchemes defined with at least one scheme (apiKey/http/oauth2/openIdConnect). Global security: [...] requirement is a bonus.
(8) Servers — servers[] array present (or Swagger 2.0 host+basePath fallback). Without a base URL, the consumer has no way to know where to call.
(9) Tag grouping with descriptions — operations use tags[], AND root-level tags: [...] defines each tag with a description. Undescribed tags = the OpenAPI docs UI shows no structure.
(10) Request body schema quality — POST/PUT/PATCH operations define requestBody.content.<media-type>.schema (a JSON schema or $ref), not just a string-typed body.
(11) Example presence — at least one example (component-level or in-schema) anywhere in the spec. Without examples, AI agents guess at the shape.
(12) Deprecated Swagger 2.0 fields in 3.x spec — flags host/basePath/schemes/consumes/produces/definitions if present in a openapi: 3.x spec. These are deprecated in OAS 3 and indicate a stale spec.
Scoring
- 10 points: spec version compliant
- 5 points:
info.title+info.versionpresent - 15 points: info completeness (3 per field x 5 fields)
- 20 points: operationId coverage (full at 95%+, partial below)
- 5 points: no duplicate operationIds
- 15 points: response documentation coverage
- 10 points: reusable error schema
- 10 points: security (full if schemes + global requirement)
- 5 points: servers present
- 5 points: tag definitions with descriptions
- 5 points: request body schema quality
- 5 points: example presence
- -5 points: deprecated fields
Total: 100. A-F grade bands at 90/75/60/40.
Why this matters for AI agents
When an LLM is given an OpenAPI spec and asked to call the API:
- No operationId = the LLM has to invent names -> unreliable tool calls
- No documented 4xx = the LLM can't tell if a 401 means "retry" or "give up"
- No error schema = every error response is a different shape -> LLM has to parse 5 different JSON structures
- No security scheme = the LLM doesn't know it needs to send a Bearer token
-
No
servers[]= the LLM has no base URL to call
The endpoint is live at $0.0005 per call via x402 on Base mainnet — POST a spec, get a 0-100 score + per-dimension breakdown in ~50ms.
curl "https://periodically-february-medieval-responsibility.trycloudflare.com/api/openapi-quality?url=https://petstore3.openapitools.org/openapi.json" \
-H "X-PAYMENT: <base64-encoded-payment-payload>"
Returns a JSON object with quality_score, grade, spec_format, operations_count, operation_ids.coverage_pct, info_completeness.score, error_schema.has_reusable_error_schema, findings[] (specific issues), and per-dimension sub-scores. The full source is in a 36K-line Flask service — happy to share patterns if you're building something similar.
Top comments (0)