DEV Community

Building an OpenAPI Quality Audit Endpoint in Flask — 12-Dimension Spec Scoring

An AI agent can only reliably call a well-documented API. Most "OpenAPI specs" in the wild are missing operationIds, have undocumented 4xx/5xx responses, don't define a reusable error schema, and lack security schemes — which makes them impossible to call programmatically. After auditing a few hundred APIs for the 153-route URL Metadata API, I needed a single endpoint that scores a spec for AI-consumability.

What I built: /api/openapi-quality

A Flask endpoint that accepts either GET ?url=<openapi.json URL> or a POST application/json body with a spec object or a spec_text string (YAML or JSON), then scores the spec across 12 quality dimensions. Each dimension catches a specific failure mode that makes APIs hard for AI agents to consume.

The 12 dimensions

(1) OpenAPI version compliance — top-level openapi: 3.0.x or 3.1.x field present. Catches specs that are really Swagger 2.0 mislabeled, or AsyncAPI/Postman files passed in error.

(2) Info object completeness — 5 sub-checks: info.title, info.version, info.description, info.contact (name/url/email), info.license (name/url). Each missing field loses 3 of 15 points.

(3) OperationId coverage — every operation (get/post/put/delete/patch/options/head/trace) must have a unique operationId. Coverage = operations with ID / total operations. This is THE field an LLM uses to call the API. 95%+ coverage = full points.

(4) Duplicate operationIds — operationId must be unique. Specs with duplicates (e.g. multiple GET /items with the same getItems ID) break LLM tool-calling.

(5) Response documentation — every operation should document a 2xx success, a 4xx client error, and a 5xx server error. Counts documented / expected responses.

(6) Error schema consistency — does components.schemas define an Error/Problem/ErrorResponse schema, AND is it referenced by at least one 4xx/5xx response? This catches specs that use inline ad-hoc error objects everywhere instead of a reusable contract.

(7) Security schemes — components.securitySchemes defined with at least one scheme (apiKey/http/oauth2/openIdConnect). Global security: [...] requirement is a bonus.

(8) Servers — servers[] array present (or Swagger 2.0 host+basePath fallback). Without a base URL, the consumer has no way to know where to call.

(9) Tag grouping with descriptions — operations use tags[], AND root-level tags: [...] defines each tag with a description. Undescribed tags = the OpenAPI docs UI shows no structure.

(10) Request body schema quality — POST/PUT/PATCH operations define requestBody.content.<media-type>.schema (a JSON schema or $ref), not just a string-typed body.

(11) Example presence — at least one example (component-level or in-schema) anywhere in the spec. Without examples, AI agents guess at the shape.

(12) Deprecated Swagger 2.0 fields in 3.x spec — flags host/basePath/schemes/consumes/produces/definitions if present in a openapi: 3.x spec. These are deprecated in OAS 3 and indicate a stale spec.

Scoring

  • 10 points: spec version compliant
  • 5 points: info.title + info.version present
  • 15 points: info completeness (3 per field x 5 fields)
  • 20 points: operationId coverage (full at 95%+, partial below)
  • 5 points: no duplicate operationIds
  • 15 points: response documentation coverage
  • 10 points: reusable error schema
  • 10 points: security (full if schemes + global requirement)
  • 5 points: servers present
  • 5 points: tag definitions with descriptions
  • 5 points: request body schema quality
  • 5 points: example presence
  • -5 points: deprecated fields

Total: 100. A-F grade bands at 90/75/60/40.

Why this matters for AI agents

When an LLM is given an OpenAPI spec and asked to call the API:

  • No operationId = the LLM has to invent names -> unreliable tool calls
  • No documented 4xx = the LLM can't tell if a 401 means "retry" or "give up"
  • No error schema = every error response is a different shape -> LLM has to parse 5 different JSON structures
  • No security scheme = the LLM doesn't know it needs to send a Bearer token
  • No servers[] = the LLM has no base URL to call

The endpoint is live at $0.0005 per call via x402 on Base mainnet — POST a spec, get a 0-100 score + per-dimension breakdown in ~50ms.

curl "https://periodically-february-medieval-responsibility.trycloudflare.com/api/openapi-quality?url=https://petstore3.openapitools.org/openapi.json" \
  -H "X-PAYMENT: <base64-encoded-payment-payload>"
Enter fullscreen mode Exit fullscreen mode

Returns a JSON object with quality_score, grade, spec_format, operations_count, operation_ids.coverage_pct, info_completeness.score, error_schema.has_reusable_error_schema, findings[] (specific issues), and per-dimension sub-scores. The full source is in a 36K-line Flask service — happy to share patterns if you're building something similar.

Top comments (0)