DEV Community

Ozan Dikici for Sitelemetry

Posted on Fully Autonomous

Website audits as MCP tools: OAuth sign-in and what one finding has to contain

The problem

After a coding agent ships a site, the checks that should follow are scattered across separate tools: headers and TLS in one, technical SEO in another, then accessibility, performance and analytics tags. Each has its own report, usually written for a person. The agent that will make the change sees none of it unless I paste a summary in.

I'm Ozan. I built Sitelemetry to put those checks where the agent works. It audits sites you own or are authorized to test, in six areas: security, technical SEO, AI visibility, accessibility, performance and integrations (analytics and tag setup).

The intended loop

  1. Audit. The agent calls an audit tool against the site.
  2. Fix list. The tool result carries the evidence and the remediation for every finding, so the agent works from that list.
  3. Change. The agent edits the code or configuration.
  4. Re-audit. The same audit runs again on the same target. A finding that is missing from the new list is fixed; one that is still there is not.

The audit tools are read-only: they report, and the agent that has the repository open makes the change. Step 4 is a new run, not a summary of what was edited.

Why OAuth for the MCP connection, not an API key in a config file

The audits are seven tools on a remote MCP server over Streamable HTTP: audit_security, audit_seo, audit_ai_visibility, audit_accessibility, audit_performance, audit_integrations and audit_full.

They are MCP tools because the agent is the one that needs the output. Findings arrive as a tool result, so the evidence and the remediation text are already in context when it starts the change.

The usual alternative is an API key in a JSON config file. I did not want that. Those files get committed, synced between machines and pasted into bug reports, and the key is a long-lived secret in plain text.

So authorization is OAuth 2.1 with PKCE (S256) and dynamic client registration. You configure one URL. In Cursor's mcp.json, for example:

{
  "mcpServers": {
    "sitelemetry": {
      "url": "https://sitelemetry.com/mcp"
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

That is the whole entry. The client reads /.well-known/oauth-protected-resource and /.well-known/oauth-authorization-server, registers itself, and requests one scope, audit. The config file holds a URL and nothing secret. The cost is that the client has to support this flow.

Long audits

A long audit returns status: running with a jobId. The client calls the same tool again with that jobId until the result is final. Polling does not start a second scan.

What one finding looks like

This is a real finding from audit_security, passive profile, run against my own production site on 2026-09-30:

severity:    medium
title:       CSP permits inline styles
evidence:    Affected location: https://sitelemetry.com/ Content-Security-Policy: style-src: 'unsafe-inline'; hardening points: 13/22 (deficit: 9).
impact:      Inline style permissions weaken restrictions on injected CSS and page appearance.
remediation: Remove 'unsafe-inline' from the reported style directives by moving inline styles into approved stylesheets; test in Report-Only mode first.
Enter fullscreen mode Exit fullscreen mode

It is not fixed yet, so I have no before and after to show.

Give an agent only "medium: CSP" and it has to guess what is wrong and where. Each field has a job:

  • Evidence is what was observed: the URL, the header and the directive. The agent does not have to rediscover the problem.
  • Impact is one sentence on why it matters, so I can decide whether to do it now.
  • Remediation is the change to make, including the caution to test in Report-Only mode first. A CSP edit can break a page's styling, so that clause belongs in the finding.
  • Severity orders the work.

Passing checks carry evidence too. In the same run, "HSTS active" came back with max-age=31536000; includeSubDomains and "Response compression enabled" with content-encoding: gzip; body bytes 109605 → 26480. A pass without the observed value is only an assertion.

A check that did not run is not a pass

That run returned a security score of 96, grade A, with 17 passing checks and 0 critical, 0 high, 1 medium, 3 low and 2 info findings. It also returned coverage status "partial", because the passive profile does not run active checks such as ports or exposure paths. Each of those modules says so:

module: ports
status: unavailable
reason: The selected Passive profile does not run this active check.
Enter fullscreen mode Exit fullscreen mode

An agent that sees no port findings is likely to conclude the ports are fine. So the status is reported per module, with the reason written as a sentence. The 96 describes the checks that ran, and "partial" says that was not everything.

Limits

  • It is not a pentest.
  • A score describes what the checks returned on that run, not whether the site is secure.
  • AI visibility means readiness for AI crawlers and answers, not tracking of chatbot mentions.
  • It is a closed-source hosted service. The public repository holds client manifests only.
  • This post covers only the MCP server. Also live: a GitHub Action, a GitLab CI/CD component, an n8n node, a WordPress plugin, a Chrome extension, and a loopback-only Local Agent for localhost targets.
  • The snapshot needs no signup. The Free plan is $0 and covers security checks only, 10 scans per month. The other five areas and scheduled audits are on paid plans, from $49/month.
  • I build and support it myself.
  • My own site still has the medium finding and three low ones open: no CAA record, no MTA-STS policy, and server: nginx disclosed.

Links

Questions for you

  1. When coverage is partial, is a per-module unavailable with a reason enough, or would you withhold or cap the score?
  2. If you have connected remote MCP servers that use OAuth with dynamic client registration, which clients failed, and at which step?
  3. Would you rather have the agent work through the whole finding list in one pass, or one finding at a time so each change can be committed and re-audited on its own?

Top comments (1)

Collapse
 
marcusykim profile image
Marcus Kim •

The way you split 'evidence' and 'remediation' in the CSP finding is spot-on - the agent doesn't have to guess what's broken. I noticed the 'impact' field gives a clear why it matters without jargon. For the partial coverage, I'd want the score to stay at 96 instead of dropping to 90 when some checks aren't run - that's a real-world tradeoff between completeness and actionable feedback.