I spent a day building a Yandex SEO MCP server: wiring Yandex Webmaster into the Model Context Protocol so an AI assistant could read the Yandex side of organic search. The tools themselves are unremarkable: list sites, site summary, query stats, indexing history.
Getting there cost six surprises, most of which the English documentation either omits or states incorrectly. This is the writeup I wanted before I started.
1. The scope you need is not the one the names suggest
Yandex Webmaster has two OAuth scopes:
-
webmaster:hostinfo, which reads like the read scope -
webmaster:verify, which reads like a write scope for adding and verifying sites
I requested hostinfo alone. Least privilege, every report is a read, a token that cannot touch a user's properties is a strictly smaller blast radius. Site listing worked. Every data endpoint returned this:
403 ACCESS_FORBIDDEN
Access to this resource is not allowed with scopes available for this application.
Required scope: COMMON, application scopes: [ALL_SCOPES, HOST_LIST, EXTERNAL_LINKS]
webmaster:hostinfo buys the site list and nothing else. Summary, queries, indexing, site quality index: all of them need webmaster:verify. There is no read-only configuration of this API that returns data.
Two things follow. First, if you promise users read-only access in your UI copy, you cannot. The grant that reads their numbers is the same grant that can add and verify sites, and your consent screen will say so. Say it honestly instead.
Second, scope changes are not retroactive. Widening the scope on the OAuth app does nothing for tokens already issued: existing users must disconnect and reconnect. Plan for that migration before you ship, not after.
The scope strings themselves appear only in the Russian documentation. The English pages never print them. When the two disagree, the Russian pages are the source and the English ones are a stale translation. The English terms page still claims you need a partner agreement with Yandex to use the API; the Russian version of the same document says to register an app and get a client id, and registration is self-service.
2. The app type is permanent, and the documented happy path is the wrong one
Yandex ID offers two application types, chosen at creation, immutable afterwards:
| Type | Constraint |
|---|---|
| For user login on a third-party service | Free choice of redirect_uri, capped at three permission groups |
| For API access / debugging |
redirect_uri locked to https://oauth.yandex.ru/verification_code, cannot be edited |
Yandex's own walkthrough points you at the second one, with response_type=token. That is the single-developer recipe: grab a token for yourself and paste it into a script.
A multi-tenant integration needs the first one, with response_type=code and a callback on your own domain. The API-access type structurally cannot do it, and you cannot convert the app afterwards. Pick wrong and you register again from scratch.
3. Authorization: OAuth, not Bearer, and the token response lies to you
The token endpoint returns this:
{
"access_token": "...",
"refresh_token": "...",
"token_type": "bearer",
"expires_in": 31536000
}
token_type: "bearer". Send Authorization: Bearer <token> to the Webmaster API and it rejects you. The correct header is:
Authorization: OAuth <token>
Ignore token_type. The payload actively points you at the wrong header.
4. The refresh token dies with the access token
This is the one with operational consequences.
With Google, a refresh token effectively lives forever: a connection can sit idle for a year and still come back. Yandex states plainly that the refresh token's lifetime is the same as the OAuth token's. When a grant expires, the refresh token expired at the same instant, and there is no recovery path. The user must authorise again by hand.
So refreshing lazily on the next request is not a strategy, it is a countdown. You need a proactive sweep:
// Renew everything expiring inside the margin, daily.
// The margin is deliberately wide (7 days) so a cron outage,
// a deploy, or a bad night still leaves several chances before
// a grant becomes unrecoverable.
const cutoff = Math.floor(Date.now() / 1000) + REFRESH_MARGIN_SECONDS
Two more details from the docs that bite later: a refresh returns a new refresh token, so persist it on every refresh and not just the first; and if two refreshers race, they invalidate each other's token. Pick one owner for renewal. In my case the app owns it and the MCP server deliberately does not refresh at all, which also keeps the OAuth client secret out of the MCP deployment.
5. The dates do work, whatever the docs imply
The documentation describes the popular-queries pool as the top 3,000 queries of the past week. I read that as "date_from and date_to select inside a rolling week and cannot reach further back", wrote it into my caveat copy, and left the date filter off the queries report because it would have been decoration.
Then I measured it. Same property, three windows, nothing else changed:
| Requested range | Queries returned | Impressions for the top query | Range echoed back |
|---|---|---|---|
| default (no dates) | 10 | 15 | 7 days |
| 28 days | 25 | 52 | 28 days, as asked |
| 86 days | 32 | 62 | 86 days, as asked |
The dates select a real window. A longer range returns more queries and larger totals, and the requested start date comes back unchanged. Omitting the dates gives you the last week, which is probably where the documentation's framing comes from.
What the API does clamp is the end of the range. A request ending on the 9th came back ending on the 6th: Yandex finishes a day roughly four days late, against two or three for Search Console. Which leads to a rule worth coding: render the range the response reports, never the range you requested. The response carries date_from and date_to for exactly this reason.
Measure the API. The docs are a hypothesis.
6. Paging, and the subtle bug hiding inside it
Yandex serves at most 500 rows per request and holds up to 3,000. One call therefore returns the top sixth of what is available.
The missing rows are the obvious problem. The subtle one is worse: those 500 are the top 500 by order_by, which defaults to impressions. Render them in a sortable table and a user clicking "sort by clicks" is ranking a set that was selected by impressions. The sort is honest about the rows it has and silent about the ones it does not.
So walk the pool:
for (let fetched = 0; fetched < budget; ) {
const pageSize = Math.min(PAGE_SIZE, budget - fetched)
let data
try {
data = await fetchPage(startOffset + fetched, pageSize)
} catch (err) {
// A failure on page one is fatal: there is nothing to report.
// A failure on page four is not: keep what we have and let
// totalCount/truncated describe the shortfall honestly.
if (rows.length === 0) throw err
break
}
const page = parse(data)
rows.push(...page)
fetched += page.length
// Short page means the pool is exhausted.
if (page.length < pageSize) break
if (totalCount > 0 && startOffset + fetched >= totalCount) break
}
Two things worth copying. The loop stops on a short page, so a site with 300 queries costs one request and not six. And a failure partway through keeps the rows already collected rather than throwing the whole report away, because the response already reports how much of the pool is missing.
Keep the calls serial. Yandex publishes no rate-limit numbers, only a 429 contract, so there is nothing to parallelise against.
Two smaller things that will cost you an hour each
Host ids are opaque. A Yandex host id looks like https:www.gscwizard.com:443. You cannot construct it from a URL; you list the account's hosts and match. Yandex also resolves www/non-www and http/https mirrors itself and tells you which one is canonical, so follow the mirror it nominates rather than the one that matched.
Dates come in two spellings, one of which JavaScript rejects. Several endpoints return 2016-01-01T00:00:00,000+0300, with a comma before the milliseconds. new Date() returns Invalid Date on that, silently, in the middle of a chart. Normalise before parsing.
The MCP-specific bit: ship the limits with the data
Everything above is specific to the Yandex Webmaster MCP work. This part is not.
Yandex's data comes with a lot of "you cannot conclude that from this": totals are floors, there is no query-by-page join, CTR is derived rather than reported, positions are not comparable with Google's, there is no previous-period series to compare against. All of that normally lives in documentation, which the model never reads.
So every tool response carries them, as data:
{
"siteUrl": "sc-domain:example.com",
"range": { "startDate": "2026-09-30", "endDate": "2026-10-06" },
"totalAvailable": 10,
"truncated": false,
"limitations": [
"Yandex returns only its top queries for the period, at most 3,000 of them...",
"Yandex measures position differently from Google and splits it in two...",
"Yandex has no query-by-page report, so its queries cannot be attributed to URLs..."
],
"queries": [ ... ]
}
The strings live in one module, which the app's report pages render and the MCP server vendors, so a dashboard and an assistant answering about the same numbers give the same qualifications. One definition, two surfaces, no drift.
It costs a few hundred bytes per response and it changes the answers. An assistant that can see "this is a floor" in the payload says so. One that cannot will tell your user they got 62 impressions, full stop, and be wrong.
If you are building data tools for models, this is the cheapest correctness win available: stop documenting what your numbers do not mean, and start returning it.
Checkout the full integration by trying out GSC Wizard


Top comments (0)