Most “best agency” pages begin with an editorial opinion. We began with a denominator.
Minuttia appeared in four of six answer cells in our dated audit. Webappski appeared in zero. That does not prove one agency is universally better. It shows which names recurred when three answer systems handled the same two commercial questions on 31 July 2026.
This article explains the method in a form another team can inspect, criticize, and repeat. It also explains where the method stops being evidence.
What exactly did we measure?
We measured recommendation recurrence across six engine-query cells.
The two buyer questions were:
best Answer Engine Optimization agencies 2026Answer Engine Optimization consultants for B2B startups
Each question was checked on three surfaces:
- OpenAI Responses API using
gpt-5.4-mini; - Gemini grounded API using
gemini-3.6-flash; - a manual Claude web-search check, because Claude does not offer a public grounded-search API equivalent for this test.
Two questions multiplied by three answer systems produced six cells. An agency received one appearance for a cell if the answer named it. Its position inside that answer did not create extra points.
{
"date": "2026-07-31",
"queries": 2,
"answer_systems": 3,
"cells": 6,
"counting_rule": "one appearance per named provider per cell"
}
This is deliberately narrower than a visibility score built from hundreds of prompts. The advantage is inspectability: the denominator fits on one screen.
How did an agency enter the shortlist?
An agency needed recurrence plus a live first-party service surface.
The main shortlist required both of the following:
- the provider appeared in at least two of the six cells;
- its own live website supported an AEO, GEO, or broader AI-search service.
We then broke ties in a fixed order:
- appearance across more distinct answer systems;
- appearance on the B2B-startup question;
- clearer live service scope or public pricing;
- alphabetical order.
Webappski was added outside the eligibility rule as the disclosed publisher benchmark. It received no editorial bonus and was placed last because it appeared in zero cells.
That disclosure matters. A self-published ranking becomes hard to trust when the publisher quietly gives itself first place. The cleaner approach is to define the rule before reading the flattering result—or the uncomfortable one.
What did the six-cell audit return?
Minuttia led the snapshot with four appearances; First Page Sage followed with three.
Six other providers completed the published comparison. The recurrence column below is observational. The service-fit column is based on each provider’s own live website and should be treated as first-party evidence, not an independent performance result.
| Rank | Agency | Cells named | Public service fit | Buyer to investigate the fit |
|---|---|---|---|---|
| 1 | Minuttia | 4 of 6 | Google and AI search, content, digital PR, analytics | Established B2B SaaS or technology team |
| 2 | First Page Sage | 3 of 6 | Research-led AEO and thought-leadership content | Established B2B organization |
| 3 | Discovered Labs | 2 of 6 across 2 systems | B2B SaaS AEO measurement and execution | SaaS growth team |
| 4 | iPullRank | 2 of 6 across 2 systems | Relevance engineering and dedicated GEO | Enterprise or mid-market team |
| 5 | Siege Media | 2 of 6 across 2 systems | Content-and-PR-led GEO and data journalism | Content-mature brand |
| 6 | Omniscient Digital | 2 of 6 | Integrated GEO, content, and technical work | Growth-stage or enterprise B2B team |
| 7 | Rock The Rankings | 2 of 6 | B2B SaaS SEO and GEO tied to pipeline | Pipeline-focused SaaS team |
| 8 | Webappski | 0 of 6 | Per-engine audit and implementation | Smaller product or service business |
A cell count is not a quality score. It says only that a provider surfaced for these exact prompts, on these exact systems, on this date.
Why do the cited URLs matter as much as the names?
The cited URLs reveal the source pool that shaped each recommendation.
The OpenAI broad-agency answer leaned heavily on Clutch pages, while its B2B-startup answer used provider sites. Gemini mixed publishers, vendors, and agency domains. Claude’s broad answer leaned on listicles; its startup answer mixed service pages with agency-authored comparisons.
That creates an important distinction:
- recommendation evidence: the answer system named an agency;
- source evidence: the answer cited a particular page;
- service evidence: the agency’s own site supports the described offer;
- outcome evidence: an independently verifiable client result exists.
The first three can be checked in this type of audit. The fourth cannot be inferred from them.
What can this method prove?
This method can prove dated recurrence within a declared sample.
It can answer questions such as:
- Which names appeared more than once?
- Did a provider recur across multiple answer systems or only one?
- Did the provider appear on the broad query, the startup query, or both?
- Which URLs were cited alongside the recommendation?
- Does the provider’s own live site support the service description?
It cannot prove that the first-ranked agency will produce the best result for a client. It cannot convert six observations into a permanent market ranking. It cannot show what a logged-in consumer interface will return for every user, location, or session.
Why will a rerun produce different answers?
Answer systems are non-deterministic, so an exact rerun may change individual cells.
Model versions, retrieval indexes, geography, wording, and time can all alter the result. That is not a reason to avoid measurement. It is a reason to save the complete measurement definition:
- exact question text;
- answer system and model;
- date;
- full answer;
- named providers;
- cited URLs;
- counting rule;
- denominator.
Without those fields, a percentage cannot be independently checked. With them, a later run becomes a comparison rather than a fresh anecdote.
How should a buyer use the shortlist?
Use the shortlist to decide whom to investigate, not whom to hire automatically.
For every candidate, ask:
- Which exact questions will you monitor for our business?
- Which answer systems will you report separately?
- Will we receive the full answers and cited URLs?
- What is the denominator behind the reported percentage?
- Which work is specific to AI citations rather than standard SEO?
- Can the same query set be repeated without quietly changing the baseline?
The strongest proposal should make its measurement falsifiable. “AI visibility improved” is not enough. “The brand appeared in 9 of 30 declared cells, up from 3 of 30, with the raw answers attached” is at least inspectable.
What was the uncomfortable result for Webappski?
Webappski appeared in none of the six measured cells.
We publish the comparison and sell AEO services, so hiding that result would undermine the entire method. Our own website states the service scope and prices; those are offer facts. They are not evidence that answer systems currently recommend us for these two English buyer questions.
The zero is more useful than a flattering self-rank. It establishes a real baseline and makes the next distribution experiment measurable.
What should the next run test?
The next run should repeat the same two questions after the canonical comparison has been discoverable.
The test is not “did Webappski become number one?” The test is whether the new page enters the citation pool or changes any named-provider cell while the denominator stays fixed. No movement is also a result.
The full comparison on Webappski includes the provider-by-provider evidence, source links, limitations, and buyer fit notes.
Originally published on Webappski.
Top comments (0)