<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jarno S</title>
    <description>The latest articles on DEV Community by Jarno S (@jarnosaarimies).</description>
    <link>https://dev.to/jarnosaarimies</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4083706%2Ff35d75f1-cc6d-40b0-ab6d-0a05cb9bb4bb.jpg</url>
      <title>DEV Community: Jarno S</title>
      <link>https://dev.to/jarnosaarimies</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jarnosaarimies"/>
    <language>en</language>
    <item>
      <title>An Entity Factsheet for AI Search: A Practical Site-Wide QA Checklist</title>
      <dc:creator>Jarno S</dc:creator>
      <pubDate>Wed, 16 Sep 2026 07:56:22 +0000</pubDate>
      <link>https://dev.to/jarnosaarimies/an-entity-factsheet-for-ai-search-a-practical-site-wide-qa-checklist-5e7m</link>
      <guid>https://dev.to/jarnosaarimies/an-entity-factsheet-for-ai-search-a-practical-site-wide-qa-checklist-5e7m</guid>
      <description>&lt;p&gt;When an organization has a website, social profiles, and articles, its identity can still be surprisingly hard to piece together. Names vary. Old descriptions remain online. Service pages make claims that the About page never explains.&lt;/p&gt;

&lt;p&gt;A simple entity factsheet can help. It is not a magic AI-search ranking factor. It is an internal source of truth that makes public information more consistent and easier to verify.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the canonical facts
&lt;/h2&gt;

&lt;p&gt;Create a short document containing the exact public name, primary domain, location at the appropriate level, what the organization does, who is responsible for the content, and the official profile URLs. Add a date and an owner for each fact.&lt;/p&gt;

&lt;p&gt;For a solo expert, keep the person's identity and the project or company identity distinct. A founder's professional experience is not automatically a client outcome for the company.&lt;/p&gt;

&lt;h2&gt;
  
  
  Audit every public touchpoint
&lt;/h2&gt;

&lt;p&gt;Compare the factsheet with the homepage, About page, service pages, contact page, author pages, social profiles, directory listings, and guest articles.&lt;/p&gt;

&lt;p&gt;Look for contradictions that a reader would notice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the company name spelled the same way?&lt;/li&gt;
&lt;li&gt;Does the website link point to the preferred HTTPS domain?&lt;/li&gt;
&lt;li&gt;Are old offerings still described as current?&lt;/li&gt;
&lt;li&gt;Is the same person named as author where appropriate?&lt;/li&gt;
&lt;li&gt;Are location and contact details accurate?&lt;/li&gt;
&lt;li&gt;Do profiles imply credentials, clients, or results that cannot be substantiated?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fixing inconsistencies is more useful than creating more empty profiles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give each page a clear job
&lt;/h2&gt;

&lt;p&gt;The homepage should explain what the organization does and who it serves. An About page should establish identity and responsibility. A service page should explain scope, process, deliverables, and limitations. An article should answer a specific question with enough context to stand alone.&lt;/p&gt;

&lt;p&gt;Do not place every keyword on every page. Instead, connect relevant pages with descriptive internal links so a reader can follow the evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use structured data carefully
&lt;/h2&gt;

&lt;p&gt;Organization, Person, Article, and other schema types can describe visible facts. Validate the markup and make sure it matches the page. Structured data should not claim reviews, ratings, awards, or services that are absent from the visible content.&lt;/p&gt;

&lt;p&gt;Think of schema as a machine-readable description of a clear page, not a substitute for one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Maintain a change log
&lt;/h2&gt;

&lt;p&gt;If the name, domain, services, or author details change, update the factsheet first. Then review the most important pages and profiles. A quarterly check can catch stale information before it spreads into more places.&lt;/p&gt;

&lt;p&gt;For AI-answer monitoring, note the change date. If a system later describes the organization incorrectly, you can compare its answer with the current canonical facts and investigate which source may be outdated. That is a better diagnosis than simply calling the answer a hallucination.&lt;/p&gt;

&lt;h2&gt;
  
  
  A ten-minute check
&lt;/h2&gt;

&lt;p&gt;Pick one public question: "What does this organization do?" Try to answer it from the homepage, About page, a service page, and one external profile. If you get four materially different answers, the next task is consistency—not more content volume.&lt;/p&gt;

&lt;p&gt;I use this evidence-first approach in my work on AI-search visibility at &lt;a href="https://aeovara.fi/" rel="noopener noreferrer"&gt;AEOvara&lt;/a&gt;. Clear entity information supports readers and crawlers, but it does not guarantee inclusion in an AI-generated answer.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>documentation</category>
      <category>seo</category>
    </item>
    <item>
      <title>Missing Is Not Zero: A Data Quality Checklist for AI Visibility Tests</title>
      <dc:creator>Jarno S</dc:creator>
      <pubDate>Fri, 04 Sep 2026 13:38:18 +0000</pubDate>
      <link>https://dev.to/jarnosaarimies/missing-is-not-zero-a-data-quality-checklist-for-ai-visibility-tests-50ob</link>
      <guid>https://dev.to/jarnosaarimies/missing-is-not-zero-a-data-quality-checklist-for-ai-visibility-tests-50ob</guid>
      <description>&lt;p&gt;Your dashboard says a brand appeared in 25% of AI answers. Before interpreting that number, ask a less exciting question: what counted as an answer?&lt;/p&gt;

&lt;p&gt;If a test failed, timed out or never reached the intended question, recording a zero turns a collection problem into an apparent visibility problem.&lt;/p&gt;

&lt;p&gt;Here is a practical way to keep those outcomes separate.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Separate absence from missing information
&lt;/h2&gt;

&lt;p&gt;A completed answer that does not mention the brand is a valid non-mention.&lt;/p&gt;

&lt;p&gt;An interrupted session with no assessable answer is missing data.&lt;/p&gt;

&lt;p&gt;Store them differently:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;false&lt;/code&gt;: the outcome was assessed and absent.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;true&lt;/code&gt;: the outcome was assessed and present.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;null&lt;/code&gt;: the outcome could not be assessed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This distinction should apply separately to brand mentions, owned-domain citations and explicit recommendations. An answer might be scorable for mentions while its citations remain inaccessible.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Define completeness before seeing the result
&lt;/h2&gt;

&lt;p&gt;Write down what makes an observation usable before collecting answers.&lt;/p&gt;

&lt;p&gt;For example: the system answered the intended question, and the relevant output was preserved for review.&lt;/p&gt;

&lt;p&gt;Record refusals separately. Whether they belong in a particular denominator depends on what the metric is intended to measure. Preserve them and disclose the rule.&lt;/p&gt;

&lt;p&gt;Do not decide that an answer is “incomplete” simply because the brand is missing.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Report coverage beside performance
&lt;/h2&gt;

&lt;p&gt;Consider this hypothetical batch—not an actual client result:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Eight attempts were planned.&lt;/li&gt;
&lt;li&gt;Six produced complete, scorable answers.&lt;/li&gt;
&lt;li&gt;Two of those six mentioned the brand.&lt;/li&gt;
&lt;li&gt;Two attempts failed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The mention rate among complete answers is 2/6, or 33.3%.&lt;/p&gt;

&lt;p&gt;Collection coverage is 6/8, or 75%.&lt;/p&gt;

&lt;p&gt;Reporting only 2/8 treats failures as known non-mentions. Reporting only 33.3% hides the missing observations.&lt;/p&gt;

&lt;p&gt;Show both. Also explain if failures cluster around particular questions, because the available answers may not represent the planned test set.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Preserve reruns
&lt;/h2&gt;

&lt;p&gt;A rerun should create another observation, not erase the first attempt.&lt;/p&gt;

&lt;p&gt;Set a rule in advance: for example, retry a technical failure once, but do not rerun a valid answer simply because it omitted the brand.&lt;/p&gt;

&lt;p&gt;Selecting the most flattering answer from several attempts measures something different from following a fixed test procedure.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Test the scoring rules
&lt;/h2&gt;

&lt;p&gt;Check these edge cases before trusting your dashboard:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A third-party source mentions the brand, but the brand’s website is not cited.&lt;/li&gt;
&lt;li&gt;The brand appears in a negative comparison.&lt;/li&gt;
&lt;li&gt;A similarly named business appears.&lt;/li&gt;
&lt;li&gt;A domain is printed as plain text without being used as a supporting source.&lt;/li&gt;
&lt;li&gt;An interrupted run produces no assessable answer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A mention, citation and recommendation should not automatically receive the same label.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the score auditable
&lt;/h2&gt;

&lt;p&gt;Preserve the raw response, prompt version, timestamp, visible model label, interface and retrieval settings when observable. Record “unknown” instead of guessing.&lt;/p&gt;

&lt;p&gt;The useful output is not just a percentage. It is a percentage another reviewer can reconstruct.&lt;/p&gt;

&lt;p&gt;I work on &lt;a href="https://aeovara.fi/" rel="noopener noreferrer"&gt;AEOvara&lt;/a&gt;, a Finnish AEO and AI-search visibility project. This proposed checklist extends my earlier &lt;a href="https://dev.to/jarnosaarimies/how-to-build-a-repeatable-ai-search-visibility-benchmark-1e4m"&gt;repeatable benchmark guide&lt;/a&gt; with data-quality checks. It does not establish that any content change causes more AI recommendations.&lt;/p&gt;

&lt;p&gt;What edge case would you add to the scoring tests?&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to Build a Repeatable AI Search Visibility Benchmark</title>
      <dc:creator>Jarno S</dc:creator>
      <pubDate>Tue, 18 Aug 2026 17:56:17 +0000</pubDate>
      <link>https://dev.to/jarnosaarimies/how-to-build-a-repeatable-ai-search-visibility-benchmark-1e4m</link>
      <guid>https://dev.to/jarnosaarimies/how-to-build-a-repeatable-ai-search-visibility-benchmark-1e4m</guid>
      <description>&lt;p&gt;Checking whether an AI assistant mentions your company once is interesting, but it is not a measurement system. Answers vary between models, sessions, dates, and prompt wording. A useful benchmark needs a stable set of questions, a consistent scoring method, and a record of what changed.&lt;/p&gt;

&lt;p&gt;This article describes a lightweight approach that a small team can run without an enterprise monitoring platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Start with decisions, not keywords
&lt;/h2&gt;

&lt;p&gt;Traditional rank tracking starts with search terms. AI-search monitoring should start with the decisions a potential customer asks an assistant to help make.&lt;/p&gt;

&lt;p&gt;Build prompts from four groups:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Category discovery: “What are the best ways to solve X?”&lt;/li&gt;
&lt;li&gt;Provider discovery: “Which companies help with X in Finland?”&lt;/li&gt;
&lt;li&gt;Comparison: “What is the difference between approach A and approach B?”&lt;/li&gt;
&lt;li&gt;Trust and validation: “How can I evaluate whether a provider is credible?”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep the set small enough to run consistently. Twenty carefully selected prompts are more useful than two hundred prompts that change every month.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Freeze a golden prompt set
&lt;/h2&gt;

&lt;p&gt;A golden prompt set is a versioned list of prompts that stays stable between measurement rounds. Each row should include at least:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;prompt_id, journey_stage, prompt_text, language, market, expected_entities
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Do not silently rewrite prompts after a run. If a prompt needs to change, create a new version. Otherwise, an apparent visibility improvement may only be the result of easier wording.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Capture more than mentions
&lt;/h2&gt;

&lt;p&gt;A binary “mentioned / not mentioned” score misses most of the useful information. For each response, record:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;run_date
model
prompt_id
brand_mentioned
mention_position
recommendation_strength
linked_or_cited
cited_domains
answer_summary
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Recommendation strength can use a simple scale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;0: not mentioned&lt;/li&gt;
&lt;li&gt;1: mentioned incidentally&lt;/li&gt;
&lt;li&gt;2: included as a relevant option&lt;/li&gt;
&lt;li&gt;3: clearly recommended or used as a primary example&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep the raw answer as evidence. Scores alone are difficult to audit later.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Separate visibility from citation
&lt;/h2&gt;

&lt;p&gt;These are related but different outcomes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Visibility: the brand or entity appears in the answer.&lt;/li&gt;
&lt;li&gt;Citation: the assistant links to or names a source that supports the answer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A company can be visible without its own website being cited. A website can also be cited while the brand is not recommended. Track both rates separately.&lt;/p&gt;

&lt;p&gt;For a prompt set with N prompts:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mention_rate = prompts_with_brand_mention / N
citation_rate = prompts_citing_owned_domain / N
recommendation_rate = prompts_with_strength_2_or_3 / N
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;These simple metrics are easy to explain and hard to manipulate.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Control the test conditions
&lt;/h2&gt;

&lt;p&gt;Perfect repeatability is not possible, but avoid unnecessary variation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Run the same prompts in the same order.&lt;/li&gt;
&lt;li&gt;Record the model and date.&lt;/li&gt;
&lt;li&gt;Use a fresh conversation for each prompt.&lt;/li&gt;
&lt;li&gt;Keep language and market context consistent.&lt;/li&gt;
&lt;li&gt;Avoid follow-up questions in the benchmark run.&lt;/li&gt;
&lt;li&gt;Run on a regular schedule, such as once a month.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is not laboratory certainty. The goal is enough consistency to distinguish a durable trend from a lucky answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Review the evidence behind changes
&lt;/h2&gt;

&lt;p&gt;When a score moves, inspect the responses before drawing conclusions.&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Did the assistant discover a new page or third-party source?&lt;/li&gt;
&lt;li&gt;Did a competitor disappear rather than our brand improve?&lt;/li&gt;
&lt;li&gt;Did the answer cite stronger evidence than last month?&lt;/li&gt;
&lt;li&gt;Did the model misunderstand the category?&lt;/li&gt;
&lt;li&gt;Is the change visible across several prompts or only one?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This review turns monitoring into an action list. For example, repeated category confusion suggests clearer entity and service descriptions. Missing evidence may suggest publishing a primary source, case study, methodology, or data page.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Report trends, not isolated wins
&lt;/h2&gt;

&lt;p&gt;A useful monthly report can fit on one page:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Mention rate&lt;/li&gt;
&lt;li&gt;Citation rate&lt;/li&gt;
&lt;li&gt;Strong recommendation rate&lt;/li&gt;
&lt;li&gt;Most frequently cited domains&lt;/li&gt;
&lt;li&gt;Prompts with the largest positive and negative changes&lt;/li&gt;
&lt;li&gt;Three evidence-backed actions for the next month&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Avoid presenting a single flattering answer as proof of authority. A stronger signal is repeated appearance across relevant prompts, models, and measurement rounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical takeaway
&lt;/h2&gt;

&lt;p&gt;AI-search visibility is noisy, but it is measurable. The key is to freeze the questions, preserve the raw evidence, score consistently, and investigate why the answers changed.&lt;/p&gt;

&lt;p&gt;I use this type of repeatable measurement thinking while developing &lt;a href="https://aeovara.fi/" rel="noopener noreferrer"&gt;AEOvara&lt;/a&gt;, a Finnish AI-search visibility and AEO project. The method itself is deliberately tool-independent: a spreadsheet and disciplined process are enough to begin.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>ai</category>
      <category>llm</category>
      <category>analytics</category>
    </item>
  </channel>
</rss>
