<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aniruddha Garje</title>
    <description>The latest articles on DEV Community by Aniruddha Garje (@garje).</description>
    <link>https://dev.to/garje</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172431%2F2ee74ce3-0d4d-468c-a2e4-68160967389b.png</url>
      <title>DEV Community: Aniruddha Garje</title>
      <link>https://dev.to/garje</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/garje"/>
    <language>en</language>
    <item>
      <title>I searched Google Hotels for a city that does not exist. It found five hotels.</title>
      <dc:creator>Aniruddha Garje</dc:creator>
      <pubDate>Fri, 09 Oct 2026 13:44:10 +0000</pubDate>
      <link>https://dev.to/garje/i-searched-google-hotels-for-a-city-that-does-not-exist-it-found-five-hotels-2el</link>
      <guid>https://dev.to/garje/i-searched-google-hotels-for-a-city-that-does-not-exist-it-found-five-hotels-2el</guid>
      <description>&lt;p&gt;I typed "hotels in Qwzxvbnmplk" into my Google Hotels scraper to watch it fail. It did not fail. It came back with five hotels near San Antonio, Texas, each with a price, a rating and a map position.&lt;/p&gt;

&lt;p&gt;This is the story of that answer and the fix. It is also how my fix broke a real capital city, and the small check I now run before I trust any search result.&lt;/p&gt;

&lt;h2&gt;
  
  
  A test that was supposed to fail
&lt;/h2&gt;

&lt;p&gt;On 8 October I was testing inputs that should go wrong. One of them was a search for a city that does not exist. I expected an error, or an empty table.&lt;/p&gt;

&lt;p&gt;Here is what the run returned (run pLMLM09FNCR1f0Tft):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;Price per night&lt;/th&gt;
&lt;th&gt;Rating&lt;/th&gt;
&lt;th&gt;Reviews&lt;/th&gt;
&lt;th&gt;Latitude, longitude&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Flat with pool &amp;amp; barbeque, Pet-friendly&lt;/td&gt;
&lt;td&gt;$75.77&lt;/td&gt;
&lt;td&gt;4.665&lt;/td&gt;
&lt;td&gt;49&lt;/td&gt;
&lt;td&gt;29.404, -98.626&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Holiday Inn San Antonio Seaworld by IHG&lt;/td&gt;
&lt;td&gt;$83.09&lt;/td&gt;
&lt;td&gt;4.4&lt;/td&gt;
&lt;td&gt;2,198&lt;/td&gt;
&lt;td&gt;29.448, -98.680&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Texas Sage One Room Cabin&lt;/td&gt;
&lt;td&gt;$113.00&lt;/td&gt;
&lt;td&gt;4.725&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;29.777, -98.950&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Luxurious Canyon Lake glamping with hot tub + pool&lt;/td&gt;
&lt;td&gt;$254.00&lt;/td&gt;
&lt;td&gt;4.4&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;29.869, -98.302&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apartment with 1 bedroom for 2 persons&lt;/td&gt;
&lt;td&gt;$64.79&lt;/td&gt;
&lt;td&gt;4.805&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;td&gt;29.475, -98.536&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every field was filled. Every price was a real price for a real place. My scraper charges per priced hotel, so each of those five rows would have been a paid result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why nothing looked wrong
&lt;/h2&gt;

&lt;p&gt;The same run had two other searches. "hotels in Paris" returned hotels in Paris. "hotels in Springfield" returned five hotels around 39.75, -89.6, which is Springfield, Illinois. Google picked one of the many Springfields and did not say so.&lt;/p&gt;

&lt;p&gt;That is the pattern. A search engine always tries to answer. When it cannot match your words to a place, it still shows you hotels somewhere. I do not know why it chose San Antonio.&lt;/p&gt;

&lt;p&gt;My scraper checked the shape of the answer. Prices were numbers, ratings were between 1 and 5, every row had coordinates. All of that was true. None of it said whether the answer belonged to the question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The signal was there all along
&lt;/h2&gt;

&lt;p&gt;Google's results page carries the place it understood. For Paris, the page held a place ID and the name "Paris". For Springfield, a place ID and "Springfield".&lt;/p&gt;

&lt;p&gt;For "hotels in Qwzxvbnmplk" it held nothing. The run's own statistics record the place for the nonsense search as null. The source had told me it did not understand. I was not reading that part.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;I changed the order of work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Resolve the place first. If Google returns no place, skip the search, charge nothing and name it in the run's status message.&lt;/li&gt;
&lt;li&gt;Put the place Google understood on every row, in a field called &lt;code&gt;resolvedPlace&lt;/code&gt;, so you can see which Springfield you got.&lt;/li&gt;
&lt;li&gt;If the resolved place shares no word with the search, log a warning. Only a warning: it never drops results.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The next run said: "Finished with 10 priced hotels. Not recognised by Google, so skipped and not charged: "hotels in Qwzxvbnmplk"." (run 5jrjtwZ72Z9SFK7tY).&lt;/p&gt;

&lt;h2&gt;
  
  
  Then it skipped Prague
&lt;/h2&gt;

&lt;p&gt;Later that day another run finished with this message: "Finished with 0 priced hotels. Not recognised by Google, so skipped and not charged: "hotels in Prague"." (run YTSa6YvgqYp6QrGc5).&lt;/p&gt;

&lt;p&gt;Google had sent the page for Prague without its place data. My new check did exactly what I told it to. It treated a real capital city like a made-up one.&lt;/p&gt;

&lt;p&gt;The second fix: when the place is missing, retry the page up to 3 times on a new IP before skipping the search. The next run resolved Prague and still skipped the nonsense search (run WiNWSzkUnEGFirBjx).&lt;/p&gt;

&lt;p&gt;A strict check needs a way to be wrong safely. Mine had none until Prague showed me.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check you can use tomorrow
&lt;/h2&gt;

&lt;p&gt;This is not my scraper's code. It is the idea in a form you can paste into anything that asks a source a question and gets a confident answer back. Think of a geocoder, a search API, an entity lookup or a tool call from a language model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;words&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\w+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()))&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_answer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resolved&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;resolved is whatever the source says it understood, for example a place name.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;resolved&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;skip: the source did not say what it understood&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;words&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="nf"&gt;words&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resolved&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;warn: the source may have answered a different question&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ok&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;check_answer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hotels in Qwzxvbnmplk&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;          &lt;span class="c1"&gt;# skip
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;check_answer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hotels in Springfield&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Springfield&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;  &lt;span class="c1"&gt;# ok, but which Springfield?
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;check_answer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hotels in Paris&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Paris&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;              &lt;span class="c1"&gt;# ok
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two rules hide in those lines. Before you count rows, check that the source echoed your question. And keep what it understood next to every row, so a person can see it later.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same bug in different clothes
&lt;/h2&gt;

&lt;p&gt;On 9 October, while setting up public examples, I ran one that asked for 2 hotels with each booking site's offer. The run delivered 2 hotels and 478 offers (run 3BgiBhR7DZifI3D2P).&lt;/p&gt;

&lt;p&gt;The scraper had fetched offers for every hotel it kept on the page, not only for the two it delivered. Each offer is a paid event. Again every row was complete and plausible, and the total was wrong.&lt;/p&gt;

&lt;p&gt;After the fix, the same kind of request gave 2 hotels and 49 offers (run BNCMJr19AnUWUQUt4). The check here is simple: compare what you bill against what you delivered, every run.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you test a source that disagrees with itself?
&lt;/h2&gt;

&lt;p&gt;After this, I wanted a number for how often my hotel results match what a person sees. So I had a real browser read the live Google Hotels pages, ran the scraper straight after, and compared the two.&lt;/p&gt;

&lt;p&gt;The first surprise was the browser. Two browser visits 12 minutes apart had only 87% of their hotels in common. Google reshuffles its list between visits.&lt;/p&gt;

&lt;p&gt;A test that demands 95% agreement with a source that agrees with itself 87% of the time will fail forever. Worse, it teaches you to ignore it. So my pass mark for the hotel list is the browser's own agreement, 87.2%, minus 3 points: 84.2%.&lt;/p&gt;

&lt;p&gt;Prices get a strict bar, because a price should not drift between two looks at the same hotel. For hotels that appeared in both, the price matched within 10% in 99% of cases.&lt;/p&gt;

&lt;p&gt;Separate what the source keeps changing from what it should never get wrong, and test each one against its own bar.&lt;/p&gt;

&lt;h2&gt;
  
  
  The alarm that fired because the data got better
&lt;/h2&gt;

&lt;p&gt;On 9 October my daily check on Paris hotels failed. It watches how often each field comes back empty, and alarms when that share moves more than 5 points from a saved baseline.&lt;/p&gt;

&lt;p&gt;What moved was good news. Fewer hotels than usual were missing a star class: 15%, against 30% in the baseline (run J3MVimtp32Xg6WCHp). On 20 hotels, 5 points is a single hotel, and the list reshuffles between visits.&lt;/p&gt;

&lt;p&gt;I changed the alarm to fire only when empty values rise, and only when 3 of the 20 hotels change. The rerun passed (run pql2nTclNeiU8o8sn). A check that cries wolf gets switched off, and then it protects nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I still cannot know
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Google's hotel list changes between visits. In my test, two browser visits 12 minutes apart had 87% of their hotels in common. "Correct" for a list means "matches what Google showed at that moment", not a fixed truth.&lt;/li&gt;
&lt;li&gt;My warning is a word check. A search for a landmark can resolve to a city whose name shares no word with it, and then the warning fires for nothing. That is why it only warns.&lt;/li&gt;
&lt;li&gt;I can tell you which Springfield Google picked. I cannot tell you whether it is the one you meant. Adding the state to the search, as in "hotels in Springfield, Massachusetts", is what fixes that.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Over to you
&lt;/h2&gt;

&lt;p&gt;Where in your stack does a source answer confidently when it did not understand the question? I would like to hear the strangest example you have found. Reply below.&lt;/p&gt;

&lt;p&gt;These checks now live in the Google Hotels scraper I built on Apify: every row carries &lt;code&gt;resolvedPlace&lt;/code&gt;, and a search Google does not recognise is skipped and not charged. It is here if you want to see them work: &lt;a href="https://apify.com/garje/google-hotels-scraper?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=story&amp;amp;utm_content=city-that-does-not-exist" rel="noopener noreferrer"&gt;https://apify.com/garje/google-hotels-scraper?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=story&amp;amp;utm_content=city-that-does-not-exist&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Fact or number&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"hotels in Qwzxvbnmplk" returned 5 hotels near San Antonio, with the names, prices, ratings, review counts and coordinates in the table&lt;/td&gt;
&lt;td&gt;Run pLMLM09FNCR1f0Tft (8 Oct 2026, 08:46 UTC, build 0.1.9), dataset rows; reports/fixes/FIXES.md item (b)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Each priced hotel is a paid result&lt;/td&gt;
&lt;td&gt;claims.csv row 70 (pricing); reports/fixes/FIXES.md item (c)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"hotels in Springfield" returned hotels around 39.75, -89.6&lt;/td&gt;
&lt;td&gt;Run pLMLM09FNCR1f0Tft dataset rows (coordinates of Springfield, Illinois)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Place recorded as null for the nonsense search; place IDs and names for Paris and Springfield&lt;/td&gt;
&lt;td&gt;Run pLMLM09FNCR1f0Tft, RUN_STATS.places (reports/fixes/before.json)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;resolvedPlace on every row; unrecognised search skipped, not charged, named in the status message; warning when no word is shared&lt;/td&gt;
&lt;td&gt;claims.csv rows 64 and 66 (row 66 UNVERIFIED for the log message wording, so the article only says it logs a warning, as in the code); reports/fixes/FIXES.md item (b); google-hotels-scraper/src/main.py place_matches&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Status message after the fix, 10 priced hotels&lt;/td&gt;
&lt;td&gt;Run 5jrjtwZ72Z9SFK7tY (build 0.1.10)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prague skipped, 0 priced hotels&lt;/td&gt;
&lt;td&gt;Run YTSa6YvgqYp6QrGc5 (8 Oct 2026, 09:18 UTC, build 0.1.10)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retry up to 3 times on a new IP; Prague then resolved&lt;/td&gt;
&lt;td&gt;reports/fixes/FIXES.md item (b); run WiNWSzkUnEGFirBjx (build 0.1.11, 12 priced hotels, nonsense search skipped)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 hotels and 478 offers&lt;/td&gt;
&lt;td&gt;Run 3BgiBhR7DZifI3D2P (9 Oct 2026, build 0.1.13), chargedEventCounts hotel 2, offer 478&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Each offer is a paid event&lt;/td&gt;
&lt;td&gt;claims.csv row 70&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 hotels and 49 offers after the fix&lt;/td&gt;
&lt;td&gt;Run BNCMJr19AnUWUQUt4 (build 0.1.14); reports/decisions.md row 47&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Two browser visits 12 minutes apart had 87% of hotels in common&lt;/td&gt;
&lt;td&gt;claims.csv row 67&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Adding the state chooses the Springfield&lt;/td&gt;
&lt;td&gt;claims.csv row 65 (runs klsxwgZlpaG94y8Bh, 5jrjtwZ72Z9SFK7tY)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

</description>
      <category>debugging</category>
      <category>python</category>
      <category>testing</category>
      <category>webscraping</category>
    </item>
    <item>
      <title>I searched for Nike and my scraper never found Nike</title>
      <dc:creator>Aniruddha Garje</dc:creator>
      <pubDate>Fri, 09 Oct 2026 13:44:07 +0000</pubDate>
      <link>https://dev.to/garje/i-searched-for-nike-and-my-scraper-never-found-nike-12h9</link>
      <guid>https://dev.to/garje/i-searched-for-nike-and-my-scraper-never-found-nike-12h9</guid>
      <description>&lt;p&gt;I asked my Google Ads Transparency scraper for the ads of "nike". It returned ads from two advertisers called Nikena and nikey. It returned nothing from Nike, Inc., which Google counts at 9,000 to 10,000 ads.&lt;/p&gt;

&lt;p&gt;Every row in that output was valid. Each ad had an ID, a format, dates and a link that opened. This is about two bugs like that, where nothing was missing and the answer was still wrong. It is also about the time the bug turned out to be in my own test.&lt;/p&gt;

&lt;h2&gt;
  
  
  The search that found the wrong company
&lt;/h2&gt;

&lt;p&gt;The Google Ads Transparency Center lets anyone look up the ads an advertiser has run. You can search by advertiser name. When you type a name, Google suggests a list of advertisers, and you pick one.&lt;/p&gt;

&lt;p&gt;My scraper picked for you. It took the first suggestions Google gave. For "nike", Nike, Inc. was eighth on that list. The run asked for 2 ads per advertiser, and the four ads it returned for "nike" came from Nikena and nikey (run w3qudNilMcBOHnIh4).&lt;/p&gt;

&lt;p&gt;Here is what Google's suggestion list held, with the number of ads Google reports for each (run T2yQUXYCf9kV2Bl9A):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Advertiser&lt;/th&gt;
&lt;th&gt;Country&lt;/th&gt;
&lt;th&gt;Ads Google reports&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Nike, Inc.&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;9,000 to 10,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NIKE SRL&lt;/td&gt;
&lt;td&gt;IT&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nike&lt;/td&gt;
&lt;td&gt;KE&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;nikey&lt;/td&gt;
&lt;td&gt;BG&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nikena&lt;/td&gt;
&lt;td&gt;BG&lt;/td&gt;
&lt;td&gt;27&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  A score instead of a guess
&lt;/h2&gt;

&lt;p&gt;The fix was to stop guessing silently. Every suggested advertiser now gets a match score, and every candidate is saved, so you can see what was picked and what was skipped.&lt;/p&gt;

&lt;p&gt;This is the actual scoring function from the scraper, shortened only in the list of legal forms:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;difflib&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;

&lt;span class="n"&gt;LEGAL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inc&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llc&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ltd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;srl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gmbh&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ag&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sa&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;co&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;plc&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;corp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\w+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;LEGAL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;match_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;term&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;term&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;difflib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;SequenceMatcher&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;ratio&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Nike, Inc.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NIKE SRL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nikey&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Nikena&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;match_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nike&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="c1"&gt;# Nike, Inc. 1.0 / NIKE SRL 1.0 / nikey 0.89 / Nikena 0.8
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By default only exact names, a score of 1, are scraped, the advertiser with the most ads first. The next run scraped Nike, Inc. with a score of 1 (run T2yQUXYCf9kV2Bl9A).&lt;/p&gt;

&lt;p&gt;Here is the honest part. NIKE SRL in Italy and an advertiser named Nike in Kenya also score 1. The score tells you the names match. It does not tell you they are the brand you meant. That is why every candidate, picked or not, is saved with its score, and why searching by website domain is the safer route when you know it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eleven dates, each one day late
&lt;/h2&gt;

&lt;p&gt;The second bug came from an accuracy check. A real browser read the live Transparency Center pages, the scraper ran straight after, and a script compared the two.&lt;/p&gt;

&lt;p&gt;25 of 25 ads were found. But the last-shown date matched for only 14 of them (run HM4pQz7xlm09D9wRV). The other 11 were all exactly one day late. For one image ad, the site showed 6 October and my output said 7 October.&lt;/p&gt;

&lt;p&gt;The pattern gave it away: every one of the 11 was last shown between midnight and about 8:00 UTC. The Transparency Center displays dates in US Pacific time. My scraper took the UTC timestamp and printed its date. Between midnight UTC and the end of the day in California, those are different dates.&lt;/p&gt;

&lt;p&gt;Why only 11? The UTC date and the Pacific date differ only from midnight UTC until midnight in California. An ad last shown in the afternoon UTC gets the same date either way. That is why 14 dates were right and hid the bug: the error appeared only for ads that stopped running in those early UTC hours.&lt;/p&gt;

&lt;p&gt;You can see it in three lines (the timestamp is an example, not from the data):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;zoneinfo&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ZoneInfo&lt;/span&gt;

&lt;span class="n"&gt;ts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromisoformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-10-07T03:00:00+00:00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;date&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;                                              &lt;span class="c1"&gt;# 2026-10-07
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;astimezone&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ZoneInfo&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;America/Los_Angeles&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;date&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;  &lt;span class="c1"&gt;# 2026-10-06
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fix kept the UTC timestamps exactly as they were and added &lt;code&gt;firstShownDate&lt;/code&gt; and &lt;code&gt;lastShownDate&lt;/code&gt; as the site displays them. The next check matched 25 of 25 (run 6erVdpbUYI1O7psAY).&lt;/p&gt;

&lt;p&gt;If your data comes from a page people read, store the raw timestamp and also the date the page shows. People will compare your output with the page, not with UTC.&lt;/p&gt;

&lt;h2&gt;
  
  
  The time the bug was mine
&lt;/h2&gt;

&lt;p&gt;A later accuracy check found 135 of 146 ads (run DNTVboAQ1sdphxKl1). All 11 misses belonged to one advertiser. My audit had capped the scraper at 30 ads per advertiser, and that advertiser's browser ads sat hundreds of places deep.&lt;/p&gt;

&lt;p&gt;With the cap at 1,000, the same audit found 146 of 146 (run G0zl0lGm3dQc33FVz). Nothing in the scraper changed. Before you fix the code, check that the test asked the same question as the browser.&lt;/p&gt;

&lt;h2&gt;
  
  
  One more thing I did not expect
&lt;/h2&gt;

&lt;p&gt;If you plan to analyse ad copy from this source, know this first. In one run of 4,578 ads, 2,866 of 2,958 text ads, 96.9%, came back as a single image rather than as words (run CmLk2DJUK1igwGucM). You get an image URL, not the headline text.&lt;/p&gt;

&lt;p&gt;That raised a quieter problem. A blank headline can mean "this ad has no headline" or "I could not read it". Those are different facts, and a null field hides the difference.&lt;/p&gt;

&lt;p&gt;So every ad now carries a &lt;code&gt;contentStatus&lt;/code&gt;: ok, imageOnly, externallyHosted, notAvailable, previewFailed or none. In that same run, 79.6% were imageOnly, 14.3% were ok and 6.1% were notAvailable. Image URLs came back for 83% of ads. When a field is empty, the status says why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three checks for your own pipeline
&lt;/h2&gt;

&lt;p&gt;If you pull data by name from any source that suggests matches, these came out of the two bugs above:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Never take the first suggestion silently.&lt;/strong&gt; Score every candidate against what was asked, keep the scores, and pick by rule. Save the candidates you skipped, so a person can see the choice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store two times, not one.&lt;/strong&gt; Keep the raw timestamp for arithmetic, and the date exactly as the page shows it for people. They will compare your output with the page, not with UTC.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the test.&lt;/strong&gt; When a check fails, first confirm it asked the same question as the browser: the same limits, the same time zone, the same page depth.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The safest name is often no name at all. If you know the brand's website, search the Transparency Center by domain. A domain such as nike.com has no spelling to guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I still have not solved
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;An exact name is not proof of identity. The score narrows the guess; a person still decides.&lt;/li&gt;
&lt;li&gt;My scraper does not read text out of images, so for most text ads you get the image, not the words.&lt;/li&gt;
&lt;li&gt;Outside the EU, the Transparency Center shows only a last-shown date per region, not a first-shown date or impression ranges. I cannot return what the source does not show.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Over to you
&lt;/h2&gt;

&lt;p&gt;Two habits came out of this: score your matches instead of guessing, and keep both the raw timestamp and the date your users will see. What is the worst silent match you have found in a data pipeline? Reply below, I read every one.&lt;/p&gt;

&lt;p&gt;These fixes now live in the Google Ads Transparency scraper I built on Apify. Every row carries the search term and its match score, and dates come both as UTC timestamps and as the site shows them. It is here if you want to try it: &lt;a href="https://apify.com/garje/google-ads-transparency-scraper?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=story&amp;amp;utm_content=searching-for-nike" rel="noopener noreferrer"&gt;https://apify.com/garje/google-ads-transparency-scraper?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=story&amp;amp;utm_content=searching-for-nike&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Fact or number&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Search "nike" returned ads from Nikena and nikey and none from Nike, Inc.; 2 ads per advertiser asked, 8 ads delivered across "nike" and "hubspot"&lt;/td&gt;
&lt;td&gt;Run w3qudNilMcBOHnIh4 (build 0.1.10), sample rows and RUN_STATS (reports/fixes/before.json); reports/fixes/FIXES.md item (d)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nike, Inc. was eighth in Google's suggestions&lt;/td&gt;
&lt;td&gt;reports/fixes/FIXES.md item (d)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Advertisers, countries and ad counts in the table&lt;/td&gt;
&lt;td&gt;Run T2yQUXYCf9kV2Bl9A, SEARCH_MATCHES record (reports/fixes/after.json)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scoring function and the scores 1.0, 1.0, 0.89, 0.8&lt;/td&gt;
&lt;td&gt;google-ads-transparency-scraper/src/parsers.py match_score (legal-form list shortened in the article); scores reproduced locally 9 Oct 2026 and stored in run T2yQUXYCf9kV2Bl9A&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Only exact names scraped by default, biggest first; Nike, Inc. scraped with score 1&lt;/td&gt;
&lt;td&gt;claims.csv row 27; reports/fixes/FIXES.md item (d); run T2yQUXYCf9kV2Bl9A&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25 of 25 found, last-shown date right for 14 of 25, 11 one day late, all last shown between midnight and about 8:00 UTC&lt;/td&gt;
&lt;td&gt;reports/ACCURACY_REPORT.md (run HM4pQz7xlm09D9wRV, build 0.1.6)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Example ad shown 6 October, output 7 October&lt;/td&gt;
&lt;td&gt;reports/ACCURACY_REPORT.md mismatch table, CR11595459565080018945&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dates displayed in US Pacific time&lt;/td&gt;
&lt;td&gt;claims.csv row 26&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25 of 25 after the fix&lt;/td&gt;
&lt;td&gt;Run 6erVdpbUYI1O7psAY (build 0.1.7)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;135 of 146 with a cap of 30, all 11 misses one advertiser; 146 of 146 with a cap of 1,000&lt;/td&gt;
&lt;td&gt;reports/launch/gate_sheet.md; reports/decisions.md row 25; runs DNTVboAQ1sdphxKl1 and G0zl0lGm3dQc33FVz&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2,866 of 2,958 text ads (96.9%) came back as images, in a run of 4,578 ads&lt;/td&gt;
&lt;td&gt;claims.csv rows 31 and 76 (run CmLk2DJUK1igwGucM)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The scraper does not read text out of images&lt;/td&gt;
&lt;td&gt;claims.csv row 33&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-region first-shown dates and impression ranges only for EU countries&lt;/td&gt;
&lt;td&gt;claims.csv row 36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UTC and Pacific dates differ only from midnight UTC until midnight in California; 14 of 25 right&lt;/td&gt;
&lt;td&gt;Time zone arithmetic; reports/ACCURACY_REPORT.md (all 11 misses last shown between midnight and about 8:00 UTC)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;contentStatus values; imageOnly 79.6%, ok 14.3%, notAvailable 6.1%; image URLs for 83%&lt;/td&gt;
&lt;td&gt;claims.csv rows 29 and 32 (run CmLk2DJUK1igwGucM)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ads can be looked up by website domain&lt;/td&gt;
&lt;td&gt;claims.csv row 23&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

</description>
      <category>debugging</category>
      <category>programming</category>
      <category>webscraping</category>
    </item>
    <item>
      <title>Apple's review feed went quiet, and my scraper believed it</title>
      <dc:creator>Aniruddha Garje</dc:creator>
      <pubDate>Fri, 09 Oct 2026 13:43:48 +0000</pubDate>
      <link>https://dev.to/garje/apples-review-feed-went-quiet-and-my-scraper-believed-it-1c25</link>
      <guid>https://dev.to/garje/apples-review-feed-went-quiet-and-my-scraper-believed-it-1c25</guid>
      <description>&lt;p&gt;Apple's public App Store review feed sometimes answers with an empty page that still has reviews behind it. My scraper read the first empty page as the end of the data. For WhatsApp in four countries it collected 50 reviews, and after the fix the same request collected 1,050.&lt;/p&gt;

&lt;p&gt;This is how I found it, why my first fix was not enough, and the one question to ask of any paged source before you trust an empty page.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the feed works
&lt;/h2&gt;

&lt;p&gt;Apple publishes App Store reviews through a public JSON feed. You ask for an app, a country and a sort order, and you page through the results 50 at a time.&lt;/p&gt;

&lt;p&gt;The feed has a hard end. It gives 10 pages, 500 reviews, per country and sort order. Ask for page 11 and it answers HTTP 400.&lt;/p&gt;

&lt;p&gt;So the natural rule is: walk the pages, and stop when a page comes back empty. That rule is in a lot of scrapers. It was in mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The page that stayed empty
&lt;/h2&gt;

&lt;p&gt;On 8 October the feed started returning some pages empty. They were not the end. Reviews existed past them.&lt;/p&gt;

&lt;p&gt;One page stayed empty on 45 tries over 105 seconds. My scraper never made 45 tries. It saw the first empty page and stopped.&lt;/p&gt;

&lt;p&gt;For WhatsApp across four countries, that meant 50 App Store reviews (run 44E0AdNn8IGPY6454). Nothing failed. No error was raised. The run simply finished early and looked complete.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an empty page looks like
&lt;/h2&gt;

&lt;p&gt;I saved one of those empty pages from 8 October. It is a perfectly normal-looking answer. Here it is, shortened:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"feed"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"author"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"iTunes Store"&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"updated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-10-08T01:54:23-07:00"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"iTunes Store: Customer Reviews"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://itunes.apple.com/us/rss/customerreviews/page=2/id=310633997/sortby=mostrecent/json"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The title, the author, the timestamp and the page address are all there. Only the &lt;code&gt;entry&lt;/code&gt; list is missing, and &lt;code&gt;entry&lt;/code&gt; is where the reviews live. The saved file is 895 bytes. A normal page for the same app, with its 50 reviews, is 37,971 bytes.&lt;/p&gt;

&lt;p&gt;Nothing about that response says "error". If your code asks "did the request succeed?" the answer is yes. If it asks "is this the last page?" the honest answer is "I cannot tell from this".&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 1: do not stop at an empty page
&lt;/h2&gt;

&lt;p&gt;The first fix changed three things. Retry an empty page. If it stays empty, carry on to the next page instead of stopping. And list every page that stayed empty in the run summary, so a person can see what may be missing.&lt;/p&gt;

&lt;p&gt;The same request then collected 1,050 reviews (run fnfG0xKFneKM5RKp1). I thought that was the end of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check that failed anyway
&lt;/h2&gt;

&lt;p&gt;Before launch I ran an accuracy check. A real browser read the live review pages, the scraper ran straight after, and a script compared the two.&lt;/p&gt;

&lt;p&gt;Recall was 61.2%: 101 of 165 reviews (run bf4Uu4WnKX4qwbqNv). All 64 misses were App Store reviews, and every one sat on a page Apple had served empty to the scraper.&lt;/p&gt;

&lt;p&gt;The run's own message said: "Apple returned empty pages for 12 app and country pairs even after retries". It would have been easy to stop there. Apple was flaky, the scraper said so, nothing to fix.&lt;/p&gt;

&lt;p&gt;But the browser had seen those reviews. The data existed. My retries came one straight after another, so they all saw the same empty page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 2: come back later
&lt;/h2&gt;

&lt;p&gt;The second fix spreads the tries out in time. When a page stays empty, the scraper carries on with the rest of the walk and comes back to it: 3 more tries, 15 seconds apart.&lt;/p&gt;

&lt;p&gt;The next check found 165 of 165 (run Sr0gowqUd4rynW5tH). That run still reported 3 app and country pairs where a page stayed empty after every try. Those are listed in the run summary, not hidden.&lt;/p&gt;

&lt;p&gt;Here is the idea as a sketch you can adapt. It is not the scraper's code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;walk&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;last_page&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;late_tries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;fetch(page) returns a list of rows, or [] when the page comes back empty.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stuck&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;page&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;last_page&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;got&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tries&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;got&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;got&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;break&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;got&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;got&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;stuck&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="c1"&gt;# an empty page is not the end: keep walking
&lt;/span&gt;    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;page&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stuck&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;          &lt;span class="c1"&gt;# come back later, when the page may have filled in
&lt;/span&gt;        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;late_tries&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;got&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;got&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;got&lt;/span&gt;
                &lt;span class="n"&gt;stuck&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;remove&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stuck&lt;/span&gt;                &lt;span class="c1"&gt;# report what stayed empty instead of hiding it
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The question behind it: what is this source's real end signal? For Apple's feed it is page 10, and an HTTP 400 on page 11. An empty page is not that signal. If you only know "empty means done", you will stop early the day the source has a bad minute.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two stores, two ways to say "the end"
&lt;/h2&gt;

&lt;p&gt;The same scraper reads Google Play, and Google Play ends differently. Each page comes with a token for the next page. When there is no token, there are no more reviews.&lt;/p&gt;

&lt;p&gt;Apple says it another way: 10 pages, then an HTTP 400 on page 11. Neither store uses an empty page to mean the end. I had invented that rule myself, because it is usually true.&lt;/p&gt;

&lt;h2&gt;
  
  
  How much the quiet pages cost
&lt;/h2&gt;

&lt;p&gt;In the failed check, 64 of 165 reviews were missing. That is 38.8% of the reviews a person could see in the browser, and every one sat on a page the scraper was told was empty.&lt;/p&gt;

&lt;p&gt;A scraper that stops early is harder to catch than one that crashes. A crash shows up in red. An early finish shows up as a smaller, tidy dataset that nobody questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I watch for it now
&lt;/h2&gt;

&lt;p&gt;Every run lists the pages that stayed empty after all tries, in its summary and in its status message. A small fixed run checks the scraper every morning. Once a week a real browser reads the live pages again and the scraper is compared against it.&lt;/p&gt;

&lt;p&gt;None of that stops Apple from serving an empty page. It makes sure I hear about it the day it happens, not the day a user notices.&lt;/p&gt;

&lt;h2&gt;
  
  
  The time the bug was mine
&lt;/h2&gt;

&lt;p&gt;The same kind of check, on Google Play, first showed 13 reviews with a date one day off. That one was not the scraper. Google Play displays the UTC date, and all 13 reviews were written after 18:30 UTC, which is midnight in India, where my check ran. My comparison used local time.&lt;/p&gt;

&lt;p&gt;I changed the comparison to UTC. No expected value was edited to make the numbers pass. On that earlier build, the check then matched 165 of 165 on every compared field (run X6yvF3mZldZch9wZt).&lt;/p&gt;

&lt;h2&gt;
  
  
  What I could not fix
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Apple's public feed stops at 500 reviews per country and sort order. Page 11 answers HTTP 400. The App Store website loads more through an access token issued to Apple's own web app, which is not a public method, so I do not use it. Adding countries is the honest way to get more.&lt;/li&gt;
&lt;li&gt;Some pages stay empty even after every try. I cannot make Apple serve them. I can only tell you exactly which ones.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Over to you
&lt;/h2&gt;

&lt;p&gt;Look at the paged sources in your own code. What does each one use to say "this is the end"? If the answer is "an empty page", I would like to know whether you have seen it lie. Reply below.&lt;/p&gt;

&lt;p&gt;These retries now live in the App Store and Google Play reviews scraper I built on Apify, and every run lists any page that stayed empty. It is here if you want to try it: &lt;a href="https://apify.com/garje/app-store-play-reviews-scraper?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=story&amp;amp;utm_content=feed-that-went-quiet" rel="noopener noreferrer"&gt;https://apify.com/garje/app-store-play-reviews-scraper?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=story&amp;amp;utm_content=feed-that-went-quiet&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Fact or number&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10 pages of 50 (500) per country and sort order; page 11 answers HTTP 400&lt;/td&gt;
&lt;td&gt;claims.csv row 11; reports/fixes/FIXES.md item (f)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apple's feed returned some pages empty although they held reviews&lt;/td&gt;
&lt;td&gt;claims.csv row 12; reports/fixes/FIXES.md item (f) extra&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One page stayed empty on 45 tries over 105 seconds&lt;/td&gt;
&lt;td&gt;reports/fixes/FIXES.md item (f) extra&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50 App Store reviews for WhatsApp in 4 countries before the fix&lt;/td&gt;
&lt;td&gt;Run 44E0AdNn8IGPY6454 (8 Oct 2026, build 0.1.5); reports/fixes/FIXES.md&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1,050 reviews after fix 1&lt;/td&gt;
&lt;td&gt;Run fnfG0xKFneKM5RKp1 (build 0.1.6)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recall 61.2%, 101 of 165; all 64 misses App Store reviews on pages served empty&lt;/td&gt;
&lt;td&gt;reports/decisions.md row 4; reports/launch/gate_sheet.md (run bf4Uu4WnKX4qwbqNv, build 0.1.10)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Status message "Apple returned empty pages for 12 app and country pairs even after retries"&lt;/td&gt;
&lt;td&gt;Run bf4Uu4WnKX4qwbqNv status message&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Late retries: 3 more tries, 15 seconds apart, while the walk carries on&lt;/td&gt;
&lt;td&gt;reports/decisions.md row 4; app-store-play-reviews-scraper/src/main.py APPLE_LATE_TRIES, APPLE_LATE_WAIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;165 of 165 after fix 2; 3 pairs still empty after every try&lt;/td&gt;
&lt;td&gt;reports/decisions.md row 5; run Sr0gowqUd4rynW5tH (build 0.1.11) status message&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;13 Google Play dates one day off; Google Play displays the UTC date; all 13 written after 18:30 UTC; no expected value changed&lt;/td&gt;
&lt;td&gt;reports/FINAL_ACCURACY.md section 1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The check ran on a browser in India&lt;/td&gt;
&lt;td&gt;reports/ACCURACY_REPORT.md ("a real Chromium browser on this computer, in India")&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;165 of 165 on every compared field&lt;/td&gt;
&lt;td&gt;claims.csv row 49; reports/FINAL_ACCURACY.md (run X6yvF3mZldZch9wZt)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;More reviews need an access token issued to Apple's web app, not a public method&lt;/td&gt;
&lt;td&gt;reports/fixes/FIXES.md item (f)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Play pages run until the store has no more reviews (a next-page token)&lt;/td&gt;
&lt;td&gt;claims.csv row 10 (run bpwfcFd4HmPVSohLv; test_play_page_parses_reviews_token_and_replies)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;64 of 165 missing is 38.8%&lt;/td&gt;
&lt;td&gt;Computed from reports/decisions.md row 4 (64 / 165)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Empty pages listed in the run summary and status message&lt;/td&gt;
&lt;td&gt;claims.csv row 12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A fixed check runs every morning; a browser comparison runs every week&lt;/td&gt;
&lt;td&gt;Daily canary task canary-reviews (06:00 UTC, reports/decisions.md row 40 context); claims.csv row 50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The empty page: fields present, no &lt;code&gt;entry&lt;/code&gt; list, 895 bytes; a 50-review page is 37,971 bytes&lt;/td&gt;
&lt;td&gt;Saved responses app-store-play-reviews-scraper/tests/fixtures/apple_rss_us_310633997_p2_empty_flaky.json and apple_rss_us_310633997_p1.json (8 Oct 2026); the article shows the empty page shortened&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

</description>
      <category>api</category>
      <category>debugging</category>
      <category>programming</category>
      <category>webscraping</category>
    </item>
  </channel>
</rss>
