Apple's public App Store review feed sometimes answers with an empty page that still has reviews behind it. My scraper read the first empty page as the end of the data. For WhatsApp in four countries it collected 50 reviews, and after the fix the same request collected 1,050.
This is how I found it, why my first fix was not enough, and the one question to ask of any paged source before you trust an empty page.
How the feed works
Apple publishes App Store reviews through a public JSON feed. You ask for an app, a country and a sort order, and you page through the results 50 at a time.
The feed has a hard end. It gives 10 pages, 500 reviews, per country and sort order. Ask for page 11 and it answers HTTP 400.
So the natural rule is: walk the pages, and stop when a page comes back empty. That rule is in a lot of scrapers. It was in mine.
The page that stayed empty
On 8 October the feed started returning some pages empty. They were not the end. Reviews existed past them.
One page stayed empty on 45 tries over 105 seconds. My scraper never made 45 tries. It saw the first empty page and stopped.
For WhatsApp across four countries, that meant 50 App Store reviews (run 44E0AdNn8IGPY6454). Nothing failed. No error was raised. The run simply finished early and looked complete.
What an empty page looks like
I saved one of those empty pages from 8 October. It is a perfectly normal-looking answer. Here it is, shortened:
{
"feed": {
"author": {"name": {"label": "iTunes Store"}},
"updated": {"label": "2026-10-08T01:54:23-07:00"},
"title": {"label": "iTunes Store: Customer Reviews"},
"id": {"label": "https://itunes.apple.com/us/rss/customerreviews/page=2/id=310633997/sortby=mostrecent/json"}
}
}
The title, the author, the timestamp and the page address are all there. Only the entry list is missing, and entry is where the reviews live. The saved file is 895 bytes. A normal page for the same app, with its 50 reviews, is 37,971 bytes.
Nothing about that response says "error". If your code asks "did the request succeed?" the answer is yes. If it asks "is this the last page?" the honest answer is "I cannot tell from this".
Fix 1: do not stop at an empty page
The first fix changed three things. Retry an empty page. If it stays empty, carry on to the next page instead of stopping. And list every page that stayed empty in the run summary, so a person can see what may be missing.
The same request then collected 1,050 reviews (run fnfG0xKFneKM5RKp1). I thought that was the end of it.
The check that failed anyway
Before launch I ran an accuracy check. A real browser read the live review pages, the scraper ran straight after, and a script compared the two.
Recall was 61.2%: 101 of 165 reviews (run bf4Uu4WnKX4qwbqNv). All 64 misses were App Store reviews, and every one sat on a page Apple had served empty to the scraper.
The run's own message said: "Apple returned empty pages for 12 app and country pairs even after retries". It would have been easy to stop there. Apple was flaky, the scraper said so, nothing to fix.
But the browser had seen those reviews. The data existed. My retries came one straight after another, so they all saw the same empty page.
Fix 2: come back later
The second fix spreads the tries out in time. When a page stays empty, the scraper carries on with the rest of the walk and comes back to it: 3 more tries, 15 seconds apart.
The next check found 165 of 165 (run Sr0gowqUd4rynW5tH). That run still reported 3 app and country pairs where a page stayed empty after every try. Those are listed in the run summary, not hidden.
Here is the idea as a sketch you can adapt. It is not the scraper's code.
import time
def walk(fetch, last_page=10, tries=4, late_tries=3, wait=15):
"""fetch(page) returns a list of rows, or [] when the page comes back empty."""
rows, stuck = [], []
for page in range(1, last_page + 1):
got = []
for _ in range(tries):
got = fetch(page)
if got:
break
if got:
rows += got
else:
stuck.append(page) # an empty page is not the end: keep walking
for page in list(stuck): # come back later, when the page may have filled in
for _ in range(late_tries):
time.sleep(wait)
got = fetch(page)
if got:
rows += got
stuck.remove(page)
break
return rows, stuck # report what stayed empty instead of hiding it
The question behind it: what is this source's real end signal? For Apple's feed it is page 10, and an HTTP 400 on page 11. An empty page is not that signal. If you only know "empty means done", you will stop early the day the source has a bad minute.
Two stores, two ways to say "the end"
The same scraper reads Google Play, and Google Play ends differently. Each page comes with a token for the next page. When there is no token, there are no more reviews.
Apple says it another way: 10 pages, then an HTTP 400 on page 11. Neither store uses an empty page to mean the end. I had invented that rule myself, because it is usually true.
How much the quiet pages cost
In the failed check, 64 of 165 reviews were missing. That is 38.8% of the reviews a person could see in the browser, and every one sat on a page the scraper was told was empty.
A scraper that stops early is harder to catch than one that crashes. A crash shows up in red. An early finish shows up as a smaller, tidy dataset that nobody questions.
How I watch for it now
Every run lists the pages that stayed empty after all tries, in its summary and in its status message. A small fixed run checks the scraper every morning. Once a week a real browser reads the live pages again and the scraper is compared against it.
None of that stops Apple from serving an empty page. It makes sure I hear about it the day it happens, not the day a user notices.
The time the bug was mine
The same kind of check, on Google Play, first showed 13 reviews with a date one day off. That one was not the scraper. Google Play displays the UTC date, and all 13 reviews were written after 18:30 UTC, which is midnight in India, where my check ran. My comparison used local time.
I changed the comparison to UTC. No expected value was edited to make the numbers pass. On that earlier build, the check then matched 165 of 165 on every compared field (run X6yvF3mZldZch9wZt).
What I could not fix
- Apple's public feed stops at 500 reviews per country and sort order. Page 11 answers HTTP 400. The App Store website loads more through an access token issued to Apple's own web app, which is not a public method, so I do not use it. Adding countries is the honest way to get more.
- Some pages stay empty even after every try. I cannot make Apple serve them. I can only tell you exactly which ones.
Over to you
Look at the paged sources in your own code. What does each one use to say "this is the end"? If the answer is "an empty page", I would like to know whether you have seen it lie. Reply below.
These retries now live in the App Store and Google Play reviews scraper I built on Apify, and every run lists any page that stayed empty. It is here if you want to try it: https://apify.com/garje/app-store-play-reviews-scraper?utm_source=devto&utm_medium=article&utm_campaign=story&utm_content=feed-that-went-quiet
Sources
| Fact or number | Source |
|---|---|
| 10 pages of 50 (500) per country and sort order; page 11 answers HTTP 400 | claims.csv row 11; reports/fixes/FIXES.md item (f) |
| Apple's feed returned some pages empty although they held reviews | claims.csv row 12; reports/fixes/FIXES.md item (f) extra |
| One page stayed empty on 45 tries over 105 seconds | reports/fixes/FIXES.md item (f) extra |
| 50 App Store reviews for WhatsApp in 4 countries before the fix | Run 44E0AdNn8IGPY6454 (8 Oct 2026, build 0.1.5); reports/fixes/FIXES.md |
| 1,050 reviews after fix 1 | Run fnfG0xKFneKM5RKp1 (build 0.1.6) |
| Recall 61.2%, 101 of 165; all 64 misses App Store reviews on pages served empty | reports/decisions.md row 4; reports/launch/gate_sheet.md (run bf4Uu4WnKX4qwbqNv, build 0.1.10) |
| Status message "Apple returned empty pages for 12 app and country pairs even after retries" | Run bf4Uu4WnKX4qwbqNv status message |
| Late retries: 3 more tries, 15 seconds apart, while the walk carries on | reports/decisions.md row 4; app-store-play-reviews-scraper/src/main.py APPLE_LATE_TRIES, APPLE_LATE_WAIT |
| 165 of 165 after fix 2; 3 pairs still empty after every try | reports/decisions.md row 5; run Sr0gowqUd4rynW5tH (build 0.1.11) status message |
| 13 Google Play dates one day off; Google Play displays the UTC date; all 13 written after 18:30 UTC; no expected value changed | reports/FINAL_ACCURACY.md section 1 |
| The check ran on a browser in India | reports/ACCURACY_REPORT.md ("a real Chromium browser on this computer, in India") |
| 165 of 165 on every compared field | claims.csv row 49; reports/FINAL_ACCURACY.md (run X6yvF3mZldZch9wZt) |
| More reviews need an access token issued to Apple's web app, not a public method | reports/fixes/FIXES.md item (f) |
| Google Play pages run until the store has no more reviews (a next-page token) | claims.csv row 10 (run bpwfcFd4HmPVSohLv; test_play_page_parses_reviews_token_and_replies) |
| 64 of 165 missing is 38.8% | Computed from reports/decisions.md row 4 (64 / 165) |
| Empty pages listed in the run summary and status message | claims.csv row 12 |
| A fixed check runs every morning; a browser comparison runs every week | Daily canary task canary-reviews (06:00 UTC, reports/decisions.md row 40 context); claims.csv row 50 |
The empty page: fields present, no entry list, 895 bytes; a 50-review page is 37,971 bytes |
Saved responses app-store-play-reviews-scraper/tests/fixtures/apple_rss_us_310633997_p2_empty_flaky.json and apple_rss_us_310633997_p1.json (8 Oct 2026); the article shows the empty page shortened |
Top comments (0)