I run a visa agency in Bali, and the side of my work that has nothing to do with visas is building scrapers on Apify. One of them reads Zillow: homes for sale, homes for rent, homes recently sold. It is the Actor I have spent the most time measuring rather than writing, because Zillow's data looks far more uniform from the outside than it is.
Everything below is counted from real runs — dates and dataset sizes included — not from the docs.
The 820 ceiling is the first thing to design around
A Zillow search page serves 41 results per page and stops at page 20. That product is the hard ceiling of any single Zillow search:
RESULTS_PER_PAGE = 41
MAX_PAGES = 20
RESULT_CAP = RESULTS_PER_PAGE * MAX_PAGES # 820
The page also tells you, honestly, how many homes it is hiding from you. searchPageState.cat1.searchList.totalResultCount is the real match count. Three searches for one city, straight from my run logs:
| Search (Austin, TX) | Date | totalResultCount |
Times over the 820 cap |
|---|---|---|---|
| For sale | 6 Oct 2026 | 5,589 | 6.8× |
| Recently sold | 7 Oct 2026 | 14,290 | 17.4× |
| For rent | 4 Oct 2026 | 23,994 | 29.3× |
So for any city-sized query, paginating to the end gets you 3–15% of the market and no warning that anything is missing. That is the failure mode worth worrying about: not an error, just a short answer.
The way out is the map. A Zillow query state carries mapBounds — north, south, east, west — and the cap applies per query, not per region. Read the total first; if it is over 820, cut the box into four and recurse:
def split_bounds(b):
mid_lat = (b['north'] + b['south']) / 2
mid_lng = (b['east'] + b['west']) / 2
return [
{'north': b['north'], 'south': mid_lat, 'west': b['west'], 'east': mid_lng},
{'north': b['north'], 'south': mid_lat, 'west': mid_lng, 'east': b['east']},
{'north': mid_lat, 'south': b['south'], 'west': b['west'], 'east': mid_lng},
{'north': mid_lat, 'south': b['south'], 'west': mid_lng, 'east': b['east']},
]
Two details decide whether this works. First, keep the region id in the query state while you narrow the box, or quadrants will pull in homes from the next town over. Second, dedupe on zpid as rows come in: quadrant edges overlap, and a home on the line comes back twice. My implementation allows up to six levels of recursion, which is more headroom than any US metro I have tried needs: 23,994 Austin rentals come apart in two or three levels, because the split follows density — a box under 820 stops immediately and never splits again.
Same schema, two completely different datasets
Now the part I did not expect. I ran two neighbouring Napa County ZIP codes on 24 September 2026 with full property details on: 94515 (Calistoga, for sale, 106 homes, 111 requests, 35 seconds) and 94558 (Napa, for rent, 122 rows). Every row in both runs has the same 91 keys — the schema is identical, by construction. What is filled is almost mutually exclusive:
| Field | For sale (106 rows) | For rent (88 non-building rows) |
|---|---|---|
mlsId |
89% | 0% |
propertyTaxRate |
82% | 0% |
lotSizeSqft |
81% | 2% |
yearBuilt |
74% | 7% |
agentPhone |
98% | 39% |
priceHistory |
2% | 100% |
schools |
0% | 100% |
taxHistory |
0% | 60% |
county |
16% | 100% |
parcelNumber |
0% | 55% |
description, photos
|
100% | 100% |
Both page types embed a JSON property object in __NEXT_DATA__ (under gdpClientCache), and both are called the same thing, but they are assembled by different page templates with different data loaders. A for-sale page is built around the MLS record: listing id, agent, tax rate, lot. A rental page is built around the renter's questions: who the schools are, what the place rented for before, who owns the parcel.
The same field can even change shape between them. A price-history entry from a rental page:
{"date": "2026-08-30", "event": "Price change", "price": 4498,
"priceChangeRate": -0.011, "pricePerSqft": 2, "source": "Zillow Rentals"}
And from one of the two for-sale pages that had any history at all:
{"date": "2026-05-23", "event": null, "price": 1799000,
"priceChangeRate": null, "pricePerSqft": null, "source": null}
Dates and prices, no event labels. If your downstream logic branches on event == "Price change" to count price cuts, it silently counts zero on for-sale inventory while working perfectly on rentals.
The practical rule: do not validate a Zillow pipeline on one listing type. A completeness check that passes on rentals will flag for-sale data as broken, and vice versa. I now keep per-listing-type expectations, and the Actor's daily self-test runs all three.
A rent search does not return apartments
The second trap is a counting one. Ask for rentals in Austin and you do not get apartments — you get buildings. All 10 rows of my 4 October rental run came back with isBuilding: true:
{
"buildingName": "Hawthorn Village",
"address": "3663 Solano Ave, Napa, CA 94558",
"price": 2350, "minRent": 2350, "maxRent": 3147,
"availableUnits": 11, "phone": "(707) 394-2529",
"beds": null, "baths": null, "sqft": null,
"homeType": "apartment"
}
Note what price is. Across all 34 buildings in the Napa run, price equalled minRent in 34 cases out of 34 — it is the cheapest available unit, every time. Sixteen of those buildings advertise a rent range, with a median spread of 24% between cheapest and dearest unit and a maximum of 56%. Average the price column and you get $2,459; average each building's midpoint instead and you get $2,602. That is a 6% error in a number people put in investment decks, and it is biased in one direction, always low.
Buildings also explain a number that confused me for a day. The 94558 rent search reported totalResultCount: 229 and my Actor wrote 122 rows. Nothing was lost: 88 of those rows were individual homes, and the other 34 were buildings carrying 134 available units between them. 88 + 134 = 222, against Zillow's reported 229. Zillow counts units; its own result list returns buildings. Switching on one-row-per-unit turns a building into rows like this, which is what a rent comp actually needs:
{"buildingName": "Hawthorn Village", "unitNumber": "Unit 76", "floorPlan": "Adriatic",
"beds": 2, "baths": 1, "sqft": 932, "price": 2600, "availableFrom": "2026-09-11"}
Sold prices have a hole the size of Texas
Recently-sold data is the reason most people want Zillow at all — comps. Here is an Austin sold search from 7 October 2026, 10 rows:
-
soldDate: 10 of 10 -
zestimate: 10 of 10 -
soldPrice: 1 of 10
Texas is one of roughly a dozen non-disclosure states, where sale prices are not public record. Zillow shows the sold date and nothing else, so that is what comes out. The single row with a price had closed that same day, and its last asking price was still attached to the listing.
The temptation here is obvious and wrong: zestimate is filled on every one of those rows, so a pipeline that quietly falls back to it produces a complete-looking comps table made of estimates. I return null instead. If you are building sold-price analytics, filter by state first and know which markets simply cannot answer.
Three smaller things that cost me time
"$0" means "price on request". One row in 140 Austin for-sale rows: price: null, priceText: "$0", a 7-bed, 7,700 sqft house with a $4,645,600 Zestimate and 185 days on Zillow. Rare, but a naive int(priceText) puts a $0 luxury home at the top of every "cheapest listings" report.
Land has zero bedrooms. 12 of those 140 rows were homeType: "lot". Eleven had beds: 0, one had beds: null, and all 12 had sqft: null — so pricePerSqft is null too. Any city-wide average of beds or price per sqft needs lots excluded, and any bedsMin: 1 filter silently drops the land market.
My own bug, for symmetry. Zillow's homeInsights block repeats itself, and I concatenated it: 76 of 76 for-sale rows and 84 of 84 rental rows had every highlight phrase twice. Nobody reported it because duplicated marketing phrases look like marketing phrases. The fix is one line — list(dict.fromkeys(phrases)), order preserved — and it went in the day I counted it. The lesson is that a field nobody filters on is a field nobody validates.
Calling it
The Actor is Zillow Scraper on Apify. One request, rows back:
curl -X POST "https://api.apify.com/v2/acts/lergassy~zillow-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"listingType":"sold","locations":["Zilker, Austin, TX"],"daysOnZillow":"90","maxItems":200}'
Search rows cost $1.30 per 1,000 and full property details $2.50 per 1,000, so the 106-home Calistoga run with details on was about $0.40. No login, no API key, no cookies; Zillow answers datacenter addresses directly and the residential fallback only engages when a request is actually refused.
When it is the wrong tool: tax history and school ratings on for-sale pages come from a separate signed request that a logged-out visitor cannot reproduce, so I leave them empty rather than guess. Sold prices in non-disclosure states are not recoverable from Zillow at any price. And if you need one specific home's full record, hand it a URL, a ZPID or a street address — running a whole city search to find one house is paying by the thousand for a single row.
If you work with Zillow data, the one thing I would take from this: check completeness per listing type, not per Actor. The schema promises uniformity that the source does not have.
Top comments (0)