DEV Community

Aniruddha Garje
Aniruddha Garje

Posted on Originally published at garje-data-notes.laude--pify.workers.dev

I searched Google Hotels for a city that does not exist. It found five hotels.

I typed "hotels in Qwzxvbnmplk" into my Google Hotels scraper to watch it fail. It did not fail. It came back with five hotels near San Antonio, Texas, each with a price, a rating and a map position.

This is the story of that answer and the fix. It is also how my fix broke a real capital city, and the small check I now run before I trust any search result.

A test that was supposed to fail

On 8 October I was testing inputs that should go wrong. One of them was a search for a city that does not exist. I expected an error, or an empty table.

Here is what the run returned (run pLMLM09FNCR1f0Tft):

Name Price per night Rating Reviews Latitude, longitude
Flat with pool & barbeque, Pet-friendly $75.77 4.665 49 29.404, -98.626
Holiday Inn San Antonio Seaworld by IHG $83.09 4.4 2,198 29.448, -98.680
Texas Sage One Room Cabin $113.00 4.725 32 29.777, -98.950
Luxurious Canyon Lake glamping with hot tub + pool $254.00 4.4 22 29.869, -98.302
Apartment with 1 bedroom for 2 persons $64.79 4.805 64 29.475, -98.536

Every field was filled. Every price was a real price for a real place. My scraper charges per priced hotel, so each of those five rows would have been a paid result.

Why nothing looked wrong

The same run had two other searches. "hotels in Paris" returned hotels in Paris. "hotels in Springfield" returned five hotels around 39.75, -89.6, which is Springfield, Illinois. Google picked one of the many Springfields and did not say so.

That is the pattern. A search engine always tries to answer. When it cannot match your words to a place, it still shows you hotels somewhere. I do not know why it chose San Antonio.

My scraper checked the shape of the answer. Prices were numbers, ratings were between 1 and 5, every row had coordinates. All of that was true. None of it said whether the answer belonged to the question.

The signal was there all along

Google's results page carries the place it understood. For Paris, the page held a place ID and the name "Paris". For Springfield, a place ID and "Springfield".

For "hotels in Qwzxvbnmplk" it held nothing. The run's own statistics record the place for the nonsense search as null. The source had told me it did not understand. I was not reading that part.

The fix

I changed the order of work:

  1. Resolve the place first. If Google returns no place, skip the search, charge nothing and name it in the run's status message.
  2. Put the place Google understood on every row, in a field called resolvedPlace, so you can see which Springfield you got.
  3. If the resolved place shares no word with the search, log a warning. Only a warning: it never drops results.

The next run said: "Finished with 10 priced hotels. Not recognised by Google, so skipped and not charged: "hotels in Qwzxvbnmplk"." (run 5jrjtwZ72Z9SFK7tY).

Then it skipped Prague

Later that day another run finished with this message: "Finished with 0 priced hotels. Not recognised by Google, so skipped and not charged: "hotels in Prague"." (run YTSa6YvgqYp6QrGc5).

Google had sent the page for Prague without its place data. My new check did exactly what I told it to. It treated a real capital city like a made-up one.

The second fix: when the place is missing, retry the page up to 3 times on a new IP before skipping the search. The next run resolved Prague and still skipped the nonsense search (run WiNWSzkUnEGFirBjx).

A strict check needs a way to be wrong safely. Mine had none until Prague showed me.

The check you can use tomorrow

This is not my scraper's code. It is the idea in a form you can paste into anything that asks a source a question and gets a confident answer back. Think of a geocoder, a search API, an entity lookup or a tool call from a language model.

import re

def words(text):
    return set(re.findall(r"\w+", text.lower()))

def check_answer(question, resolved):
    """resolved is whatever the source says it understood, for example a place name."""
    if not resolved:
        return "skip: the source did not say what it understood"
    if not words(question) & words(resolved):
        return "warn: the source may have answered a different question"
    return "ok"

print(check_answer("hotels in Qwzxvbnmplk", None))          # skip
print(check_answer("hotels in Springfield", "Springfield"))  # ok, but which Springfield?
print(check_answer("hotels in Paris", "Paris"))              # ok
Enter fullscreen mode Exit fullscreen mode

Two rules hide in those lines. Before you count rows, check that the source echoed your question. And keep what it understood next to every row, so a person can see it later.

The same bug in different clothes

On 9 October, while setting up public examples, I ran one that asked for 2 hotels with each booking site's offer. The run delivered 2 hotels and 478 offers (run 3BgiBhR7DZifI3D2P).

The scraper had fetched offers for every hotel it kept on the page, not only for the two it delivered. Each offer is a paid event. Again every row was complete and plausible, and the total was wrong.

After the fix, the same kind of request gave 2 hotels and 49 offers (run BNCMJr19AnUWUQUt4). The check here is simple: compare what you bill against what you delivered, every run.

How do you test a source that disagrees with itself?

After this, I wanted a number for how often my hotel results match what a person sees. So I had a real browser read the live Google Hotels pages, ran the scraper straight after, and compared the two.

The first surprise was the browser. Two browser visits 12 minutes apart had only 87% of their hotels in common. Google reshuffles its list between visits.

A test that demands 95% agreement with a source that agrees with itself 87% of the time will fail forever. Worse, it teaches you to ignore it. So my pass mark for the hotel list is the browser's own agreement, 87.2%, minus 3 points: 84.2%.

Prices get a strict bar, because a price should not drift between two looks at the same hotel. For hotels that appeared in both, the price matched within 10% in 99% of cases.

Separate what the source keeps changing from what it should never get wrong, and test each one against its own bar.

The alarm that fired because the data got better

On 9 October my daily check on Paris hotels failed. It watches how often each field comes back empty, and alarms when that share moves more than 5 points from a saved baseline.

What moved was good news. Fewer hotels than usual were missing a star class: 15%, against 30% in the baseline (run J3MVimtp32Xg6WCHp). On 20 hotels, 5 points is a single hotel, and the list reshuffles between visits.

I changed the alarm to fire only when empty values rise, and only when 3 of the 20 hotels change. The rerun passed (run pql2nTclNeiU8o8sn). A check that cries wolf gets switched off, and then it protects nothing.

What I still cannot know

  • Google's hotel list changes between visits. In my test, two browser visits 12 minutes apart had 87% of their hotels in common. "Correct" for a list means "matches what Google showed at that moment", not a fixed truth.
  • My warning is a word check. A search for a landmark can resolve to a city whose name shares no word with it, and then the warning fires for nothing. That is why it only warns.
  • I can tell you which Springfield Google picked. I cannot tell you whether it is the one you meant. Adding the state to the search, as in "hotels in Springfield, Massachusetts", is what fixes that.

Over to you

Where in your stack does a source answer confidently when it did not understand the question? I would like to hear the strangest example you have found. Reply below.

These checks now live in the Google Hotels scraper I built on Apify: every row carries resolvedPlace, and a search Google does not recognise is skipped and not charged. It is here if you want to see them work: https://apify.com/garje/google-hotels-scraper?utm_source=devto&utm_medium=article&utm_campaign=story&utm_content=city-that-does-not-exist

Sources

Fact or number Source
"hotels in Qwzxvbnmplk" returned 5 hotels near San Antonio, with the names, prices, ratings, review counts and coordinates in the table Run pLMLM09FNCR1f0Tft (8 Oct 2026, 08:46 UTC, build 0.1.9), dataset rows; reports/fixes/FIXES.md item (b)
Each priced hotel is a paid result claims.csv row 70 (pricing); reports/fixes/FIXES.md item (c)
"hotels in Springfield" returned hotels around 39.75, -89.6 Run pLMLM09FNCR1f0Tft dataset rows (coordinates of Springfield, Illinois)
Place recorded as null for the nonsense search; place IDs and names for Paris and Springfield Run pLMLM09FNCR1f0Tft, RUN_STATS.places (reports/fixes/before.json)
resolvedPlace on every row; unrecognised search skipped, not charged, named in the status message; warning when no word is shared claims.csv rows 64 and 66 (row 66 UNVERIFIED for the log message wording, so the article only says it logs a warning, as in the code); reports/fixes/FIXES.md item (b); google-hotels-scraper/src/main.py place_matches
Status message after the fix, 10 priced hotels Run 5jrjtwZ72Z9SFK7tY (build 0.1.10)
Prague skipped, 0 priced hotels Run YTSa6YvgqYp6QrGc5 (8 Oct 2026, 09:18 UTC, build 0.1.10)
Retry up to 3 times on a new IP; Prague then resolved reports/fixes/FIXES.md item (b); run WiNWSzkUnEGFirBjx (build 0.1.11, 12 priced hotels, nonsense search skipped)
2 hotels and 478 offers Run 3BgiBhR7DZifI3D2P (9 Oct 2026, build 0.1.13), chargedEventCounts hotel 2, offer 478
Each offer is a paid event claims.csv row 70
2 hotels and 49 offers after the fix Run BNCMJr19AnUWUQUt4 (build 0.1.14); reports/decisions.md row 47
Two browser visits 12 minutes apart had 87% of hotels in common claims.csv row 67
Adding the state chooses the Springfield claims.csv row 65 (runs klsxwgZlpaG94y8Bh, 5jrjtwZ72Z9SFK7tY)

Top comments (0)