DEV Community

Aniruddha Garje
Aniruddha Garje

Posted on Originally published at garje-data-notes.laude--pify.workers.dev

I searched for Nike and my scraper never found Nike

I asked my Google Ads Transparency scraper for the ads of "nike". It returned ads from two advertisers called Nikena and nikey. It returned nothing from Nike, Inc., which Google counts at 9,000 to 10,000 ads.

Every row in that output was valid. Each ad had an ID, a format, dates and a link that opened. This is about two bugs like that, where nothing was missing and the answer was still wrong. It is also about the time the bug turned out to be in my own test.

The search that found the wrong company

The Google Ads Transparency Center lets anyone look up the ads an advertiser has run. You can search by advertiser name. When you type a name, Google suggests a list of advertisers, and you pick one.

My scraper picked for you. It took the first suggestions Google gave. For "nike", Nike, Inc. was eighth on that list. The run asked for 2 ads per advertiser, and the four ads it returned for "nike" came from Nikena and nikey (run w3qudNilMcBOHnIh4).

Here is what Google's suggestion list held, with the number of ads Google reports for each (run T2yQUXYCf9kV2Bl9A):

Advertiser Country Ads Google reports
Nike, Inc. US 9,000 to 10,000
NIKE SRL IT 8
Nike KE 1
nikey BG 10
Nikena BG 27

A score instead of a guess

The fix was to stop guessing silently. Every suggested advertiser now gets a match score, and every candidate is saved, so you can see what was picked and what was skipped.

This is the actual scoring function from the scraper, shortened only in the list of legal forms:

import difflib, re

LEGAL = {"inc", "llc", "ltd", "srl", "gmbh", "ag", "bv", "sa", "co", "plc", "corp"}

def norm(text):
    return " ".join(w for w in re.findall(r"\w+", text.lower()) if w not in LEGAL)

def match_score(term, name):
    t, n = norm(term), norm(name)
    if not t or not n:
        return 0.0
    if t == n:
        return 1.0
    return round(difflib.SequenceMatcher(None, t, n).ratio(), 2)

for name in ["Nike, Inc.", "NIKE SRL", "nikey", "Nikena"]:
    print(name, match_score("nike", name))
# Nike, Inc. 1.0 / NIKE SRL 1.0 / nikey 0.89 / Nikena 0.8
Enter fullscreen mode Exit fullscreen mode

By default only exact names, a score of 1, are scraped, the advertiser with the most ads first. The next run scraped Nike, Inc. with a score of 1 (run T2yQUXYCf9kV2Bl9A).

Here is the honest part. NIKE SRL in Italy and an advertiser named Nike in Kenya also score 1. The score tells you the names match. It does not tell you they are the brand you meant. That is why every candidate, picked or not, is saved with its score, and why searching by website domain is the safer route when you know it.

Eleven dates, each one day late

The second bug came from an accuracy check. A real browser read the live Transparency Center pages, the scraper ran straight after, and a script compared the two.

25 of 25 ads were found. But the last-shown date matched for only 14 of them (run HM4pQz7xlm09D9wRV). The other 11 were all exactly one day late. For one image ad, the site showed 6 October and my output said 7 October.

The pattern gave it away: every one of the 11 was last shown between midnight and about 8:00 UTC. The Transparency Center displays dates in US Pacific time. My scraper took the UTC timestamp and printed its date. Between midnight UTC and the end of the day in California, those are different dates.

Why only 11? The UTC date and the Pacific date differ only from midnight UTC until midnight in California. An ad last shown in the afternoon UTC gets the same date either way. That is why 14 dates were right and hid the bug: the error appeared only for ads that stopped running in those early UTC hours.

You can see it in three lines (the timestamp is an example, not from the data):

from datetime import datetime
from zoneinfo import ZoneInfo

ts = datetime.fromisoformat("2026-10-07T03:00:00+00:00")
print(ts.date())                                              # 2026-10-07
print(ts.astimezone(ZoneInfo("America/Los_Angeles")).date())  # 2026-10-06
Enter fullscreen mode Exit fullscreen mode

The fix kept the UTC timestamps exactly as they were and added firstShownDate and lastShownDate as the site displays them. The next check matched 25 of 25 (run 6erVdpbUYI1O7psAY).

If your data comes from a page people read, store the raw timestamp and also the date the page shows. People will compare your output with the page, not with UTC.

The time the bug was mine

A later accuracy check found 135 of 146 ads (run DNTVboAQ1sdphxKl1). All 11 misses belonged to one advertiser. My audit had capped the scraper at 30 ads per advertiser, and that advertiser's browser ads sat hundreds of places deep.

With the cap at 1,000, the same audit found 146 of 146 (run G0zl0lGm3dQc33FVz). Nothing in the scraper changed. Before you fix the code, check that the test asked the same question as the browser.

One more thing I did not expect

If you plan to analyse ad copy from this source, know this first. In one run of 4,578 ads, 2,866 of 2,958 text ads, 96.9%, came back as a single image rather than as words (run CmLk2DJUK1igwGucM). You get an image URL, not the headline text.

That raised a quieter problem. A blank headline can mean "this ad has no headline" or "I could not read it". Those are different facts, and a null field hides the difference.

So every ad now carries a contentStatus: ok, imageOnly, externallyHosted, notAvailable, previewFailed or none. In that same run, 79.6% were imageOnly, 14.3% were ok and 6.1% were notAvailable. Image URLs came back for 83% of ads. When a field is empty, the status says why.

Three checks for your own pipeline

If you pull data by name from any source that suggests matches, these came out of the two bugs above:

  1. Never take the first suggestion silently. Score every candidate against what was asked, keep the scores, and pick by rule. Save the candidates you skipped, so a person can see the choice.
  2. Store two times, not one. Keep the raw timestamp for arithmetic, and the date exactly as the page shows it for people. They will compare your output with the page, not with UTC.
  3. Test the test. When a check fails, first confirm it asked the same question as the browser: the same limits, the same time zone, the same page depth.

The safest name is often no name at all. If you know the brand's website, search the Transparency Center by domain. A domain such as nike.com has no spelling to guess.

What I still have not solved

  • An exact name is not proof of identity. The score narrows the guess; a person still decides.
  • My scraper does not read text out of images, so for most text ads you get the image, not the words.
  • Outside the EU, the Transparency Center shows only a last-shown date per region, not a first-shown date or impression ranges. I cannot return what the source does not show.

Over to you

Two habits came out of this: score your matches instead of guessing, and keep both the raw timestamp and the date your users will see. What is the worst silent match you have found in a data pipeline? Reply below, I read every one.

These fixes now live in the Google Ads Transparency scraper I built on Apify. Every row carries the search term and its match score, and dates come both as UTC timestamps and as the site shows them. It is here if you want to try it: https://apify.com/garje/google-ads-transparency-scraper?utm_source=devto&utm_medium=article&utm_campaign=story&utm_content=searching-for-nike

Sources

Fact or number Source
Search "nike" returned ads from Nikena and nikey and none from Nike, Inc.; 2 ads per advertiser asked, 8 ads delivered across "nike" and "hubspot" Run w3qudNilMcBOHnIh4 (build 0.1.10), sample rows and RUN_STATS (reports/fixes/before.json); reports/fixes/FIXES.md item (d)
Nike, Inc. was eighth in Google's suggestions reports/fixes/FIXES.md item (d)
Advertisers, countries and ad counts in the table Run T2yQUXYCf9kV2Bl9A, SEARCH_MATCHES record (reports/fixes/after.json)
Scoring function and the scores 1.0, 1.0, 0.89, 0.8 google-ads-transparency-scraper/src/parsers.py match_score (legal-form list shortened in the article); scores reproduced locally 9 Oct 2026 and stored in run T2yQUXYCf9kV2Bl9A
Only exact names scraped by default, biggest first; Nike, Inc. scraped with score 1 claims.csv row 27; reports/fixes/FIXES.md item (d); run T2yQUXYCf9kV2Bl9A
25 of 25 found, last-shown date right for 14 of 25, 11 one day late, all last shown between midnight and about 8:00 UTC reports/ACCURACY_REPORT.md (run HM4pQz7xlm09D9wRV, build 0.1.6)
Example ad shown 6 October, output 7 October reports/ACCURACY_REPORT.md mismatch table, CR11595459565080018945
Dates displayed in US Pacific time claims.csv row 26
25 of 25 after the fix Run 6erVdpbUYI1O7psAY (build 0.1.7)
135 of 146 with a cap of 30, all 11 misses one advertiser; 146 of 146 with a cap of 1,000 reports/launch/gate_sheet.md; reports/decisions.md row 25; runs DNTVboAQ1sdphxKl1 and G0zl0lGm3dQc33FVz
2,866 of 2,958 text ads (96.9%) came back as images, in a run of 4,578 ads claims.csv rows 31 and 76 (run CmLk2DJUK1igwGucM)
The scraper does not read text out of images claims.csv row 33
Per-region first-shown dates and impression ranges only for EU countries claims.csv row 36
UTC and Pacific dates differ only from midnight UTC until midnight in California; 14 of 25 right Time zone arithmetic; reports/ACCURACY_REPORT.md (all 11 misses last shown between midnight and about 8:00 UTC)
contentStatus values; imageOnly 79.6%, ok 14.3%, notAvailable 6.1%; image URLs for 83% claims.csv rows 29 and 32 (run CmLk2DJUK1igwGucM)
Ads can be looked up by website domain claims.csv row 23

Top comments (0)