DEV Community

dodou
dodou

Posted on

Geo-Fenced Collection with lat/lng/zoom: A Reproducible Grid

Geo-Fenced Collection with lat/lng/zoom: A Reproducible Grid

The problem with "collect local business data" scripts is that they're not reproducible. You point at a city, get 20 results, and tomorrow the same query returns a different 20 — because you never said where you were looking. The maps endpoint takes optional lat/lng/zoom parameters that turn a vague query into a defined geographic window. This post is about using them properly: building a grid over an area, collecting with a fixed recipe, and being able to re-run the exact same collection next month.

I'll use the field contract from the SerpBase Maps API documentation. Each successful maps request costs 2 credits, so the grid design matters for your budget as much as for correctness.

The three parameters

Per the docs, POST /google/maps/search takes q (required), hl, gl, page (all optional), plus the geo trio:

  • lat — map center latitude, optional. "Must be sent together with lng."
  • lng — map center longitude, optional. "Must be sent together with lat."
  • zoom — "Map zoom from 1 to 21", and here's the part most people miss: "Defaults to 14 when coordinates are sent."

That default is a trap. If you send lat/lng without zoom, you get zoom 14 — roughly a city-district-scale view. If you assumed a tight radius, your grid cells silently overlap and you pay twice for the same businesses. Always set zoom explicitly, even when you think you want the default.

One more contract detail from the docs: sending lat without lng (or vice versa) is invalid — the pair is the unit. And page still works with coordinates, but each page is a separate request (2 credits), so paginating a grid multiplies your cost by the number of pages you walk.

What comes back, and what doesn't

The result array is places. Each MapsPlace has only three required fields: name, title (an alias for name), and feature_id. Everything else is optional, and the optional list is where the surprises live:

  • Coordinates come back as latitude/longitude — not lat/lng. The request takes lat, the response gives latitude. That asymmetry has burned me twice.
  • There is no review_count field in the schema. rating exists; review counts don't. If your analysis needs "how many reviews", the documented payload won't give it to you — plan around rating, or enrich from another source.
  • position is documented as "Maps Search only" — a rank alias within the search response.
  • Phone, website, hours, open_status, plus_code, photos and friends are all "when present" fields. Presence varies by business type: a coffee shop has hours and a phone; a park has neither.

So a robust grid collector treats name/title/feature_id as guaranteed and everything else with .get(). The feature_id is also your join key to maps/detail later (same 2 credits) if you need full hours or phone data.

The grid script

A reproducible grid needs four things: a defined area, a fixed cell size, deterministic query order, and a persisted record of what was collected. Here's the collector:

import csv, json, pathlib, requests

API = "https://api.serpbase.dev/google/maps/search"
KEY = "your_api_key"   # X-API-Key header, same as every other endpoint

def maps_search(q: str, lat: float, lng: float, zoom: int, page: int = 1) -> list:
    resp = requests.post(API,
        headers={"X-API-Key": KEY, "Content-Type": "application/json"},
        json={"q": q, "lat": lat, "lng": lng, "zoom": zoom, "page": page},
        timeout=30)
    data = resp.json()
    if data.get("status") != 0:
        raise RuntimeError(f"{data.get('status')}: {data.get('error')}")
    return data.get("places", [])

def grid(center_lat: float, center_lng: float, cells: int, step_deg: float):
    """Deterministic square grid around a center point."""
    half = (cells - 1) / 2
    for i in range(cells):
        for j in range(cells):
            yield (round(center_lat + (i - half) * step_deg, 6),
                   round(center_lng + (j - half) * step_deg, 6))

def collect(query: str, center_lat: float, center_lng: float,
            cells: int = 3, step_deg: float = 0.02, zoom: int = 14) -> list:
    """Collect the query across a grid; dedupe by feature_id; cache to disk."""
    cache = pathlib.Path(f"grid_{query}_{center_lat}_{center_lng}_{cells}_{zoom}.json")
    if cache.exists():
        return json.loads(cache.read_text("utf-8"))
    seen, rows = {}, []
    for lat, lng in grid(center_lat, center_lng, cells, step_deg):
        for p in maps_search(query, lat, lng, zoom):
            fid = p.get("feature_id")
            if not fid or fid in seen:
                continue
            seen[fid] = True
            rows.append({
                "cell": f"{lat},{lng}",
                "name": p.get("name", ""),
                "feature_id": fid,
                "rating": p.get("rating", ""),
                "address": p.get("address", ""),
                "latitude": p.get("latitude", ""),
                "longitude": p.get("longitude", ""),
                "website": p.get("website", ""),
            })
    cache.write_text(json.dumps(rows, ensure_ascii=False, indent=1), encoding="utf-8")
    return rows

if __name__ == "__main__":
    # 3x3 grid around central London at zoom 14, one query
    rows = collect("coffee shop", 51.5074, -0.1278, cells=3, step_deg=0.02, zoom=14)
    with open("grid_coffee_london.csv", "w", newline="", encoding="utf-8-sig") as f:
        w = csv.DictWriter(f, fieldnames=list(rows[0].keys()) if rows else
                           ["cell", "name", "feature_id", "rating", "address",
                            "latitude", "longitude", "website"])
        w.writeheader()
        w.writerows(rows)
    print(f"{len(rows)} unique places from {9} cells → grid_coffee_london.csv")
Enter fullscreen mode Exit fullscreen mode

Before the big run, cost it out: a 3×3 grid is 9 requests = 18 credits. A 5×5 grid over five queries is 125 requests = 250 credits. With the $10/20k standard pack that's cents, but if you walk page=2 on every cell you double it — pagination on a grid should be the exception, not the default.

Three rules for a grid you can re-run

  1. Freeze the recipe in code, not in your head. The grid center, cell count, step, zoom and query list belong in constants or a config file, and the output filename should encode them (as the cache above does). "Same grid next month" only means something if the parameters are stored, and if the cached JSON is keyed by them you literally cannot accidentally re-collect with different settings.
  2. Dedupe by feature_id, never by name or coordinates. feature_id is required and stable; names carry branch suffixes ("Coffee (Soho)") and latitude/longitude can shift slightly between runs. Coordinate-based dedupe with a distance threshold sounds rigorous but silently merges two cafés on opposite corners of an intersection.
  3. Treat title as name, and don't expect review counts. If your pipeline stores both, you're storing the same string twice — pick name and move on. And if a stakeholder asks for review volumes, the documented maps payload has no review_count field; say so before they build a dashboard on it.

FAQ

Why not just send one query without coordinates? You can — the coordinates are optional — but then the result set is whatever Google decides to center on, and it changes between runs. The grid is what makes the collection reproducible.

Does page work with lat/lng? Yes, page is a normal optional parameter. But each page costs 2 credits, so a grid × pagination multiplies fast. Collect pages only when a single page is demonstrably short for your cell.

What's the right zoom for a grid? Zoom 14 is the default for a reason: it's roughly district-scale, a good match for a grid with a small step. If you're scanning a whole metro area, larger steps at zoom 13; for a single neighborhood, zoom 15-16 with smaller steps. The point is that zoom and step are chosen together, deliberately.

Can I get phone numbers or opening hours? Not reliably from search results — they're "when present" fields. For a specific business you need, take its feature_id and call maps/detail (also 2 credits) with that exact 0x...:0x...-formatted id.

Take the script, replace the center with your own city, and run the smallest grid first — one cell, one query — to see how dense your category is before you scale to 25 cells and a 250-credit run.

Top comments (0)