DEV Community

Devil Scrapes
Devil Scrapes

Posted on

Meetup Events Scraper: the search page ships its whole Apollo cache, and that is the API

Quick answer

The Meetup Events Scraper pulls public Meetup.com events by keyword and city into JSON, CSV or Excel: title, start time with timezone, in-person or online, venue with coordinates, the aggregate "going" count, capacity, fee, and the hosting group with its Meetup URL. No login and no API key. Meetup's own API is positioned for Meetup Pro customers, so we read the data the public search page already embeds instead: a normalized Apollo GraphQL cache inside the page's __NEXT_DATA__ script tag. Pricing is pay-per-result, $0.20 per run plus $0.004 per event, $4.20 per 1,000 events.

Does Meetup have a public API?

Not for this. Meetup's help centre describes its API as a way for Meetup Pro customers to build integrations and automate the networks they manage. For the question most people actually have, "which AI meetups are happening in New York next month and how many people are going", that means a paid organiser subscription and an OAuth flow to reach data that is already printed on a public web page. The search page at meetup.com/find/?keywords=ai&location=New%20York%2C%20NY renders without a session, and the server sends more than the HTML.

Where does the data actually come from?

Meetup's site is a Next.js app, and like every Next.js page it carries a <script id="__NEXT_DATA__"> tag with the props the page was rendered from. Under props.pageProps.__APOLLO_STATE__ sits the Apollo Client cache for that render: the GraphQL result of the search, flattened into entities keyed by type and id.

In our golden fixture that cache has 11 keys: one ROOT_QUERY, five Event:* entities and five Group:* entities. The search result itself lives under a ROOT_QUERY key that embeds the filter in its name, something like eventSearch:{"filter":{"city":"New York","country":"us",...}}, and each edge's node is not an event but a pointer: {"__ref": "Event:315229748"}. The event entity points at its group the same way. Venue and RSVP data are nested inline.

Meetup's public search page embeds the GraphQL response it was rendered from, as a normalized Apollo cache, in a script tag anyone can read.

So the parser is a small dereferencer, not a DOM walker:

def _make_deref(apollo: dict) -> Callable[[Any], dict]:
    def deref(value):
        if isinstance(value, dict) and "__ref" in value:
            return apollo.get(value["__ref"]) or {}
        return value if isinstance(value, dict) else {}
    return deref

event = deref(edge["node"])          # Event:315229748
group = deref(event.get("group"))    # Group:29431902
venue = deref(event.get("venue"))    # inline dict, passes through
Enter fullscreen mode Exit fullscreen mode

What goes wrong if you treat the cache as a flat list?

Four things, each of which we hit once and now test for.

  • The search key is not stable. It contains the serialized filter, so a hardcoded key breaks the moment the location changes. We match on the eventSearch prefix and take the first hit.
  • Nodes are pointers. Read edge["node"]["title"] and you get a KeyError, because the node is {"__ref": ...}. Every __ref has to be resolved against the cache, including the group inside the event.
  • "Free" is the absence of a field. There is no isFree boolean. A free event has feeSettings: null; a paid one carries {amount, currency}. We derive is_free from that and emit fee_amount and fee_currency only when a fee exists.
  • Online events still have a venue. It is a placeholder with blank strings. We collapse blanks to null, so an online event reads venue_name: null rather than "".

The RSVP count comes from rsvps.totalCount, an aggregate. The Actor emits that number and nothing about who is going. Member identities are never requested, which keeps the dataset GDPR-safe by construction rather than by a filter you have to trust.

What the Actor gives you

One row per public event, 23 fields: search_keyword, event_id, title, event_url, date_time (ISO-8601 with offset, e.g. 2026-07-21T08:30:00-04:00), event_type and is_online, going_count, max_tickets, is_free, fee_amount, fee_currency, venue_name, venue_address, venue_city, venue_state, venue_country, latitude, longitude, group_name, group_url, description and scraped_at. Pass several keywords in one run, optionally a location, and set eventType to any, physical or online.

Honest limitations 🚧

One search page is one request and yields roughly 27 events per keyword, in Meetup's own relevance order. We do not paginate past it, so breadth comes from more keywords, not deeper pages. maxEventsPerKeyword caps at 500. A keyword with no matches is a successful run and costs only the start fee. Descriptions are Meetup's markdown, kept verbatim; switch includeDescription off for a leaner row.

FAQ

Do I need a Meetup account or API key?
No. The Actor reads public search pages only.

What about blocks and rate limits?
That is our job. Every request replays a real browser TLS handshake through curl-cffi, rotating between Chrome, Firefox and Safari profiles, runs through Apify Proxy, and retries with exponential backoff on 403, 408, 429 and 5xx responses, up to 5 attempts per page with the delay capped at 30 seconds.

Does it collect attendee data?
No. Only the aggregate going count. No names, no profiles, no member IDs.

Can I get a specific group's events rather than a keyword search?
Not in this version. Use the group's name as a keyword and filter on group_url in the output.

How does pricing work?
$0.20 per run plus $0.004 per event row: $4.20 per 1,000 events. No data, no charge beyond the start fee.

→ Meetup Events Scraper on Apify


Built by Devil Scrapes. We handle the fingerprints, the proxies and the retries, and we test the parser against a real search payload before a customer ever runs it. We do the dirty work so your dataset stays clean. 😈

Top comments (1)

Collapse
 
ispyhumanfly profile image
Dan Stephenson •

The Apollo-cache-as-API approach is clever. Discovering that the search page ships its whole cache means you skip a lot of fragile DOM parsing. Have you seen the cache key structure change across Meetup's A/B experiments, or has the payload stayed stable enough that you haven't needed a fallback?