DEV Community

Siftwright
Siftwright

Posted on Originally published at apify.com

Scraping Google News search results with the public RSS feed (no API key, no browser)

Google News doesn't have a public search API. If you want structured articles for a keyword — title, source, link, publish date — the usual options are a paid SERP API or scraping the HTML search results page (fragile, breaks often, and against Google's ToS for the main search UI).

There's a third option that's easy to miss: Google News publishes an unauthenticated RSS feed for any search query, at news.google.com/rss/search?q=.... It's meant for feed readers, but it's public, stable, and returns clean structured data. This tutorial builds a small scraper around it.

The endpoint

GET https://news.google.com/rss/search?q=<query>&hl=en-US&gl=US&ceid=US:en
Enter fullscreen mode Exit fullscreen mode

No API key, no auth, no proxy needed — it's a plain public RSS/XML endpoint. Each <item> in the feed has a title, link, source, and publish date, which is really all "Google News search results" means in structured form.

Parsing it

import { parseStringPromise } from 'xml2js';

async function searchGoogleNews(query) {
    const url = `https://news.google.com/rss/search?q=${encodeURIComponent(query)}&hl=en-US&gl=US&ceid=US:en`;
    const res = await fetch(url);
    const xml = await res.text();
    const parsed = await parseStringPromise(xml, { explicitArray: false });

    const items = parsed?.rss?.channel?.item;
    if (!items) return [];

    const list = Array.isArray(items) ? items : [items];
    return list.map((item) => ({
        title: item.title,
        link: item.link,
        source: item.source?._ ?? item.source,
        publishedAt: item.pubDate,
        snippet: item.title, // Google's feed doesn't include a separate summary
    }));
}
Enter fullscreen mode Exit fullscreen mode

That's genuinely the whole scraper. No headless browser, no proxy rotation, no CAPTCHA handling — it's a public feed, so it behaves like one.

Real output

Running it for "artificial intelligence regulation" today returns real, current articles like:

{
    "title": "Bill Gates says unchecked AI could 'cause a billion deaths' in call for regulation",
    "link": "https://news.google.com/rss/articles/CBMi...",
    "source": "The Guardian",
    "publishedAt": "2026-09-27T09:00:00.000Z",
    "snippet": "Bill Gates says unchecked AI could 'cause a billion deaths' in call for regulation The Guardian"
}
Enter fullscreen mode Exit fullscreen mode
{
    "title": "Trump rejects AI regulation, citing parallels with climate change, in U.N. address",
    "link": "https://news.google.com/rss/articles/CBMi...",
    "source": "Scientific American",
    "publishedAt": "2026-09-22T16:00:00.000Z",
    "snippet": "Trump rejects AI regulation, citing parallels with climate change, in U.N. address Scientific American"
}
Enter fullscreen mode Exit fullscreen mode

(Links are truncated here for length — the real output includes the full Google News redirect URL, which resolves to the original article.)

What to watch for

  • No summary field. Google's RSS feed repeats the title as the description, so don't expect a real snippet/abstract — if you need one, you'd have to fetch the linked article separately.
  • Links are Google redirect URLs, not the publisher's direct URL. They resolve fine in a browser or with a follow-redirects fetch, but store them as-is if you just need a working link.
  • Rate limits are unpublished. It's a public feed, not a documented API — keep request volume reasonable and add backoff if you see errors.
  • Empty results are normal, not an error. An obscure or misspelled query can legitimately return zero items — don't treat that as a failure.

Try it hosted

I turned this exact pattern into a small pay-per-event Apify Actor — Google News Scraper — so you don't have to run and maintain the scraper yourself. It charges $0.0015 per article returned and nothing for empty or failed searches, and it's callable directly by AI agents through Apify's MCP server if you're wiring this into an agent pipeline instead of a script.

Source and more real examples: github.com/siftwright/siftwright-tools

Top comments (0)