DEV Community

quiethand098
quiethand098

Posted on

How to get real article URLs from Google News RSS (decoding the CBMi links)

Google News has a free RSS feed, but every <link> in it points to news.google.com/rss/articles/CBMi... instead of the publisher. If you feed those into a crawler or an LLM pipeline you get Google's redirect page, not the article. Here is how the links work and how to decode them.

The feed

https://news.google.com/rss/search?q=openai+when:7d&hl=en-US&gl=US&ceid=US:en
Enter fullscreen mode Exit fullscreen mode

when:7d limits the time window, site:reuters.com and quoted phrases work as in normal search. Topics live at /rss/headlines/section/topic/TECHNOLOGY. You get up to about 100 items per query.

Decoding the link

The last path segment is URL-safe base64. Decode it and you get a protobuf-like blob.

  1. Strip the \x08\x13\x22 prefix and a \xd2\x01\x00 suffix if present.
  2. Read the length byte (two bytes if it is 0x80 or more) and take that many characters.
  3. If the string does not start with AU_yqL, it is the publisher URL. Done, no network call.
  4. Otherwise (newer links) request https://news.google.com/rss/articles/<id>, read data-n-a-sg and data-n-a-ts from the HTML, and POST a Fbv4je call to /_/DotsSplashUi/data/batchexecute. The second JSON array in the response holds the real URL.
const id = new URL(src).pathname.split('/').pop();
let b = Buffer.from(id.replace(/-/g, '+').replace(/_/g, '/'), 'base64').toString('latin1');
if (b.startsWith('\x08\x13\x22')) b = b.slice(3);
if (b.endsWith('\xd2\x01\x00')) b = b.slice(0, -3);
let len = b.charCodeAt(0), start = 1;
if (len >= 0x80) { len = (len & 0x7f) | (b.charCodeAt(1) << 7); start = 2; }
const url = b.slice(start, start + len);
if (!url.startsWith('AU_yqL')) console.log(url); // else: batchexecute call
Enter fullscreen mode Exit fullscreen mode

Google changes this occasionally, so keep a fallback to the redirect link.

Don't want to maintain it?

I packaged the whole thing as an Apify actor: multiple queries per run, topics, country and language, time ranges, dedupe, real URLs. No proxy and no API key needed. Pay per article (about $3 per 1,000).

Same author, two more small actors if you build data pipelines: a YouTube transcript extractor with RAG chunks and a remote jobs aggregator.

Top comments (0)