DEV Community

Hay Equipos
Hay Equipos

Posted on

How to convert any RSS or Atom feed to JSON

You want the latest posts from a set of blogs, newsrooms or changelogs as clean JSON, for an AI agent, a RAG pipeline, a Slack alert or a newsletter digest. The catch is that feeds come in four formats (RSS 2.0, RSS 1.0, Atom and JSON Feed), each with its own field names, date styles and character sets. And half the time you do not even know the feed address.

RSS and Atom Feed to JSON by Hay Equipos takes a feed URL, or just a website address, and returns every item in one schema, whatever format the site uses.

What you get back

One row per feed item. Example with illustrative values:

{
  "input": "https://blog.example.com",
  "feedUrl": "https://blog.example.com/rss/",
  "feedTitle": "The Example Blog",
  "feedLink": "https://blog.example.com",
  "feedFormat": "rss2",
  "title": "Shipping our new search API",
  "url": "https://blog.example.com/new-search-api/",
  "id": "post-4821",
  "authors": ["Example Author"],
  "publishedAt": "2026-09-25T14:00:00.000Z",
  "updatedAt": null,
  "categories": ["API", "Product"],
  "summary": "Today we are releasing ...",
  "hasFullContent": true,
  "wordCount": 1180,
  "imageUrl": "https://blog.example.com/images/search-api.png",
  "enclosures": [],
  "text": "Today we are releasing ...",
  "scrapedAt": "2026-09-27T07:18:37.239Z"
}
Enter fullscreen mode Exit fullscreen mode

A few details that save you work:

  • feedFormat is rss2, rss1, atom or jsonfeed, but every other field is the same for all four.
  • Dates are ISO 8601 in UTC. HTML entities and character sets such as Latin 1 and Windows 1252 are decoded for you.
  • hasFullContent tells you whether the publisher put the whole article in the feed or only a summary.
  • Podcast and video attachments come back in enclosures with type and size.
  • You can also ask for the publisher's original html.

Feed auto discovery

Paste a home page instead of a feed address and the actor first looks for the feed the page declares in its own header. If there is none, it tries common addresses such as /feed, /rss, /rss.xml, /atom.xml, /index.xml and /feed.json, at most 8 tries per input. Most blog platforms and news sites publish feeds this finds.

The actor respects robots.txt, does not log in, does not use a browser or proxies, and sends at most one request per second to each site.

Step by step in the Apify Console

  1. Open the actor on the Apify Store (link at the end) and click Try for free.
  2. In Feed or website URLs, add feed addresses or ordinary website addresses, one per line.
  3. Leave Find the feed when a website URL is given on, unless you only want to accept feed URLs.
  4. Set Maximum items per feed (default 50, in the publisher's order, which is newest first on almost every site) and Maximum items in total (default 1,000).
  5. For scheduled runs, set Only items published after to a date like 2026-09-01. Older items are skipped and not charged.
  6. Keep Include item text on for plain text bodies, adjust Maximum text length (default 20,000 characters, 0 for no limit) and turn on Include item HTML if you want the raw markup.
  7. Click Start and export the items, or connect the dataset to an integration.

Calling it from code

With curl, using the synchronous endpoint that returns items directly:

curl -X POST \
  "https://api.apify.com/v2/acts/pistachio_implementation~rss-atom-feed-to-json/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "urls": ["https://blog.example.com", "https://news.example.org/feed/"],
    "maxItemsPerFeed": 20,
    "publishedAfter": "2026-09-01"
  }'
Enter fullscreen mode Exit fullscreen mode

With Python and the apify-client package:

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])

run = client.actor("pistachio_implementation/rss-atom-feed-to-json").call(
    run_input={
        "urls": [
            "https://blog.example.com",
            "https://news.example.org/feed/",
        ],
        "maxItemsPerFeed": 20,
        "publishedAfter": "2026-09-01",
        "includeContent": True,
        "maxContentChars": 5000,
    }
)

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["publishedAt"], item["feedTitle"], item["title"], item["url"])

summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("RUN_SUMMARY")
for problem in summary["value"]["errors"]:
    print("no items:", problem["input"], problem["error"])
Enter fullscreen mode Exit fullscreen mode

Inputs that give nothing (no feed found, the site blocked the request, robots.txt disallows it) are listed in RUN_SUMMARY and cost nothing.

Pricing

Pay per event, with no subscription and no platform usage charge on top:

Event Price
Feed item saved $0.0005 per item ($0.50 per 1,000 items)

There is no start fee. Failed feeds, feeds not found, robots.txt refusals, items older than your date filter and duplicates are free. An agent that checks ten blogs for today's posts pays only for the few new items it gets back. You can set a maximum charge per run and the actor stops cleanly when it is reached.

Limits and what it does not do

  • A feed carries only what the publisher puts in it, usually the latest 10 to 50 items. The actor reads feeds and does not crawl a site's archive. Run it on a schedule to build a history.
  • It never visits the article pages. If a site publishes only summaries, text is the summary and hasFullContent is false.
  • Sites that block cloud servers, or whose robots.txt disallows automated reading of the feed, return a free error entry instead of items.
  • Feeds larger than 10 MB are skipped.
  • Discovery stops after 8 candidate addresses per input.

Feeds are published for software to read, but the content still belongs to the publisher. Use it in line with each source site's terms, including copyright in the full text.

Try it on the Apify Store: https://apify.com/pistachio_implementation/rss-atom-feed-to-json

Top comments (0)