If you have ever tried to feed news articles to an LLM, you know the boring part is not the summary. It is everything before it: cookie banners, "related stories", share buttons, newsletter boxes, author bios, and pages that are not articles at all (a home page, a section front, a login wall).
I built a small API that does the boring part and then, only if you ask, the summary. Below are real calls I made on 1 October 2026, with the actual responses.
Step 1: just the article (no AI)
curl --request POST \
--url 'https://article-extractor-ai-summarizer-api.p.rapidapi.com/api/v1/article' \
--header 'X-RapidAPI-Key: YOUR_RAPIDAPI_KEY' \
--header 'X-RapidAPI-Host: article-extractor-ai-summarizer-api.p.rapidapi.com' \
--header 'Content-Type: application/json' \
--data '{"url": "https://www.theguardian.com/technology/2026/sep/29/oura-ring-public-offering", "maxChars": 400}'
The response came back in 0.48 s (plain HTTP, no browser needed). Trimmed to the fields that matter:
{
"success": true,
"isArticle": true,
"pageType": "article",
"articleType": "NewsArticle",
"title": "Smart ring maker Oura puts off initial public offering due to market ‘uncertainty’",
"authors": ["Guardian staff reporter", "Associated Press"],
"publishedAt": "2026-09-29T17:57:01.000Z",
"modifiedAt": "2026-09-29T18:20:35.000Z",
"publisher": "The Guardian",
"language": "en",
"tags": ["Technology", "IPOs", "AI (artificial intelligence)", "Business", "US news"],
"paywalled": false,
"wordCount": 296,
"readingTimeMinutes": 1,
"tokensEstimate": 445,
"format": "markdown",
"markdown": "…(body as Markdown, cut at 400 characters because of maxChars)…",
"truncated": true,
"summaryGenerated": false
}
A few details I cared about:
-
authorsis an array, not a joined string, and the dates are ISO 8601, so you can filter by freshness without parsing. -
wordCountandtokensEstimatedescribe the whole article even whenmaxCharscuts the body. Handy for deciding whether a page fits your context window before you download all of it. -
formatcan bemarkdown,textorboth.
Step 2: the same page, with a TL;DR
The summary lives on a separate endpoint, /api/v1/article/summary, so extraction-only calls never touch the model (and never use your summary quota).
curl --request POST \
--url 'https://article-extractor-ai-summarizer-api.p.rapidapi.com/api/v1/article/summary' \
--header 'X-RapidAPI-Key: YOUR_RAPIDAPI_KEY' \
--header 'X-RapidAPI-Host: article-extractor-ai-summarizer-api.p.rapidapi.com' \
--header 'Content-Type: application/json' \
--data '{"url": "https://www.theguardian.com/technology/2026/sep/29/oura-ring-public-offering", "summary": "tldr"}'
1.74 s, and the summary block looked like this:
"summary": {
"style": "tldr",
"language": "same as source",
"maxWords": 30,
"text": "Oura postpones its initial public offering due to market uncertainty despite strong demand and revenue growth.",
"inputChars": 1866,
"inputTruncated": false,
"tokens": { "input": 575, "output": 20 }
}
You get the token counts back, so you can see what the model actually read.
Bullets, in another language
Same article, "summary": "bullets", "maxWords": 80, "language": "Spanish" — 3.92 s, 7 bullets. The first three:
- Oura Inc está retrasando su oferta pública inicial debido a la incertidumbre del mercado.
- La empresa había planeado vender 50m acciones en la oferta pública inicial a un precio entre $40 y $44.
- Casi tres cuartas partes de las acciones serían vendidas por accionistas actuales.
There is also "summary": "paragraph" and a focus option (for example "focus": "numbers and dates"). If you already have the text, send text instead of url.
Step 3: what happens with a page that is not an article
This is the part I wanted most. Summarising a home page gives you a confident paragraph about nothing. So the API checks the page type first and refuses before any AI runs:
--data '{"url": "https://www.bbc.com/news"}'
{
"success": false,
"error": {
"code": "not_an_article",
"message": "This page is not an article: this page lists articles (a section, category or tag page) rather than being one.",
"details": { "pageType": "listing", "finalUrl": "https://www.bbc.com/news" }
}
}
That came back in 0.67 s with HTTP 422. If you do want the page anyway, pass "allowNonArticle": true and you will get it with isArticle: false.
Section fronts were the hard case. My first test run let the Guardian's technology section through as an "article" (no author, no date). The check now also looks at the page's structured data (a CollectionPage is not an article, even when og:type says so), at how much of the page is headline cards with short teasers, and at URL shapes like /section/, /category/ and /tag/. In a re-run, 15 section, category and tag pages (Guardian, BBC, Reuters, Spiegel, TechCrunch, The Verge, Medium, Cloudflare's blog and others) all came back 422 with pageType: listing, and 13 real articles (news in English, German and Japanese, Wikipedia in English and Japanese, a Paul Graham essay, Medium, Substack, Cloudflare and personal blog posts) were all still accepted. It is still a heuristic, so if you feed it mixed URLs, pageType and publishedAt are worth a glance.
And some sites simply refuse automated requests. One news site I tried answered the first extraction and then returned HTTP 405 to the next calls; the API reported that as 422 target_blocked instead of making something up.
Python: summarise a list of links
import requests
HOST = "article-extractor-ai-summarizer-api.p.rapidapi.com"
HEADERS = {"X-RapidAPI-Key": "YOUR_RAPIDAPI_KEY", "X-RapidAPI-Host": HOST}
def tldr(url: str) -> str | None:
r = requests.post(f"https://{HOST}/api/v1/article/summary",
json={"url": url, "summary": "tldr"}, headers=HEADERS, timeout=60)
body = r.json()
if not body.get("success"):
print(url, "->", body["error"]["code"]) # not_an_article, target_blocked, timeout...
return None
return body["summary"]["text"]
for link in ["https://www.theguardian.com/technology/2026/sep/29/oura-ring-public-offering",
"https://www.bbc.com/news"]:
print(link, "=>", tldr(link))
Limits
- Web pages only (for PDFs and Word files use a document converter).
- The model reads up to about 6,000 tokens of the article; longer articles are summarised from the start, and
inputTruncatedtells you. - At most one real browser render per call, for sites that need JavaScript; most news pages do not.
- 30 s timeout.
Price
On RapidAPI there is a free plan (200 requests and 30 summaries a month, hard limit; RapidAPI may ask for a card even for free plans). The plans as of 1 October 2026:
| Plan | Per month | Extraction requests | Summaries |
|---|---|---|---|
| Basic | $0 | 200 | 30 |
| Pro | $12.99 | 10,000 | 1,500 |
| Ultra | $39.99 | 50,000 | 6,000 |
| Mega | $99.99 | 150,000 | 15,000 |
Every call counts as a request; only /article/summary calls count as summaries. RapidAPI also has its own bandwidth fee above 10 GB a month, which is theirs, not mine.
Try it: https://rapidapi.com/tidytools/api/article-extractor-ai-summarizer-api
If you would rather run it in bulk without code, the same summariser is an Apify Actor, Article Summarizer & Web Page Summarizer - AI Text Summarizer ($6 per 1,000 pages).
Disclosure: I built this API and the Apify Actor, and I earn money when people use them. All outputs above are real responses from 1 October 2026; the article quoted is Associated Press/Guardian reporting, used here only as a test input.
Top comments (0)