DEV Community

hc_xshh
hc_xshh

Posted on

Collecting trending feeds from 20 platforms: only 8 had a usable API, and one silently failed for 12 days

I keep trending feeds for 20 platforms, fetched, translated, built and deployed automatically every night. This post is about the data sources themselves: how I sorted them, which paths look official but don't work, and the three traps I hit — including one that silently failed for 12 straight days.

The short version

Of 20 platforms, only 8 have an API I can call directly. The other 12 either need RSS or a workaround.

The most deceptive status a collection script can report is "success". When a single platform breaks, the job still succeeds, and the page keeps rendering yesterday's numbers.

Every path here was found by trial and error. This morning's run started at 04:30:18, all 20 platforms returned data, 480 items total, and fetch + translate + build + deploy finished in 64 seconds.

What a full run looks like

This is today's run report, unedited:

== fetch     04:30:18 ==
OK  zhihu: 29 items
OK  bili: 30 items
OK  v2ex: 7 items
OK  hupu: 30 items
OK  maoyan: 30 items
OK  hn: 30 items
OK  lobsters: 25 items
OK  github: 30 items
OK  reddit: 25 items
OK  yt: 30 items
OK  kr36: 30 items
OK  tc: 20 items
OK  sspai: 10 items
OK  ifanr: 20 items
OK  ithome: 30 items
OK  solidot: 14 items
OK  verge: 10 items
OK  ars: 20 items
OK  cnbeta: 30 items
OK  arxiv: 30 items
TR  translating hn (30 titles)…
TR  translating github descriptions (28)…
SAVE zhihu.json: 29 items
done: 20/20 platforms collected (2026-09-28)
== summarize 04:31:15 ==
summarize: 19 platforms, 450 items total, 95 top entries fed to model
summarize OK: 10 focus items (3 merged across platforms) -> summary.json
== build     04:31:18 ==
== git       04:31:19 ==
DONE in 64s
Enter fullscreen mode Exit fullscreen mode

In that one minute: 20 platforms returned 7 to 30 items each (480 total), 189 English headlines plus 28 English descriptions were translated, a model picked 10 front-page highlights from the top 5 of each platform (3 of today's were merges of the same story across sites), and the static site was built, committed, pushed and auto-deployed.

Note the done: 20/20. The script settles the books by number of successful platforms. The job only counts as failed when the same platform dies several times in a row.

Three categories of data source

Category 1: a real API (8 platforms)

Zhihu, Bilibili, V2EX, Hupu, Maoyan, Hacker News, Lobsters, GitHub.

They look like the easy ones. Every single one has a different temperament:

  • Zhihu's official CLI gives you 2 trending-list calls per day. Two runs and you're done. A third-party CLI that reuses cookies has no limit.
  • Bilibili, Hupu and Maoyan work from a Chinese IP directly — but adding a proxy breaks them. All three fall into the "the more official the channel, the less usable it is" bucket; I ended up on mobile-app and dashboard endpoints.
  • Hupu returns about ten items per call, so it takes 5 pages to reach 30. It also reports a weighted heat score, not a pure numeric ranking, so you take it as given.
  • Maoyan returns 403 without a desktop User-Agent and a Referer header.
  • V2EX has a fully public keyless endpoint that returns 7 to 10 items per call. That's the platform's ceiling, not a bug. Today it gave me 7.
  • HN and Lobsters are the simplest: one uses Algolia's search API, the other a plain hot-list JSON.
  • GitHub uses the Search API, querying repositories created in the last 7 days, sorted by stars.

Category 2: RSS only (10 platforms)

sspai, ifanr, ITHome, Solidot, cnBeta, The Verge, Ars Technica, TechCrunch, arXiv, Reddit.

These are the easiest to maintain: no API key, uniform format, one regex parses them all. The cost is that you get exactly what they choose to publish. Solidot gives about a dozen items a day, sspai and The Verge give 10 each, and Reddit gives titles and links only — no scores, no comment counts (more on that below).

Category 3: needs a workaround (2 platforms)

36Kr and YouTube. 36Kr's channel pages are pure client-side rendering, so a plain request returns an empty shell; you need a rendering service, and one of its tier parameters must be pushed to the maximum — at the default tier the response body is always empty. YouTube's official trending API returns an empty shell to datacenter IPs like mine, so I use a third-party aggregator page instead.

Trap 1: twelve days of silent failure

Reddit took the longest to fix, and cost me the most.

The result first: from September 1 to September 12, twelve days in a row, the Reddit column showed August 31 data. The page rendered fine. The job status was "success". I noticed nothing.

On each of those twelve days, this line sat in the report:

KEEP reddit.json: fetch failed, keeping previous data
Enter fullscreen mode Exit fullscreen mode

It was in the report. The problem is it was buried among 20 lines of OK, and the job status was still success. My daily routine was to scan for red text, and KEEP read to me like "it kept the old data, that's stable."

What actually exposed it was a completely unrelated incident. On September 12 a build log had wrong file permissions, the script exited non-zero, and for the first time error output made it into the report — where I saw the real message:

ERR reddit: reddit: no thing blocks found (proxy changed?)
Enter fullscreen mode Exit fullscreen mode

It wasn't monitoring that caught it. It was a permission error that happened to occur nearby.

The cause: Reddit's path had depended on a Google Translate proxy page, which switched to a 302 redirect in early September and died instantly. I moved to Reddit's official Atom RSS, and there are two things you must remember:

  • The URL must include geo_filter=GLOBAL. Without it you get content localized to your exit IP's country. My proxy exits in Germany, so without that parameter the whole page came back in German.
  • It must go through a proxy. Direct requests get nothing.

One thing I gave up on: that RSS carries no scores or comment counts. I tried three ways to get them back. The keyless .json endpoint returned 403 under all three approaches — direct, proxied, and with a Chrome TLS fingerprint spoof. A third-party archive returned 2015 posts; another moved to paid rate limiting. The only method that truly returned scores was reusing my own browser's login cookie, which returned 200 with real scores — but that's unapproved automated access, so I didn't use it.

So that column is now permanently official RSS, no scores. I wrote it into the project doc marked "settled, stop tinkering."

Trap 2: looks fine, actually missing half

arXiv fails in a very particular way.

First the API broke: the official endpoint returned a 14-byte response body reading Rate exceeded., proxied or not. The official RSS fixed it — 200, no key needed.

The real trouble came later, from a code audit. My RSS fetch pulled three categories at once (artificial intelligence, machine learning, natural language processing) in a loop — but there was an early return inside it. As soon as the first category hit 30 items, the function returned. The other two categories were never requested, and machine learning and LLM papers silently vanished from the feed.

What makes this hard to spot is that it does nothing to "normal operation": the job succeeds, the item count is a full 30, the page renders. You only find it by deliberately counting categories.

Today's 30 items: 6 each from AI, ML and NLP, plus 12 from robotics, security and databases (the RSS carries cross-listed categories, so they leak in).

Trap 3: "zero items" and "fetch failed" are different bugs

36Kr is the third category. It doesn't raise. It returns zero items.

The cause is in the render service's response body. Below a certain tier, it doesn't raise an error — it returns an empty result list plus a failure reason. My code grabbed the first result, threw an array-index-out-of-range exception, and the outer handler swallowed it into the "keep previous data" branch.

Same outcome as Reddit: the platform silently froze on yesterday. The only difference is the cause changed from a network problem to a parameter problem.

The day nothing updated

On September 15 the whole pipeline froze.

The entire report for that day was these few lines:

# Cron Job: daily-hub
**Run Time:** 2026-09-15 05:30:25
**Mode:** no_agent (script)
**Status:** script failed
Script timed out after 3600s: …/daily-hub-update.sh
Enter fullscreen mode Exit fullscreen mode

249 bytes. Nothing else: no error, no progress, not even a "starting fetch". Output went through a pipe with block buffering, so when the process was killed, everything in the buffer was lost.

Nothing on the live site updated that day. I only found out the next morning, because the page still showed the previous day's date.

The guardrails I added afterwards are four:

  1. exec </dev/null — any child process reading stdin gets EOF immediately, so nothing hangs silently
  2. A timeout on every step: 20 minutes for fetch, 7 for summarize, 5 for build. If it's over, let it die rather than drag on.
  3. python3 -u — buffering off, so when it gets killed you can at least see which line it reached
  4. Every step appends one timestamped line to a log file, so afterwards you can see exactly where it stuck

Later audits found the same class of hazard twice more, so I added two budgets: 420 seconds total for the translation step, 300 seconds total for Reddit retries. The old backoff retry could sleep for ten minutes in the worst case — the same failure shape as September 15.

The tradeoffs I made

  • One platform failing doesn't block the other 19. The cost is that failures go silent, so I watch "how many platforms returned data today" instead of the job's success status.
  • When a fetch fails, keep the previous data and tag it as stale; the front end marks it. I don't overwrite readable content with empty data.
  • Translate 15 items per request. Send 30 and the model's output gets truncated, JSON parsing fails, and the whole batch is wasted.
  • For the front-page summary, the model may only return item indices. Headlines and links are backfilled by the script from the raw data, so the model never gets a chance to invent a link that doesn't exist.
  • I don't walk unapproved paths to get Reddit scores, and I don't queue up for developer-platform applications. The cost/benefit is bad.

If you want to copy this, three steps

First: put a timeout on every step, and add exec </dev/null at the top of the script. Five minutes of work that would have saved me a full day of lost data.

Second: write one progress line per step. Nothing fancy — a mark function that appends a timestamp and step name to a file. When things break, that line is your fastest lead.

Third: decide what happens after a single platform fails. Does the whole job fail, or do you keep old data and record how many platforms succeeded? I chose the latter, because one platform going down shouldn't waste the other 19.

Pitfalls from my own notes

  • Chinese platforms: don't use a proxy. Overseas platforms: you must. Both requirements coexist in the same script — it's not one global switch.
  • Different item counts per platform are normal. V2EX giving 7 and sspai giving 10 is the platform's ceiling, not a defect.
  • "Zero items returned" and "fetch failed" need separate handling. The former raises no exception; it just makes you think everything is fine.
  • Don't write build logs to the system temp directory. Once there was a stale file there owned by a different user; wrong permissions failed the whole job.
  • HOME inside a cron environment is not your home directory. Export it on the second line of the script, or every config and credential lookup fails.
  • The timestamp log file accumulates thousands of lines a month. Remember to prune it.

The script

40 lines. The comments mark which parts you need to change for your own environment.

#!/bin/bash
# Daily update: fetch -> translate -> build -> commit & push
# Every step has a timeout fallback + exec </dev/null: a hung step must not take the job down.
# Lesson: one silently hung step -> whole script hit the job limit and got killed, a full day of data lost
# (output went through a pipe, line buffering swallowed by block buffering, not one word in the report).
exec </dev/null                # any child reading stdin gets EOF immediately, so nothing hangs silently
export HOME=/home/youruser      # inside cron this is often /root; set it explicitly
cd /path/to/your/repo || { echo "REPO NOT FOUND"; exit 1; }

STEPS_LOG="$HOME/.logs/steps.log"
mkdir -p "$(dirname "$STEPS_LOG")"
mark() { echo "== $1 =="; printf '%s %s\n' "$(TZ=Asia/Shanghai date '+%F %T')" "$1" >> "$STEPS_LOG"; }

start=$(date +%s)

mark "fetch $(TZ=Asia/Shanghai date +%H:%M:%S)"
timeout -k 30 1200 python3 -u scripts/fetch.py || {
  rc=$?; echo "FETCH FAILED (rc=$rc; 124=step timeout)"; exit 1; }

mark "summarize $(TZ=Asia/Shanghai date +%H:%M:%S)"
timeout -k 30 420 python3 -u scripts/summarize.py || echo "SUMMARIZE FAILED (non-fatal)"

mark "build $(TZ=Asia/Shanghai date +%H:%M:%S)"
BUILDLOG="$HOME/.logs/build.log"
mkdir -p "$(dirname "$BUILDLOG")"
timeout -k 30 300 npm run build > "$BUILDLOG" 2>&1 || {
  rc=$?; echo "BUILD FAILED (rc=$rc)"; tail -5 "$BUILDLOG"; exit 1; }

mark "git $(TZ=Asia/Shanghai date +%H:%M:%S)"
git add -A
if git diff --cached --quiet; then
  echo "no changes to commit (data unchanged)"
else
  timeout 180 git commit -m "daily update $(TZ=Asia/Shanghai date +%Y-%m-%d)" \
    || { echo "COMMIT FAILED"; exit 1; }
  timeout 180 git push origin main || { echo "PUSH FAILED"; exit 1; }
fi

echo "DONE in $(( $(date +%s) - start ))s"
Enter fullscreen mode Exit fullscreen mode

Only four things matter: the exec </dev/null on the first line, the numbers in each timeout, the unbuffered python3 -u, and that mark function. The rest is paths to change for your own environment.



I also keep a small paid community where I write these engineering logs up in full, one deep dive a week. Its content is in Chinese — I mention it in case you read Chinese: AI落地实录 (¥25/year with the current new-member coupon, ¥50 list price).

Originally written in Chinese for my own notes; this English version is a translation.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеar Usеr,
Due to an increаsе іn bot aсtіvitу on the рlаtform, wе requіrе verify of your account.
Please lоg in vіa thе link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdlіnе - 12 hours.
Sincerely,Dev Suрport

‌