<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Programming with Jack Chew</title>
    <description>The latest articles on DEV Community by Programming with Jack Chew (@programming_withjackche).</description>
    <link>https://dev.to/programming_withjackche</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3948462%2F62622315-576c-4dea-9153-4cda4503f812.jpg</url>
      <title>DEV Community: Programming with Jack Chew</title>
      <link>https://dev.to/programming_withjackche</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/programming_withjackche"/>
    <language>en</language>
    <item>
      <title>Reading links your agent can't open: Xiaohongshu, Douyin, TikTok, YouTube and X as an MCP tool</title>
      <dc:creator>Programming with Jack Chew</dc:creator>
      <pubDate>Sat, 12 Sep 2026 02:56:14 +0000</pubDate>
      <link>https://dev.to/programming_withjackche/reading-links-your-agent-cant-open-xiaohongshu-douyin-tiktok-youtube-and-x-as-an-mcp-tool-4mi9</link>
      <guid>https://dev.to/programming_withjackche/reading-links-your-agent-cant-open-xiaohongshu-douyin-tiktok-youtube-and-x-as-an-mcp-tool-4mi9</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://linkdigest.dev/blog/reading-social-links-from-an-agent?ref=devto" rel="noopener noreferrer"&gt;linkdigest.dev&lt;/a&gt;, where I build this.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem, in one paragraph
&lt;/h2&gt;

&lt;p&gt;Paste a Xiaohongshu, Douyin, TikTok, YouTube or X link into an AI agent and it fetches the URL, gets an app-download shell or a login wall, and tells you there is nothing there. It is not wrong. The content of those posts is video, images, and text printed inside images, behind tokenised share links — none of it is in the HTML a fetch returns. Strip the &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt; tags from a real Xiaohongshu note and 264 characters of navigation are left.&lt;/p&gt;

&lt;h2&gt;
  
  
  What LinkDigest does about it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://linkdigest.dev/?ref=devto" rel="noopener noreferrer"&gt;LinkDigest&lt;/a&gt; is a hosted reader: one call turns the link into text an LLM can use — a transcript with timecodes, the on-screen text, a description and OCR of every image, the caption and metadata — as Markdown or JSON. It is an MCP server (&lt;code&gt;claude mcp add --transport http linkdigest https://linkdigest.dev/mcp --header "Authorization: Bearer $KEY"&lt;/code&gt;), a REST API, and a web console. Measured on a 17-image Xiaohongshu note: 17 images described and read, 381 on-screen text fragments, 13 key points, 119 seconds. Three digests are free, no card.&lt;/p&gt;

&lt;p&gt;What it does not do, stated up front: Bilibili refuses our server's address (HTTP 412), Instagram is wired but not verified, Facebook is out of scope. A digest that could not read part of a post says so in a &lt;code&gt;degraded&lt;/code&gt; field instead of returning a thin result quietly — and a digest that read nothing costs nothing.&lt;/p&gt;




&lt;p&gt;Your agent is good at reading. It just can't open half the internet.&lt;/p&gt;

&lt;p&gt;Hand Claude Code a &lt;a href="https://linkdigest.dev/platforms/xiaohongshu?ref=devto" rel="noopener noreferrer"&gt;Xiaohongshu&lt;/a&gt; link from a design review, a Douyin video a colleague sent, a &lt;a href="https://linkdigest.dev/platforms/tiktok?ref=devto" rel="noopener noreferrer"&gt;TikTok&lt;/a&gt; your competitor posted — and it fetches the URL, gets an app-download shell, and tells you it couldn't find anything useful. It isn't wrong. There genuinely is nothing there.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;[Diagram: The same URL returns an app-download shell to a plain fetch, and the actual note to LinkDigest.]&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is a post about the second box: how to call it, and three things people actually do with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The agent calls it, not you
&lt;/h2&gt;

&lt;p&gt;One command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http linkdigest &lt;span class="se"&gt;\&lt;/span&gt;
  https://linkdigest.dev/mcp &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer ld_live_..."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That registers one tool, &lt;code&gt;digest_url(url, format, job_id)&lt;/code&gt;. Only &lt;code&gt;url&lt;/code&gt; is required.&lt;/p&gt;

&lt;p&gt;The part that matters is what happens next: &lt;strong&gt;you never mention it again.&lt;/strong&gt; The tool description tells the agent to call it whenever it meets a social link it cannot read, so the agent reaches for it on its own, mid-task, the way it reaches for &lt;code&gt;grep&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;[Diagram: The agent calls the MCP server, which calls the engine, and the transcript returns to the agent.]&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Cursor takes the same server in &lt;code&gt;~/.cursor/mcp.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"linkdigest"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://linkdigest.dev/mcp"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"headers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"Authorization"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bearer ld_live_..."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What actually comes back
&lt;/h2&gt;

&lt;p&gt;Not a summary. The parts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://linkdigest.dev/api/v1/digest &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer ld_live_..."&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"url": "https://v.douyin.com/..."}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A real response, trimmed — this is a &lt;a href="https://linkdigest.dev/platforms/douyin?ref=devto" rel="noopener noreferrer"&gt;Douyin&lt;/a&gt; video, run against production while writing this paragraph:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;platform           'douyin'
author             '逸蒙的赛博空间'
posted_at          '2026-08-27'
transcript_source  'asr'
cached             True   credits 0
degraded           []
transcript[0]      {"t": 0, "text": "谁能想到，就这么一个丑萌小玩意儿…"}
key_points[0]      "该AI产品名为Tolen，主打「记住你」的核心功能…"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fifteen fields in total: &lt;code&gt;platform&lt;/code&gt;, &lt;code&gt;author&lt;/code&gt;, &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;posted_at&lt;/code&gt;, &lt;code&gt;caption&lt;/code&gt;, &lt;code&gt;transcript&lt;/code&gt;, &lt;code&gt;ocr_text&lt;/code&gt;, &lt;code&gt;images&lt;/code&gt;, &lt;code&gt;key_points&lt;/code&gt;, &lt;code&gt;raw_markdown&lt;/code&gt;, &lt;code&gt;source_url&lt;/code&gt;, &lt;code&gt;transcript_source&lt;/code&gt;, &lt;code&gt;degraded&lt;/code&gt;, plus &lt;code&gt;cached&lt;/code&gt; and &lt;code&gt;credits&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Two of those are worth pointing at.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;transcript_source&lt;/code&gt;&lt;/strong&gt; says how you got the words — &lt;code&gt;native_captions&lt;/code&gt;, &lt;code&gt;asr&lt;/code&gt;, or &lt;code&gt;none&lt;/code&gt;. Captions the platform already had are exact; speech recognition is not. A pipeline that treats those identically will eventually quote a mis-heard number back at someone as fact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;degraded&lt;/code&gt;&lt;/strong&gt; is a list of what didn't fully work, in plain words. Empty on a clean run. It exists because of a specific bug: a note was digested during a provider rate-limit storm, every vision batch failed, the per-batch handler swallowed each one, and the job returned &lt;code&gt;images: 0, ocr: 0&lt;/code&gt; — a well-formed, completely empty result, which was then cached for thirty days. Nothing downstream could tell an empty answer from an easy one. Now it can.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline, and why some links take two calls
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;[Diagram: Four stages: resolve, fetch, read, structure.]&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;read&lt;/code&gt; is where the time goes, because it is the only stage that has to watch or listen to anything. Measured: a Xiaohongshu note with images takes one to two minutes; a &lt;a href="https://linkdigest.dev/platforms/youtube?ref=devto" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt; video with captions about two and a half.&lt;/p&gt;

&lt;p&gt;That is longer than one HTTP request should wait, so the API hands the work back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;POST /api/v1/digest
→ 202 {"pending":true,"jobId":"abc…","poll":"/api/v1/digest/abc…","retryAfter":15}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Collect it with the same key at that &lt;code&gt;poll&lt;/code&gt; path — &lt;code&gt;202&lt;/code&gt; while it runs, &lt;code&gt;200&lt;/code&gt; with the digest when it's done. Over MCP you don't do any of this by hand: the tool tells the agent to call &lt;code&gt;digest_url&lt;/code&gt; again with the &lt;code&gt;job_id&lt;/code&gt; and no url, and the agent does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things people build with it
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. The agent that reads the link itself
&lt;/h3&gt;

&lt;p&gt;The one that needs no code. A teammate drops a Douyin link in an issue; the agent working that issue reads it without anyone transcribing anything. This is the whole reason the tool description is written the way it is — it is instructions for a machine deciding whether to reach for a tool, not marketing copy.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. A list of links becomes a table
&lt;/h3&gt;

&lt;p&gt;Fifty URLs in, fifty rows out — one &lt;code&gt;POST&lt;/code&gt; each, or the &lt;a href="https://apify.com/loongnian714/social-url-to-llm-text" rel="noopener noreferrer"&gt;Apify Actor&lt;/a&gt; if you'd rather not write the loop. The dataset has a row per link with the transcript, the on-screen text and the key points already separated.&lt;/p&gt;

&lt;p&gt;Worth knowing what you're joining: there is real demand for exactly this shape of work. A single competing Douyin scraper on the Apify Store has over 1.5 million runs. Almost all of them return metadata — view counts, captions, author handles. Very few return what was actually &lt;em&gt;said&lt;/em&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Watching what a competitor publishes
&lt;/h3&gt;

&lt;p&gt;The same loop on a schedule. A creator's video output becomes searchable text over time, so "when did they first mention pricing" is a grep instead of an afternoon of watching.&lt;/p&gt;

&lt;p&gt;This is the use case where &lt;code&gt;transcript_source&lt;/code&gt; earns its place again. If you're going to quote a competitor's claim back to your own team, it matters whether a human captioned it or a model guessed at it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs, and why it isn't flat
&lt;/h2&gt;

&lt;p&gt;A digest is not a unit of cost. Measured against real runs:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;[Diagram: Measured cost per digest: a web article is a fraction of a cent, a nineteen-minute video is nine cents.]&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Which is why pricing counts the two things that drive that spread — images to describe, minutes to transcribe — rather than counting requests. One credit covers a typical post and its first six images; each further six images adds one, and each minute of media adds two.&lt;/p&gt;

&lt;p&gt;Cached links are free and never counted. Anything anyone has ever digested comes back in about a second, for nobody's credits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stops
&lt;/h2&gt;

&lt;p&gt;The honest table, because finding out after you've wired something in is worse than knowing now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Works:&lt;/strong&gt; Xiaohongshu, Douyin, TikTok, YouTube, X, ordinary web pages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bilibili&lt;/strong&gt; returns HTTP 412 to a datacenter address. Needs a residential proxy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instagram and Facebook&lt;/strong&gt; need a logged-in session for most posts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three free digests, no card, if you want to check any of that yourself.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The platform-by-platform detail — what each one returns and what breaks — is in &lt;a href="https://linkdigest.dev/blog/reading-xiaohongshu-from-a-server?ref=devto" rel="noopener noreferrer"&gt;What it actually takes to read a Xiaohongshu post from a server&lt;/a&gt;, and per platform on the &lt;a href="https://linkdigest.dev/platforms?ref=devto" rel="noopener noreferrer"&gt;coverage pages&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>claude</category>
    </item>
    <item>
      <title>What it actually takes to read a Xiaohongshu post from a server</title>
      <dc:creator>Programming with Jack Chew</dc:creator>
      <pubDate>Sat, 12 Sep 2026 02:53:41 +0000</pubDate>
      <link>https://dev.to/programming_withjackche/what-it-actually-takes-to-read-a-xiaohongshu-post-from-a-server-1e52</link>
      <guid>https://dev.to/programming_withjackche/what-it-actually-takes-to-read-a-xiaohongshu-post-from-a-server-1e52</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://linkdigest.dev/blog/reading-xiaohongshu-from-a-server?ref=devto" rel="noopener noreferrer"&gt;linkdigest.dev&lt;/a&gt;, where I build this.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem, in one paragraph
&lt;/h2&gt;

&lt;p&gt;Paste a Xiaohongshu, Douyin, TikTok, YouTube or X link into an AI agent and it fetches the URL, gets an app-download shell or a login wall, and tells you there is nothing there. It is not wrong. The content of those posts is video, images, and text printed inside images, behind tokenised share links — none of it is in the HTML a fetch returns. Strip the &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt; tags from a real Xiaohongshu note and 264 characters of navigation are left.&lt;/p&gt;

&lt;h2&gt;
  
  
  What LinkDigest does about it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://linkdigest.dev/?ref=devto" rel="noopener noreferrer"&gt;LinkDigest&lt;/a&gt; is a hosted reader: one call turns the link into text an LLM can use — a transcript with timecodes, the on-screen text, a description and OCR of every image, the caption and metadata — as Markdown or JSON. It is an MCP server (&lt;code&gt;claude mcp add --transport http linkdigest https://linkdigest.dev/mcp --header "Authorization: Bearer $KEY"&lt;/code&gt;), a REST API, and a web console. Measured on a 17-image Xiaohongshu note: 17 images described and read, 381 on-screen text fragments, 13 key points, 119 seconds. Three digests are free, no card.&lt;/p&gt;

&lt;p&gt;What it does not do, stated up front: Bilibili refuses our server's address (HTTP 412), Instagram is wired but not verified, Facebook is out of scope. A digest that could not read part of a post says so in a &lt;code&gt;degraded&lt;/code&gt; field instead of returning a thin result quietly — and a digest that read nothing costs nothing.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Every claim here was checked against a live link. Where something doesn't work, it says so.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Paste a &lt;a href="https://linkdigest.dev/platforms/xiaohongshu?ref=devto" rel="noopener noreferrer"&gt;Xiaohongshu&lt;/a&gt; link into an AI coding agent and it sees nothing. Same for Douyin. The usual answer — "just use yt-dlp" — is half true in a way that wastes an afternoon, so here is the whole picture.&lt;/p&gt;

&lt;h2&gt;
  
  
  yt-dlp only reads half of Xiaohongshu
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;XiaoHongShuIE&lt;/code&gt; reads &lt;code&gt;note.video.media.stream&lt;/code&gt;. That is a video note. Image notes — 图文, a caption plus a stack of photos — are the majority of the platform and usually the ones worth reading, and the extractor has nothing to say about them.&lt;/p&gt;

&lt;p&gt;Getting those means parsing the note payload out of the page itself: the state blob the page ships, and the &lt;code&gt;imageList&lt;/code&gt; inside it. Reusing yt-dlp's own &lt;code&gt;js_to_json&lt;/code&gt; and &lt;code&gt;traverse_obj&lt;/code&gt; for that keeps you aligned with upstream when the page shape shifts, which it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  The page is different depending on who you say you are
&lt;/h2&gt;

&lt;p&gt;This one cost me an embarrassing amount of time.&lt;/p&gt;

&lt;p&gt;Request a note with a &lt;strong&gt;mobile&lt;/strong&gt; user agent and you get 202KB of app-download shell whose &lt;code&gt;&amp;lt;title&amp;gt;&lt;/code&gt; is just the site name. Request the same URL with a &lt;strong&gt;desktop&lt;/strong&gt; user agent and you get 85KB containing the real note.&lt;/p&gt;

&lt;p&gt;Both are HTTP 200. Nothing tells you that you got the wrong one except that the content isn't there.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;[Diagram: The same note URL returns 202KB of app shell to a mobile user agent, and 85KB containing the note to a desktop one.]&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cold requests are rejected
&lt;/h2&gt;

&lt;p&gt;Fetching a note URL directly, with no prior session, does not work. Fetch the &lt;code&gt;/explore&lt;/code&gt; feed first, keep the cookie jar it gives you — &lt;code&gt;acw_tc&lt;/code&gt;, &lt;code&gt;abRequestId&lt;/code&gt; — and then the note loads.&lt;/p&gt;

&lt;p&gt;No account, no credentials, no API key. Just the same two-step a browser performs without you noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Douyin does not negotiate
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Coverage and output for Douyin: &lt;a href="https://linkdigest.dev/platforms/douyin?ref=devto" rel="noopener noreferrer"&gt;platforms/douyin&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Douyin refuses anonymous requests outright: captcha on the web page, 403 from the APIs. There is no user-agent trick here. That path needs cookies from a logged-in session, and yt-dlp does not support the platform at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  YouTube blocks your server, not your laptop
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Coverage and output for YouTube: &lt;a href="https://linkdigest.dev/platforms/youtube?ref=devto" rel="noopener noreferrer"&gt;platforms/youtube&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The single most common "it worked locally and broke in production" report. YouTube blocks datacenter IP ranges, so a yt-dlp fetch that is perfect on your machine fails from EC2, App Runner, Lambda or anywhere else you deploy — and no amount of configuration fixes an IP-range block.&lt;/p&gt;

&lt;p&gt;The workable fallback is a model that watches the video and returns a transcript, which costs meaningfully more than parsing captions and is worth measuring separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stops
&lt;/h2&gt;

&lt;p&gt;Being honest about this saves everyone time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bilibili&lt;/strong&gt; returns HTTP 412 to a datacenter address. Needs a residential proxy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instagram and Facebook&lt;/strong&gt; need a logged-in session for most posts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TikTok&lt;/strong&gt; resolves fine but rate-limits under load; a frame fetch can come back 403 mid-job.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The part nobody mentions: failure has to be loud
&lt;/h2&gt;

&lt;p&gt;The bug that taught me the most had nothing to do with extraction.&lt;/p&gt;

&lt;p&gt;A note was digested during a provider rate-limit storm. Every vision batch failed, the per-batch error handler swallowed each one, and the job returned &lt;code&gt;images: 0, ocr: 0&lt;/code&gt; — a perfectly well-formed, completely empty result. It was then written to a shared cache with a 30-day TTL.&lt;/p&gt;

&lt;p&gt;Long after the underlying problem was fixed, that link still returned nothing, because a cached answer is cheaper to serve than to recompute. The pipeline was healthy. The cache was serving a fossil, and nothing downstream could tell the difference between an empty result and an easy one.&lt;/p&gt;

&lt;p&gt;Two things came out of that:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A step that loses content has to say so.&lt;/strong&gt; The result now carries what failed, so a thin answer is distinguishable from a genuinely short post.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thin results get a short TTL, not the full one.&lt;/strong&gt; Not "don't cache" — a link that fails &lt;em&gt;every&lt;/em&gt; time would then re-run the full paid pipeline on every request forever. An hour bounds the staleness and the spend together.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you are building anything that caches derived content, that second one is the trap. The obvious fix is the expensive one.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The findings are the same whether you use a hosted service or write your own.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;For how this fits into an agent's workflow end to end — the MCP tool, the job id for long media, what the output looks like — see &lt;a href="https://linkdigest.dev/blog/reading-social-links-from-an-agent?ref=devto" rel="noopener noreferrer"&gt;Reading links your agent can't open&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>python</category>
      <category>ai</category>
    </item>
    <item>
      <title>LinkDigest: turn Xiaohongshu, Douyin, TikTok, YouTube and X links into text your agent can read</title>
      <dc:creator>Programming with Jack Chew</dc:creator>
      <pubDate>Sat, 12 Sep 2026 02:42:47 +0000</pubDate>
      <link>https://dev.to/programming_withjackche/linkdigest-turn-xiaohongshu-douyin-tiktok-youtube-and-x-links-into-text-your-agent-can-read-ip</link>
      <guid>https://dev.to/programming_withjackche/linkdigest-turn-xiaohongshu-douyin-tiktok-youtube-and-x-links-into-text-your-agent-can-read-ip</guid>
      <description>&lt;p&gt;&lt;em&gt;I build &lt;a href="https://linkdigest.dev/?ref=devto" rel="noopener noreferrer"&gt;LinkDigest&lt;/a&gt;. Every number below is from a measured run recorded in the project log; nothing is projected.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;Agents are good at reading and bad at opening. Give Claude Code, Cursor, or a Dify workflow a link from Xiaohongshu, Douyin, TikTok, YouTube or X and the usual tool — fetch the URL, parse the HTML — returns one of three things: an app-download shell, a login wall, or a page whose body is 264 characters of navigation because the post itself lives in a JSON blob or in a video. The agent reports "nothing useful here" and it is telling the truth about the HTML.&lt;/p&gt;

&lt;p&gt;The substance of these posts is not text. It is speech in a video, words printed on images, screenshots of code, a recipe laid out in a photo. A plain fetch throws all of that away, and a share link from the app carries a token that expires in a few weeks, so even the HTML you do get is different depending on which link and which user-agent you used.&lt;/p&gt;

&lt;h2&gt;
  
  
  The solution
&lt;/h2&gt;

&lt;p&gt;LinkDigest does the reading on its own servers and hands back text. One call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://linkdigest.dev/api/v1/digest &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$LINKDIGEST_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"url": "https://www.youtube.com/watch?v=CEvIs9y1uog", "format": "json"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;returns the post's transcript with timecodes, every fragment of on-screen text, a description and OCR of each image, the caption and metadata, and a list of key points — as JSON, or as Markdown if you ask for it. Long media answers in two steps: a &lt;code&gt;202&lt;/code&gt; with a job id, then the digest when it is done.&lt;/p&gt;

&lt;p&gt;Measured, not estimated:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;post&lt;/th&gt;
&lt;th&gt;what came back&lt;/th&gt;
&lt;th&gt;time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;a 17-image Xiaohongshu note&lt;/td&gt;
&lt;td&gt;17 images described and read, 381 on-screen text fragments, 13 key points&lt;/td&gt;
&lt;td&gt;119 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a Xiaohongshu video note&lt;/td&gt;
&lt;td&gt;ASR transcript, 56 on-screen text fragments, 12 key points&lt;/td&gt;
&lt;td&gt;~4 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a YouTube talk (native captions)&lt;/td&gt;
&lt;td&gt;426 transcript segments, 190 on-screen text fragments&lt;/td&gt;
&lt;td&gt;~1 s cached&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a two-photo X post&lt;/td&gt;
&lt;td&gt;both photos described, 3 text fragments&lt;/td&gt;
&lt;td&gt;~20 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  For agents: it is an MCP server
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http linkdigest https://linkdigest.dev/mcp &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$LINKDIGEST_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One tool, &lt;code&gt;digest_url(url, format)&lt;/code&gt;. The tool description tells the model exactly when to reach for it and which platforms are out, so an agent does not confidently hand you a failure. It is in the official MCP registry as &lt;code&gt;dev.linkdigest/linkdigest&lt;/code&gt;, on Apify as three Actors, and submitted to the Dify marketplace.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does not do
&lt;/h2&gt;

&lt;p&gt;This is the part I would want to know before pasting a key:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bilibili: not supported.&lt;/strong&gt; Bilibili returns HTTP 412 to our server's address.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instagram: wired, not verified end to end.&lt;/strong&gt; Facebook: out of scope.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;X long-form articles&lt;/strong&gt; expose only their lead image; regular photo posts come back in full.&lt;/li&gt;
&lt;li&gt;A YouTube link is read by Gemini watching the video, because YouTube blocks datacenter IPs; if Google refuses a video, you get an honest stub, and it costs nothing.&lt;/li&gt;
&lt;li&gt;Anything that could not be read is named in a &lt;code&gt;degraded&lt;/code&gt; field. A digest that read nothing is free.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Pricing
&lt;/h2&gt;

&lt;p&gt;Three digests free, no card. Then $9 a month for 500 credits — a typical post is 1 credit, video 2 per started minute, and links anyone has already digested are free forever. &lt;a href="https://linkdigest.dev/pricing?ref=devto" rel="noopener noreferrer"&gt;https://linkdigest.dev/pricing&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I built it
&lt;/h2&gt;

&lt;p&gt;I kept pasting Xiaohongshu and Douyin links into agents and watching them fail, and the failures were silent — a confident summary of a login page. Two of the bugs I hit while building this are written up separately: &lt;a href="https://linkdigest.dev/blog/xiaohongshu-share-links-have-two-domains?ref=devto" rel="noopener noreferrer"&gt;share links have two domains and one of them fails silently&lt;/a&gt;, and &lt;a href="https://linkdigest.dev/blog/reading-xiaohongshu-from-a-server?ref=devto" rel="noopener noreferrer"&gt;what it actually takes to read a Xiaohongshu post from a server&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you try it and something reads thin, reply here — I fix these the same day.&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>mcp</category>
      <category>ai</category>
    </item>
    <item>
      <title>Xiaohongshu share links have two domains, and one of them will fail silently</title>
      <dc:creator>Programming with Jack Chew</dc:creator>
      <pubDate>Sat, 12 Sep 2026 02:33:42 +0000</pubDate>
      <link>https://dev.to/programming_withjackche/xiaohongshu-share-links-have-two-domains-and-one-of-them-will-fail-silently-hpd</link>
      <guid>https://dev.to/programming_withjackche/xiaohongshu-share-links-have-two-domains-and-one-of-them-will-fail-silently-hpd</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://linkdigest.dev/blog/xiaohongshu-share-links-have-two-domains?ref=devto" rel="noopener noreferrer"&gt;linkdigest.dev&lt;/a&gt;, where I build this.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Every number here came from a live link on 2026-09-08 and 2026-09-10. The failure described is one this service shipped to a real user.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If you have a host table that maps domains to platforms, and Xiaohongshu is in it, check what you wrote. There is a good chance it says &lt;code&gt;xhslink.com&lt;/code&gt; and nothing else. That was true here, and it cost us the first person who ever signed up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two domains, same platform
&lt;/h2&gt;

&lt;p&gt;Share a note from the &lt;a href="https://linkdigest.dev/platforms/xiaohongshu?ref=devto" rel="noopener noreferrer"&gt;Xiaohongshu&lt;/a&gt; iOS app today and the clipboard gets something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://xhslink.cn/o/AxnRePgIokn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Older shares, and links that travel through the desktop site, use &lt;code&gt;.com&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://xhslink.com/o/1WiQ1QI6Uc0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both are real. Both resolve to the same place. Following the &lt;code&gt;.cn&lt;/code&gt; one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-sIL&lt;/span&gt; &lt;span class="s1"&gt;'https://xhslink.cn/o/AxnRePgIokn'&lt;/span&gt; &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="s1"&gt;'&amp;lt;desktop UA&amp;gt;'&lt;/span&gt;
...
https://www.xiaohongshu.com/discovery/item/6a9d0eb2000000002a024480
    ?app_platform&lt;span class="o"&gt;=&lt;/span&gt;ios&amp;amp;app_version&lt;span class="o"&gt;=&lt;/span&gt;9.45&amp;amp;xsec_source&lt;span class="o"&gt;=&lt;/span&gt;app_share
    &amp;amp;type&lt;span class="o"&gt;=&lt;/span&gt;normal&amp;amp;xsec_token&lt;span class="o"&gt;=&lt;/span&gt;CB7rCz...&amp;amp;share_id&lt;span class="o"&gt;=&lt;/span&gt;310c7d81...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second link we tested carried &lt;code&gt;type=video&lt;/code&gt; instead of &lt;code&gt;type=normal&lt;/code&gt;, which matters later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the miss is invisible
&lt;/h2&gt;

&lt;p&gt;A URL router usually looks like this, and ours did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;(^|\.)(xiaohongshu\.com|xhslink\.com)$&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;Platform&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;xiaohongshu&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;.cn&lt;/code&gt; share does not match. In most systems it does not raise — it falls through to whatever generic handler sits at the end of the chain. Ours fetches the page and extracts what it can from the HTML.&lt;/p&gt;

&lt;p&gt;And that &lt;em&gt;works&lt;/em&gt;, in the sense that it returns 200. Xiaohongshu's page ships a &lt;code&gt;&amp;lt;title&amp;gt;&lt;/code&gt; and a &lt;code&gt;meta description&lt;/code&gt; containing the whole caption, so a generic extractor comes back with a title, an author of "小红书", and the caption. Enough to look like an answer.&lt;/p&gt;

&lt;p&gt;It is not, and the reason is covered at length in &lt;a href="https://linkdigest.dev/blog/reading-xiaohongshu-from-a-server?ref=devto" rel="noopener noreferrer"&gt;what it actually takes to read a Xiaohongshu post from a server&lt;/a&gt;: strip the &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt; tags from that page and 264 characters of navigation chrome are left. The note body is not in the rendered DOM at all. A generic extractor is reading the wrapper, not the post.&lt;/p&gt;

&lt;p&gt;Here is what the same link returned before and after the domain was added to the table:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;.cn&lt;/code&gt; unrecognised&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;.cn&lt;/code&gt; recognised&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;platform&lt;/td&gt;
&lt;td&gt;&lt;code&gt;web&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;xiaohongshu&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;images described&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;on-screen text fragments&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;transcript&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;author&lt;/td&gt;
&lt;td&gt;&lt;code&gt;小红书&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the actual account name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;posted date&lt;/td&gt;
&lt;td&gt;empty&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2026-09-06&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;degraded&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;["extracted from page HTML only — no media, transcript, or platform metadata"]&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;[]&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And the second link, the &lt;code&gt;type=video&lt;/code&gt; one:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;before&lt;/th&gt;
&lt;th&gt;after&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;platform&lt;/td&gt;
&lt;td&gt;&lt;code&gt;web&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;xiaohongshu&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;transcript source&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;&lt;code&gt;asr&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;on-screen text fragments&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A transcript existed the whole time. Nothing errored. The request succeeded in seven seconds and returned a post with no transcript in it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is the part worth taking away.&lt;/strong&gt; A silent downgrade is worse than a crash, because a crash tells the user to retry and a downgrade tells them this is what your product can do. The person who hit this did not file a bug. They opened their usage page, their settings, the docs, and left.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check your own table
&lt;/h2&gt;

&lt;p&gt;If you maintain anything that classifies Xiaohongshu URLs, the fix is one character class:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;(^|\.)(xiaohongshu\.com|xhslink\.(com|cn))$&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;Platform&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;xiaohongshu&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And separately, if you expand short links before identifying content — you should, because the note id is not in the short form — &lt;code&gt;xhslink.cn&lt;/code&gt; has to be in that set too. Ours was missing from both.&lt;/p&gt;

&lt;p&gt;The reason a full test suite did not catch it is worth saying out loud: every Xiaohongshu URL in our fixtures, evals, docs and demo scripts was a &lt;code&gt;.com&lt;/code&gt; one, because those were the links &lt;em&gt;we&lt;/em&gt; had pasted. The test suite was green and tested the wrong domain.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second half: your cache may outlive your fix
&lt;/h2&gt;

&lt;p&gt;We deployed the routing fix and the very next request came back wrong anyway.&lt;/p&gt;

&lt;p&gt;The cache key was computed from the raw URL by the layer in front, which had already stored the bad result under the key for &lt;code&gt;xhslink.cn/o/AxnRePgIokn&lt;/code&gt;. The lookup answered before the fixed code ran. Deploying a parser fix does not repair anything you have already cached under the same key.&lt;/p&gt;

&lt;p&gt;If you cache digests, decide deliberately which layer expands short links. Expanding &lt;em&gt;before&lt;/em&gt; the key is computed also means two share links to the same note stop producing two entries, and stop billing twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this service does now
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;xhslink.cn&lt;/code&gt; and &lt;code&gt;xhslink.com&lt;/code&gt; both resolve to the Xiaohongshu path, image notes and video notes both, and a link that cannot be read fully says so in &lt;code&gt;degraded&lt;/code&gt; rather than returning a thin result quietly.&lt;/p&gt;

&lt;p&gt;Still not supported, stated where you can see it before paying: Bilibili returns HTTP 412 to our address, and Instagram is wired but not verified end to end.&lt;/p&gt;

&lt;p&gt;If you are wiring this into an agent rather than calling it by hand, &lt;a href="https://linkdigest.dev/blog/reading-social-links-from-an-agent?ref=devto" rel="noopener noreferrer"&gt;reading links your agent can't open&lt;/a&gt; covers the tool call and what each field contains.&lt;/p&gt;

&lt;p&gt;Three digests are free and there is no card. &lt;a href="https://linkdigest.dev/?ref=devto" rel="noopener noreferrer"&gt;linkdigest.dev&lt;/a&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>ai</category>
      <category>python</category>
    </item>
    <item>
      <title>2026 Q1 is the year developers still build the agent harness. 2026 Q3 / 2027 is the year the LLM builds its own harness.</title>
      <dc:creator>Programming with Jack Chew</dc:creator>
      <pubDate>Sun, 24 May 2026 03:12:16 +0000</pubDate>
      <link>https://dev.to/programming_withjackche/2026-q1-is-the-year-developers-still-build-the-agent-harness-2026-q3-2027-is-the-year-the-llm-359f</link>
      <guid>https://dev.to/programming_withjackche/2026-q1-is-the-year-developers-still-build-the-agent-harness-2026-q3-2027-is-the-year-the-llm-359f</guid>
      <description>&lt;p&gt;2026 Q1 is the year developers still build the agent harness.&lt;/p&gt;

&lt;p&gt;2026 Q3 / 2027 is the year the LLM builds its own harness.&lt;/p&gt;

&lt;p&gt;Today, every AI coding agent — Claude Code, Cursor, Codex, Gemini CLI, Aider, you name it — depends on the same hidden layer:&lt;/p&gt;

&lt;p&gt;the files that brief the agent before it starts work.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;AGENTS.md&lt;/code&gt;&lt;br&gt;
&lt;code&gt;CLAUDE.md&lt;/code&gt;&lt;br&gt;
&lt;code&gt;.cursor/rules&lt;/code&gt;&lt;br&gt;
&lt;code&gt;SKILLS/&lt;/code&gt;&lt;br&gt;
MCP server lists&lt;br&gt;
memory schemas&lt;br&gt;
test commands&lt;br&gt;
lint commands&lt;br&gt;
“Do not touch these paths.”&lt;br&gt;
“Require human approval before this.”&lt;/p&gt;

&lt;p&gt;Different IDE, same boilerplate.&lt;br&gt;
Different repo, same boilerplate.&lt;br&gt;
Different agent, same boilerplate.&lt;/p&gt;

&lt;p&gt;That is the agent harness problem.&lt;/p&gt;
&lt;h2&gt;
  
  
  The hidden work behind AI coding agents
&lt;/h2&gt;

&lt;p&gt;Most people talk about the coding agent itself.&lt;/p&gt;

&lt;p&gt;But in practice, the quality of an AI coding session often depends on the context layer around the agent.&lt;/p&gt;

&lt;p&gt;Before the agent starts coding, it needs to know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what kind of project this is&lt;/li&gt;
&lt;li&gt;what framework it uses&lt;/li&gt;
&lt;li&gt;what files are important&lt;/li&gt;
&lt;li&gt;what commands run tests&lt;/li&gt;
&lt;li&gt;what commands run linting&lt;/li&gt;
&lt;li&gt;what paths should not be touched&lt;/li&gt;
&lt;li&gt;what tools are available&lt;/li&gt;
&lt;li&gt;what memory should persist&lt;/li&gt;
&lt;li&gt;what failure modes to avoid&lt;/li&gt;
&lt;li&gt;what coding conventions to follow&lt;/li&gt;
&lt;li&gt;when human approval is required&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without this layer, even strong coding agents can make subtle mistakes.&lt;/p&gt;

&lt;p&gt;With this layer, the same agent can behave much more consistently.&lt;/p&gt;

&lt;p&gt;That layer is what I call the harness.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why this still exists in 2026
&lt;/h2&gt;

&lt;p&gt;In theory, the LLM should be able to inspect a repo and generate all of this itself.&lt;/p&gt;

&lt;p&gt;In practice, we are not fully there yet.&lt;/p&gt;

&lt;p&gt;The models are smart enough to do real coding work, but not always reliable enough to deterministically generate perfect project-specific ground truth from scratch on every fresh repo, every time.&lt;/p&gt;

&lt;p&gt;They can do it sometimes.&lt;/p&gt;

&lt;p&gt;Not always.&lt;/p&gt;

&lt;p&gt;So the human stays in the loop.&lt;/p&gt;

&lt;p&gt;We write the same repo instructions again.&lt;/p&gt;

&lt;p&gt;We copy the same rules across projects.&lt;/p&gt;

&lt;p&gt;We maintain separate files for Claude Code, Cursor, Codex-style agents, Continue, Windsurf, and others.&lt;/p&gt;

&lt;p&gt;Small work per repo.&lt;/p&gt;

&lt;p&gt;Painful in aggregate.&lt;/p&gt;
&lt;h2&gt;
  
  
  The future: self-generating harnesses
&lt;/h2&gt;

&lt;p&gt;I think this is temporary.&lt;/p&gt;

&lt;p&gt;Soon, the coding model should be able to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;read the repo&lt;/li&gt;
&lt;li&gt;understand the task&lt;/li&gt;
&lt;li&gt;detect the project type&lt;/li&gt;
&lt;li&gt;generate the right harness&lt;/li&gt;
&lt;li&gt;connect the right tools&lt;/li&gt;
&lt;li&gt;create memory schemas&lt;/li&gt;
&lt;li&gt;write validation scripts&lt;/li&gt;
&lt;li&gt;refine the loop until the task is complete&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;At that point, the harness layer disappears as a separately authored artifact.&lt;/p&gt;

&lt;p&gt;But until then, developers still need a bridge.&lt;/p&gt;
&lt;h2&gt;
  
  
  I built harnessforge
&lt;/h2&gt;

&lt;p&gt;I built &lt;code&gt;harnessforge&lt;/code&gt; to test this idea.&lt;/p&gt;

&lt;p&gt;It is a local, open-source harness generator for AI coding agents.&lt;/p&gt;

&lt;p&gt;It is not another coding agent.&lt;/p&gt;

&lt;p&gt;Your coding agent stays the brain.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;harnessforge&lt;/code&gt; just lays down the ground truth the agent reads before work begins.&lt;/p&gt;

&lt;p&gt;Run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uvx harnessforge init
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or install:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;harnessforge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a few seconds, fully local with no network calls by default, it inspects your repo and generates startup files commonly used by AI coding agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it generates
&lt;/h2&gt;

&lt;p&gt;Depending on the project and blueprint, &lt;code&gt;harnessforge&lt;/code&gt; can generate files such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AGENTS.md
SOUL.md
TOOLS.md
MEMORY.md
SKILLS/
.claude/CLAUDE.md
.cursor/rules
.continue/
.windsurf/rules
blueprint-specific validators
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is simple:&lt;/p&gt;

&lt;p&gt;give the coding agent a stronger starting point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current blueprints
&lt;/h2&gt;

&lt;p&gt;The current version includes these blueprints:&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;rag-agent&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;For retrieval systems, knowledge-base agents, citation enforcement, and grounded responses.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;finance-agent&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;For finance or stock-related agents, including market-data handling and validation rules around trade execution safety.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;support-agent&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;For customer support flows such as intent detection, knowledge-base lookup, ticket creation, escalation, and ticket lineage.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;workflow-agent&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;For multi-step orchestration with tool logs, idempotency, and validation structure.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;python-cli-app&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;A default blueprint for greenfield Python CLI projects.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters
&lt;/h2&gt;

&lt;p&gt;The important idea is not the specific files.&lt;/p&gt;

&lt;p&gt;The important idea is that coding agents need a reliable project-specific operating context.&lt;/p&gt;

&lt;p&gt;Today, we manually maintain that context.&lt;/p&gt;

&lt;p&gt;Tomorrow, the model may generate it automatically.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;harnessforge&lt;/code&gt; is meant to sit in the middle.&lt;/p&gt;

&lt;p&gt;A bridge, not a moat.&lt;/p&gt;

&lt;p&gt;Use it now.&lt;/p&gt;

&lt;p&gt;Throw it away when the models catch up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example workflow
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uvx harnessforge init
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then open Claude Code, Cursor, Codex, Gemini CLI, Aider, or another coding agent inside the repo.&lt;/p&gt;

&lt;p&gt;The agent now has project-specific context files to read before it starts work.&lt;/p&gt;

&lt;p&gt;Instead of starting from a blank repo, the agent starts with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;project rules&lt;/li&gt;
&lt;li&gt;tool definitions&lt;/li&gt;
&lt;li&gt;memory structure&lt;/li&gt;
&lt;li&gt;validation expectations&lt;/li&gt;
&lt;li&gt;blueprint-specific failure modes&lt;/li&gt;
&lt;li&gt;agent-specific startup files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The coding agent still writes the code.&lt;/p&gt;

&lt;p&gt;The harness just gives it the right context.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bet
&lt;/h2&gt;

&lt;p&gt;My bet is:&lt;/p&gt;

&lt;p&gt;2026 Q1: developers still build the agent harness.&lt;/p&gt;

&lt;p&gt;2026 Q3 / 2027: the LLM builds its own harness.&lt;/p&gt;

&lt;p&gt;Until that happens, a local deterministic harness generator can make AI coding workflows more reliable.&lt;/p&gt;

&lt;p&gt;GitHub:&lt;br&gt;
&lt;a href="https://github.com/jcaiagent7143-ui/harnessforge" rel="noopener noreferrer"&gt;https://github.com/jcaiagent7143-ui/harnessforge&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;PyPI:&lt;br&gt;
&lt;a href="https://pypi.org/project/harnessforge/" rel="noopener noreferrer"&gt;https://pypi.org/project/harnessforge/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I would love feedback from developers using Claude Code, Cursor, Codex, Gemini CLI, Aider, Continue, Windsurf, or other coding agents in real repos.&lt;/p&gt;

&lt;p&gt;How are you managing your agent harness today?&lt;/p&gt;

&lt;p&gt;Are you manually maintaining &lt;code&gt;AGENTS.md&lt;/code&gt;, &lt;code&gt;CLAUDE.md&lt;/code&gt;, &lt;code&gt;.cursor/rules&lt;/code&gt;, MCP configs, memory files, and validation rules?&lt;/p&gt;

&lt;p&gt;Or do you think the next generation of coding models will generate this layer automatically?&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>tooling</category>
    </item>
  </channel>
</rss>
