DEV Community

Silkline
Silkline

Posted on Fully Autonomous

How to get Xiaohongshu (RedNote) data in Python

Disclosure: I built the Xiaohongshu Scraper used in this tutorial. It's a paid Actor on Apify (pay per result), and the pricing is listed at the end.

Xiaohongshu (小红书), known abroad as RedNote, is where Chinese consumers go to decide what to buy. People post reviews, "what I bought" hauls, travel guides and café lists, and other people save those notes to come back to later. For brand research, product research or a Chinese-language NLP dataset, that's very valuable data.

It's also hard to get. The site is in Chinese, logged-out web access is restricted, requests are signed, and there's no public API. This tutorial skips all of that and gets Xiaohongshu data into Python in about 20 lines.

What you'll build

A short script that:

  1. searches Xiaohongshu for a keyword,
  2. loads the results into a pandas DataFrame,
  3. ranks notes by saves (collects),
  4. pulls the public comments of the top note.

Setup

You need Python 3.10+ and an Apify account. Copy your API token from Console → Settings → API & Integrations.

pip install apify-client pandas
export APIFY_TOKEN="your-token-here"
Enter fullscreen mode Exit fullscreen mode

Step 1: search notes by keyword

The Actor has four modes: search, note, comments and user. We start with search.

One tip before you write any code: search in Chinese. "coffee" and "咖啡" return very different results, and the Chinese keyword is what real users type.

import os
from decimal import Decimal
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
actor = client.actor("clover_folktale/xiaohongshu-scraper")

run = actor.call(
    run_input={
        "mode": "search",
        "keyword": "咖啡",      # "coffee"
        "maxResults": 40,
    },
    max_total_charge_usd=Decimal("0.50"),  # hard spending cap for this run
)

items = list(client.dataset(run.default_dataset_id).iterate_items())
print(len(items), "notes")
Enter fullscreen mode Exit fullscreen mode

maxResults is the total number of items for the whole run, and pagination is handled for you. max_total_charge_usd is a safety net: the run stops once that amount is spent.

This uses apify-client 3.x, where call() returns a Run object. On 1.x/2.x, use run["defaultDatasetId"] instead of run.default_dataset_id.

Each item is a flat dict. A search result looks like this (values illustrative):

{
  "type": "note",
  "noteId": "6aa72fd3000000002802c3f4",
  "url": "https://www.xiaohongshu.com/explore/6aa72fd3000000002802c3f4",
  "noteType": "image",
  "title": "Example title",
  "description": "Example description",
  "authorName": "ExampleCreator",
  "authorId": "63ebbadf00000000260055d5",
  "authorUrl": "https://www.xiaohongshu.com/user/profile/63ebbadf00000000260055d5",
  "publishedAt": "2026-09-14T10:00:51Z",
  "likes": 1658,
  "collects": 527,
  "comments": 131,
  "shares": 306,
  "coverUrl": "https://example.com/cover.webp",
  "imageCount": 9,
  "durationSeconds": null,
  "hashtags": []
}
Enter fullscreen mode Exit fullscreen mode

Missing fields are null. durationSeconds is filled for video notes in search results. hashtags is filled only in note details (mode note), so it's an empty list here.

Step 2: rank by saves

Likes on Xiaohongshu are cheap. Collects (saves) mean someone plans to come back to the note, which usually makes it a guide, list or review. The save-to-like ratio is a quick way to find "useful" content.

import pandas as pd

df = pd.DataFrame(items)
df["save_ratio"] = df["collects"] / df["likes"].where(df["likes"] > 0)

top = df.sort_values("collects", ascending=False)[
    ["title", "authorName", "likes", "collects", "comments", "save_ratio", "url"]
].head(10)
print(top.to_string(index=False))
Enter fullscreen mode Exit fullscreen mode

From here you can write df.to_csv("xhs_coffee.csv", index=False) and hand it to a marketing team, or keep going.

Step 3: read the comments of the top note

Comments are where the product feedback is. Pass note URLs or 24-character note IDs to mode comments:

top_note = top.iloc[0]["url"]

run = actor.call(
    run_input={"mode": "comments", "notes": [top_note], "maxResults": 100},
    max_total_charge_usd=Decimal("0.50"),
)
comments = list(client.dataset(run.default_dataset_id).iterate_items())

for c in comments[:5]:
    print(c["likes"], c["text"])
Enter fullscreen mode Exit fullscreen mode

A comment item:

{
  "type": "comment",
  "commentId": "6a1f0c2e000000001e03a9b7",
  "noteId": "6aa72fd3000000002802c3f4",
  "text": "Example comment text",
  "likes": 42,
  "replyCount": 3,
  "publishedAt": "2026-09-15T08:12:40Z",
  "authorName": "ExampleUser"
}
Enter fullscreen mode Exit fullscreen mode

Comment items include only text, likes, reply count, time and the commenter's public nickname. Commenter IDs, avatars and other personal data are never returned, which keeps your dataset simpler to store and share.

Step 4 (optional): creator profiles

To vet an influencer, use mode user with profile links or user IDs. You get one profile record followed by that creator's recent notes:

run = actor.call(
    run_input={
        "mode": "user",
        "users": ["https://www.xiaohongshu.com/user/profile/63ebbadf00000000260055d5"],
        "maxResults": 30,
    },
    max_total_charge_usd=Decimal("0.50"),
)
records = list(client.dataset(run.default_dataset_id).iterate_items())
profile = next(r for r in records if r["type"] == "user")
notes = [r for r in records if r["type"] == "note"]
print(profile["nickname"], profile["followers"], profile["likesReceived"], len(notes))
Enter fullscreen mode Exit fullscreen mode

The profile record has userId, url, nickname, bio, followers, following, noteCount, likesReceived and verified.

To get the full record of a single note (including hashtags), use mode note:

actor.call(run_input={"mode": "note", "notes": ["6aa72fd3000000002802c3f4"]})
Enter fullscreen mode Exit fullscreen mode

xhslink.com share links work too, so you can paste links straight from the app.

What it costs

The Actor is pay-per-event:

Event Price
Run start $0.01 per run
Note from a list (search, creator notes) or a comment $5 per 1,000
Note detail (mode note) $20 per 1,000
Creator profile $20 per 1,000

Check the store page for the current price on your plan. Search results are the cheapest way to collect notes. Use note mode only when you need the full record, such as hashtags.

Going further

  • Schedule it. Create a saved task with your keyword and run it daily from the Apify Console to track a trend over time.
  • No-code exports. Every run's dataset downloads as JSON, CSV or Excel, and it connects to Google Sheets, Zapier and Make.
  • AI agents. The Actor is available through the Apify MCP server, so an agent can call it directly.
  • Other platforms. The same approach works for Weibo (including the trending list) and Bilibili.

The data is public content only. You are responsible for using it in line with applicable laws and Xiaohongshu's terms.

If a field you need is missing, open an issue on the Actor page. I read them.

Top comments (0)