DEV Community

leonmch
leonmch

Posted on

How to Add Chinese NLP to Your AI Agent in 5 Minutes

Most LLMs read Chinese well at the sentence level. What they handle badly is the
mechanical layer underneath: cutting a paragraph into words the way a search
index would, romanising names so they sort correctly, pulling keywords with
weights, or scanning text against a list you maintain yourself.

Those four operations are cheap to run locally. I packaged them as an MCP server.

The problem

Chinese has no spaces between words, so "我喜欢北京天安门" has to be cut
somewhere, and where you cut it changes what everything downstream sees. Segment
it wrong and your keyword extractor returns fragments, your index misses
matches, your slug generator emits gibberish.

LLMs guess at this. Ask one to split a paragraph into words and you get something
plausible that does not match any tokenizer on your stack. Ask it to convert
names to pinyin for sorting and it will be right most of the time, which is
harder to catch than a consistent error.

The usual fix is a cloud API. That works, and it means every document your agent
processes leaves your machine, plus a latency dependency and a cost line.

The solution

chinese-nlp-mcp exposes four working tools over stdio, all in-process:

  • segment_chinese — jieba segmentation, default / search / index modes
  • convert_pinyin — pypinyin, tone / tone2 / initials / first_letter
  • extract_keywords — TF-IDF extraction with weights
  • detect_sensitive_words — Aho-Corasick against a caller-supplied word list

Plus hello_world for health checks.

Install

uvx chinese-nlp-mcp
Enter fullscreen mode Exit fullscreen mode

Or with pip:

pip install chinese-nlp-mcp
chinese-nlp-mcp
Enter fullscreen mode Exit fullscreen mode

Python 3.10+. Also on PyPI and the Official MCP Registry as
io.github.leonmch-byte/chinese-nlp-mcp.

Setup

Claude Desktop (claude_desktop_config.json):

{
  "mcpServers": {
    "chinese-nlp-mcp": {
      "command": "uvx",
      "args": ["chinese-nlp-mcp"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Cursor (.cursor/mcp.json):

{
  "mcpServers": {
    "chinese-nlp-mcp": {
      "command": "uvx",
      "args": ["chinese-nlp-mcp"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

VS Code (.vscode/mcp.json) uses servers, not mcpServers:

{
  "servers": {
    "chinese-nlp-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["chinese-nlp-mcp"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Demo

Four product reviews, the kind you get pasted into a ticket:

物流很快,包装完好,质量不错。
客服态度很好,解决问题很耐心。
发货太慢了,等了五天才到。
质量太差了,收到就是坏的,申请退款。
Enter fullscreen mode Exit fullscreen mode

On its own this is fine for an LLM to summarise. But if you want to route it,
group it, or track themes over time, you need keywords first. Ask Claude to call
extract_keywords:

质量   0.610
发货   0.536
退款   0.531
太慢   0.510
太差   0.500
客服   0.456
五天   0.450
完好   0.431
Enter fullscreen mode Exit fullscreen mode

Weights are raw TF-IDF, unnormalised, so they exceed 1.0 on longer documents.
质量 and 发货 are the themes; 退款, 太慢, 太差 are the complaints. That
is a routing decision the LLM can now make on numbers instead of vibes.

Two more calls make it useful. Scan against your own escalation terms:

detect_sensitive_words(text=reviews, words=["退款", "太差", "坏的"])
→ {'matches': [{'word': '太差', 'index': 48},
               {'word': '坏的', 'index': 56},
               {'word': '退款', 'index': 61}], 'clean': False}
Enter fullscreen mode Exit fullscreen mode

And segment when you need offsets rather than topics:

segment_chinese("物流很快,包装完好,质量不错。", mode="search")
→ ['物流', '很快', ',', '包装', '完好', ',', '质量', '不错', '。']
Enter fullscreen mode Exit fullscreen mode

(jieba keeps punctuation as its own token; filter it out if you don't want it.)

One gotcha worth knowing: initials and first_letter are both per
character
. 中国 is zhong guo, so you get zh g and z g, not zh and z.
Take [0] yourself if you want one letter per word.

Free vs Pro

The free tier is 500 segment_chinese, 300 convert_pinyin, 100
extract_keywords and 200 detect_sensitive_words calls per day, resetting at
00:00 UTC. That covers individual use comfortably. hello_world is never
throttled, so a health check can never look like a broken server. Hitting a limit
returns an error naming the limit, the reset time and what Pro adds.

Pro is $19 once, not a subscription. It adds batch processing (100 texts per
call), custom dictionaries and priority support. Batch is the one that matters
if you are running this over a corpus rather than a ticket queue — without it
you pay one round trip per document, and the per-call overhead dominates on
short texts. Custom dictionaries matter if your domain has vocabulary jieba has
never seen, which for most Chinese technical corpora it has.

Privacy

No license key check, no network call, no telemetry. Your text never leaves the
machine. FastMCP's update check is disabled at startup; jieba, pypinyin and the
automaton all run in-process.

Daily counters live in a local file (%LOCALAPPDATA% on Windows,
~/Library/Application Support on macOS, $XDG_STATE_HOME on Linux), signed
with HMAC-SHA256.

Two things we state plainly. First, the daily counters are tamper-resistant but
not tamper-proof — this is MIT-licensed local software and anyone determined
enough can patch it. Second, there is no DRM. Device binding records a hashed
machine ID and warns past three devices, but never blocks. Losing a little
revenue beats locking out someone who paid.

Try it

https://github.com/leonmch-byte/chinese-nlp-mcp

Segmentation bugs are the interesting reports. If jieba gets a name, a jargon
term or a domain term wrong, that is the bug I want to hear about. So is
feedback on whether search should be the default mode, and on whether
initials needs a word-level variant.

Top comments (0)