I'm building a side project, WeChat Source. It's a directory that helps Chinese readers find WeChat Official Accounts (the newsletter-style publications inside WeChat). The site is in Chinese, but the design problem behind it isn't: how do you show LLM-generated ratings without them turning into confident-sounding noise?
Here are the rules I ended up with. All of them are visible on every account page.
1. Fix the evidence window
Each account is rated on up to its 20 most recent articles with readable text, using at most the first 6,000 characters of each. If an account has fewer than 20 usable articles, it gets no overall score. A small sample produces a confident-looking number with little behind it, so I'd rather show nothing.
2. Every rating carries quotes and a limitation
There are five dimensions: depth of argument, source transparency, information gain, clarity, and caution in stating opinions. Each gets a 1–5 score plus:
- a short explanation,
- a "limitations" paragraph saying what the account does badly on that dimension, and
- quotes from specific articles (with article IDs) as evidence.
The page also states the scale: 3 is acceptable, 4 is good, and 5 requires strong evidence.
3. Verify the quotes, and degrade gracefully
LLMs sometimes "quote" text that doesn't exist. Quotes are checked against the stored article text, and any that don't match are removed. If that leaves a dimension without solid evidence, it's displayed as "can't judge", with a note that the quotes failed verification and the item needs re-analysis or a human review.
Importantly, "can't judge" is not counted as zero. Treating "unknown" as "bad" would quietly punish accounts for the model's mistakes.
4. Show provenance
Each profile shows how many articles it's based on, the model name, a rubric version string, and the analysis date. There's also a line saying the score is the platform's AI analysis, not an official evaluation. If the rubric changes, older profiles are at least identifiable.
5. Don't let similarity pose as quality
At the bottom of each account page there's a "keep exploring" list. It's a rule-based match on shared categories and tags, and the page says exactly that, including the similarity percentage and the shared category. It would be easy to let people assume these are quality recommendations. They aren't, so the label says so.
What it doesn't solve
- Coverage is the real bottleneck. About 180 accounts are curated so far. The ~9,900 total includes many that have only basic info and no articles collected yet.
- The model can still be wrong about tone or intent. Quotes make that visible, but they don't prevent it.
Stack
Next.js front end (served under a /wechat base path), a Python back end, and Cloudflare in front.
If you've shipped LLM-generated ratings or summaries, I'm curious how you handle evidence and "unknown" states. And if you read Chinese, the site is here. Feedback is very welcome.
Top comments (0)