DEV Community

Cover image for Pīzhù 批注: a writing tutor that never rewrites your essay, built for my friend Joy (Gemma 4)
Muhammad Murtuza Hussain
Muhammad Murtuza Hussain

Posted on AI-assisted

Pīzhù 批注: a writing tutor that never rewrites your essay, built for my friend Joy (Gemma 4)

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My friend Joy is doing an MA in Global Mass Communication in Ireland. Her first language is Mandarin. She reads media theory in two languages, argues well, and still gets essays back with comments like "expression needs work" and "be careful with academic tone".

Me and Joy met at the library of our university, both of us spend most of our day working in the reading room. This one particular afternoon, Joy invited me for a drink break at our student bar, where she also had her informal meeting with her EAP (english for academic purposes) faculty, where they discussed her dissertation, and I noticed how Joy received the feedback from faculty, but couldn't get the reasoning behind it, the faculty all could say was,

It's Just bad english

she said.

This moment had me to tinkering, but also I recalled that, I have seen the same thing with a lot of friends from China and Taiwan. The marks they lose are rarely about ideas. They come from a handful of habits that make perfect sense in Chinese and read badly to a British marker:

  • No articles. Mandarin has no a or the, so "The society is changing" feels natural.
  • 一逗到底, "one comma all the way". A Chinese sentence can run through commas until the thought is complete. In English that is a comma splice.
  • Exam-template English. "With the development of society…", "In a word…" were taught as good style for 高考 and IELTS. A UK marker reads them as filler.
  • Confident claims. "This proves…" and "Everyone knows…" (众所周知) persuade in Chinese argumentative writing. In a UK essay they are unsupported.

The usual tools make this worse. Grammar checkers and chatbots rewrite the paragraph. It sounds fluent, it stops sounding like the student, she learns nothing, and in UK and Irish universities an AI-rewritten essay is an academic-integrity problem.

So I built Pīzhù (批注). The word means margin notes: the comments a good teacher writes beside your work. That is the whole product:

  • It underlines the exact words a UK marker would pause on.
  • It explains why, in English and in Chinese (简体 or 繁體), including why a Mandarin speaker tends to write it that way, framed as a habit, never a mistake in thinking.
  • It gives a hint, and at most a tiny nudge of a few words.
  • It never rewrites. Joy makes every change herself.

Pīzhù's live demo: a Chinglish sentence gets marker-pen underlines and bilingual margin notes

Demo

Try it: pizhu.onrender.com. Press Try a sample, then Check my writing. The sample essay is one I wrote for testing; it is full of these habits and includes one reference I invented on purpose, so you can watch the reference checker catch it. (The demo runs on Render's free tier, so the very first load can take a few seconds to wake up.)

The essay on grid paper, with notes in the margin. Every note sits beside its own line, and the open one draws an ink line to its words.

Margin notes beside the essay

Lesson mode. One note at a time, with tapioca pearls as progress. Enter for "got it", the arrow keys to move.

Lesson mode

Back to Word. Students write in Word, and Word comments are literally margin notes, so one click gives Joy her own document back, every word unchanged, with each note as a comment from Mòmo.

Reference check. Google Scholar via SerpApi: three real books found, the invented one flagged.

Reference check

On a phone, in dark mode, in 简体中文.

Mobile, dark mode, Simplified Chinese

Code

Everything is open source (MIT) at github.com/MuhammadMurtuzaHussain/pizhu. If you only read four files, read these:

File What it does
src/lib/guard.ts The "never rewrite" rule: verbatim spans and a word-level edit budget for nudges
src/mastra/agents/prompts.ts The tutor brief: twelve Mandarin-to-English habits and the checklist Gemma follows
src/mastra/workflows/references.ts The reference check: Gemma extracts, code matches, SerpApi verifies
src/components/MarginPaper.tsx Margin notes positioned beside each line, with the ink connector

Run it privately on your own laptop in three commands:

ollama pull gemma4:e4b
git clone https://github.com/MuhammadMurtuzaHussain/pizhu && cd pizhu && pnpm install
pnpm dev   # http://localhost:3000
Enter fullscreen mode Exit fullscreen mode

Pīzhù 批注: margin notes, not rewrites. Mòmo the ink panda holding a bubble tea.

Margin notes, not rewrites. An open-source writing tutor for Mandarin-speaking students at UK and Irish universities, powered by open-weight Gemma 4

Live demo · DEV post · Run locally · Built for the DEV Hacktoberfest Weekend Challenge

A Chinglish sentence gets marker-pen underlines and bilingual margin notes

Why

My friend Joy is doing an MA in Global Mass Communication in an Irish University. Her ideas are sharp; her marks lose points to a small set of habits that make perfect sense in Chinese and read badly to a British marker: no articles (中文没有冠词), comma splices (一逗到底), exam-template phrases ("With the development of society…"), and confident claims (众所周知).

Grammar tools and chatbots rewrite the paragraph. It sounds fluent, it stops sounding like her, she learns nothing, and in an Irish university it is an academic-integrity risk.

Pīzhù does what a good tutor does in the margin: it points at the exact words, explains why in English and Chinese (简体 or…




How I Built It

Gemma 4 is the only model in Pīzhù. It is open weights under Apache 2.0, and the same family runs in two places:

Where Model Why
Joy's laptop (the default) gemma4:e4b or gemma4:12b via Ollama Her coursework never leaves her machine
Hosted demo gemma-4-26b-a4b-it via Google AI Studio So anyone can try it without installing anything

One environment variable switches between them.

Next.js UI ──► /api/analyze  (streams one paragraph at a time)
                 │
                 ▼
          Mastra tutorAgent ──► Gemma 4  (structured output, Zod schema)
                 │
                 ▼
          guard.ts  ← the "never rewrite" rule, enforced in code
                 │
                 ▼
          + rule-based British spelling  →  margin notes

Reference tab ──► Mastra workflow:
          extract (Gemma) → match citations (code) → verify (SerpApi Google Scholar)
Enter fullscreen mode Exit fullscreen mode

1. Teaching Gemma to annotate instead of rewrite

The tutor is a Mastra agent with a system prompt written like a brief for a university writing-centre tutor. It describes twelve habits, each pairing a UK convention with the Mandarin pattern behind it, and gives Gemma a checklist: check every comma, every "As for…", every bare countable noun.

Gemma returns structured JSON validated with Zod. The schema has no field for a rewritten paragraph. The only text Gemma may propose is an optional nudge.

2. Not trusting the prompt

Prompts are requests, not guarantees. guard.ts checks every note before Joy sees it:

  1. The highlighted span must exist in her text (tolerating curly quotes and spacing). If Gemma "quotes" something she never wrote, the note is dropped.
  2. A nudge may change at most max(4, 30%) of the span's words, measured by word-level edit distance. Anything bigger is stripped back to a hint, and the card says "Try this one yourself."
  3. Spans that cross a sentence boundary never get a nudge. Restructuring is the student's job.
export function nudgeAllowed(span: string, nudge: string): boolean {
  if (!nudge.trim() || nudge.trim() === span.trim()) return false;
  if (words(span).length > 25) return false;
  if (crossesSentence(span)) return false;
  return wordEditDistance(span, nudge) <= Math.max(4, Math.ceil(words(span).length * 0.3));
}
Enter fullscreen mode Exit fullscreen mode

Across my evaluation runs, Gemma proposed 89 nudges and the guard turned 8 of them back into hints. Most came from the smallest model, which is exactly where you want the safety net.

3. Rules for what rules do best

The small model was unreliable at spotting American spellings, so I stopped asking it. A word list catches analyze → analyse and behavior → behaviour every time, and those notes merge with Gemma's. Gemma handles the judgement; the rules handle the lookups.

4. Checking references with SerpApi

Fabricated or garbled references are a growing problem, so the reference tab runs a three-step Mastra workflow:

  1. Gemma extracts each entry (authors, year, title) and offers Harvard formatting hints, never a corrected entry.
  2. Plain code matches in-text citations like (Hall, 1980) against the list: cited but missing, or listed but never cited.
  3. SerpApi's Google Scholar engine looks up every title: ✅ found, ≈ details differ, or ❓ couldn't find it, so check it exists.

Two real-world snags made it better. Scholar often drops subtitles ("Understanding media" for Understanding Media: The Extensions of Man), and its top hit is sometimes a book review rather than the book. Pīzhù compares main titles, searches with Scholar's author: operator first, and accepts a one-year drift for reprints. Only titles and first authors are sent, never essay text.

5. Designed for Joy, not for a generic user

This is where I spent the second half of the weekend, because a tool you feel judged by is a tool you stop opening.

  • Mòmo 墨墨, the ink panda. Black and white is literally 墨 ink on paper. Pandas are loved in both mainland China and Taiwan, and a gentle mascot takes the sting out of being corrected. Mòmo reads with you, sips bubble tea, cheers when you fix things, and presses a 好 seal when you finish. A reduplicated name like 墨墨 feels affectionate, the way 团团圆圆 does.
  • A bubble-tea palette. Taro purple, milk white, strawberry for must-fix, mango for worth-fixing, matcha for all good. Red is the colour of luck and celebration in Chinese culture, so errors are a soft strawberry, never alarm red.
  • Feedback that protects 面子 (face). Notes are private, never graded, and called habits. The "Marker's eye" summary under each paragraph uses UK marking dimensions (argument, structure, language, referencing) and never gives a mark.
  • 集章, collecting stamps. Anyone who grew up with 7-Eleven point cards or station stamp books knows the pull. Work through three notes of a habit and you earn its stamp.
  • Handwriting-style Chinese. Explanations are set in 霞鹜文楷 (LXGW WenKai), an open-source kai typeface many Chinese and Taiwanese students already love, so they read like a teacher's note.
  • The interface in 简体 and 繁體, with explanations in English, Chinese or both.

The stamp book

Why Does Open Innovation Matter?

This project only makes sense with open models.

1. Privacy and academic integrity. Pasting unsubmitted coursework into a third-party chatbot means handing it to a server you do not control, and universities increasingly tell students to be careful with exactly that. With Gemma in Ollama, Joy's drafts stay on her laptop, and the app says so in plain words: 🔒 Local. Your essay stays on this laptop.

2. Control over behaviour. Closed chat products are built to produce the answer. I needed the opposite: a model that points and explains but does not fix. With open weights inside my own pipeline, I could enforce that rule in code instead of hoping a prompt holds.

3. Being able to swap models, and measure it. I wrote fifteen test paragraphs seeded with known habits, plus two clean ones to catch over-flagging, and ran the same pipeline on three sizes of Gemma:

Model (where it runs) Recall on seeded habits Must-fix notes on clean paragraphs Seconds per paragraph
Gemma 4 E4B (laptop, 16 GB) 84.2% 1 12.8
Gemma 4 12B (laptop, 16 GB) 94.7% 0 34.1
Gemma 4 26B-A4B (AI Studio) 89.5% 0 25.4

Joy can choose her trade-off: E4B for quick checks while drafting, 12B before she submits. Tuning the hosted model taught me something too: Gemma 4 "thinks" at length by default, and one paragraph took 335 seconds. Setting the thinking level to minimal brought it to about 30 seconds with no visible loss in quality.

4. Cost. Running locally costs nothing. International students already pay some of the highest fees in UK higher education; a tool behind a subscription is a tool many of them will not use.

Handing it to Joy

I sent Joy the link on WeChat with a slightly nervous "i made this app for u". She didn't test it on my sample essay. She tested it on what she is actually writing right now, her dissertation, and later sent back a review I could not have scripted.

What worked for her:

The general feeling of this app: it's very useful, especially for our module assignments. The Chinese explanation is easier to understand than the criteria written in our module 😂

Another advantage of it is that it explains why I should correct my sentence. As grammarly could only give me a red label but sometimes I don't know that's the problem and how to fix it. The explanation really helps me to understand how to write properly.

That second message is the reason Pīzhù exists, put better than I managed: a red underline tells you that something is wrong, and Joy needed to know why. It is the same gap as "It's just bad English".

She was just as honest about what didn't work:

As my dissertation is a long paper, when I tried to ask advice from momo, I have to send my paper separately, even need to split each chapter to 2-3 parts. Thus, momo cannot get the full understanding of my argument. If the app could do longer articles, it would be great!

(She calls it "momo". That made my day.)

And she had two requests:

Another advice I would like to share it that, if it could help you analyse the assignment requirements or criteria, and have similarity detection like turntin, it will help students understand how to do their assignments and worry less before submitting.

What I shipped the same evening

  • Longer papers. The hosted demo now accepts 6,000 words instead of 1,500. The cap only exists because the demo runs on free tiers; running Pīzhù locally has no limit at all.
  • A whole-paper overview. Mòmo now reads the entire paper in one pass, alongside the line-by-line notes, and tells you what it thinks your argument is, what works, and the three most important whole-paper fixes. If Mòmo's summary of your argument isn't what you meant, that is the most useful feedback of all.
  • Your assignment brief, explained. Paste the task and the marking criteria, and Mòmo restates each criterion in plain English and Chinese, then marks it met, partly or missing against your paper. Joy's line about the Chinese being easier than her module criteria was the whole spec.
  • It travels. The overview sits at the top of the Word export, above her own unchanged text.

Whole-paper overview checked against an assignment brief

Turnitin-style similarity checking is the one request I'm deliberately not building. It only works by comparing your essay against a database of other people's papers, which means uploading and keeping your writing: exactly what Pīzhù is designed never to do. It does the honest half of that job instead: it flags claims that need a citation and checks that every reference actually exists.

Her last message: "No problem, it's already a brilliant app!" I've promised her half the prize if it wins.

My Agent Session

I built Pīzhù with Claude Code over the weekend, from reading the challenge rules, to the PRD, to debugging the reference matcher and the margin-note layout:

Building Pīzhù 批注 with Claude Code: a margin-note writing tutor on Gemma 4
You

Today's Hacktoberfest challenge is Build for a Friend. My idea: a tool that helps Mandarin-native students write better academic English (EAP). Lots of my friends from China and Taiwan lose marks and confidence when they're so smart. I do NOT want the model to rewrite their work, just tell them where the English is off and why. Focus on UK and Irish academic writing, UI relevant to Mandarin speakers, with a Chinese UI option. Plan it against the submission template and partners, build a PRD, and ask me questions.

Agent

Read the challenge rules (deadline 5 Oct 06:59 UTC; writing quality weighted most) and checked partner tech. Gemma 4 is open-weight under Apache 2.0, multilingual and runs locally, so it is the core model; Tinker can't fine-tune Gemma, so I left it out. Asked four questions: the friend, where Gemma runs, partner categories, and stack.

You

Named friend who will test (Joy, Global Mass Communication, Simplified Chinese). Hybrid inference. Gemma with SerpApi. TypeScript: Next.js + Mastra. 16 GB RAM.

Agent

Wrote the PRD: 'annotate, don't rewrite'. Built a Mastra tutor agent with a Zod schema that has no field for a rewritten paragraph, plus guard.ts: every span must exist verbatim in the student's text, and any nudge changing more than max(4, 30%) of the words is stripped back to a hint. First local run on gemma4:e4b: six notes on paragraph 1 with natural Chinese explanations, one oversized fix stripped by the guard.

Agent

Wrote 15 hand-made test paragraphs seeded with Mandarin-L1 habits plus 2 clean ones. E4B recall 73.7%, weakest on American spellings and missing citations. Added a deterministic US→UK spelling list (rules for lookups, Gemma for judgement) and a sharper checklist: recall rose to 84.2%. 12B: 94.7% with zero false must-fix notes.

Agent

Hosted path: one paragraph took 335 s on Gemma 4 31B because it thinks at length. Probed the API: 31B was returning 500s, and 26B-A4B with thinkingLevel 'minimal' answered in ~7 s. Switched the hosted demo to 26B-A4B; paragraphs now take ~30 s with no visible quality loss (89.5% recall).

Agent

Reference checker as a Mastra workflow: Gemma extracts entries, plain code matches in-text citations, SerpApi Google Scholar verifies. Two real snags fixed: Scholar drops subtitles ('Understanding media'), and its top hit is sometimes a book review. Compare main titles, search with author: first, allow ±1 year for reprints. Result: three real books found, the invented one flagged.

You

Deploy failed (Render: EROFS read-only file system, unlink /usr/bin/pnpm).

Agent

corepack enable tries to write a pnpm link into Render's read-only /usr/bin. Switched render.yaml to corepack pnpm install/build/start, verified a local production build, pushed. It went live at pizhu.onrender.com.

You

Load the design skills we used for my website and enhance the UI/UX and colours. It's for a Chinese audience, so maybe a dragon or panda avatar, chosen by the consumer psychology of Chinese and Taiwanese students.

Agent

Recommended 墨墨 Mòmo, an ink panda: black and white is literally ink on paper, pandas are loved in both the PRC and Taiwan, and a dragon reads as authority, intimidating for something that corrects you. Design principles: protect 面子 (private, never graded, 'habits' not errors), 萌 mascot culture, 集章 stamp collecting, and 霞鹜文楷 for Chinese so notes read like a teacher's handwriting.

You

It's still lackluster, and the palette looks like the Claude colour scheme. Can we think of something much more?

Agent

Offered four palette worlds; you chose 珍奶 taro & boba, plus all four signature features. Built: a live demo hero (sentence types in, underlines draw, notes pop in); true margin notes positioned beside each line with an ink connector; lesson mode with tapioca-pearl progress and a share card; and a Mòmo companion dock that narrates, cheers and stamps 好.

You

Add a light/dark toggle. Bug: the line from words to note stays while I edit a paragraph. The hero demo grows and shrinks so the page jumps. Export my habits gives JSON, improve it. Credit the challenge and me in the footer. Make this a great finished product that would make us win!

Agent

Fixed all five, each verified in the browser: the toggle applies before first paint; editing hides that paragraph's notes and connector; the demo card reserves its final size (height constant across a full loop); a habit-report image and 集章 stamp book replace raw JSON; and the footer has credits. Extras: Word export with notes as real Word comments, friendly errors with retry, draft autosave, WCAG contrast fixes, OG card. A headless-Chrome script ran the whole flow on the live site and captured the README screenshots.

Prize Categories

  • Best Use of Gemma: Gemma 4 is the only model, running locally (E4B and 12B via Ollama) and hosted (26B-A4B via AI Studio), with a measured size comparison and a guardrail built around its structured output.
  • Best Use of SerpApi: the Google Scholar engine powers the reference checker that catches sources which may not exist, using the author: operator to find the work itself rather than reviews of it.

Built with AI assistance (Claude Code). I reviewed, tested and edited everything, and the evaluation numbers come from real runs of the code in the repo.

Top comments (0)