I am developing a Rohingya OCR tool: software that turns an image of Hanifi Rohingya writing into editable text. My wider work includes collaboration with 100 Hanifi teaching schools in camps and Saudi Arabia, dictionary development, and an ongoing translation project targeting 300,000 sentences.
For OCR, those activities raise a specific engineering question: what should count as a trustworthy training example?
In my previous article, I described a browser extension that changes the writing system of text already on a webpage. OCR starts earlier. The input is pixels, and the correct text is something we need to establish.
This post describes the data preparation work in my local prototype and a review workflow I propose for the next stage. The OCR tool is still in development. I am not reporting a production recognizer or an accuracy result from the school collaboration.
Start with the right unit of annotation
For a line recognizer, one image should correspond to one transcribed line. Tesseract's tesstrain workflow uses paired line images and single-line text files, for example:
page-017-line-03.png
page-017-line-03.gt.txt
The .gt.txt file contains the exact text visible in that crop. Matching filenames establish the pairing; they do not establish that the transcription is correct.
One real example in my local project contains two visual lines, while its supplied transcription is a paragraph. The training notes explicitly keep it as a whole-image example until each line has an exact separate transcription. Assigning the entire paragraph to either crop would teach the recognizer the wrong correspondence.
A practical next step is to keep a manifest alongside the pairs. This is an illustrative schema, rather than a claim that all these fields have already been collected:
{
"sample_id": "page-017-line-03",
"source_document_id": "doc-017",
"image": "page-017-line-03.png",
"label": "page-017-line-03.gt.txt",
"origin": "real_scan",
"review_status": "needs_second_review"
}
The document ID matters when we create an evaluation split. Review status matters when deciding which labels are ready for training. Permission to use an image also needs to be recorded separately from the transcription.
Preserve the Unicode text and make review readable
The Hanifi Rohingya Unicode block contains letters, vowels, marks, tone signs and digits. Checking only the large letter shapes would overlook some of the information that the label needs to preserve.
For a browser review screen, I would use explicit text direction and language metadata:
<textarea dir="rtl" lang="rhg" spellcheck="false"></textarea>
dir="rtl" tells the browser how to display right-to-left text. It does not mean we should reverse the stored string. The local prototype already uses this pattern for its transcription box.
A font that covers Hanifi is also necessary for visual review. A correct Unicode label that appears as boxes is difficult to check; a readable font alone cannot establish that a label matches the image. My Hanifi Unicode reference is useful for inspecting characters while reviewing labels.
Before comparison, choose and document a normalization policy. My evaluation helper uses NFC and collapses whitespace:
import unicodedata
def normalize(text):
return " ".join(unicodedata.normalize("NFC", text).split())
Python's unicodedata.normalize supports canonical normalization. This helper does not remove tone or nasalization marks. Collapsing whitespace is a deliberate evaluation convention, so the original transcriptions should remain available when spacing fidelity matters.
Use synthetic images for a defined experiment
My local generator creates 1,800 word sequences from dictionary forms converted to Hanifi, followed by 240 character-focused sequences. These are synthetic text lines, not 2,040 independently authored Rohingya sentences.
The renderer draws them using available Hanifi fonts, varies font size and padding, and sometimes adds slight blur or rotation. Each image is paired with the string used to render it.
That makes the label traceable to the rendering operation. It does not independently validate the spelling produced by the converter. A rule-based script converter is useful for bootstrapping an experiment, but its output still needs linguistic review if it will be treated as authoritative text.
Synthetic images also have a narrower visual distribution than real photographs. Their fonts, backgrounds and distortions come from the rendering script. Phone photos can contain lighting gradients, perspective changes, compression and damaged print that the generator does not reproduce.
I would therefore keep synthetic evaluation and real-image evaluation separate, and report exactly which one produced each result.
Split real data by source document
Randomly splitting line crops can put different lines from the same page into training and evaluation. The recognizer may encounter the same printing, scanning conditions and document layout in both sets.
For the proposed real-data workflow, I would assign a whole source document to one split before distributing its line crops. Another useful check is to reserve documents from an unseen acquisition batch. The aim is to test a stated kind of generalization, rather than assume that one split covers every deployment condition.
Document grouping does not automatically remove repeated passages across books. Shared source content and duplicate images still need their own checks. Handwriting, if included later, would also need a split that considers the writer.
Define the denominator before quoting an error rate
My evaluation helper uses edit distance between normalized Unicode strings. Here is a self-contained version of the core calculation:
def edit_distance(reference, prediction):
row = list(range(len(prediction) + 1))
for i, expected in enumerate(reference, 1):
next_row = [i]
for j, actual in enumerate(prediction, 1):
next_row.append(min(
next_row[-1] + 1,
row[j] + 1,
row[j - 1] + (expected != actual),
))
row = next_row
return row[-1]
def codepoint_cer(reference, prediction):
reference = normalize(reference)
prediction = normalize(prediction)
if not reference:
raise ValueError("CER needs a nonempty reference")
return edit_distance(reference, prediction) / len(reference)
This counts Unicode code points, not grapheme clusters. That choice must accompany the number. Insertions can make CER exceed 100%, and a small mark error can affect a word even when the total rate is low.
For a collection of lines, sum the edit distances and divide by the total number of reference code points. Averaging each line's percentage would give short and long lines equal weight. I would also report exact line matches and show examples of the remaining errors, including missing marks and digits.
Make human review part of the data pipeline
The next stage I propose is a small real-document pilot: establish permission to use the material, crop lines, obtain transcriptions, have a second reader review them, and resolve disagreements before freezing an evaluation set. The manifest should distinguish verified labels from unresolved examples.
My collaboration with teaching schools offers a route to discussing that process with people who use Hanifi. It does not establish that every school has agreed to contribute OCR data, or that their materials are ready for publication.
The technical work is useful when another person can inspect the image, trace its label, understand the evaluation split and reproduce the reported check. That is the foundation I want in place before presenting a Hanifi OCR model as ready for everyday use.
AI assistance was used to help draft this article. The described prototype behavior was checked against the local implementation; the proposed review workflow is identified as future work.
Top comments (0)