DEV Community

Wadifa Info
Wadifa Info

Posted on

OCR for 1,700 scanned exam papers in Arabic and French with Tesseract (and the traps we hit)

We run Wadifa Info, a Moroccan site that lists public-sector recruitment exams in Arabic and French. One part of the site is a free library of past exam papers: more than 1,700 of them, from police and customs exams to teacher-training and engineering-school entrance exams (French index).

Most of these papers arrive as photos or scans. A page that is only an image can't be searched, can't be quoted in a page description, and tells a search engine almost nothing. So every page needs its text. This post covers how we OCR'd the whole library on one Windows machine with Tesseract, and the four traps that made "done" turn out to be not done.

Setup: Tesseract 5 with the right language data

Tesseract 5.4 installs on Windows with one command:

winget install --id UB-Mannheim.TesseractOCR -e
Enter fullscreen mode Exit fullscreen mode

It ships with English only. Arabic and French data come from the tessdata_fast and tessdata_best repositories, and you point Tesseract at them with --tessdata-dir (handy when you can't write to Program Files).

We measured the options on a real bilingual page, counting recognised Arabic and Latin characters:

Setting Arabic chars French chars
ara+fra (fast models), raw image 674 914
ara_best+fra, preprocessed 805 998
ara_best alone good French destroyed

Two lessons: the best Arabic model is clearly better on photographed pages, and a single-language model mangles the other script, so bilingual pages need both (-l ara_best+fra).

Preprocessing that actually helped

Three cheap steps, in Pillow, before every page:

def prep(src):
    im = Image.open(src).convert("L")                  # grey
    if im.width < 1800:                                # small phone photos
        im = im.resize((im.width * 2, im.height * 2), Image.LANCZOS)
    return ImageOps.autocontrast(im)
Enter fullscreen mode Exit fullscreen mode

Then tesseract page.jpg stdout -l ara_best+fra --psm 6. Page segmentation mode 6 ("a single uniform block of text") beat the automatic mode on exam sheets, which are mostly dense paragraphs and numbered questions.

On the first pass this filled 286 of 289 papers that had no text, median about 2,900 characters per paper.

Trap 1: photographed pages that return only our own footer

A few phone photos came back with exactly one line: the footer we add to every page. Tesseract found no text blocks at all, at any --psm, at any scale. The page wasn't blurry; it had uneven lighting, so a global threshold turned half the sheet into one dark blob.

The fix was an adaptive threshold, done in Pillow without OpenCV: compare each pixel with a heavily blurred copy of the page, and keep only what is darker than its local background.

bg   = im.filter(ImageFilter.GaussianBlur(20))
diff = ImageChops.subtract(bg, im)            # ink is darker than the local background
th   = diff.point(lambda v: 0 if v > 12 else 255)
Enter fullscreen mode Exit fullscreen mode

We now try four variants per page (plain or adaptive, --psm 4 or 6) and keep the one with the most real words. "Real word" is deliberately crude: Arabic words of three letters or more, and Latin words of three letters or more that contain a vowel. Our worst three papers went from a single footer line to 4,800–5,600 characters each.

Scoring trap inside the trap: we first stripped footer lines before scoring the old text. A one-line OCR result is the footer line, so stripping it scored the old text as zero and every comparison looked like a huge win. Strip the footer phrase, not the line.

Trap 2: a page cap we forgot about

The first script had MAX_PAGES = 8 to keep test runs short. It never came out. Fifty-eight papers longer than eight pages had pages 9+ silently skipped, about 400 pages in total.

What caught it was a simple health check: for each paper, compare the number of images we publish with the highest page marker in its text. Our OCR writes a marker between pages (— صفحة N —), so a 12-page paper whose text stops at page 8 is easy to spot. We now run that check after every import.

Trap 3: papers that are only half there

This one wasn't an OCR bug; OCR found it. Many official papers print their own footer: Page 2 sur 5, or الصفحة 2 من 5. If the paper says "page 1 of 9" and we publish 4 images, we are missing pages.

Matching that pattern across the library found nine incomplete papers. One of them showed only "Page 2 sur 2": visitors were opening the paper halfway through, at exercise 3.

Two things made the detector usable:

  1. Never match a bare N/M. Moroccan exams are marked out of 20 or 10, so 14/20 is a score, not a page number. The naive pattern produced 34 candidates, almost all wrong. Only explicit forms count: Page N sur M, Page N | M, الصفحة N من M.
  2. Read the context. "Répondre sur le document réponse (page 06/10)" is a cross-reference, not a footer.

Trap 4: duplicates that look different, and different papers that look the same

The same paper often reaches us twice: a clean PDF and a phone photo of a printout, or the same engineering-school subject adapted for two streams. Image hashing on a page region gave false positives (shared headers), and comparing only the first 3,000 characters did too: the adapted version shares the opening, then diverges.

What worked was comparing the full OCR text at the word level:

from difflib import SequenceMatcher
same = SequenceMatcher(None, a.split(), b.split()).ratio() >= 0.84
Enter fullscreen mode Exit fullscreen mode

Above 0.84 it was always a true copy in our checks; below it, a genuinely different paper. Seven copies were hidden this way.

Where it ended up

  • Every published paper has text, except a handful of genuinely unreadable photos that we leave empty on purpose rather than fill with noise.
  • Each paper page shows the transcription in a collapsed "text of the questions" box, with a note that it was extracted automatically and the images remain the reference. Candidates can search and copy the questions, and search engines can finally read what each paper covers.
  • The same text drives our duplicate check when new papers come in.
  • The health checks (page count vs. last marker, very short texts, printed "Page N sur M") run after every batch.

If you're digitising scanned documents in Arabic, or any mixed-script material, the short version is: use the best model, always pair the scripts, preprocess, keep an adaptive-threshold fallback, and never trust "100% done" until a check you didn't write for the happy path agrees.

The papers themselves are free to browse at wadifa-info.com.


Written with the help of an AI assistant from our own code and logs; all figures are from our production data.

Top comments (0)