This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Basira (بصيرة, "insight" in Arabic) is a private, offline document reader for people with low vision. You take a photo of a paper, a bill, or a medicine leaflet, and Basira reads it aloud and explains it in simple words. Everything runs on your own computer, with no cloud, no subscription, and no API keys.
I built it for my grandmother, who has low vision. [Add one real detail here, e.g. "She can no longer read the small print on her medicine boxes, so she has to wait for someone in the family to read them to her," or "Every bill that arrives is a mystery until someone has time to read it."] I wanted her to be able to read her own mail, understand it, and not feel like a burden when she needs to know what a paper says.
Documents like bills and medicine leaflets are also private and sensitive, which is why I wanted everything to stay on her computer and never go to a cloud service.
Here is how Basira helps:
- Capture: she uploads or photographs a document
- Read: OCR turns the image into text
- Explain: a small local AI model rewrites it in simple language and highlights amounts, dates, deadlines, and what she needs to do
- Speak: a local neural voice reads the original text or the simplified explanation aloud
It works in Arabic and English, has a very large-text, high-contrast interface, and lets her choose between the exact original text and a simple explanation.
Demo
https://drive.google.com/file/d/1m0cCmfTIbDJIFtRY_ezC6SslruRhPIjK/view?usp=sharing
Basira runs locally by design, so there is no hosted demo link. Setup instructions are in the repo.
Code
aelaraby6
/
Basira
Offline document-to-speech reader for Arabic and English powered by Tesseract, Ollama, and Piper TTS
Basira
A private, offline document reader for people with low vision. Take a photo of a paper, bill, or medicine leaflet. Basira reads it aloud and explains it in simple words, all on your own computer, with no cloud and no subscription.
Table of Contents
- What Basira Does
- Why Open Source
- How It Works
- Tech Stack
- Requirements
- Step-by-Step Setup
- Project Structure
- Run It
- Use It From a Phone
- Troubleshooting
- Safety Notes
- Future Ideas
- License
1. What Basira Does
| Step | What happens |
|---|---|
| 1. Capture | The user uploads or photographs a document. |
| 2. Read | OCR (text recognition) turns the image into text. |
| 3. Explain | A small local AI model rewrites the text in simple, clear language and highlights key information: amounts, dates, deadlines, and required actions. |
| 4. Speak | A local text-to-speech voice reads the text or explanation aloud. |
Main features
- Works in Arabic and English
- Very large text and high-contrast interface
- Option…
How I Built It
The whole pipeline is open source and runs locally:
| Job | Tool |
|---|---|
| Language | Python 3.10+ |
| Web interface | Gradio |
| OCR | Tesseract (pytesseract), with Arabic and English data |
| Image cleanup | Pillow |
| Local AI model | Ollama running qwen2.5:3b
|
| Text to speech | Piper (piper-tts), voices en_US-lessac-medium and ar_JO-kareem-medium
|
The flow is: photo → Tesseract → local LLM via Ollama → Piper → audio + text, wrapped in a Gradio UI.
The code is split into small modules so each part can be swapped independently:
-
services/ocr.py: image preprocessing and OCR -
services/llm.py: prompts and explanation through Ollama -
services/tts.py: voice synthesis and caching -
core/processor.py: the pipeline that connects them -
ui/app_ui.py: the accessible interface
It needs about 8 GB of RAM (16 GB recommended) and no GPU. Since Gradio serves the UI over the local network, my grandmother can also open it on her phone, take a photo with the camera, and hear the result.
[Optional: mention one real challenge you hit, e.g. Arabic OCR quality on blurry photos, getting the LLM to give short, simple explanations in Arabic, or choosing the right 3B model. Judges like honest details.]
Why Does Open Innovation Matter?
Bills, medical papers, and ID documents are some of the most sensitive things a person owns. A closed cloud API would mean sending those private documents to someone else's servers, paying per page, and needing a stable internet connection. For many people, that is a dealbreaker.
Open-weight models and open-source tools made it possible to build something that:
- Keeps every document on the user's device, so nothing leaves the computer
- Works offline after the initial setup
- Is free forever, with no subscriptions or rate limits
- Supports Arabic properly, with a modular design where the model, voice, or OCR language can be changed without rewriting anything
Most of all, open tools let me shape Basira around one specific person. I could tune the prompts to explain things the way my grandmother needs, in simple Arabic and without overwhelming detail, something no generic closed product would do for her.
Accessibility tools should not depend on a company's pricing or an internet connection.
Top comments (0)