🎥 The 6-minute demo video below shows a full Tamil interview, the record filling, doctor approval and prescription reading. It's the best way to see what KuralMD does.
The friend
Dharini is a final-year MBBS student. On outpatient days in Community Medicine she sees one patient after another, and for each one she does the same two jobs: she takes the history, then she types it into the hospital's web portal.
Taking the history is where the time goes. "Fever" is never just fever. How many days? Does it go up in the evening? Any chills? Did you check the temperature? Burning when you pass urine? A cough needs the same treatment: dry or with phlegm, what colour, any breathlessness, worse at night? After that come past illnesses, medicines, allergies, family, diet, sleep, smoking, work. Most of her patients answer more comfortably in Tamil, so she translates in her head while typing English into the form.
When I asked what would actually help, she did not want an "AI doctor". She wanted someone to ask the standard questions in the patient's language and put the answers in the right boxes, so she could check them and move on.
That is KuralMD. Kural (குரல்) means "voice" in Tamil.
What I Built
The patient talks; the form fills. The patient presses a mic button and answers in Tamil, English, or the mix most people actually speak ("எனக்கு two days ஆ cold இருக்கு"). KuralMD asks one short question at a time, out loud, in everyday spoken Tamil rather than textbook Tamil. The next question depends on the last answer. Mention a cold and it asks about fever and body pain. Mention a cough and it asks whether there is phlegm and what colour it is. Watch the video to know how the full interview works and the follow up questions are asked.
The doctor gets a draft, not a decision. When the interview ends, KuralMD fills the history section of the outpatient (OP) record: chief complaint, history of presenting illness, past, drug, allergy, family, personal and occupational history. Every filled field carries a small "from patient" chip. Tap it and you see the patient's exact words in Tamil next to the English translation. If the model could not point to something the patient actually said, the field stays empty and turns amber.
The Patient can also upload old prescriptions and Gemma 4 with langextract extracts medicine names and other parameters.
The doctor edits whatever is wrong, can send the patient a targeted follow-up ("ask the exact duration of the fever"), and approves the record.
What KuralMD will never do is diagnose, suggest a medicine, or touch the examination section. Those fields stay empty and say "Doctor only". Dharini drew that line before I wrote any code, and the app keeps to it.
Demo
Code: https://github.com/VidhiyaSB/KuralMD
How I Built it
mic → ElevenLabs Scribe (speech to text)
→ Gemma 4: translate the answer and choose the next question, in one call
→ guardrails: red flags, no clinical advice
→ ElevenLabs text to speech → the patient hears the next question in Tamil
finish → LangExtract on Gemma 4 → grounding check → OP record → doctor approves
Gemma 4 through Ollama does the reasoning: translating, choosing the next question, pulling fields out of the transcript, and reading photos of old prescriptions. It is an open model and it runs on my own laptop.
ElevenLabs handles Tamil speech both ways. Its speech-to-text read my test Tamil at 0.99 confidence. A voice app that stops when one API call fails is useless in a clinic, so there are fallbacks: Google TTS if ElevenLabs speech fails or hits a rate limit, and Chrome's built-in recognition if transcription fails. If everything fails, the patient can type.
Google's LangExtract turns the transcript into fields. I chose it because each extraction comes back with the exact span of text it was taken from. In a medical record, that link back to the source matters more than anything else.
Langfuse gets one trace per encounter. Patient names and phone numbers are masked before anything leaves the machine. The same steps are also stored in the app, so the portal has its own timeline even when Langfuse is down.
The in-app pipeline for one encounter: speech-to-text, agent, guardrails, extraction, with timings : screenshot from the demo video

The rest is deliberately boring: FastAPI, SQLite, Jinja templates, Tailwind and Alpine.js, with no build step.
Guardrails, because this is medical
Every value needs a quote. The extractor has to say which words of the patient's each value came from, and I check that those words really appear in the transcript. Speech-to-text and translation are never word-perfect, so the check is fuzzy:
def grounding_score(quote, transcript):
q, t = normalise(quote), normalise(transcript)
return 100.0 if q in t else fuzz.partial_ratio(q, t)
# below 80: the value is dropped, the field is marked "unclear", the doctor sees it in amber
If the model invents an allergy the patient never mentioned, it does not reach the form.
Red flags stop everything. The red flags are chest pain, severe breathlessness, fainting, blood in vomit or sputum, and suicidal thoughts. They are checked twice: once by a plain keyword list in both English and Tamil (with a rough negation check, so "no chest pain" does not trigger it), and once by the model. Either check can end the interview. KuralMD then tells the patient in Tamil to inform the nurse immediately and puts the encounter at the top of the review queue marked urgent.
No advice slips through. Every question is checked for diagnosis or treatment language before it is spoken. If it fails, a safe pre-written question is asked instead.
Returning patients aren't asked everything again. If someone has an approved record from an earlier visit, the assistant reads it back ("last time you said no sugar or BP, has anything changed?") instead of asking ten questions. Anything the patient does not re-confirm is marked for the doctor to verify.
Old prescriptions.
Patients often show a photo of an old prescription instead of remembering medicine names. In the video I upload one, a handwritten-style sample, and Gemma 4 reads the medicines into drug history, each tagged "from prescription" so the doctor knows where it came from. My test images are printed or neatly handwritten. A real doctor's handwriting is a harder problem that I have not solved yet.
Please watch the video for the complete demonstration
Local first, and what the video shows
KuralMD is built to run on local models: Gemma 4 e4b on Ollama, on the same laptop that serves the portal.
Voice audio goes to ElevenLabs, and the masked traces go to Langfuse Cloud. For a fully private setup you would self-host Langfuse and switch speech-to-text to a local Whisper model. The speech-to-text layer is already behind an interface for that.
The seeded patients are fictional. The record in the screenshots is my own test interview.
Watch the demo first
If you only have six minutes, spend them here. The video shows one complete visit, start to finish: a patient answering in Tamil, KuralMD following up the way a doctor would, the OP record filling itself, the doctor checking every field against the patient's own words and approving it, and an old prescription photo being read into the drug history. Then it opens the AI pipeline and the Langfuse trace, so you can see each step that happened along the way.
The whole history took about six minutes. Done by hand, that is a conversation, a translation in your head, and a form typed out field by field, for every patient in the queue. The screenshots below are from the same session, but the video shows how it actually feels to use.
Code
Code: https://github.com/VidhiyaSB/KuralMD
Why Does Open Innovation Matter?
Because this is patient data, and because it is a government-hospital OPD in Tamil Nadu, not a startup with a cloud budget.
- The data can stay in the building. With open weights, the model that hears a patient talk about their symptoms can run on a machine the hospital owns. With a closed API, every patient's history leaves the hospital by default, and "we promise not to train on it" is the best you can get.
- It can run where the internet doesn't. A rural primary health centre with patchy connectivity can still run Gemma on a local box. The speech layer is already behind an interface so it can be swapped for a local Whisper model. That is the part I would make fully offline next.
- I could see why it went wrong. Every bug in "What went wrong" was found by reading what the open pieces actually did: Gemma's raw output, LangExtract's parser, Langfuse's spans. With a closed API, the empty record would have stayed a mystery.
- It can learn Tamil medical speech. Open weights can be fine-tuned on how patients here actually describe their symptoms. With a closed API you take whatever Tamil the vendor ships.
- No per-patient bill. A clinic seeing hundreds of patients a day cannot pay per token forever. Open models make the cost a machine, not a meter.
To be honest about it, KuralMD is not fully open yet. Speech goes through ElevenLabs, and the video used hosted Gemma for speed. But every piece that touches clinical reasoning can run locally today, and the remaining pieces are behind interfaces so they can be replaced.
My Agent Session
I built KuralMD with Claude Code as a pair programmer, over one long weekend session: planning the scope from a one-page brief, writing the FastAPI app and prompts, and then debugging the live Tamil interviews (the timeouts, the repeated question, the silent LangExtract failure and the SQLite locks).
What Dharini said

She felt that this integration to their existing system would drastically reduce the manual effort and would let her examine the patient physically saving a lot of time and She was very sure that this system should not produce diagnostic results or medical suggestions.
Prize Categories
- Best Use of Gemma: Gemma 4 handles translation, the interview agent, extraction (through LangExtract) and prescription vision.
- Best Use of ElevenLabs: Tamil speech-to-text (Scribe) and Tamil text-to-speech for the full voice interview.
What's next
- Test with more real Tamil speech: different accents, a noisy OPD, older patients who speak softly.
- Get the local speech-to-text path good enough that the whole thing can run offline in a rural primary health centre.
Thanks to Dharini for the idea, for checking the questions, and for saying "never let it diagnose" before I had written a single line.







Top comments (0)