My friend understands English.
They can read it.
They can watch English videos.
They know a decent number of English words.
But ask them to speak.
There is a pause.
They first think in Hindi, then translate the sentence into English, then worry whether the grammar is correct, and sometimes the sentence never comes out.
And when they do make a mistake, they often don't know why it was wrong.
I noticed something else too.
The problem wasn't a lack of English content.
There are thousands of videos, grammar lessons, vocabulary apps and courses available.
The problem was practice.
There wasn't someone available every day who would patiently listen, correct the important mistakes, explain them, and then say:
"Okay. Now say it again."
So I decided to build that person as an AI.
That's how TalkSaathi AI started.
What TalkSaathi does
TalkSaathi AI is an English-speaking companion designed for someone who understands English but struggles to speak confidently.
Its core idea is simple:
Listen → Understand → Recall → Speak → Correct → Repeat → Converse → Track → Adapt
Instead of completing a grammar exercise and moving on, TalkSaathi turns a mistake into the next practice opportunity.
For example, imagine my friend says:
"I take breakfast at 8 and I go college."
TalkSaathi doesn't just say "wrong."
You said
I take breakfast at 8 and I go college.
A better version
I have breakfast at 8 and I go to college.
Why?
We usually say "have breakfast" in natural English.
We also use "to" before a destination.
More natural
I usually have breakfast at 8 and then go to college.
And then comes the important part:
🎤 Now say it again.
The learner repeats the sentence.
The AI checks the new attempt.
That creates a much more useful loop:
┌─────────────┐
│ Mistake │
└──────┬──────┘
│
▼
┌─────────────┐
│ Correction │
└──────┬──────┘
│
▼
┌─────────────┐
│ Explanation │
└──────┬──────┘
│
▼
┌─────────────┐
│ Repeat │
└──────┬──────┘
│
▼
┌─────────────┐
│ Re-check │
└──────┬──────┘
│
▼
┌─────────────┐
│ Improvement │
└─────────────┘
│
└──────────────► 🔄 Next Practice
The problem I actually wanted to solve
While building this, I kept coming back to one question:
Why can someone understand English but still struggle to speak it?
For my friend, several things were connected.
They were:
- translating from Hindi before speaking;
- afraid of making grammar mistakes;
- unsure how to form sentences naturally;
- using the same simple vocabulary repeatedly;
- not getting enough speaking practice;
- lacking a conversation partner.
So TalkSaathi isn't designed to make someone memorize thousands of words.
It's designed to give them a place where they can actually speak.
The first practice session
When a learner opens TalkSaathi for the first time, there is no complicated registration.
They create a small profile:
- English level;
- learning goal;
- interests;
- preferred topics;
- daily practice time;
- speaking goal.
They can choose:
10, 15, 20 or 30 minutes.
TalkSaathi then creates a practice session based on their level and previous performance.
A typical session might look like:
┌────────────┐
│ Listen │
└─────┬──────┘
↓
┌────────────┐
│ Understand │
└─────┬──────┘
↓
┌────────────┐
│ Recall │
└─────┬──────┘
↓
┌────────────┐
│ Speak │
└─────┬──────┘
↓
┌────────────┐
│ Correct │
└─────┬──────┘
↓
┌────────────┐
│ Repeat │
└─────┬──────┘
↓
┌────────────┐
│Conversation │
└────────────┘
The important part is that the practice doesn't end after the first answer.
The learner gets another chance.
I didn't want Hindi to become a dependency
One feature I particularly wanted was a Hindi → English speaking bridge.
Because when someone is learning English, sometimes they know exactly what they want to say—but they don't know how to say it in English.
For example:
"मुझे कल कॉलेज जाना है।"
TalkSaathi can turn that into:
I have to go to college tomorrow.
Then it explains the sentence and asks:
🎤 Now say it in English.
After the learner repeats it, TalkSaathi doesn't stop at the translation.
It continues:
What do you usually do at college?
This is important.
Hindi is not supposed to become the final destination.
The goal is to gradually move from:
Hindi thought
↓
English translation
↓
English speaking
towards:
English thought
↓
English speaking
It also remembers mistakes
Another problem with normal practice is that the same mistake can happen again tomorrow.
And again next week.
TalkSaathi keeps a lightweight local history of recurring mistakes.
For example:
Past tense 8 mistakes
Prepositions 6 mistakes
Articles 4 mistakes
Vocabulary 3 mistakes
If the learner repeatedly says:
"I didn't went."
TalkSaathi can recognize that as a recurring pattern.
The next practice can then include:
Past Tense Challenge
So the system isn't only asking:
"What did you say?"
It's also asking:
"What does this learner need to practice next?"
That's where the personalization comes from.
Conversation mode
Sometimes you don't want a lesson.
You just want to talk.
TalkSaathi has a free conversation mode with topics such as:
Daily life
Family, food, college, hobbies, friends.
Professional
Job interviews, meetings, presentations and workplace situations.
Social
Shopping, restaurants, travel and asking directions.
Education & Technology
Programming, AI, college and learning.
Or you can simply enter your own topic.
The AI doesn't try to correct every single word while you're talking.
That would make the conversation feel like an exam.
Instead, it keeps the conversation natural and saves deeper feedback for the end.
A conversation report can show:
Grammar 78%
Fluency 81%
Vocabulary 74%
Important corrections: 4
Focus next:
1. Past tense
2. Articles
3. Prepositions
I wanted the AI to feel different in different moments
TalkSaathi has four personalities, depending on what the learner is doing.
During learning:
Friendly teacher
During conversation:
Natural friend
During correction:
Supportive tutor
During speaking practice:
Encouraging coach
The AI should never make the learner feel embarrassed about making mistakes.
A mistake is not a failure.
It's the next thing to practice.
What happens when you don't know the sentence at all?
I also wanted TalkSaathi to work outside a normal lesson.
So I added a photo-based grammar feature.
A learner can upload a photo containing English text.
For example:
"I am going market yesterday."
TalkSaathi extracts the text and analyzes it.
It returns:
Original
I am going market yesterday.
Correct
I went to the market yesterday.
Why?
"Yesterday" tells us that the action happened in the past, so we use went.
Natural version
I went to the market yesterday.
And then:
🎤 Try saying the corrected sentence.
So even a photo becomes a speaking exercise.
How I built it
The architecture is intentionally split into two important layers:
The open-weight model is the brain.
ElevenLabs is the voice interface.
╭─────────────────────────────────────────────────────────────╮
│ 🗣️ TALKSAATHI AI │
│ Your AI Saathi for English Speaking │
╰─────────────────────────────┬───────────────────────────────╯
│
▼
┌───────────────────┐
│ Next.js Frontend │
└─────────┬─────────┘
│
┌───────────────┼────────────────┐
│ │ │
▼ ▼ ▼
🎤 VOICE INPUT ⌨️ TEXT INPUT 🖼️ IMAGE INPUT
│ │ │
▼ │ ▼
┌─────────────┐ │ ┌─────────────┐
│ ElevenLabs │ │ │ OCR │
│ STT │ │ └──────┬──────┘
└──────┬──────┘ │ │
└───────────────┼────────────────┘
▼
┌───────────────────┐
│ 🧠 GEMMA │
│ Open-weight AI │
│ Brain │
└─────────┬─────────┘
│
┌────────────────┼────────────────┐
│ │ │
▼ ▼ ▼
✨ Correction 💬 Conversation 🎯 Personalization
│ │ │
└────────────────┼────────────────┘
▼
┌───────────────────┐
│ ElevenLabs TTS │
│ AI Voice Output │
└─────────┬─────────┘
│
▼
🔊 SPEAK • PRACTICE
│
▼
👤 THE LEARNER
│
└───────────────┐
│
🔄 Repeat & Improve
│
└───► 🧠 Gemma
The application is built with:
- Next.js
- TypeScript
- Tailwind CSS
- shadcn/ui
- Gemma / compatible open-weight inference
- Hugging Face / compatible inference provider
- ElevenLabs Speech-to-Text
- ElevenLabs Text-to-Speech
- OCR
- LocalStorage
- Vercel
The backend uses Next.js API routes, so API credentials stay on the server.
For the MVP, I deliberately avoided forcing users to create an account.
Learning data such as progress, mistakes, phrases and preferences can be stored locally.
Raw voice recordings and uploaded photos don't need to be permanently stored.
Why an open-weight model?
This was an important choice for me.
I didn't want to build:
Frontend → closed AI API → response
and call that the whole product.
The open-weight model is responsible for the actual language-learning intelligence:
- grammar correction;
- explanations;
- Hindi → English conversion;
- conversation;
- lesson generation;
- mistake analysis;
- personalized exercises;
- phrase suggestions;
- image-text language analysis.
That gives TalkSaathi a separate reasoning layer from its voice layer.
The architecture becomes:
🧠 GEMMA
│
▼
┌───────────┐
│ BRAIN │
└─────┬─────┘
│
│ Understand
│ Correct
│ Explain
│ Personalize
│ Converse
│ Generate
▼
┌─────────────────┐
│ ElevenLabs │
│ VOICE │
└────────┬────────┘
│
▼
🎤 Listen & Speak
This also means the AI layer can evolve independently.
A different open model can be tested later without rebuilding the entire product.
Why not just use Hindi → English translation?
Because translation isn't the goal.
If the app simply translated everything for the learner, the learner could become dependent on translation.
That's why TalkSaathi uses translation as a temporary bridge.
The sequence is:
"I’m stuck"
↓
Hindi
↓
English sentence
↓
Understand why
↓
Say it yourself
↓
Continue in English
The final goal is not better translation.
It's less need for translation.
What building this taught me
One thing became obvious while designing the product:
The hardest part isn't getting an AI model to generate a correction.
The harder question is:
What should happen after the correction?
If the app simply says:
"Correct sentence: ..."
the learning moment is over.
But if it says:
"Now you try."
the mistake becomes practice.
Then another question appears:
What if the learner makes the same mistake tomorrow?
That's why TalkSaathi has mistake memory.
And then:
What should they practice tomorrow?
That's why the system uses the learner's history to recommend the next activity.
The product became much more interesting once I stopped thinking about individual AI responses and started thinking about the learning loop.
The part I'm most excited about
It's not the dashboard.
It's not the number of features.
It's this tiny interaction:
"Now say it again."
Because that's the difference between an AI that answers questions and an AI that actually helps someone practice a skill.
TalkSaathi is designed around that idea.
Handing it to my friend
The real test isn't whether I can make the demo work.
The real test is whether my friend can open it tomorrow, press Start Practice, speak without feeling embarrassed, make mistakes, and want to come back again.
That's what I wanted to build.
Not another English course.
Not another chatbot.
A speaking companion.
What's next?
The current version focuses on the core speaking experience.
There are still things I want to improve:
- more advanced pronunciation analysis;
- better adaptive lesson generation;
- cloud sync;
- user accounts;
- PostgreSQL + Prisma;
- mobile support;
- more languages;
- teacher dashboards;
- community speaking challenges.
But I want to validate the most important thing first:
Does regular speaking practice with TalkSaathi actually make my friend more confident?
Prize Category
Best Use of ElevenLabs
TalkSaathi uses ElevenLabs for:
- Speech-to-Text;
- Text-to-Speech;
- spoken learning content;
- AI conversation voice.
The reasoning layer remains an open-weight model, while ElevenLabs handles the voice interaction.
Try TalkSaathi
Source Code: https://github.com/Rishabh-KSingh/TalkSaathi-AI
Demo Video: https://drive.google.com/file/d/17Pd4YNAB7QNAFNTbSCoJyoxhqDwOKSQ3/view?usp=sharing
I started with a simple observation:
Knowing English and speaking English are two different skills.
So I built something for someone who needed more opportunities to speak.
TalkSaathi AI — Your AI Saathi for English Speaking.
Don't just learn English. Practice speaking it.
Top comments (0)