Hi, I'm Swayam Verma, a backend developer from Bhopal, India, working with Java and Spring Boot. This post is about Virtual AI Doctor, an AI health assistant I built, and the biggest change I made to it: moving from prompt concatenation to agentic tool calling with the Model Context Protocol (MCP).
Virtual AI Doctor is a learning and portfolio project. It does not replace a real doctor.
What the app does
Users describe their symptoms in Hindi or English and get medical guidance back in the same language. The stack:
- Java 17, Spring Boot 3.5, Spring Security with JWT
-
Spring AI for LLM orchestration (
ChatClient,ChatMemory, tool calling) - Groq (LLaMA 3.1), accessed through Spring AI's OpenAI-compatible client
- MySQL with Spring Data JPA and Hibernate
- Deployed on Render (backend), Railway (database) and Vercel (frontend)
The problem with my first design
My first version was linear. For every message, the backend built one big prompt: instructions, the user's health profile (age, blood group, allergies, medical history), the conversation so far, and the new symptoms. Then it asked the model to answer and parsed the result.
That worked, but it had three problems:
- Every request carried the full patient profile, even when the question didn't need it.
- Prompts grew with every turn.
- Detecting emergencies meant parsing a severity tag out of plain text, which was fragile.
The fix: let the model call tools
Spring AI supports an MCP server/client architecture, built on the official MCP Java SDK. I turned two parts of the app into tools the model can call when it decides it needs them:
-
getPatientProfile: fetches the patient's health profile on demand, instead of injecting it into every request -
flagHighSeverity: called by the model itself when it judges a consultation to be an emergency
The model now decides when it needs the profile, and it decides when something is serious enough to flag. The backend no longer scrapes severity out of text. A tool call is a structured action, not a guess at a string.
When a high-severity flag comes in, the backend sends an email alert through the Brevo API and the user can download a PDF report of the consultation.
Making tools work per user
Tool calls need to know whose session they belong to. I used a request-scoped context (ConsultationContext) so the tools can see the current session without exposing it to the model. Getting this right matters, because a tool that returns the wrong patient's profile would be a serious bug.
Memory across turns
Multi-turn conversations use Spring AI's ChatMemory, backed by Spring AI's JDBC chat-memory repository on MySQL. The assistant keeps the context of the conversation, and past sessions are stored in the consultation history with diagnosis and severity.
Other features
- JWT authentication for signup and login, with secured APIs
- Nearby pharmacy finder using OpenStreetMap and the Overpass API (no API key needed)
- Health profile with age, blood group, allergies and medical history
- Dark and light mode
What I learned
- Tool calling beats prompt stuffing. The model asks for context when it needs it, which keeps prompts smaller and data access explicit.
-
Structured actions beat text parsing. Replacing "find the severity tag in the reply" with a
flagHighSeveritytool made emergency handling cleaner. - Security is part of the design. In a health app, authentication and per-session context come before features.
What's next
- A MEDIUM-severity tool to go with the existing HIGH-severity flag
- A React frontend
- WebSocket-based real-time chat
- Voice input for symptoms
Try it and see the code
- Live demo: https://virtual-ai-doctor-frontend.vercel.app
- Source code: https://github.com/Swayamverma8/Virtual-Ai-Doctor
- My portfolio: https://swayam-portfolio-eta.vercel.app
I'm a final-year CSE student at TIT Bhopal, a Google Student Ambassador (Gemini AI), and open to backend and AI-focused roles. Connect with me on LinkedIn or GitHub.
If you've built something similar with Spring AI or MCP, I'd love to hear what you ran into.


Top comments (0)