<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sham Prakash K</title>
    <description>The latest articles on DEV Community by Sham Prakash K (@shamprakash2000).</description>
    <link>https://dev.to/shamprakash2000</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122412%2F8c5d1550-eebf-4bd2-abb7-e1179032f56b.png</url>
      <title>DEV Community: Sham Prakash K</title>
      <link>https://dev.to/shamprakash2000</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shamprakash2000"/>
    <language>en</language>
    <item>
      <title>What Is RAG — And Why Every AI App That Touches Real Data Needs It</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Sat, 03 Oct 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/what-is-rag-and-why-every-ai-app-that-touches-real-data-needs-it-146l</link>
      <guid>https://dev.to/shamprakash2000/what-is-rag-and-why-every-ai-app-that-touches-real-data-needs-it-146l</guid>
      <description>&lt;p&gt;The model doesn't know your data. That's the sentence most AI tutorials skip.&lt;/p&gt;

&lt;p&gt;You ask it about your product catalog — it hallucinates one. You ask it about your internal policy — it describes something it saw during training. You ask it what happened last week — it has no idea. Not because the model is bad. Because your data was never in the training set.&lt;/p&gt;

&lt;p&gt;This is the problem RAG solves. And when the answer genuinely isn't in your data, RAG makes the model say so — instead of making something up. That's the other half of what it fixes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the model doesn't know your data
&lt;/h2&gt;

&lt;p&gt;An LLM learns by reading an enormous amount of public text — web pages, books, code, articles. After training, those patterns are baked into the model's weights. That's how it "knows" things.&lt;/p&gt;

&lt;p&gt;But training has two hard limits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A cutoff date.&lt;/strong&gt; Anything published after the training cutoff doesn't exist for the model. It doesn't know what happened last month. It doesn't know what your team shipped last week.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Only what was in the training data.&lt;/strong&gt; Your company's internal docs were never there. Your product catalog, your customer records, your policies — none of it was public text the model could learn from.&lt;/p&gt;

&lt;p&gt;So when you ask about your specific domain, the model doesn't retrieve from a database. It generates text that fits the pattern based on what it learned during training. For general questions, that's fine. For questions about your data, it guesses — and guesses confidently.&lt;/p&gt;




&lt;h2&gt;
  
  
  The naive fix and why it fails
&lt;/h2&gt;

&lt;p&gt;The first thought most people have: just include everything in the prompt.&lt;/p&gt;

&lt;p&gt;Before the model replies, paste your entire document library into the context. All the product docs, all the policies, all the records. Give it everything — it'll find what it needs.&lt;/p&gt;

&lt;p&gt;This fails for three reasons.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Token limits.&lt;/strong&gt; A 50-page document is roughly 35,000 tokens. Even with a million-token context window, you can't dump an entire document library into every request. Real systems have thousands of documents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The "lost in the middle" problem.&lt;/strong&gt; We covered this in article 03. When you send a very long prompt, the model pays less attention to content buried in the middle. Relevant information sitting in paragraph 47 gets deprioritised. More context is not always better — it can actively hurt response quality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost.&lt;/strong&gt; If your context is 50,000 tokens and you're paying per million tokens, that's real money per API call — multiplied across every request from every user.&lt;/p&gt;

&lt;p&gt;Dumping everything in the prompt doesn't scale. You need something smarter.&lt;/p&gt;




&lt;h2&gt;
  
  
  "Why not just fine-tune the model on my data?"
&lt;/h2&gt;

&lt;p&gt;This is the first question most engineers ask. It sounds cleaner — train the model once on your data, and it knows everything.&lt;/p&gt;

&lt;p&gt;Fine-tuning doesn't solve this problem. Here's why.&lt;/p&gt;

&lt;p&gt;Fine-tuning teaches a model a &lt;em&gt;style&lt;/em&gt; or a &lt;em&gt;behaviour&lt;/em&gt;, not facts. You can fine-tune a model to always respond in JSON, or to follow a specific tone, or to focus on a domain. But it doesn't reliably store factual knowledge in a way you can trust. Fine-tuned models still hallucinate — they just hallucinate in your style.&lt;/p&gt;

&lt;p&gt;More practically: your data changes. Products get updated, policies change, new documents are added. Every time your data changes, you'd need to re-run fine-tuning — a process that takes hours and costs significant compute. RAG retrieves from your current data on every query. Add a document today, it's searchable immediately.&lt;/p&gt;

&lt;p&gt;Fine-tuning and RAG solve different problems. Fine-tuning for behaviour and style. RAG for grounding the model in specific, up-to-date facts. They're not alternatives — production systems often use both.&lt;/p&gt;




&lt;h2&gt;
  
  
  What RAG is
&lt;/h2&gt;

&lt;p&gt;RAG stands for Retrieval Augmented Generation.&lt;/p&gt;

&lt;p&gt;The word order matters. &lt;strong&gt;Retrieval comes first. Generation comes second.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of sending everything to the model and hoping it finds the relevant piece — you find the relevant piece first, then send only that.&lt;/p&gt;

&lt;p&gt;Here's the idea with a concrete example.&lt;/p&gt;

&lt;p&gt;You have 500 pages of product documentation. A user asks: "What's the return policy for electronics?"&lt;/p&gt;

&lt;p&gt;Without RAG: send all 500 pages to the model, hope it finds the answer somewhere.&lt;/p&gt;

&lt;p&gt;With RAG:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Search your 500 pages for the sections most relevant to "return policy for electronics"&lt;/li&gt;
&lt;li&gt;Find 3–4 paragraphs that actually contain that information&lt;/li&gt;
&lt;li&gt;Send only those paragraphs to the model, along with the question&lt;/li&gt;
&lt;li&gt;The model reads what you gave it and answers from that&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The model never searches anything. It just reads what you put in front of it. You are the one doing the retrieval — and then giving the model exactly what it needs.&lt;/p&gt;

&lt;p&gt;That's RAG. Retrieve the relevant pieces. Generate the answer from them.&lt;/p&gt;




&lt;h2&gt;
  
  
  Two phases: ingest and query
&lt;/h2&gt;

&lt;p&gt;Every RAG system has two distinct phases. Understanding this split is the key to understanding how it works.&lt;/p&gt;




&lt;h3&gt;
  
  
  Phase 1 — Ingest (once per document)
&lt;/h3&gt;

&lt;p&gt;Before any user can ask questions, you prepare your data.&lt;/p&gt;

&lt;p&gt;You take your documents — PDFs, text files, markdown, whatever — and process them:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Split&lt;/strong&gt; each document into small chunks (a few hundred words each)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embed&lt;/strong&gt; each chunk — convert it into a list of numbers called a vector that captures its meaning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store&lt;/strong&gt; those vectors in a database built for vector search&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is a one-time operation per document. Once ingested, a document is ready to be retrieved.&lt;/p&gt;




&lt;h3&gt;
  
  
  Phase 2 — Query (every user request)
&lt;/h3&gt;

&lt;p&gt;When a user asks a question:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Embed&lt;/strong&gt; the question — convert it into a vector using the &lt;strong&gt;same embedding model used during ingest&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Search&lt;/strong&gt; the vector database for stored chunks that are semantically similar to the question vector&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieve&lt;/strong&gt; the top 3–5 most relevant chunks (more on why this range in a moment)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inject&lt;/strong&gt; those chunks into the prompt alongside the question&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generate&lt;/strong&gt; — the model reads the context and answers&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step by step, every time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why 3–5 chunks?&lt;/strong&gt; Retrieve just 1 and you might miss the answer — the single most similar chunk may not contain everything the question needs, especially if the answer spans related sections. Retrieve 20 and you've recreated the "lost in the middle" problem — relevant details buried in a wall of context. 3–5 is the sweet spot: enough coverage to find the answer, focused enough that the model can use it. You'll tune this number based on your chunk size and your data.&lt;/p&gt;

&lt;p&gt;One constraint matters here: the embedding model used at query time must be the same one used during ingest. Every embedding model has its own internal "map" of meaning. A vector produced by one model lives in a completely different space than a vector from another model — similarity scores between them are meaningless. If you embed your documents with Gemini's embedding model and then embed questions with OpenAI's, your similarity search will return garbage. Same model, both ways, always.&lt;/p&gt;




&lt;h2&gt;
  
  
  How the prompt changes
&lt;/h2&gt;

&lt;p&gt;Without RAG, a prompt looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;System: You are a helpful assistant.
User: What's the return policy for electronics?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model guesses from its training data.&lt;/p&gt;

&lt;p&gt;With RAG, the prompt looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;System: You are a helpful assistant. Answer ONLY based on the provided context.
        If the context lacks enough information, say so clearly.

Context:
---
[Chunk 1: "Electronics purchased at full price may be returned within 30 days
with original packaging and receipt. Items must be in original condition..."]

[Chunk 2: "Exceptions to the standard return policy: laptops and tablets
cannot be returned after the seal is broken unless defective..."]

[Chunk 3: "To initiate a return, contact customer support with your order number.
Refunds are processed within 5–7 business days..."]
---

User: What's the return policy for electronics?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the model isn't guessing. It's reading actual policy text and summarising it for the user.&lt;/p&gt;

&lt;p&gt;Notice the system prompt instruction: &lt;code&gt;Answer ONLY based on the provided context&lt;/code&gt;. This is critical. Without it, the model might supplement your retrieved context with its training knowledge — mixing real data with guesses. That instruction keeps it grounded in what you provided.&lt;/p&gt;

&lt;p&gt;And when the answer isn't in the retrieved chunks? The model says: "I don't have enough information to answer this." That's not a failure — that's the system working correctly. A confident wrong answer is far more dangerous than an honest "I don't know." RAG gives you the second option.&lt;/p&gt;




&lt;h2&gt;
  
  
  RAG is not a search engine
&lt;/h2&gt;

&lt;p&gt;A keyword search finds documents that contain your words. RAG finds documents that are semantically related to your question — even when they share no keywords.&lt;/p&gt;

&lt;p&gt;Search for "fix broken leg" — a keyword search returns results with those exact words. A RAG embedding model understands that "broken leg," "bone fracture," and "orthopedic injury" are close in meaning — and ranks results by semantic similarity, not by word overlap.&lt;/p&gt;

&lt;p&gt;This is what makes retrieval in RAG powerful. It understands meaning. How it does that — through embeddings — is what the next article covers in detail.&lt;/p&gt;




&lt;h2&gt;
  
  
  Two types of questions RAG answers
&lt;/h2&gt;

&lt;p&gt;Before building, be clear which problem you're solving. RAG works for two different use cases that look similar from the outside.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Document questions&lt;/strong&gt; — "what does the policy say about X?"&lt;br&gt;
You're searching for information inside documents. Product specs, how-to guides, FAQs, policies, uploaded notes. Semantic retrieval finds the relevant section.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Record questions&lt;/strong&gt; — "what happened last week?"&lt;br&gt;
You're looking up structured data. Order history, user activity, database records. Semantic search isn't the right tool here — structured queries are. This is where a database agent (covered later in the series) takes over.&lt;/p&gt;

&lt;p&gt;RAG handles documents. A SQL agent handles records. A complete AI backend needs both — but they're separate problems.&lt;/p&gt;


&lt;h2&gt;
  
  
  The architecture after adding RAG
&lt;/h2&gt;

&lt;p&gt;Before RAG, the architecture is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User → API → LLM
         ↓
   PostgreSQL (chat history)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After adding RAG, two new components appear:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User → API → LLM
         ↓        ↑
   PostgreSQL   Context (retrieved chunks)
                    ↑
              Vector DB ← Embedding Model
                    ↑
              Ingest pipeline (split → embed → store)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Embedding model&lt;/strong&gt; — converts text to vectors. For Gemini users, Google provides an embedding API. Same provider, no extra dependency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vector database&lt;/strong&gt; — stores and searches vectors by similarity. Pinecone is the most common choice. Free tier covers everything you need to learn.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ingest pipeline&lt;/strong&gt; — when a document is uploaded, it runs through: split → embed → store. One operation, runs once per document.&lt;/p&gt;

&lt;p&gt;On every user query, the API: embeds the question → searches the vector DB → retrieves the top chunks → injects them into the prompt → calls the LLM.&lt;/p&gt;




&lt;h2&gt;
  
  
  A mistake worth knowing before you start
&lt;/h2&gt;

&lt;p&gt;When I first built the ingest pipeline, I chunked documents into 2,000-word pieces — roughly half a page of text.&lt;/p&gt;

&lt;p&gt;Retrieval looked fine. Top chunks came back, all semantically relevant. But responses were vague. The model had the right section but couldn't extract specific answers cleanly.&lt;/p&gt;

&lt;p&gt;The problem: each chunk was too large. A 2,000-word chunk gets embedded as one vector representing the average meaning of the whole section — not any specific fact inside it. When retrieved, you get a chunk that contains the answer somewhere inside 400 other words. The model has to work harder and often glosses over the detail.&lt;/p&gt;

&lt;p&gt;Dropping to 300-word chunks fixed it. Retrieval precision jumped. Specific questions started getting specific answers.&lt;/p&gt;

&lt;p&gt;Chunk size is one of the biggest practical levers in RAG quality. It's easy to get wrong the first time — and it fails silently. We'll go deep on exactly why in upcoming articles.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;RAG needs two things: a way to measure semantic similarity so you can find relevant chunks, and a database that can search by that similarity. Both depend on embeddings.&lt;/p&gt;

&lt;p&gt;What is an embedding? Why does converting text into a list of numbers let you search by meaning? Why does "cricket" end up close to "sports" in vector space even though they share no characters?&lt;/p&gt;

&lt;p&gt;That's what the next article is about.&lt;/p&gt;




&lt;p&gt;Building something where the model needs to answer from your own data? Drop in the comments what the use case is — I'm curious what people are solving.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>java</category>
      <category>rag</category>
      <category>llm</category>
    </item>
    <item>
      <title>The System Prompt Is Not a Description — It's a Contract</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/the-system-prompt-is-not-a-description-its-a-contract-fp2</link>
      <guid>https://dev.to/shamprakash2000/the-system-prompt-is-not-a-description-its-a-contract-fp2</guid>
      <description>&lt;p&gt;My first system prompt was this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a helpful and friendly AI assistant. You are knowledgeable about many topics 
and always try to give thorough, well-explained answers. You are patient and 
understanding. You never refuse to help with reasonable questions. You always 
maintain a positive and encouraging tone in all your responses.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It did nothing. The model behaved exactly the same with it as without it. I'd written five sentences that described what an LLM already does by default.&lt;/p&gt;

&lt;p&gt;This article is about writing system prompts that actually change how the model behaves — with real examples from the app we've been building.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a system prompt actually is
&lt;/h2&gt;

&lt;p&gt;When you call the Gemini API, your request has three parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;System prompt&lt;/strong&gt; — instructions the model reads before anything else&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversation history&lt;/strong&gt; — the previous messages&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User message&lt;/strong&gt; — what the user just typed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The system prompt is processed first, before any user input. The model uses it to understand its role, constraints, and expected behaviour for the entire conversation.&lt;/p&gt;

&lt;p&gt;Think of it as the briefing you give an employee before their first day. The more specific and well-structured the briefing, the less they have to guess.&lt;/p&gt;




&lt;h2&gt;
  
  
  What breaks a system prompt
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Too vague&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a helpful assistant.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This tells the model nothing it doesn't already know. It'll just do what it would do anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Describing what the model IS instead of what it should DO&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are an expert travel planner who knows everything about destinations worldwide.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Describing the model's identity doesn't constrain its behaviour. The model can still say anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Too long with no structure&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A wall of text is hard for the model to parse. Instructions buried in paragraph 4 often get ignored or deprioritised. Important rules need to stand out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contradictory instructions&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Be concise. Always give thorough, detailed answers with examples.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When instructions conflict, the model picks one — and you won't know which.&lt;/p&gt;




&lt;h2&gt;
  
  
  What actually works
&lt;/h2&gt;

&lt;p&gt;Good system prompts have clear building blocks. Not every prompt needs all of them, but knowing what each one does helps you write deliberately.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Persona — who the model is in this context&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not a description of the model's general nature, but a specific role with scope:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a database assistant. You help users query business data 
using natural language.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Persona alone isn't enough, but it sets the frame for everything that follows.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Constraints — what the model will not do&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is where system prompts earn their keep. Hard limits that override the model's defaults:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Only write SELECT queries — no INSERT, UPDATE, DELETE, DROP, or DDL.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Answer ONLY based on the provided context. If the context lacks enough 
information, say so clearly.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;If the user asks about any destination outside India, politely decline.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Constraints are the most powerful thing in a system prompt. They make the model predictable. Without them, the model will try to be helpful in ways you didn't intend.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Output format — how to structure the response&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ALWAYS return query results as a full markdown table showing ALL rows 
and ALL columns — do NOT summarize, abbreviate, or omit rows.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Answer only YES or NO, nothing else.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output format instructions are surprisingly effective. When you tell the model exactly what shape the response should be, it follows it consistently.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Few-shot examples — show the model exactly what you want&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instructions tell the model what to do. Examples show it. When format consistency matters, one or two examples in the system prompt are more reliable than a paragraph of instructions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a helpful assistant for a Java backend engineer learning AI development.
Be concise and practical.

Example:
User: How do I check if a string is empty in Java?
Assistant: Use `str == null || str.isEmpty()`. Prefer `str.isBlank()` in Java 11+ to also catch whitespace-only strings.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model picks up the response style — length, tone, format — from the example and applies it consistently. One example is usually enough. Two if the format is unusual or strict.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sequence — the order things should happen&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For tool-calling agents, the order of operations matters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;For DATABASE questions, ALWAYS follow this sequence:
1. Call listTables to see what tables are available
2. Call getTableSchema for every table you need — never guess column names
3. Write a safe SELECT query and call executeQuery
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without this, the model might try to write a query without checking the schema first — and guess column names that don't exist.&lt;/p&gt;




&lt;h2&gt;
  
  
  Real examples from this app
&lt;/h2&gt;

&lt;p&gt;Here's how these principles look in practice, using the actual system prompts from the app we've been building.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;The simplest case — lean and specific&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="no"&gt;CHAT_AI_SYSTEM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
    &lt;span class="s"&gt;"You are a helpful assistant for a Java backend engineer learning AI development. "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"Be concise and practical."&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two sentences. But they're specific. The model knows the audience (Java backend engineers learning AI) and the output style (concise and practical). It won't give long theoretical explanations when a code snippet would do.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Single-purpose validator — extremely tight&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="no"&gt;INDIA_VALIDATION_SYSTEM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
    &lt;span class="s"&gt;"You are a geography validator. Answer only YES or NO, nothing else."&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prompt is used to check whether a destination is in India before the travel agent runs. The output constraint (&lt;code&gt;only YES or NO, nothing else&lt;/code&gt;) is absolute. No explanation, no uncertainty — just the answer the code needs to branch on.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;RAG assistant — context-bound&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="no"&gt;RAG_SYSTEM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
    &lt;span class="s"&gt;"You are a helpful assistant. Answer ONLY based on the provided context. "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"If the context lacks enough information, say so clearly."&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key instruction is &lt;code&gt;ONLY based on the provided context&lt;/code&gt;. Without this, the model uses its general knowledge to fill in gaps — which means it might answer with information that isn't in your documents. For a RAG system, that's a hallucination problem.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Complex agent — structured with rules&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="no"&gt;DATABASE_AGENT_SYSTEM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
    &lt;span class="s"&gt;"You are a database and knowledge assistant. "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"You have five tools: listTables, getTableSchema, executeQuery, askDocuments, ingestDocument.\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"For DATABASE questions, ALWAYS follow this sequence:\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"1. Call listTables to see what tables are available\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"2. Call getTableSchema for every table you need — never guess column names\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"3. Write a safe SELECT query and call executeQuery\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"DATABASE RULES:\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"- Only write SELECT queries — no INSERT, UPDATE, DELETE, DROP, or DDL\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"- Always check the schema before querying — column names must come from getTableSchema, not guesses\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"- ALWAYS return query results as a full markdown table showing ALL rows and ALL columns\n"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;"- If the user asks something the data cannot answer, say so honestly"&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prompt is longer — but it's structured. The sequence is numbered. The rules are bulleted. Important words are capitalised (&lt;code&gt;ALWAYS&lt;/code&gt;, &lt;code&gt;ONLY&lt;/code&gt;). The model can scan this prompt and find what applies to the current situation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Dynamic system prompts
&lt;/h2&gt;

&lt;p&gt;System prompts don't have to be static strings. In Spring AI you can build the system prompt at call time and inject runtime context — current date, user preferences, session data — before sending the request.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/chat-ai"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;userId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"userId"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;systemPrompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;format&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;"You are a helpful assistant for a Java backend engineer learning AI development. "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;"Be concise and practical. Today's date is %s."&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;
        &lt;span class="nc"&gt;LocalDate&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;now&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;system&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;systemPrompt&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;// ← override the default system prompt per request&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"message"&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;advisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;param&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"chat_memory_conversation_id"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;)))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.system(systemPrompt)&lt;/code&gt; on the prompt overrides the &lt;code&gt;defaultSystem(...)&lt;/code&gt; you set in the &lt;code&gt;ChatClient&lt;/code&gt; builder for that one call. The builder default is still used for everything else.&lt;/p&gt;

&lt;p&gt;This is also how production applications inject per-user context: user tier, preferences, retrieved memory facts, or feature flags — all assembled into the system prompt at request time rather than stored as a static constant.&lt;/p&gt;

&lt;p&gt;Keep the static parts in &lt;code&gt;Prompts.java&lt;/code&gt; as constants and string-format in the runtime values. That way the static structure stays readable and the dynamic parts are clearly marked.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to test if your prompt is working
&lt;/h2&gt;

&lt;p&gt;Write test cases the same way you'd write unit tests. For each instruction in your prompt, write a user message that should trigger it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Instruction&lt;/th&gt;
&lt;th&gt;Test input&lt;/th&gt;
&lt;th&gt;Expected behaviour&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;India-only&lt;/td&gt;
&lt;td&gt;"Plan a trip to Paris"&lt;/td&gt;
&lt;td&gt;Politely declines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SELECT only&lt;/td&gt;
&lt;td&gt;"Delete all orders"&lt;/td&gt;
&lt;td&gt;Refuses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;YES or NO only&lt;/td&gt;
&lt;td&gt;"Is Mumbai in India?"&lt;/td&gt;
&lt;td&gt;Returns &lt;code&gt;YES&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concise&lt;/td&gt;
&lt;td&gt;"What is Spring Boot?"&lt;/td&gt;
&lt;td&gt;Short answer, no long essay&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If the model doesn't behave as expected on your test inputs, the instruction is either too vague, missing, or buried where the model deprioritises it. Move critical rules earlier in the prompt and make them more explicit.&lt;/p&gt;




&lt;h2&gt;
  
  
  Practical rules I follow
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Lead with constraints, not with personality.&lt;/strong&gt; The model's personality is fine by default. Your constraints are what matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use ALL CAPS for non-negotiable rules.&lt;/strong&gt; &lt;code&gt;ALWAYS&lt;/code&gt;, &lt;code&gt;NEVER&lt;/code&gt;, &lt;code&gt;ONLY&lt;/code&gt;. It works.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Number sequences.&lt;/strong&gt; If order matters, number the steps. Bullet points suggest optional items. Numbered lists suggest mandatory sequence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Be specific about output format.&lt;/strong&gt; Don't say "give a structured answer." Say "return a markdown table with all rows and columns."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep it under 300 tokens where possible.&lt;/strong&gt; Every token in your system prompt runs on every API call. A bloated system prompt costs money and attention.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;System prompts shape how the model thinks. Next: structured output — making the model return JSON instead of prose, so your backend can parse and act on the response programmatically.&lt;/p&gt;




&lt;p&gt;Have a system prompt that finally started working after you changed something specific? Drop it in the comments.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>java</category>
      <category>springboot</category>
      <category>ai</category>
      <category>beginners</category>
    </item>
    <item>
      <title>I Hit a 429 Before I Got a Bill — Managing Gemini API Limits in Spring Boot</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/i-hit-a-429-before-i-got-a-bill-managing-gemini-api-limits-in-spring-boot-5658</link>
      <guid>https://dev.to/shamprakash2000/i-hit-a-429-before-i-got-a-bill-managing-gemini-api-limits-in-spring-boot-5658</guid>
      <description>&lt;p&gt;I didn't get a surprise bill. I got a 429.&lt;/p&gt;

&lt;p&gt;Requests started failing. The app was returning errors instead of responses. I opened the Gemini AI Studio rate limit page and saw it — quota exhausted. I'd burned through my daily request limit without realising it.&lt;/p&gt;

&lt;p&gt;That's when I understood that token usage isn't abstract. It's a real counter ticking up with every API call, and if you're not watching it, you'll hit the wall.&lt;/p&gt;

&lt;p&gt;This article is about understanding what you're spending, where it goes, and how to stay within limits — especially on the free tier.&lt;/p&gt;

&lt;h2&gt;
  
  
  What rate limits actually are
&lt;/h2&gt;

&lt;p&gt;When you use the Gemini API on the free tier, you're not billed — but you're not unlimited either. Google enforces three types of limits:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RPM — Requests Per Minute.&lt;/strong&gt; How many API calls you can make in a 60-second window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TPM — Tokens Per Minute.&lt;/strong&gt; How many tokens (input + output combined) you can send and receive per minute.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RPD — Requests Per Day.&lt;/strong&gt; Total API calls allowed in a 24-hour window.&lt;/p&gt;

&lt;p&gt;When you exceed any of these, the API returns a &lt;code&gt;429 Too Many Requests&lt;/code&gt; error. Your app stops working until the window resets.&lt;/p&gt;

&lt;p&gt;The free tier limits are lower than you'd expect — and they go fast when you're testing. Check the current limits for your model on the Gemini API rate limits page in AI Studio since they change as Google updates the free tier.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why the limits go fast
&lt;/h2&gt;

&lt;p&gt;When I first hit the limit I thought I hadn't made that many calls. I was wrong — I just hadn't understood what each call actually costs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem 1: Every call sends the full history.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In Article 6 we built conversation memory. Every time you send a message, Spring AI includes the last 20 messages in the request. That means message 20 doesn't cost 1 message worth of tokens — it costs up to 20 messages worth.&lt;/p&gt;

&lt;p&gt;A conversation with 10 back-and-forth exchanges might look like 10 API calls. But the token cost is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Call 1: 1 message&lt;/li&gt;
&lt;li&gt;Call 2: 3 messages (1+2)&lt;/li&gt;
&lt;li&gt;Call 3: 6 messages (1+2+3)&lt;/li&gt;
&lt;li&gt;...&lt;/li&gt;
&lt;li&gt;Call 10: 55 messages worth of tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Problem 2: Your system prompt runs on every call.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every request includes the system prompt. If your system prompt is 200 words, that's ~270 tokens added to every single API call — silently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem 3: Output tokens cost too.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The model's reply counts against your token budget. A long detailed answer costs more than a short one. By default, the model will be as verbose as it wants to be.&lt;/p&gt;




&lt;h2&gt;
  
  
  See exactly what you're spending
&lt;/h2&gt;

&lt;p&gt;The fix starts with visibility. Spring AI exposes token usage on every response — you just have to ask for it.&lt;/p&gt;

&lt;p&gt;Change your chat method to use &lt;code&gt;chatResponse()&lt;/code&gt; instead of &lt;code&gt;content()&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/chat-ai"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nc"&gt;ChatResponse&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"message"&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;advisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;param&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"chat_memory_conversation_id"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;)))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatResponse&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;   &lt;span class="c1"&gt;// ← get the full response object, not just the text&lt;/span&gt;

    &lt;span class="c1"&gt;// Log token usage&lt;/span&gt;
    &lt;span class="nc"&gt;Usage&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getMetadata&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getUsage&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;info&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Tokens — input: {}, output: {}, total: {}"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getPromptTokens&lt;/span&gt;&lt;span class="o"&gt;(),&lt;/span&gt;
        &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getGenerationTokens&lt;/span&gt;&lt;span class="o"&gt;(),&lt;/span&gt;
        &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getTotalTokens&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getResult&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getOutput&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getText&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now every API call logs something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Tokens — input: 847, output: 312, total: 1159
Tokens — input: 1203, output: 198, total: 1401
Tokens — input: 1589, output: 445, total: 2034
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Watch the input token count grow with each message in a conversation. That's the history accumulating.&lt;/p&gt;




&lt;h2&gt;
  
  
  Track a running total per session
&lt;/h2&gt;

&lt;p&gt;Once you can see individual call usage, it's useful to track cumulative usage per session. A simple in-memory counter works fine for this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Integer&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;sessionTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ConcurrentHashMap&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;();&lt;/span&gt;

&lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/chat-ai"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;conversationId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

    &lt;span class="nc"&gt;ChatResponse&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"message"&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;advisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;param&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"chat_memory_conversation_id"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatResponse&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="nc"&gt;Usage&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getMetadata&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getUsage&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;callTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getTotalTokens&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;runningTotal&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sessionTokens&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;merge&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;callTokens&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nl"&gt;Integer:&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="n"&gt;sum&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

    &lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;info&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId: {} | this call: {} tokens | session total: {} tokens"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;callTokens&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;runningTotal&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getResult&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getOutput&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getText&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After a few messages you'll see the session total climbing. This makes the accumulation visible and gives you something to act on.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to reduce token usage
&lt;/h2&gt;

&lt;p&gt;Now that you can see what you're spending, here's where to trim.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Keep your system prompt short.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every extra word in your system prompt costs tokens on every single call. Go through it and cut anything that isn't load-bearing.&lt;/p&gt;

&lt;p&gt;Before:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a helpful and friendly AI assistant. You are knowledgeable about many topics 
and always try to give thorough, well-explained answers. You are patient and 
understanding. You never refuse to help with reasonable questions. You always 
maintain a positive and encouraging tone in all your responses.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a helpful assistant. Be concise and practical.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same behaviour, a fraction of the tokens — multiplied across every API call you make.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Limit conversation history.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We already set &lt;code&gt;maxMessages(20)&lt;/code&gt; in Article 6. That's a good default. For development and testing, drop it lower:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nc"&gt;MessageWindowChatMemory&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MessageWindowChatMemory&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatMemoryRepository&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryRepository&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;maxMessages&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;// ← 10 messages instead of 20 while testing&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Older messages in a long conversation rarely affect the current reply anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Tell the model to be concise.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The model's default verbosity is tunable via the system prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a helpful assistant. Be concise — answer in 2-3 sentences unless the user 
asks for more detail.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A shorter reply means fewer output tokens. Output tokens count against the same limits as input tokens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Don't test with long conversations.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;During development, start a new session for every test instead of continuing an existing conversation. A fresh session has zero history — each call costs only the system prompt + one message, not the accumulated history of 20 exchanges.&lt;/p&gt;




&lt;h2&gt;
  
  
  Beyond maxMessages — how production apps handle long conversations
&lt;/h2&gt;

&lt;p&gt;You've been using &lt;code&gt;maxMessages(20)&lt;/code&gt; since Article 6. That's the simplest strategy — and it has a name: &lt;strong&gt;sliding window&lt;/strong&gt;. But it's not the only approach. Here's how the strategies stack up as conversations get longer and requirements get harder.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sliding window&lt;/strong&gt; — what you already have&lt;/p&gt;

&lt;p&gt;Keep only the last N messages. Older messages are dropped. Simple, predictable, zero extra API calls.&lt;/p&gt;

&lt;p&gt;Works well for most chat apps where conversations are short-to-medium. The downside: the model loses context from early in the conversation. If the user mentioned their name in message 1 and you're now on message 25, the model has forgotten it.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Summarization&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When history hits a threshold, ask the model to summarize the older messages into a compact paragraph. Replace those messages with the summary. Then continue with summary + recent messages.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;size&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// Ask the model to summarize older messages&lt;/span&gt;
    &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;oldMessages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;formatMessages&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;subList&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;

    &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Summarize this conversation in 4-5 sentences, keeping key facts and decisions: "&lt;/span&gt;
              &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;oldMessages&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="c1"&gt;// Replace 10 old messages with one summary message&lt;/span&gt;
    &lt;span class="nc"&gt;List&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Message&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;compressed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ArrayList&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;();&lt;/span&gt;
    &lt;span class="n"&gt;compressed&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;add&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;SystemMessage&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Earlier conversation summary: "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;
    &lt;span class="n"&gt;compressed&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;addAll&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;subList&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;size&lt;/span&gt;&lt;span class="o"&gt;()));&lt;/span&gt;
    &lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;compressed&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;2000 tokens of raw history becomes 150 tokens of summary. The model keeps the gist without the full detail. This is what Claude Code does — the summary you see at the start of a long session is exactly this pattern applied automatically.&lt;/p&gt;

&lt;p&gt;The tradeoff: one extra API call to summarize, and fine-grained detail from early messages is lost. For most apps that's acceptable.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;RAG-based memory&lt;/strong&gt; &lt;em&gt;(coming later in this series)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Instead of sending recent messages, store every message as an embedding in a vector database. When a new message arrives, search for the most &lt;em&gt;relevant&lt;/em&gt; past messages — not just the most recent ones.&lt;/p&gt;

&lt;p&gt;This means a conversation from last week about a specific topic gets retrieved when relevant, even if it's buried under hundreds of other messages. We'll build this when we cover RAG.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Structured memory extraction&lt;/strong&gt; &lt;em&gt;(coming later in this series)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Instead of storing raw messages, the system extracts facts from the conversation and stores them separately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"user_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Sham"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"project"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Spring Boot AI chat app"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"decisions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Gemini API"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Neon PostgreSQL"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"deployed on Render"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model receives this fact store + recent messages — not raw history at all. It's how products like ChatGPT's memory feature work. Much more token-efficient for long-running applications.&lt;/p&gt;




&lt;p&gt;For now, sliding window (&lt;code&gt;maxMessages(20)&lt;/code&gt;) handles most cases. Add summarization when your conversations consistently hit the limit. RAG and structured memory come later in the series when we have the foundations for them.&lt;/p&gt;




&lt;h2&gt;
  
  
  When to upgrade
&lt;/h2&gt;

&lt;p&gt;The free tier is enough to learn and build demos. You'll hit limits when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're testing heavily (many short test sessions in the same day)&lt;/li&gt;
&lt;li&gt;Your conversations get long (history accumulates fast)&lt;/li&gt;
&lt;li&gt;Multiple people are using the app simultaneously&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When you do upgrade to a paid tier, the rate limits increase significantly and you're charged per million tokens instead. At that point the logging you built here becomes essential — it's how you know what you're actually paying for.&lt;/p&gt;

&lt;p&gt;For now, the key habits are: log every call, watch the input token growth, keep your system prompt lean, and use fresh sessions during testing.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The app is efficient now. Next: writing system prompts that actually shape the model's behaviour — not just what it says, but how it thinks about your problem.&lt;/p&gt;




&lt;p&gt;Hit a 429 while building? Drop it in the comments — you're definitely not the only one.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>java</category>
      <category>ai</category>
      <category>springboot</category>
      <category>beginners</category>
    </item>
    <item>
      <title>How I Made My AI App Stream Like ChatGPT — Server-Sent Events in Spring Boot</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Sun, 27 Sep 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/how-i-made-my-ai-app-stream-like-chatgpt-server-sent-events-in-spring-boot-593m</link>
      <guid>https://dev.to/shamprakash2000/how-i-made-my-ai-app-stream-like-chatgpt-server-sent-events-in-spring-boot-593m</guid>
      <description>&lt;p&gt;The app is live. Users can chat with the AI and it remembers the conversation.&lt;/p&gt;

&lt;p&gt;But there's one problem I noticed immediately after deploying: a user sends a message, and then they stare at a completely blank screen for three to five seconds. Then — bam — the full reply appears all at once.&lt;/p&gt;

&lt;p&gt;Compare that to ChatGPT or Gemini's own interface. You type a question and the response starts flowing immediately, word by word. It feels alive.&lt;/p&gt;

&lt;p&gt;That difference is streaming. This article is about adding it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the delay exists
&lt;/h2&gt;

&lt;p&gt;When you call &lt;code&gt;.call().content()&lt;/code&gt; in Spring AI, here's what actually happens:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Your message goes to the Gemini API&lt;/li&gt;
&lt;li&gt;Gemini starts generating the response — one token at a time&lt;/li&gt;
&lt;li&gt;Gemini keeps generating until the response is complete&lt;/li&gt;
&lt;li&gt;Gemini sends the &lt;strong&gt;entire response&lt;/strong&gt; back to your server in one HTTP response&lt;/li&gt;
&lt;li&gt;Your server returns it to the client&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The model isn't slow — it's generating tokens quickly. But you're waiting for the last token before you get the first one. All that generation time appears as a blank screen.&lt;/p&gt;

&lt;p&gt;Streaming flips this. Instead of waiting for everything, you receive each token the moment the model generates it — and send it to the client immediately.&lt;/p&gt;




&lt;h2&gt;
  
  
  What SSE is
&lt;/h2&gt;

&lt;p&gt;Server-Sent Events (SSE) is a simple HTTP mechanism for the server to push data to the client over a single connection that stays open.&lt;/p&gt;

&lt;p&gt;Normal HTTP: client sends request → server sends response → connection closes.&lt;/p&gt;

&lt;p&gt;SSE: client sends request → server keeps the connection open and sends data in chunks as it becomes available → connection closes when the stream ends.&lt;/p&gt;

&lt;p&gt;It's one-directional (server → client only), which is exactly what we need for streaming AI responses. The client asks once, and the server streams the answer back token by token.&lt;/p&gt;




&lt;h2&gt;
  
  
  What changes in the code
&lt;/h2&gt;

&lt;p&gt;Right now the chat endpoint looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/chat-ai"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;ResponseEntity&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"message"&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;advisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;param&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"chat_memory_conversation_id"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;)))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;       &lt;span class="c1"&gt;// ← waits for the full response&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;ResponseEntity&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"answer"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The streaming version changes two things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;.call()&lt;/code&gt; becomes &lt;code&gt;.stream()&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The return type becomes &lt;code&gt;Flux&amp;lt;String&amp;gt;&lt;/code&gt; instead of &lt;code&gt;ResponseEntity&amp;lt;String&amp;gt;&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"/chat-ai/stream"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;produces&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MediaType&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;TEXT_EVENT_STREAM_VALUE&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;Flux&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;streamChat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"message"&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;advisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;param&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"chat_memory_conversation_id"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;)))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;     &lt;span class="c1"&gt;// ← returns a Flux, not a String&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. The same &lt;code&gt;chatClient&lt;/code&gt;, the same memory advisor, the same conversation ID — just &lt;code&gt;.stream()&lt;/code&gt; instead of &lt;code&gt;.call()&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  What is Flux?
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;Flux&amp;lt;String&amp;gt;&lt;/code&gt; is from Project Reactor — the reactive library that Spring WebFlux is built on. Think of it as a sequence of values that arrive over time, not all at once.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;.call().content()&lt;/code&gt; gives you one &lt;code&gt;String&lt;/code&gt; — the complete reply.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;.stream().content()&lt;/code&gt; gives you a &lt;code&gt;Flux&amp;lt;String&amp;gt;&lt;/code&gt; — a stream of token chunks that arrive one by one as Gemini generates them.&lt;/p&gt;

&lt;p&gt;Spring automatically serialises a &lt;code&gt;Flux&amp;lt;String&amp;gt;&lt;/code&gt; return type as SSE when you set &lt;code&gt;produces = MediaType.TEXT_EVENT_STREAM_VALUE&lt;/code&gt;. Each item in the Flux becomes a &lt;code&gt;data:&lt;/code&gt; event in the SSE stream.&lt;/p&gt;

&lt;p&gt;You don't need to add any WebFlux dependency — Spring AI's streaming support works in a regular Spring Boot MVC application. The &lt;code&gt;Flux&lt;/code&gt; return type is handled automatically.&lt;/p&gt;




&lt;h2&gt;
  
  
  Add the dependency
&lt;/h2&gt;

&lt;p&gt;Make sure you have the Reactor dependency. If you're using &lt;code&gt;spring-boot-starter-web&lt;/code&gt;, add:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;dependency&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;groupId&amp;gt;&lt;/span&gt;io.projectreactor&lt;span class="nt"&gt;&amp;lt;/groupId&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;artifactId&amp;gt;&lt;/span&gt;reactor-core&lt;span class="nt"&gt;&amp;lt;/artifactId&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/dependency&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Spring Boot manages the version through the BOM, so no version needed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Full streaming controller
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@RestController&lt;/span&gt;
&lt;span class="nd"&gt;@RequestMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;StreamingChatController&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;ChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;StreamingChatController&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ChatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;Builder&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;JdbcTemplate&lt;/span&gt; &lt;span class="n"&gt;jdbcTemplate&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;JdbcChatMemoryRepository&lt;/span&gt; &lt;span class="n"&gt;memoryRepository&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;JdbcChatMemoryRepository&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;jdbcTemplate&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;jdbcTemplate&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;dialect&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ChatHistoryDialect&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="nc"&gt;MessageWindowChatMemory&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MessageWindowChatMemory&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatMemoryRepository&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryRepository&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;maxMessages&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatClient&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;defaultSystem&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"You are a helpful assistant."&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;defaultAdvisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;MessageChatMemoryAdvisor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// Non-streaming — returns the full reply at once&lt;/span&gt;
    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/chat-ai"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"message"&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;advisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;param&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"chat_memory_conversation_id"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;)))&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// Streaming — returns tokens as they arrive&lt;/span&gt;
    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"/chat-ai/stream"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;produces&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MediaType&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;TEXT_EVENT_STREAM_VALUE&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;Flux&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;streamChat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"message"&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;advisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;param&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"chat_memory_conversation_id"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;)))&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;record&lt;/span&gt; &lt;span class="nf"&gt;ChatRequest&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both endpoints share the same &lt;code&gt;chatClient&lt;/code&gt; with the same memory. The streaming endpoint is just an extra route — your existing non-streaming endpoint keeps working.&lt;/p&gt;




&lt;h2&gt;
  
  
  Test it with curl
&lt;/h2&gt;

&lt;p&gt;Start the app, then run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-N&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST http://localhost:8080/api/chat-ai/stream &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"conversationId":"test-123","message":"Explain what a Docker container is in simple terms"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;-N&lt;/code&gt; flag disables buffering — without it curl would collect everything before displaying it, which defeats the purpose.&lt;/p&gt;

&lt;p&gt;You should see tokens arriving one by one in your terminal, not all at once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;data:A

data: Docker

data: container

data: is

data: like

data: a

data: lightweight

data: virtual

data: machine
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each &lt;code&gt;data:&lt;/code&gt; line is one SSE event — one token chunk from the model.&lt;/p&gt;




&lt;h2&gt;
  
  
  What about the frontend?
&lt;/h2&gt;

&lt;p&gt;The browser has a built-in &lt;code&gt;EventSource&lt;/code&gt; API for consuming SSE, but it only supports GET requests. Since your endpoint is a POST (because you're sending a JSON body), you use the Fetch API with a readable stream instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/api/chat-ai/stream&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;conversationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getReader&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextDecoder&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;done&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;done&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="c1"&gt;// each chunk is "data: token\n\n" — parse and append to your UI&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/^data: /&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="nf"&gt;appendToUI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each chunk you read from the stream is a token. Append it to the message div as it arrives — same effect as ChatGPT's typing animation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Streaming vs non-streaming — when to use which
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Non-streaming&lt;/th&gt;
&lt;th&gt;Streaming&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Response time to first byte&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Slow (waits for full reply)&lt;/td&gt;
&lt;td&gt;Instant (first token arrives immediately)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;UX&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Blank screen then full reply&lt;/td&gt;
&lt;td&gt;Tokens appear as they're generated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Frontend complexity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Simple fetch + JSON&lt;/td&gt;
&lt;td&gt;Fetch + ReadableStream parsing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Good for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Internal APIs, background processing&lt;/td&gt;
&lt;td&gt;Any user-facing chat interface&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Keep both endpoints. Use the streaming one for your frontend chat UI. Use the non-streaming one for internal calls where you need the complete reply as a string — tool calling, post-processing, logging the full response.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The app now streams responses like ChatGPT. But streaming means more API calls, and more API calls means more tokens — and tokens cost money. Next: understanding and managing API cost before your first surprise bill.&lt;/p&gt;




&lt;p&gt;Noticed a difference in feel between streaming and waiting for the full reply? Drop it in the comments.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>java</category>
      <category>springboot</category>
      <category>ai</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Stop Running Your AI App on Your Laptop — Deploy It to the Cloud for Free</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Fri, 25 Sep 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/stop-running-your-ai-app-on-your-laptop-deploy-it-to-the-cloud-for-free-7e2</link>
      <guid>https://dev.to/shamprakash2000/stop-running-your-ai-app-on-your-laptop-deploy-it-to-the-cloud-for-free-7e2</guid>
      <description>&lt;p&gt;The chat app works. It remembers conversations. It stores history in PostgreSQL.&lt;/p&gt;

&lt;p&gt;And it runs on my laptop.&lt;/p&gt;

&lt;p&gt;That's not a chat app. That's a script. The moment I close the terminal, it's gone. Nobody else can use it.&lt;/p&gt;

&lt;p&gt;This article is about fixing that. We're taking the app and deploying it to the cloud so it runs 24/7 at a real URL.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to deploy — a quick tour
&lt;/h2&gt;

&lt;p&gt;There are several platforms that make deploying a Spring Boot app straightforward. Here's the honest overview:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Railway&lt;/strong&gt; — fast setup, generous free tier, great developer experience. The free tier has a monthly usage limit that can catch you off guard if the app gets traffic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fly.io&lt;/strong&gt; — powerful, runs real VMs in regions close to your users, excellent for performance-sensitive apps. The free tier is there but the setup is more involved — you configure regions, vm sizes, and scaling yourself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Koyeb&lt;/strong&gt; — solid option, deploys from Docker or GitHub, decent free tier. Less documentation than the others.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Render&lt;/strong&gt; — this is what I used. Free tier, no credit card required for basic deployment, deploys directly from a GitHub repo and rebuilds on every push. The free tier does spin down after 15 minutes of inactivity, but for learning and demos it's exactly what you need.&lt;/p&gt;

&lt;p&gt;We're going with Render.&lt;/p&gt;




&lt;h2&gt;
  
  
  What you need before deploying
&lt;/h2&gt;

&lt;p&gt;Three things need to be true before deployment works:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Your code is in a GitHub repo&lt;/li&gt;
&lt;li&gt;Your secrets are in environment variables, not in the code&lt;/li&gt;
&lt;li&gt;Your app is packaged as a Docker image&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You probably have 1 and 2 already. Let's talk about 3.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Docker?
&lt;/h2&gt;

&lt;p&gt;When you run the app locally, Java is installed on your machine. Maven is installed. The right version of everything is there.&lt;/p&gt;

&lt;p&gt;Render's servers don't have your setup. They don't know you're using Java 17. They don't know your Maven version. They don't know your dependencies.&lt;/p&gt;

&lt;p&gt;Docker solves this by packaging your app and everything it needs into one self-contained image. The image runs the same way everywhere — your laptop, Render, AWS, anywhere.&lt;/p&gt;

&lt;p&gt;Think of it like a shipping container. Before shipping containers, every port had to know how to handle every type of cargo. After containers, you just move the box — the contents are someone else's problem.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;Dockerfile&lt;/code&gt; is the recipe for building that box.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Dockerfile
&lt;/h2&gt;

&lt;p&gt;Create a file called &lt;code&gt;Dockerfile&lt;/code&gt; in the root of your project (same level as &lt;code&gt;pom.xml&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;maven:3.9-eclipse-temurin-17&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;AS&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;build&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; pom.xml .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;mvn dependency:go-offline &lt;span class="nt"&gt;-q&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; src ./src&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;mvn package &lt;span class="nt"&gt;-DskipTests&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt;

&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; eclipse-temurin:17-jre&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --from=build /app/target/gemini-chat-0.0.1-SNAPSHOT.jar app.jar&lt;/span&gt;
&lt;span class="k"&gt;EXPOSE&lt;/span&gt;&lt;span class="s"&gt; 8080&lt;/span&gt;
&lt;span class="k"&gt;ENTRYPOINT&lt;/span&gt;&lt;span class="s"&gt; ["java", "-jar", "app.jar"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a &lt;strong&gt;multi-stage build&lt;/strong&gt;. Two stages, one file.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage 1 — build the app:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;FROM maven:3.9-eclipse-temurin-17 AS build&lt;/code&gt; starts from a base image that already has Maven and Java 17 installed. You name this stage &lt;code&gt;build&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;COPY pom.xml .&lt;/code&gt; then &lt;code&gt;RUN mvn dependency:go-offline&lt;/code&gt; — this is a Docker trick. By copying just &lt;code&gt;pom.xml&lt;/code&gt; first and downloading dependencies before copying source code, Docker can cache that layer. If you only change Java files later, Docker skips the dependency download step entirely. Builds get much faster.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;COPY src ./src&lt;/code&gt; then &lt;code&gt;RUN mvn package -DskipTests&lt;/code&gt; — now copy the source and compile. &lt;code&gt;-DskipTests&lt;/code&gt; is intentional here; you'd run tests in CI separately.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage 2 — run the app:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;FROM eclipse-temurin:17-jre&lt;/code&gt; — this is a different base image. Just the JRE (Java Runtime Environment), not the full JDK with Maven. Much smaller — a JDK image can be 500MB+, a JRE is ~200MB.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;COPY --from=build /app/target/gemini-chat-0.0.1-SNAPSHOT.jar app.jar&lt;/code&gt; — copy just the compiled jar from stage 1. The Maven installation, all the source code, the intermediate build files — none of that comes along.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;EXPOSE 8080&lt;/code&gt; — tells Docker (and Render) what port the app listens on.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ENTRYPOINT ["java", "-jar", "app.jar"]&lt;/code&gt; — the command that runs when the container starts.&lt;/p&gt;

&lt;p&gt;The final image contains only the JRE and your jar. Everything else stays behind in stage 1 and is discarded.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where does the Docker image go?
&lt;/h2&gt;

&lt;p&gt;This is something I wasn't clear on at first.&lt;/p&gt;

&lt;p&gt;You never push a Docker image anywhere. You don't need Docker Hub. You don't need Docker installed on your machine at all.&lt;/p&gt;

&lt;p&gt;Here's what actually happens when you deploy to Render:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You push your source code to GitHub (including the &lt;code&gt;Dockerfile&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Render pulls that code from GitHub&lt;/li&gt;
&lt;li&gt;Render runs your &lt;code&gt;Dockerfile&lt;/code&gt; on their own build servers — Maven downloads dependencies, compiles your Java code, packages the jar&lt;/li&gt;
&lt;li&gt;The resulting Docker image is stored in Render's internal container registry&lt;/li&gt;
&lt;li&gt;Render starts a container from that image and gives you a URL&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The image lives on Render's infrastructure. Maven is only involved inside stage 1 of the build — it's just the tool that compiles the Java code. Once the image is built, Maven is gone.&lt;/p&gt;

&lt;p&gt;The only reason to run Docker locally is if you want to test the image before pushing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker build &lt;span class="nt"&gt;-t&lt;/span&gt; gemini-chat &lt;span class="nb"&gt;.&lt;/span&gt;
docker run &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:8080 &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nv"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;your_key gemini-chat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But for deploying — GitHub to Render is all you need.&lt;/p&gt;




&lt;h2&gt;
  
  
  Check your environment variables
&lt;/h2&gt;

&lt;p&gt;Before pushing, verify your &lt;code&gt;application.properties&lt;/code&gt; uses environment variables, not hardcoded values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="c"&gt;# Gemini
&lt;/span&gt;&lt;span class="py"&gt;spring.ai.google.gemini.api-key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${GEMINI_API_KEY}&lt;/span&gt;
&lt;span class="py"&gt;spring.ai.google.gemini.chat.options.model&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;gemini-2.5-flash&lt;/span&gt;

&lt;span class="c"&gt;# Database
&lt;/span&gt;&lt;span class="py"&gt;spring.datasource.url&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${DATABASE_URL}&lt;/span&gt;
&lt;span class="py"&gt;spring.datasource.username&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${DATABASE_USERNAME}&lt;/span&gt;
&lt;span class="py"&gt;spring.datasource.password&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${DATABASE_PASSWORD}&lt;/span&gt;

&lt;span class="py"&gt;spring.sql.init.schema-locations&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;classpath:chat-memory-schema.sql&lt;/span&gt;
&lt;span class="py"&gt;spring.sql.init.mode&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;always&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every secret is read from the environment at runtime. If someone opens your GitHub repo, they see variable names, not actual values.&lt;/p&gt;

&lt;p&gt;Also check your &lt;code&gt;.gitignore&lt;/code&gt; — this should already be there, but verify:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="c"&gt;# Never commit this
&lt;/span&gt;&lt;span class="err"&gt;.env&lt;/span&gt;
&lt;span class="err"&gt;*.env&lt;/span&gt;
&lt;span class="err"&gt;application-local.properties&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Deploy to Render
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Step 1 — Push your code to GitHub&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Your project needs to be in a GitHub repo. If it's not already:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git init
git add &lt;span class="nb"&gt;.&lt;/span&gt;
git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"initial commit"&lt;/span&gt;
git remote add origin https://github.com/yourusername/gemini-chat.git
git push &lt;span class="nt"&gt;-u&lt;/span&gt; origin main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Make absolutely sure your API keys and database credentials are not in any committed file.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2 — Create a Render account&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Go to &lt;a href="https://render.com" rel="noopener noreferrer"&gt;render.com&lt;/a&gt; and sign up. You can use GitHub to sign in — that's the easiest option since Render will need access to your repos anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3 — Create a new Web Service&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In the Render dashboard:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Click &lt;strong&gt;New&lt;/strong&gt; → &lt;strong&gt;Web Service&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Connect your GitHub account and select your repository&lt;/li&gt;
&lt;li&gt;Render will detect the &lt;code&gt;Dockerfile&lt;/code&gt; automatically&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Set the basics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Name:&lt;/strong&gt; gemini-chat (or whatever you like)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Region:&lt;/strong&gt; Pick the closest one to you&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Branch:&lt;/strong&gt; main&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instance type:&lt;/strong&gt; Free&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 4 — Add environment variables&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In the service settings, scroll to &lt;strong&gt;Environment Variables&lt;/strong&gt; and add:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Key&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GEMINI_API_KEY&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;your Gemini API key&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DATABASE_URL&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;your Neon connection string&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DATABASE_USERNAME&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;your Neon username&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DATABASE_PASSWORD&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;your Neon password&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is where secrets live on Render — not in your code, not in your repo. Render injects these as environment variables when the container starts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 5 — Deploy&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Click &lt;strong&gt;Create Web Service&lt;/strong&gt;. Render will:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pull your code from GitHub&lt;/li&gt;
&lt;li&gt;Build the Docker image (this runs your Dockerfile — Maven downloads dependencies, compiles the app, packages the jar)&lt;/li&gt;
&lt;li&gt;Start the container&lt;/li&gt;
&lt;li&gt;Give you a URL like &lt;code&gt;https://gemini-chat-xxxx.onrender.com&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first build takes a few minutes. After that, every push to main triggers a new build automatically.&lt;/p&gt;




&lt;h2&gt;
  
  
  Test it
&lt;/h2&gt;

&lt;p&gt;Once the service is live, send a request to your Render URL:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://gemini-chat-xxxx.onrender.com/api/chat-ai/session
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should get back a conversation ID. Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://gemini-chat-xxxx.onrender.com/api/chat-ai/chat &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"conversationId":"your-id-here","message":"Hello, are you running in the cloud?"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you get a response — it's live. Running on Render's servers, hitting Gemini's API, storing history in Neon PostgreSQL. Not on your laptop anymore.&lt;/p&gt;




&lt;h2&gt;
  
  
  The free tier caveat
&lt;/h2&gt;

&lt;p&gt;Render's free tier spins the service down after 15 minutes of no traffic. The first request after a sleep wakes it up, which takes about 30-60 seconds. For production you'd pay for an always-on instance, but for learning and demos the free tier is fine.&lt;/p&gt;

&lt;p&gt;If you want to keep it warm, you can set up a simple ping from a service like UptimeRobot — free, hits your URL every 5 minutes, keeps the service awake.&lt;/p&gt;




&lt;h2&gt;
  
  
  What just happened
&lt;/h2&gt;

&lt;p&gt;Let's step back and look at what you've built:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Spring Boot REST API that calls Gemini&lt;/li&gt;
&lt;li&gt;Conversation history that persists in PostgreSQL&lt;/li&gt;
&lt;li&gt;Deployed to the cloud with Docker&lt;/li&gt;
&lt;li&gt;Running at a real URL, 24/7&lt;/li&gt;
&lt;li&gt;All secrets safely in environment variables, never in code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's a real AI backend. Not a tutorial project — a deployed service.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The app is live. Next: making the responses stream in real time — instead of waiting for the full reply, you see tokens appear as the model generates them. Same way ChatGPT works.&lt;/p&gt;




&lt;p&gt;Deployed your first AI app? Drop the URL in the comments.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>java</category>
      <category>docker</category>
      <category>beginners</category>
      <category>devops</category>
    </item>
    <item>
      <title>Why Your AI Chatbot Forgets Everything — And How to Fix It</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Wed, 23 Sep 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/why-your-ai-chatbot-forgets-everything-and-how-to-fix-it-26je</link>
      <guid>https://dev.to/shamprakash2000/why-your-ai-chatbot-forgets-everything-and-how-to-fix-it-26je</guid>
      <description>&lt;p&gt;In the last article we built a working chat endpoint. Send a message, get a reply. It felt like magic.&lt;/p&gt;

&lt;p&gt;Then I tried to have an actual conversation.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Me: "My name is Sham."&lt;br&gt;
AI: "Hi Sham! How can I help you?"&lt;br&gt;
Me: "What's my name?"&lt;br&gt;
AI: "I don't have access to personal information about you."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model had completely forgotten who I was. Not because it was broken — because of something fundamental about how LLMs work. Every API call is completely independent. The model has no memory between calls.&lt;/p&gt;

&lt;p&gt;If you want it to remember anything, that's your problem to solve.&lt;/p&gt;

&lt;p&gt;This article shows how — starting from the simplest possible solution, hitting its limits, then building the real one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the model forgets
&lt;/h2&gt;

&lt;p&gt;When you call the Gemini API, you send a list of messages. The model reads them, generates a reply, and the call ends. The next call starts completely fresh — the model has no idea the previous call ever happened.&lt;/p&gt;

&lt;p&gt;So when the user sends message 5, the model only sees message 5. It has no knowledge of messages 1 through 4.&lt;/p&gt;

&lt;p&gt;The fix is simple in concept: &lt;strong&gt;include all previous messages in every call.&lt;/strong&gt; Send the full conversation history every time, so the model always has context.&lt;/p&gt;

&lt;p&gt;Let's build that.&lt;/p&gt;




&lt;h2&gt;
  
  
  Solution 1 — A simple Map in memory
&lt;/h2&gt;

&lt;p&gt;The simplest fix: a &lt;code&gt;Map&lt;/code&gt; where the key is a session ID and the value is the list of messages for that session.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@RestController&lt;/span&gt;
&lt;span class="nd"&gt;@RequestMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api/chat"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatController&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;ChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="c1"&gt;// session ID → list of messages for that session&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;List&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Message&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;sessions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ConcurrentHashMap&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;();&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;ChatController&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ChatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;Builder&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatClient&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;defaultSystem&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"You are a helpful assistant."&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/session"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;startSession&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;sessionId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;UUID&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;randomUUID&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;toString&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;sessions&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;put&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sessionId&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ArrayList&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;());&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"sessionId"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sessionId&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;ChatRequest&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;List&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Message&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sessions&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getOrDefault&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sessionId&lt;/span&gt;&lt;span class="o"&gt;(),&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ArrayList&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;gt;());&lt;/span&gt;

        &lt;span class="c1"&gt;// Add user message to history&lt;/span&gt;
        &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;add&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;UserMessage&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="o"&gt;()));&lt;/span&gt;

        &lt;span class="c1"&gt;// Send full history to the model&lt;/span&gt;
        &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="c1"&gt;// Add model reply to history&lt;/span&gt;
        &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;add&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;AssistantMessage&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;
        &lt;span class="n"&gt;sessions&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;put&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;sessionId&lt;/span&gt;&lt;span class="o"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;record&lt;/span&gt; &lt;span class="nf"&gt;ChatRequest&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;sessionId&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every user gets their own session ID. Their messages are stored in their own list. Every API call sends the full history for that session — so the model has context.&lt;/p&gt;

&lt;p&gt;Now try the conversation:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Me: "My name is Sham."&lt;br&gt;
AI: "Hi Sham! How can I help you?"&lt;br&gt;
Me: "What's my name?"&lt;br&gt;
AI: "Your name is Sham."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It works. Two different users with two different session IDs — completely separate conversations.&lt;/p&gt;

&lt;p&gt;The code is simple. Every Java developer knows what a &lt;code&gt;Map&lt;/code&gt; and a &lt;code&gt;List&lt;/code&gt; are. No framework magic, just plain Java.&lt;/p&gt;




&lt;h2&gt;
  
  
  The problem with in-memory history
&lt;/h2&gt;

&lt;p&gt;This works great — until you restart the server. All history is gone. Everyone's conversations, gone.&lt;/p&gt;

&lt;p&gt;There's another problem: this is a single &lt;code&gt;ArrayList&lt;/code&gt; shared across all users. User A and User B are in the same conversation. Not great.&lt;/p&gt;

&lt;p&gt;And there's the token problem: a long conversation becomes thousands of tokens on every call, whether those old messages are relevant or not.&lt;/p&gt;

&lt;p&gt;In-memory works for a quick demo. For anything real, you need persistent storage with session isolation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where to get a free PostgreSQL database
&lt;/h2&gt;

&lt;p&gt;Before writing any code, you need a database. The easiest free option is &lt;a href="https://neon.tech" rel="noopener noreferrer"&gt;Neon&lt;/a&gt; — serverless PostgreSQL, free tier, no credit card required.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Go to &lt;a href="https://neon.tech" rel="noopener noreferrer"&gt;neon.tech&lt;/a&gt; and sign up&lt;/li&gt;
&lt;li&gt;Create a new project — Neon gives you a PostgreSQL database instantly&lt;/li&gt;
&lt;li&gt;Copy the connection string from the dashboard — it looks like:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;postgresql&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="n"&gt;username&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;password&lt;/span&gt;&lt;span class="o"&gt;@&lt;/span&gt;&lt;span class="n"&gt;ep&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;xxx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;us&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;east&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;aws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;neon&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tech&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;dbname&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="n"&gt;sslmode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;require&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Set it as an environment variable:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;postgresql://username:password@...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. Free, no setup, no local PostgreSQL installation needed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Solution 2 — PostgreSQL with Spring AI
&lt;/h2&gt;

&lt;p&gt;Spring AI has a built-in &lt;code&gt;JdbcChatMemoryRepository&lt;/code&gt; that stores conversation history in a database. Each conversation gets a unique ID — so different users are completely isolated from each other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1 — Add the dependency&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;dependency&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;groupId&amp;gt;&lt;/span&gt;org.springframework.ai&lt;span class="nt"&gt;&amp;lt;/groupId&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;artifactId&amp;gt;&lt;/span&gt;spring-ai-starter-model-chat-memory-repository-jdbc&lt;span class="nt"&gt;&amp;lt;/artifactId&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/dependency&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;dependency&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;groupId&amp;gt;&lt;/span&gt;org.postgresql&lt;span class="nt"&gt;&amp;lt;/groupId&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;artifactId&amp;gt;&lt;/span&gt;postgresql&lt;span class="nt"&gt;&amp;lt;/artifactId&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;scope&amp;gt;&lt;/span&gt;runtime&lt;span class="nt"&gt;&amp;lt;/scope&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/dependency&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2 — Create the table&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Create &lt;code&gt;src/main/resources/chat-memory-schema.sql&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;chat_history&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;conversation_id&lt;/span&gt; &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;         &lt;span class="nb"&gt;TEXT&lt;/span&gt;         &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;type&lt;/span&gt;            &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;64&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nb"&gt;timestamp&lt;/span&gt;       &lt;span class="nb"&gt;TIMESTAMP&lt;/span&gt;    &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 3 — Tell Spring AI to use your table&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;By default Spring AI uses a table called &lt;code&gt;SPRING_AI_CHAT_MEMORY&lt;/code&gt;. To use your own table name, implement &lt;code&gt;JdbcChatMemoryRepositoryDialect&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatHistoryDialect&lt;/span&gt; &lt;span class="kd"&gt;implements&lt;/span&gt; &lt;span class="nc"&gt;JdbcChatMemoryRepositoryDialect&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="no"&gt;TABLE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"chat_history"&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="nd"&gt;@Override&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;getSelectMessagesSql&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s"&gt;"SELECT content, type FROM "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="no"&gt;TABLE&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
               &lt;span class="s"&gt;" WHERE conversation_id = ? ORDER BY timestamp"&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="nd"&gt;@Override&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;getInsertMessageSql&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s"&gt;"INSERT INTO "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="no"&gt;TABLE&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
               &lt;span class="s"&gt;" (conversation_id, content, type, timestamp) VALUES (?, ?, ?, ?)"&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="nd"&gt;@Override&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;getSelectConversationIdsSql&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s"&gt;"SELECT DISTINCT conversation_id FROM "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="no"&gt;TABLE&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="nd"&gt;@Override&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;getDeleteMessagesSql&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s"&gt;"DELETE FROM "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="no"&gt;TABLE&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" WHERE conversation_id = ?"&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 4 — Wire it up in the controller&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@RestController&lt;/span&gt;
&lt;span class="nd"&gt;@RequestMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api/chat-ai"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SpringAiChatController&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;ChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;JdbcChatMemoryRepository&lt;/span&gt; &lt;span class="n"&gt;memoryRepository&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;SpringAiChatController&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ChatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;Builder&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;JdbcTemplate&lt;/span&gt; &lt;span class="n"&gt;jdbcTemplate&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;memoryRepository&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;JdbcChatMemoryRepository&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;jdbcTemplate&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;jdbcTemplate&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;dialect&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ChatHistoryDialect&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="c1"&gt;// Keep last 20 messages — older ones are evicted automatically&lt;/span&gt;
        &lt;span class="nc"&gt;MessageWindowChatMemory&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MessageWindowChatMemory&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatMemoryRepository&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryRepository&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;maxMessages&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatClient&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;defaultSystem&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"You are a helpful assistant."&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;defaultAdvisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;MessageChatMemoryAdvisor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// Create a new session — returns a unique conversation ID&lt;/span&gt;
    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/session"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;startSession&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;conversationId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;UUID&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;randomUUID&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;toString&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// Chat — pass the conversation ID with every message&lt;/span&gt;
    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/chat"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;ChatRequest&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;advisors&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;param&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"chat_memory_conversation_id"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;()))&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// Delete a conversation&lt;/span&gt;
    &lt;span class="nd"&gt;@DeleteMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/session/{conversationId}"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;deleteSession&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@PathVariable&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;memoryRepository&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;deleteByConversationId&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"status"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"deleted"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"conversationId"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;record&lt;/span&gt; &lt;span class="nf"&gt;ChatRequest&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;conversationId&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 5 — Configure &lt;code&gt;application.properties&lt;/code&gt;&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;spring.datasource.url&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${DATABASE_URL}&lt;/span&gt;
&lt;span class="py"&gt;spring.datasource.username&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${DATABASE_USERNAME}&lt;/span&gt;
&lt;span class="py"&gt;spring.datasource.password&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${DATABASE_PASSWORD}&lt;/span&gt;

&lt;span class="py"&gt;spring.sql.init.schema-locations&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;classpath:chat-memory-schema.sql&lt;/span&gt;
&lt;span class="py"&gt;spring.sql.init.mode&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;always&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;mode=always&lt;/code&gt; forces the schema script to run on every startup. Without it, PostgreSQL (being a non-embedded database) won't run the script at all.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works now
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Starting a conversation:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /api/chat-ai/session
→ { "conversationId": "a3f9b2c1-..." }
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Sending messages — pass the ID every time:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /api/chat-ai/chat
{ "conversationId": "a3f9b2c1-...", "message": "My name is Sham." }
→ "Hi Sham! How can I help you?"

POST /api/chat-ai/chat
{ "conversationId": "a3f9b2c1-...", "message": "What's my name?" }
→ "Your name is Sham."
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Spring AI fetches the last 20 messages for that conversation ID from PostgreSQL, includes them in the API call, saves the new exchange, and returns the response. You wrote none of that logic yourself.&lt;/p&gt;

&lt;p&gt;Two users, two different conversation IDs — completely isolated. Server restarts — history survives. Long conversation — only the last 20 messages are sent, keeping tokens under control.&lt;/p&gt;




&lt;h2&gt;
  
  
  In-memory vs PostgreSQL — when to use which
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Spring AI In-Memory&lt;/th&gt;
&lt;th&gt;Spring AI PostgreSQL&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Setup&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Zero&lt;/td&gt;
&lt;td&gt;Database + dependency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Survives restart&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Multi-user&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — isolated by session ID&lt;/td&gt;
&lt;td&gt;Yes — isolated by session ID&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Token control&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Automatic (maxMessages)&lt;/td&gt;
&lt;td&gt;Automatic (maxMessages)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Good for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Local dev, quick demos&lt;/td&gt;
&lt;td&gt;Production, anything real&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Start with in-memory locally. Switch to PostgreSQL before you deploy.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The chat app has memory now. Next up: deploying it to the cloud — Render, Docker, environment variables. Because a chat app that only runs on your laptop isn't a chat app, it's a script.&lt;/p&gt;




&lt;p&gt;Have you hit the "it forgot everything" problem before understanding why? Drop it in the comments.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>java</category>
      <category>springboot</category>
      <category>beginners</category>
    </item>
    <item>
      <title>From 20 Lines to 4: My First AI Endpoint in Spring Boot</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/from-20-lines-to-4-my-first-ai-endpoint-in-spring-boot-3010</link>
      <guid>https://dev.to/shamprakash2000/from-20-lines-to-4-my-first-ai-endpoint-in-spring-boot-3010</guid>
      <description>&lt;p&gt;Before you install any framework, before you add any dependency — you should make one raw HTTP call to the model yourself. Just to see what it actually is.&lt;/p&gt;

&lt;p&gt;That's how I started. And it's how I'd recommend everyone starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1 — Get a Gemini API key
&lt;/h2&gt;

&lt;p&gt;Go to &lt;a href="https://aistudio.google.com/app/apikey" rel="noopener noreferrer"&gt;Google AI Studio&lt;/a&gt; and create a free API key. No credit card needed. Gemini has a generous free tier — more than enough to learn and build.&lt;/p&gt;

&lt;p&gt;Once you have the key, store it as an environment variable. Never paste it directly into your code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;your_key_here&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Step 2 — Understand what you're calling
&lt;/h2&gt;

&lt;p&gt;Gemini exposes a REST API. The endpoint looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent?key=YOUR_API_KEY
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You send a JSON body, you get a JSON response. That's it. No magic.&lt;/p&gt;

&lt;p&gt;The request body looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"parts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"What is a token in LLMs?"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the response comes back like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"candidates"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"parts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"A token is a piece of text..."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"model"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your answer is at &lt;code&gt;candidates[0].content.parts[0].text&lt;/code&gt;. Navigate that JSON tree and you have your response.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3 — Build it with plain Java, no frameworks
&lt;/h2&gt;

&lt;p&gt;Here's a working Spring Boot controller that calls Gemini using nothing but Java's built-in &lt;code&gt;HttpClient&lt;/code&gt; and Jackson (which Spring Boot already includes). No Spring AI. No extra dependencies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;pom.xml&lt;/code&gt;&lt;/strong&gt; — intentionally minimal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;dependency&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;groupId&amp;gt;&lt;/span&gt;org.springframework.boot&lt;span class="nt"&gt;&amp;lt;/groupId&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;artifactId&amp;gt;&lt;/span&gt;spring-boot-starter-web&lt;span class="nt"&gt;&amp;lt;/artifactId&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/dependency&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the only dependency. Spring Boot 3.x includes Jackson. Java 11+ includes &lt;code&gt;HttpClient&lt;/code&gt;. Nothing else needed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;application.properties&lt;/code&gt;&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;gemini.api.key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${GEMINI_API_KEY}&lt;/span&gt;
&lt;span class="py"&gt;gemini.api.url&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;&lt;code&gt;ChatController.java&lt;/code&gt;&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@RestController&lt;/span&gt;
&lt;span class="nd"&gt;@RequestMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api/chat"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatController&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

    &lt;span class="nd"&gt;@Value&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"${gemini.api.key}"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;apiKey&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="nd"&gt;@Value&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"${gemini.api.url}"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;apiUrl&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;ObjectMapper&lt;/span&gt; &lt;span class="n"&gt;objectMapper&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ObjectMapper&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;HttpClient&lt;/span&gt; &lt;span class="n"&gt;httpClient&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;HttpClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;newHttpClient&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;ResponseEntity&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;userMessage&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;Exception&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

        &lt;span class="c1"&gt;// Build the request body&lt;/span&gt;
        &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;requestBody&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""
            {
              "contents": [
                {
                  "role": "user",
                  "parts": [{ "text": "%s" }]
                }
              ]
            }
        """&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;formatted&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;escape&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;userMessage&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;

        &lt;span class="c1"&gt;// Make the HTTP call&lt;/span&gt;
        &lt;span class="nc"&gt;HttpRequest&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;HttpRequest&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;newBuilder&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;uri&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="no"&gt;URI&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;create&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;apiUrl&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;"?key="&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;apiKey&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;header&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Content-Type"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"application/json"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;POST&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;HttpRequest&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;BodyPublishers&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;ofString&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;requestBody&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="nc"&gt;HttpResponse&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;httpClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;send&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;
            &lt;span class="nc"&gt;HttpResponse&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;BodyHandlers&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;ofString&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;

        &lt;span class="c1"&gt;// Parse the response — answer is at candidates[0].content.parts[0].text&lt;/span&gt;
        &lt;span class="nc"&gt;JsonNode&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;objectMapper&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;readTree&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;has&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"error"&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;ResponseEntity&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
                &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"error"&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"message"&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;asText&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;

        &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"candidates"&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"content"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"parts"&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"text"&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;asText&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;ResponseEntity&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;escape&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;replace&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"\\"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"\\\\"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;replace&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"\""&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"\\\""&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;replace&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"\n"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"\\n"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;replace&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"\r"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"\\r"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start the app, send a POST request to &lt;code&gt;/api/chat&lt;/code&gt; with a message body, and you'll get a response from Gemini.&lt;/p&gt;

&lt;p&gt;That's a fully working AI chat endpoint. No AI framework. Just HTTP.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Full source code for this phase: &lt;a href="https://github.com/shamprakash2000/gemini-chat/tree/phase-1-2-plain-http" rel="noopener noreferrer"&gt;github.com/shamprakash2000/gemini-chat/tree/phase-1-2-plain-http&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What you just built — and what's missing
&lt;/h2&gt;

&lt;p&gt;This works. Send a message, get a response. That's a real AI endpoint.&lt;/p&gt;

&lt;p&gt;But look at what you're doing manually:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Building JSON by hand&lt;/strong&gt; — string formatting, escaping characters, constructing the payload yourself&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parsing JSON by hand&lt;/strong&gt; — navigating &lt;code&gt;.path().get().path()&lt;/code&gt; to extract the answer&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No model switching&lt;/strong&gt; — the URL and response format are Gemini-specific. Switching to Claude means rewriting everything&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For one simple call, this is fine. But the moment you want conversation memory, tool calling, streaming, or RAG — you're building a framework from scratch.&lt;/p&gt;

&lt;p&gt;That's where Spring AI comes in.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Spring AI actually is
&lt;/h2&gt;

&lt;p&gt;Spring AI is the Spring team's answer to: "we keep writing the same boilerplate to call LLMs — let's standardise it."&lt;/p&gt;

&lt;p&gt;It gives you:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A common interface across models.&lt;/strong&gt; The same &lt;code&gt;ChatClient&lt;/code&gt; code works for Gemini, OpenAI, Claude, Ollama. You swap the dependency and config — the Java code stays the same.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No more manual JSON.&lt;/strong&gt; The framework builds the request payload and parses the response for you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Conversation memory, tool calling, structured output, streaming&lt;/strong&gt; — all built in. We'll get to each of these in the articles ahead. For now, just understand what the foundation gives you.&lt;/p&gt;




&lt;h2&gt;
  
  
  The same endpoint, rewritten with Spring AI
&lt;/h2&gt;

&lt;p&gt;Add the dependency:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;dependencyManagement&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;dependencies&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;dependency&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;groupId&amp;gt;&lt;/span&gt;org.springframework.ai&lt;span class="nt"&gt;&amp;lt;/groupId&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;artifactId&amp;gt;&lt;/span&gt;spring-ai-bom&lt;span class="nt"&gt;&amp;lt;/artifactId&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;version&amp;gt;&lt;/span&gt;1.1.8&lt;span class="nt"&gt;&amp;lt;/version&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;type&amp;gt;&lt;/span&gt;pom&lt;span class="nt"&gt;&amp;lt;/type&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;scope&amp;gt;&lt;/span&gt;import&lt;span class="nt"&gt;&amp;lt;/scope&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/dependency&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/dependencies&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/dependencyManagement&amp;gt;&lt;/span&gt;

&lt;span class="nt"&gt;&amp;lt;dependency&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;groupId&amp;gt;&lt;/span&gt;org.springframework.ai&lt;span class="nt"&gt;&amp;lt;/groupId&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;artifactId&amp;gt;&lt;/span&gt;spring-ai-starter-model-google-gemini&lt;span class="nt"&gt;&amp;lt;/artifactId&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/dependency&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Update &lt;code&gt;application.properties&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;spring.ai.google.gemini.api-key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;${GEMINI_API_KEY}&lt;/span&gt;
&lt;span class="py"&gt;spring.ai.google.gemini.chat.options.model&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;gemini-2.5-flash&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the controller:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@RestController&lt;/span&gt;
&lt;span class="nd"&gt;@RequestMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api/chat"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatController&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;ChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;ChatController&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ChatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;Builder&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatClient&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;defaultSystem&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"You are a helpful assistant."&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;userMessage&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;userMessage&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
            &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same functionality. A fraction of the code.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Spring AI does not protect you from
&lt;/h2&gt;

&lt;p&gt;I want to be honest here because most tutorials aren't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The API changes between minor versions.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I was on Spring AI 1.0.x and upgraded to 1.1.x. Two things broke silently — no compiler errors, no deprecation warnings. Just runtime failures. The &lt;code&gt;ChatClient&lt;/code&gt; builder API had changed, and the way conversation memory was configured had moved to a different class entirely.&lt;/p&gt;

&lt;p&gt;Lesson: when you upgrade Spring AI, read the full changelog and test every integration point. Don't assume a minor version bump is safe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The abstraction hides internals.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The plain HTTP version showed you exactly what was happening — every byte in, every byte out. Spring AI hides that. When something goes wrong deep in the framework, you need to know what's underneath to debug it. That's another reason to start with plain HTTP first — so you have that mental model when the abstraction breaks.&lt;/p&gt;




&lt;h2&gt;
  
  
  Plain HTTP vs Spring AI — the honest comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Plain HTTP&lt;/th&gt;
&lt;th&gt;Spring AI&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Dependencies&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None beyond Spring Web&lt;/td&gt;
&lt;td&gt;Spring AI starter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Code volume&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;JSON handling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Manual&lt;/td&gt;
&lt;td&gt;Automatic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model switching&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full rewrite&lt;/td&gt;
&lt;td&gt;Swap config&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tool calling / agents&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Build from scratch&lt;/td&gt;
&lt;td&gt;Built in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Debugging&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full visibility&lt;/td&gt;
&lt;td&gt;Abstraction hides internals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Version stability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stable (it's just HTTP)&lt;/td&gt;
&lt;td&gt;Breaking changes between minors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Good for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Learning, full control&lt;/td&gt;
&lt;td&gt;Production Spring Boot apps&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Start with plain HTTP. Understand what you're abstracting. Then move to Spring AI.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Spring AI is set up and working. Next: building a proper chat application — with persistent conversation history in PostgreSQL, session management, and a system prompt that shapes the model's behaviour.&lt;/p&gt;




&lt;p&gt;Did you hit a breaking change in Spring AI between versions? Drop it in the comments — you're not the only one.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>java</category>
      <category>beginners</category>
      <category>springboot</category>
    </item>
    <item>
      <title>How Your Message Actually Reaches the AI Model — And What You're Really Sending</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Sat, 19 Sep 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/how-your-message-actually-reaches-the-ai-model-and-what-youre-really-sending-1nlb</link>
      <guid>https://dev.to/shamprakash2000/how-your-message-actually-reaches-the-ai-model-and-what-youre-really-sending-1nlb</guid>
      <description>&lt;p&gt;You've used ChatGPT. You've used Gemini. You type something, it replies.&lt;/p&gt;

&lt;p&gt;But have you ever stopped and wondered — what is actually happening when you hit send?&lt;/p&gt;

&lt;p&gt;Because when you're building AI systems, that question matters. A lot.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when you type a message
&lt;/h2&gt;

&lt;p&gt;Let's take the simplest example. You open ChatGPT and type:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Explain what a token is in simple terms."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You hit enter. A few seconds later, you get a response.&lt;/p&gt;

&lt;p&gt;What just happened?&lt;/p&gt;

&lt;p&gt;Your message travelled over the internet to a server — a very powerful one, running a very large model — where it was processed, and text was generated back to you. That's it. At the core, it is a request and a response. Text goes in, text comes out.&lt;/p&gt;

&lt;p&gt;Now here's the important part: &lt;strong&gt;that model lives somewhere.&lt;/strong&gt; It doesn't run in your browser. It doesn't run on your laptop. It runs on a server, hosted by OpenAI in this case. You're connected to it over the internet without even thinking about it.&lt;/p&gt;

&lt;p&gt;This is the same thing that happens when you build an AI backend. You send text to a model. You get text back. The difference is — you're the one writing the code that sends and receives it, not a chat UI.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where can a model live?
&lt;/h2&gt;

&lt;p&gt;Here's something most tutorials skip: the model doesn't have to be on someone else's server.&lt;/p&gt;

&lt;p&gt;There are three places a model can live:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Cloud — hosted by the company (Gemini, ChatGPT, Claude)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google hosts Gemini on their servers. OpenAI hosts GPT. Anthropic hosts Claude. You connect to them over the internet using an API key — a secret key they give you that proves you're allowed to use their model.&lt;/p&gt;

&lt;p&gt;You pay per use. Every token you send and receive costs a small amount. The models are powerful, always up to date, and you don't manage any infrastructure.&lt;/p&gt;

&lt;p&gt;This is what we use in this series — Gemini, hosted by Google.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Local — running on your own machine (Ollama, LM Studio)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Yes, you can download a model and run it on your laptop.&lt;/p&gt;

&lt;p&gt;Tools like &lt;a href="https://ollama.com" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; let you pull a model — Llama 3, Mistral, Phi — and run it locally. No internet needed. No API key. No cost per call. The model runs on your CPU or GPU and responds to requests at &lt;code&gt;http://localhost:11434&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The tradeoff: smaller models, slower responses, and your laptop fans will make their presence known. But for learning, experimentation, or privacy-sensitive use cases — it's a great option.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Self-hosted — your own server, your own model&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Companies with strict data privacy requirements sometimes host their own models on their own infrastructure. Same idea as Ollama but on a cloud server they control. No data ever leaves their network.&lt;/p&gt;

&lt;p&gt;This is less common for individual developers but worth knowing exists.&lt;/p&gt;




&lt;h2&gt;
  
  
  Cloud vs Local — which should you use?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Cloud (Gemini, GPT, Claude)&lt;/th&gt;
&lt;th&gt;Local (Ollama)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Pay per token&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Setup&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;API key, done&lt;/td&gt;
&lt;td&gt;Download model, install Ollama&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model quality&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;State of the art&lt;/td&gt;
&lt;td&gt;Smaller, less capable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Speed&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fast (their hardware)&lt;/td&gt;
&lt;td&gt;Depends on your machine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Privacy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Data goes to their servers&lt;/td&gt;
&lt;td&gt;Stays on your machine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Internet needed&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Good for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Production, serious projects&lt;/td&gt;
&lt;td&gt;Learning, experiments, privacy&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For this series, we use Gemini — cloud hosted, API key, fast, capable. But if you want to experiment without spending money, Ollama is a great companion tool. The concepts are identical. Only the URL and key change.&lt;/p&gt;




&lt;h2&gt;
  
  
  How you connect — the API key
&lt;/h2&gt;

&lt;p&gt;When you use a cloud-hosted model, you authenticate with an API key.&lt;/p&gt;

&lt;p&gt;Think of it like a password. You create one in the provider's dashboard, store it securely, and include it in every request you make. The server checks the key, confirms you're allowed to use the model, and processes your request.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Your app  ──── request + API key ────▶  Gemini servers
Your app  ◀─── response ────────────   Gemini servers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For local models with Ollama, there's no key. You just call &lt;code&gt;http://localhost:11434&lt;/code&gt; directly. The model is on your machine — no authentication needed.&lt;/p&gt;

&lt;p&gt;The structure of what you send is the same either way. Only the destination changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One rule about API keys — never break this:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Never hardcode your API key directly in your source code. The moment you push that code to GitHub — even a private repo — the key is at risk. Always store it in an environment variable and read it from there.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Never do this&lt;/span&gt;
&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;apiKey&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"AIzaSyD-xxxxxxxxxxxxxxxxxxxxxxxx"&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// Always do this&lt;/span&gt;
&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;apiKey&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;System&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getenv&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"GEMINI_API_KEY"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not optional. API keys get scraped from public repos within minutes. Some providers will suspend your account automatically if they detect a leaked key.&lt;/p&gt;




&lt;h2&gt;
  
  
  What you actually send
&lt;/h2&gt;

&lt;p&gt;Now that you know where the model lives and how you connect — what do you actually put in the request?&lt;/p&gt;

&lt;p&gt;This is where most people expect something complicated. It's not.&lt;/p&gt;

&lt;p&gt;Every request to an LLM is a list of messages. Each message has two things: who sent it, and what they said.&lt;/p&gt;

&lt;p&gt;There are three senders — three &lt;strong&gt;roles&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;system&lt;/strong&gt; — you, the developer, giving the model its instructions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;user&lt;/strong&gt; — the person using your app, asking a question&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;assistant&lt;/strong&gt; — the model's previous replies&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A real request looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"system"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"You are a travel assistant for India. Only help with travel questions."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"What are the best places to visit in Rajasthan?"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Rajasthan has many beautiful destinations — Jaipur, Jodhpur, Udaipur..."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Which one is best for a 3-day trip?"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model reads this list top to bottom and writes the next assistant message.&lt;/p&gt;

&lt;p&gt;That's it. Three roles, in a list, sent on every call.&lt;/p&gt;




&lt;h2&gt;
  
  
  The system message — your instructions to the model
&lt;/h2&gt;

&lt;p&gt;The system message is where you tell the model who it is and what it should do. It's the first thing the model reads, and it sets the frame for everything else.&lt;/p&gt;

&lt;p&gt;Without a system message, the model behaves like a general-purpose assistant. With one, it becomes your travel assistant, your coding helper, your customer support agent — whatever you define.&lt;/p&gt;

&lt;p&gt;A few things I learned about writing them:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Be specific, not general.&lt;/strong&gt; "Be helpful" tells the model nothing. "Keep all responses under 5 sentences and always suggest at least one specific hotel" gives the model something concrete to follow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Put critical rules at the top.&lt;/strong&gt; The model pays more attention to what comes first. If there's one rule you really need followed — put it first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep it short.&lt;/strong&gt; I once wrote a 400-word system prompt. The model started contradicting itself. Five focused sentences beat twenty vague paragraphs.&lt;/p&gt;




&lt;h2&gt;
  
  
  The conversation history — why the model "remembers"
&lt;/h2&gt;

&lt;p&gt;Remember from earlier articles — the model is stateless. No memory between calls.&lt;/p&gt;

&lt;p&gt;So how does a chatbot remember that your name is Sham from three messages ago?&lt;/p&gt;

&lt;p&gt;You sent it. Every time the user sends a new message, you include all the previous messages too. The model sees the full conversation and can answer in context.&lt;/p&gt;

&lt;p&gt;By message 5, your request looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"system"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"You are a travel assistant."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Hi, I'm Sham."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Hi Sham! Where are you planning to travel?"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"I want to go to Goa."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Great choice! When are you planning to go?"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Next month. Any hotel recommendations?"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model knows your name, knows you're going to Goa, knows when — all because you sent the history.&lt;/p&gt;

&lt;p&gt;And because history = tokens, you manage how much you send. In production, most apps keep the last 10-15 messages. Beyond that, old messages eat tokens without adding much value.&lt;/p&gt;




&lt;h2&gt;
  
  
  Temperature — one setting worth knowing now
&lt;/h2&gt;

&lt;p&gt;Along with messages, you pass a few settings. The one you'll tune most is &lt;strong&gt;temperature&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It controls how creative or consistent the model is with its responses.&lt;/p&gt;

&lt;p&gt;Imagine a dial:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Turn it to 0&lt;/strong&gt; — the model always picks the safest, most expected word. Consistent. Predictable. Same input, same output every time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turn it to 1&lt;/strong&gt; — the model makes more adventurous word choices. More creative. More varied.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use &lt;strong&gt;0&lt;/strong&gt; when you need structure — generating SQL, extracting data, filling templates.&lt;br&gt;
Use &lt;strong&gt;0.5–0.7&lt;/strong&gt; for natural conversations — enough variation to sound human, enough consistency to be reliable.&lt;/p&gt;

&lt;p&gt;I left temperature at default for my database agent and got slightly different SQL queries for the same question on different runs. Setting it to 0 fixed it immediately.&lt;/p&gt;




&lt;h2&gt;
  
  
  When something goes wrong — how to debug
&lt;/h2&gt;

&lt;p&gt;When the model ignores your instructions or gives a strange response, there is one move that fixes 90% of issues:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Log the full messages array. Read it as if you are the model.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most of the time the bug is right there. A user message that contradicts the system prompt. A history so long the instructions are buried. A missing piece of context that makes the model guess.&lt;/p&gt;

&lt;p&gt;The model isn't broken. It responded to exactly what you sent it. Reading the full payload shows you what you accidentally told it to do.&lt;/p&gt;

&lt;p&gt;Build this habit from day one. It saves hours.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;You now understand what goes into every LLM call — where the model lives, how you connect, and the structure of what you send. The concepts are the same whether you use Gemini, GPT, Claude, or a local Ollama model.&lt;/p&gt;

&lt;p&gt;Time to write actual code. Next up: Spring AI — what it gives you over plain HTTP, and how to build your first chat endpoint in Spring Boot.&lt;/p&gt;




&lt;p&gt;If you're just starting out — try Ollama first. Free, no API key, runs on your machine. Get comfortable with the concepts, then move to a cloud model when you're ready to build something real. Drop in the comments which one you picked and what you're building.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>java</category>
      <category>programming</category>
      <category>beginners</category>
    </item>
    <item>
      <title>What Is a Token? The Concept That Unlocks Everything About LLMs</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Thu, 17 Sep 2026 14:30:00 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/what-is-a-token-the-concept-that-unlocks-everything-about-llms-1jc7</link>
      <guid>https://dev.to/shamprakash2000/what-is-a-token-the-concept-that-unlocks-everything-about-llms-1jc7</guid>
      <description>&lt;p&gt;When I first heard the word "token" in the context of LLMs, I assumed it meant words. Made sense — "I love cricket" is 3 words, so 3 tokens, right?&lt;/p&gt;

&lt;p&gt;Wrong. And that misunderstanding cost me real money before I figured out why.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is a token — really
&lt;/h2&gt;

&lt;p&gt;Let's start from the beginning.&lt;/p&gt;

&lt;p&gt;A model doesn't read text the way you do. Before it processes a single character, it breaks your text into pieces called tokens. These pieces are not words. They are not characters. They are something in between — and the exact split depends on a vocabulary the model was trained with.&lt;/p&gt;

&lt;p&gt;Let me show you with a real example. Take this sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"I love cricket"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You might expect: 3 tokens (one per word). Here's what actually happens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Text&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;I&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1 token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;love&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1 token (note: the space is part of the token)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cricket&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1 token&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Okay, so 3 tokens here. Your instinct was right this time. But try this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Microservices architecture"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Text&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Micro&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1 token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;services&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1 token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;architecture&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1 token&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's 3 tokens for 2 words. The model split "Microservices" into two pieces because it learned that &lt;code&gt;Micro&lt;/code&gt; and &lt;code&gt;services&lt;/code&gt; are more common building blocks than &lt;code&gt;Microservices&lt;/code&gt; as a whole.&lt;/p&gt;

&lt;p&gt;Now try a number:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"2024"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's 1 token. But:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"20241215"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That might be 3-4 tokens depending on the model. Long numbers get split unpredictably.&lt;/p&gt;

&lt;p&gt;And emojis? One emoji can be 2-4 tokens. A single 🚀 costs more tokens than the word "rocket."&lt;/p&gt;




&lt;h2&gt;
  
  
  How tokenisation actually works
&lt;/h2&gt;

&lt;p&gt;Every LLM uses something called a &lt;strong&gt;tokenizer&lt;/strong&gt; — a fixed vocabulary of text pieces, built during training. GPT models use a tokenizer called BPE (Byte Pair Encoding). Gemini has its own. Claude has its own.&lt;/p&gt;

&lt;p&gt;Here's how it works conceptually:&lt;/p&gt;

&lt;p&gt;The tokenizer has a vocabulary of ~50,000 to ~100,000 text pieces. Common words like "the", "is", "and" are single tokens. Less common words get split into smaller pieces. Very rare words or made-up words get broken down to almost character level.&lt;/p&gt;

&lt;p&gt;This means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Common English words → usually 1 token&lt;/li&gt;
&lt;li&gt;Long or technical words → 2 or more tokens&lt;/li&gt;
&lt;li&gt;Code → often more tokens than equivalent English (special characters, indentation)&lt;/li&gt;
&lt;li&gt;Non-English languages → often 2-3x more tokens than English for the same meaning&lt;/li&gt;
&lt;li&gt;Emojis, special characters → unpredictable, often expensive&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Is it different for GPT vs Gemini vs Claude?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Yes. Each model has its own tokenizer with its own vocabulary. The same text will produce a different token count on GPT vs Gemini vs Claude. Not dramatically different — but different. If you're comparing costs across models, you can't just copy one model's token count to another.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this matters — the context window
&lt;/h2&gt;

&lt;p&gt;Now that you know what a token is, here's why it matters.&lt;/p&gt;

&lt;p&gt;Every LLM has a &lt;strong&gt;context window&lt;/strong&gt; — a hard limit on how many tokens it can process in a single call. Think of it like RAM. The model can only "see" and work with whatever fits inside that window at once.&lt;/p&gt;

&lt;p&gt;When you make an API call, you're not just sending the user's message. You're sending:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your system prompt ("You are a helpful assistant that only answers about...")&lt;/li&gt;
&lt;li&gt;The conversation history (every message back and forth so far)&lt;/li&gt;
&lt;li&gt;Any documents you want the model to read and use&lt;/li&gt;
&lt;li&gt;The user's current message&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of that together must fit within the context window. In tokens.&lt;/p&gt;

&lt;p&gt;Let's put real numbers on this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you're sending&lt;/th&gt;
&lt;th&gt;Approximate tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;System prompt&lt;/td&gt;
&lt;td&gt;300 – 800&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Last 10 messages of conversation&lt;/td&gt;
&lt;td&gt;1,500 – 3,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 retrieved document chunks (RAG)&lt;/td&gt;
&lt;td&gt;1,500 – 3,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User's current message&lt;/td&gt;
&lt;td&gt;20 – 100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~3,300 – 6,900&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Gemini Flash has a 1 million token context window. So 6,900 tokens seems fine — and it is, for one call. But you pay for every token. And if you're not careful about what you send, those numbers grow fast.&lt;/p&gt;




&lt;h2&gt;
  
  
  What happens when you exceed the context window
&lt;/h2&gt;

&lt;p&gt;The model doesn't silently ignore the extra text. It throws an error. Your API call fails.&lt;/p&gt;

&lt;p&gt;But here's the sneaky part — you often don't hit the hard limit. Instead, you approach it gradually, and the model's quality degrades before you ever get an error.&lt;/p&gt;

&lt;p&gt;Before we get into that, let's define something you'll hear constantly when building AI systems: &lt;strong&gt;chunks&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Imagine you have a 50-page product manual as a text file. You can't send all 50 pages to the model every time a user asks a question — that's too many tokens, too slow, too expensive. So instead, you split that document into smaller pieces. Each piece is a chunk. Maybe 200-300 words each. You store all those chunks, and when a user asks something, you find the 3-5 chunks most relevant to their question and send only those to the model.&lt;/p&gt;

&lt;p&gt;That's what a chunk is — a small slice of a larger document, sized to fit comfortably inside the context window alongside everything else you're sending.&lt;/p&gt;

&lt;p&gt;Now, back to the problem.&lt;/p&gt;

&lt;p&gt;Researchers found something called the &lt;strong&gt;"lost in the middle" problem&lt;/strong&gt;. Picture this: you send the model a long prompt — system instructions at the top, then 20 document chunks in the middle, then the user's question at the bottom. The model reads all of it. But it turns out models pay more attention to what's at the very beginning and the very end of the input. The stuff buried deep in the middle? It gets less attention.&lt;/p&gt;

&lt;p&gt;Think of it like reading a very long email. You remember the opening line and the closing ask. The three paragraphs in the middle? Fuzzy.&lt;/p&gt;

&lt;p&gt;So if you send 20 chunks hoping the model finds the right answer somewhere in them — it probably won't. The answer sitting in chunk 11 might as well not exist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More context is not always better.&lt;/strong&gt; 3 highly relevant chunks will give you a better answer than 20 loosely relevant ones. Quality of what you send matters more than quantity.&lt;/p&gt;




&lt;h2&gt;
  
  
  Code costs more tokens than English
&lt;/h2&gt;

&lt;p&gt;This one surprises a lot of backend engineers.&lt;/p&gt;

&lt;p&gt;If you're building a system that processes code — code review, documentation generation, code explanation — be aware that code is significantly more expensive in tokens than regular English.&lt;/p&gt;

&lt;p&gt;Why? Because code has a lot of characters that aren't common in English — braces, semicolons, indentation spaces, underscores, camelCase names. The tokenizer wasn't primarily trained on code, so it breaks these down into smaller pieces.&lt;/p&gt;

&lt;p&gt;A rough comparison:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Content type&lt;/th&gt;
&lt;th&gt;Words&lt;/th&gt;
&lt;th&gt;Approximate tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Plain English&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;~75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Java code&lt;/td&gt;
&lt;td&gt;100 "words"&lt;/td&gt;
&lt;td&gt;~150-200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JSON payload&lt;/td&gt;
&lt;td&gt;100 "words"&lt;/td&gt;
&lt;td&gt;~120-160&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SQL query&lt;/td&gt;
&lt;td&gt;100 "words"&lt;/td&gt;
&lt;td&gt;~100-130&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So if you're sending a 500-line Java file to the model for review, you're looking at significantly more tokens than you'd expect from the line count alone. Factor this in when designing systems that handle code.&lt;/p&gt;




&lt;h2&gt;
  
  
  Non-English languages cost more too
&lt;/h2&gt;

&lt;p&gt;If you're building for users who write in Hindi, Tamil, Arabic, Chinese, or most non-English languages — tokens will cost more.&lt;/p&gt;

&lt;p&gt;The reason is the same: the tokenizer vocabulary was built primarily from English text. English words map efficiently to tokens. Non-English scripts — especially those with their own character sets — break down into many smaller pieces.&lt;/p&gt;

&lt;p&gt;Hindi text can cost 2-3x more tokens than the equivalent English meaning. This matters if you're building a multilingual product and estimating API costs.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to actually count tokens in your code
&lt;/h2&gt;

&lt;p&gt;Don't guess. Measure.&lt;/p&gt;

&lt;p&gt;Most SDKs give you a way to count tokens before making the call. In Spring AI with Gemini, you can log token usage from the response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nc"&gt;ChatResponse&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;userMessage&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;chatResponse&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// Token usage is in the metadata&lt;/span&gt;
&lt;span class="nc"&gt;Usage&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getMetadata&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getUsage&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
&lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;info&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Input tokens: {}, Output tokens: {}"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getPromptTokens&lt;/span&gt;&lt;span class="o"&gt;(),&lt;/span&gt; 
    &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getGenerationTokens&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Log this for every call in development. You'll immediately see where your tokens are going. It's the same as profiling slow SQL — you can't optimise what you haven't measured.&lt;/p&gt;




&lt;h2&gt;
  
  
  One number to remember
&lt;/h2&gt;

&lt;p&gt;If you remember nothing else from this article:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1,000 tokens ≈ 750 words ≈ a page and a half of text.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every time you construct a prompt, think in pages. If your prompt is 5,000 tokens, you're handing the model 7-8 pages of text to read before it writes a single word back to you. Is all of that necessary? That's the question to ask.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;We've covered what a token is and how the context window works. Next: what does the full payload you send to the model actually look like? System prompt, conversation history, user message — how are these structured, and how does the model use them? That's what the next article covers.&lt;/p&gt;




&lt;p&gt;Still confused about how tokenisation works for a specific case? Drop it in the comments — happy to break it down.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>java</category>
      <category>ai</category>
      <category>beginners</category>
      <category>programming</category>
    </item>
    <item>
      <title>What Is an LLM? The Foundation Every AI Backend Engineer Needs</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:39:05 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/what-is-an-llm-the-foundation-every-ai-backend-engineer-needs-50jj</link>
      <guid>https://dev.to/shamprakash2000/what-is-an-llm-the-foundation-every-ai-backend-engineer-needs-50jj</guid>
      <description>&lt;p&gt;Before I built anything with AI, I kept seeing the term LLM everywhere — in articles, in job descriptions, in GitHub repos. I nodded along like I understood it. I didn't. Not really.&lt;/p&gt;

&lt;p&gt;I knew it stood for Large Language Model. I knew ChatGPT was one. But when someone asked me "how does it actually work" — I couldn't explain it. I had the label, not the understanding.&lt;/p&gt;

&lt;p&gt;This article is the explanation I wish I had before I started building.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does "Large Language Model" actually mean
&lt;/h2&gt;

&lt;p&gt;Let's break the name down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Language&lt;/strong&gt; — it works with language. Text in, text out. It reads your words and writes words back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model&lt;/strong&gt; — it is a mathematical model. Millions (actually billions) of numbers, carefully calculated during training, that together capture patterns in language.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Large&lt;/strong&gt; — the model is enormous. GPT-4 has an estimated 1.8 trillion parameters. Gemini, Claude, Llama — all in the hundreds of billions. These numbers are what make them capable. The "large" isn't marketing — it's what separates these models from earlier, simpler ones.&lt;/p&gt;

&lt;p&gt;So: a Large Language Model is a very large mathematical model trained to understand and generate language.&lt;/p&gt;

&lt;p&gt;You've already used one. When you typed something into ChatGPT, you used GPT-4o — an LLM made by OpenAI. When you used Gemini on Google, that's an LLM. Claude (made by Anthropic), Llama (made by Meta, open source), Mistral — all LLMs. Different companies, different training data, different sizes — but the same core idea.&lt;/p&gt;

&lt;p&gt;That's the definition. But definitions don't give you intuition. Let's go deeper.&lt;/p&gt;




&lt;h2&gt;
  
  
  What is the model actually doing
&lt;/h2&gt;

&lt;p&gt;Here is the single most important thing to understand about how an LLM works:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is predicting the next word.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's it. That's the core task. You give it text, it predicts what word (technically token, but we'll get to that in the next article) comes next. Then it predicts the next one. Then the next. It keeps going until it decides to stop.&lt;/p&gt;

&lt;p&gt;Let me show you what I mean. Say you type:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The capital of France is"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model looks at those five words and predicts what comes next. "Paris" is overwhelmingly the most likely next word based on everything it learned during training. So it outputs "Paris."&lt;/p&gt;

&lt;p&gt;Now it has: "The capital of France is Paris"&lt;/p&gt;

&lt;p&gt;It predicts what comes next again. Maybe a period. Maybe "and" followed by something else. It continues predicting, one token at a time, until the response is complete.&lt;/p&gt;

&lt;p&gt;This is why LLMs can write code, answer questions, summarise documents, translate languages — all with the same underlying mechanism. It is always just: given this text, what comes next?&lt;/p&gt;




&lt;h2&gt;
  
  
  How does it "know" things
&lt;/h2&gt;

&lt;p&gt;The model didn't come with knowledge pre-loaded. It learned by reading.&lt;/p&gt;

&lt;p&gt;During training, the model was fed an enormous amount of text — web pages, books, code repositories, research papers, articles. Trillions of words. For each piece of text, it practiced the same task: predict the next word. When it got it wrong, the training process adjusted its internal numbers slightly to do better next time. This happened billions of times.&lt;/p&gt;

&lt;p&gt;After training, those billions of numbers encode patterns from all that text. When you ask it "what is the capital of France" — it doesn't look it up in a database. It pattern-matches against what it saw during training and produces the most likely answer.&lt;/p&gt;

&lt;p&gt;This is why it feels like the model "knows" things. It has absorbed patterns from an enormous amount of human writing. But it's not retrieving facts — it's generating text that fits the pattern.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why it gets things wrong — hallucination
&lt;/h2&gt;

&lt;p&gt;Here's the uncomfortable part.&lt;/p&gt;

&lt;p&gt;Because the model is predicting text rather than retrieving facts, it can generate text that sounds completely confident and is completely wrong.&lt;/p&gt;

&lt;p&gt;Ask it about a paper that doesn't exist — it might describe one in detail, with a made-up author and made-up conclusions. Ask it a question whose answer wasn't well-represented in its training data — it'll give you something that sounds right but isn't.&lt;/p&gt;

&lt;p&gt;This is called &lt;strong&gt;hallucination&lt;/strong&gt;. The model isn't lying. It's doing exactly what it was trained to do — predict plausible text. But "plausible" and "true" are not the same thing.&lt;/p&gt;

&lt;p&gt;For a backend engineer, this is critical to understand because it shapes every architectural decision you make:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You can't trust the model to know your company's internal data — it was never in the training set&lt;/li&gt;
&lt;li&gt;You can't trust it for real-time information — training has a cutoff date&lt;/li&gt;
&lt;li&gt;You can't trust it for precise facts without verification&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is exactly why RAG (Retrieval Augmented Generation) exists — you give the model the facts it needs in the prompt, so it's not guessing from training memory. But that's a later article. For now, just internalise this: &lt;strong&gt;the model generates, it doesn't retrieve.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What the model is not
&lt;/h2&gt;

&lt;p&gt;These comparisons help me think clearly when designing systems:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not a database.&lt;/strong&gt; You can't query it for a specific record. It doesn't store facts — it stores patterns. You can't ask it "what did user 123 order last week" and expect an answer unless you tell it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not a search engine.&lt;/strong&gt; A search engine finds documents that match your query. An LLM generates new text. Completely different mechanisms.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not a person.&lt;/strong&gt; It has no opinions, no emotions, no beliefs. When it says "I think" or "I feel" — that's pattern matching on how humans write, not an actual inner state. It produces text that sounds like a person because it learned from text written by people.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not always right.&lt;/strong&gt; Even when it sounds certain. Especially when it sounds certain.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this matters for what we're building
&lt;/h2&gt;

&lt;p&gt;Everything we build in this series — chat apps, RAG pipelines, AI agents — sits on top of this one thing: a model that predicts text, one token at a time, based on everything you give it in the prompt.&lt;/p&gt;

&lt;p&gt;Understanding this changes how you design:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You know you need to give it facts explicitly, not assume it knows them&lt;/li&gt;
&lt;li&gt;You know the quality of your output depends heavily on the quality of your input&lt;/li&gt;
&lt;li&gt;You know it has no state, no memory, no awareness beyond what you send in a single call&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The model is a very powerful, very fast text predictor. Your job as a backend engineer is to engineer what it predicts — by carefully constructing what you give it.&lt;/p&gt;

&lt;p&gt;That's the whole game.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Now that you know what an LLM is and how it works, the next question is: how does it measure and process the text you give it? That's where tokens and context windows come in — and that's exactly what the next article covers.&lt;/p&gt;




&lt;p&gt;What surprised you most about how LLMs actually work? Drop it in the comments.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>java</category>
      <category>beginners</category>
      <category>programming</category>
    </item>
    <item>
      <title>The Wake-Up Call: Why I Decided to Become an AI Backend Engineer</title>
      <dc:creator>Sham Prakash K</dc:creator>
      <pubDate>Sun, 13 Sep 2026 17:44:49 +0000</pubDate>
      <link>https://dev.to/shamprakash2000/the-wake-up-call-why-i-decided-to-become-an-ai-backend-engineer-1f8o</link>
      <guid>https://dev.to/shamprakash2000/the-wake-up-call-why-i-decided-to-become-an-ai-backend-engineer-1f8o</guid>
      <description>&lt;p&gt;A year ago, I was the senior engineer on the team. Today, a junior with Copilot can match my output on most tasks. Here's what I did about it.&lt;/p&gt;

&lt;p&gt;A year ago I was doing well. Four years into backend engineering — Java, Spring Boot, GraphQL, migrating monoliths to microservices, distributed systems handling 30 million requests a day with p95 latency under 10 milliseconds. I had led a team of five engineers, mentored over ten juniors, been handed performance awards consistently. By every measure I was growing.&lt;/p&gt;

&lt;p&gt;Then AI coding tools arrived. And something shifted — not dramatically, not overnight, but quietly and persistently — and I started feeling like the ground under my feet was slowly moving.&lt;/p&gt;

&lt;p&gt;This article is about that feeling, what I did about it, and the series that came out of it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The shift I noticed but couldn't name at first
&lt;/h2&gt;

&lt;p&gt;It didn't hit me as a single moment. It was a slow accumulation.&lt;/p&gt;

&lt;p&gt;Watching a junior engineer use GitHub Copilot to scaffold a Spring Boot controller in 45 seconds that would have taken me 10 minutes. Watching Claude Code refactor a service I had spent a week on — restructure the layers, rename the methods, fix the edge cases — in a single session. Seeing that the boilerplate I had spent years internalising, the patterns I knew instinctively, the "senior engineer intuition" I had built up — a lot of it was now in the model's weights.&lt;/p&gt;

&lt;p&gt;The fear most people talk about is "AI will replace developers." That's not what I felt. What I felt was more specific and more uncomfortable:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The experience gap between a senior and a junior engineer just compressed dramatically.&lt;/strong&gt; The things that took years to get fast at — scaffolding, structuring, pattern recognition — were being automated. And that trend wasn't going to reverse.&lt;/p&gt;

&lt;p&gt;I was still doing well. The systems I built were still running, the latency numbers were still good, the team was delivering. But I could see that doing the same thing for the next five years was a different bet than it had been five years ago.&lt;/p&gt;




&lt;h2&gt;
  
  
  The answers that didn't fit
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;"Learn prompt engineering."&lt;/strong&gt;&lt;br&gt;
Real skill, thin skill. It optimises how you talk to AI. The better models get at understanding natural language, the less prompt engineering matters. It doesn't compound the way technical depth does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Learn ML, become a data scientist."&lt;/strong&gt;&lt;br&gt;
Genuine option for some people. Not right for me. I like distributed systems, API design, Java, production backend work. Pivoting to Python, model training, Jupyter notebooks — that's not an upgrade, that's a restart from zero in a field where people have PhDs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Keep being a backend engineer, just use AI tools."&lt;/strong&gt;&lt;br&gt;
The most tempting option, and the one that felt like denial. Using AI tools well is already table stakes by 2026. Every backend engineer uses Copilot or Claude Code. That's not a differentiator.&lt;/p&gt;




&lt;h2&gt;
  
  
  The question that changed the framing
&lt;/h2&gt;

&lt;p&gt;At some point I stopped asking &lt;em&gt;"how do I survive AI?"&lt;/em&gt; and asked a different question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What does the backend of an AI-powered system actually look like? Who builds that?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every AI product you use — every chatbot, every AI-powered search, every document Q&amp;amp;A tool — runs on a backend. Something stores the documents and retrieves them semantically. Something routes the user's question to the right context. Something decides which tool the AI should call. Something tracks tokens and cost. Something enforces guardrails that stop the model from leaking sensitive data.&lt;/p&gt;

&lt;p&gt;That backend is not built by a data scientist. Data scientists train models. They don't design production Java APIs, tune connection pools, build MCP servers in Spring Boot, or implement rate limiting per user session.&lt;/p&gt;

&lt;p&gt;The person who builds AI backend infrastructure is a &lt;strong&gt;backend engineer who understands how AI systems work&lt;/strong&gt; — not at the model level, but at the infrastructure layer. The retrieval pipelines. The agent loops. The tool call protocols. The observability, the guardrails, the cost controls.&lt;/p&gt;

&lt;p&gt;This is called AI Backend Engineering. In 2026, the supply of people who can do it is genuinely small. Most backend engineers haven't built it. Most AI engineers don't have the production systems background. The intersection is narrow — and that's exactly where I decided to position myself.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I actually built
&lt;/h2&gt;

&lt;p&gt;Over several months — while working full-time — I built a complete AI backend system from scratch. Not a tutorial project. A real multi-component system:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/shamprakash2000/gemini-chat" rel="noopener noreferrer"&gt;gemini-chat&lt;/a&gt;&lt;/strong&gt; — Spring Boot backend with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Gemini API integration (plain HTTP first, then Spring AI — both approaches documented)&lt;/li&gt;
&lt;li&gt;Full RAG pipeline: chunking, embeddings, Pinecone vector search&lt;/li&gt;
&lt;li&gt;AI agent with ReAct loop and real tool use&lt;/li&gt;
&lt;li&gt;Database agent: natural language → SQL with safety enforcement&lt;/li&gt;
&lt;li&gt;MCP client connecting to the knowledge server&lt;/li&gt;
&lt;li&gt;Token tracking, cost monitoring, input and output guardrails&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/shamprakash2000/gemini-knowledge-mcp-server" rel="noopener noreferrer"&gt;gemini-knowledge-mcp-server&lt;/a&gt;&lt;/strong&gt; — Standalone MCP server exposing 5 tools: &lt;code&gt;listTables&lt;/code&gt;, &lt;code&gt;getTableSchema&lt;/code&gt;, &lt;code&gt;executeQuery&lt;/code&gt;, &lt;code&gt;askDocuments&lt;/code&gt;, &lt;code&gt;ingestDocument&lt;/code&gt;. LLM-agnostic — swap Gemini for Claude, the server doesn't change.&lt;/p&gt;

&lt;p&gt;Everything deployed on Render with Docker. Vector store: Pinecone. Database: Neon PostgreSQL.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this series is different
&lt;/h2&gt;

&lt;p&gt;Most AI tutorials are written in Python, use OpenAI, and are written by people who work at AI companies. They show the happy path.&lt;/p&gt;

&lt;p&gt;I'm a Java backend engineer who built this on the side. I hit Spring AI API changes between minor versions that broke my code. I debugged Pinecone dimension mismatches that silently returned wrong results with no error. I dealt with a chunking algorithm that looked correct but produced one chunk from a 500-word document. I got an 82-second response latency from the LLM and had to trace it back to 4,715 tokens of bloated context.&lt;/p&gt;

&lt;p&gt;None of that is in the tutorials. All of it is in this series.&lt;/p&gt;

&lt;p&gt;Each article covers one concept, one implementation, and the real mistakes I made. The mistakes are not edited out — they are the content.&lt;/p&gt;




&lt;h2&gt;
  
  
  Who this is for
&lt;/h2&gt;

&lt;p&gt;Backend engineers — Java, Spring Boot, any JVM language — who feel that same quiet unease and want a concrete path into AI backend work.&lt;/p&gt;

&lt;p&gt;You don't need ML knowledge. You don't need Python. You need to know how to build production APIs and be willing to learn how AI systems work at the infrastructure layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Upcoming articles covering:&lt;/strong&gt;&lt;br&gt;
Tokens, embeddings, RAG pipelines, Spring AI internals, AI agents, MCP servers in Java, observability, and production guardrails.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Next up: how LLMs actually work — not the math, not the diagrams, just the mental model that makes every other concept in this series click. If you've ever wondered what a token is, why context size matters, or why the model sometimes "forgets" things mid-conversation — that's the next one.&lt;/p&gt;




&lt;p&gt;Are you a backend engineer feeling the same shift? What's your plan? Drop it in the comments — I'm curious what other Java devs are doing.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>java</category>
      <category>ai</category>
      <category>backend</category>
      <category>career</category>
    </item>
  </channel>
</rss>
