DEV Community

Cover image for I Hit a 429 Before I Got a Bill — Managing Gemini API Limits in Spring Boot
Sham Prakash K
Sham Prakash K

Posted on AI-assisted

I Hit a 429 Before I Got a Bill — Managing Gemini API Limits in Spring Boot

I didn't get a surprise bill. I got a 429.

Requests started failing. The app was returning errors instead of responses. I opened the Gemini AI Studio rate limit page and saw it — quota exhausted. I'd burned through my daily request limit without realising it.

That's when I understood that token usage isn't abstract. It's a real counter ticking up with every API call, and if you're not watching it, you'll hit the wall.

This article is about understanding what you're spending, where it goes, and how to stay within limits — especially on the free tier.

What rate limits actually are

When you use the Gemini API on the free tier, you're not billed — but you're not unlimited either. Google enforces three types of limits:

RPM — Requests Per Minute. How many API calls you can make in a 60-second window.

TPM — Tokens Per Minute. How many tokens (input + output combined) you can send and receive per minute.

RPD — Requests Per Day. Total API calls allowed in a 24-hour window.

When you exceed any of these, the API returns a 429 Too Many Requests error. Your app stops working until the window resets.

The free tier limits are lower than you'd expect — and they go fast when you're testing. Check the current limits for your model on the Gemini API rate limits page in AI Studio since they change as Google updates the free tier.


Why the limits go fast

When I first hit the limit I thought I hadn't made that many calls. I was wrong — I just hadn't understood what each call actually costs.

Problem 1: Every call sends the full history.

In Article 6 we built conversation memory. Every time you send a message, Spring AI includes the last 20 messages in the request. That means message 20 doesn't cost 1 message worth of tokens — it costs up to 20 messages worth.

A conversation with 10 back-and-forth exchanges might look like 10 API calls. But the token cost is:

  • Call 1: 1 message
  • Call 2: 3 messages (1+2)
  • Call 3: 6 messages (1+2+3)
  • ...
  • Call 10: 55 messages worth of tokens

Problem 2: Your system prompt runs on every call.

Every request includes the system prompt. If your system prompt is 200 words, that's ~270 tokens added to every single API call — silently.

Problem 3: Output tokens cost too.

The model's reply counts against your token budget. A long detailed answer costs more than a short one. By default, the model will be as verbose as it wants to be.


See exactly what you're spending

The fix starts with visibility. Spring AI exposes token usage on every response — you just have to ask for it.

Change your chat method to use chatResponse() instead of content():

@PostMapping("/chat-ai")
public String chat(@RequestBody Map<String, String> request) {
    ChatResponse response = chatClient.prompt()
        .user(request.get("message"))
        .advisors(a -> a.param("chat_memory_conversation_id", request.get("conversationId")))
        .call()
        .chatResponse();   // ← get the full response object, not just the text

    // Log token usage
    Usage usage = response.getMetadata().getUsage();
    log.info("Tokens — input: {}, output: {}, total: {}",
        usage.getPromptTokens(),
        usage.getGenerationTokens(),
        usage.getTotalTokens());

    return response.getResult().getOutput().getText();
}
Enter fullscreen mode Exit fullscreen mode

Now every API call logs something like:

Tokens — input: 847, output: 312, total: 1159
Tokens — input: 1203, output: 198, total: 1401
Tokens — input: 1589, output: 445, total: 2034
Enter fullscreen mode Exit fullscreen mode

Watch the input token count grow with each message in a conversation. That's the history accumulating.


Track a running total per session

Once you can see individual call usage, it's useful to track cumulative usage per session. A simple in-memory counter works fine for this:

private final Map<String, Integer> sessionTokens = new ConcurrentHashMap<>();

@PostMapping("/chat-ai")
public String chat(@RequestBody Map<String, String> request) {
    String conversationId = request.get("conversationId");

    ChatResponse response = chatClient.prompt()
        .user(request.get("message"))
        .advisors(a -> a.param("chat_memory_conversation_id", conversationId))
        .call()
        .chatResponse();

    Usage usage = response.getMetadata().getUsage();
    int callTokens = usage.getTotalTokens();

    int runningTotal = sessionTokens.merge(conversationId, callTokens, Integer::sum);

    log.info("conversationId: {} | this call: {} tokens | session total: {} tokens",
        conversationId, callTokens, runningTotal);

    return response.getResult().getOutput().getText();
}
Enter fullscreen mode Exit fullscreen mode

After a few messages you'll see the session total climbing. This makes the accumulation visible and gives you something to act on.


How to reduce token usage

Now that you can see what you're spending, here's where to trim.

1. Keep your system prompt short.

Every extra word in your system prompt costs tokens on every single call. Go through it and cut anything that isn't load-bearing.

Before:

You are a helpful and friendly AI assistant. You are knowledgeable about many topics 
and always try to give thorough, well-explained answers. You are patient and 
understanding. You never refuse to help with reasonable questions. You always 
maintain a positive and encouraging tone in all your responses.
Enter fullscreen mode Exit fullscreen mode

After:

You are a helpful assistant. Be concise and practical.
Enter fullscreen mode Exit fullscreen mode

Same behaviour, a fraction of the tokens — multiplied across every API call you make.

2. Limit conversation history.

We already set maxMessages(20) in Article 6. That's a good default. For development and testing, drop it lower:

MessageWindowChatMemory memory = MessageWindowChatMemory.builder()
    .chatMemoryRepository(memoryRepository)
    .maxMessages(10)   // ← 10 messages instead of 20 while testing
    .build();
Enter fullscreen mode Exit fullscreen mode

Older messages in a long conversation rarely affect the current reply anyway.

3. Tell the model to be concise.

The model's default verbosity is tunable via the system prompt:

You are a helpful assistant. Be concise — answer in 2-3 sentences unless the user 
asks for more detail.
Enter fullscreen mode Exit fullscreen mode

A shorter reply means fewer output tokens. Output tokens count against the same limits as input tokens.

4. Don't test with long conversations.

During development, start a new session for every test instead of continuing an existing conversation. A fresh session has zero history — each call costs only the system prompt + one message, not the accumulated history of 20 exchanges.


Beyond maxMessages — how production apps handle long conversations

You've been using maxMessages(20) since Article 6. That's the simplest strategy — and it has a name: sliding window. But it's not the only approach. Here's how the strategies stack up as conversations get longer and requirements get harder.


Sliding window — what you already have

Keep only the last N messages. Older messages are dropped. Simple, predictable, zero extra API calls.

Works well for most chat apps where conversations are short-to-medium. The downside: the model loses context from early in the conversation. If the user mentioned their name in message 1 and you're now on message 25, the model has forgotten it.


Summarization

When history hits a threshold, ask the model to summarize the older messages into a compact paragraph. Replace those messages with the summary. Then continue with summary + recent messages.

if (history.size() > 15) {
    // Ask the model to summarize older messages
    String oldMessages = formatMessages(history.subList(0, 10));

    String summary = chatClient.prompt()
        .user("Summarize this conversation in 4-5 sentences, keeping key facts and decisions: "
              + oldMessages)
        .call()
        .content();

    // Replace 10 old messages with one summary message
    List<Message> compressed = new ArrayList<>();
    compressed.add(new SystemMessage("Earlier conversation summary: " + summary));
    compressed.addAll(history.subList(10, history.size()));
    history = compressed;
}
Enter fullscreen mode Exit fullscreen mode

2000 tokens of raw history becomes 150 tokens of summary. The model keeps the gist without the full detail. This is what Claude Code does — the summary you see at the start of a long session is exactly this pattern applied automatically.

The tradeoff: one extra API call to summarize, and fine-grained detail from early messages is lost. For most apps that's acceptable.


RAG-based memory (coming later in this series)

Instead of sending recent messages, store every message as an embedding in a vector database. When a new message arrives, search for the most relevant past messages — not just the most recent ones.

This means a conversation from last week about a specific topic gets retrieved when relevant, even if it's buried under hundreds of other messages. We'll build this when we cover RAG.


Structured memory extraction (coming later in this series)

Instead of storing raw messages, the system extracts facts from the conversation and stores them separately:

{
  "user_name": "Sham",
  "project": "Spring Boot AI chat app",
  "decisions": ["Gemini API", "Neon PostgreSQL", "deployed on Render"]
}
Enter fullscreen mode Exit fullscreen mode

The model receives this fact store + recent messages — not raw history at all. It's how products like ChatGPT's memory feature work. Much more token-efficient for long-running applications.


For now, sliding window (maxMessages(20)) handles most cases. Add summarization when your conversations consistently hit the limit. RAG and structured memory come later in the series when we have the foundations for them.


When to upgrade

The free tier is enough to learn and build demos. You'll hit limits when:

  • You're testing heavily (many short test sessions in the same day)
  • Your conversations get long (history accumulates fast)
  • Multiple people are using the app simultaneously

When you do upgrade to a paid tier, the rate limits increase significantly and you're charged per million tokens instead. At that point the logging you built here becomes essential — it's how you know what you're actually paying for.

For now, the key habits are: log every call, watch the input token growth, keep your system prompt lean, and use fresh sessions during testing.


What's next

The app is efficient now. Next: writing system prompts that actually shape the model's behaviour — not just what it says, but how it thinks about your problem.


Hit a 429 while building? Drop it in the comments — you're definitely not the only one.

Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.

Top comments (0)