Google is retiring gemini-2.5-pro on Vertex AI on October 20, 2026. If your product runs on it, you need to migrate before then!
I recently went through this migration practice on a legal document extraction AI-native product that runs on production. I knew going in that, it would be more than a config change, but I did not expect how many layers it would touch.
Changing the model in an AI product is very different from pointing a client at a new endpoint or upgrading a library. Each model has its own strengths and weaknesses, and its own way of filling the gaps when the input is not clear. When you swap the model, your product starts behaving differently. If you do not have the right harnesses, evals, and monitoring in place, you find out through accuracy drops, and the worst of them show up exactly where you were counting on the model's thinking capability on inferring data.
Below are the issues I ran into, in order.
Layer 1: Where the model lives
One of the most consequential decisions you have to make (not as a developer but as an organization) is where the model is being provisioned and how this impacts your data residency requirements. Gemini 2.5 Pro could be called from a single region, such as us-central1. Gemini 3.8 Flash cannot. It is offered only as a multi-region deployment (us or eu) or a global deployment, and each one uses a different endpoint shape:
| Deployment type | Example location | Endpoint |
|---|---|---|
| Single-region | us-central1 |
{region}-aiplatform.googleapis.com |
| Multi-region |
us, eu
|
aiplatform.{loc}.rep.googleapis.com |
| Global | global |
aiplatform.googleapis.com |
This difference is significant for a multinational company. A regional deployment gave you a clear answer to "where is our customer data processed?" Multi-region still keeps processing inside a jurisdiction such as the US or the EU. Global routes each request to wherever Google has capacity, so there is no data residency guarantee. If your contracts, regulators, or compliance team care about residency, you need to discuss the model location with them before you change it.
The caveats of moving to multi-regional endpoint
If your application is deployed on a cloud run with private VPC sharing and choose to move forward with a multi-region deployment option, There are chances that your first call fails with the following error:
[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: Hostname mismatch,
certificate is not valid for 'aiplatform.us.rep.googleapis.com'. (_ssl.c:1016)
The code is fine but the networking rules.. not so much. The Cloud Run service sent its egress through a centrally managed Shared VPC subnet with Private Google Access enabled. That setup works for regional and global endpoints, but not for multi-region ones. Multi-region endpoints only became generally available on May 15, 2026, so many existing network setups were never designed with them in mind. Google documents the limitation in About accessing the Vertex AI API:
Private Google Access isn't supported for multi-region endpoints. If you attempt to connect to a multi-region endpoint using Private Google Access, you might experience connectivity issues, SSL/TLS handshake errors, or certificate mismatch warnings.
To establish private connectivity to multi-region endpoints, you must configure Private Service Connect endpoints for regional Google APIs.
The proper fix is a Private Service Connect endpoint for the regional Google APIs. When a central team owns networking, that is a medium-to-large change with its own timeline.
An interim workaround is to use the global deployment, which its endpoint works with the existing network. It has two trade-offs you should know up front:
- There is no data residency guarantee, for the reasons above.
- There is no context caching. The global endpoint does not support it, so make sure your code is not relying on the caching mechanism or if it does currently, you have implemented fail-safe approaches.
If you are planning this migration, check your egress path before you write any code.
Layer 2: The API contract changed
The syntax for passing tools is the same between the two models, which makes it easy to assume nothing else changed. Several things did.
Thinking configuration. Gemini 2.5 Pro takes a numeric thinking_budget. Gemini 3.8 Flash rejects it and expects a thinking_level enum: LOW, MEDIUM (default), or HIGH. MINIMAL is not supported on 3.8 Flash and fails API validation (unlike Gemini 3.5 Flash).
Sampling parameters. The Gemini 3 family drops the legacy sampling options. Remove temperature, top_p, and top_k from your GenerateContentConfig. candidate_count is not supported at all from Gemini 3.x onward. A lot of extraction code hardcodes a low temperature to get "deterministic" output, and Google calls out that exact pattern in its Gemini 3 prompting guide:
If your existing code explicitly sets temperature (especially to low values for deterministic outputs), we recommend removing this parameter and using the Gemini 3 default of 1.0... Changing the temperature (setting it below 1.0) may lead to unexpected behavior, such as looping or degraded performance, particularly in complex mathematical or reasoning tasks.
In code, the change looks like this:
from google.genai import types
# Before: gemini-2.5-pro
config = types.GenerateContentConfig(
temperature=0.1,
top_p=0.95,
thinking_config=types.ThinkingConfig(thinking_budget=8192),
)
# After: gemini-3.8-flash
config = types.GenerateContentConfig(
thinking_config=types.ThinkingConfig(thinking_level="MEDIUM"),
)
Changing the syntax is the easy part. A budget is a number you tune, but a level is a product decision. You now have to decide which flows deserve HIGH and which are fine on MEDIUM or LOW. Each choice changes model behavior, so each one needs its own accuracy validation.
Function calling. If your application relies on them, there are three differences to check:
- If you orchestrate multi-turn tool use yourself, every
FunctionResponseyou send back must include both acall_idand aname. - If your prompt places inline instructions directly before a tool call, 2.5 Pro handles it, but 3.8 Flash can return
MALFORMED_FUNCTION_CALL. Separate those instructions with clean blank lines (\n\n). - If your tools return images or audio, 3.8 Flash expects the assets inside the response payload rather than referenced through an external pointer.
Multi-turn chats. 3.8 Flash enforces server-side state tracking through previous_interaction_id, and it restricts manually pre-filling model turns in the payload history more than 2.5 did.
Context caching minimum. Once you are on a regional or multi-region endpoint where caching works, note that the minimum cache size goes from 1,024 tokens on 2.5 Pro to 4,096 tokens on 3.8 Flash. Short inputs, such as a single small document, will no longer trigger caching. There is no code change, but your cost profile changes, so measure it on real traffic instead of estimating it. See Create a context cache.
Layer 3: The model thinks differently
This part is not in any release note, and it is where the feel the migration can change your product (or its taste).
Both models are strong thinkers, but they approach the same problem differently, much like two experienced people would. Part of that difference is the Pro and Flash tiers being designed for different jobs, and part of it is the generation. From the product's point of view, the cause matters less than the effect.
Gemini 2.5 Pro is the more verbose of the two. It is comfortable with long reasoning chains, and it likes to connect dots, sometimes ones that are not there. In one case it inferred a transaction date from the date of the email chain the document came from. In another, a contract's start date depended on a condition under a separate agreement being met. 2.5 Pro took the start date at face value, added the contract term, and returned a confident end date.
Output format was the other difference. Our response schema is referenced in the prompt rather than enforced through function calling. With 2.5 Pro, this approach gave better results than enforcing the schema, which left us with many null records in the output. The cost was that 2.5 Pro sometimes drifted from the requested JSON shape.
Gemini 3.8 Flash on MEDIUM is more concise and less inclined to overthink. In the same conditional-contract case, it noticed that the start date depended on an event that had not happened yet, and it did not assume either date. When the evidence was not clear, it left the field null, which is exactly what our instructions asked for. It followed the output conventions in the prompt much more cleanly, kept to the JSON shape from the prompt-referenced schema without being forced by function calling, and did not do more than it was asked.
Which behavior is better depends on the product. For legal extraction, an inferred date that looks authoritative is worse than an honest null. For another product, the 2.5 Pro habit of connecting dots might be exactly what you want.
What this taught me about choosing a model
Every thinking model has its own personality. A prompt tuned for one model is not tuned for the next, and the habits you build into your prompts, such as how firmly you state the output format or how you tell the model to handle missing data, can have a different effect on another model. In my experience, even the same model can behave differently depending on which cloud serves it, since providers can wrap your input with their own system-level instructions before it reaches the model.
That changes how I think about model selection. Tuning prompts and validation around a model takes real effort, so choose a model with enough lifespan to make that effort worth it. Be careful with model routers in precision work like legal document extraction, because the same document can get a different interpretation depending on which model answered it. This is also why some teams invest in smaller models they host themselves and fine-tuned for specific task.
How to actually learn a model
Release notes list what changed in the API, but they cannot show you how the model handles your own documents.
One of the most effective ways I have found is to build an eval set for your application along with a set of standard baselines. With that in place, you can run different models, and the same model at different thinking levels, against the same tasks and compare the results directly. You can also see where each model is weak and where you can add quality through context management, prompt management, or tool calling, instead of guessing.
I will write about my approach to eval testing in my next post.



Top comments (0)