Google Challenges NVIDIA’s AI Business with In-House Chips*Why TPUs are not simply faster GPUs, why Google is both a customer and a competitor to NVIDIA, and how the market for AI data centers could evolve.*
NVIDIA reported revenue of $215.9 billion for its fiscal year ending late January 2026.
$193.7 billion of that total was generated by NVIDIA’s Data Center segment. In this context, the name describes the target market, not the operator: NVIDIA does not run its customers’ data centers, but instead provides a major portion of their underlying hardware and technical infrastructure. This includes GPUs, networking components, full server systems, and associated software. Operators and customers include Google, Amazon, Microsoft, Oracle, Meta, as well as specialized cloud providers. [1]
Google Cloud also offers NVIDIA systems to its customers. At the same time, Google has been developing its own proprietary AI accelerator for over a decade: the Tensor Processing Unit (TPU).
This makes Google both a customer and a competitor.
The company needs NVIDIA hardware to meet current demand for AI compute capacity. However, with each new TPU generation, Google can offload a larger share of that workload onto its own platform — initially for Search, YouTube, Ads, and Gemini, and now increasingly for external cloud customers as well.
This dual path illustrates how the market could shift. NVIDIA won’t disappear overnight from data centers, but new AI capacity no longer has to be built using NVIDIA GPUs by default.
Why AI Runs on GPUsA computer chip consists of billions of transistors — tiny electronic switches that form computing circuits and memory. How these circuits are arranged determines which tasks a processor excels at.
The Central Processing Unit (CPU) is a computer’s versatile primary processor. It launches applications, manages memory, responds to user inputs, and handles complex logic with many branch decisions and varied execution steps.
A Graphics Processing Unit (GPU) was originally designed for graphics rendering.
When generating an image, positions, colors, and lighting must be calculated for millions of pixels. Many of these calculations are virtually identical and can be executed simultaneously.
While a CPU has a relatively small number of high-performance cores, a GPU distributes work across thousands of smaller processing units. This allows it to process massive volumes of similar numerical operations in parallel.
This exact capability is what neural networks require.
The learned knowledge of a neural network is stored in vast arrays of numerical values known as weights. During training and inference, large matrices of numbers are multiplied together and their products summed up.
These matrix operations map naturally to the parallel computing units of a GPU.
Thus, a graphics processor evolved into an AI accelerator.
Training and Inference Are Two Distinct TasksDuring training, an AI model learns. It processes training data, compares its output against a desired result, and subsequently updates its weights. This cycle repeats billions of times.
Consequently, training involves much more than generating an answer. The system must track errors backward through the network, compute gradient updates for weights, and exchange these updates across hundreds or thousands of interconnected chips. Large training runs therefore require flexible compute units, ultra-fast interconnects, and a mature software ecosystem.
During inference, the trained AI model is applied to real-world inputs. For instance, a language model receives a query and calculates a response using its existing weights. The weights themselves are not updated during this process.
In an active production environment, this workflow repeats for millions of users:
Receive input
→ Calculate output using existing weights
→ Return responseThis repetitive and highly predictable workload is ideally suited for custom, domain-specific hardware. When serving millions or billions of requests, every incremental reduction in power consumption and cost per response matters immensely.
This is why inference in particular could shift faster toward TPUs and other custom accelerators, whereas training may rely longer on NVIDIA’s greater flexibility and mature ecosystem.
CUDA Turned the GPU into a PlatformHardware alone would not have made NVIDIA the undisputed market leader.
With CUDA, NVIDIA created a software platform that allowed programs to execute non-graphics workloads directly on GPU cores.
CUDA is far more than a basic API; it encompasses compilers, optimized mathematical libraries, debugging tools, and tailored implementations of common algorithmic routines.
AI frameworks like PyTorch leverage this underlying foundation, sparing developers from having to program low-level matrix multiplications directly for the hardware.
Over many years, a powerful ecosystem emerged:
NVIDIA GPU
→ CUDA
→ Optimized AI libraries
→ PyTorch & other frameworks
→ Tooling for development & ops
→ Experienced developer poolModern NVIDIA GPUs also feature dedicated Tensor Cores — specialized hardware modules explicitly designed for the mathematical operations most common in neural networks. [2]
As a result, NVIDIA no longer simply sells graphics chips that happen to process AI workloads. It delivers specialized AI GPUs, high-speed interconnects, advanced networking tech, and fully integrated server racks.
Google Buys NVIDIA Tech — and Builds the ReplacementGoogle Cloud offers both NVIDIA GPUs and Google’s in-house TPUs. As recently as late 2025, Google announced new cloud clusters featuring NVIDIA’s GB300 platform while simultaneously preparing Ironwood, its seventh-generation TPU. [3]
This is not a contradiction.
Google demands immense volumes of compute capacity. Its internal chips cannot instantly absorb every workload, nor are they available in unlimited quantities. NVIDIA provides a mature, turnkey platform that many enterprise customers and models already support natively.
At the same time, Google has a compelling financial incentive to build its own alternative.
Search, YouTube, Ads, and Gemini generate continuous, large-scale AI workloads. If a proprietary chip makes an individual calculation even marginally cheaper or more energy-efficient, those tiny savings compound across hundreds of billions of queries.
Developing custom silicon also reduces vendor lock-in. Google can co-design chips, software, and data centers in tandem, tailoring future hardware directly to its proprietary model architectures.
The TPU did not begin as a direct head-to-head assault on NVIDIA; it started as an internal engineering solution to manage Google’s own cost structure and compute requirements.
Through Google Cloud, it evolved into an enterprise product for external customers.
What Makes a TPU Different*TPU* stands for Tensor Processing Unit.
A tensor is a multi-dimensional array of numbers — a vector is a one-dimensional tensor, while a matrix is a two-dimensional tensor.
The name highlights the chip’s core focus. Google designed the TPU from day one specifically for machine learning. Massive matrix multiplication units, high-bandwidth memory access, and inter-chip communications are custom-engineered for these exact workloads. [4]
A TPU does not contain hardcoded AI models. It is a programmable hardware accelerator capable of training and running a wide variety of supported neural architectures.
The fundamental difference between GPUs and TPUs lies in their origin stories:
GPU:
Flexible parallel processor
→ Later heavily specialized for AI
TPU:
Purpose-built for machine learning from day one
→ Later broadened for more models & workloadsToday, this boundary is blurring.
NVIDIA builds dedicated AI modules (Tensor Cores) into its GPUs, while Google expands its TPUs to handle wider ranges of model architectures, precision formats, and mathematical operations.
One platform is becoming more specialized, while the other is becoming more versatile.
TPUs Live in Racks TooA TPU chip does not operate in isolation.
A complete TPU system requires standard host CPUs, system memory, network interconnects, power delivery, and specialized cooling infrastructure. Google links large numbers of TPU chips into clusters known as “Pods,” which span multiple physical server racks.
In practice, a large AI model is split across hundreds or thousands of interconnected chips. During computation, these chips must continuously exchange intermediate results. Consequently, memory bandwidth and interconnect speed dictate overall performance just as much as raw compute power.
Google provides TPU capacity via Google Cloud, meaning customers do not need to purchase or physically host the hardware themselves. Instead, they provision specific cluster configurations for discrete training runs or ongoing inference services. [4][5]
Because of this model, Google isn’t competing with NVIDIA on a chip-by-chip basis.
Both companies deliver end-to-end platforms:
Accelerators
→ Memory architecture
→ High-speed interconnects
→ Server racks & infrastructure
→ Liquid/air cooling
→ Compilers & framework libraries
→ Cloud operational layersModels Don’t Need to “Re-learn” When Switching HardwareThe weights of a neural network are simply stored numeric parameters. They are not permanently bound to the specific GPUs on which the model was originally trained.
Therefore, a model can be trained on GPUs and subsequently deployed for inference on TPUs. Likewise, training state from a TPU run can be saved and resumed on alternative supported hardware.
A model does not need to be retrained from scratch simply because the underlying chip vendor changes.
The critical distinction is between model parameters and the surrounding software environment:
Model weights are portable. The surrounding software stack may not be.High-level AI frameworks like PyTorch and JAX define mathematical operations abstractly. Underlying compilers and runtime engines then translate those operations for specific hardware backends.
For TPUs, JAX and PyTorch/XLA serve as these translation layers. If a model relies strictly on standard, well-supported operations, porting it can be relatively straightforward. [6]
Migration becomes significantly more complex if a codebase:
- Uses custom CUDA kernels,
- Relies on NVIDIA-specific proprietary libraries,
- Has been hand-tuned for a specific GPU cache or interconnect architecture,
- Or ties its monitoring and telemetry pipeline directly to NVIDIA tooling. In those scenarios, portions of the software stack must be rewritten or re-architected.
An AI model doesn’t need to learn again — but engineers may need to reconfigure the infrastructure supporting it.
Migration Happens Workload by WorkloadNo enterprise data center swaps out its entire GPU footprint for TPUs over a single weekend.
Instead, an operator typically spins up a parallel model instance on TPUs to evaluate key metrics:
- Does output quality remain identical?
- What is the end-to-end response latency for users?
- How many concurrent requests can the system handle?
- How stable is the service under sustained load?
- What is the total cost per million input and output tokens?
- What are the real-world energy and cooling requirements? If performance and unit economics prove compelling, a small fraction of live traffic is routed to the TPU cluster. As confidence grows, that share expands over time.
GPU and TPU infrastructures can coexist peacefully within the same organization for years. Furthermore, training and inference do not need to run on identical hardware.
What cannot easily be done is swapping individual GPUs for TPUs inside an active, running training cluster. The underlying chips within a single cluster must be tightly synchronized and share an identical software and network environment.
Migration therefore occurs service by service or model by model — not component by component.
A TPU Isn’t Faster by DefaultThe real-world speed of an AI system depends on far more than nominal chip specs.
In training, the key metric is how long a model takes to reach a target validation loss or quality threshold. In inference, user latency and concurrent throughput take center stage.
A given cluster might deliver high aggregate throughput while exhibiting higher latency on individual requests. Another setup might offer lower latency but carry a higher hourly rental cost.
Theoretical compute figures (e.g., peak FLOPS) can also be deceptive. Data must be streamed efficiently from memory, and hundreds of chips must exchange intermediate tensors without bottlenecking. If memory bandwidth or network interconnects saturate, raw compute units sit idle.
A meaningful performance benchmark requires a strictly defined task:
Identical model architecture
→ Identical output quality threshold
→ Identical maximum latency allowance
→ Total end-to-end system cost
→ Energy consumed per valid outputWithout these constraints, claims like “3x faster” are largely meaningless.
For enterprise buyers, peak theoretical compute capacity is irrelevant compared to the actual cost per useful training hour or production inference call.
NVIDIA’s Business Is Transparent — Google’s TPU Revenue Is ObscuredNVIDIA’s $193.7 billion in annual Data Center revenue provides a clear metric of its market dominance and commercial reach. [1]
Google, by contrast, does not break out standalone TPU revenue in its financial reporting.
TPUs generate value across multiple channels:
- Reducing operational costs for internal Google services,
- Generating direct cloud rental revenue via Google Cloud Platform,
- Enabling Google to sell fully integrated AI infrastructure bundles,
- And, as of 2026, selling turnkey TPU systems directly as product offerings. [7] Google Cloud reported total revenue of $58.7 billion for 2025 — a figure that combines AI compute with Workspace subscriptions, databases, storage, and general cloud services.
Parent company Alphabet reported $91.4 billion in capital expenditures for 2025, directed primarily toward technical infrastructure. This sum encompasses GPUs, TPUs, servers, networking gear, real estate, and data center facilities combined. [8]
As a result, precise TPU market share figures cannot be calculated directly from public filings.
This financial structure complicates direct comparisons: NVIDIA profits primarily from selling hardware infrastructure, whereas Google can capture ROI simply by lowering the operational unit economics of Search, Ads, or Gemini.
Anthropic Turns TPUs into a Major Non-Google BusinessFor years, TPUs were viewed primarily as an internal Google technology stack.
Anthropic is changing that narrative.
The creator of Claude finalized multi-gigawatt agreements with Google and Broadcom to secure next-generation TPU capacity starting in 2027. [9]
According to Reuters, Anthropic’s financial disclosures detail a five-year commitment totaling $125.2 billion for TPU compute capacity. [10]
This does not mean Anthropic is abandoning NVIDIA entirely — leading AI laboratories routinely diversify their compute footprints across multiple hardware vendors and cloud providers.
However, a commitment of this scale confirms that an external AI company is anchoring a massive share of its core future infrastructure on Google’s TPU platform.
Google has successfully transitioned the TPU from an internal custom chip into a mainstream enterprise platform upon which leading AI firms build multi-year strategy.
The Real Competition Happens at the Next Expansion PhaseGoogle does not need to rip and replace existing NVIDIA server racks in established data centers.
Demand for AI compute continues to grow rapidly. Companies are building new facilities, expanding existing footprint, and scaling out capacity to handle rising inference volumes.
Every infrastructure expansion presents a fresh purchasing decision.
An operator can maintain its legacy GPU footprint for core training while routing new inference expansion onto TPU pods. A new model generation can be trained on a different hardware platform than its predecessor. A cloud customer can shift specific microservices without migrating its entire IT stack.
Consequently, market evolution will likely follow this pattern:
Existing NVIDIA clusters remain in production
→ New hardware platforms are introduced
→ Workloads are distributed dynamically
→ Relative market share shifts gradually over timeThis dynamic is critical for NVIDIA, even if its top-line revenue continues to grow in absolute terms.
In a rapidly expanding TAM (Total Addressable Market), two things can happen simultaneously:
- NVIDIA sells more total hardware year over year.
- NVIDIA’s percentage share of global AI compute capacity declines. For this to happen, overall market demand simply needs to grow faster than NVIDIA’s supply capacity.
Inference Could Accelerate Market FragmentationTraining frontier models requires maximum hardware flexibility, high-bandwidth interconnects, and mature software tooling — areas where NVIDIA retains substantial structural advantages.
Inference, however, works with fixed weights. Production services execute highly predictable, repetitive matrix math across massive volumes of incoming requests. A specialized architecture can optimize hardware utilization for these specific pipelines while shedding unnecessary operational overhead.
Even a minor percentage reduction in cost-per-query yields massive savings when applied to global Search, advertising, translation, or chat operations.
Google holds a unique structural advantage here: it can stress-test and refine new TPU generations against its own internal production services before offering them externally. It does not need to wait for enterprise cloud customers to validate first-run silicon.
This makes the following trajectory likely:
Frontier model training may remain GPU-dominant for longer, while the expanding inference market fragmenting more rapidly among custom accelerators.This is not a guaranteed outcome. New TPU iterations continue to close the gap in training performance, and NVIDIA continues to aggressively optimize its architectures for inference efficiency.
However, the economic return on specialized silicon is highest in high-volume, repetitive inference tasks.
Google Is Not the Only Challenger“TPU” is simply Google’s brand name for its proprietary silicon.
Amazon continues to iterate on Trainium (for training) and Inferentia (for inference). Microsoft is scaling its Azure Maia accelerators. Other hyperscalers and semiconductor startups are pursuing similar custom silicon strategies. [11][12]
Their underlying motivations are identical:
- Reducing reliance on a single vendor,
- Gaining control over hardware costs and supply chain lead times,
- Co-designing silicon alongside custom data center power and cooling systems,
- And offering proprietary hardware that locks customers deeper into their cloud platforms. NVIDIA’s long-term competition comes not just from competing semiconductor foundries, but from its largest end customers.
Hyperscalers are increasingly designing their own replacements for the hardware they previously purchased off the shelf.
This shift can erode NVIDIA’s pricing power even before unit sales drop. A buyer with a viable in-house alternative enters price and supply negotiations with significantly more leverage.
Why CUDA Remains NVIDIA’s Strongest MoatMigrating away from NVIDIA involves costs far beyond hardware rental rates.
Engineering teams must refactor code, validate numerical precision, rebuild monitoring pipelines, and potentially run dual-stack infrastructures during transitional phases.
In this respect, CUDA protects NVIDIA much like a dominant desktop operating system protects its application ecosystem.
The more heavily an organization relies on custom CUDA code and low-level NVIDIA primitives, the higher its switching costs. Conversely, as high-level compilers, PyTorch, and JAX abstract away hardware specifics, those switching costs decline.
Competition between hardware platforms ultimately plays out across three dimensions:
Raw chip performance
→ Total system efficiency & unit economics
→ Developer ergonomics & migration frictionA TPU cluster may offer superior theoretical cost-efficiency yet still lose an enterprise deal if porting the customer’s software requires excessive engineering hours.
Conversely, a slightly slower chip can win if it is readily available in volume, consumes less power, or delivers lower total cost of ownership (TCO) for a specific production service.
Will NVIDIA Build Its Own TPU?NVIDIA does not need to clone Google’s exact hardware architecture.
Through Tensor Cores, NVIDIA has already embedded specialized matrix units into its GPUs while preserving their broader general-purpose programmability.
NVIDIA’s architectural roadmap will likely continue refining this hybrid design:
- Increasingly specialized matrix execution units,
- Higher memory bandwidth architectures,
- Faster chip-to-chip interconnect topologies,
- Enhanced energy efficiency per FLOP,
- And tighter physical integration into full-rack systems. NVIDIA’s answer to the TPU is unlikely to be a standalone, stripped-down “TPU clone.”
Instead, it will be a GPU platform that grows increasingly specialized for AI workloads without abandoning the general programmability that created its ecosystem in the first place.
A Balanced OutlookNVIDIA will not lose its market leadership position overnight. Its massive installed base, CUDA ecosystem, developer mindshare, and end-to-end product catalog form a formidable competitive moat.
Furthermore, Google must demonstrate that TPUs can be deployed reliably, predictably, and easily for a wide range of third-party enterprise customers outside its own controlled environment.
Nevertheless, the era of a near-monolithic NVIDIA market is drawing to a close.
Hyperscalers are committing tens of billions of dollars to custom AI infrastructure — investments that pay off even if internal chips are only deployed across targeted internal workloads. Custom silicon doesn’t need to outperform NVIDIA across every metric; it only needs to run a growing share of internal and rented workloads more cost-effectively.
For the coming years, the following market trajectory is highly plausible:
- NVIDIA remains the leading independent provider of AI accelerators.
- Google, Amazon, and Microsoft deploy an increasing share of custom silicon across their clouds.
- Existing GPU clusters remain active, but new compute capacity is distributed across a broader mix of platforms.
- Inference markets open up to specialized silicon faster than the market for training top-tier frontier models.
- NVIDIA can continue to grow total revenue while losing overall market share in global AI compute.
- Increased competition will impact hardware pricing margins before it impacts absolute shipment volumes. Ultimately, no single chip architecture will displace all others.
The market is shifting toward a heterogeneous model where organizations deploy different hardware platforms depending on model architecture, latency requirements, and unit economics.
NVIDIA currently sets the baseline against which all alternative hardware is measured.
With its TPU program, Google has built a mature alternative that is no longer restricted to internal workloads. As major AI firms commit tens of billions of dollars to this platform, technical competition is translating into direct commercial competition.
NVIDIA does not own the data centers.
Historically, however, it supplied the fundamental hardware without which modern AI infrastructure could not be built.
Google’s goal is not to replace every NVIDIA GPU on the market.
It is enough if, during the next expansion cycle, buyers choose a TPU a little more often.
Sources- NVIDIA: Fiscal 2026 Financial Results
- NVIDIA: Tensor Cores Architecture
- Google: Alphabet Earnings Q3 2025 — NVIDIA GB300 and Ironwood
- Google Cloud: TPU System Architecture Documentation
- Google Cloud: Tensor Processing Units Overview
- PyTorch: PyTorch on XLA Devices Guide
- Alphabet: Quarterly Report Q2 2026 (SEC Form 10-Q)
- Alphabet: Annual Report 2025 (SEC Form 10-K)
- Anthropic: Expanded Partnership with Google and Broadcom for Next-Generation TPU Capacity
- Reuters: Broadcom to Lend Anthropic Up to $42 Billion to Lease Its Chips
- Amazon Web Services: AWS Trainium Accelerators
- Microsoft: Azure Maia Custom AI Silicon
Top comments (0)