DEV Community

Manu Shukla
Manu Shukla

Posted on • Originally published at ecorpit.com

AMD EPYC Venice: 256 cores, 3.30x rack throughput and the 2027 refresh call

AMD EPYC Venice: 256 cores, 3.30x rack throughput and the 2027 refresh call

Summary. AMD announced on 21 May 2026 that its 6th Gen EPYC processor, codenamed "Venice," had entered production ramp on TSMC's 2nm process in Taiwan, making it the first high-performance computing product in the industry to do so. Two months later the landing zones arrived: on 20 July 2026 Microsoft said Azure would add two new VM series, HDv2 and HXv2, built on 6th Gen EPYC "Venice" processors, and on 22 July 2026 AMD and Anthropic announced a deployment of up to 2 gigawatts of Instinct MI450 Series GPUs in Helios racks paired with Venice CPUs, alongside an AMD equity investment of up to $5 billion in Anthropic. AMD's own modelling puts a 256-core Venice at 3.30x the rack-level throughput of NVIDIA's 88-core Vera inside a 100 kW envelope, and beyond 36,000 cores per rack against 22,500 for Vera. Those are AMD projections, and AMD says so in its own footnotes. Nothing published so far gives Venice a general availability date or a price.

That gap between what is announced and what you can buy is the whole story for anyone holding a refresh budget.

What AMD has actually confirmed, and on which date

Server refresh decisions get made off roadmaps, and roadmaps get reported loosely. Here is the confirmed record, each item traceable to an AMD page.

21 May 2026. AMD announced the production ramp of its 6th Gen EPYC CPUs, codenamed "Venice," in Taiwan on TSMC's 2nm process, with plans to also ramp at TSMC's Arizona fab. AMD called Venice "the first high-performance computing (HPC) product in the industry to enter production on TSMC's advanced 2nm process technology."

"Ramping 'Venice' on TSMC 2nm process technology marks an important step forward in accelerating the next generation of AI infrastructure," said Dr. Lisa Su, chair and CEO, AMD, in that release. "As AI and agentic workloads scale rapidly, customers need platforms that can move from innovation to production faster."

The same release named a follow-on part, "Verano," a 6th Gen EPYC processor AMD describes as optimised for performance-per-dollar-per-watt, with LPDDR memory support aimed at power-constrained workloads. Verano matters to planning because it signals that the 2nm EPYC line is a family, not a single socket.

9 June 2026. AMD published rack-level modelling in a data centre blog by Raghu Nambiar, Corp VP, Datacenter Ecosystems and Application Engineering, Server BU. This is where the headline multipliers come from.

20 July 2026. Microsoft and AMD expanded their partnership. Azure will deploy the AMD Helios Rackscale Solution, which combines Instinct MI455X GPUs, EPYC "Venice" CPUs, Pensando networking and ROCm software. Azure will add two CPU VM series on Venice: HDv2, positioned for agentic AI and data pipelines, and HXv2, positioned for semiconductor design. AMD said it "will begin shipping Helios to customers, including Microsoft, in the second half of 2026."

22 July 2026. AMD and Anthropic announced a strategic partnership to deploy up to 2 gigawatts of Instinct MI450 Series GPUs in Helios rack-scale solutions, with the first gigawatt beginning in the first half of 2027. Anthropic will run Helios racks with MI455X GPUs, EPYC "Venice" CPUs, Pensando networking and ROCm. AMD committed to a strategic equity investment of up to $5 billion in Anthropic.

22 to 23 July 2026. AMD held Advancing AI 2026 in San Francisco, its annual gathering for developers, customers and channel partners.

What is absent from every one of those pages: a Venice general availability date, a Venice price, a Venice SKU list, and any published memory bandwidth figure. Several widely shared numbers circulating this week, including a specific per-socket memory bandwidth claim, do not appear on any AMD page we could retrieve. Treat them as unconfirmed until AMD publishes a spec sheet.

The rack numbers, read the way a buyer should read them

AMD's June modelling is the most useful thing it has published, because it frames the comparison the right way. The constraint in a modern data hall is not socket count. It is the power envelope. AMD normalises everything to a modelled 100 kW rack built on two-socket platforms, then asks how much work fits inside it.

Platform (as AMD lists it) Cores per socket Normalised rack throughput vs Vera Cores or sandboxes per rack
NVIDIA Vera ("Olympus") 88 1.0 (baseline) 22,500
Intel Xeon 6980P ("Granite Rapids-AP") 128 1.46x Not published
AMD EPYC 9965 ("Turin") 192 2.37x More than 27,000
AMD EPYC "Venice" 256 3.30x (projected) More than 36,000

Source: AMD, "Agentic AI Needs Rack-Scale CPU Performance," 9 June 2026. AMD's footnote states the results "are based on modeled rack-level configurations using publicly available and internal benchmark data" and "may not reflect actual deployed system performance."

Three things in that table deserve more attention than the 3.30x does.

First, the Turin number is the one you can act on now. EPYC 9965 at 2.37x is a shipping part on shipping platforms. If your business case survives at 2.37x, you do not need Venice, and you do not need to wait for it. AMD makes this point itself in the same blog, writing that this is infrastructure customers can build today on standard x86 platforms.

Second, the core-per-rack jump from 27,000 to 36,000 is roughly 33%, while the throughput multiplier moves from 2.37x to 3.30x, roughly 39%. The extra gain comes from per-core efficiency, not just density. AMD estimates a 64-core Venice part at a 27% performance-per-core advantage over the 88-core Vera, and projects a 96-core Venice part at 11% higher performance-per-core than the same Vera. Those are the figures to interrogate in a vendor briefing, because per-core performance is what your per-core licensed software bills against.

Third, and this trips up a lot of comparisons: the 1.46x figure belongs to Intel Xeon 6980P measured against Vera, not against AMD. AMD separately describes EPYC 9965 as roughly 1.6x the Intel part. Those are two different comparisons against two different baselines. Mixing them produces a number nobody published. We have seen the same baseline confusion before in Arm server CPU claims where Cobalt, Graviton and Axion figures each used a different reference point.

Where Venice lands first, and what that tells you

The Microsoft release is the most concrete signal in the set, because it names products a customer can eventually select in a portal rather than a partnership in gigawatts.

Announcement Date What is committed What is not committed
Azure HDv2 VM series 20 Jul 2026 Built on 6th Gen EPYC "Venice"; positioned for agentic AI and data pipelines No GA date, no region list, no pricing
Azure HXv2 VM series 20 Jul 2026 Built on 6th Gen EPYC "Venice"; positioned for semiconductor design No GA date, no region list, no pricing
AMD Helios on Azure 20 Jul 2026 MI455X GPUs, Venice CPUs, Pensando networking, ROCm; used for inference across frontier models and Azure AI services Shipping to customers in H2 2026; no Azure service date
Anthropic Helios deployment 22 Jul 2026 Up to 2 GW of MI450 Series GPUs; first gigawatt from H1 2027 No per-rack count, no site list
AMD equity investment in Anthropic 22 Jul 2026 Up to $5 billion Timing described only as "in the future"

"Customers are looking for AI infrastructure that is optimized for a wide range of workloads, from training and inference to data preparation, search, and reinforcement learning," said Satya Nadella, Chairman and CEO, Microsoft, in the 20 July release.

Read the two Azure VM series carefully. HDv2 is aimed at agentic AI and data pipelines; HXv2 at semiconductor design. Neither is a general-purpose replacement for the D or E families. That is a meaningful hint about where a 256-core socket earns its keep: workloads that are memory- and thread-hungry and that scale horizontally inside one box. If your fleet is mostly small, latency-sensitive web and API tiers, a 256-core part is the wrong shape, and no rack multiplier changes that.

The Anthropic deal points the other way. It is a GPU deal in which the CPU is the supporting cast. "By partnering with AMD across the stack, we are securing the capacity we need and optimizing it for training and serving Claude," said Tom Brown, co-founder and chief compute officer, Anthropic. "Running across a diversified range of hardware lets us map the right workloads to the right hardware." That last sentence is the honest description of what large AI operators are doing in 2026, and it is worth copying: they are not standardising on one vendor, they are matching workload shapes to silicon. Our own read on the rack-scale contest between AMD Helios and NVIDIA's Vera Rubin racks reaches the same place.

What 2nm buys, and what it does not

TSMC's 2nm node is the reason Venice exists in this shape, and AMD leaned on the partnership publicly.

"We are pleased to see AMD continue to make strong progress with its next-generation EPYC processor on our advanced 2nm process technology," said Dr. C.C. Wei, Chairman and CEO, TSMC, in the 21 May release.

For a buyer, a node shrink converts into three practical things and one non-thing.

It converts into more cores inside the same socket power budget, which is what produces the 27,000 to 36,000 core-per-rack move. It converts into better performance per watt, which is what makes the 100 kW rack the right unit of comparison. And it converts into supply risk, because leading-edge capacity is contended: AMD ramping in Taiwan with future plans for TSMC Arizona is a supply-chain statement as much as a technical one.

What it does not convert into is a lower bill by itself. Denser racks concentrate power and cooling demand into fewer floor tiles. If your colocation contract is priced per rack with a power cap well under 100 kW, a 256-core part does not help you; it strands cores you are paying for. Check your contracted kW per rack before you check the core count. The real cost is usually the facility, not the CPU.

A refresh decision framework

Most teams asking "should we wait for Venice?" are really asking three separate questions. Separate them.

Situation Wait for Venice? What to do instead
Hardware out of support in the next 9 months No Buy Turin-class now; no Venice GA date has been published
Rack power cap below roughly 15 kW No Density gains are unusable; optimise per-core licensing instead
Per-core licensed software dominates the bill Evaluate carefully Ask for per-core figures, not rack multipliers
Buying cloud instances, not servers Partly Track Azure HDv2 and HXv2; no GA date announced yet
Building GPU training or inference capacity Different question The decision is Helios versus alternatives, not Venice versus Turin
Greenfield hall being designed for 2027 or later Yes, design for it Specify 2nm-class density and revisit when AMD publishes SKUs and prices

The pattern underneath the table: Venice is a density and efficiency play, so its value to you is proportional to how power-constrained you are. Teams whose bottleneck is licensing, storage, network egress or engineering time will not feel it. That is also the honest answer to why so many 2026 refresh cycles land on the previous generation.

If your spend problem is GPUs rather than CPUs, the analysis is different again, and the levers are commitment structures and utilisation rather than silicon generation. We have written about why GPU spend became the top FinOps concern in 2026 and about how capacity block pricing changes the maths on reserved GPU capacity.

The scheduling problem nobody puts in the press release

A 256-core two-socket node is 512 cores in one failure domain. That changes how you schedule.

On Kubernetes, bin-packing behaviour that was fine on 64-core nodes starts producing noisy-neighbour effects and long tail latencies at 256 cores per socket, and a single node drain becomes a much larger disruption. Teams running mixed AI and service workloads on the same fleet typically need topology-aware scheduling, NUMA-aware placement and explicit gang scheduling for distributed jobs before very large nodes pay off. We covered the mechanics of that in our note on gang scheduling and the Kubernetes Workload API for GPU jobs.

Budget engineering time for this. A fleet that moves from 96-core to 256-core sockets without revisiting requests, limits, pod disruption budgets and NUMA policy usually gives back a large share of the theoretical gain within the first quarter.

India-specific considerations

Three things shape this decision differently for teams building in India.

Power density is the binding constraint in most Indian colocation halls, not silicon availability. Racks contracted at conventional densities cannot absorb a 100 kW envelope, and liquid cooling retrofits are a facility project with a lead time measured in quarters. Confirm the contracted kW per rack and the cooling method before specifying any 2nm-class part; that number, not the core count, decides whether Venice is buyable for you this cycle.

Cloud availability lags announcement by regions, not weeks. Neither AMD nor Microsoft has published a region list for Azure HDv2 or HXv2, so Indian teams should assume the usual pattern: US and European regions first, India regions later. If your workload has data residency obligations under the Digital Personal Data Protection Act 2023, a VM series that is unavailable in an Indian region is unavailable to you regardless of its benchmark numbers. Plan against what is generally available in Central India and South India today.

Pricing in rupees is unknowable right now. AMD has published no list price for Venice and Microsoft has published no rate for HDv2 or HXv2. Any rupee-per-hour figure you see for these parts today is extrapolation. Build the business case on the instance families you can actually price, and treat Venice as an option to re-evaluate at GA. Our broader guidance on managing AWS, Azure and GCP costs from India applies unchanged here.

What we would do with a refresh budget this quarter

Buy Turin-class capacity for anything that expires within nine months, because 2.37x against the Vera baseline is available on shipping platforms and a Venice GA date is not published.

Ask AMD and your OEM for per-core numbers under NDA rather than rack multipliers, because per-core performance is what your database and per-core licensed software bills against, and it is the figure that survives contact with a power-capped rack.

For cloud-first teams, do nothing yet beyond adding Azure HDv2 and HXv2 to a watch list. There is no GA date, no region list and no price. Re-evaluate when Microsoft publishes those three things.

For anyone designing a hall for 2027 occupancy, specify for 2nm-class density now, because retrofitting cooling later is the expensive path.

And treat every multiplier in this cycle as a modelled projection until an independent benchmark lands. AMD labelled its own numbers that way. The vendors that are honest in their footnotes deserve to be read there.

FAQ

Has AMD EPYC Venice launched?

AMD announced on 21 May 2026 that Venice had entered production ramp on TSMC's 2nm process in Taiwan. Production ramp is not the same as general availability. As of 23 July 2026, AMD has published no general availability date, no SKU list and no price for the 6th Gen EPYC Venice family.

How many cores does EPYC Venice have?

AMD lists Venice at 256 cores per socket in its June 2026 rack-scale modelling, compared with 192 cores for the current EPYC 9965 Turin part and 88 cores for NVIDIA Vera. AMD states that a two-socket Venice platform is architected to scale beyond 36,000 cores in a modelled 100 kW rack class.

Is the 3.30x performance claim measured or projected?

Projected. AMD's 9 June 2026 blog presents 3.30x as a normalised rack-level figure against an NVIDIA Vera baseline, and its footnote states the results are based on modelled rack-level configurations using publicly available and internal benchmark data and may not reflect actual deployed system performance.

Which Azure VM series will run on EPYC Venice?

Microsoft said on 20 July 2026 that Azure will add two VM series built on 6th Gen AMD EPYC Venice processors: HDv2, positioned for agentic AI and data pipelines, and HXv2, positioned for semiconductor design. Microsoft has not published a general availability date, a region list or pricing for either series.

What is AMD Helios and when does it ship?

Helios is AMD's rack-scale platform combining Instinct MI455X GPUs, EPYC Venice CPUs, Pensando networking and ROCm software. AMD said on 20 July 2026 that it will begin shipping Helios to customers, including Microsoft, in the second half of 2026. Anthropic's first gigawatt deployment begins in the first half of 2027.

How large is the AMD and Anthropic deal?

AMD and Anthropic announced on 22 July 2026 a partnership to deploy up to 2 gigawatts of AMD Instinct MI450 Series GPUs in Helios rack-scale solutions, with the first gigawatt beginning in the first half of 2027. AMD separately committed to a strategic equity investment of up to $5 billion in Anthropic.

Should we delay a server refresh to wait for Venice?

If your hardware falls out of support within nine months, no. AMD has published no Venice availability date, and the shipping EPYC 9965 Turin part already models at 2.37x the Vera rack baseline. Waiting makes sense only for greenfield halls being designed for 2027 occupancy or later.

Does higher core density automatically lower our cloud bill?

No. Density converts into savings only when your bottleneck is rack power. If your colocation contract caps power well below the 100 kW envelope AMD models against, a 256-core socket strands capacity you are paying for. Check contracted kilowatts per rack before comparing core counts.

How eCorpIT can help

eCorpIT is a Gurugram-based technology consultancy that helps engineering teams turn infrastructure announcements into defensible capacity plans. We model refresh options against your actual rack power, licensing terms and workload shapes rather than vendor multipliers, and we run the Kubernetes and scheduling work that large-socket nodes require before the density pays off. If you are weighing a 2026 or 2027 refresh, or sizing GPU capacity alongside it, talk to our team. We also run managed cloud FinOps engagements and managed Kubernetes AI platform work for teams that would rather not build the practice in-house.

References

  1. AMD Announces Production Ramp of Next-Generation AMD EPYC Processor "Venice" on TSMC 2nm Process Technology, AMD Newsroom, 21 May 2026.
  2. Agentic AI Needs Rack-Scale CPU Performance, AMD EPYC Delivers It Today, Raghu Nambiar, AMD Data Center Blog, 9 June 2026.
  3. Microsoft to Deploy Next-Gen AMD Instinct and AMD EPYC Processors as the Companies Expand Their Long-Term Strategic Partnership, AMD Investor Relations, 20 July 2026.
  4. AMD and Anthropic Announce Strategic Partnership to Deploy Up to 2 Gigawatts of AMD Instinct MI450 Series GPUs, AMD Investor Relations, 22 July 2026.
  5. AMD EPYC Server CPU Specifications, AMD.
  6. AMD Announces Production Ramp of Next-Generation AMD EPYC Processor "Venice" (investor version), AMD Investor Relations, 21 May 2026.
  7. AMD Achieves First TSMC N2 Product Silicon Milestone, AMD Newsroom, 14 April 2025.
  8. AMD EPYC Processors product page, AMD.
  9. AMD Advancing AI 2026 event page, AMD.
  10. AMD and Meta Announce Expanded Strategic Partnership to Deploy 6 Gigawatts of AMD GPUs, AMD Newsroom, 24 February 2026.
  11. Announcing the ROCm Certification Program, AMD Developer Resources, 13 July 2026.
  12. AMD Reports First Quarter 2026 Financial Results, AMD Newsroom, 5 May 2026.

Last updated: 23 July 2026.

Top comments (0)