<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Amit</title>
    <description>The latest articles on DEV Community by Amit (@amitrix).</description>
    <link>https://dev.to/amitrix</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3962358%2F978a8f18-68b0-409b-9b3a-2156d0be550c.png</url>
      <title>DEV Community: Amit</title>
      <link>https://dev.to/amitrix</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/amitrix"/>
    <language>en</language>
    <item>
      <title>Most Enterprises Are Chasing AI Autonomy Before Defining the Outcome</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Thu, 24 Sep 2026 18:00:35 +0000</pubDate>
      <link>https://dev.to/amitrix/most-enterprises-are-chasing-ai-autonomy-before-defining-the-outcome-2c1d</link>
      <guid>https://dev.to/amitrix/most-enterprises-are-chasing-ai-autonomy-before-defining-the-outcome-2c1d</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Enterprises are adopting agents faster than they are redesigning the work those agents will perform.&lt;/li&gt;
&lt;li&gt;The unit of autonomy is the &lt;strong&gt;delegation envelope&lt;/strong&gt;: one actor, one goal, one allowed action set, one evidence contract, and one recovery boundary.&lt;/li&gt;
&lt;li&gt;Capability, delegated authority, and operating maturity are separate scorecards. A system that can perform an action has not automatically earned permission to take it.&lt;/li&gt;
&lt;li&gt;The goal is the minimum authority needed to improve a defined outcome, expanded only when the workflow can prove and recover from its work.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;a href="https://www.ibm.com/thought-leadership/institute-business-value/report/2025-ceo" rel="noopener noreferrer"&gt;IBM found that 64% of 2,000 surveyed CEOs&lt;/a&gt; said fear of falling behind drives technology investment before value is clear. Another 61% were adopting AI agents and preparing to scale them, while only 25% of AI initiatives had delivered the expected return. Enterprises are chasing autonomy faster than outcomes.&lt;/p&gt;

&lt;p&gt;That is why &lt;em&gt;How autonomous should our AI be?&lt;/em&gt; is the wrong opening question. An enterprise does not have one autonomy level. A software-maintenance workflow can complete bounded work while a payment workflow requires approval for every consequential action. Customer service can resolve a routine request while legal review remains advisory.&lt;/p&gt;

&lt;p&gt;The useful question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What outcome are we trying to achieve, how does the work produce it today, and what authority has each part of that workflow earned?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Start With the Outcome, Not the Agent
&lt;/h2&gt;

&lt;p&gt;“Deploy an agent” is not a business outcome.&lt;/p&gt;

&lt;p&gt;Reduce account-opening time. Resolve routine service requests without repeat contact. Cut invoice exceptions. Detect fraud before money moves. These are measurable outcomes. An agent is one possible actor alongside deterministic software, analytical models, policies, systems of record, and human judgment.&lt;/p&gt;

&lt;p&gt;Task productivity is not workflow performance. An agent may draft an answer in seconds while the case waits two days for approval, or generate more code while review becomes the bottleneck.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-2025" rel="noopener noreferrer"&gt;McKinsey’s 2025 global survey&lt;/a&gt; found that 62% of respondents’ organizations were experimenting with agents, but only 23% were scaling an agentic system somewhere in the enterprise. In any individual function, no more than 10% reported scaling agents. Adoption, workflow penetration, and business value are different scorecards.&lt;/p&gt;

&lt;p&gt;The outcome defines what “done” means, which decisions matter, what the agent may change, which errors are tolerable, when a person intervenes, and how success is measured. Without that sequence, autonomy becomes a technology objective detached from the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Delegation Envelope Is the Real Unit
&lt;/h2&gt;

&lt;p&gt;Authority lives in a &lt;strong&gt;delegation envelope&lt;/strong&gt;: one actor, one goal, one allowed action set, one evidence contract, and one recovery boundary. Change any part and you have a different envelope. This is the unit that can be designed, audited, graduated, and revoked.&lt;/p&gt;

&lt;p&gt;A workflow contains multiple envelopes at different levels. The same agent may hold advisory authority for a legal conclusion, bounded authority for assembling its evidence, and no authority to send the result externally. Calling the whole system “L3” hides the consequential boundaries.&lt;/p&gt;

&lt;p&gt;Three scorecards must stay separate:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Capability:&lt;/strong&gt; What can the system accomplish, and with what reliability?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delegated authority:&lt;/strong&gt; What effects may it cause without further approval?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operating maturity:&lt;/strong&gt; Can the organization observe, constrain, audit, interrupt, and recover that delegation?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;“The model can do it” answers only the first question. The levels below describe envelope scope, not corporate progress. There is no universal L1–L5 standard for enterprise agents; this is my design model.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;What the envelope permits&lt;/th&gt;
&lt;th&gt;Human responsibility&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;L1 — Advise&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Analyze, draft, predict, or recommend. No material external action.&lt;/td&gt;
&lt;td&gt;Performs the work and makes the decision&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;L2 — Per-action approval&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One discrete action whose exact occurrence and consequential parameters receive explicit approval&lt;/td&gt;
&lt;td&gt;Directs the action and reviews consequential parameters&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;L3 — Pre-authorized bounded workflow&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A multi-step workflow whose class, boundaries, and evidence contract were approved &lt;em&gt;in advance&lt;/em&gt;; individual runs are not separately approved&lt;/td&gt;
&lt;td&gt;Approves the workflow class and its exceptions—not each step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;L4 — Coordinating envelope&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Sequences and selects among several already-bounded L2/L3 envelopes toward a defined outcome. Composed authority, not new authority.&lt;/td&gt;
&lt;td&gt;Sets constraints, owns material deviations and cross-boundary exceptions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;L5 — Continuing objective&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Creates and revises its own plans toward a standing objective under policy, budget, and time limits&lt;/td&gt;
&lt;td&gt;Sets strategy, prohibited actions, and revocation authority&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The L2/L3 distinction is when approval happens: each occurrence at L2, the workflow class at L3. L4 composes bounded envelopes; if the coordinator gains authority none of them had, the boundary has leaked. L5 is a horizon, not a target. The public sources reviewed for this article did not document a standing-objective envelope with self-revised plans under material consequence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enterprise Workflows Are Landing at Different Levels
&lt;/h2&gt;

&lt;p&gt;The ranges below are &lt;strong&gt;my assessment of public evidence available as of September 2026&lt;/strong&gt;—not a market standard or benchmark. I use ranges because the evidence does not support point estimates, and I label vendor claims, regulatory boundaries, and thin evidence.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Enterprise workflow&lt;/th&gt;
&lt;th&gt;Typical range&lt;/th&gt;
&lt;th&gt;Basis, and its limits&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Software engineering&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L2–L3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Vendor product claim:&lt;/em&gt; &lt;a href="https://docs.github.com/en/copilot/concepts/coding-agent/coding-agent" rel="noopener noreferrer"&gt;GitHub documents&lt;/a&gt; that its coding agent implements bounded changes and opens pull requests; humans retain merge authority. Documented capability, not measured adoption depth.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;IT and employee service&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L2–L3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Vendor product claim:&lt;/em&gt; &lt;a href="https://newsroom.servicenow.com/press-releases/details/2026/ServiceNow-opens-its-full-system-of-action-to-every-AI-Agent-in-the-enterprise/default.aspx" rel="noopener noreferrer"&gt;ServiceNow's platform announcement&lt;/a&gt; describes routine service workflows completing inside existing approvals. Marketing material; no independent outcome data.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Customer service&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L2–L3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Routine resolution is the most commonly reported bounded-completion case in the surveys cited above; exceptions stay human-owned. Vendor deflection figures are not comparable across firms; I do not rely on them.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Sales&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L1–L2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Weak evidence.&lt;/strong&gt; No function-specific study cited. Inference from the ≤10% per-function scaling ceiling in &lt;a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-2025" rel="noopener noreferrer"&gt;McKinsey's survey&lt;/a&gt; plus the structural point that pricing and commitments are contractually binding. Hypothesis.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Marketing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L1–L2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Observed adoption:&lt;/em&gt; &lt;a href="https://www.bcg.com/publications/2026/making-the-agentic-marketing-transformation-a-reality" rel="noopener noreferrer"&gt;BCG reports agent-led workflows remain a minority&lt;/a&gt;; campaigns operate inside approved briefs and budgets.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;HR&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L1–L2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Observed adoption:&lt;/em&gt; employee service and onboarding advance. &lt;em&gt;Regulatory boundary:&lt;/em&gt; employment decisions face &lt;a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai" rel="noopener noreferrer"&gt;EU AI Act high-risk obligations&lt;/a&gt;—a legal ceiling, not a measurement of practice.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Finance and accounting&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L1–L2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Observed adoption:&lt;/em&gt; reconciliation and close preparation advance; assurance readiness and material approvals remain limiting (&lt;a href="https://kpmg.com/content/dam/kpmgsites/uk/pdf/2026/05/ai-in-finance-report.pdf.coredownload.inline.pdf" rel="noopener noreferrer"&gt;KPMG&lt;/a&gt;).&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Procurement&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L1–L2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Observed adoption:&lt;/em&gt; supplier analysis and low-value transactions advance; trust in autonomous decisions is the leading reported barrier (&lt;a href="https://www.bcg.com/publications/2026/scaling-agentic-ai-in-tech-procurement" rel="noopener noreferrer"&gt;BCG&lt;/a&gt;).&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Supply chain and logistics&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L1–L2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Observed adoption:&lt;/em&gt; agents generate and execute narrow plans; cross-network trade-offs typically remain planner-approved (&lt;a href="https://www.bcg.com/publications/2026/how-ai-agents-are-transforming-supply-chains" rel="noopener noreferrer"&gt;BCG&lt;/a&gt;).&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Risk, legal, and compliance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Observed adoption:&lt;/em&gt; research, review, monitoring, drafting, and evidence collection dominate (&lt;a href="https://legal.thomsonreuters.com/blog/highlights-from-the-2026-ai-in-professional-services-report-and-what-it-means-for-legal-teams-tri" rel="noopener noreferrer"&gt;Thomson Reuters&lt;/a&gt;). Also subject to professional-responsibility limits.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Physical operations&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;L2–L3 in engineered environments only&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Weak evidence.&lt;/strong&gt; No study cited here covers robotics. Inference from the engineered-environment argument below: bounded robots operate where state and recovery are physically constrained. General-purpose physical agents are earlier. Treat as hypothesis.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three patterns matter. Engineering leads where evidence and rollback are executable. Regulated functions move more slowly because their actions carry legal or financial consequence. Maturity varies within every function because workflows expose different boundaries.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Software Engineering Moved First
&lt;/h2&gt;

&lt;p&gt;Software engineering already had the control system for bounded delegation: machine-readable state, tests, sandboxes, granular permissions, reviewable diffs, independent approval, and rollback.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.github.com/en/copilot/concepts/coding-agent/coding-agent" rel="noopener noreferrer"&gt;GitHub's coding-agent documentation&lt;/a&gt;—a vendor product description—maps cleanly onto the envelope. The agent evaluates an issue, changes code, runs tests, and opens a pull request. The pull request is the evidence contract; merge authority is the boundary. The agent completes a pre-authorized workflow without inheriting authority over the software outcome.&lt;/p&gt;

&lt;p&gt;Other functions need equivalent controls: simulation instead of a branch, deterministic business rules instead of tests, action diffs instead of code diffs, task-scoped credentials instead of deployment access, and compensating transactions instead of rollback.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://research.ibm.com/publications/measuring-agents-in-production" rel="noopener noreferrer"&gt;ICLR 2026 &lt;em&gt;Measuring Agents in Production&lt;/em&gt; study&lt;/a&gt; found that across 306 practitioners and 20 case studies, 68% of production agents executed at most ten steps before human intervention, 74% relied primarily on human evaluation, and reliability remained the leading challenge. Production maturity currently comes from bounded authority, not unrestricted action.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Worked Example: Invoice Exceptions
&lt;/h2&gt;

&lt;p&gt;The outcome is not “process invoices with an agent.” It is “resolve valid exceptions within one business day without an unauthorized posting or payment.”&lt;/p&gt;

&lt;p&gt;An L1 envelope classifies the exception and assembles evidence. L2 requests missing documentation under a pre-approved rule. L3 resolves named exception classes below a financial threshold, while payment release remains separately gated. Each step records the source data, rule, and proposed change. Recovery means reopening the exception, reversing a provisional posting, and revoking authority when error thresholds are exceeded.&lt;/p&gt;

&lt;p&gt;The envelope earns wider scope only when its goal, action set, evidence contract, authority history, and recovery boundary strengthen together. Widening authority without strengthening those controls is weak governance wearing a level number.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Executives Are Actually Deciding
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.deloitte.com/us/en/pages/consulting/articles/state-of-generative-ai-in-enterprise.html" rel="noopener noreferrer"&gt;Deloitte’s 2026 State of AI research&lt;/a&gt; found that 74% of surveyed leaders expected moderate or extensive agent use within two years. Yet only 21% reported mature governance for autonomous agents, 30% were redesigning key processes around AI, and 84% had not redesigned jobs around it.&lt;/p&gt;

&lt;p&gt;The gap is concrete: organizations want class-level autonomy before they have evidence contracts that make class-level approval safe. L4 compounds the problem across decision rights, data definitions, permissions, budgets, and accountability. The composition itself needs an owner.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-2025" rel="noopener noreferrer"&gt;McKinsey found AI high performers were 2.8 times more likely&lt;/a&gt; to report fundamental workflow redesign than other respondents—55% versus 20%. That is an association, not proof of causation, but it is a stronger maturity signal than agent count.&lt;/p&gt;

&lt;p&gt;The operating sequence is simple: observe the workflow, let the system advise, approve individual actions, pre-authorize proven workflow classes, then coordinate bounded workflows. At each step, collect outcome and exception evidence, monitor drift, and demote or revoke authority when the evidence deteriorates.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf" rel="noopener noreferrer"&gt;NIST AI Risk Management Framework&lt;/a&gt; asks organizations to map an AI system’s intended purpose, context, affected actors, objectives, and legal requirements. NIST’s &lt;a href="https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd" rel="noopener noreferrer"&gt;February 2026 draft on agent identity&lt;/a&gt; asks how standards can support identification, authorization, auditing, and non-repudiation. My architectural implication is that agents should receive scoped, auditable delegation rather than silently inherit a person’s full access.&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s Missing
&lt;/h2&gt;

&lt;p&gt;The market has plenty of agent frameworks. It has fewer mechanisms for operating authority.&lt;/p&gt;

&lt;p&gt;Business workflows still lack standard equivalents for tests, action diffs, merge protection, and rollback. Evaluations score answers while production risk lives in tool use, permissions, side effects, and exceptions. Weak human review becomes an approval button without enough context or power to refuse.&lt;/p&gt;

&lt;p&gt;The missing layer is not another autonomy label. It is an authority lifecycle that defines the envelope, gathers evidence in simulation and production, graduates scope within explicit thresholds, monitors drift, and revokes authority when performance deteriorates.&lt;/p&gt;

&lt;h2&gt;
  
  
  So What
&lt;/h2&gt;

&lt;p&gt;Ask which envelope you are designing: whose goal, which actions, what evidence, what recovery, and whether approval happens per action or per workflow class. Some envelopes should remain advisory permanently. Maturity is knowing which can move—and having a mechanism for moving them back.&lt;/p&gt;

&lt;p&gt;Here is a bet that can be checked. &lt;strong&gt;By September 2028,&lt;/strong&gt; enterprises publicly reporting durable financial outcomes from agents will be distinguished by workflow-level evidence infrastructure—action diffs, scoped credentials, compensating transactions, and named envelope owners—rather than by the breadth or autonomy level of their deployments. I also expect no credible public case of a sustained L5 envelope operating over material consequence by that date.&lt;/p&gt;

&lt;p&gt;I am wrong if enterprises scale agent authority broadly and profitably without those controls, or if a standing-objective agent runs a consequential business objective with self-revised plans and holds up under audit.&lt;/p&gt;

&lt;p&gt;The question I have not resolved is how much evidence a workflow should accumulate before it moves up one rung—and who should set that threshold when the benefits and the cost of failure fall on different functions.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>enterpriseai</category>
      <category>decisionmaking</category>
      <category>aigovernance</category>
    </item>
    <item>
      <title>Wiring Codex and ChatGPT Desktop to OpenAI Models on Amazon Bedrock Runtime</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:54:39 +0000</pubDate>
      <link>https://dev.to/amitrix/wiring-codex-and-chatgpt-desktop-to-openai-models-on-amazon-bedrock-runtime-5b6e</link>
      <guid>https://dev.to/amitrix/wiring-codex-and-chatgpt-desktop-to-openai-models-on-amazon-bedrock-runtime-5b6e</guid>
      <description>&lt;p&gt;Codex CLI and the ChatGPT desktop app can use the same native Amazon Bedrock Runtime configuration. The working pattern has three parts: the built-in Runtime provider, an AWS profile and Region, and a global model profile ID.&lt;/p&gt;

&lt;p&gt;I verified this pattern on September 11, 2026 with four OpenAI model profiles: GPT-5.6 Sol, Terra, Luna, and GPT-6 Astra. All four completed headless runs through the native Runtime provider in the desktop app's bundled Codex 0.153.4. Astra also passed with Codex CLI 0.154.0 using the same saved configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Codex CLI and the ChatGPT desktop app read the same local Codex configuration, so one provider definition can serve both surfaces.&lt;/li&gt;
&lt;li&gt;The endpoint remains regional, while a &lt;code&gt;global.openai.*&lt;/code&gt; model ID selects global cross-Region inference.&lt;/li&gt;
&lt;li&gt;The built-in &lt;code&gt;amazon-bedrock-runtime&lt;/code&gt; provider uses an AWS profile directly; no static API key, token generator, or local proxy is required.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;amazon-bedrock&lt;/code&gt; and &lt;code&gt;amazon-bedrock-runtime&lt;/code&gt; are different providers. The first selects Mantle in the tested clients; the second selects Bedrock Runtime.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The working architecture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://learn.chatgpt.com/docs/amazon-bedrock" rel="noopener noreferrer"&gt;OpenAI documents shared local configuration across Codex surfaces&lt;/a&gt;. &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/endpoints.html" rel="noopener noreferrer"&gt;Amazon Bedrock documents the Runtime endpoint&lt;/a&gt; as the recommended endpoint for new applications, including OpenAI-compatible Responses and Chat Completions paths under &lt;code&gt;/openai/v1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The resulting request path is direct:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Codex CLI or ChatGPT desktop
        |
        | reads ~/.codex/config.toml
        v
Native amazon-bedrock-runtime provider
        |
        | uses AWS profile and Region
        v
bedrock-runtime.&amp;lt;region&amp;gt;.amazonaws.com/openai/v1
        |
        v
global.openai.&amp;lt;model&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no gateway in this path. AWS credentials stay in the normal local profile chain, and Codex sends Responses API traffic directly to Bedrock Runtime.&lt;/p&gt;

&lt;p&gt;The word &lt;code&gt;global&lt;/code&gt; applies to the model profile, not the hostname. The network endpoint still contains a Region such as &lt;code&gt;us-west-2&lt;/code&gt;; the &lt;code&gt;global.openai.*&lt;/code&gt; model ID tells Bedrock to use its global cross-Region inference profile. &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-mantle.html" rel="noopener noreferrer"&gt;The Bedrock Responses API documentation&lt;/a&gt; shows the Runtime base URL and global profile naming pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;p&gt;The setup needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A current Codex CLI installation&lt;/li&gt;
&lt;li&gt;The ChatGPT desktop app with Codex support&lt;/li&gt;
&lt;li&gt;An AWS profile with permission to invoke the selected Bedrock models&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Confirm the AWS session before changing Codex:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sts get-caller-identity &lt;span class="nt"&gt;--profile&lt;/span&gt; YOUR_AWS_PROFILE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The shared Codex configuration
&lt;/h2&gt;

&lt;p&gt;Add the following provider to &lt;code&gt;~/.codex/config.toml&lt;/code&gt;. Replace &lt;code&gt;YOUR_AWS_PROFILE&lt;/code&gt; and select a Bedrock Region where the models are available.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="py"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"global.openai.gpt-6-astra"&lt;/span&gt;
&lt;span class="py"&gt;model_provider&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"amazon-bedrock-runtime"&lt;/span&gt;
&lt;span class="py"&gt;model_reasoning_effort&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"medium"&lt;/span&gt;

&lt;span class="nn"&gt;[model_providers.amazon-bedrock-runtime.aws]&lt;/span&gt;
&lt;span class="py"&gt;region&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"us-west-2"&lt;/span&gt;
&lt;span class="py"&gt;profile&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"YOUR_AWS_PROFILE"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://learn.chatgpt.com/docs/config-file/config-sample" rel="noopener noreferrer"&gt;Codex configuration sample&lt;/a&gt; documents AWS profile and Region settings for built-in Amazon Bedrock providers. This configuration references the selected profile rather than embedding access keys. The profile can still use AWS IAM Identity Center, role assumption, environment credentials, or another supported AWS credential source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Global OpenAI model profiles
&lt;/h2&gt;

&lt;p&gt;These are the four global profile IDs I verified:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Codex model ID&lt;/th&gt;
&lt;th&gt;Verified result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;&lt;code&gt;global.openai.gpt-5.6-sol&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Native Runtime passed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Terra&lt;/td&gt;
&lt;td&gt;&lt;code&gt;global.openai.gpt-5.6-terra&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Native Runtime passed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Luna&lt;/td&gt;
&lt;td&gt;&lt;code&gt;global.openai.gpt-5.6-luna&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Native Runtime passed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;&lt;code&gt;global.openai.gpt-6-astra&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Native Runtime passed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/machine-learning/get-started-with-openai-gpt-5-6-sol-terra-and-luna-on-amazon-bedrock/" rel="noopener noreferrer"&gt;AWS introduced Sol, Terra, and Luna on Amazon Bedrock&lt;/a&gt; in July 2026 and &lt;a href="https://aws.amazon.com/blogs/machine-learning/take-on-your-most-ambitious-work-with-gpt-6-astra-on-amazon-bedrock/" rel="noopener noreferrer"&gt;announced GPT-6 Astra availability&lt;/a&gt; in September 2026. Model access still depends on the account, selected Region, and current Bedrock catalog, so discovery in one account is evidence for that environment rather than a universal availability guarantee.&lt;/p&gt;

&lt;p&gt;Set Astra as the default in the configuration, or select a model for an individual Codex run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex &lt;span class="nt"&gt;-m&lt;/span&gt; global.openai.gpt-5.6-sol
codex &lt;span class="nt"&gt;-m&lt;/span&gt; global.openai.gpt-5.6-terra
codex &lt;span class="nt"&gt;-m&lt;/span&gt; global.openai.gpt-5.6-luna
codex &lt;span class="nt"&gt;-m&lt;/span&gt; global.openai.gpt-6-astra
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Verify the Runtime path
&lt;/h2&gt;

&lt;p&gt;A headless run provides a clean test because it removes the interactive interface from the request path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ephemeral&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--skip-git-repo-check&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-m&lt;/span&gt; global.openai.gpt-5.6-sol &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"Reply with exactly SOL_RUNTIME_OK"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Repeat the probe with the other model IDs and a distinct expected string. My four probes returned:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOL_RUNTIME_OK
TERRA_RUNTIME_OK
LUNA_RUNTIME_OK
ASTRA_RUNTIME_OK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This verifies more than model discovery. It confirms that Codex can load the provider, resolve the named AWS profile, reach the Runtime endpoint, select the global profile, and receive a model response.&lt;/p&gt;

&lt;p&gt;For the desktop app, fully quit and reopen the application after changing &lt;code&gt;~/.codex/config.toml&lt;/code&gt;. The model picker should then expose the same provider catalog. Existing conversations may retain their original model selection, so use a new task when validating a newly selected model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The provider name selects the endpoint
&lt;/h2&gt;

&lt;p&gt;The two built-in Bedrock provider names do not select the same endpoint.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Endpoint selected in the tested clients&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;amazon-bedrock&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Bedrock Mantle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;amazon-bedrock-runtime&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Bedrock Runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This distinction is visible in an actual call. With Codex 0.153.4, &lt;code&gt;amazon-bedrock&lt;/code&gt; sent Astra to a &lt;code&gt;bedrock-mantle&lt;/code&gt; URL and received a model-not-found response. The same binary, model ID, AWS profile, and Region succeeded after changing only the provider to &lt;code&gt;amazon-bedrock-runtime&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The native provider passed across both tested Codex versions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Desktop-bundled Codex 0.153.4 returned &lt;code&gt;SOL_NATIVE_OK&lt;/code&gt;, &lt;code&gt;TERRA_NATIVE_OK&lt;/code&gt;, &lt;code&gt;LUNA_NATIVE_OK&lt;/code&gt;, and &lt;code&gt;ASTRA_NATIVE_OK&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Standalone Codex CLI 0.154.0 returned &lt;code&gt;CLI_NATIVE_RUNTIME_OK&lt;/code&gt; from the same saved provider configuration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Inspect the versions when diagnosing behavior:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The desktop app's bundled version is visible in its diagnostic output. Version differences still matter for model catalogs and client features, but they do not require a custom provider for this Runtime path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Direct Runtime first, gateway when required
&lt;/h2&gt;

&lt;p&gt;Direct Bedrock Runtime is the lower-complexity path when AWS identity, model access, logging, and billing controls meet the requirement. An &lt;a href="https://aws.amazon.com/blogs/machine-learning/set-up-openai-chatgpt-codex-with-litellm-on-amazon-ecs-and-amazon-bedrock/" rel="noopener noreferrer"&gt;AWS reference architecture for Codex with LiteLLM&lt;/a&gt; positions a gateway as an additional layer for centralized routing, budgets, or policy controls—not as a requirement for basic Codex connectivity.&lt;/p&gt;

&lt;p&gt;That distinction matters. A local proxy or hosted gateway adds another credential boundary, another failure point, and another component to operate. Add one when it supplies a control the direct path does not provide.&lt;/p&gt;

&lt;h2&gt;
  
  
  What remains open
&lt;/h2&gt;

&lt;p&gt;The native Runtime provider passed direct Codex execution in both binaries. The desktop app also restarted with successful configuration and model-catalog reads. I have not yet tested every desktop-only helper action against all four global model profiles; those flows can select models independently from the main conversation.&lt;/p&gt;

&lt;p&gt;The configuration boundary is the part to keep under test: one file, two client surfaces, four global model profiles, and one AWS profile. When any client version changes, rerun one headless probe and one fresh desktop task before treating the wiring as stable.&lt;/p&gt;

&lt;p&gt;So what: Bedrock model access is only one layer of the setup. The complete contract is endpoint, API format, authentication, model profile, and client configuration—and all five need to work together.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>infrastructure</category>
      <category>patterns</category>
    </item>
    <item>
      <title>Alternative AI Hardware Is a Systems Problem</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:54:04 +0000</pubDate>
      <link>https://dev.to/amitrix/alternative-ai-hardware-is-a-systems-problem-2p08</link>
      <guid>https://dev.to/amitrix/alternative-ai-hardware-is-a-systems-problem-2p08</guid>
      <description>&lt;p&gt;Alternative AI hardware is not one product category.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.huawei.com/en/news/2025/9/hc-lingqu-ai-superpod" rel="noopener noreferrer"&gt;Huawei Ascend&lt;/a&gt; is becoming a complete compute, interconnect, systems, and software platform. &lt;a href="https://furiosa.ai/rngd" rel="noopener noreferrer"&gt;FuriosaAI RNGD&lt;/a&gt; is a power-conscious inference accelerator that fits into standard servers. &lt;a href="https://rebellions.ai/resources/" rel="noopener noreferrer"&gt;Rebellions REBEL&lt;/a&gt; starts with a PCIe NPU and scales into multi-card and rack-level systems.&lt;/p&gt;

&lt;p&gt;All three reduce dependence on the dominant GPU stack. They do it at different boundaries, with different software commitments and different evidence. This is an architecture comparison; public evidence is not yet consistent enough for a normalized performance ranking.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Huawei competes at the platform boundary: accelerator, scale-up fabric, clusters, compiler, framework, and cloud infrastructure.&lt;/li&gt;
&lt;li&gt;FuriosaAI narrows the problem to efficient transformer inference in a familiar PCIe server form factor.&lt;/li&gt;
&lt;li&gt;Rebellions combines a data-center NPU with PCIe, CXL-oriented memory expansion, scale-out Ethernet, and rack-level system designs.&lt;/li&gt;
&lt;li&gt;The strategic test is not whether an NPU has competitive peak arithmetic. It is whether the surrounding system can run supported models reliably, at the required quality, with enough software and operational evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Three architectures, three ownership boundaries
&lt;/h2&gt;

&lt;p&gt;The phrase “alternative accelerator” hides the most important distinction: who owns the system around the silicon?&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;Primary boundary&lt;/th&gt;
&lt;th&gt;Physical unit&lt;/th&gt;
&lt;th&gt;Software commitment&lt;/th&gt;
&lt;th&gt;Strategic objective&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Huawei Ascend&lt;/td&gt;
&lt;td&gt;Full platform&lt;/td&gt;
&lt;td&gt;Chip, server, SuperPod, SuperCluster, and cloud&lt;/td&gt;
&lt;td&gt;CANN, MindSpore, framework adapters, cluster software&lt;/td&gt;
&lt;td&gt;Build an integrated sovereign compute stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FuriosaAI RNGD&lt;/td&gt;
&lt;td&gt;Efficient inference accelerator&lt;/td&gt;
&lt;td&gt;PCIe card, server, and partner system&lt;/td&gt;
&lt;td&gt;Furiosa SDK, compiler, runtime, and model support&lt;/td&gt;
&lt;td&gt;Add efficient inference capacity to standard data centers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rebellions REBEL&lt;/td&gt;
&lt;td&gt;Scalable inference NPU&lt;/td&gt;
&lt;td&gt;PCIe card, multi-card system, rack, and cloud service&lt;/td&gt;
&lt;td&gt;RBLN SDK, compiler, runtime, serving integration&lt;/td&gt;
&lt;td&gt;Scale transformer inference from card to rack&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These systems should not be collapsed into one benchmark row. Huawei is attempting to reproduce much of the platform surrounding a GPU. Furiosa and Rebellions can coexist beside GPU pools because their first deployment boundary remains a server accelerator.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    APP["Models and applications"]

    subgraph H["Huawei platform boundary"]
        HC["Ascend accelerator"]
        HF["UnifiedBus / system fabric"]
        HS["SuperPod and SuperCluster"]
        HSW["CANN, MindSpore, cluster software"]
        HC --&amp;gt; HF --&amp;gt; HS
        HSW --&amp;gt; HS
    end

    subgraph F["Furiosa deployment boundary"]
        FC["RNGD PCIe cards"]
        FS["Standard inference server"]
        FSW["Furiosa SDK and runtime"]
        FC --&amp;gt; FS
        FSW --&amp;gt; FS
    end

    subgraph R["Rebellions deployment boundary"]
        RC["REBEL PCIe NPUs"]
        RS["Multi-card server or rack"]
        RSW["RBLN SDK and serving stack"]
        RC --&amp;gt; RS
        RSW --&amp;gt; RS
    end

    APP --&amp;gt; HSW
    APP --&amp;gt; FSW
    APP --&amp;gt; RSW&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The operator accepts the largest platform commitment with Huawei. The smaller vendors ask for a narrower model and serving commitment, but depend more heavily on host CPUs, merchant networking, server partners, and external orchestration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Huawei is building the whole stack
&lt;/h2&gt;

&lt;p&gt;Huawei’s architecture has moved far beyond a single Ascend card.&lt;/p&gt;

&lt;p&gt;The company’s &lt;a href="https://e.huawei.com/en/products/computing/ascend/atlas-900-a3-superpod" rel="noopener noreferrer"&gt;Atlas 900 A3 SuperPoD&lt;/a&gt; combines Ascend compute, high-speed interconnect, storage, management, and liquid-cooled cabinets into a scale-up system. Huawei describes a 384-card configuration with 300 TB of HBM and 122.8 TB/s of scale-up bandwidth.&lt;/p&gt;

&lt;p&gt;The published &lt;a href="https://www.huawei.com/en/news/2025/9/hc-lingqu-ai-superpod" rel="noopener noreferrer"&gt;SuperCluster architecture&lt;/a&gt; expands that boundary again. Multiple SuperPods connect through the company’s UnifiedBus fabric, with separate scale-up and scale-out layers. Huawei announced Atlas 950 and Atlas 960 SuperPods based on future Ascend generations and described SuperClusters containing hundreds of thousands to more than one million cards.&lt;/p&gt;

&lt;p&gt;Those figures are roadmap and vendor architecture claims. They establish direction and intended scale, not measured production efficiency.&lt;/p&gt;

&lt;p&gt;The software boundary is equally broad. Huawei’s &lt;a href="https://www.hiascend.com/en/software/cann" rel="noopener noreferrer"&gt;CANN platform&lt;/a&gt; provides compilers, libraries, runtime components, development tools, and framework adapters for Ascend hardware. MindSpore supplies a Huawei-backed framework path, while PyTorch and other frameworks rely on integration layers.&lt;/p&gt;

&lt;p&gt;That creates an integrated stack:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model and framework
  → CANN graph compilation and kernels
  → Ascend runtime and collectives
  → UnifiedBus scale-up fabric
  → SuperPod
  → SuperCluster and cloud operations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The benefit is one architecture across silicon, fabric, server, and software. The cost is a larger migration and validation boundary. Kernel coverage, numerical behavior, framework compatibility, tooling, and distributed execution all need to match the model portfolio.&lt;/p&gt;

&lt;p&gt;Huawei is therefore the closest of these three to a sovereign platform strategy. The objective is not merely efficient inference. It is control over the complete compute system and its supply path.&lt;/p&gt;

&lt;h2&gt;
  
  
  FuriosaAI keeps the facility boundary familiar
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://furiosa.ai/rngd" rel="noopener noreferrer"&gt;RNGD&lt;/a&gt; takes a narrower approach.&lt;/p&gt;

&lt;p&gt;The second-generation FuriosaAI accelerator is a 180-watt PCIe card with 48 GB of HBM3. The company positions it for large-language and multimodal inference, with two cards fitting inside a conventional server power and cooling envelope.&lt;/p&gt;

&lt;p&gt;RNGD’s Tensor Contraction Processor uses a programmable tensor engine rather than exposing a GPU-style collection of fixed matrix units. Furiosa argues that the architecture can map a wider set of tensor shapes while keeping control overhead and data movement low.&lt;/p&gt;

&lt;p&gt;The data-center contract remains recognizable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PCIe attachment;&lt;/li&gt;
&lt;li&gt;standard x86 host servers;&lt;/li&gt;
&lt;li&gt;air-cooled operation;&lt;/li&gt;
&lt;li&gt;HBM on each card;&lt;/li&gt;
&lt;li&gt;Kubernetes and container-based deployment;&lt;/li&gt;
&lt;li&gt;a vendor compiler and runtime behind supported frameworks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That matters because adopting a specialist involves two separate changes. The first is physical: power, cooling, network, server qualification, and sparing. The second is logical: model compilation, operator support, quantization, serving, and observability. A 180-watt PCIe card reduces the first change. It does not remove the second.&lt;/p&gt;

&lt;p&gt;Furiosa’s &lt;a href="https://furiosa.ai/software" rel="noopener noreferrer"&gt;software stack&lt;/a&gt; includes a compiler, runtime, model tooling, and serving integration. The company publishes supported models and performance material, but the usable boundary remains model-specific. An architecture can execute transformers in principle while still lacking a production-ready path for a particular operator, attention variant, quantization method, or serving engine.&lt;/p&gt;

&lt;p&gt;Vendor and partner evidence now includes &lt;a href="https://furiosa.ai/blog/rngd-enters-mass-production-the-high-performance-ai-accelerator-for-any-data-center" rel="noopener noreferrer"&gt;RNGD volume shipments&lt;/a&gt;, a live &lt;a href="https://furiosa.ai/blog/furiosaai-and-samsung-sds" rel="noopener noreferrer"&gt;Samsung SDS NPU-as-a-Service offering&lt;/a&gt;, and announced data-center installations. These sources establish commercial activity, but they are not independently normalized production measurements. The key missing figures are sustained utilization, failure rates, model onboarding time, power at the complete server boundary, and independently comparable service cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rebellions scales a PCIe NPU toward the rack
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://rebellions.ai/resources/" rel="noopener noreferrer"&gt;REBEL&lt;/a&gt; also starts as a standard data-center accelerator.&lt;/p&gt;

&lt;p&gt;The card provides 144 GB of HBM3E and connects through PCIe Gen5. Rebellions describes native support for several low-precision formats and positions the design for large-model inference. A REBEL-Quad system combines four accelerators behind one host, creating a larger local memory and compute domain.&lt;/p&gt;

&lt;p&gt;The architecture then extends through merchant data-center interfaces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PCIe for host attachment;&lt;/li&gt;
&lt;li&gt;CXL-oriented memory expansion in system designs;&lt;/li&gt;
&lt;li&gt;Ethernet for scale-out;&lt;/li&gt;
&lt;li&gt;liquid-cooled rack configurations for higher density;&lt;/li&gt;
&lt;li&gt;a Kubernetes-based operational path.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Rebellions’ &lt;a href="https://docs.rbln.ai/" rel="noopener noreferrer"&gt;RBLN software stack&lt;/a&gt; provides the compiler, runtime, framework integration, model zoo, and serving tools required to turn checkpoints into executable artifacts. As with Corsair and RNGD, that compiler is part of the product. The hardware cannot be evaluated separately from supported operators, model conversion, quantization, and distributed-serving behavior.&lt;/p&gt;

&lt;p&gt;The company has announced cloud and telecom infrastructure work, including a &lt;a href="https://news.sktelecom.com/en/2951" rel="noopener noreferrer"&gt;2026 collaboration with SK Telecom and Arm&lt;/a&gt;, along with server-system partnerships. These announcements show an operating ecosystem. They do not yet create a common benchmark boundary across REBEL, RNGD, Ascend, and GPUs.&lt;/p&gt;

&lt;p&gt;Rebellions occupies the middle of the three strategies. It does not own a full sovereign software and fabric stack like Huawei, but it aims beyond a single add-in card by defining multi-card and rack-scale inference systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  Efficiency claims need a service boundary
&lt;/h2&gt;

&lt;p&gt;Each vendor presents performance-per-watt or cost advantages. Those claims can be true within their test boundary and still fail to answer a production question.&lt;/p&gt;

&lt;p&gt;The useful unit is not peak TOPS. It is a completed unit of model service:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Required boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Prompt length, output length, batch policy, streaming behavior, and percentile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput&lt;/td&gt;
&lt;td&gt;Concurrent requests, quality target, scheduler policy, and complete system&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy&lt;/td&gt;
&lt;td&gt;Model revision, dataset, quantization, and numerical tolerance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power&lt;/td&gt;
&lt;td&gt;Card, host server, network, cooling, or facility boundary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Hardware utilization, software work, support, spares, and facility life&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Portability&lt;/td&gt;
&lt;td&gt;Time and engineering required to qualify another model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Availability&lt;/td&gt;
&lt;td&gt;Card, server, rack, compiler, and service failure rates&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A PCIe NPU can draw less power than a high-end GPU and still produce a worse service if the model does not compile cleanly, the cards remain underutilized, or requests spill back to a costly compatibility pool.&lt;/p&gt;

&lt;p&gt;The reverse is also true. A specialist does not need universal model coverage to create value. It needs enough stable demand for supported models to keep the pool utilized.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two meanings of sovereignty
&lt;/h2&gt;

&lt;p&gt;These architectures separate two meanings of sovereignty.&lt;/p&gt;

&lt;p&gt;One is &lt;strong&gt;supply sovereignty&lt;/strong&gt;: control over silicon, manufacturing access, networking, system design, and software. Huawei is pursuing that broader boundary.&lt;/p&gt;

&lt;p&gt;The other is &lt;strong&gt;operational choice&lt;/strong&gt;: adding accelerators from more than one supplier so stable workloads are not tied to one runtime and pricing model. Furiosa and Rebellions contribute to that boundary without replacing the rest of the data-center stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  What remains unproven
&lt;/h2&gt;

&lt;p&gt;The vendors have demonstrated supported transformer models through published systems, model catalogs, partner deployments, or cloud endpoints. That establishes execution capability for those documented configurations, not broad model coverage or equal production economics.&lt;/p&gt;

&lt;p&gt;The gap is whether the surrounding systems can sustain a broad production workload with predictable onboarding time, quality, utilization, and availability. Public material still lacks a consistent comparison of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the same current model and serving software;&lt;/li&gt;
&lt;li&gt;the same quantization and accuracy target;&lt;/li&gt;
&lt;li&gt;identical latency percentiles and batch policy;&lt;/li&gt;
&lt;li&gt;card, server, rack, and facility power;&lt;/li&gt;
&lt;li&gt;compiler and model-porting effort;&lt;/li&gt;
&lt;li&gt;system availability and replacement time;&lt;/li&gt;
&lt;li&gt;software support across successive model generations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That missing evidence affects each architecture differently. Huawei must prove the breadth and operability of an integrated platform. Furiosa must prove that a narrow, efficient pool can stay busy across enough models. Rebellions must prove that its card-to-rack expansion preserves the efficiency and simplicity promised at the card boundary.&lt;/p&gt;

&lt;p&gt;Alternative AI hardware becomes credible when the complete operating system around the chip is visible. The unresolved question is who will publish enough production evidence to compare that system honestly—without reducing architectural choice to a peak-performance table or treating every alternative as the same kind of accelerator.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>inference</category>
      <category>semiconductors</category>
      <category>sovereignai</category>
    </item>
    <item>
      <title>SambaNova Compiles Models Into a Rack</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:53:28 +0000</pubDate>
      <link>https://dev.to/amitrix/sambanova-compiles-models-into-a-rack-3ob9</link>
      <guid>https://dev.to/amitrix/sambanova-compiles-models-into-a-rack-3ob9</guid>
      <description>&lt;p&gt;SambaNova's most important product is not its accelerator chip. It is a compiled model deployment that spans the chip, three memory tiers, a 16-RDU rack, and the software routing requests across them.&lt;/p&gt;

&lt;p&gt;That system boundary separates SambaNova from a conventional accelerator card. A model configuration becomes a Processor Executable Format binary, its weights occupy a planned memory hierarchy, and SambaStack deploys the result as a model service rather than treating every request as an arbitrary program arriving at a general-purpose processor.&lt;/p&gt;

&lt;p&gt;The thesis: SambaNova compiles models into a rack. The Reconfigurable Dataflow Unit, or RDU, is the execution engine, but the useful commercial unit is the complete inference system around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The RDU maps a model's dataflow graph onto programmable compute, memory, and communication resources instead of repeatedly launching isolated kernels.&lt;/li&gt;
&lt;li&gt;SN40L combines distributed SRAM, HBM, and attached DDR; the memory system is part of model placement, not merely storage behind the processor.&lt;/li&gt;
&lt;li&gt;The mature deployment boundary is SambaRack SN40-16 plus SambaStack on-premises, or the same architecture exposed through SambaCloud.&lt;/li&gt;
&lt;li&gt;SambaNova's software compiles specific model, sequence-length, batch-size, and speculative-decoding configurations into deployable binaries and validates whether a deployment fits available RDU memory.&lt;/li&gt;
&lt;li&gt;SN50 extends the thesis toward larger scale and disaggregated inference, but its strongest public evidence remains preview benchmarks and vendor roadmap claims rather than broad production measurements.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Dataflow changes what the compiler owns
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.nvidia.com/cuda/cuda-programming-guide/" rel="noopener noreferrer"&gt;CUDA's programming model&lt;/a&gt; presents a GPU as a device executing kernels. Each kernel performs a bounded operation, while software schedules launches and moves tensors through the memory hierarchy. Fusion can reduce that traffic, but the execution model remains kernel-oriented.&lt;/p&gt;

&lt;p&gt;SambaNova makes a different cut. Its current &lt;a href="https://sambanova.ai/products/dataflow-architecture" rel="noopener noreferrer"&gt;Dataflow Architecture description&lt;/a&gt; shows a grid of programmable compute units and SRAM-based programmable memory units connected as a streaming pipeline. While one operator computes, the system can fetch data for the next. Intermediate activations stay near the operators consuming them.&lt;/p&gt;

&lt;p&gt;The underlying idea predates the current product naming. SambaNova's architecture material describes an RDU as a tiled array of compute and memory units joined by a programmable communication fabric. The &lt;a href="https://arxiv.org/abs/2405.07518" rel="noopener noreferrer"&gt;SN40L paper presented at MICRO 2024&lt;/a&gt; adds the current system details: dataflow cores contain compute units, memory units, address-generation units, and an on-chip reconfigurable network.&lt;/p&gt;

&lt;p&gt;This does not mean the RDU executes without instructions or memory movement. It means the compiler can plan more of the communication path and keep a longer section of the model graph resident as a spatial pipeline.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Model checkpoint] --&amp;gt; B[Conversion and compilation]
    B --&amp;gt; C[PEF: compiled model configuration]
    A --&amp;gt; D[Model weights]
    C --&amp;gt; E[ModelProfile: supported runtime shapes]
    D --&amp;gt; F[Model: checkpoint and tokenizer]
    E --&amp;gt; G[Model and profile pair]
    F --&amp;gt; G
    G --&amp;gt; H[Optional ModelBundle]
    G --&amp;gt; I[ModelDeployment]
    H --&amp;gt; I
    I --&amp;gt; J[Memory-fit legalizer]
    J --&amp;gt; K[Inference router and runtime]
    K --&amp;gt; L[Attached DDR: model catalog and cache]
    L --&amp;gt; M[HBM: weight streaming]
    M --&amp;gt; N[Distributed SRAM: active data]
    N --&amp;gt; O[PCU and PMU dataflow pipeline]
    O --&amp;gt; P[OpenAI-compatible API response]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The diagram exposes the adoption trade. A familiar checkpoint enters the system, but it does not remain an arbitrary framework program at runtime. The current &lt;a href="https://docs.sambanova.ai/docs/en/sambastack/service-administration/model-deployment/deploying-model-bundles" rel="noopener noreferrer"&gt;SambaStack deployment model&lt;/a&gt; defines a PEF as a compiled model binary for an RDU. A &lt;code&gt;ModelProfile&lt;/code&gt; groups compatible PEFs and the runtime shapes they support. A &lt;code&gt;Model&lt;/code&gt; identifies checkpoints and tokenizers, while a &lt;code&gt;ModelDeployment&lt;/code&gt; creates replicas and a routable endpoint. &lt;code&gt;ModelBundle&lt;/code&gt; is now optional: it groups several model/profile pairs when they need to be deployed, validated, or shared as one unit.&lt;/p&gt;

&lt;p&gt;That compilation boundary is where SambaNova gains control over data movement. It is also where model support, graph coverage, compiler maturity, and artifact lifecycle become commercial constraints.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three memory tiers make the rack the machine
&lt;/h2&gt;

&lt;p&gt;The RDU's second architectural choice is capacity placement. The current &lt;a href="https://sambanova.ai/hubfs/SambaRack%20data%20sheet%20template%2007%2001%2025%20(1).pdf" rel="noopener noreferrer"&gt;SambaRack SN40-16 datasheet&lt;/a&gt; specifies 16 SN40L RDUs with three memory tiers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Memory tier&lt;/th&gt;
&lt;th&gt;SN40-16 published capacity&lt;/th&gt;
&lt;th&gt;System role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Distributed on-RDU SRAM&lt;/td&gt;
&lt;td&gt;8 GB total&lt;/td&gt;
&lt;td&gt;Keeps active operands and intermediate data close to compute&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HBM&lt;/td&gt;
&lt;td&gt;1 TB total&lt;/td&gt;
&lt;td&gt;Streams model weights at high bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RDU-attached DDR&lt;/td&gt;
&lt;td&gt;4 TB total&lt;/td&gt;
&lt;td&gt;Holds checkpoints, prompt cache, and a catalog of model configurations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The same rack contains host CPUs, NVMe storage, management components, and 400/200 GbE networking. SambaNova publishes 10 kW typical and 16.4 kW maximum inference power for the air-cooled rack. Those are rack specifications, not measured energy per token.&lt;/p&gt;

&lt;p&gt;This hierarchy supports a different capacity strategy from placing every active model entirely in HBM. Large DDR capacity can hold more weights and configurations close to the RDUs. HBM provides the high-bandwidth path, while SRAM feeds the active graph. SambaNova's claim is that this enables fast switching among resident models and reduces repeated transfers from external storage.&lt;/p&gt;

&lt;p&gt;The software reveals how tightly memory and deployment are coupled. SambaStack runs a &lt;strong&gt;legalizer&lt;/strong&gt; before deployment to verify that the selected PEFs and checkpoints fit the node's memory constraints. A deployment can include multiple models, sequence-length profiles, batch sizes, and a target/draft pair for speculative decoding. The runtime then selects among compiled configurations according to the request.&lt;/p&gt;

&lt;p&gt;The deployable object is therefore not "a model on an RDU." It is a memory-valid deployment distributed across a rack.&lt;/p&gt;

&lt;h2&gt;
  
  
  The commercial boundary has three forms
&lt;/h2&gt;

&lt;p&gt;SambaNova now exposes the same architectural thesis through different operational boundaries.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Offering&lt;/th&gt;
&lt;th&gt;Customer operates&lt;/th&gt;
&lt;th&gt;SambaNova supplies&lt;/th&gt;
&lt;th&gt;Practical boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://sambanova.ai/products/sambastack" rel="noopener noreferrer"&gt;SambaStack&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data center, network, Kubernetes integration, identity, storage, and service policy&lt;/td&gt;
&lt;td&gt;RDU rack, runtime, management software, compiled artifacts, and serving stack&lt;/td&gt;
&lt;td&gt;Customer-operated private or sovereign inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://sambanova.ai/products/sambamanaged" rel="noopener noreferrer"&gt;SambaManaged&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Facility boundary and application integration&lt;/td&gt;
&lt;td&gt;Fully managed inference infrastructure installed in the customer's data center&lt;/td&gt;
&lt;td&gt;Dedicated capacity without operating the full stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://sambanova.ai/products/sambacloud" rel="noopener noreferrer"&gt;SambaCloud&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;API client, model selection, and application&lt;/td&gt;
&lt;td&gt;Infrastructure, model deployment, routing, and service operation&lt;/td&gt;
&lt;td&gt;Multi-model inference API&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The on-premises path is not a simple appliance with one network cable. SambaNova's &lt;a href="https://docs.sambanova.ai/docs/en/sambastack/getting-started/on-prem" rel="noopener noreferrer"&gt;current installation guide&lt;/a&gt; requires a front-end data network, a private inter-rack network, a management network, customer identity and DNS services, storage, load balancing, and a Kubernetes control plane. SambaStack installs through Helm and exposes deployed models through an OpenAI-compatible API.&lt;/p&gt;

&lt;p&gt;At the other end, &lt;a href="https://docs.sambanova.ai/docs/en/release-notes/sambacloud" rel="noopener noreferrer"&gt;SambaCloud&lt;/a&gt; hides the compiled artifacts and rack operations. Its July 2026 release added prompt caching and continued expanding API compatibility, tool use, structured output, and model support. That surface is easier to adopt, but it offers the curated models and configurations SambaNova has prepared rather than arbitrary code execution.&lt;/p&gt;

&lt;p&gt;This is the commercial read: the buyer chooses how much of the system boundary to own. Infrastructure operators can run a private RDU system, let SambaNova operate dedicated capacity in their facility, or consume the architecture as an API without seeing the compiler boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  SN40L is deployed; SN50 is the next claim
&lt;/h2&gt;

&lt;p&gt;The generation boundary matters because SambaNova's public material now discusses SN40 and SN50 together.&lt;/p&gt;

&lt;p&gt;SN40L is the fourth-generation RDU behind the documented SambaRack and the current on-premises software guides. Independent deployment evidence exists. The &lt;a href="https://www.alcf.anl.gov/support-center/facility-updates/cerebras-cs-3-and-sambanova-sn40l-systems-now-available-users" rel="noopener noreferrer"&gt;Argonne Leadership Computing Facility&lt;/a&gt; opened its Metis system to researchers in October 2025 with two 16-RDU SambaRacks. Its &lt;a href="https://docs.alcf.anl.gov/ai-testbed/sn40l_inference/" rel="noopener noreferrer"&gt;current user guide&lt;/a&gt; describes a larger six-rack SN40L cluster serving multiple models through OpenAI-compatible endpoints.&lt;/p&gt;

&lt;p&gt;Earlier systems also operated alongside conventional supercomputers. &lt;a href="https://www.llnl.gov/article/49821/llnl-sambanova-systems-announce-additional-ai-hardware-support-labs-cognitive-simulation-efforts" rel="noopener noreferrer"&gt;Lawrence Livermore National Laboratory&lt;/a&gt; integrated SambaNova dataflow hardware into its computing environment for cognitive simulation research. That establishes real installation history, although the older system used an earlier RDU generation and does not prove current LLM-serving economics.&lt;/p&gt;

&lt;p&gt;SN50 is SambaNova's fifth-generation inference processor. The company says a 16-chip SambaRack averages 20 kW, offers five times the compute and four times the network bandwidth of SN40, and can connect as many as 256 RDUs across racks. Its &lt;a href="https://sambanova.ai/blog/introducing-the-sn50-rdu-purpose-built-for-agentic-inference" rel="noopener noreferrer"&gt;February 2026 announcement&lt;/a&gt; said customer shipments would begin in the second half of 2026.&lt;/p&gt;

&lt;p&gt;The most interesting SN50 direction is not a larger standalone rack. It is disaggregated inference. SambaNova demonstrated NVIDIA GPUs handling compute-heavy prefill while RDUs handled memory-bound decode. A &lt;a href="https://sambanova.ai/blog/sn50-runs-fastest-minimax-speeds-in-the-world" rel="noopener noreferrer"&gt;July 2026 demonstration&lt;/a&gt; used four H200 GPUs for prefill and one 16-RDU SN50 rack for decode on MiniMax M2.7.&lt;/p&gt;

&lt;p&gt;The evidence has advanced beyond first-party performance charts, but it remains early. SambaNova's &lt;a href="https://sambanova.ai/blog/semianalysis-benchmarks-sambarack-sn50-with-fast-inference-on-minimax-m2.7" rel="noopener noreferrer"&gt;summary of a SemiAnalysis-run benchmark&lt;/a&gt; reports roughly 800 tokens per second at the highest-interactivity point and roughly 400 tokens per second as throughput approached the fastest B200 configuration tested on the same model. The available public summary does not provide the power, utilization, availability, or economic data required for a complete rack comparison.&lt;/p&gt;

&lt;p&gt;That changes the competitive framing. SambaNova does not need to replace every GPU to become useful. It can become a specialized decode tier behind a serving layer that sends each phase to different hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the architecture fits
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Architectural fit&lt;/th&gt;
&lt;th&gt;Evidence boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Interactive LLM decode&lt;/td&gt;
&lt;td&gt;Strong: sequential generation rewards memory bandwidth, locality, and planned execution&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://artificialanalysis.ai/providers/sambanova" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; independently measures high output speed on several SambaCloud models, but endpoint tests do not reveal rack utilization or energy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-model and agent workflows&lt;/td&gt;
&lt;td&gt;Strong in principle: large DDR capacity and deployments keep multiple models and configurations resident&lt;/td&gt;
&lt;td&gt;Fast switching and model bundling are documented primarily by SambaNova&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Private or sovereign inference&lt;/td&gt;
&lt;td&gt;Strong: complete air-cooled racks can run on-premises and expose familiar APIs&lt;/td&gt;
&lt;td&gt;Installation requirements still include Kubernetes, networking, identity, storage, and vendor-provided artifacts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disaggregated prefill and decode&lt;/td&gt;
&lt;td&gt;Promising: GPUs retain prefill while RDUs specialize in decode&lt;/td&gt;
&lt;td&gt;Public SN40/SN50 demonstrations exist; broad production measurements do not&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom fine-tuned models&lt;/td&gt;
&lt;td&gt;Supported through checkpoint conversion, PEFs, and custom bundles&lt;/td&gt;
&lt;td&gt;Support depends on compiler coverage and validated configurations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large-scale model training&lt;/td&gt;
&lt;td&gt;Historically supported by earlier SambaNova systems&lt;/td&gt;
&lt;td&gt;Current SN40/SN50 positioning, software, and evidence concentrate on inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arbitrary or rapidly changing research graphs&lt;/td&gt;
&lt;td&gt;Weaker fit&lt;/td&gt;
&lt;td&gt;Compilation and validated artifacts add friction when operators change models or shapes frequently&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Artificial Analysis currently reports SambaCloud output rates above 700 tokens per second for some &lt;code&gt;gpt-oss-120b&lt;/code&gt; configurations and above 400 tokens per second for MiniMax M2.7. Those measurements establish that the public service can deliver fast single-request generation. They do not establish aggregate rack throughput, sustained concurrency, power efficiency, availability, or total service cost.&lt;/p&gt;

&lt;p&gt;That distinction matters. The architecture is optimized around planned model execution, but a production service also contains queues, routers, caches, network paths, replicas, failures, and model-specific quality choices. Tokens per second from one endpoint are evidence about the service experience—not a complete system benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's missing
&lt;/h2&gt;

&lt;p&gt;SambaNova publishes enough material to understand the shape of the system. It does not publish enough comparable data to close the economic argument.&lt;/p&gt;

&lt;p&gt;There is no current, public evidence set that places SN40 or SN50 beside competing systems with the same model, precision, context distribution, batch policy, concurrency, accuracy target, measured wall power, and availability window. The SN40L MICRO paper is detailed, but its headline comparisons focus on SambaNova's Composition of Experts workload rather than a neutral industry benchmark. The SN50 material is earlier still.&lt;/p&gt;

&lt;p&gt;Compiler coverage is the other open boundary. SambaStack documents stable, preview, deprecated, and removed PEF and checkpoint versions. That is healthy operational machinery, but it also shows that model support is a maintained product matrix. A framework checkpoint being available does not mean every shape, feature, quantization, or serving optimization is immediately deployable on an RDU.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what
&lt;/h2&gt;

&lt;p&gt;SambaNova is best understood as a system for controlling model movement.&lt;/p&gt;

&lt;p&gt;The RDU turns a model graph into a spatial dataflow pipeline. The memory hierarchy decides what remains resident and what streams into active execution. The compiler turns framework artifacts into hardware-specific configurations. SambaStack packages those configurations into a rack-level service, while SambaCloud hides the machinery behind an API.&lt;/p&gt;

&lt;p&gt;That full boundary is the product. Comparing one RDU with one GPU misses it.&lt;/p&gt;

&lt;p&gt;The unresolved question is whether SambaNova can keep compiler and artifact delivery ahead of changing model architectures while proving that phase-specific routing still wins after network, utilization, power, and operational costs are included.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>sambanova</category>
      <category>inference</category>
      <category>dataflow</category>
    </item>
    <item>
      <title>The GPU Stopped Being the Product</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:52:53 +0000</pubDate>
      <link>https://dev.to/amitrix/the-gpu-stopped-being-the-product-34fd</link>
      <guid>https://dev.to/amitrix/the-gpu-stopped-being-the-product-34fd</guid>
      <description>&lt;p&gt;NVIDIA's most important architectural change is not a faster GPU. It is the expansion of what counts as the computer.&lt;/p&gt;

&lt;p&gt;The Tesla P100 arrived as an accelerator inside a server. The A100 made the eight-GPU system a repeatable unit. Hopper joined the CPU and GPU more tightly. Blackwell added a top-end configuration whose scale-up boundary is a liquid-cooled rack. Vera Rubin extends the design across compute racks, networking, power, cooling, and inference orchestration.&lt;/p&gt;

&lt;p&gt;At the high end, the product is no longer only a chip. It is a large part of the AI data center assembled as one machine: Tensor Cores change the math, HBM feeds it, NVLink expands the scale-up domain, networking connects racks, and the software stack makes each generation usable.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;P100 established the modern foundation with HBM2, NVLink, and the first DGX system; V100 added Tensor Cores.&lt;/li&gt;
&lt;li&gt;A100 turned the GPU into elastic infrastructure through TF32, structured sparsity, Multi-Instance GPU, faster NVLink, and a standardized eight-GPU NVSwitch system.&lt;/li&gt;
&lt;li&gt;Hopper optimized transformers directly through FP8 and Transformer Engine, while Grace Hopper introduced coherent CPU-GPU coupling.&lt;/li&gt;
&lt;li&gt;Blackwell introduced a top-end 72-GPU liquid-cooled rack with 130 TB/s of summed bidirectional NVLink endpoint bandwidth, while B200 also remained available in server configurations.&lt;/li&gt;
&lt;li&gt;Vera Rubin doubles NVLink bandwidth again and raises HBM bandwidth to 22 TB/s per GPU. Groq-derived LPX racks add specialized inference paths, while Rubin remains in production ramp rather than broad availability.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Start before A100
&lt;/h2&gt;

&lt;p&gt;The useful starting point is the &lt;a href="https://images.nvidia.com/content/pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;Tesla P100 and Pascal GP100 architecture&lt;/a&gt; in 2016. P100 introduced three pieces that still define NVIDIA's systems:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;HBM on the GPU package&lt;/strong&gt; for high-bandwidth access to model data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NVLink&lt;/strong&gt; for GPU-to-GPU communication beyond PCIe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DGX-1&lt;/strong&gt;, an integrated eight-GPU server with a tested software stack.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;P100 did not have Tensor Cores. Its AI acceleration came from running FP16 arithmetic at twice the rate of FP32 on ordinary CUDA cores. The &lt;a href="https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;Tesla V100 Volta architecture&lt;/a&gt; made the next decisive move in 2017: 640 first-generation Tensor Cores dedicated to matrix multiplication.&lt;/p&gt;

&lt;p&gt;Pascal made the GPU a better parallel processor. Volta began shaping its numerical machinery around neural networks.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    P["P100 Pascal&amp;lt;br/&amp;gt;2016&amp;lt;br/&amp;gt;HBM2 + NVLink + DGX-1"] --&amp;gt; V["V100 Volta&amp;lt;br/&amp;gt;2017&amp;lt;br/&amp;gt;Tensor Cores"]
    V --&amp;gt; A["A100 Ampere&amp;lt;br/&amp;gt;2020&amp;lt;br/&amp;gt;TF32 + sparsity + MIG"]
    A --&amp;gt; H["H100/H200 Hopper&amp;lt;br/&amp;gt;2022–2024&amp;lt;br/&amp;gt;FP8 + Transformer Engine"]
    H --&amp;gt; B["B200 Blackwell&amp;lt;br/&amp;gt;2024 architecture&amp;lt;br/&amp;gt;dual-die GPU + NVL72"]
    B --&amp;gt; BU["B300 Blackwell Ultra&amp;lt;br/&amp;gt;2025 architecture&amp;lt;br/&amp;gt;288 GB HBM3e"]
    BU --&amp;gt; R["Vera Rubin&amp;lt;br/&amp;gt;2026 production ramp&amp;lt;br/&amp;gt;HBM4 + NVLink 6"]
    R -. roadmap .-&amp;gt; RU["Rubin Ultra&amp;lt;br/&amp;gt;2027"]
    RU -. roadmap .-&amp;gt; F["Feynman&amp;lt;br/&amp;gt;2028"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The names hide overlapping product lines. H200 is a Hopper memory upgrade, while Blackwell Ultra expands Blackwell's memory and low-precision throughput. Grace Hopper, GB200, and Vera Rubin are platforms spanning CPUs, GPUs, memory, interconnects, and systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  The generation map
&lt;/h2&gt;

&lt;p&gt;The table uses representative high-end data-center configurations. SXM and rack-scale modules run at different power and bandwidth levels from PCIe cards with the same architecture. Peak Tensor Core figures also change numerical format across generations, so they are not one continuous performance benchmark.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;Architectural step&lt;/th&gt;
&lt;th&gt;Representative memory&lt;/th&gt;
&lt;th&gt;Scale-up fabric per GPU&lt;/th&gt;
&lt;th&gt;Deployment unit&lt;/th&gt;
&lt;th&gt;Status on Aug. 30, 2026&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://images.nvidia.com/content/pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;P100 Pascal&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;HBM2, NVLink 1, high-rate FP16&lt;/td&gt;
&lt;td&gt;16 GB, 0.72 TB/s&lt;/td&gt;
&lt;td&gt;0.16 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;Eight-GPU DGX-1 with a direct NVLink mesh&lt;/td&gt;
&lt;td&gt;Legacy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;V100 Volta&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;First Tensor Cores, NVLink 2&lt;/td&gt;
&lt;td&gt;16–32 GB, 0.9 TB/s&lt;/td&gt;
&lt;td&gt;0.30 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;DGX-1 mesh; DGX-2 introduced NVSwitch&lt;/td&gt;
&lt;td&gt;Legacy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/nvidia-ampere-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;A100 Ampere&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;TF32, BF16, FP64 Tensor Cores, 2:4 sparsity, MIG&lt;/td&gt;
&lt;td&gt;40–80 GB, up to 2.0 TB/s&lt;/td&gt;
&lt;td&gt;0.60 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;Eight-GPU DGX/HGX NVSwitch baseboard&lt;/td&gt;
&lt;td&gt;Mature production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet" rel="noopener noreferrer"&gt;H100 Hopper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;FP8, Transformer Engine, Tensor Memory Accelerator&lt;/td&gt;
&lt;td&gt;80 GB, 3.35 TB/s&lt;/td&gt;
&lt;td&gt;0.90 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;Eight-GPU HGX/DGX; external NVLink Switch systems&lt;/td&gt;
&lt;td&gt;Mature production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/h200/" rel="noopener noreferrer"&gt;H200 Hopper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Hopper compute with larger, faster HBM3e&lt;/td&gt;
&lt;td&gt;141 GB, 4.8 TB/s&lt;/td&gt;
&lt;td&gt;0.90 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;H100-compatible system generation&lt;/td&gt;
&lt;td&gt;Mature production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/" rel="noopener noreferrer"&gt;B200 Blackwell&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Two reticle-sized dies act as one GPU; FP4; NVLink 5&lt;/td&gt;
&lt;td&gt;180 GB, up to 8 TB/s&lt;/td&gt;
&lt;td&gt;1.8 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;DGX/HGX server or GB200 NVL72 rack&lt;/td&gt;
&lt;td&gt;Production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/gb300-nvl72/" rel="noopener noreferrer"&gt;B300 Blackwell Ultra&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;288 GB HBM3e, more NVFP4 and attention throughput&lt;/td&gt;
&lt;td&gt;288 GB, up to 8 TB/s&lt;/td&gt;
&lt;td&gt;1.8 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;GB300 NVL72 rack&lt;/td&gt;
&lt;td&gt;Production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/" rel="noopener noreferrer"&gt;Vera Rubin&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;HBM4, sixth-generation Tensor Cores, NVLink 6, Vera CPU&lt;/td&gt;
&lt;td&gt;288 GB, 22 TB/s&lt;/td&gt;
&lt;td&gt;3.6 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;Vera Rubin NVL72 and pod-scale platform&lt;/td&gt;
&lt;td&gt;Production ramp; shipments scheduled for fall 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://nvidianews.nvidia.com/news/gtc-2025-news" rel="noopener noreferrer"&gt;Rubin Ultra and Feynman&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Larger scale-up domain, then next named architecture&lt;/td&gt;
&lt;td&gt;Not fully disclosed&lt;/td&gt;
&lt;td&gt;Not fully disclosed&lt;/td&gt;
&lt;td&gt;2027 and 2028 roadmap&lt;/td&gt;
&lt;td&gt;Roadmap&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The status boundary matters. On August 27, &lt;a href="https://blogs.nvidia.com/blog/vera-cpu-delivery/" rel="noopener noreferrer"&gt;AWS received its first Vera CPU server and Rubin GPU&lt;/a&gt; for deployment work. NVIDIA still describes &lt;a href="https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer" rel="noopener noreferrer"&gt;production shipments as beginning in fall 2026&lt;/a&gt;. Early component delivery is not general availability of complete NVL72 systems.&lt;/p&gt;

&lt;p&gt;Three caveats matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;NVIDIA's NVLink numbers add transmit and receive bandwidth. They are not one-way payload rates.&lt;/li&gt;
&lt;li&gt;Sparse throughput assumes the required structured sparsity pattern. It is not dense performance.&lt;/li&gt;
&lt;li&gt;FP16, TF32, FP8, FP4, and NVFP4 trade precision, range, storage, and compute differently. A larger headline FLOPS number at a narrower format does not mean every workload runs proportionally faster.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Component 1: Compute changed from general arithmetic to model-aware arithmetic
&lt;/h2&gt;

&lt;p&gt;P100 accelerated FP16 on CUDA cores. V100 added dedicated Tensor Cores that multiplied FP16 matrices and accumulated in FP32, giving neural-network matrix operations their own hardware.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/nvidia-ampere-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;A100's third-generation Tensor Cores&lt;/a&gt; added TF32, BF16, IEEE FP64, INT8, INT4, and 2:4 structured sparsity. One matrix engine now covered training, inference, and scientific computing.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet" rel="noopener noreferrer"&gt;H100 architecture&lt;/a&gt; added FP8 and the first Transformer Engine, coordinating lower- and higher-precision execution by transformer layer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/" rel="noopener noreferrer"&gt;Blackwell&lt;/a&gt; extended that direction to FP6 and FP4 with a second-generation Transformer Engine. &lt;a href="https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/" rel="noopener noreferrer"&gt;Rubin's third-generation Transformer Engine&lt;/a&gt; adds sixth-generation Tensor Cores and up to 50 petaflops of NVFP4 inference compute according to NVIDIA. That is a vendor peak, not a cross-generation application benchmark.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Generation&lt;/th&gt;
&lt;th&gt;Numerical shift&lt;/th&gt;
&lt;th&gt;What changed architecturally&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pascal&lt;/td&gt;
&lt;td&gt;FP16 on CUDA cores&lt;/td&gt;
&lt;td&gt;AI uses the general parallel processor more efficiently&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Volta&lt;/td&gt;
&lt;td&gt;FP16 Tensor Cores with FP32 accumulation&lt;/td&gt;
&lt;td&gt;Matrix multiplication receives dedicated hardware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ampere&lt;/td&gt;
&lt;td&gt;TF32, BF16, FP64 Tensor Cores, INT8/4, 2:4 sparsity&lt;/td&gt;
&lt;td&gt;One matrix engine spans training, inference, and HPC&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hopper&lt;/td&gt;
&lt;td&gt;FP8 plus Transformer Engine&lt;/td&gt;
&lt;td&gt;Precision selection becomes part of model execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blackwell&lt;/td&gt;
&lt;td&gt;FP6, FP4, microscaling&lt;/td&gt;
&lt;td&gt;Low-precision formats become finer-grained and model-aware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rubin&lt;/td&gt;
&lt;td&gt;Expanded NVFP4/NVFP6 and third-generation Transformer Engine&lt;/td&gt;
&lt;td&gt;Precision, attention, sparsity, and communication are co-optimized&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Component 2: Memory became as important as compute
&lt;/h2&gt;

&lt;p&gt;Neural networks repeatedly move weights, activations, and cache state. More arithmetic units do not help when those units wait for data.&lt;/p&gt;

&lt;p&gt;P100's HBM2 placed stacked memory beside the GPU on a silicon interposer. It delivered 16 GB at 720 GB/s. V100 raised bandwidth to 900 GB/s. The original A100 reached 1.555 TB/s with 40 GB, and the later &lt;a href="https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/a100-80gb-datasheet-update-nvidia-us-1521051-r2-web.pdf" rel="noopener noreferrer"&gt;A100 80 GB&lt;/a&gt; exceeded 2 TB/s.&lt;/p&gt;

&lt;p&gt;H100 increased bandwidth to 3.35 TB/s, but H200 demonstrates why capacity and bandwidth must be separated. H200 retained Hopper compute while moving to 141 GB of HBM3e at 4.8 TB/s. On memory-intensive large-model inference, that change can matter without a new Tensor Core generation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/" rel="noopener noreferrer"&gt;Blackwell&lt;/a&gt; increased the representative B200 package to 180 GB at up to 8 TB/s. &lt;a href="https://www.nvidia.com/en-us/data-center/gb300-nvl72/" rel="noopener noreferrer"&gt;Blackwell Ultra&lt;/a&gt; reached 288 GB. &lt;a href="https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/" rel="noopener noreferrer"&gt;Rubin&lt;/a&gt; keeps 288 GB but replaces HBM3e with HBM4 and raises the vendor-published peak to 22 TB/s.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;xychart-beta
    title "Selected vendor-published peak HBM bandwidth"
    x-axis ["P100", "V100", "A100 80GB", "H100", "H200", "B200", "Rubin"]
    y-axis "TB/s" 0 --&amp;gt; 24
    bar [0.72, 0.90, 2.04, 3.35, 4.80, 8.00, 22.00]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The chart keeps one physical quantity on the axis, but it combines product introductions, later memory variants, and Rubin's production-ramp specification. Every value is a published peak rather than achieved application bandwidth.&lt;/p&gt;

&lt;p&gt;Rubin's published HBM bandwidth is more than 30 times P100's. Power and cooling rose with it: &lt;a href="https://images.nvidia.com/content/pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;P100 SXM&lt;/a&gt; was rated at 300 watts, &lt;a href="https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/nvidia-ampere-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;A100 SXM&lt;/a&gt; at 400 watts, &lt;a href="https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet" rel="noopener noreferrer"&gt;H100 SXM&lt;/a&gt; at up to 700 watts, and &lt;a href="https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/" rel="noopener noreferrer"&gt;B200&lt;/a&gt; at up to 1,000 watts per GPU.&lt;/p&gt;

&lt;h2&gt;
  
  
  Component 3: Packaging and CPU coupling expanded the processor boundary
&lt;/h2&gt;

&lt;p&gt;P100, V100, A100, and H100 were large monolithic GPU dies. Each generation pushed close to the practical manufacturing limit for one piece of silicon.&lt;/p&gt;

&lt;p&gt;Blackwell changed the package. &lt;a href="https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/" rel="noopener noreferrer"&gt;Two reticle-limited dies connect through a 10 TB/s chip-to-chip interface&lt;/a&gt; and present themselves as one CUDA GPU. NVIDIA could increase transistor count to 208 billion without requiring one manufacturable die of that size.&lt;/p&gt;

&lt;p&gt;Rubin continues the dual-reticle approach with 336 billion transistors. NVIDIA calls the inter-die connection NV-HBI. From the programming layer, the goal remains one GPU abstraction. Underneath it, the package has become a tightly integrated multi-die system.&lt;/p&gt;

&lt;p&gt;Once a "GPU" contains multiple compute dies, HBM stacks, high-speed die links, NVLink interfaces, and control engines, the package is already a system.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://developer.nvidia.com/blog/nvidia-grace-hopper-superchip-architecture-in-depth/" rel="noopener noreferrer"&gt;Grace Hopper&lt;/a&gt; tightened that relationship. A Grace CPU and Hopper GPU communicate over coherent NVLink-C2C, giving the CPU and GPU a higher-bandwidth path and a coherent memory relationship inside one superchip.&lt;/p&gt;

&lt;p&gt;Blackwell's GB200 platform connects one Grace CPU to two Blackwell GPUs. The GB200 NVL72 rack contains 36 Grace CPUs and 72 Blackwell GPUs. Vera Rubin replaces Grace with the Vera CPU while retaining the 36-to-72 relationship.&lt;/p&gt;

&lt;p&gt;This creates two different fabrics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NVLink-C2C&lt;/strong&gt; connects the CPU and GPU inside a superchip.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NVLink and NVSwitch&lt;/strong&gt; connect GPUs across the scale-up domain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;PCIe remains the host and peripheral fabric, while NVLink-C2C and NVLink carry the tightly coupled processor paths.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    APP["Models and applications"]
    ORCH["Inference orchestration&amp;lt;br/&amp;gt;Dynamo, routing, cache coordination"]
    SERVE["Serving and model runtimes&amp;lt;br/&amp;gt;Triton, TensorRT-LLM"]
    LIBS["Accelerated libraries&amp;lt;br/&amp;gt;cuDNN, cuBLAS, NCCL, NVSHMEM"]
    CUDA["CUDA platform&amp;lt;br/&amp;gt;compiler, runtime, graphs, memory APIs"]
    CPU["Grace or Vera CPU&amp;lt;br/&amp;gt;system memory and control"]
    GPU["GPU package&amp;lt;br/&amp;gt;Tensor Cores, cache, HBM"]
    SCALEUP["NVLink + NVSwitch&amp;lt;br/&amp;gt;scale-up domain"]
    NODE["Server or rack I/O domain"]
    SCALEOUT["InfiniBand or Spectrum-X&amp;lt;br/&amp;gt;scale-out fabric"]
    OPS["NGC, GPU Operator, Mission Control&amp;lt;br/&amp;gt;deployment and operations"]

    APP --&amp;gt; ORCH --&amp;gt; SERVE --&amp;gt; LIBS --&amp;gt; CUDA --&amp;gt; GPU
    CPU &amp;lt;--&amp;gt;|"NVLink-C2C"| GPU
    GPU &amp;lt;--&amp;gt;|"NVLink"| SCALEUP
    SCALEUP --&amp;gt; NODE
    NODE &amp;lt;--&amp;gt;|"ConnectX / BlueField"| SCALEOUT
    OPS -. manages .-&amp;gt; SERVE
    OPS -. manages .-&amp;gt; CUDA
    OPS -. manages .-&amp;gt; SCALEOUT&lt;/code&gt;&lt;/pre&gt;



&lt;h2&gt;
  
  
  Component 4: NVLink turned multiple GPUs into one scale-up computer
&lt;/h2&gt;

&lt;p&gt;PCIe connects devices to a host. NVLink exists because tightly coupled GPU workloads need a different path.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://images.nvidia.com/content/pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;P100&lt;/a&gt; supported four first-generation NVLinks for 160 GB/s of aggregate bidirectional bandwidth per GPU. DGX-1 arranged eight GPUs in a hybrid cube mesh, so not every pair had a dedicated direct connection.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;V100&lt;/a&gt; increased per-GPU NVLink bandwidth to 300 GB/s. DGX-2 then introduced NVSwitch, replacing the fixed mesh with a switched 16-GPU fabric.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/nvidia-ampere-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;A100&lt;/a&gt; doubled per-GPU bandwidth to 600 GB/s and standardized an eight-GPU fully connected NVSwitch baseboard. &lt;a href="https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet" rel="noopener noreferrer"&gt;H100&lt;/a&gt; reached 900 GB/s, while external NVLink Switch systems stretched the scale-up fabric beyond one server.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/" rel="noopener noreferrer"&gt;Blackwell&lt;/a&gt; moved the fabric across an entire rack. GB200 and GB300 NVL72 connect 72 GPUs through fifth-generation NVLink at 1.8 TB/s per GPU, producing NVIDIA's rounded 130 TB/s aggregate endpoint figure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72" rel="noopener noreferrer"&gt;Rubin&lt;/a&gt; doubles the per-GPU figure to 3.6 TB/s and the 72-GPU aggregate to 260 TB/s.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;xychart-beta
    title "Published aggregate bidirectional NVLink bandwidth per GPU"
    x-axis ["P100", "V100", "A100", "H100", "B200", "Rubin"]
    y-axis "TB/s" 0 --&amp;gt; 4
    line [0.16, 0.30, 0.60, 0.90, 1.80, 3.60]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The directionality deserves emphasis. A100's 600 GB/s means 300 GB/s in each direction. Blackwell's 1.8 TB/s and Rubin's 3.6 TB/s are also bidirectional endpoint totals. The 130 TB/s and 260 TB/s rack numbers multiply those endpoint figures by 72 GPUs. They do not state the fabric's one-way bisection bandwidth.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;NVLink generation&lt;/th&gt;
&lt;th&gt;Introduction&lt;/th&gt;
&lt;th&gt;Per-GPU published bandwidth&lt;/th&gt;
&lt;th&gt;Representative scale-up domain&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;NVLink 1&lt;/td&gt;
&lt;td&gt;&lt;a href="https://images.nvidia.com/content/pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;P100&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;160 GB/s bidirectional&lt;/td&gt;
&lt;td&gt;Eight-GPU direct mesh&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVLink 2&lt;/td&gt;
&lt;td&gt;&lt;a href="https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;V100&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;300 GB/s bidirectional&lt;/td&gt;
&lt;td&gt;Eight-GPU mesh; 16 GPUs with NVSwitch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVLink 3&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/nvidia-ampere-architecture-whitepaper.pdf" rel="noopener noreferrer"&gt;A100&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;600 GB/s bidirectional&lt;/td&gt;
&lt;td&gt;Eight-GPU fully switched baseboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVLink 4&lt;/td&gt;
&lt;td&gt;&lt;a href="https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet" rel="noopener noreferrer"&gt;H100&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;900 GB/s bidirectional&lt;/td&gt;
&lt;td&gt;Eight-GPU HGX; larger Grace Hopper switch systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVLink 5&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/" rel="noopener noreferrer"&gt;Blackwell&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;1.8 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;72-GPU NVL72 rack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVLink 6&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72" rel="noopener noreferrer"&gt;Rubin&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;3.6 TB/s bidirectional&lt;/td&gt;
&lt;td&gt;72-GPU NVL72; larger 2027 roadmap domain&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The scale-up domain supports frequent, low-latency communication inside one model execution. Scale-out networking joins systems and racks into a larger cluster.&lt;/p&gt;

&lt;h2&gt;
  
  
  Component 5: Networking and cooling expanded the deployment boundary
&lt;/h2&gt;

&lt;p&gt;NVIDIA's &lt;a href="https://nvidianews.nvidia.com/news/nvidia-completes-acquisition-of-mellanox-creating-major-force-driving-next-gen-data-centers" rel="noopener noreferrer"&gt;2020 acquisition of Mellanox&lt;/a&gt; brought the network into the platform.&lt;/p&gt;

&lt;p&gt;Mellanox added several distinct components:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ConnectX&lt;/strong&gt; network adapters move data between GPU systems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quantum InfiniBand&lt;/strong&gt; provides a purpose-built scale-out fabric for large compute clusters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spectrum-X Ethernet&lt;/strong&gt; combines Ethernet switches, adapters, congestion control, and software for AI traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BlueField DPUs&lt;/strong&gt; handle infrastructure work including networking, storage, isolation, and security.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Successive NVIDIA systems used 100 Gb/s EDR, 200 Gb/s HDR, 400 Gb/s NDR or Spectrum-X, and then &lt;a href="https://nvidianews.nvidia.com/news/spectrum-x-ethernet-networking-platform-for-ai-expands-to-support-coreweave-gpu-cloud" rel="noopener noreferrer"&gt;800 Gb/s Quantum-X800 and Spectrum-X800&lt;/a&gt;. Rubin adds &lt;a href="https://nvidianews.nvidia.com/news/nvidia-connectx-9-supernic" rel="noopener noreferrer"&gt;ConnectX-9&lt;/a&gt; while retaining InfiniBand and Ethernet scale-out paths.&lt;/p&gt;

&lt;p&gt;NVLink and InfiniBand are not competing names for the same connection:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Fabric&lt;/th&gt;
&lt;th&gt;Primary job&lt;/th&gt;
&lt;th&gt;Typical boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PCIe&lt;/td&gt;
&lt;td&gt;Host and peripheral attachment&lt;/td&gt;
&lt;td&gt;CPU, GPU, NIC, storage inside a system&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVLink-C2C&lt;/td&gt;
&lt;td&gt;Coherent processor coupling&lt;/td&gt;
&lt;td&gt;CPU to GPU inside a superchip&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVLink + NVSwitch&lt;/td&gt;
&lt;td&gt;High-bandwidth GPU scale-up&lt;/td&gt;
&lt;td&gt;GPUs inside a server, rack, or tightly coupled domain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;InfiniBand / Spectrum-X&lt;/td&gt;
&lt;td&gt;System scale-out&lt;/td&gt;
&lt;td&gt;Servers and racks across a cluster&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The deployment unit grew with the fabric. The first DGX-1 was a 3U, air-cooled server rated at 3.2 kW. &lt;a href="https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/dgx-a100/dgxa100-system-architecture-white-paper.pdf" rel="noopener noreferrer"&gt;DGX A100&lt;/a&gt; grew to 6U and 6.5 kW, while &lt;a href="https://docs.nvidia.com/dgx/dgxh100-user-guide/introduction-to-dgxh100.html" rel="noopener noreferrer"&gt;DGX H100&lt;/a&gt; reached roughly 10.2 kW.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.nvidia.com/dgx/dgxgb200-user-guide/" rel="noopener noreferrer"&gt;GB200 NVL72&lt;/a&gt; changes the boundary: 72 GPUs, 36 CPUs, nine NVSwitch trays, power shelves, networking, and direct-to-chip liquid cooling form one rack-scale system. NVIDIA documents approximately 120 kW for that rack, while &lt;a href="https://docs.nvidia.com/dgx/dgxgb200-user-guide/hardware.html" rel="noopener noreferrer"&gt;GB300 NVL72&lt;/a&gt; raises the published envelope to about 142 kW.&lt;/p&gt;

&lt;p&gt;Those figures do not compare efficiency because the measured boundary grows from one eight-GPU server to an entire 72-GPU rack. They show the architectural transition: electrical delivery, cooling, switching, and compute are now specified together.&lt;/p&gt;

&lt;h2&gt;
  
  
  Component 6: CUDA expanded into an operating stack for AI infrastructure
&lt;/h2&gt;

&lt;p&gt;CUDA is the longest-running source of continuity. The CUDA programming model, drivers, compilers, and libraries let applications survive multiple hardware generations without starting over.&lt;/p&gt;

&lt;p&gt;The stack around CUDA kept expanding:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Representative components&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Programming platform&lt;/td&gt;
&lt;td&gt;&lt;a href="https://docs.nvidia.com/cuda/cuda-programming-guide/" rel="noopener noreferrer"&gt;CUDA&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Kernels, memory management, graphs, compilation, runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accelerated libraries&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://docs.nvidia.com/cuda/cublas/" rel="noopener noreferrer"&gt;cuBLAS&lt;/a&gt;, &lt;a href="https://docs.nvidia.com/deeplearning/cudnn/" rel="noopener noreferrer"&gt;cuDNN&lt;/a&gt;, &lt;a href="https://docs.nvidia.com/deeplearning/nccl/" rel="noopener noreferrer"&gt;NCCL&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Matrix math, neural-network operations, multi-GPU collectives&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Precision layer&lt;/td&gt;
&lt;td&gt;&lt;a href="https://docs.nvidia.com/deeplearning/transformer-engine/" rel="noopener noreferrer"&gt;Transformer Engine&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;FP8 and newer precision recipes for transformer models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model optimization&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://docs.nvidia.com/deeplearning/tensorrt/latest/" rel="noopener noreferrer"&gt;TensorRT&lt;/a&gt;, &lt;a href="https://nvidia.github.io/TensorRT-LLM/" rel="noopener noreferrer"&gt;TensorRT-LLM&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Graph optimization, quantization, kernels, KV cache, distributed LLM execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Serving&lt;/td&gt;
&lt;td&gt;&lt;a href="https://docs.nvidia.com/deeplearning/triton-inference-server/" rel="noopener noreferrer"&gt;Triton Inference Server&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Model serving, batching, ensembles, multiple backends&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fleet orchestration&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://docs.nvidia.com/dynamo/" rel="noopener noreferrer"&gt;Dynamo&lt;/a&gt; and NIXL&lt;/td&gt;
&lt;td&gt;Request routing, disaggregated prefill and decode, cache and tensor movement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Packaging and operations&lt;/td&gt;
&lt;td&gt;NGC containers, Container Toolkit, GPU Operator, Mission Control&lt;/td&gt;
&lt;td&gt;Tested software combinations, deployment, drivers, scheduling, lifecycle management&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource isolation&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://docs.nvidia.com/datacenter/tesla/mig-user-guide/" rel="noopener noreferrer"&gt;MIG&lt;/a&gt; and vGPU&lt;/td&gt;
&lt;td&gt;Partitioning and sharing accelerator capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Volta exposed Tensor Cores through CUDA 9 and accelerated libraries. Ampere added TF32, sparsity, CUDA Graphs, and MIG. Hopper paired FP8 hardware with Transformer Engine. Blackwell paired FP4 with TensorRT-LLM, NeMo, and rack-scale NVLink. Dynamo moved optimization above the model server into routing, KV-cache placement, prefill/decode separation, and worker-to-worker data movement.&lt;/p&gt;

&lt;p&gt;CUDA compatibility is valuable, but it is not magic portability. The &lt;a href="https://docs.nvidia.com/deploy/cuda-compatibility/" rel="noopener noreferrer"&gt;CUDA compatibility model&lt;/a&gt; has minimum driver requirements and version boundaries. Native cubins target specific compute capabilities. TensorRT engines can be tied to library versions and GPU assumptions. MIG profiles differ by architecture. An optimized FP8 or NVFP4 execution recipe may need rebuilding for a different generation.&lt;/p&gt;

&lt;p&gt;The application logic may travel. The highest-performance execution artifact often does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Groq adds a specialized inference path
&lt;/h2&gt;

&lt;p&gt;The Groq relationship adds one important direction to NVIDIA's architecture: the GPU does not need to execute every phase of inference.&lt;/p&gt;

&lt;p&gt;This was not an acquisition of Groq. In December 2025, &lt;a href="https://groq.com/newsroom/groq-and-nvidia-enter-non-exclusive-inference-technology-licensing-agreement-to-accelerate-ai-inference-at-global-scale" rel="noopener noreferrer"&gt;Groq announced a non-exclusive license of its inference technology to NVIDIA&lt;/a&gt;. Several Groq leaders and team members joined NVIDIA, while Groq remained an independent company and continued operating GroqCloud.&lt;/p&gt;

&lt;p&gt;NVIDIA used the licensed technology in the &lt;a href="https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/" rel="noopener noreferrer"&gt;Groq 3 LPX rack&lt;/a&gt;, which it positions beside Vera Rubin NVL72. One representative pattern separates prefill from decode: Rubin processes the prompt and creates the KV cache, LPX generates output tokens, and &lt;a href="https://developer.nvidia.com/dynamo" rel="noopener noreferrer"&gt;NVIDIA Dynamo&lt;/a&gt; coordinates routing and data movement.&lt;/p&gt;

&lt;p&gt;The pattern matters more than the product pairing. NVIDIA is extending its platform around heterogeneous inference—general-purpose GPUs for broad model execution, specialized processors for latency-sensitive phases, and software that decides where each phase runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  What measured performance can actually tell us
&lt;/h2&gt;

&lt;p&gt;Peak specifications describe hardware ceilings. Vendor platform claims describe selected workloads and configurations. Neither is a substitute for a controlled benchmark.&lt;/p&gt;

&lt;p&gt;The methodologically cleanest comparison in the available results is &lt;a href="https://github.com/mlcommons/training_results_v0.7" rel="noopener noreferrer"&gt;MLPerf Training v0.7&lt;/a&gt;. NVIDIA submitted an eight-GPU V100 DGX-1 and an eight-GPU DGX A100 to the same BERT benchmark under the same release. Applying &lt;a href="https://github.com/mlcommons/training_policies/blob/master/training_rules.adoc" rel="noopener noreferrer"&gt;MLPerf's rule of discarding the fastest and slowest runs and averaging the remaining eight&lt;/a&gt; gives roughly 166.2 minutes for V100 and 49.0 minutes for A100—about 3.4 times faster.&lt;/p&gt;

&lt;p&gt;The result combines hardware and generation-appropriate software rather than isolating the chip—which matches how users experience the system.&lt;/p&gt;

&lt;p&gt;Later comparisons are harder:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://developer.nvidia.com/blog/leading-mlperf-training-2-1-with-full-stack-optimizations-for-ai/" rel="noopener noreferrer"&gt;NVIDIA reported up to 6.7 times higher H100 performance&lt;/a&gt; than its first A100 submission in MLPerf Training v2.1, but "up to" selects the strongest workload and spans different submission vintages.&lt;/li&gt;
&lt;li&gt;In &lt;a href="https://developer.nvidia.com/blog/nvidia-h200-tensor-core-gpus-and-nvidia-tensorrt-llm-set-mlperf-llm-inference-records/" rel="noopener noreferrer"&gt;MLPerf Inference v4.0&lt;/a&gt;, NVIDIA reported about 28% greater H200 performance than H100 at the same 700-watt envelope and up to 45% at 1,000 watts. The equal-power result is the cleaner evidence for HBM3e's contribution.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://mlcommons.org/benchmarks/training/" rel="noopener noreferrer"&gt;Blackwell and Blackwell Ultra have measured MLPerf submissions&lt;/a&gt;, but accelerator count, numerical format, model, power, and software must match before calculating a ratio.&lt;/li&gt;
&lt;li&gt;Rubin had no comparable public MLPerf submission by August 30, 2026. Its performance and cost claims remain NVIDIA measurements until reproducible benchmark systems appear.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is no honest single line for "GPU performance over time." Workload, precision, sparsity, GPU count, power, and software all move.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's missing
&lt;/h2&gt;

&lt;p&gt;The stack still lacks a neutral system benchmark that matches how current AI infrastructure is purchased and operated.&lt;/p&gt;

&lt;p&gt;MLPerf is the strongest public mechanism available, but the product boundary keeps expanding faster than the benchmark boundary. A useful rack-scale comparison would hold all of these constant:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;model and quality target;&lt;/li&gt;
&lt;li&gt;prompt and output-length distribution, including concurrency;&lt;/li&gt;
&lt;li&gt;time to first token and inter-token latency;&lt;/li&gt;
&lt;li&gt;sustained throughput under concurrency;&lt;/li&gt;
&lt;li&gt;training or inference precision;&lt;/li&gt;
&lt;li&gt;dense versus sparse execution;&lt;/li&gt;
&lt;li&gt;rack, network, cooling, and facility power;&lt;/li&gt;
&lt;li&gt;software engineering required to reach the result;&lt;/li&gt;
&lt;li&gt;failure recovery and degraded-operation behavior;&lt;/li&gt;
&lt;li&gt;acquisition price and cost per completed unit of work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without that boundary, a vendor can compare one Blackwell rack with many Hopper servers, or FP4 with FP8, or projected utilization with measured throughput. The numbers may all be correct while the comparison remains unhelpful.&lt;/p&gt;

&lt;p&gt;Rubin makes the gap larger. NVIDIA now describes benefits from GPU architecture, HBM4, NVLink 6, Vera CPUs, Spectrum-X, power smoothing, liquid cooling, Dynamo, and workload scheduling in one platform claim. That is a reasonable systems argument. It also makes independent reproduction much harder.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what
&lt;/h2&gt;

&lt;p&gt;NVIDIA's evolution from P100 to Rubin is the story of bottlenecks moving outward.&lt;/p&gt;

&lt;p&gt;P100 attacked memory and PCIe. V100 attacked matrix math. A100 attacked precision, utilization, and multi-tenancy. Hopper attacked transformer execution. Blackwell attacked the server boundary. Rubin attacks the rack and pod as operating systems for continuous inference.&lt;/p&gt;

&lt;p&gt;That changes how to evaluate the company and its competitors. Comparing accelerator FLOPS is now the narrowest possible cut. The real comparison is the amount of useful model work a complete system produces within its power, latency, reliability, and software constraints.&lt;/p&gt;

&lt;p&gt;The GPU did not disappear. It became the center of a much larger machine.&lt;/p&gt;

&lt;p&gt;The open question is whether this integration remains one durable platform advantage or creates enough cost, power, and operational complexity for the stack to split again—between general-purpose GPUs, specialized inference processors, open interconnects, and software that can schedule work across all of them.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>nvidia</category>
      <category>gpu</category>
      <category>semiconductors</category>
    </item>
    <item>
      <title>There Is No Universal AI Runtime—So Build a Portable Control Plane</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:52:17 +0000</pubDate>
      <link>https://dev.to/amitrix/there-is-no-universal-ai-runtime-so-build-a-portable-control-plane-2hhb</link>
      <guid>https://dev.to/amitrix/there-is-no-universal-ai-runtime-so-build-a-portable-control-plane-2hhb</guid>
      <description>&lt;p&gt;AI accelerators are not interchangeable computers. A CUDA kernel does not become a Trainium executable because both systems run PyTorch, and an inference server supporting two GPU families does not erase their different compilers, libraries, memory models, collectives, or performance envelopes.&lt;/p&gt;

&lt;p&gt;The portable layer sits higher.&lt;/p&gt;

&lt;p&gt;Applications can call one stable service contract. A model registry can govern several hardware-specific artifacts. A router can place each request on a qualified pool. Observability can compare the resulting service, and policy can decide when moving a workload is economically justified.&lt;/p&gt;

&lt;p&gt;That is a portable control plane—not a universal runtime.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Kernels, compiled artifacts, collectives, quantization paths, and performance tuning remain tied to an accelerator stack.&lt;/li&gt;
&lt;li&gt;Serving frameworks can present common APIs across several backends, but backend support does not guarantee identical model coverage, behavior, or economics.&lt;/li&gt;
&lt;li&gt;The durable architecture keeps NVIDIA, AMD, TPU, Trainium, and specialist accelerators in separate pools behind one service contract.&lt;/li&gt;
&lt;li&gt;Model governance, workload qualification, routing policy, service-level telemetry, and cost policy are more portable than execution.&lt;/li&gt;
&lt;li&gt;Portability is an operating model: maintain a broad compatibility pool and move stable workloads only after another pool passes correctness, performance, and cost gates.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The runtime boundary is real
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.nvidia.com/cuda/cuda-programming-guide/" rel="noopener noreferrer"&gt;CUDA&lt;/a&gt; defines a GPU programming model, runtime, compiler toolchain, and libraries for NVIDIA hardware. CUDA C++ device code is compiled by &lt;code&gt;nvcc&lt;/code&gt;, which can produce architecture-specific binary images and PTX for later compilation by the CUDA runtime. That mechanism provides compatibility within the NVIDIA platform; it does not produce an executable for an unrelated accelerator.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://rocm.docs.amd.com/projects/HIP/en/develop/how-to/hip_porting_guide.html" rel="noopener noreferrer"&gt;HIP&lt;/a&gt; deliberately resembles CUDA and includes tools that translate many CUDA API calls. The same documentation also states that HIP code must be compiled for a specific AMD GPU architecture and that the resulting binaries contain code for that target. Source translation reduces migration work. It does not turn a CUDA binary into an AMD binary.&lt;/p&gt;

&lt;p&gt;The boundary becomes clearer with cloud accelerators. The &lt;a href="https://awsdocs-neuron.readthedocs-hosted.com/en/latest/compiler/index.html" rel="noopener noreferrer"&gt;Neuron compiler&lt;/a&gt; transforms framework graphs into a Neuron Executable File Format artifact that the Neuron runtime loads on Neuron devices. &lt;a href="https://docs.pytorch.org/xla/master/" rel="noopener noreferrer"&gt;PyTorch/XLA&lt;/a&gt; lets applications retain familiar PyTorch APIs, but it traces operations into an intermediate graph and sends HLO to XLA for compilation. Changes in graph shape can cause additional compilations, and hardware-specific optimization still happens below the framework.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openxla.org/xla/architecture" rel="noopener noreferrer"&gt;OpenXLA&lt;/a&gt; is a useful portability layer precisely because it separates a versioned operation set from target-specific backends. StableHLO can carry a model graph between framework and compiler, while each backend still performs its own optimization and code generation. The graph can travel farther than the executable.&lt;/p&gt;

&lt;p&gt;This produces a simple hierarchy:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What can be shared&lt;/th&gt;
&lt;th&gt;What remains hardware-specific&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Application&lt;/td&gt;
&lt;td&gt;Request schema, authentication, streaming contract, error model&lt;/td&gt;
&lt;td&gt;Model behavior under different precision and implementations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model source&lt;/td&gt;
&lt;td&gt;Framework checkpoint, tokenizer, configuration, evaluation set&lt;/td&gt;
&lt;td&gt;Converted, sharded, quantized, or compiled artifacts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Graph representation&lt;/td&gt;
&lt;td&gt;ONNX or StableHLO where the model and operators are supported&lt;/td&gt;
&lt;td&gt;Unsupported operators, custom extensions, lowering, graph partitioning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Serving&lt;/td&gt;
&lt;td&gt;Endpoint shape, admission control, rollout mechanism&lt;/td&gt;
&lt;td&gt;Engine build, kernels, cache layout, batching, parallelism&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scheduling&lt;/td&gt;
&lt;td&gt;Workload intent, pool labels, quotas, priority&lt;/td&gt;
&lt;td&gt;Device discovery, drivers, topology, allocation mechanism&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Execution&lt;/td&gt;
&lt;td&gt;Nothing universal enough to assume interchangeability&lt;/td&gt;
&lt;td&gt;Runtime, compiler, kernels, libraries, collectives, memory placement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Measurement&lt;/td&gt;
&lt;td&gt;Correctness, latency, throughput, availability, cost per outcome&lt;/td&gt;
&lt;td&gt;Vendor counters, power boundaries, utilization definitions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The mistake is not using abstraction. The mistake is placing the abstraction below the level where the evidence supports it.&lt;/p&gt;

&lt;h2&gt;
  
  
  One contract, separate hardware pools
&lt;/h2&gt;

&lt;p&gt;The stable unit should be a service, not a device. Each accelerator family operates as a separately qualified pool with its own software image, driver, runtime, serving engine, model artifacts, deployment process, and failure domain.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    A[Applications and agents] --&amp;gt; B[Stable inference or training contract]
    B --&amp;gt; C[Identity, quotas, admission control]
    C --&amp;gt; D[Model catalog and artifact registry]
    D --&amp;gt; E[Workload qualification and policy router]

    E --&amp;gt; N[NVIDIA pool&amp;lt;br/&amp;gt;CUDA artifacts and kernels]
    E --&amp;gt; M[AMD pool&amp;lt;br/&amp;gt;ROCm and HIP artifacts]
    E --&amp;gt; T[TPU pool&amp;lt;br/&amp;gt;XLA-compiled artifacts]
    E --&amp;gt; R[Trainium pool&amp;lt;br/&amp;gt;Neuron-compiled artifacts]
    E --&amp;gt; S[Specialist pool&amp;lt;br/&amp;gt;Vendor runtime and artifacts]

    N --&amp;gt; O[Normalized service telemetry]
    M --&amp;gt; O
    T --&amp;gt; O
    R --&amp;gt; O
    S --&amp;gt; O

    O --&amp;gt; P[Cost, reliability, and placement policy]
    P --&amp;gt; E&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;This design does not schedule one container indiscriminately across every device. It selects a pool that already has a tested artifact and a known operating envelope for the requested model, precision, context length, batch profile, and service objective.&lt;/p&gt;

&lt;p&gt;The model catalog therefore needs more than a model name. A useful entry binds a logical model version to a set of qualified implementations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;logical model: reasoning-model-v3
evaluation policy: eval-suite-2026-08

qualified implementations:
  - pool: nvidia-hopper
    artifact: fp8-tensor-parallel-8
    status: production
  - pool: amd-mi300
    artifact: fp8-tensor-parallel-8-rocm
    status: canary
  - pool: trainium
    artifact: neuron-compiled-seq4096
    status: batch-only
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The logical model gives applications a stable name. Each implementation retains its own artifact identity, software bill of materials, evaluation results, supported limits, and rollback history. Promotion happens per implementation, not once for the abstract model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Serving frameworks widen the interface, not the binary
&lt;/h2&gt;

&lt;p&gt;Modern serving projects make this architecture easier, but none removes the qualification work.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.vllm.ai/en/latest/getting_started/installation/" rel="noopener noreferrer"&gt;vLLM&lt;/a&gt; supports hardware plugins outside its main repository. &lt;a href="https://github.com/sgl-project/sglang" rel="noopener noreferrer"&gt;SGLang&lt;/a&gt; documents support for multiple accelerator families and offers an API compatible with common language-model request patterns. These are valuable common surfaces. Their hardware paths can still differ in kernel availability, quantization support, distributed execution, model coverage, and release maturity.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/backend/README.html" rel="noopener noreferrer"&gt;Triton Inference Server&lt;/a&gt; makes the separation visible: every model is associated with a backend, and the backend is the implementation that executes it. Triton can wrap TensorRT, PyTorch, ONNX Runtime, vLLM, or custom logic, but its own documentation warns that not every backend is supported on every platform.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://onnxruntime.ai/docs/execution-providers/" rel="noopener noreferrer"&gt;ONNX Runtime&lt;/a&gt; follows a similar pattern through Execution Providers. The application uses a consistent API while a provider claims supported nodes or subgraphs and hands them to a hardware-specific library. That is portable application integration with specialized execution underneath. It is not proof that every model will partition, compile, or perform equally on every provider.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://kserve.github.io/website/docs/model-serving/generative-inference/overview" rel="noopener noreferrer"&gt;KServe&lt;/a&gt; can standardize how a model service is declared and operated on Kubernetes. Its serving runtimes still select concrete backends, images, resource requests, and model formats. The operational object becomes portable; the implementation remains qualified for a particular pool.&lt;/p&gt;

&lt;p&gt;The practical contract should cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;logical model and version;&lt;/li&gt;
&lt;li&gt;task type and request schema;&lt;/li&gt;
&lt;li&gt;streaming and cancellation behavior;&lt;/li&gt;
&lt;li&gt;context, input, and output limits;&lt;/li&gt;
&lt;li&gt;safety and data-handling policy;&lt;/li&gt;
&lt;li&gt;latency and availability objectives;&lt;/li&gt;
&lt;li&gt;error categories and retry rules;&lt;/li&gt;
&lt;li&gt;response metadata identifying the implementation used.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It should not promise identical tokenization, numerical output, latency, or cost unless those properties were explicitly tested.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scheduler places workloads; it does not port them
&lt;/h2&gt;

&lt;p&gt;Cluster schedulers already understand heterogeneous resources. &lt;a href="https://kubernetes.io/docs/concepts/extend-kubernetes/compute-storage-net/device-plugins/" rel="noopener noreferrer"&gt;Kubernetes device plugins&lt;/a&gt; let vendors advertise specialized devices to the kubelet. A cluster can expose resources such as &lt;code&gt;nvidia.com/gpu&lt;/code&gt; and &lt;code&gt;amd.com/gpu&lt;/code&gt;, then use labels and affinity to place a workload on the intended node type.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://slurm.schedmd.com/gres.html" rel="noopener noreferrer"&gt;Slurm GRES&lt;/a&gt; similarly models GPU types as generic resources and lets a job request a specific type and count. &lt;a href="https://docs.ray.io/en/latest/ray-core/scheduling/accelerators.html" rel="noopener noreferrer"&gt;Ray&lt;/a&gt; defines accelerator resources for NVIDIA GPUs, AMD GPUs, Neuron cores, TPUs, and other devices, then assigns tasks or actors to nodes with the requested resource.&lt;/p&gt;

&lt;p&gt;These systems solve allocation. They do not determine whether a model is correct on that accelerator, whether the serving engine has the required kernel, or whether the workload meets its latency target. A portable control plane must make placement eligibility a product of qualification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;eligible pool =
  supported model
  + approved artifact
  + compatible request limits
  + passing evaluation
  + healthy capacity
  + acceptable economics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only after those gates pass should the router compare eligible pools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route on workload shape and evidence
&lt;/h2&gt;

&lt;p&gt;A lowest-cost accelerator rule is too crude. Inference cost changes with context length, prefill-to-decode ratio, output length, concurrency, batch formation, cache reuse, quantization, and the amount of idle capacity held to protect latency.&lt;/p&gt;

&lt;p&gt;The routing decision needs a qualification matrix:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision input&lt;/th&gt;
&lt;th&gt;Why it changes placement&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model and revision&lt;/td&gt;
&lt;td&gt;Operator and architecture support change between releases&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Precision and quantization&lt;/td&gt;
&lt;td&gt;Available kernels and acceptable accuracy loss differ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context length&lt;/td&gt;
&lt;td&gt;Memory consumption, compilation shapes, and cache pressure change&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefill versus decode&lt;/td&gt;
&lt;td&gt;The two phases stress compute and memory differently&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency objective&lt;/td&gt;
&lt;td&gt;A cheap saturated pool can violate the service contract&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batchability&lt;/td&gt;
&lt;td&gt;Offline work can use capacity that interactive traffic cannot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Geography and data policy&lt;/td&gt;
&lt;td&gt;Some pools may be ineligible before performance is considered&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Current queue and warm capacity&lt;/td&gt;
&lt;td&gt;A nominally efficient pool can be expensive after delay and cold start&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Measured cost per accepted output&lt;/td&gt;
&lt;td&gt;Hardware-hour price alone omits utilization and failed work&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Routing also occurs at more than one level. &lt;a href="https://docs.ray.io/en/latest/serve/llm/architecture/routing-policies.html" rel="noopener noreferrer"&gt;Ray Serve’s routing documentation&lt;/a&gt; distinguishes model-level ingress routing from replica selection. A cross-hardware control plane needs both: first select the qualified implementation and pool, then let the local serving system select a replica using queue, cache, and locality signals.&lt;/p&gt;

&lt;p&gt;The broad compatibility pool plays an important role. New models, unusual operators, rapidly changing context requirements, and workloads without enough evaluation evidence stay there. Specialist pools earn traffic when the workload becomes stable enough to compile, measure, and govern.&lt;/p&gt;

&lt;h2&gt;
  
  
  Normalize outcomes, preserve native telemetry
&lt;/h2&gt;

&lt;p&gt;Observability is portable only at the service layer. Time to first token, inter-token latency, total response time, requests per second, tokens per second, error rate, evaluation score, queue delay, and cost per accepted output can be defined consistently.&lt;/p&gt;

&lt;p&gt;The lower-level evidence remains native. GPU occupancy, NeuronCore utilization, compiler-cache behavior, HBM counters, collective timing, and device power are not one universal metric family. Flattening them into a single “accelerator utilization” number destroys the details needed to diagnose a pool.&lt;/p&gt;

&lt;p&gt;The clean design keeps two views:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Normalized service telemetry&lt;/strong&gt; compares the outcome delivered to the application.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend-native telemetry&lt;/strong&gt; explains why one implementation produced that outcome.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Economic policy should use both. A pool that has a lower device-hour price but compiles frequently, rejects more requests, requires excess standby capacity, or produces lower-quality outputs can have the higher service cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compare service cost, not accelerator price
&lt;/h2&gt;

&lt;p&gt;A portable control plane needs one economic denominator across every pool:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;service cost per accepted unit =
  reserved accelerator and host capacity
  + network, storage, and energy
  + amortized model-porting and software work
  + operator and support cost
  + failed, retried, and quality-rejected work
  ------------------------------------------------
  accepted outputs that met the service objective
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The denominator can be accepted tokens, images, training steps, or completed jobs, but it must include the quality and latency objective. Utilization belongs inside the calculation because reserved capacity accrues cost while idle. Availability belongs inside it because failed work and standby capacity reduce the accepted output produced by the same reservation.&lt;/p&gt;

&lt;p&gt;This illustrative monthly comparison shows why hardware price alone is misleading:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Monthly component&lt;/th&gt;
&lt;th&gt;Broad compatibility pool&lt;/th&gt;
&lt;th&gt;New specialist pool&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reserved compute, hosts, and fabric&lt;/td&gt;
&lt;td&gt;$72,000&lt;/td&gt;
&lt;td&gt;$45,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Energy and facility allocation&lt;/td&gt;
&lt;td&gt;$8,000&lt;/td&gt;
&lt;td&gt;$4,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amortized porting and software work&lt;/td&gt;
&lt;td&gt;$6,000&lt;/td&gt;
&lt;td&gt;$15,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operations, support, and standby capacity&lt;/td&gt;
&lt;td&gt;$8,000&lt;/td&gt;
&lt;td&gt;$12,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failed, retried, or rejected work&lt;/td&gt;
&lt;td&gt;$4,000&lt;/td&gt;
&lt;td&gt;$6,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accepted output meeting the service objective&lt;/td&gt;
&lt;td&gt;140 billion tokens&lt;/td&gt;
&lt;td&gt;90 billion tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Service cost per million accepted tokens&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.70&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.91&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The numbers are hypothetical; the mechanism is not. The specialist has lower compute and energy cost but loses the comparison because it carries porting work, idle capacity, and failed output across a smaller qualified workload. If model coverage grows, utilization rises, and software work is amortized across more accepted output, the same pool can later become the cheaper path.&lt;/p&gt;

&lt;p&gt;The control plane should therefore store both the current service cost and the evidence behind it: measurement interval, reservation model, utilization, failure rate, quality threshold, energy boundary, labor allocation, and software version. Routing on a number without that provenance recreates the same false portability problem at the economic layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is still missing
&lt;/h2&gt;

&lt;p&gt;The software projects above provide pieces of this architecture, not one finished control plane. Schedulers allocate devices. Serving engines execute models. Model registries track artifacts. Gateways route requests. Telemetry systems collect measurements.&lt;/p&gt;

&lt;p&gt;The missing layer is a durable qualification record connecting all of them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;which logical model maps to which hardware-specific artifact;&lt;/li&gt;
&lt;li&gt;which evaluation approved that artifact;&lt;/li&gt;
&lt;li&gt;which request shapes it supports;&lt;/li&gt;
&lt;li&gt;which software and driver versions produced the result;&lt;/li&gt;
&lt;li&gt;which service and economic thresholds it currently meets;&lt;/li&gt;
&lt;li&gt;which policy moved traffic and how to reverse that decision.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without that record, “multi-accelerator” becomes a collection of independent deployments sharing an endpoint. With it, hardware diversity becomes governable.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what
&lt;/h2&gt;

&lt;p&gt;Do not wait for a universal AI runtime. The compilers, kernels, collectives, and memory systems are moving in the opposite direction—toward deeper specialization.&lt;/p&gt;

&lt;p&gt;Build portability where it holds: the application contract, model identity, evaluation policy, workload router, service telemetry, and economic decision. Keep execution pools separate, qualify every implementation, and route only across proven choices.&lt;/p&gt;

&lt;p&gt;The open thread is how much control-plane standardization is possible before the policy itself becomes coupled to the dominant serving engine. If every backend exposes different cache state, queue semantics, compilation behavior, and power evidence, the router may remain portable in name while its best decisions depend on vendor-specific signals.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aicompute</category>
      <category>infrastructure</category>
      <category>inference</category>
      <category>patterns</category>
    </item>
    <item>
      <title>The Interconnect Determines How Much of the Chip You Can Use</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:51:41 +0000</pubDate>
      <link>https://dev.to/amitrix/the-interconnect-determines-how-much-of-the-chip-you-can-use-43g3</link>
      <guid>https://dev.to/amitrix/the-interconnect-determines-how-much-of-the-chip-you-can-use-43g3</guid>
      <description>&lt;p&gt;An accelerator can finish its arithmetic and still leave the system waiting.&lt;/p&gt;

&lt;p&gt;That is the architectural shift behind rack-scale AI. More processors do not automatically produce a faster computer. They produce a larger distributed system whose useful performance depends on moving weights, activations, gradients, expert tokens, and cache state between the right processors at the right time.&lt;/p&gt;

&lt;p&gt;At rack and pod scale, data movement can constrain useful accelerator performance more than arithmetic. The chip still matters. But the interconnect increasingly decides how much of the chip the system can use.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Scale-up fabrics connect accelerators into one tightly coupled compute domain; scale-out networks connect servers, racks, and pods. They solve different communication problems.&lt;/li&gt;
&lt;li&gt;NVLink and NVSwitch are a mature vertically owned scale-up stack. UALink and Ethernet-based scale-up proposals create open alternatives, but a published specification is not a deployed ecosystem.&lt;/li&gt;
&lt;li&gt;InfiniBand and Ethernet/RoCE both support large AI clusters. Their useful performance depends on NIC placement, topology, congestion control, collectives, and operations—not the link rate alone.&lt;/li&gt;
&lt;li&gt;NCCL, RCCL, and related libraries turn collective operations into traffic patterns. They are part of the performance architecture, not a thin software wrapper.&lt;/li&gt;
&lt;li&gt;Copper remains efficient over short electrical reaches. Pluggable optics dominate longer links. Co-packaged optics and optical I/O move conversion closer to the switch or accelerator as electrical reach and front-panel power become harder constraints.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Scale-up and scale-out are different systems
&lt;/h2&gt;

&lt;p&gt;The cleanest way to understand an AI fabric is to ask where the tightly coupled compute domain ends.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scale-up&lt;/strong&gt; connects accelerators that exchange data frequently enough to behave like one larger machine. Latency, bandwidth, memory semantics, ordering, and collective behavior all matter. The fabric may remain inside a server, cross a rack, or eventually span several racks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scale-out&lt;/strong&gt; connects those compute domains into a larger cluster. It carries distributed training collectives, inference traffic, model and checkpoint movement, storage access, and control traffic across a broader topology.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    subgraph R1["Rack-scale compute domain"]
        A1["Accelerator"]
        A2["Accelerator"]
        A3["Accelerator"]
        SU1["Scale-up switches&amp;lt;br/&amp;gt;NVSwitch, UALink, or another fabric"]
        A1 &amp;lt;--&amp;gt; SU1
        A2 &amp;lt;--&amp;gt; SU1
        A3 &amp;lt;--&amp;gt; SU1
    end

    NIC1["NIC or SuperNIC"]
    SO["Scale-out network&amp;lt;br/&amp;gt;InfiniBand or Ethernet"]
    NIC2["NIC or SuperNIC"]

    subgraph R2["Second compute domain"]
        SU2["Scale-up fabric"]
        B1["Accelerator"]
        B2["Accelerator"]
        SU2 &amp;lt;--&amp;gt; B1
        SU2 &amp;lt;--&amp;gt; B2
    end

    SU1 &amp;lt;--&amp;gt; NIC1
    NIC1 &amp;lt;--&amp;gt; SO
    SO &amp;lt;--&amp;gt; NIC2
    NIC2 &amp;lt;--&amp;gt; SU2&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The distinction is architectural, not merely physical. NVIDIA's &lt;a href="https://www.nvidia.com/en-us/data-center/nvlink/" rel="noopener noreferrer"&gt;NVLink&lt;/a&gt; carries tightly coupled GPU traffic, while its &lt;a href="https://www.nvidia.com/en-us/networking/" rel="noopener noreferrer"&gt;Quantum InfiniBand and Spectrum-X Ethernet platforms&lt;/a&gt; connect systems across the scale-out network. A cable does not become scale-up because it is fast, and Ethernet does not become scale-out-only because that is its conventional role.&lt;/p&gt;

&lt;p&gt;That last boundary is now moving. The &lt;a href="https://ualinkconsortium.org/specification/" rel="noopener noreferrer"&gt;UALink Consortium&lt;/a&gt; defines an open 200 Gb/s-per-lane scale-up specification for accelerator pods. Broadcom's &lt;a href="https://docs.broadcom.com/doc/scale-up-ethernet-framework" rel="noopener noreferrer"&gt;Scale-Up Ethernet framework&lt;/a&gt; applies Ethernet technology to the scale-up domain. AMD's &lt;a href="https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html" rel="noopener noreferrer"&gt;UALink over Ethernet, or UALoE&lt;/a&gt;, uses that broader direction to connect 72 MI455X GPUs in the Helios rack. The physical and link technologies can converge while the workload contract remains different.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scale-up contest is also an ownership contest
&lt;/h2&gt;

&lt;p&gt;NVLink's advantage is not only its published bandwidth. NVIDIA owns the GPU interface, NVSwitch silicon, firmware, topology, collective library, rack design, and much of the qualification process.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/nvlink/" rel="noopener noreferrer"&gt;Blackwell NVLink 5&lt;/a&gt; provides 1.8 TB/s of aggregate bidirectional bandwidth per GPU. &lt;a href="https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/" rel="noopener noreferrer"&gt;Rubin NVLink 6&lt;/a&gt; doubles that published endpoint figure to 3.6 TB/s. In NVL72 systems, NVSwitch extends the tightly coupled domain across 72 GPUs in one rack.&lt;/p&gt;

&lt;p&gt;Those figures are endpoint totals, not achieved application throughput or one-way fabric bisection bandwidth. Blackwell systems are deployed, while NVIDIA describes Vera Rubin as ramping into full production. Both belong to an integrated product stack—not an interface specification waiting for implementations.&lt;/p&gt;

&lt;p&gt;UALink takes a different path. Its &lt;a href="https://ualinkconsortium.org/specification/" rel="noopener noreferrer"&gt;200G 1.0 specification&lt;/a&gt; separates the fabric definition from any single accelerator vendor and supports scale-up domains as large as 1,024 accelerators. That creates room for accelerators, switch silicon, cables, management components, and systems from different suppliers.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://ualinkconsortium.org/specification/" rel="noopener noreferrer"&gt;UALink specification family&lt;/a&gt; now also separates common, data-link, physical-layer, chiplet, and manageability documents. The open boundary moves integration responsibility outward. Compatibility requires more than matching electrical lanes: memory semantics, error handling, management, firmware, topology discovery, collectives, and system qualification must converge across companies.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Fabric or initiative&lt;/th&gt;
&lt;th&gt;Primary boundary&lt;/th&gt;
&lt;th&gt;Ownership model&lt;/th&gt;
&lt;th&gt;Status by Aug. 30, 2026&lt;/th&gt;
&lt;th&gt;What the label does not prove&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/nvlink/" rel="noopener noreferrer"&gt;NVLink + NVSwitch&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;NVIDIA GPU scale-up&lt;/td&gt;
&lt;td&gt;Vertically controlled by NVIDIA&lt;/td&gt;
&lt;td&gt;Mature across several generations; NVL72 in production&lt;/td&gt;
&lt;td&gt;Application bandwidth, availability, or compatibility with non-NVIDIA accelerators&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://ualinkconsortium.org/specification/" rel="noopener noreferrer"&gt;UALink specifications&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Open accelerator scale-up&lt;/td&gt;
&lt;td&gt;Consortium specifications and multi-vendor implementations&lt;/td&gt;
&lt;td&gt;200G 1.0 plus newer component specifications published; implementation ecosystem emerging&lt;/td&gt;
&lt;td&gt;Shipping interoperable systems at NVLink maturity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://docs.broadcom.com/doc/scale-up-ethernet-framework" rel="noopener noreferrer"&gt;Broadcom Scale-Up Ethernet&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Ethernet-based accelerator scale-up&lt;/td&gt;
&lt;td&gt;Merchant switch silicon plus ecosystem&lt;/td&gt;
&lt;td&gt;Framework and products published; ecosystem integration under way&lt;/td&gt;
&lt;td&gt;That ordinary data-center Ethernet automatically meets scale-up semantics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html" rel="noopener noreferrer"&gt;AMD UALoE&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;72-GPU Helios scale-up domain&lt;/td&gt;
&lt;td&gt;AMD implementation using Ethernet switch silicon&lt;/td&gt;
&lt;td&gt;Helios architecture introduced in 2026&lt;/td&gt;
&lt;td&gt;Independent interoperability or NVLink-equivalent operational maturity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://ultraethernet.org/specification-history/" rel="noopener noreferrer"&gt;Ultra Ethernet 1.0.3&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;AI/HPC scale-out transport&lt;/td&gt;
&lt;td&gt;Consortium specification&lt;/td&gt;
&lt;td&gt;Current specification published; products and software arriving by vendor&lt;/td&gt;
&lt;td&gt;A replacement for scale-up fabrics or an installed multi-vendor fleet&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These rows are not benchmark peers. NVLink is a deployed proprietary stack. UALink is an open scale-up specification. Ultra Ethernet is a scale-out transport specification. Scale-Up Ethernet is a family of merchant-Ethernet approaches aimed at a new boundary.&lt;/p&gt;

&lt;p&gt;The strategic question is whether an open ecosystem can coordinate those layers quickly enough to behave like one system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Link speed is not collective performance
&lt;/h2&gt;

&lt;p&gt;Distributed models do not send a uniform stream of independent packets. They execute collective operations.&lt;/p&gt;

&lt;p&gt;An &lt;strong&gt;all-reduce&lt;/strong&gt; combines values across workers and returns the result to each worker. An &lt;strong&gt;all-gather&lt;/strong&gt; assembles shards. A &lt;strong&gt;reduce-scatter&lt;/strong&gt; combines and redistributes them. Mixture-of-experts models can produce all-to-all traffic as tokens move toward different experts.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/overview.html" rel="noopener noreferrer"&gt;NCCL&lt;/a&gt; implements these collectives for NVIDIA GPUs across PCIe, NVLink, InfiniBand, and Ethernet paths. &lt;a href="https://rocm.docs.amd.com/projects/rccl/en/latest/" rel="noopener noreferrer"&gt;RCCL&lt;/a&gt; provides the corresponding collective library in AMD's ROCm stack. Intel maintains &lt;a href="https://github.com/uxlfoundation/oneCCL" rel="noopener noreferrer"&gt;oneCCL&lt;/a&gt; for distributed communication across its supported compute environments.&lt;/p&gt;

&lt;p&gt;The library chooses algorithms and routes based on topology, message size, available transports, and system configuration. A ring can use links efficiently for large transfers. A tree can reduce latency for other message sizes. Hierarchical algorithms can reduce data inside a server or rack before crossing the scale-out network.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    OP["Model operation&amp;lt;br/&amp;gt;all-reduce, all-gather, all-to-all"]
    LIB["Collective library&amp;lt;br/&amp;gt;NCCL, RCCL, or oneCCL"]
    PLAN["Algorithm and topology choice&amp;lt;br/&amp;gt;ring, tree, hierarchy, channels"]
    LOCAL["Local movement&amp;lt;br/&amp;gt;HBM, PCIe, scale-up fabric"]
    NIC["NIC path&amp;lt;br/&amp;gt;RDMA, GPU memory access, transport offloads"]
    FABRIC["Scale-out fabric&amp;lt;br/&amp;gt;switches, routes, congestion control"]
    RESULT["Useful synchronized work"]

    OP --&amp;gt; LIB --&amp;gt; PLAN
    PLAN --&amp;gt; LOCAL
    PLAN --&amp;gt; NIC --&amp;gt; FABRIC
    LOCAL --&amp;gt; RESULT
    FABRIC --&amp;gt; RESULT&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;This stack explains why a nominally faster network can deliver less useful performance. A job can lose time through a poor rank-to-NIC mapping, oversubscribed links, uneven paths, head-of-line blocking, retransmission, collective imbalance, or one slow worker.&lt;/p&gt;

&lt;p&gt;The accelerator metric that matters is not peak FLOPS. It is the fraction of execution time spent doing useful compute rather than waiting for communication or synchronization.&lt;/p&gt;

&lt;h2&gt;
  
  
  InfiniBand and Ethernet are systems, not cable choices
&lt;/h2&gt;

&lt;p&gt;InfiniBand entered AI clusters with an integrated remote-memory and lossless-fabric model. NVIDIA's &lt;a href="https://www.nvidia.com/en-us/networking/products/infiniband/" rel="noopener noreferrer"&gt;Quantum-2 platform&lt;/a&gt; combines switches, ConnectX adapters, RDMA, congestion management, telemetry, and its SHARP in-network-reduction capability. In supported Quantum deployments, the fabric can participate in selected collective operations instead of treating every reduction as endpoint-only traffic.&lt;/p&gt;

&lt;p&gt;Ethernet starts with a broader multi-vendor base. Running AI collectives over Ethernet commonly uses &lt;a href="https://docs.nvidia.com/doca/sdk/rdma+over+converged+ethernet/index.html" rel="noopener noreferrer"&gt;RoCE&lt;/a&gt;, which carries RDMA over Ethernet. That removes repeated CPU and kernel involvement from the data path, but it does not remove the need to engineer congestion behavior.&lt;/p&gt;

&lt;p&gt;Traditional RoCE deployments can combine priority flow control, explicit congestion notification, data-center quantized congestion notification, adaptive routing, and careful buffer configuration. NVIDIA's &lt;a href="https://www.nvidia.com/en-us/networking/products/ethernet/" rel="noopener noreferrer"&gt;Spectrum-X&lt;/a&gt; packages Ethernet switches, BlueField SuperNICs, telemetry, congestion control, and adaptive routing as one AI-networking platform.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://ultraethernet.org/specification-history/" rel="noopener noreferrer"&gt;Ultra Ethernet Consortium's current 1.0.3 specification&lt;/a&gt; goes further by defining an AI/HPC transport that includes packet spraying across paths, selective retransmission, congestion management, and collective-communication support. It is an attempt to make large Ethernet fabrics less dependent on every deployment recreating the same behavior through proprietary combinations.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;InfiniBand example&lt;/th&gt;
&lt;th&gt;Ethernet example&lt;/th&gt;
&lt;th&gt;Operational owner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Link and switching&lt;/td&gt;
&lt;td&gt;Quantum switches and InfiniBand links&lt;/td&gt;
&lt;td&gt;Ethernet switches and links&lt;/td&gt;
&lt;td&gt;Network platform team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint&lt;/td&gt;
&lt;td&gt;ConnectX HCA&lt;/td&gt;
&lt;td&gt;NIC or SuperNIC&lt;/td&gt;
&lt;td&gt;Server and network teams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Remote-memory transport&lt;/td&gt;
&lt;td&gt;InfiniBand RDMA&lt;/td&gt;
&lt;td&gt;RoCE or Ultra Ethernet transport&lt;/td&gt;
&lt;td&gt;NIC, driver, and fabric stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Congestion behavior&lt;/td&gt;
&lt;td&gt;InfiniBand controls and adaptive routing&lt;/td&gt;
&lt;td&gt;PFC/ECN/DCQCN, vendor controls, or UEC mechanisms&lt;/td&gt;
&lt;td&gt;Fabric engineering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Collective layer&lt;/td&gt;
&lt;td&gt;NCCL/RCCL plus optional in-network operations&lt;/td&gt;
&lt;td&gt;NCCL/RCCL/oneCCL plus transport integration&lt;/td&gt;
&lt;td&gt;Accelerator platform team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Job placement&lt;/td&gt;
&lt;td&gt;Topology-aware scheduler and rank placement&lt;/td&gt;
&lt;td&gt;Topology-aware scheduler and rank placement&lt;/td&gt;
&lt;td&gt;Cluster platform team&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is why “InfiniBand versus Ethernet” is an incomplete comparison. The real unit is endpoint hardware, switching, transport, congestion control, collectives, topology, and operations.&lt;/p&gt;

&lt;h2&gt;
  
  
  NICs and DPUs own different work
&lt;/h2&gt;

&lt;p&gt;A NIC terminates the network path. A modern AI NIC can also expose RDMA, access accelerator memory, steer traffic across rails, timestamp packets, and offload transport work.&lt;/p&gt;

&lt;p&gt;NVIDIA uses the term &lt;strong&gt;SuperNIC&lt;/strong&gt; for a high-performance network endpoint optimized for accelerator traffic. It positions &lt;a href="https://www.nvidia.com/en-us/networking/products/data-processing-unit/" rel="noopener noreferrer"&gt;BlueField-3 SuperNIC&lt;/a&gt; around GPU-to-GPU networking, while its DPU configuration adds programmable Arm cores and infrastructure offloads.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;DPU&lt;/strong&gt; has a broader ownership boundary. It can run networking, storage, security, isolation, and management functions away from the host CPU. AMD's &lt;a href="https://www.amd.com/en/products/data-processing-units/pensando.html" rel="noopener noreferrer"&gt;Pensando infrastructure products&lt;/a&gt; follow that broader offload model.&lt;/p&gt;

&lt;p&gt;Neither device fixes a poor fabric by itself. A faster NIC attached through the wrong PCIe root, mapped to distant GPUs, or connected to an oversubscribed leaf still creates a slow communication path. Hardware topology and software placement have to agree.&lt;/p&gt;

&lt;h2&gt;
  
  
  Copper reaches a wall before optics does
&lt;/h2&gt;

&lt;p&gt;The physical medium is becoming an architectural choice.&lt;/p&gt;

&lt;p&gt;Passive copper is attractive inside a tray or across short rack distances because it avoids optical conversion and can deliver low power and low latency. Active electrical cables extend that reach by adding signal conditioning. As lane rates rise, electrical loss increases and the practical reach shrinks.&lt;/p&gt;

&lt;p&gt;Pluggable optical modules move the electrical-to-optical conversion to the switch faceplate. They provide longer reach and replaceable modules, but every port consumes front-panel area and conversion power. Linear pluggable optics remove part of the digital signal processing from the module to reduce power, while placing tighter signal-integrity requirements on the host system.&lt;/p&gt;

&lt;p&gt;Co-packaged optics moves the optical engines beside the switch ASIC. Broadcom's &lt;a href="https://www.broadcom.com/products/fiber-optic-modules-components/co-packaged-optics" rel="noopener noreferrer"&gt;co-packaged-optics portfolio&lt;/a&gt; and NVIDIA's &lt;a href="https://www.nvidia.com/en-us/networking/products/silicon-photonics/" rel="noopener noreferrer"&gt;silicon-photonics switch platforms&lt;/a&gt; shorten the high-speed electrical path and increase optical bandwidth density. They also couple optics more closely to switch packaging, cooling, yield, and service procedures.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    NEED["Required reach, bandwidth density,&amp;lt;br/&amp;gt;power, and service model"]
    TRACE["On-board traces&amp;lt;br/&amp;gt;inside board or package"]
    COPPER["Passive or active copper&amp;lt;br/&amp;gt;tray and short rack links"]
    PLUG["Pluggable optics&amp;lt;br/&amp;gt;conversion at faceplate"]
    CPO["Co-packaged optics&amp;lt;br/&amp;gt;conversion beside switch ASIC"]
    OIO["Optical I/O&amp;lt;br/&amp;gt;conversion beside compute or memory"]

    NEED --&amp;gt; TRACE
    NEED --&amp;gt; COPPER
    NEED --&amp;gt; PLUG
    NEED --&amp;gt; CPO
    NEED --&amp;gt; OIO&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Optical I/O pushes the boundary closer to compute. &lt;a href="https://ayarlabs.com/products/" rel="noopener noreferrer"&gt;Ayar Labs' TeraPHY&lt;/a&gt; packages optical chiplets for chip-to-chip communication. &lt;a href="https://lightmatter.co/products/passage/" rel="noopener noreferrer"&gt;Lightmatter Passage&lt;/a&gt; combines optical connectivity with packaging intended to link processors and memory. &lt;a href="https://www.celestial.ai/technology" rel="noopener noreferrer"&gt;Celestial AI's Photonic Fabric&lt;/a&gt; targets optical scale-up and memory connectivity.&lt;/p&gt;

&lt;p&gt;These products do not share one maturity level or deployment model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Optical boundary&lt;/th&gt;
&lt;th&gt;Public maturity by Aug. 30, 2026&lt;/th&gt;
&lt;th&gt;Evidence boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.nvidia.com/en-us/networking/interconnect/" rel="noopener noreferrer"&gt;Pluggable optics&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Shipping infrastructure across data-center networks&lt;/td&gt;
&lt;td&gt;Mature module ecosystem; power and faceplate density remain constraints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://www.broadcom.com/products/fiber-optic-modules-components/co-packaged-optics" rel="noopener noreferrer"&gt;Broadcom CPO&lt;/a&gt; and &lt;a href="https://www.nvidia.com/en-us/networking/products/silicon-photonics/" rel="noopener noreferrer"&gt;NVIDIA silicon photonics&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Commercial portfolios and announced switch platforms&lt;/td&gt;
&lt;td&gt;Product availability varies by platform and configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://ayarlabs.com/products/" rel="noopener noreferrer"&gt;Ayar Labs optical I/O&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Chiplet products and customer sampling programs&lt;/td&gt;
&lt;td&gt;Not yet a broadly interchangeable accelerator interface&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://lightmatter.co/products/passage/" rel="noopener noreferrer"&gt;Lightmatter Passage&lt;/a&gt; and &lt;a href="https://www.celestial.ai/technology" rel="noopener noreferrer"&gt;Celestial AI Photonic Fabric&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Product programs and partner integrations&lt;/td&gt;
&lt;td&gt;Deployment scale and production economics remain vendor-reported&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The direction is consistent even though the maturity is not: when electrical reach, connector density, and SerDes power become system limits, designers can move the optical conversion point inward. That choice is driven by distance and packaging boundary, not one universal copper-to-optics migration sequence.&lt;/p&gt;

&lt;h2&gt;
  
  
  The interconnect changes who owns the computer
&lt;/h2&gt;

&lt;p&gt;NVIDIA's model places the scale-up link, switch, NIC, network, collective library, and rack qualification under one platform owner. That reduces the number of boundaries a system integrator must reconcile.&lt;/p&gt;

&lt;p&gt;An open stack distributes those responsibilities. An accelerator vendor can implement UALink. A merchant-silicon supplier can provide scale-up and scale-out switches. A NIC vendor can own RDMA and congestion functions. An OEM can assemble the rack. A collective library can adapt the topology to the framework.&lt;/p&gt;

&lt;p&gt;That structure creates supplier choice. It also creates a harder acceptance test. A specification match does not establish that the assembled system sustains collectives under congestion, survives link failures, maps jobs correctly, or exposes one support boundary when performance falls.&lt;/p&gt;

&lt;p&gt;The practical evaluation unit is therefore not the accelerator card. It is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;the tightly coupled scale-up domain;&lt;/li&gt;
&lt;li&gt;the scale-out fabric joining those domains;&lt;/li&gt;
&lt;li&gt;the collective software translating model operations into traffic;&lt;/li&gt;
&lt;li&gt;the topology and congestion mechanisms preserving progress;&lt;/li&gt;
&lt;li&gt;the physical medium carrying the required bandwidth within the power and service envelope.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What remains unproven
&lt;/h2&gt;

&lt;p&gt;Vendor bandwidth figures are abundant. Comparable useful-work evidence is not.&lt;/p&gt;

&lt;p&gt;A credible system comparison needs the same model, precision, collective mix, accelerator count, topology, message sizes, failure policy, and power boundary. It also needs achieved bandwidth, communication-to-compute overlap, tail behavior under congestion, job completion time, and recovery behavior—not only a clean microbenchmark.&lt;/p&gt;

&lt;p&gt;Open fabrics face a second proof burden: interoperability. UALink and Ultra Ethernet can publish complete specifications before multiple vendors ship independently qualified systems. That is normal standards work, but it creates a gap between protocol maturity and operational maturity.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what
&lt;/h2&gt;

&lt;p&gt;The next AI platform decision cannot stop at chip throughput. The fabric determines the size of the useful computer, the software determines how traffic enters it, and optics determines how far that computer can extend before power and signal integrity pull it apart.&lt;/p&gt;

&lt;p&gt;The open question is whether the industry can create a genuinely interchangeable scale-up and scale-out ecosystem—or whether every high-performing AI cluster will remain a platform-specific computer whose compatibility ends exactly where its collective software, congestion controls, and optical packaging diverge.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>networking</category>
      <category>interconnect</category>
      <category>optics</category>
    </item>
    <item>
      <title>Groq Makes the Compiler Part of the Processor</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:51:06 +0000</pubDate>
      <link>https://dev.to/amitrix/groq-makes-the-compiler-part-of-the-processor-1ilj</link>
      <guid>https://dev.to/amitrix/groq-makes-the-compiler-part-of-the-processor-1ilj</guid>
      <description>&lt;p&gt;Groq's architectural bet is that an inference processor should execute a plan, not discover one while the model is running.&lt;/p&gt;

&lt;p&gt;The Language Processing Unit removes much of the reactive machinery found in conventional processors: hardware caches, dynamic arbitration, speculative execution, and out-of-order scheduling. Its compiler decides where operations run, where data lives, and when each value moves. The processor then follows that schedule across a large on-chip SRAM system.&lt;/p&gt;

&lt;p&gt;That changes the unit worth evaluating. One LPU does not hold a modern large language model. The useful machine is the connected system of LPUs, its compiler-generated schedule, and the cloud service around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Groq moves instruction scheduling, tensor placement, and much data movement into the compiler, leaving simpler hardware to execute a predetermined program.&lt;/li&gt;
&lt;li&gt;SRAM supplies high bandwidth and predictable access time, but limited capacity forces model weights and execution across many connected LPUs.&lt;/li&gt;
&lt;li&gt;Deterministic chip execution reduces one source of latency variation; it does not remove API networking, admission control, queueing, or shared-service contention.&lt;/li&gt;
&lt;li&gt;Groq is strongest where one user's token rate matters: voice, coding loops, agents, and other interactive inference. That is different from maximizing total batched throughput.&lt;/li&gt;
&lt;li&gt;NVIDIA did not acquire Groq. It licensed Groq's inference technology non-exclusively and hired Groq leaders and staff; Groq remains independent and continues to operate GroqCloud.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The processor begins in the compiler
&lt;/h2&gt;

&lt;p&gt;Modern GPUs are also compiler-driven machines. Framework graphs are lowered into kernels, kernels are compiled, and serving engines plan memory and batch work. The distinction is what remains unresolved at runtime.&lt;/p&gt;

&lt;p&gt;A GPU retains hardware mechanisms that react to events while a program executes: caches answer memory requests, arbiters resolve contention, and schedulers choose from ready work. Groq's original &lt;a href="https://dl.acm.org/doi/10.1109/ISCA45697.2020.00023" rel="noopener noreferrer"&gt;Tensor Streaming Processor paper&lt;/a&gt; describes the opposite cut. It eliminates reactive elements such as caches and arbiters so the compiler can reason about execution timing precisely.&lt;/p&gt;

&lt;p&gt;The hardware is functionally sliced. Separate regions perform matrix operations, vector arithmetic, memory access, data movement, and instruction control. Values stream between those regions through a software-controlled interconnect. The compiler maps the model graph onto that physical machine and emits a cycle-level schedule.&lt;/p&gt;

&lt;p&gt;Groq describes this as a &lt;a href="https://groq.com/blog/the-groq-lpu-explained" rel="noopener noreferrer"&gt;programmable assembly line&lt;/a&gt;: each functional unit receives an instruction, takes data from a named stream, performs its operation, and places the result onto another stream. Hardware does not wait to discover which path wins arbitration because the compiler has already assigned the path and time.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System decision&lt;/th&gt;
&lt;th&gt;Conventional GPU-oriented execution&lt;/th&gt;
&lt;th&gt;Groq LPU execution&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Instruction readiness&lt;/td&gt;
&lt;td&gt;Hardware and runtime mechanisms select ready work&lt;/td&gt;
&lt;td&gt;Compiler places work into a predetermined schedule&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data locality&lt;/td&gt;
&lt;td&gt;Registers, caches, HBM, and runtime-managed movement&lt;/td&gt;
&lt;td&gt;Compiler assigns tensors across distributed SRAM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource contention&lt;/td&gt;
&lt;td&gt;Hardware arbitration resolves competing requests&lt;/td&gt;
&lt;td&gt;Compiler attempts to prevent conflicts before execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timing model&lt;/td&gt;
&lt;td&gt;Performance can vary with cache behavior and contention&lt;/td&gt;
&lt;td&gt;Scheduled accelerator execution is predictable cycle by cycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Optimization boundary&lt;/td&gt;
&lt;td&gt;Kernels, runtime, serving engine, and hardware&lt;/td&gt;
&lt;td&gt;Model compiler, LPU fabric, and serving system&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is not “software instead of hardware.” The compiler works because the hardware was built to expose timing and data movement directly. Groq makes the compiler part of the processor's effective control plane.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    GRAPH["Model graph"]
    COMPILER["Groq compiler&amp;lt;br/&amp;gt;partition + place + schedule"]
    ARTIFACT["Compiled execution plan"]

    subgraph SYSTEM["The useful compute boundary"]
        LPU1["LPU&amp;lt;br/&amp;gt;compute + SRAM"]
        LPU2["LPU&amp;lt;br/&amp;gt;compute + SRAM"]
        LPUN["More LPUs&amp;lt;br/&amp;gt;compute + SRAM"]
        FABRIC["RealScale interconnect&amp;lt;br/&amp;gt;scheduled chip-to-chip movement"]

        LPU1 &amp;lt;--&amp;gt; FABRIC
        LPU2 &amp;lt;--&amp;gt; FABRIC
        LPUN &amp;lt;--&amp;gt; FABRIC
    end

    API["GroqCloud or system API"]
    REQUESTS["Inference requests"]

    GRAPH --&amp;gt; COMPILER --&amp;gt; ARTIFACT --&amp;gt; SYSTEM
    REQUESTS --&amp;gt; API --&amp;gt; SYSTEM
    SYSTEM --&amp;gt; API&lt;/code&gt;&lt;/pre&gt;



&lt;h2&gt;
  
  
  SRAM makes speed possible—and scale necessary
&lt;/h2&gt;

&lt;p&gt;Groq places the primary working memory on the processor in SRAM. SRAM is faster and more predictable than fetching weights from off-chip DRAM or HBM, but it occupies much more silicon area per stored bit.&lt;/p&gt;

&lt;p&gt;The first-generation chip described in the architecture literature contains approximately &lt;strong&gt;220 MB of SRAM&lt;/strong&gt;. Groq's current &lt;a href="https://groq.com/groqnode-server/" rel="noopener noreferrer"&gt;eight-accelerator GroqNode&lt;/a&gt; lists &lt;strong&gt;1.76 GB of aggregate on-die memory&lt;/strong&gt;, consistent with eight chips at that capacity. A 70-billion-parameter model requires roughly 35 GB merely to store weights at four bits per parameter, before accounting for scales, activations, code, and key-value cache.&lt;/p&gt;

&lt;p&gt;One chip is therefore the wrong capacity boundary. Groq partitions model weights and operations across multiple LPUs, then streams activations through the connected machine. GroqNode exposes 88 chip-to-chip connectors and is positioned for multi-server and multi-rack scaling without a separate external switch in the LPU fabric.&lt;/p&gt;

&lt;p&gt;The result resembles a pipeline more than a pool of interchangeable accelerators. Each stage owns part of the model and passes intermediate values to the next stage according to the compiled schedule. Adding LPUs expands both SRAM capacity and execution resources, but it also makes partitioning, communication, and pipeline utilization part of the performance result.&lt;/p&gt;

&lt;p&gt;Groq's latest rack specification makes that system boundary explicit. Its &lt;a href="https://groq.com/inference/" rel="noopener noreferrer"&gt;inference platform page&lt;/a&gt; lists 256 LPUs, 128 GB of on-chip SRAM, and 40 PB/s of aggregate SRAM bandwidth per rack for the Groq 3 generation. Those are vendor specifications, not independently measured application results.&lt;/p&gt;

&lt;p&gt;The architectural trade is clear: spend silicon and system complexity on low-latency SRAM and predictable movement, then recover model capacity by joining many processors into one scheduled machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deterministic execution is not deterministic service
&lt;/h2&gt;

&lt;p&gt;Groq uses “deterministic” precisely at the accelerator layer. Given the same compiled program, the LPU schedule specifies when instructions execute and when values move. The hardware does not introduce cache misses or arbitration delays that change the cycle count.&lt;/p&gt;

&lt;p&gt;A cloud request crosses a larger system:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;client network
  → API gateway
  → admission control and queue
  → model replica selection
  → prompt processing
  → scheduled LPU execution
  → streamed response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only one part of that chain is the deterministic machine. Request bursts can still create queues. Prompts have different lengths. Tool calls introduce external dependencies. Network paths vary. A shared service can also reserve, batch, or route work differently as load changes.&lt;/p&gt;

&lt;p&gt;This distinction matters because latency has several meanings:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;What it measures&lt;/th&gt;
&lt;th&gt;What it can hide&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Time to first token&lt;/td&gt;
&lt;td&gt;Delay before generation begins&lt;/td&gt;
&lt;td&gt;Prompt length, queueing, and network transit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output speed&lt;/td&gt;
&lt;td&gt;Tokens delivered per second after generation starts&lt;/td&gt;
&lt;td&gt;Queueing before the first token and aggregate capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;End-to-end response time&lt;/td&gt;
&lt;td&gt;User wait for the complete answer&lt;/td&gt;
&lt;td&gt;Different answer lengths and reasoning-token behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aggregate throughput&lt;/td&gt;
&lt;td&gt;Total work completed by the serving system&lt;/td&gt;
&lt;td&gt;Whether any individual user receives tokens quickly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tail latency&lt;/td&gt;
&lt;td&gt;Slow requests at the 95th or 99th percentile&lt;/td&gt;
&lt;td&gt;Average results that look healthy while some users wait&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Groq's strongest public evidence is output speed. Its &lt;a href="https://console.groq.com/docs/models" rel="noopener noreferrer"&gt;GroqCloud model catalog&lt;/a&gt; publishes model-specific rates ranging into hundreds of tokens per second, with smaller models reaching about 1,000 tokens per second. These are service specifications and vary by model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://artificialanalysis.ai/providers/groq" rel="noopener noreferrer"&gt;Artificial Analysis independently measures&lt;/a&gt; Groq's public endpoints. At the August 30, 2026 cutoff, it reports the fastest tracked Groq configuration near 980 output tokens per second. For the same Llama 3.3 70B model, its &lt;a href="https://artificialanalysis.ai/models/llama-3-3-instruct-70b/providers" rel="noopener noreferrer"&gt;provider comparison&lt;/a&gt; measures Groq at roughly 298 output tokens per second—more than 20 times one of the slower tracked providers.&lt;/p&gt;

&lt;p&gt;That is useful independent evidence for interactive speed. It is not a rack-throughput benchmark, a power measurement, or a total-cost comparison. Artificial Analysis measures hosted APIs; it does not expose how many chips, replicas, queued users, or reserved capacity sit behind each endpoint.&lt;/p&gt;

&lt;p&gt;The distinction also explains Groq's workload fit. Fast sequential generation is valuable for conversational voice, coding agents waiting inside a loop, interactive search, and workflows that call a model repeatedly. Offline jobs with large batches care more about total tokens per dollar and total tokens per rack than the token rate of one stream.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://newsletter.semianalysis.com/p/nvidia-the-inference-kingdom-expands" rel="noopener noreferrer"&gt;SemiAnalysis argues&lt;/a&gt; that SRAM-heavy systems trade capacity available for weights and key-value cache against high single-user token rates, while GPUs can win aggregate throughput and cost through batching. That is the right comparison boundary: latency-first and throughput-first inference are different products even when both report “tokens per second.”&lt;/p&gt;

&lt;h2&gt;
  
  
  GroqCloud is the complete product
&lt;/h2&gt;

&lt;p&gt;The LPU does not arrive as a general accelerator card with a CUDA-like programming surface. GroqCloud presents supported models through an API that is largely compatible with the OpenAI client pattern. The compiler, model artifacts, partitioning, and machine topology remain behind the service boundary.&lt;/p&gt;

&lt;p&gt;That lowers adoption work for a supported model: change the endpoint and test behavior. It also moves control to the provider. Model availability, context limits, rate limits, pricing, quantization, and rollout schedules follow the GroqCloud catalog rather than the customer's own accelerator fleet.&lt;/p&gt;

&lt;p&gt;The fit is strongest when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the model and operators are supported by Groq's compiler and service;&lt;/li&gt;
&lt;li&gt;low per-user generation latency changes the product experience;&lt;/li&gt;
&lt;li&gt;the workload is stable enough to justify a compiled deployment;&lt;/li&gt;
&lt;li&gt;API consumption is preferable to owning the hardware and compiler lifecycle.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fit weakens when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the workload requires custom or rapidly changing operators;&lt;/li&gt;
&lt;li&gt;the model is absent from the supported catalog;&lt;/li&gt;
&lt;li&gt;large dynamic key-value caches dominate memory demand;&lt;/li&gt;
&lt;li&gt;maximum batched throughput matters more than individual stream speed;&lt;/li&gt;
&lt;li&gt;infrastructure ownership or artifact control is mandatory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;GroqCloud therefore sells more than processor time. It packages a compiler-controlled distributed machine behind a familiar inference contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  NVIDIA licensed the technology; it did not acquire Groq
&lt;/h2&gt;

&lt;p&gt;The NVIDIA relationship is easy to overstate because technology and senior personnel moved together.&lt;/p&gt;

&lt;p&gt;On December 24, 2025, Groq announced a &lt;a href="https://groq.com/newsroom/groq-and-nvidia-enter-non-exclusive-inference-technology-licensing-agreement-to-accelerate-ai-inference-at-global-scale" rel="noopener noreferrer"&gt;non-exclusive inference technology licensing agreement with NVIDIA&lt;/a&gt;. Founder Jonathan Ross, president Sunny Madra, and other Groq staff joined NVIDIA. The same announcement says Groq would remain an independent company under CEO Simon Edwards and that GroqCloud would continue operating.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.spokesman.com/stories/2025/dec/24/nvidia-joining-big-tech-deal-spree-to-license-groq/" rel="noopener noreferrer"&gt;Reuters described&lt;/a&gt; the structure as a technology license plus executive hiring that stopped short of a formal acquisition. Reported deal values remain reporting, not disclosed transaction terms from either company.&lt;/p&gt;

&lt;p&gt;The distinction now appears in two product paths:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;GroqCloud&lt;/strong&gt; remains Groq's independently operated inference service.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA Groq 3 LPX&lt;/strong&gt; uses licensed Groq technology inside NVIDIA's Vera Rubin platform.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;NVIDIA's &lt;a href="https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/" rel="noopener noreferrer"&gt;Groq 3 LPX architecture description&lt;/a&gt; specifies a 256-LPU rack with 128 GB of SRAM, 40 PB/s of SRAM bandwidth, and 640 TB/s of rack-scale communication. NVIDIA announced the product in &lt;a href="https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-full-production-with-world-class-speed-for-agentic-ai" rel="noopener noreferrer"&gt;full production in August 2026&lt;/a&gt;, with initial cloud availability planned later in the year.&lt;/p&gt;

&lt;p&gt;LPX does not establish that every model should run entirely on an LPU rack. NVIDIA presents a heterogeneous design coordinated by Dynamo:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Vera Rubin GPUs can process prefill, attention, and key-value cache.&lt;/li&gt;
&lt;li&gt;LPX can execute latency-sensitive feed-forward or mixture-of-experts decode stages.&lt;/li&gt;
&lt;li&gt;The two systems exchange intermediate activations according to the selected serving pattern.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;NVIDIA calls these patterns &lt;a href="https://developer.nvidia.com/blog/how-nvidia-groq-3-lpx-unlocks-ultrafast-interactivity-at-long-context-on-nvidia-vera-rubin/" rel="noopener noreferrer"&gt;prefill-decode disaggregation, attention-FFN disaggregation, and external-drafter speculative decoding&lt;/a&gt;. The published numbers—including 3,431 output tokens per second on one 100K-context test—come from NVIDIA-operated LPX systems and vendor-selected configurations, even when Artificial Analysis supplied the test harness.&lt;/p&gt;

&lt;p&gt;The evidence floor is narrower than “Groq replaces GPU inference.” LPX shows that a latency-specialized SRAM machine can become one stage inside a larger GPU inference system. GPUs retain the memory-heavy and general-purpose work; LPUs handle selected decode paths where deterministic scheduling and SRAM bandwidth matter most.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's missing
&lt;/h2&gt;

&lt;p&gt;Groq has credible evidence for high output speed and a coherent explanation for why the architecture produces it. The public record remains thinner on aggregate system economics.&lt;/p&gt;

&lt;p&gt;A complete comparison needs the same model, quality target, prompt distribution, context length, concurrency, and service-level objective across systems. It also needs measured rack power, total output throughput, tail latency, compiler coverage, failed-compilation rate, capacity reservation, and cost under sustained load.&lt;/p&gt;

&lt;p&gt;No current public MLPerf result establishes a standardized Groq-to-GPU comparison across those boundaries. GroqCloud API benchmarks prove that users can receive tokens quickly. They do not reveal whether the underlying rack delivers the best aggregate tokens per watt or tokens per dollar.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what
&lt;/h2&gt;

&lt;p&gt;Groq's contribution is larger than a fast inference endpoint. It redraws the processor boundary around the compiler and the connected machine.&lt;/p&gt;

&lt;p&gt;SRAM makes each operation fast and predictable. The compiler turns that memory and compute into a timed execution plan. The interconnect lets many small memory domains behave like one model-serving pipeline. GroqCloud turns that system into an API product, while NVIDIA's licensed LPX path places the same architectural idea inside a heterogeneous GPU platform.&lt;/p&gt;

&lt;p&gt;The open question is operational: as inference splits into prefill, attention, feed-forward, speculative decoding, and cache management, does the winning system use one broad accelerator—or does the compiler become the place where several specialized processors are assembled into one service?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>groq</category>
      <category>inference</category>
      <category>semiconductors</category>
    </item>
    <item>
      <title>d-Matrix Moves the Math Into Memory</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:50:30 +0000</pubDate>
      <link>https://dev.to/amitrix/d-matrix-moves-the-math-into-memory-gd6</link>
      <guid>https://dev.to/amitrix/d-matrix-moves-the-math-into-memory-gd6</guid>
      <description>&lt;p&gt;Inference is becoming a memory-placement problem.&lt;/p&gt;

&lt;p&gt;A modern accelerator can perform arithmetic faster than it can repeatedly move model weights into the compute engines. Autoregressive decoding makes the mismatch visible: every generated token requires another pass through the model, often at a batch size small enough that raw matrix throughput is not the limiting resource.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.d-matrix.ai/product/" rel="noopener noreferrer"&gt;d-Matrix Corsair&lt;/a&gt; attacks that boundary by placing digital arithmetic beside a large SRAM tier. The system still needs off-chip capacity memory, PCIe, Ethernet, compilation, quantization, and distributed execution. The architectural bet is narrower and more interesting than “replace the GPU”: keep the latency-critical working set near the arithmetic and use the rest of the memory hierarchy deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Corsair combines digital in-memory compute with 2 GB of high-bandwidth SRAM per card and up to 256 GB of off-chip capacity memory.&lt;/li&gt;
&lt;li&gt;The useful deployment boundary expands from a paired PCIe card to an eight-card server and then a 64-card rack because larger models must be partitioned across multiple SRAM pools.&lt;/li&gt;
&lt;li&gt;d-Matrix reduces physical adoption friction by using standard PCIe and Ethernet, but model compilation, compression, partitioning, and numerical validation remain accelerator-specific.&lt;/li&gt;
&lt;li&gt;The clearest near-term pattern may be heterogeneous inference: use a specialist for a memory-bound phase while GPUs retain compute-heavy phases.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The bottleneck is the trip to memory
&lt;/h2&gt;

&lt;p&gt;The arithmetic in transformer inference is not uniform.&lt;/p&gt;

&lt;p&gt;Prefill processes many prompt tokens in parallel and can keep a large compute engine busy. Decode generates tokens sequentially. At small batch sizes, the accelerator repeatedly reads model weights to produce relatively little new work. The service becomes sensitive to memory bandwidth, placement, and communication latency.&lt;/p&gt;

&lt;p&gt;Corsair puts a digital in-memory compute engine beside what d-Matrix calls &lt;strong&gt;Performance Memory&lt;/strong&gt;. One card provides 2 GB of that SRAM at a published 150 TB/s, plus as much as 256 GB of &lt;strong&gt;Capacity Memory&lt;/strong&gt; off chip. A two-card pair doubles those figures and connects through the company’s DMX Bridge.&lt;/p&gt;

&lt;p&gt;Those two tiers serve different jobs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Memory tier&lt;/th&gt;
&lt;th&gt;Published card capacity&lt;/th&gt;
&lt;th&gt;Architectural job&lt;/th&gt;
&lt;th&gt;Main constraint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Performance Memory&lt;/td&gt;
&lt;td&gt;2 GB SRAM&lt;/td&gt;
&lt;td&gt;Keep latency-sensitive weights and activations next to the compute engines&lt;/td&gt;
&lt;td&gt;Capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capacity Memory&lt;/td&gt;
&lt;td&gt;Up to 256 GB&lt;/td&gt;
&lt;td&gt;Hold larger models and support throughput-oriented execution&lt;/td&gt;
&lt;td&gt;Lower bandwidth and greater movement cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Distributed rack memory&lt;/td&gt;
&lt;td&gt;128 GB aggregate SRAM and up to 16.4 TB capacity in the 64-card reference rack&lt;/td&gt;
&lt;td&gt;Partition larger models and requests across cards and servers&lt;/td&gt;
&lt;td&gt;Fabric latency, partitioning, and software coordination&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is not one flat pool. The compiler and runtime must decide what fits in the fast tier, how the model is divided, and when data moves between cards or servers.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    MODEL["Model checkpoint"]
    COMP["Aviator compiler&amp;lt;br/&amp;gt;compression and partitioning"]
    CAP["Capacity Memory&amp;lt;br/&amp;gt;larger model storage"]
    PERF["Performance Memory&amp;lt;br/&amp;gt;SRAM beside compute"]
    DIMC["Digital in-memory&amp;lt;br/&amp;gt;compute engines"]
    FABRIC["DMX Bridge, PCIe,&amp;lt;br/&amp;gt;or Ethernet"]
    NEXT["Next card or&amp;lt;br/&amp;gt;inference stage"]

    MODEL --&amp;gt; COMP
    COMP --&amp;gt; CAP
    COMP --&amp;gt; PERF
    CAP --&amp;gt; PERF
    PERF &amp;lt;--&amp;gt; DIMC
    DIMC --&amp;gt; FABRIC
    FABRIC --&amp;gt; NEXT&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The diagram captures the trade: less distance between arithmetic and the fast memory tier, followed by more responsibility for placement and distributed execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  A card is not the complete product
&lt;/h2&gt;

&lt;p&gt;Corsair uses an industry-standard full-height, full-length PCIe Gen5 form factor. That makes it physically familiar, but the deployment unit changes with the model.&lt;/p&gt;

&lt;p&gt;The company describes three reference boundaries:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Boundary&lt;/th&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;Published memory boundary&lt;/th&gt;
&lt;th&gt;Intended use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Card pair&lt;/td&gt;
&lt;td&gt;Two connected Corsair cards&lt;/td&gt;
&lt;td&gt;4 GB Performance Memory; up to 512 GB Capacity Memory&lt;/td&gt;
&lt;td&gt;Smaller models and low-latency stages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Server&lt;/td&gt;
&lt;td&gt;Eight cards&lt;/td&gt;
&lt;td&gt;16 GB Performance Memory; up to 2 TB Capacity Memory&lt;/td&gt;
&lt;td&gt;Multi-card inference within one host&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rack&lt;/td&gt;
&lt;td&gt;Eight servers, 64 cards&lt;/td&gt;
&lt;td&gt;128 GB Performance Memory; up to 16.4 TB Capacity Memory&lt;/td&gt;
&lt;td&gt;Larger models and higher request volume&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The aggregate SRAM determines which model weights can remain in the fastest tier. d-Matrix describes smaller models as an appliance-level fit and larger models as a rack-level problem. That is why the system architecture matters more than the card specification.&lt;/p&gt;

&lt;p&gt;Within a node, PCIe and DMX Bridge connect the accelerators. Across nodes, the design uses Ethernet. d-Matrix’s &lt;a href="https://www.d-matrix.ai/why-we-decoupled-execution-to-accelerate-i-o/" rel="noopener noreferrer"&gt;JetStream&lt;/a&gt; NIC is intended to let accelerators coordinate transfers without routing every step through the host control path. The company says each JetStream connects four Corsair cards to a leaf switch and extends accelerator-to-accelerator execution across servers.&lt;/p&gt;

&lt;p&gt;That is still a distributed system. The architecture reduces some host-mediated handshakes; it does not remove switches, congestion, topology, failure handling, or the need to keep model partitions synchronized.&lt;/p&gt;

&lt;h2&gt;
  
  
  The software stack carries the memory policy
&lt;/h2&gt;

&lt;p&gt;In-memory compute does not make a PyTorch model automatically executable.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.d-matrix.ai/product/" rel="noopener noreferrer"&gt;Aviator software stack&lt;/a&gt; contains model templates, compression tools, an MLIR-based compiler, a distributed inference engine, host runtime, device firmware, and operational tooling. It integrates with PyTorch and Triton DSL at the upper boundary, then produces artifacts for Corsair’s spatial architecture and numerical formats.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Software layer&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model Factory&lt;/td&gt;
&lt;td&gt;Provides distributed model templates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compressor&lt;/td&gt;
&lt;td&gt;Converts weights into supported block floating-point formats&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compiler&lt;/td&gt;
&lt;td&gt;Places and schedules operations for the accelerator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference Engine&lt;/td&gt;
&lt;td&gt;Coordinates execution across multiple cards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Host Runtime&lt;/td&gt;
&lt;td&gt;Manages interaction between server CPUs and cards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chip Runtime&lt;/td&gt;
&lt;td&gt;Launches device work and manages on-chip execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operational tools&lt;/td&gt;
&lt;td&gt;Monitor cards and analyze workload performance&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is the portability boundary. A framework checkpoint may be shared with another accelerator, but the compiled artifact, quantization choices, partitioning plan, and performance profile are hardware-specific.&lt;/p&gt;

&lt;p&gt;The numerical choice matters. Corsair supports microscaling block floating-point formats, including MXINT8 and MXINT4 modes. Lower precision reduces storage and movement, but every model still needs accuracy evaluation. Peak low-precision throughput is not proof that a particular model preserves quality or meets a production latency target.&lt;/p&gt;

&lt;h2&gt;
  
  
  Standard racks reduce one kind of risk
&lt;/h2&gt;

&lt;p&gt;d-Matrix does not require a proprietary wafer-scale chassis or a new liquid-cooled megawatt rack. Corsair is a PCIe card, and the company positions air-cooled server and rack designs that use Ethernet for scale-out.&lt;/p&gt;

&lt;p&gt;That can reduce facility integration work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;standard server mechanics;&lt;/li&gt;
&lt;li&gt;familiar PCIe enumeration;&lt;/li&gt;
&lt;li&gt;Ethernet-based rack networking;&lt;/li&gt;
&lt;li&gt;incremental deployment beside existing accelerator pools;&lt;/li&gt;
&lt;li&gt;independent scaling of specialist and GPU capacity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It does not remove system integration. A 64-card model partition still needs validated servers, network topology, firmware, cooling, power, observability, and replacement procedures. The operator also needs a serving layer that knows which model artifact can run on which pool.&lt;/p&gt;

&lt;p&gt;The deployment pattern is therefore less “one new universal accelerator” and more “one specialized pool behind a stable inference contract.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Heterogeneous inference is the stronger first use
&lt;/h2&gt;

&lt;p&gt;The most useful evidence for Corsair is not a vendor comparison claiming that one rack replaces a GPU fleet. It is a narrower experiment that gives the specialist one phase of the request.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://gimletlabs.ai/blog/low-latency-spec-decode-corsair" rel="noopener noreferrer"&gt;partner engineering study from Gimlet Labs&lt;/a&gt; evaluated Corsair for speculative decoding while leaving prefill and verification on GPUs. The target model was &lt;code&gt;gpt-oss-120b&lt;/code&gt;; a 1.6-billion-parameter draft model ran on two Corsair cards. Gimlet reported a 2–10× improvement in interactivity over its GPU-only speculative-decoding configuration at matched energy-efficiency points.&lt;/p&gt;

&lt;p&gt;The boundary matters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the analysis combined measured and modeled data;&lt;/li&gt;
&lt;li&gt;it modeled an 8K-input, 1K-output coding workload;&lt;/li&gt;
&lt;li&gt;token-acceptance assumptions materially affected the result;&lt;/li&gt;
&lt;li&gt;the exact comparison GPU was not disclosed;&lt;/li&gt;
&lt;li&gt;the test evaluated one inference phase, not complete replacement of the target-model hardware.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That caveat does not weaken the result. It makes the result useful. A memory-bound draft phase fits SRAM-centric hardware differently from compute-heavy prefill and batched verification. The experiment supports a heterogeneous architecture where each phase runs on the hardware that matches its bottleneck.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    REQ["Request"]
    PREFILL["Prefill&amp;lt;br/&amp;gt;GPU pool"]
    DRAFT["Draft tokens&amp;lt;br/&amp;gt;Corsair pool"]
    VERIFY["Verify tokens&amp;lt;br/&amp;gt;GPU pool"]
    ACCEPT{"Accepted?"}
    OUTPUT["Stream output"]

    REQ --&amp;gt; PREFILL
    PREFILL --&amp;gt; DRAFT
    DRAFT --&amp;gt; VERIFY
    VERIFY --&amp;gt; ACCEPT
    ACCEPT --&amp;gt;|yes| OUTPUT
    ACCEPT --&amp;gt;|continue| DRAFT&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;This is disaggregated inference in concrete form. It adds network hops and orchestration, so the specialist must save more time than the handoff consumes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The evidence is still uneven
&lt;/h2&gt;

&lt;p&gt;d-Matrix publishes large claims for Corsair, including tens of thousands of tokens per second and substantial improvements in latency, power efficiency, and total cost. Its &lt;a href="https://www.d-matrix.ai/announcements/d-matrix-unveils-corsair-the-worlds-most-efficient-ai-computing-platform-for-inference-in-datacenters/" rel="noopener noreferrer"&gt;launch announcement&lt;/a&gt; labels the performance and cost estimates preliminary and subject to change.&lt;/p&gt;

&lt;p&gt;The product page also mixes several boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;card-level peak arithmetic;&lt;/li&gt;
&lt;li&gt;server and rack aggregate memory;&lt;/li&gt;
&lt;li&gt;model-specific projected performance;&lt;/li&gt;
&lt;li&gt;preliminary comparisons against H100;&lt;/li&gt;
&lt;li&gt;capacity-mode and performance-mode configurations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those numbers are not interchangeable. A useful independent comparison needs the same model, precision, quality threshold, prompt and output lengths, batching policy, service-level objective, power boundary, and complete system cost.&lt;/p&gt;

&lt;p&gt;No Corsair result appeared in the &lt;a href="https://mlcommons.org/benchmarks/inference-datacenter/" rel="noopener noreferrer"&gt;MLCommons Inference results available during this August 30, 2026 review&lt;/a&gt;. This is a bounded research finding, not proof that a submission cannot exist elsewhere. Public partner and integration announcements establish ecosystem activity, but they do not provide the same evidence as measured production utilization, availability, or cost per accepted token.&lt;/p&gt;

&lt;p&gt;The roadmap adds another evidence boundary. d-Matrix says its &lt;a href="https://www.d-matrix.ai/scaling-ai-inference-with-3dimc/" rel="noopener noreferrer"&gt;Pavehawk 3DIMC test silicon&lt;/a&gt; is operating in the lab and that a future Raptor architecture will use stacked memory. Those are development milestones and targets, not shipping-product measurements.&lt;/p&gt;

&lt;h2&gt;
  
  
  What d-Matrix changes in the landscape
&lt;/h2&gt;

&lt;p&gt;d-Matrix is not trying to recreate the entire NVIDIA stack.&lt;/p&gt;

&lt;p&gt;It attacks one expensive path: inference workloads where moving weights dominates useful arithmetic. The physical interface remains PCIe and Ethernet. The differentiated layer is the memory-compute complex, followed by the compiler and runtime required to use it.&lt;/p&gt;

&lt;p&gt;That makes d-Matrix easier to combine with other accelerators than a closed, facility-scale replacement architecture. It also means the surrounding control plane becomes critical. The model registry needs multiple compiled artifacts. The router needs to understand workload phases. Observability needs to compare latency, throughput, quality, energy, and handoff overhead across dissimilar pools.&lt;/p&gt;

&lt;p&gt;The strategic question is not whether every token should run on in-memory compute. It is whether enough stable, memory-bound work can be isolated to keep the specialist busy without making the serving system harder to operate than the savings justify.&lt;/p&gt;

&lt;p&gt;If heterogeneous inference becomes the normal design, the unresolved boundary moves upward: who owns the optimizer that decides when a request should cross from a GPU into a specialist—and what evidence proves that decision still wins after network, compilation, capacity, and operational costs are included?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>inference</category>
      <category>semiconductors</category>
      <category>memory</category>
    </item>
    <item>
      <title>When the Platform Operator Designs the Chip</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:49:55 +0000</pubDate>
      <link>https://dev.to/amitrix/when-the-platform-operator-designs-the-chip-2c85</link>
      <guid>https://dev.to/amitrix/when-the-platform-operator-designs-the-chip-2c85</guid>
      <description>&lt;p&gt;When the operator of a cloud, model platform, or captive fleet designs the accelerator, the chip stops being the product. The product becomes a vertically integrated compute system: silicon, memory, interconnect, compiler, scheduler, and service assembled behind one operating boundary.&lt;/p&gt;

&lt;p&gt;The thesis: when the platform operator owns those layers together, its infrastructure becomes the computer.&lt;/p&gt;

&lt;p&gt;This is a different competitive model from selling an accelerator card. AWS Trainium, Google TPU, Microsoft Maia, Meta MTIA, and OpenAI Jalapeño all pursue it, but they are not at the same maturity or aimed at the same customer. Trainium and TPU are external cloud platforms. Maia is entering Azure infrastructure. MTIA is a captive production system. Jalapeño is working silicon still moving through qualification.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The real unit of platform-designed silicon is a managed system, not a chip: the operator controls compilation, topology, scheduling, capacity, and service delivery.&lt;/li&gt;
&lt;li&gt;Trainium exposes the cleanest cloud hierarchy—NeuronCore to chip to UltraServer to UltraCluster—while TPU has the deepest published history of co-designing matrix engines, XLA, optical fabrics, and pod scheduling.&lt;/li&gt;
&lt;li&gt;Maia, MTIA, and Jalapeño validate the same vertical-integration thesis at different stages, but their public evidence does not yet support treating them as interchangeable cloud products.&lt;/li&gt;
&lt;li&gt;The trade is architectural: tighter co-design can improve utilization and economics, while moving portability upward from device binaries to models, frameworks, APIs, and workload-routing policy.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The platform is the system boundary
&lt;/h2&gt;

&lt;p&gt;A merchant accelerator has to fit into systems that other companies assemble. A platform-designed accelerator starts with a different constraint: it can assume the operator also controls the host, network, rack topology, compiler fleet, workload scheduler, and service API.&lt;/p&gt;

&lt;p&gt;That changes where optimization happens. A compiler can target known memory layouts. A scheduler can place a job into a known accelerator topology. The network can expose collectives that match the chip. Capacity planning can treat thousands of devices as one managed pool. The customer requests instances, slices, reservations, pods, or a managed model service rather than buying the underlying board.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    A["Model and framework&amp;lt;br/&amp;gt;PyTorch, JAX, service API"] --&amp;gt; B["Compiler and graph optimizer&amp;lt;br/&amp;gt;Neuron or XLA-class layer"]
    B --&amp;gt; D["Capacity scheduler&amp;lt;br/&amp;gt;selects an available topology"]
    D --&amp;gt; E["Selected cluster, pod,&amp;lt;br/&amp;gt;or accelerator slice"]
    E --&amp;gt; C["Device runtime and&amp;lt;br/&amp;gt;collective libraries"]
    C --&amp;gt; F["Scale-up server or&amp;lt;br/&amp;gt;accelerator domain"]
    F --&amp;gt; G["Accelerator + HBM +&amp;lt;br/&amp;gt;local interconnect"]

    H["Cloud control plane"] --&amp;gt; B
    H --&amp;gt; D
    H --&amp;gt; E
    H --&amp;gt; F&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The diagram is the important part of the argument. The platform operator is not merely substituting one arithmetic engine for another. It is closing the control loop from model graph to physical placement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trainium turns EC2 into the machine
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://aws.amazon.com/ec2/instance-types/trn3/" rel="noopener noreferrer"&gt;Trainium3&lt;/a&gt; makes the hierarchy unusually explicit. One chip contains eight NeuronCores and 144 GB of HBM3e with 4.9 TB/s of memory bandwidth. A &lt;code&gt;trn3.48xlarge&lt;/code&gt; instance contains 16 chips. A Trn3 UltraServer joins nine instances—144 chips—through NeuronSwitch-v1 and a 28.8 TB/s NeuronFabric scale-up domain. UltraCluster 3.0 then connects UltraServers through Elastic Fabric Adapter networking.&lt;/p&gt;

&lt;p&gt;That hierarchy matters more than any isolated TOPS number:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NeuronCore
  → Trainium3 chip
  → EC2 instance
  → Trn3 UltraServer
  → UltraCluster 3.0
  → managed training or inference service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The software path is equally deliberate. The &lt;a href="https://aws.amazon.com/machine-learning/neuron/" rel="noopener noreferrer"&gt;AWS Neuron SDK&lt;/a&gt; connects PyTorch and JAX to Neuron compilers, runtimes, collective communication libraries, profiling, and monitoring. &lt;a href="https://awsdocs-neuron.readthedocs-hosted.com/en/latest/nki/index.html" rel="noopener noreferrer"&gt;Neuron Kernel Interface&lt;/a&gt; gives developers a lower-level route for custom kernels without exposing the hardware as a generic CUDA-compatible device.&lt;/p&gt;

&lt;p&gt;Trainium3 also shows why headline generational numbers need careful reading. AWS reports 2.52 PFLOPS of FP8 dense compute and 10.1 PFLOPS at FP4, alongside 1.2 PFLOPS of BF16 dense compute. The largest gains come from lower precision, sparsity, memory bandwidth, and a larger connected system—not an equal multiplier for every model and numerical format. The usable gain depends on whether the compiler maps a workload onto those features without losing model quality or creating excess communication.&lt;/p&gt;

&lt;p&gt;The cloud integration reaches beyond instances. Trainium is exposed through EC2 and integrated into managed paths including &lt;a href="https://aws.amazon.com/ai/machine-learning/trainium/" rel="noopener noreferrer"&gt;Amazon SageMaker AI and Amazon EKS&lt;/a&gt;. AWS also offers dedicated &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/ai-factories/" rel="noopener noreferrer"&gt;AI Factories&lt;/a&gt; that combine accelerator infrastructure, networking, services, and customer-selected data-center locations.&lt;/p&gt;

&lt;p&gt;The maturity signal is customer access: Trainium2 and Trainium3 are purchasable cloud infrastructure rather than lab prototypes. The evidence gap is comparability. AWS publishes specifications, workload claims, and customer adoption, but the public record offers less standardized third-party accelerator benchmarking than the NVIDIA and Google ecosystems. Trainium is a real cloud platform whose strongest evidence is production availability and disclosed usage, not a broad set of independently normalized chip measurements.&lt;/p&gt;

&lt;h2&gt;
  
  
  TPU makes the pod programmable
&lt;/h2&gt;

&lt;p&gt;Google’s TPU architecture has followed this systems thesis for a decade. The original &lt;a href="https://research.google/pubs/in-datacenter-performance-analysis-of-a-tensor-processing-unit/" rel="noopener noreferrer"&gt;in-datacenter TPU paper&lt;/a&gt; described a domain-specific processor built around matrix multiplication rather than the instruction and cache machinery of a general-purpose CPU or GPU. The architectural through-line is systolic execution: data flows through an array of multiply-accumulate units while software arranges computation and movement around the array.&lt;/p&gt;

&lt;p&gt;The more consequential transition arrived at pod scale. &lt;a href="https://arxiv.org/abs/2304.01433" rel="noopener noreferrer"&gt;TPU v4&lt;/a&gt; combined 4,096 chips with optical circuit switches that could reconfigure the interconnect topology. Optical switching let Google route around unavailable components, change topology for different workloads, and improve system availability without wiring every connection through a fixed electrical fabric.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.cloud.google.com/tpu/docs/tpu7x" rel="noopener noreferrer"&gt;Ironwood, or TPU7x&lt;/a&gt;, extends that model for inference. Each chip provides 192 GiB of HBM3e and 7.38 TB/s of memory bandwidth. A cube contains 64 chips, and a full pod reaches 9,216 chips through the Inter-Chip Interconnect and optical fabric. Google exposes smaller slices through its multislice software, so the physical pod becomes a capacity pool rather than one indivisible machine.&lt;/p&gt;

&lt;p&gt;The software boundary is XLA. &lt;a href="https://docs.cloud.google.com/tpu/docs/system-architecture-tpu-vm" rel="noopener noreferrer"&gt;Cloud TPU system architecture&lt;/a&gt; places the user program on a TPU virtual machine while XLA compiles framework graphs into TPU executables. JAX is the native-feeling path because its transformations and array semantics align with XLA, while &lt;a href="https://docs.pytorch.org/xla/master/" rel="noopener noreferrer"&gt;PyTorch/XLA&lt;/a&gt; lets PyTorch programs target the same compiler stack.&lt;/p&gt;

&lt;p&gt;The scheduler is part of the architecture. &lt;a href="https://cloud.google.com/tpu/docs/queued-resources" rel="noopener noreferrer"&gt;queued resources&lt;/a&gt; let users request TPU capacity that Google Cloud provisions when the topology is available. &lt;a href="https://docs.cloud.google.com/tpu/docs/request-using-flex-start" rel="noopener noreferrer"&gt;flex-start VMs&lt;/a&gt; add a managed queue for workloads that can wait for capacity. Reserved capacity and calendar modes serve workloads that need predictable start times. These are not billing decorations; they decide whether a distributed executable receives the physical shape it was compiled to use.&lt;/p&gt;

&lt;p&gt;The next generation makes workload specialization more explicit. Google describes &lt;a href="https://cloud.google.com/blog/products/compute/tpu8-technical-deep-dive" rel="noopener noreferrer"&gt;TPU 8&lt;/a&gt; as two designs: TPU 8t for training and TPU 8i for inference. TPU 8t pairs two compute dies with eight HBM3e stacks and introduces an expanded optical scale-up domain. TPU 8i uses four HBM3e stacks and targets inference economics. At this cutoff, TPU 8 remains an early-access system, while Ironwood has the clearer public cloud operating record.&lt;/p&gt;

&lt;p&gt;Google also has stronger standardized evidence than the other cloud ASIC programs. Its published &lt;a href="https://cloud.google.com/blog/products/compute/google-cloud-mlperf-training-v6-results" rel="noopener noreferrer"&gt;MLPerf Training v6.0 results&lt;/a&gt; cover TPU v6e, TPU7x, and NVIDIA systems. That does not make every TPU claim independently proven—vendor submissions remain optimized submissions—but it creates a defined benchmark boundary that other stacks often lack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five vertically integrated stacks, five maturity levels
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stack&lt;/th&gt;
&lt;th&gt;Real system boundary&lt;/th&gt;
&lt;th&gt;Compiler and software path&lt;/th&gt;
&lt;th&gt;Availability by Aug. 30, 2026&lt;/th&gt;
&lt;th&gt;Public evidence boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://aws.amazon.com/ec2/instance-types/trn3/" rel="noopener noreferrer"&gt;AWS Trainium3&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;NeuronCore → chip → instance → 144-chip UltraServer → UltraCluster&lt;/td&gt;
&lt;td&gt;PyTorch/JAX → Neuron compiler, runtime, collectives, NKI&lt;/td&gt;
&lt;td&gt;Customer-accessible EC2 infrastructure&lt;/td&gt;
&lt;td&gt;Detailed vendor specifications and production adoption; limited standardized cross-vendor results&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://docs.cloud.google.com/tpu/docs/system-architecture-tpu-vm" rel="noopener noreferrer"&gt;Google TPU&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;TensorCore → chip → cube/slice → optically connected pod&lt;/td&gt;
&lt;td&gt;JAX/PyTorch → XLA → TPU runtime and ICI collectives&lt;/td&gt;
&lt;td&gt;Ironwood available; TPU 8 in early access&lt;/td&gt;
&lt;td&gt;Architecture papers, cloud documentation, and MLPerf submissions; TPU 8 evidence remains early&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://blogs.microsoft.com/blog/2026/01/26/maia-200-the-ai-accelerator-built-for-inference/" rel="noopener noreferrer"&gt;Microsoft Maia 200&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Accelerator → Azure server and network → regional inference fleet&lt;/td&gt;
&lt;td&gt;PyTorch models → Maia SDK, compiler, Triton language, firmware and runtime&lt;/td&gt;
&lt;td&gt;Deployed in Azure regions; SDK preview&lt;/td&gt;
&lt;td&gt;Microsoft specifications and workload claims; external customer and benchmark evidence remains narrow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://engineering.fb.com/2026/08/24/networking-traffic/mtia-300-meta-training-chip-built-in-nics/" rel="noopener noreferrer"&gt;Meta MTIA 300&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Accelerator → eight-chip host → Ethernet scale-up/scale-out fleet&lt;/td&gt;
&lt;td&gt;PyTorch → in-house compiler, runtime, PyTorch Distributed, TorchTitan&lt;/td&gt;
&lt;td&gt;Production use inside Meta&lt;/td&gt;
&lt;td&gt;Concrete production workload and fleet design; not offered as a general cloud service&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip" rel="noopener noreferrer"&gt;OpenAI Jalapeño&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Accelerator → server → rack → OpenAI inference fleet&lt;/td&gt;
&lt;td&gt;PyTorch → Triton-based kernels, compiler, runtime, observability&lt;/td&gt;
&lt;td&gt;Engineering samples; deployment targeted for late 2026&lt;/td&gt;
&lt;td&gt;Vendor-published silicon measurements; production qualification still underway&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The table is not a ranking. Availability, public programmability, and benchmark disclosure answer different questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Maia, MTIA, and Jalapeño narrow the target
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://azure.microsoft.com/en-us/blog/maia-200-the-ai-accelerator-built-for-inference/" rel="noopener noreferrer"&gt;Maia 200&lt;/a&gt; is an inference accelerator with 216 GB of HBM3e and 7 TB/s of memory bandwidth. Microsoft says it is deployed in its U.S. Central data-center region, with a second region planned, and that its workloads include models from the Microsoft Foundry catalog. The Maia SDK includes a Triton compiler, PyTorch integration, kernel libraries, a low-level programming language, and simulation tools.&lt;/p&gt;

&lt;p&gt;That is credible system construction, but the public service boundary is still forming. Maia 200 is not documented as a broadly selectable Azure VM family comparable to established GPU instances or current TPU and Trainium offerings. The evidence establishes deployed hardware and an SDK preview, not an open production ecosystem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://engineering.fb.com/2026/08/24/networking-traffic/mtia-300-meta-training-chip-built-in-nics/" rel="noopener noreferrer"&gt;MTIA 300&lt;/a&gt; is farther along as captive infrastructure. Meta reports production use training a recommendation model with more than 150 billion parameters. Each accelerator includes two network interfaces, and an eight-accelerator host uses Ethernet for both scale-up and scale-out communication. The system plugs into PyTorch Distributed and TorchTitan while Meta’s compiler and runtime manage the device-specific path.&lt;/p&gt;

&lt;p&gt;MTIA therefore supplies strong evidence for vertical integration and weak evidence for external portability. It is optimized around Meta’s models, data centers, network, and fleet software. That is the point. A captive accelerator does not need to become a universal platform if it can remove cost or capacity pressure from a large, stable workload.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip" rel="noopener noreferrer"&gt;OpenAI Jalapeño&lt;/a&gt; takes the same idea into the model-provider layer. OpenAI and Broadcom co-designed an inference accelerator with 128 GB of HBM3e and 3.6 TB/s of bandwidth, then built a PyTorch-oriented compiler, runtime, kernel library, observability layer, server, and rack around it. OpenAI reports first-silicon measurements from engineering samples and plans deployment in its inference fleet by the end of 2026.&lt;/p&gt;

&lt;p&gt;Jalapeño is the least mature system in this comparison. The silicon is operating and the software stack exists, but production qualification, manufacturing scale, and fleet behavior remain open. Its significance is strategic rather than proven economic superiority: a model provider is designing the execution substrate around the distribution of its own workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Portability moves above the device
&lt;/h2&gt;

&lt;p&gt;Platform-designed silicon relocates portability rather than eliminating lock-in. Device binaries, kernels, collectives, topology assumptions, and compiled artifacts remain hardware-specific even when several stacks accept PyTorch.&lt;/p&gt;

&lt;p&gt;The portable boundary sits at model identity, evaluation, service contracts, and outcome telemetry. &lt;a href="https://artificialcuriositylabs.ai/posts/2026-08-30-no-universal-ai-runtime-portable-control-plane/" rel="noopener noreferrer"&gt;There Is No Universal AI Runtime—So Build a Portable Control Plane&lt;/a&gt; develops that pattern; the important point here is that vertical integration makes execution less interchangeable while making the complete service easier for one operator to optimize.&lt;/p&gt;

&lt;h2&gt;
  
  
  What remains unproven
&lt;/h2&gt;

&lt;p&gt;Vertical integration creates a measurement problem. The provider can optimize hardware, software, placement, and pricing together, but outsiders rarely observe every boundary. Peak arithmetic says little about model quality. Chip power says little about facility energy. Tokens per second say little without batch size, latency distribution, context length, model version, and service utilization.&lt;/p&gt;

&lt;p&gt;Trainium needs more standardized independent results. TPU 8 needs production evidence beyond early access. Maia needs a clearer customer-facing service boundary. MTIA needs no external ecosystem to succeed, but that makes cross-platform comparison difficult. Jalapeño needs manufacturing and fleet evidence after qualification.&lt;/p&gt;

&lt;p&gt;The strongest evaluation unit is therefore the complete service: supported model, required code changes, compilation reliability, useful throughput, latency, accuracy, availability, energy boundary, and total cost over a sustained workload. Any comparison that stops at the chip cuts the system at the wrong layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what
&lt;/h2&gt;

&lt;p&gt;Platform-designed silicon changes the question from “Which chip is fastest?” to “Which operator can turn its workload, compiler, network, scheduler, and capacity into the best service?”&lt;/p&gt;

&lt;p&gt;Trainium and TPU show the mature forms of that strategy. Maia and MTIA show how large operators narrow the architecture around workloads they already understand. Jalapeño pushes the boundary one layer higher: the model provider starts designing the machine that serves the model.&lt;/p&gt;

&lt;p&gt;The open thread is control. If the cloud becomes the computer, how much of that computer must remain visible—topology, compilation, queueing, power, failure domains, and pricing—for builders to make an informed infrastructure choice rather than accept a black-box service?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aicompute</category>
      <category>accelerators</category>
      <category>cloudinfrastructure</category>
      <category>systemsarchitecture</category>
    </item>
    <item>
      <title>Cerebras Moved the Cluster Boundary Onto a Wafer</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:49:19 +0000</pubDate>
      <link>https://dev.to/amitrix/cerebras-moved-the-cluster-boundary-onto-a-wafer-38g</link>
      <guid>https://dev.to/amitrix/cerebras-moved-the-cluster-boundary-onto-a-wafer-38g</guid>
      <description>&lt;p&gt;Cerebras makes one of the largest boundaries in computing disappear: the edge of an individual accelerator package.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.cerebras.ai/chip" rel="noopener noreferrer"&gt;Wafer-Scale Engine&lt;/a&gt; keeps an entire silicon wafer intact, connecting 900,000 compute cores and 44 GB of distributed SRAM through an on-wafer mesh. Computation that would cross packages, boards, switches, and cables on a conventional cluster can remain on silicon.&lt;/p&gt;

&lt;p&gt;That does not make distributed systems disappear. It moves their first boundary outward.&lt;/p&gt;

&lt;p&gt;The useful Cerebras hierarchy is now &lt;strong&gt;WSE → CS system → Nexus rack → multi-system deployment&lt;/strong&gt;. Training adds MemoryX and SwarmX outside the wafer. Large inference deployments add Direct Wafer Links, Ethernet, rack power, liquid cooling, and sometimes a separate prefill engine. The wafer becomes a larger processor, but the processor still lives inside a system.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Wafer-scale integration replaces thousands of package-level connections with a 2D mesh spanning 900,000 cores, but the 44 GB SRAM capacity still forces a memory hierarchy above the wafer.&lt;/li&gt;
&lt;li&gt;Cerebras training stores model weights in MemoryX, broadcasts them through SwarmX, and keeps activations on the wafer; this simplifies model placement without eliminating cluster communication.&lt;/li&gt;
&lt;li&gt;CS-4 puts three WSE-3 Turbo processors into an active-power, liquid-cooled Nexus rack and connects racks through Direct Wafer Links or RoCE v2 Ethernet.&lt;/li&gt;
&lt;li&gt;The compiler is part of the architecture because it maps a model onto distributed cores, SRAM, routes, and data movement rather than targeting a conventional shared-memory accelerator.&lt;/li&gt;
&lt;li&gt;Independent evidence supports high output speed on specific hosted inference models and useful scientific-computing results. It does not establish universal training cost, aggregate throughput, or energy leadership.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Start with the boundary Cerebras removes
&lt;/h2&gt;

&lt;p&gt;A conventional accelerator begins as a die, becomes a package with external memory, joins other packages on a board, and then communicates across servers through network adapters and switches. Every boundary adds serialization, protocol overhead, distance, and power.&lt;/p&gt;

&lt;p&gt;Cerebras uses the wafer as the compute package. The current &lt;a href="https://www.cerebras.ai/chip" rel="noopener noreferrer"&gt;WSE-3 architecture&lt;/a&gt; contains four trillion transistors across 46,225 mm² of silicon. Its 900,000 cores each sit beside local SRAM and a router. The routers form a 2D mesh that extends across the wafer.&lt;/p&gt;

&lt;p&gt;The important architectural difference is not transistor count. It is locality.&lt;/p&gt;

&lt;p&gt;The company’s &lt;a href="https://www.cerebras.ai/chip" rel="noopener noreferrer"&gt;WSE architecture material&lt;/a&gt; describes 48 KB of independently addressed SRAM per core and explicit communication through the fabric. There is no large shared memory behind a cache hierarchy. Data moves through compiler-configured routes between small compute-and-memory tiles.&lt;/p&gt;

&lt;p&gt;That creates an unusual trade:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What Cerebras collapses&lt;/th&gt;
&lt;th&gt;What remains outside&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;WSE-3&lt;/td&gt;
&lt;td&gt;Package-to-package links among 900,000 cores&lt;/td&gt;
&lt;td&gt;Model weights larger than 44 GB, host I/O, power, cooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CS-3 system&lt;/td&gt;
&lt;td&gt;Wafer, control, and I/O in one appliance&lt;/td&gt;
&lt;td&gt;MemoryX, input servers, management, multi-system fabric&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training cluster&lt;/td&gt;
&lt;td&gt;Model placement across accelerators through weight streaming&lt;/td&gt;
&lt;td&gt;Weight storage, gradient reduction, data ingestion, scheduling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CS-4 on Nexus&lt;/td&gt;
&lt;td&gt;Three wafer processors, rack power, cooling, and I/O&lt;/td&gt;
&lt;td&gt;Other racks, facility power and water, heterogeneous prefill&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosted inference&lt;/td&gt;
&lt;td&gt;Hardware exposed through a familiar API&lt;/td&gt;
&lt;td&gt;Model availability, queueing, concurrency, service economics&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Wafer scale removes the smallest expensive boundaries. It does not remove memory capacity, network scale-out, or data-center operations.&lt;/p&gt;

&lt;h2&gt;
  
  
  SRAM changes the memory problem
&lt;/h2&gt;

&lt;p&gt;HBM-based accelerators place large, fast DRAM stacks beside the processor. Cerebras takes the opposite path: distribute SRAM across the compute fabric and place every small memory bank close to its core.&lt;/p&gt;

&lt;p&gt;HBM and SRAM solve different problems. HBM is the large warehouse: it holds far more data, but cores reach it through shared memory controllers and an off-die interface. SRAM is the workbench: it consumes much more silicon area per byte, but puts the active data beside the arithmetic that uses it.&lt;/p&gt;

&lt;p&gt;GPUs already use both. NVIDIA calls each repeated compute neighborhood a &lt;strong&gt;streaming multiprocessor&lt;/strong&gt;, or SM; it contains compute cores, registers, cache, and programmable shared SRAM. The &lt;a href="https://docs.nvidia.com/cuda/blackwell-tuning-guide/" rel="noopener noreferrer"&gt;Blackwell tuning guide&lt;/a&gt; lists up to 228 KB of shared memory per B200 SM and 126 MB of shared L2 cache on GB200, backed by as much as 180 GB of HBM. The distinction is not SRAM versus no SRAM. It is how much SRAM exists, where it sits, and whether software can keep the active working set there.&lt;/p&gt;

&lt;p&gt;Cerebras gives each of its 900,000 cores 48 KB of private SRAM. One bank cannot hold a model or even a large layer, but it can hold the small tile of instructions, activations, and intermediate values that its core is processing now. Multiplied across the wafer, those small banks become 43.2 GB—marketed as 44 GB—and can all serve their nearby cores in parallel. The &lt;a href="https://www.cerebras.ai/blog/supercharge-your-hpc-research-with-the-cerebras-sdk" rel="noopener noreferrer"&gt;Cerebras SDK description&lt;/a&gt; makes that locality explicit: each processing element owns its 48 KB bank rather than sharing it with the rest of the wafer.&lt;/p&gt;

&lt;p&gt;That is why a small amount beside each core creates a large system effect. The compiler can move a tile into local SRAM once, reuse it across many operations, and exchange nearby values through the mesh instead of repeatedly returning to HBM. The advantage depends on locality: when the active data does not fit or cannot be reused, external memory and communication become the constraint again.&lt;/p&gt;

&lt;p&gt;The result is high local bandwidth and low movement distance. The &lt;a href="https://www.cerebras.ai/chip" rel="noopener noreferrer"&gt;WSE-3&lt;/a&gt; publishes 21 PB/s of memory bandwidth and 44 GB of on-wafer SRAM. The newer &lt;a href="https://www.cerebras.ai/cs4" rel="noopener noreferrer"&gt;WSE-3 Turbo inside CS-4&lt;/a&gt; doubles the vendor-published figures to 43.2 PB/s and 250 sparse FP16 PFLOPS per wafer by running the existing architecture at a higher operating point.&lt;/p&gt;

&lt;p&gt;Those are peak hardware specifications, not application results. They still reveal the architecture’s bet: spend silicon area on distributed memory and communication rather than attach a larger external HBM pool to each compute die.&lt;/p&gt;

&lt;p&gt;SRAM solves bandwidth and creates a capacity constraint. Forty-four gigabytes cannot hold the weights of a frontier model, optimizer state, activations, and temporary data at once. Cerebras therefore separates two kinds of state:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Activations and intermediate values remain in distributed SRAM on the wafer.&lt;/li&gt;
&lt;li&gt;Model weights live in external MemoryX storage and stream through the system layer by layer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The wafer is not a self-contained memory island. It is a compute surface fed by a separate weight-memory system.&lt;/p&gt;

&lt;h2&gt;
  
  
  MemoryX and SwarmX bring the cluster back
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://training-docs.cerebras.ai/rel-2.5.0/concepts/cerebras-wafer-scale-cluster" rel="noopener noreferrer"&gt;Cerebras Wafer-Scale Cluster documentation&lt;/a&gt; makes the distributed architecture explicit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MemoryX&lt;/strong&gt; stores model weights and streams them toward the compute systems. During training, updated gradients travel back toward the weight-memory layer. &lt;strong&gt;SwarmX&lt;/strong&gt; broadcasts weights to multiple CS systems and reduces gradients in the opposite direction. Input servers prepare training data, while management servers schedule and coordinate the cluster.&lt;/p&gt;

&lt;p&gt;This avoids one class of complexity. Developers do not manually split individual layers across a mesh of accelerator memories in the same way required by many tensor- and pipeline-parallel GPU configurations. Cerebras presents weight streaming and data-parallel scale-out as the main abstraction.&lt;/p&gt;

&lt;p&gt;The communication has not vanished. It has become more structured:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    CODE["PyTorch model or CSL program"]
    COMPILER["Cerebras compiler&amp;lt;br/&amp;gt;graph placement + routes + memory"]
    MANAGER["Cluster management&amp;lt;br/&amp;gt;scheduling + input services"]

    subgraph TRAIN["Training path"]
        MX["MemoryX&amp;lt;br/&amp;gt;weights + optimizer state"]
        SX["SwarmX&amp;lt;br/&amp;gt;weight broadcast + gradient reduction"]
        C1["CS-3&amp;lt;br/&amp;gt;WSE-3 + SRAM"]
        C2["CS-3&amp;lt;br/&amp;gt;WSE-3 + SRAM"]
        C3["Additional CS systems"]
        MX --&amp;gt; SX
        SX --&amp;gt; C1
        SX --&amp;gt; C2
        SX --&amp;gt; C3
        C1 --&amp;gt;|"gradients"| SX
        C2 --&amp;gt;|"gradients"| SX
        C3 --&amp;gt;|"gradients"| SX
    end

    subgraph INFER["Current rack-scale inference path"]
        PREFILL["Optional external prefill&amp;lt;br/&amp;gt;GPU or ASIC"]
        NEXUS["CS-4 Nexus rack&amp;lt;br/&amp;gt;power + liquid cooling + I/O"]
        W1["WSE-3 Turbo"]
        W2["WSE-3 Turbo"]
        W3["WSE-3 Turbo"]
        SCALE["Direct Wafer Links&amp;lt;br/&amp;gt;or RoCE v2 Ethernet"]
        PREFILL --&amp;gt; NEXUS
        NEXUS --&amp;gt; W1
        NEXUS --&amp;gt; W2
        NEXUS --&amp;gt; W3
        NEXUS --&amp;gt; SCALE
    end

    CODE --&amp;gt; COMPILER
    COMPILER --&amp;gt; MANAGER
    MANAGER --&amp;gt; MX
    COMPILER --&amp;gt; NEXUS&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The processor boundary grows from one die to one wafer. The training boundary still includes storage, fabric, preprocessing, and management.&lt;/p&gt;

&lt;h2&gt;
  
  
  The compiler is part of the machine
&lt;/h2&gt;

&lt;p&gt;A wafer containing 900,000 small cores is not useful if software treats it as a large conventional CPU or GPU.&lt;/p&gt;

&lt;p&gt;The high-level path begins with supported &lt;a href="https://www.cerebras.ai/product-software" rel="noopener noreferrer"&gt;PyTorch models and the Cerebras software stack&lt;/a&gt;. The Cerebras Graph Compiler converts the model into an executable placement across cores, local SRAM, and communication routes. The lower-level &lt;a href="https://sdk.cerebras.ai/csl/csl-compiler" rel="noopener noreferrer"&gt;Cerebras Software Language compiler&lt;/a&gt; exposes the wafer as a programmable fabric for custom kernels and scientific workloads.&lt;/p&gt;

&lt;p&gt;This compiler role is deeper than translating arithmetic instructions. It determines where tensors live, how data flows, and which paths connect computation across the wafer. Hardware and compiler share responsibility for making the wafer behave like one processor.&lt;/p&gt;

&lt;p&gt;The cost is a narrower compatibility boundary. A PyTorch model is source material, not proof that every operator, dynamic graph, custom kernel, numerical mode, and training recipe maps without changes. The relevant portability question is not “Does it support PyTorch?” It is “Does this exact model compile, converge, and meet its service target on the supported stack?”&lt;/p&gt;

&lt;p&gt;That distinction matters for both training and inference. A broad GPU runtime can absorb rapidly changing model architectures through a large library and kernel ecosystem. Cerebras can win when the compiler has a good mapping and the wafer’s locality matches the workload. Unsupported or inefficient mappings remain a software problem rather than a silicon problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  CS-4 turns wafer scale into rack scale
&lt;/h2&gt;

&lt;p&gt;The August 2026 &lt;a href="https://www.cerebras.ai/cs4" rel="noopener noreferrer"&gt;CS-4 system&lt;/a&gt; changes the deployment unit. It is not a new WSE-4 chip. CS-4 uses three WSE-3 Turbo processors inside the first implementation of the &lt;strong&gt;Nexus&lt;/strong&gt; rack-scale platform.&lt;/p&gt;

&lt;p&gt;Each wafer sits in a rear-mounted “Wafer-Scale Backpack” containing local power conversion, direct liquid cooling, I/O, and control electronics. The rack provides active power infrastructure. Cerebras publishes 750 sparse FP16 PFLOPS, 129.6 PB/s of memory bandwidth, and 7.2 Tbit/s of I/O for the three-wafer system. These are vendor peak specifications.&lt;/p&gt;

&lt;p&gt;Nexus supports two scale-out paths:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Direct Wafer Links&lt;/strong&gt; connect Cerebras systems without an intervening switch and target low-latency wafer-to-wafer communication.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RoCE v2 RDMA over Ethernet&lt;/strong&gt; connects the rack to standard data-center and heterogeneous infrastructure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;CS-4 also formalizes disaggregated inference. An external GPU or ASIC can process prefill—the prompt-heavy first phase—then transfer model state to Cerebras for latency-sensitive token generation. The &lt;a href="https://www.cerebras.ai/blog/introducing-cerebras-cs-4" rel="noopener noreferrer"&gt;CS-4 announcement&lt;/a&gt; names AMD and Trainium as potential prefill platforms.&lt;/p&gt;

&lt;p&gt;This is an important correction to the “cluster became a processor” thesis. Cerebras can make decode behave like a tightly coupled wafer-scale computation while still depending on another processor, a state-transfer protocol, Ethernet, and service-level routing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Power and cooling are the least transparent boundary
&lt;/h2&gt;

&lt;p&gt;Wafer scale shortens communication distance, but it concentrates heat.&lt;/p&gt;

&lt;p&gt;CS-4 is an active-power rack with three direct-liquid-cooled compute backpacks. Cerebras says moving power conversion within 0.5 millimeters of the processor reduces board-level loss and permits twice the power delivery to WSE-3 Turbo. The public &lt;a href="https://www.cerebras.ai/cs4" rel="noopener noreferrer"&gt;CS-4 specifications&lt;/a&gt; do not publish rack input power, coolant temperatures, flow requirements, or facility heat-rejection assumptions.&lt;/p&gt;

&lt;p&gt;That missing number matters. Cerebras claims up to ten times more throughput per watt than CS-3, but the comparison combines a new operating point, system design, workload selection, and projected results. Without measured rack input, cooling overhead, model, precision, concurrency, and quality target, it cannot support a facility-level efficiency conclusion.&lt;/p&gt;

&lt;p&gt;The deployment comparison is therefore architectural rather than economic:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment&lt;/th&gt;
&lt;th&gt;Compute unit&lt;/th&gt;
&lt;th&gt;Memory and fabric&lt;/th&gt;
&lt;th&gt;Facility boundary&lt;/th&gt;
&lt;th&gt;Evidence available&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CS-3 appliance&lt;/td&gt;
&lt;td&gt;One WSE-3&lt;/td&gt;
&lt;td&gt;44 GB SRAM; external MemoryX and SwarmX for training&lt;/td&gt;
&lt;td&gt;System plus external memory, fabric, input, and management components&lt;/td&gt;
&lt;td&gt;Shipping systems, public docs, customer installations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-CS-3 training cluster&lt;/td&gt;
&lt;td&gt;Multiple WSE-3 systems&lt;/td&gt;
&lt;td&gt;MemoryX weight storage; SwarmX broadcast and reduction&lt;/td&gt;
&lt;td&gt;Multiple systems, input and management servers, network and storage&lt;/td&gt;
&lt;td&gt;Vendor scaling claims; selected research publications&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CS-4 Nexus rack&lt;/td&gt;
&lt;td&gt;Three WSE-3 Turbo processors&lt;/td&gt;
&lt;td&gt;Direct Wafer Links and RoCE v2&lt;/td&gt;
&lt;td&gt;Active rack power and direct liquid cooling; public power draw absent&lt;/td&gt;
&lt;td&gt;Announced Aug. 2026; shipments stated to begin during the quarter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cerebras cloud inference&lt;/td&gt;
&lt;td&gt;Provider-operated wafer systems&lt;/td&gt;
&lt;td&gt;API-hidden placement and service fabric&lt;/td&gt;
&lt;td&gt;Operator-owned power, cooling, capacity, and queueing&lt;/td&gt;
&lt;td&gt;Independently observable endpoint speed; limited facility data&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The rack is easier to reason about mechanically than economically. The modules, cooling approach, and I/O modes are visible. The watts, water, sustained utilization, and failure behavior are not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Training and inference are different evidence problems
&lt;/h2&gt;

&lt;p&gt;Cerebras has credible evidence, but it is fragmented across different workload boundaries.&lt;/p&gt;

&lt;p&gt;For hosted inference, &lt;a href="https://artificialanalysis.ai/providers/cerebras" rel="noopener noreferrer"&gt;Artificial Analysis tracks Cerebras endpoints&lt;/a&gt; and has measured high output speeds on specific available models. This is stronger than a vendor slide because an external evaluator sends requests to a live service. It still does not establish aggregate throughput at high concurrency, queueing behavior during capacity pressure, power consumption, or performance on unsupported models.&lt;/p&gt;

&lt;p&gt;For scientific computing, a &lt;a href="https://www.nature.com/articles/s41467-025-63798-0" rel="noopener noreferrer"&gt;2025 Nature Communications paper&lt;/a&gt; reports large molecular-dynamics simulations on the WSE. That result demonstrates that the fabric and local memory can accelerate structured nearest-neighbor computation outside neural networks. It does not translate directly into transformer training economics.&lt;/p&gt;

&lt;p&gt;For large-model training, Cerebras publishes cluster scale, parameter capacity, and time-to-train results. Those claims are useful architecture evidence when labeled as vendor measurements. Public, independently reproduced comparisons with matched model, token count, convergence target, precision, complete cluster power, and software effort remain scarce.&lt;/p&gt;

&lt;p&gt;The honest evidence ladder is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Strongest:&lt;/strong&gt; externally tested live inference endpoints and peer-reviewed workload-specific papers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Useful but bounded:&lt;/strong&gt; customer or laboratory deployment announcements confirming that systems exist and run real work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vendor measured:&lt;/strong&gt; model speed, scaling, and throughput claims with a disclosed configuration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forward-looking:&lt;/strong&gt; CS-4 extrapolations, projected efficiency, and announced capacity not yet demonstrated in sustained operation.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No single rung proves that Cerebras is faster or cheaper for AI as a category.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the architecture gets right
&lt;/h2&gt;

&lt;p&gt;Cerebras attacks a real cost in distributed computing: moving data among small processors because manufacturing and packaging force the machine to be divided.&lt;/p&gt;

&lt;p&gt;The wafer-scale approach buys a larger low-latency domain. Distributed SRAM keeps the active working set from crossing an HBM interface for every operation. The compiler can treat hundreds of thousands of tiles as one placement surface. Weight streaming allows models larger than the wafer’s SRAM to use that compute without slicing every layer across device memories.&lt;/p&gt;

&lt;p&gt;The design also makes its remaining boundaries clearer. MemoryX owns weights. SwarmX owns training broadcast and reduction. Nexus owns rack power, cooling, and I/O. Direct Wafer Links or Ethernet own scale-out. An optional heterogeneous processor owns prefill.&lt;/p&gt;

&lt;p&gt;That is not the elimination of distributed systems. It is a different decomposition of one.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what
&lt;/h2&gt;

&lt;p&gt;Cerebras should not be evaluated as an oversized GPU. Its architecture changes which links are on silicon, which state remains local, which state streams from outside, and which decisions move into the compiler.&lt;/p&gt;

&lt;p&gt;The useful comparison is therefore not wafer versus chip. It is complete system versus complete system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How much model state can remain local?&lt;/li&gt;
&lt;li&gt;Which communication occurs on silicon, inside the rack, and across the network?&lt;/li&gt;
&lt;li&gt;What compilation and model changes are required?&lt;/li&gt;
&lt;li&gt;How does performance change under concurrency rather than one fast stream?&lt;/li&gt;
&lt;li&gt;What is the measured rack and facility power for completed work?&lt;/li&gt;
&lt;li&gt;How does the system recover when a wafer, memory service, network path, or prefill engine becomes unavailable?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cerebras has shown that a wafer can behave like a processor. The unresolved question is whether the layers above it—weight memory, heterogeneous inference, rack power, scale-out networking, and software compatibility—can become as repeatable and measurable as the wafer itself.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>cerebras</category>
      <category>semiconductors</category>
      <category>inference</category>
    </item>
    <item>
      <title>AMD Is Rebuilding the GPU Stack in the Open</title>
      <dc:creator>Amit</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:48:43 +0000</pubDate>
      <link>https://dev.to/amitrix/amd-is-rebuilding-the-gpu-stack-in-the-open-19df</link>
      <guid>https://dev.to/amitrix/amd-is-rebuilding-the-gpu-stack-in-the-open-19df</guid>
      <description>&lt;p&gt;The series' read is that AMD is the closest merchant alternative to NVIDIA because it is rebuilding more than a GPU. Instinct accelerators now sit inside eight-GPU platforms, liquid-cooled racks, Ethernet fabrics, open scale-up standards, and a software stack that reaches from PyTorch kernels to cluster validation.&lt;/p&gt;

&lt;p&gt;The architecture is deliberately less vertically owned. AMD supplies critical pieces, but ROCm, UALink, UALoE, Ultra Ethernet, OEM systems, merchant switches, optics, and facility integration cross organizational boundaries. That creates choice at each layer—and makes the operator responsible for proving that those layers work as one system.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Instinct evolved from a 32 GB accelerator in MI100 to 288 GB of HBM3E in MI355X, while the deployment unit expanded from a PCIe card to eight-GPU platforms and a planned 72-GPU Helios rack.&lt;/li&gt;
&lt;li&gt;AMD's current shipping scale-up path still centers on Infinity Fabric inside an eight-GPU platform; Helios shifts the rack-scale domain toward the open UALink ecosystem and AMD's Ethernet-based UALoE fabric.&lt;/li&gt;
&lt;li&gt;ROCm is open and increasingly production-oriented, but openness does not remove qualification work. AMD documents exact operating-system, driver, firmware, container, collective-library, and network requirements.&lt;/li&gt;
&lt;li&gt;AMD has credible evidence: production availability through major cloud and OEM channels, MLPerf submissions, and named deployments. That evidence still does not establish interchangeable performance across models, precisions, software versions, or rack designs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The accelerator became a memory system
&lt;/h2&gt;

&lt;p&gt;The clearest Instinct trend is not a single compute multiplier. It is the steady expansion of the memory domain surrounding the accelerator.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Generation&lt;/th&gt;
&lt;th&gt;Public status by Aug. 30, 2026&lt;/th&gt;
&lt;th&gt;HBM capacity&lt;/th&gt;
&lt;th&gt;Peak memory bandwidth&lt;/th&gt;
&lt;th&gt;Board or module power&lt;/th&gt;
&lt;th&gt;System boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.amd.com/content/dam/amd/en/documents/instinct-business-docs/white-papers/amd-cdna-white-paper.pdf" rel="noopener noreferrer"&gt;MI100&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Previous generation&lt;/td&gt;
&lt;td&gt;32 GB HBM2&lt;/td&gt;
&lt;td&gt;1.23 TB/s&lt;/td&gt;
&lt;td&gt;300 W&lt;/td&gt;
&lt;td&gt;PCIe accelerator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.amd.com/content/dam/amd/en/documents/instinct-business-docs/white-papers/amd-cdna2-white-paper.pdf" rel="noopener noreferrer"&gt;MI250X&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Previous generation&lt;/td&gt;
&lt;td&gt;128 GB HBM2e&lt;/td&gt;
&lt;td&gt;3.2 TB/s&lt;/td&gt;
&lt;td&gt;560 W peak&lt;/td&gt;
&lt;td&gt;Dual-die accelerator used in multi-GPU systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/data-sheets/amd-instinct-mi300x-data-sheet.pdf" rel="noopener noreferrer"&gt;MI300X&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Shipping&lt;/td&gt;
&lt;td&gt;192 GB HBM3&lt;/td&gt;
&lt;td&gt;5.3 TB/s&lt;/td&gt;
&lt;td&gt;750 W&lt;/td&gt;
&lt;td&gt;Eight-GPU platform with 1.5 TB aggregate HBM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi325x.html" rel="noopener noreferrer"&gt;MI325X&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Shipping&lt;/td&gt;
&lt;td&gt;256 GB HBM3E&lt;/td&gt;
&lt;td&gt;6.0 TB/s&lt;/td&gt;
&lt;td&gt;1,000 W&lt;/td&gt;
&lt;td&gt;Eight-GPU platform with 2 TB aggregate HBM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html" rel="noopener noreferrer"&gt;MI355X&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Shipping&lt;/td&gt;
&lt;td&gt;288 GB HBM3E&lt;/td&gt;
&lt;td&gt;8.0 TB/s&lt;/td&gt;
&lt;td&gt;1,400 W&lt;/td&gt;
&lt;td&gt;Liquid-cooled eight-GPU platform with 2.3 TB aggregate HBM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.amd.com/en/products/rackscale-solutions/helios.html" rel="noopener noreferrer"&gt;MI455X in Helios&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Launched; shipments planned for the second half of 2026&lt;/td&gt;
&lt;td&gt;432 GB HBM4 per GPU&lt;/td&gt;
&lt;td&gt;23.3 TB/s per GPU&lt;/td&gt;
&lt;td&gt;Rack specification is the relevant boundary&lt;/td&gt;
&lt;td&gt;72-GPU rack with 31 TB aggregate HBM&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;MI200 made chiplets part of the accelerator design. MI300X pushed the idea further: eight compute dies sit over four I/O dies, connected through Infinity Fabric and surrounded by eight HBM stacks. AMD's &lt;a href="https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/white-papers/amd-cdna-3-white-paper.pdf" rel="noopener noreferrer"&gt;CDNA 3 architecture documentation&lt;/a&gt; describes the package as one logical accelerator even though compute, I/O, and memory are physically disaggregated.&lt;/p&gt;

&lt;p&gt;MI350 continues that direction with CDNA 4 and denser low-precision formats. The more durable change is 288 GB of HBM3E per accelerator and 2.3 TB across an eight-GPU platform. Large models can retain more weights and key-value cache close to compute, reducing how often capacity pressure forces work across slower boundaries.&lt;/p&gt;

&lt;p&gt;Helios changes the boundary again. AMD's current public design pairs 72 MI455X GPUs with EPYC host processors, Pensando networking, 31 TB of HBM4, and 1.7 PB/s of aggregate HBM bandwidth. Those are announced system specifications, not a claim about generally available production capacity or measured application performance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scale-up is moving from a platform to a rack
&lt;/h2&gt;

&lt;p&gt;MI300X, MI325X, and MI355X platforms connect eight GPUs through Infinity Fabric. At that scale, AMD controls the accelerator package and much of the GPU-to-GPU communication path. The OEM still owns the server implementation, firmware bundle, power delivery, cooling, and qualification.&lt;/p&gt;

&lt;p&gt;Helios expands the scale-up domain to 72 GPUs. Its architecture uses AMD's &lt;a href="https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html" rel="noopener noreferrer"&gt;UALoE&lt;/a&gt;—an Ethernet-based fabric using UALink-aligned semantics—rather than treating Ethernet only as the network between independent servers. AMD develops this path alongside the broader &lt;a href="https://ualinkconsortium.org/" rel="noopener noreferrer"&gt;UALink Consortium&lt;/a&gt; effort, where accelerator, switch, CPU, and systems companies define an open scale-up interconnect.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    app["Models and frameworks"]
    rocm["ROCm: HIP, kernels, RCCL, tools"]
    rack["72-GPU Helios rack"]
    gpu["AMD Instinct MI455X"]
    scaleup["UALink software model + UALoE fabric"]
    scaleout["Ultra Ethernet / RoCE scale-out"]
    network["Pensando NICs and merchant Ethernet switches"]
    oem["OEM / ODM rack integration"]
    facility["Operator: power, liquid cooling, firmware, validation"]

    app --&amp;gt; rocm --&amp;gt; rack
    rack --&amp;gt; gpu
    rack --&amp;gt; scaleup
    rack --&amp;gt; scaleout --&amp;gt; network
    gpu --&amp;gt; oem
    scaleup --&amp;gt; oem
    network --&amp;gt; oem
    oem --&amp;gt; facility&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The diagram exposes the trade. An open interface can let multiple switch, accelerator, cable, and system suppliers participate. It does not create one owner for congestion control, firmware compatibility, optics, cabling, collectives, telemetry, or failure recovery.&lt;/p&gt;

&lt;p&gt;Scale-out follows the same pattern. AMD is a founding member of the &lt;a href="https://ultraethernet.org/" rel="noopener noreferrer"&gt;Ultra Ethernet Consortium&lt;/a&gt;, while Pensando supplies NIC and DPU technology. The resulting cluster can use familiar Ethernet operations and merchant switching, but the final behavior depends on the complete path: GPU memory, PCIe attachment, NIC firmware, switch configuration, routing, loss management, RCCL, and the model's communication pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  ROCm makes portability possible, not automatic
&lt;/h2&gt;

&lt;p&gt;ROCm is no longer accurately described as a compiler shim around a GPU. It includes the HIP programming model, math and kernel libraries, RCCL collectives, profiling and debugging tools, container images, model recipes, validation utilities, and integrations with major frameworks.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Software layer&lt;/th&gt;
&lt;th&gt;AMD path&lt;/th&gt;
&lt;th&gt;What remains operator-owned&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Framework&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://rocm.docs.amd.com/projects/install-on-linux/en/latest/install/3rd-party/pytorch-install.html" rel="noopener noreferrer"&gt;PyTorch for ROCm&lt;/a&gt;, TensorFlow, JAX&lt;/td&gt;
&lt;td&gt;Pinning a supported framework, Python, driver, and ROCm combination&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Portability&lt;/td&gt;
&lt;td&gt;&lt;a href="https://rocm.docs.amd.com/projects/HIP/en/latest/how-to/hip_porting_guide.html" rel="noopener noreferrer"&gt;HIP porting model&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Reworking CUDA-specific extensions, intrinsics, build logic, and performance assumptions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kernels&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://github.com/ROCm/composable_kernel" rel="noopener noreferrer"&gt;Composable Kernel&lt;/a&gt;, hipBLASLt, MIOpen, rocBLAS&lt;/td&gt;
&lt;td&gt;Selecting or tuning kernels for the exact model, shape, precision, and GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Distributed execution&lt;/td&gt;
&lt;td&gt;&lt;a href="https://rocm.docs.amd.com/projects/rccl/en/latest/" rel="noopener noreferrer"&gt;RCCL&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Fabric configuration, topology, collective tuning, fault handling, and version alignment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://rocm.docs.amd.com/projects/ai-ecosystem/en/latest/inference/vllm.html" rel="noopener noreferrer"&gt;vLLM on ROCm&lt;/a&gt; and &lt;a href="https://rocm.docs.amd.com/projects/ai-ecosystem/en/latest/inference/sglang.html" rel="noopener noreferrer"&gt;SGLang on ROCm&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Testing model features, quantization paths, custom operators, latency, and capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cluster acceptance&lt;/td&gt;
&lt;td&gt;&lt;a href="https://rocm.docs.amd.com/en/docs-7.2.3/how-to/rocm-for-ai/system-setup/system-health-check.html" rel="noopener noreferrer"&gt;ROCm validation and system-health tooling&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Establishing a known-good baseline and revalidating it after firmware, driver, network, or container changes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;HIP can translate a large amount of CUDA-shaped source code and preserve a familiar kernel programming model. It cannot guarantee that a CUDA extension compiles unchanged, that the same fused kernel exists, or that identical scheduling choices remain optimal. Portability is strongest at the framework and model-contract layers; it weakens around custom kernels, collectives, memory placement, quantization, and compilation.&lt;/p&gt;

&lt;p&gt;AMD's own compatibility matrices make this responsibility explicit. Supported configurations pair specific GPU architectures with operating systems, kernel drivers, firmware, libraries, and framework releases. The &lt;a href="https://instinct.docs.amd.com/projects/gpu-cluster-networking/en/latest/" rel="noopener noreferrer"&gt;Instinct cluster deployment guides&lt;/a&gt; add validated NIC, RoCE, topology, and collective requirements. A container freezes user-space dependencies, but it does not freeze host drivers, GPU firmware, NIC firmware, switch behavior, or cooling performance.&lt;/p&gt;

&lt;p&gt;That is the operational difference between open source and an integrated product boundary. Source access improves inspection, contribution, and supplier choice. It does not eliminate systems engineering.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deployment evidence is real but uneven
&lt;/h2&gt;

&lt;p&gt;AMD has moved beyond reference designs. MI300X and later Instinct platforms are available through major server vendors and cloud providers. Microsoft's &lt;a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/gpu-accelerated/ndmi300xv5-series" rel="noopener noreferrer"&gt;Azure ND MI300X v5 virtual machines&lt;/a&gt; expose eight MI300X GPUs with 1.5 TB of aggregate HBM. Oracle has published &lt;a href="https://blogs.oracle.com/cloud-infrastructure/post/oci-amd-instinct-mi355x-mlperf-training-v60" rel="noopener noreferrer"&gt;MI355X cluster results and deployment details&lt;/a&gt;. Microsoft has separately announced that &lt;a href="https://newsroom.amd.com/news/microsoft-azure-ai-infrastructure/" rel="noopener noreferrer"&gt;Azure will deploy Helios for inference&lt;/a&gt;, with AMD stating that shipments begin in the second half of 2026.&lt;/p&gt;

&lt;p&gt;The evidence ladder still matters:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evidence level&lt;/th&gt;
&lt;th&gt;What AMD has&lt;/th&gt;
&lt;th&gt;What it establishes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Product specification&lt;/td&gt;
&lt;td&gt;Datasheets for Instinct accelerators and platforms&lt;/td&gt;
&lt;td&gt;Capacity, interfaces, rated bandwidth, and thermal boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vendor benchmark&lt;/td&gt;
&lt;td&gt;AMD model and platform results&lt;/td&gt;
&lt;td&gt;Performance under the vendor's disclosed configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard benchmark&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://mlcommons.org/benchmarks/training/" rel="noopener noreferrer"&gt;MLPerf Training results&lt;/a&gt; from AMD and partners&lt;/td&gt;
&lt;td&gt;Reproducible benchmark submissions within one MLPerf version and workload&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud or OEM availability&lt;/td&gt;
&lt;td&gt;Azure, Oracle, Dell, HPE, Lenovo, Supermicro and others&lt;/td&gt;
&lt;td&gt;A purchasable or consumable deployment path&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Production disclosure&lt;/td&gt;
&lt;td&gt;Named operator announcements and engineering reports&lt;/td&gt;
&lt;td&gt;Workload use at a stated boundary, when configuration detail is sufficient&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of these rows alone proves lower service cost or better application performance. A useful comparison must align model, precision, sequence length, batch policy, software version, system count, networking, power boundary, and availability target. A result from eight MI300X GPUs cannot be casually compared with a rack-scale result, and an announced 72-GPU design is not equivalent to a measured production fleet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The open stack changes the buying decision
&lt;/h2&gt;

&lt;p&gt;AMD's advantage is not that every layer is interchangeable. It is that more layers can be selected, inspected, and sourced independently: accelerator, host CPU, NIC, switch, rack integrator, framework build, inference engine, and cloud or on-premises operator.&lt;/p&gt;

&lt;p&gt;That changes the evaluation from a GPU comparison into an ownership map:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Which party certifies the complete server and rack?&lt;/li&gt;
&lt;li&gt;Who owns failures that cross GPU, NIC, switch, firmware, and collective-library boundaries?&lt;/li&gt;
&lt;li&gt;Which model features depend on backend-specific kernels?&lt;/li&gt;
&lt;li&gt;Can the operator reproduce benchmark results after changing containers or firmware?&lt;/li&gt;
&lt;li&gt;Does supplier choice reduce cost enough to fund the added integration and qualification work?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The thesis is therefore narrower than "open beats closed." AMD is building the closest merchant alternative because it now spans credible accelerators, large HBM domains, scale-up and scale-out networking, rack designs, and a usable software ecosystem. Its structure also places more responsibility at the interfaces between those layers.&lt;/p&gt;

&lt;p&gt;The unresolved question is whether the UALink and UALoE ecosystem can make a multi-company rack behave like one supported product without converging on a single de facto integrator. If operators still need one company to certify every firmware, fabric, cooling, and software combination, how much of the architectural openness will survive contact with production?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of the &lt;a href="https://dev.to/series/ai-compute/"&gt;AI Compute Landscape&lt;/a&gt; — an ongoing exploration of accelerator architectures, software stacks, and data-center systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aicompute</category>
      <category>amd</category>
      <category>gpu</category>
      <category>infrastructure</category>
    </item>
  </channel>
</rss>
