The NVIDIA Empire: How One Chip Company Became the Backbone of the AI Age | NeuralWired
Deep DiveNeuralWired · May 2026 · 14 min read
NVIDIA Built the Machine That Runs the AI Age, And Nobody Saw It Coming
From a scrappy Santa Clara startup fighting pixel wars in 1993, NVIDIA has become the most strategically indispensable company in modern technology. Here is every secret, every bet, every decision that turned a graphics chip maker into the architect of the world’s artificial intelligence infrastructure.
The Origin Story Nobody Tells Correctly
NVIDIA didn’t set out to rule artificial intelligence. It set out to make video games look better. Jensen Huang, Chris Malachowsky, and Curtis Priem founded the company in 1993 with a single obsession: real-time graphics acceleration for the personal computer. The industry barely noticed. Competition came from everywhere, 3dfx, ATI, and the ever-present shadow of Intel, and NVIDIA spent its early years in genuine financial peril, one bad product cycle from extinction.
What saved them wasn’t luck. It was a culture of making bets most executives wouldn’t dare write in a boardroom presentation. Huang, an engineer who’d come up through AMD and LSI Logic, had an instinct for long-horizon thinking that bordered on irrational to anyone watching quarterly earnings. The company nearly went under multiple times before its first major hit. That formative near-death experience, embedded into NVIDIA’s DNA, explains almost everything that came after.
Company Snapshot: Founded 1993, Santa Clara, California. Founders: Jensen Huang, Chris Malachowsky, Curtis Priem. Employees: 30,000+. Market cap as of 2026: approximately $2.8 to $3.0 trillion. Core segments: Data Center & AI, Gaming, Professional Visualization, Automotive & Robotics.
The Moment NVIDIA Invented the GPU, and Changed Everything
1999 is the inflection point. NVIDIA released the GeForce 256 and, simultaneously, coined the term “GPU”, Graphics Processing Unit. This wasn’t marketing. It was a genuine architectural claim: here was a processor purpose-built for the massively parallel math that real-time rendering demands. Central processors handled tasks sequentially. GPUs handled thousands of calculations at once. The difference, as it turned out, would matter enormously beyond gaming.
The GeForce architecture gave NVIDIA a product that sold in volume and funded everything else. Gaming revenues became the war chest Huang needed to take bigger, stranger bets. And the biggest, strangest bet was still seven years away.
“The GPU is a massively parallel processor. It turns out that the computation of intelligence is a lot like the computation of graphics.”
Jensen Huang, CEO, NVIDIA, GTC 2024 Keynote
That insight, that graphics math and AI math are structurally identical, wasn’t obvious to anyone in 1999. It took another decade of basic research before the academic community would confirm it. NVIDIA got there first not because it predicted deep learning, but because it built the hardware that made deep learning possible by accident, and then moved aggressively to own that accident.
CUDA: The Secret Weapon That Competitors Still Can’t Copy
In 2006, NVIDIA launched CUDA, Compute Unified Device Architecture. The idea was simple and audacious: let developers program GPUs directly for general-purpose computing, not just graphics. Write code in a familiar C-like language, run it on massively parallel GPU hardware, and suddenly the chip inside a gaming PC becomes a scientific supercomputer.
Nobody wanted it at first. The early adopters were a handful of academic researchers running physics simulations and protein-folding experiments. NVIDIA subsidized developer adoption, gave away toolkits, built documentation, ran workshops at universities. For years, CUDA generated no meaningful revenue. It was an investment in a future that wasn’t guaranteed.
The CUDA Moat Explained: CUDA isn’t just software, it’s 20 years of accumulated developer workflows, pre-built libraries (cuDNN, cuBLAS, TensorRT), and a community of millions of engineers who learned AI on NVIDIA hardware. AMD and Intel have competing frameworks (ROCm, oneAPI), but they lack CUDA’s maturity, breadth, and ecosystem gravity. Switching costs are enormous. This is not a moat competitors can buy their way across.
Then 2012 happened. A team at the University of Toronto, led by Geoffrey Hinton, entered a deep learning model called AlexNet into the ImageNet Large Scale Visual Recognition Challenge. AlexNet was trained on two NVIDIA GTX 580 GPUs using CUDA. It didn’t just win, it demolished the competition by a margin so large the entire machine learning field snapped to attention. CUDA was suddenly not a curiosity. It was infrastructure.
NVIDIA had planted a flag in 2006 and spent six years waiting for the world to catch up. When it did, nobody else had a flag anywhere nearby.
What CUDA Actually Controls
The largest GPU developer ecosystem on the planet, with millions of active CUDA programmers
Pre-built AI libraries, cuDNN (deep neural networks), cuBLAS (linear algebra), TensorRT (inference optimization), that underpin every major AI framework
Native support baked into PyTorch, TensorFlow, JAX, and every significant AI research tool
20 years of optimized code that researchers, engineers, and enterprises depend on daily
Switching friction so high that even well-funded competitors struggle to peel away users
How NVIDIA Saw the AI Wave Before the AI Wave Existed
By 2017, NVIDIA’s data center revenue surpassed gaming revenue for the first time. Inside the company, this was confirmation of a thesis Huang had been running since the early CUDA days: the future of computing was parallel, and parallel computing was NVIDIA’s territory. He’d said it in interviews, said it in shareholder letters, said it to skeptical analysts. Most assumed it was boosterism.
It wasn’t. The 2020s AI explosion — ChatGPT, large language models, generative AI, inference at scale, required exactly the kind of hardware NVIDIA had spent two decades building. When OpenAI needed to train GPT-3, they turned to NVIDIA A100s. When Google, Microsoft, Amazon, and Meta began building out their own AI infrastructure, the bill of materials had NVIDIA at the top. Every serious AI model trained between 2020 and 2026 ran on NVIDIA hardware.
The Hopper architecture, introduced in 2022, was purpose-designed for transformer-based AI workloads. The H100 GPU became the most sought-after piece of silicon in history. Lead times stretched to 52 weeks. Cloud providers paid billions for allocation. Startups structured their entire fundraising strategies around securing H100 access. This was not a supply chain story. It was a story about irreplaceability.
“We are no longer a chip company. We are an AI infrastructure company. We sell AI factories.”
Jensen Huang, CEO, NVIDIA, Annual Investor Day 2025
Jensen Huang’s Execution Playbook: What Actually Makes This Work
Jensen Huang is one of the few trillion-dollar CEOs who still understands every layer of his own product. He writes code. He reads chip specs. He can speak in detail about interconnect bandwidth, memory hierarchy, and power delivery in the same breath as competitive strategy and developer ecosystems. That technical depth isn’t incidental to NVIDIA’s success. It’s structural to it.
Huang runs NVIDIA with a flat management philosophy that concentrates decision-making at the top and moves fast when it matters. He’s known for “betting the company” repeatedly. CUDA was a bet. The data center pivot was a bet. The automotive AI investment was a bet. None had guaranteed payoffs. All required sustaining investment through years when the returns weren’t visible.
The Culture He Built
Engineering culture above all, product decisions are made by people who understand the silicon
Kill weak products early and double down on winners — no sentimentality about legacy lines
Developer-first mindset, CUDA’s early free distribution was a deliberate market seeding strategy
Speed as a cultural value, rapid architecture cycles are not just technical achievements, they’re cultural ones
Long-horizon thinking, investments that won’t pay off for 5 to 10 years are normal operating procedure
The 2022 attempted acquisition of ARM is instructive even in failure. NVIDIA offered $40 billion for the chip architecture that runs nearly every mobile device on earth. Regulators blocked it after 18 months of scrutiny. Huang didn’t waver publicly. The lesson he took wasn’t “don’t attempt ambitious acquisitions”, it was “build what you can’t buy.” The Blackwell architecture and NVLink networking infrastructure that followed were direct responses to that lesson.
NVIDIA vs. Everyone Else: An Honest Scorecard
AMD makes competitive GPUs. Intel has poured billions into accelerators. Qualcomm owns automotive and mobile AI edge cases. Amazon, Google, and Microsoft build custom chips for their own clouds. Huawei serves the Chinese market with domestic alternatives. On paper, NVIDIA faces genuine competition from every direction. In practice, the competitive dynamic is less symmetric than it appears.
Company
Primary AI Chip Offering
CUDA Equivalent
Data Center Presence
Core Weakness vs. NVIDIA
AMD
Instinct MI300X
ROCm (maturing)
Growing
Ecosystem depth, CUDA lock-in
Intel
Gaudi 3
oneAPI
Limited
Software maturity, market share
Google
TPU v5 (internal)
XLA (TF-focused)
Google Cloud only
Not sold externally; framework-specific
Amazon
Trainium 2 / Inferentia
Neuron SDK
AWS only
Locked to one cloud; limited ecosystem
Huawei
Ascend 910B
CANN
China-focused
Export restrictions limit global reach
The table above shows the structural problem for every competitor: none has CUDA. ROCm, oneAPI, and the rest are catching up, but the gap is measured in decades of ecosystem maturity, not months of engineering. An enterprise that has spent five years building AI pipelines on CUDA libraries doesn’t switch platforms because a rival chip scored 10% better on a benchmark. The total cost of migration, retraining teams, rewriting code, re-validating models, is prohibitive.
The Architecture Arms Race NVIDIA Keeps Winning
NVIDIA’s hardware cadence is relentless. Pascal gave way to Volta, Volta to Turing, Turing to Ampere, Ampere to Hopper, Hopper to Blackwell. Each generation delivers meaningful performance leaps, not incremental tweaks, but wholesale redesigns tuned to the demands of whatever AI workload the market is building toward. By the time competitors have productized a response to Hopper, NVIDIA is already shipping Blackwell.
The 2025 Blackwell architecture represents a step-change in how NVIDIA thinks about scale. Rather than optimizing individual GPUs, Blackwell is designed around rack-scale systems. The GB200 NVL72 configuration packs 72 Blackwell GPUs into a single rack, connected by NVLink 5 with 1.8 terabytes per second of bandwidth between chips. This is not a GPU. This is a distributed compute fabric that happens to fit in a data center cabinet.
Why Rack-Scale Matters: Training frontier AI models now requires moving petabytes of data between thousands of chips simultaneously. The limiting factor isn’t raw compute, it’s the bandwidth between chips. NVLink collapses that bottleneck. Competitors selling individual GPUs are competing in a category NVIDIA is moving away from.
The Mellanox acquisition, completed in 2020 for $6.9 billion, was the move that made this possible. Mellanox owned InfiniBand, the high-speed networking fabric used in supercomputers worldwide. Owning the networking layer meant NVIDIA could co-design chips and interconnects together, something no GPU competitor can do. AMD sells GPUs. Intel sells accelerators. NVIDIA sells the entire compute stack, from silicon to software to network.
The Financial Engine Behind the Empire
NVIDIA’s revenue mix has inverted entirely since the early 2010s. Data center now drives the largest share of income by a wide margin, with gaming remaining significant but no longer defining. Professional visualization, automotive, and licensing round out the portfolio. The growth trajectory is steep enough that financial analysts have struggled to model it accurately, NVIDIA consistently beats consensus estimates by margins that suggest the AI infrastructure buildout is larger and faster than any outside observer predicted.
🏭
Data Center
Largest revenue segment. Driven by AI training, inference, and hyperscaler GPU purchases. Growth has been explosive since 2022.
🎮
Gaming
Still a major business. GeForce RTX cards dominate the discrete GPU market. AI-enhanced features like DLSS add new value.
🚗
Automotive
DRIVE platform powers autonomous vehicle development. Long-horizon bet with multi-year design cycles and growing pipeline.
🔬
Pro Visualization
Quadro/RTX workstation GPUs for designers, engineers, and digital artists. Steady, high-margin business.
The global AI infrastructure buildout projected through 2030 sits at $3 to $4 trillion across cloud providers, enterprises, and governments. NVIDIA doesn’t capture all of it, but it captures the part every other participant depends on. Even the hyperscalers building custom chips still buy NVIDIA GPUs for workloads where CUDA’s ecosystem is irreplaceable. That’s the tell. When your competitors are also your customers, your competitive position is not merely strong. It’s structural.
The Real Risks: What Could Actually Hurt NVIDIA
NVIDIA faces challenges that can’t be dismissed. China export restrictions, tightened progressively since 2022, have cut off a significant portion of a market that once represented meaningful revenue. The company has released export-compliant variants of its chips (A800, H800, H20) but these occupy a different performance tier, and the regulatory environment remains unpredictable. Any further tightening hits the top line directly.
Supply chain constraints are real and persistent. TSMC manufactures NVIDIA’s most advanced chips on leading-edge process nodes. That dependency on a single foundry, in a geopolitically sensitive geography, creates concentration risk that no amount of procurement strategy can fully eliminate. When demand surged in 2023 and 2024, NVIDIA could not produce H100s fast enough. Revenue was limited by manufacturing, not by demand.
China export restrictions have cut NVIDIA off from one of the world’s fastest-growing AI markets
TSMC dependency creates geopolitical supply risk that is structural, not easily hedged
Rising competition from AMD’s MI300X, particularly for inference workloads, is closing the gap in specific use cases
Custom silicon from Google (TPU), Amazon (Trainium), and Microsoft (Maia) reduces these hyperscalers’ dependency on external GPU suppliers over time
Regulatory scrutiny is intensifying globally, NVIDIA’s market position is large enough to attract antitrust attention
Energy consumption of AI data centers faces political and environmental pushback that could reshape demand curves
The custom chip threat from hyperscalers deserves particular attention. Google’s TPUs have been in production for over a decade and continue to improve. Amazon’s Trainium 2 is targeting training workloads at scale. Microsoft’s Maia chip is in deployment. These chips are purpose-built for specific workloads and don’t need to match NVIDIA’s general-purpose performance, they need only to be good enough for their owner’s most common tasks, at a lower cost per compute unit. Over a long enough horizon, this erodes NVIDIA’s share of hyperscaler spend, even if it doesn’t displace NVIDIA entirely.
Where NVIDIA Goes Next: The 2026 and Beyond Strategy
NVIDIA’s stated future is not a product roadmap. It’s a platform vision. Huang has positioned the company as the architect of “AI factories”, full-stack systems that enterprises and governments buy the way they once bought data centers, complete with GPUs, networking, software, and management infrastructure. The GB200 NVL72 rack is the current physical embodiment of this vision. Future iterations will scale further.
Robotics is the next major frontier. NVIDIA’s Isaac robotics platform and its Omniverse simulation environment give it tools to train physical AI systems, robots that operate in the real world rather than in data centers. The automotive DRIVE platform feeds into this strategy: every autonomous vehicle is, from NVIDIA’s perspective, a mobile robot. The data it generates, the simulation environments needed to train it, and the compute required to run inference all flow through NVIDIA’s stack.
Edge AI is the third vector. As AI models get smaller and more efficient, inference moves toward devices, industrial sensors, medical equipment, consumer electronics, network infrastructure. NVIDIA’s Jetson platform competes in this space. It’s a smaller market today, but the installed base of AI-capable edge devices is expected to exceed the installed base of data center nodes by a wide margin within this decade.
Five Things to Watch
01Blackwell successor architecture, when NVIDIA announces the next generation, watch the NVLink bandwidth and memory specs for signals about model-scale ambitions.
02China policy, any easing or further tightening of US export controls directly affects NVIDIA’s addressable market by tens of billions of dollars.
03Hyperscaler custom chip adoption rates, if Google or Amazon meaningfully reduces external GPU purchases, that signals the beginning of a structural share shift.
04AMD ROCm ecosystem maturity, if ROCm closes the gap on CUDA for mainstream PyTorch workflows, the switching barrier drops significantly.
05NVIDIA software revenue, as the company expands NIM microservices and AI Enterprise licensing, watch the software revenue line as a percentage of total revenue.
NVIDIA’s Real Secret: The Moat Is Time, Not Technology
Strip away the marketing and the narrative, and NVIDIA’s competitive position comes down to a single uncomfortable truth for its rivals: the company got there first and invested in the right things for twenty years before those things were worth investing in. CUDA launched in 2006. AlexNet vindicated it in 2012. The H100 dominated in 2023. That’s a 17-year arc from investment to dominance.
Jensen Huang didn’t predict the AI boom with precision. Nobody did. What he did was build an architecture, hardware, software, ecosystem, culture, that was positioned to win regardless of which specific AI application took off first. Deep learning? CUDA was ready. Large language models? Hopper was designed for transformers. Inference at edge? Jetson was already in production. The strategy wasn’t prediction. It was preparation.
NVIDIA’s story is fundamentally about the compounding value of technical bets made early and sustained through years of uncertain returns. Its competitors face the task of not just building better chips, but building richer ecosystems, deeper developer communities, and more complete full-stack offerings, all while NVIDIA continues advancing at the same pace. The lead is large. The moat is real. And the company that started by making video games look pretty now runs the machines that are reshaping civilization.
Frequently Asked Questions
Why does NVIDIA dominate AI chips so completely?
Three compounding advantages: the H100 and Blackwell GPUs deliver leading compute performance for AI workloads; CUDA is the developer ecosystem every major AI framework is built on; and NVIDIA sells full-stack systems, GPUs, networking, software, and management tools together. No competitor matches all three simultaneously.
What is CUDA and why can’t competitors replicate it?
CUDA is NVIDIA’s GPU programming platform, launched in 2006. It includes a programming model, compiler, libraries (cuDNN, cuBLAS, TensorRT), and a developer ecosystem built over 20 years. Competing platforms like AMD’s ROCm exist but lack the library depth, documentation maturity, and universal framework support CUDA has accumulated. Switching costs for enterprises are enormous.
How does NVIDIA make money?
Primary revenue comes from data center GPU sales to hyperscalers, cloud providers, and enterprises. Gaming GPUs remain a large secondary business. Professional visualization, automotive (DRIVE platform), and a growing software licensing business round out the portfolio. Data center now dominates the revenue mix by a significant margin.
What is the Blackwell architecture?
Blackwell is NVIDIA’s 2025 GPU architecture, designed for rack-scale AI systems. The GB200 NVL72 configuration packs 72 Blackwell GPUs into a single rack with NVLink 5 interconnect running at 1.8 TB/s between chips. It’s designed for training and inference of frontier AI models at scales that previous GPU generations couldn’t support efficiently.
What are the biggest risks facing NVIDIA?
US export restrictions limiting sales to China represent the most immediate revenue risk. TSMC manufacturing dependency creates geopolitical supply risk. Long-term, hyperscaler custom chips (Google TPU, Amazon Trainium, Microsoft Maia) could reduce external GPU demand. AMD’s ROCm ecosystem improving is a slower-moving but real competitive threat.
Will NVIDIA remain the AI chip leader?
The CUDA ecosystem and full-stack integration give NVIDIA structural advantages that are difficult to displace quickly. However, at a $3 trillion market cap, the company already prices in continued dominance. The scenarios where NVIDIA loses meaningful share, rapid ROCm adoption, aggressive hyperscaler insourcing, geopolitical disruption, are low-probability but not zero. Sustained leadership is likely; guaranteed leadership is not.
Stay ahead of the AI hardware race.
NeuralWired covers chip architecture, AI infrastructure, and the companies building the computational future.
Why 80% of AI Pilots Fail in 2026: The 7-Step CTO Playbook That Actually Scales | NeuralWired
AI Strategy
Most AI projects collapse between pilot and production. Here is the data-backed strategy for CTOs who need to move from experiments to enterprise-grade ROI, before competitors close the gap.
NeuralWired EditorialMarch 2026
Eighty percent of AI pilots launched in 2025 will not scale. Not because the models were wrong. Not because the vendors overpromised. But because CTOs built the roof before the foundation.
That is the hard finding emerging from enterprise analysis heading into 2026. While boards push for AI returns and engineering teams prototype agents at record pace, most organizations are hitting the same wall: demos do not equal deployments, and pilots do not equal platforms.
The CTOs winning this race are not the ones who moved fastest. They are the ones who moved correctly. They audited maturity, built governance infrastructure, matched risk to capability, and measured outcomes against real benchmarks. This article delivers that exact framework: a 7-step AI strategy for CTOs built from current research, practitioner data, and competitive analysis of what separates the 20% who scale from the 80% who stall.
80%of AI pilots fail to reach production scale
50%cost reduction achievable through proper AI governance
30%of enterprises will automate over half of network activities by 2026
2025 Was the Year of the Pilot. 2026 Is the Year of the Foundation.
Last year’s AI investments were largely exploratory. Teams tested tools, ran proofs of concept, and shipped demos to stakeholders. That phase is closing fast.
“2025 was the year of the AI pilot,” wrote tech leader Kaustav Mohanta in a December 2025 analysis. “2026 is the year of the AI foundation.” The distinction matters enormously. Foundations require different investments, different governance structures, and different success criteria than pilots do.
The board-level pressure is intensifying. As analysts at CXO India noted in February 2026, “CTOs must balance innovation with pragmatism, as boards demand ROI from AI investments.” That balance, between speed and sustainability, is exactly where most AI strategies currently break.
Post-mortem analysis of failed AI rollouts consistently surfaces three root causes. Understanding them is the prerequisite for everything that follows.
Gap 1: Data readiness is assumed, not verified. Teams launch agents against unstructured, poorly governed data and wonder why outputs are unreliable. The model is rarely the problem. The data pipeline almost always is.
Gap 2: Governance is bolted on after deployment, or skipped entirely. Roughly 70% of CTOs ignore governance during the pilot phase, according to CTO interview data compiled by Accedia’s AI strategy blueprint. That omission becomes catastrophic at scale when compliance, security, and audit requirements arrive.
Gap 3: Infrastructure does not match ambition. There is a significant difference between infrastructure that supports 5 pilots and infrastructure that supports 50 production use cases. Most organizations optimize for the former, then wonder why scaling fails.
“Match risk to capability. Your CRUD endpoints can be at level 7 while payment processing stays at level 3.”
Schmidt’s point is counterintuitive but critical. The right AI strategy is not uniform across an organization. Different systems warrant different levels of AI integration based on risk tolerance, regulatory exposure, and the cost of errors. Treating everything as equally ready for automation is how organizations create catastrophic failure points.
Before deploying anything new, assess honestly where your organization sits. Use AmazingCTO’s 9-level adoption model as a diagnostic. Level 3 (daily AI use across engineering teams) is the first meaningful milestone. Many organizations claiming AI adoption have not reached it. Crucially, identify your level per system, not per organization. Payment processing and internal tooling do not share a risk profile.
2
Build the Data and AI Factory First
Structured pipelines, clean data governance, and observable model behavior are not features. They are prerequisites. Infrastructure that handles 5 pilots will fail at 50 production use cases. This is where most CTOs underinvest, and where scaling failures originate. Budget 20 to 30% of tech spend on this layer before any agent deployment.
3
Prioritize Use Cases by Risk Profile
Not all automation candidates are equal. Map each use case against business value and risk-to-error. High-value, low-risk systems should be accelerated to higher AI integration levels. High-stakes systems (payments, compliance, patient data) should progress more deliberately. Mixing these risk profiles into one deployment timeline is a governance failure waiting to happen.
4
Integrate With Cloud and Security Stacks From Day One
AI deployments that ignore existing cloud and security architecture create technical debt that compounds fast. Zero-trust principles, API gateway management, and identity-aware access controls should be applied to AI workloads from the first production deployment, not retrofitted post-incident. This integration also unlocks the 30% supply chain downtime reductions that mature agentic AI deployments are delivering right now.
5
Define Pilot-to-Scale Criteria Before You Pilot
Most pilots fail not in the pilot phase but in the transition. Set explicit success criteria before launch: daily active usage rates, latency benchmarks, error thresholds, and business impact metrics. If a pilot cannot articulate how it becomes production in 90 days, do not start it. The near-term milestone to target: consistent daily AI use across the relevant team, which is Level 3 in AmazingCTO’s framework.
6
Establish an AI Governance Council
Genpact’s client data shows that proper governance cuts AI project costs by 50% while accelerating time-to-value. The council should own decision rights for model deployment, data usage policies, vendor selection, and incident response. Track these KPIs: time-to-value per use case, model performance drift rates, and compliance audit pass rates. Without this structure, every AI deployment becomes an ad hoc negotiation.
7
Measure ROI With the Right Denominator
Success metrics should include automation percentage (target: 30% or more of eligible operations), cost reduction per use case, and time saved per workflow. But measure ROI against total cost of ownership, which includes governance infrastructure, talent upskilling, and ongoing model maintenance. Organizations reporting 2x or 3x returns are measuring this correctly. Skeptics often are not counting hidden costs, or hidden benefits.
Build vs. Buy: The Decision CTOs Most Often Get Wrong
One of the most expensive AI strategy mistakes is applying a uniform build-or-buy policy across an entire technology stack. The financial implications are significant, and the right answer varies by use case.
Factor
Custom AI Build
Off-the-Shelf (COTS)
ROI in Edge Cases
Up to 2x higher
Median performance
Time to Deploy
2x longer to build
Fast initial deployment
Vendor Lock-in Risk
Low
High
Domain Specificity
High, tuned to your data
Generalist, may miss nuance
Best For
Core differentiating workflows
Commodity tasks, rapid prototyping
Industry analysis from Kaustav Mohanta suggests custom AI delivers up to 2x ROI over off-the-shelf in edge cases, but takes twice as long to build. The answer is not one or the other. Build custom AI where differentiation matters (core product logic, proprietary data workflows). Buy commodity AI everywhere else. Organizations that try to build everything burn capital. Those that buy everything give up their competitive moat.
As the Kanerika guide for CTOs and CIOs frames it: build what creates sustainable competitive advantage, and buy what speeds up everything else. Apply that filter to every AI investment decision in 2026.
Pre-Deployment Readiness: The Integration Checklist
Before any AI system goes into production, the following should be verified, not assumed. This checklist covers the integration gaps that most commonly kill AI deployments between pilot approval and go-live.
AI Production Readiness
Data governance framework documented and approved by legal and compliance
Zero-trust access controls applied to all AI-adjacent APIs
Model observability tools integrated (logging, alerting, drift detection)
Rollback protocol defined and tested before go-live
Pilot-to-scale success criteria written and agreed upon before launch
AI governance council notified and in the decision loop
18-month total cost of ownership modeled, including talent and maintenance
Security incident response plan updated for AI-specific scenarios
“Organizations that master these elements don’t just launch pilots. They build a repeatable engine for growth.”
Understanding where AI infrastructure is headed helps CTOs make investments today that will not require costly rewrites in 18 months. Current trend analysis points to three distinct phases ahead.
26
2026: Infrastructure and Foundation Year
The year of governance councils, data factories, and scaling pilots to production. Gartner ranks AI-native platforms as a top 2026 technology trend. Organizations that build this foundation correctly will have a durable competitive advantage through the rest of the decade.
27
2027: Agentic AI Moves from Hype to Deployment
Multi-agent systems that coordinate autonomously across workflows are in Gartner’s hype cycle now. By 2027, organizations that built clean infrastructure in 2026 will deploy agents that genuinely handle complex, multi-step operations. Those that did not will be playing catch-up.
28
2028: Mature Agentic Operations at Scale
The full vision of AI-augmented engineering and operations becomes operational reality for prepared organizations. Barriers between now and then: data quality, talent availability, and governance discipline. All of which get built in 2026.
The CTO Strategy OS 2026 deck, designed for board-level communication, projects 20 to 30% of annual tech spend shifting to AI infrastructure over this period. CTOs who can frame that investment in ROI language, not just engineering metrics, will secure the budgets to execute this roadmap.
Frequently Asked Questions
What should a CTO prioritize in AI for 2026?
Infrastructure and governance over features. Before expanding AI capabilities, CTOs should audit their organization’s current adoption maturity, targeting at least Level 3 daily use, establish data pipelines that can support 50 or more production use cases rather than 5 pilots, and create AI governance councils with clear decision rights. Gartner’s 2026 trends place AI-native platforms at the top of the priority list, which means foundational investment before new capability development.
How do you measure AI ROI for enterprises?
Track time-to-value per use case, automation percentage targeting 30% or more of eligible workflows, and cost reduction against a total cost of ownership baseline that includes governance, talent, and maintenance. Agentic AI systems in supply chain contexts are delivering 30% reductions in downtime. Use sector benchmarks like these as calibration points for your own expectations.
What are AI governance best practices in 2026?
Establish a cross-functional AI council with documented decision rights over deployment, data access, vendor selection, and incident response. Define KPIs including time-to-value, drift rates, and compliance pass rates before deploying any system. Genpact’s client data shows organizations with proper governance cut AI project costs by 50% compared to those that govern reactively.
What are the biggest AI integration challenges for legacy systems?
Three challenges dominate: unstructured or poorly governed data that degrades model outputs, security architectures not designed for API-heavy AI workloads, and organizational resistance to changing long-established workflows. The tactical approach: start with API wrappers around legacy systems to isolate them from AI agents, apply zero-trust controls from day one, and sequence deployments by risk profile, beginning with low-risk, high-value operations first.
What are the top AI risks CTOs should plan for?
The pilot-to-scale gap is the most immediate risk. Roughly 80% of pilots fail to reach production, primarily due to data and governance deficits identified too late. Beyond that: hype-driven investment that outpaces infrastructure readiness, vendor lock-in from premature COTS adoption, and talent shortages in AI infrastructure and governance roles. Mitigate through maturity audits before new initiatives, explicit build-vs-buy criteria, and upskilling plans that run parallel to deployments.
Should CTOs build custom AI or buy off-the-shelf solutions?
Both, applied selectively. Build custom AI for core differentiating workflows where proprietary data creates competitive advantage. Custom solutions can deliver up to 2x ROI over off-the-shelf in these use cases, though they take longer to build. Buy commodity AI for standardized tasks where speed matters more than differentiation. Apply this filter per use case, not as an organization-wide policy.
What does a CTO AI adoption roadmap look like in practice?
AmazingCTO’s 9-level adoption framework provides the most actionable map available: from basic tooling replacement at Level 1 to AI-only engineering at Level 9. The near-term goal for most organizations is Level 3, which is consistent daily AI use across engineering teams. From there, the playbook sequences risk-matched use cases, builds governance infrastructure, and scales toward agentic operations by 2027 and 2028.
The Bottom Line
The pattern across failed AI deployments is consistent. Organizations that skip foundations, including data governance, observability, and risk-matched deployment sequencing, do not scale. The 7-step AI strategy for CTOs outlined here is not a shortcut. It is the actual path. And it is considerably shorter than the detour most organizations take through pilot purgatory.
What is at stake extends beyond this year’s budget cycle. As agentic AI matures from hype to infrastructure between 2026 and 2028, the gap between organizations that built proper foundations and those that did not will widen. The competitive advantage in AI is shifting from access to technology, which commoditizes rapidly, to organizational readiness. That readiness gets built in 2026.
Three things to watch: vendor consolidation around AI governance platforms, regulatory requirements for model observability, and an accelerating talent shortage in AI infrastructure roles. CTOs who start building toward all three now will find themselves in the 20% that scales, not the 80% that stalls.
Nvidia NemoClaw: The Open-Source AI Agent Play That Could Reshape Enterprise — NeuralWired
AI AgentsEnterpriseNeuralWired Staff · March 13, 2026 · 6 min read
Days before GTC 2026, Nvidia has quietly pitched a new open-source AI agent platform to Salesforce, Google, Cisco, Adobe, and CrowdStrike. Here’s why it matters far beyond the chip wars.
Jensen Huang once called OpenClaw “the single most important release of software probably ever.” Now Nvidia is building its answer. And it wants Salesforce, Google, Cisco, Adobe, and CrowdStrike along for the ride.
According to reports first published by WIRED on March 9, 2026, Nvidia is developing NemoClaw: an open-source platform for deploying AI agents across enterprise workflows. Pre-announcement pitches from Huang’s team are already underway. The formal unveiling is expected at Nvidia’s GTC 2026 keynote on March 16 in San Jose.
This isn’t just another AI announcement. It’s Nvidia making its most explicit move yet into enterprise software, territory historically owned by Microsoft, Salesforce, and ServiceNow. For CTOs deciding their agentic infrastructure strategy, founders building on top of emerging platforms, and investors watching Nvidia’s margin story evolve, NemoClaw deserves close attention now, before the hype cycle distorts the signal.
This analysis covers what NemoClaw is, why Nvidia is building it, how it compares to OpenClaw and proprietary alternatives, what the genuine security risks are, and what decisions enterprise leaders should be making right now.
What NemoClaw Actually Is (And Where It Comes From)
NemoClaw is best understood as an extension of Nvidia’s existing NeMo platform, which already handles the AI model lifecycle: data curation, fine-tuning, reinforcement learning, and deployment via microservices. NeMo gave enterprises the infrastructure to build and run models. NemoClaw adds the orchestration layer: coordinating AI agents that can autonomously complete multi-step workforce tasks.
The key architectural details confirmed so far:
Open source: Unlike most enterprise AI agent frameworks, NemoClaw will be publicly available, inviting community contributions and third-party integrations.
Hardware-agnostic: A deliberate departure from Nvidia’s CUDA lock-in philosophy. NemoClaw is designed to run on any hardware, a significant strategic concession meant to accelerate enterprise adoption.
Built-in security and privacy layers: The platform includes native security controls, directly addressing what cybersecurity experts describe as OpenClaw’s “lethal trifecta”: private data access, external communications, and potential for harmful content generation.
Local execution: Agents can run on-premises or in hybrid configurations, meeting enterprise data sovereignty requirements that cloud-only solutions can’t satisfy.
The name itself signals lineage. “Nemo” from the NeMo suite; “Claw” borrowed from the agentic framing popularized by OpenClaw. Nvidia is positioning this as both a technical successor and a market response.
Why Nvidia Is Moving Into Software, Explained Honestly
The obvious question: why does a chip company need an agent platform?
The honest answer is that Nvidia doesn’t need one for revenue. It needs one for survival.
“The single most important release of software probably ever.”
Jensen Huang, CEO, Nvidia — on OpenClaw, the framework NemoClaw now aims to rival
Huang’s effusive praise for a competitor’s software wasn’t mere politeness. It was a recognition that agentic frameworks are becoming the new platform layer in enterprise AI. Whoever controls the orchestration layer controls the deployment roadmap, the security model, the integration patterns, and ultimately the hardware purchasing decisions that follow.
Three specific pressures are driving this:
1. Chip competition is intensifying. AMD, Intel, and a wave of custom silicon startups (Google’s TPUs, Amazon’s Trainium, Meta’s MTIA) are narrowing Nvidia’s GPU performance gap. Nvidia can’t defend $130B+ in annual revenue on silicon alone indefinitely.
2. Software creates lock-in that hardware can’t. Once enterprises build workflows on NemoClaw’s agent orchestration model, switching costs multiply. That’s the Microsoft Azure playbook, applied to AI infrastructure.
3. OpenClaw exposed the gap. When OpenClaw went viral and was reportedly acquired by OpenAI last month, it demonstrated real enterprise demand for open, composable agent frameworks. Nvidia, with its existing NeMo infrastructure and deep enterprise relationships, saw the opening.
This is a platform play, not a product launch. The distinction matters enormously for how enterprises should evaluate it.
NemoClaw vs. OpenClaw vs. Proprietary: A CTO’s Trade-off Map
Enterprise AI agent decisions in 2026 essentially come down to three buckets. Here’s an honest comparison based on what’s confirmed today, with appropriate caveats for what remains unverified pre-GTC.
The table above reflects reality as of March 13, 2026. Many NemoClaw entries carry significant uncertainty. “Built-in security layers” is a marketing claim until independent audits confirm it. “Hardware agnostic” is architecturally sound given NeMo’s existing design but untested at enterprise scale for NemoClaw specifically.
For CTOs in regulated industries (financial services, healthcare, defense), the governance maturity gap is real and won’t close at GTC. Proprietary solutions with documented compliance frameworks will remain the safer near-term choice. For CTOs in less regulated sectors building internal automation, NemoClaw’s open-source model and local execution story could be compelling by Q3 2026, assuming the security claims hold up.
The Security Question No One Is Answering Yet
Every serious discussion of AI agents eventually arrives at the same problem: agents that can act autonomously, access private data, communicate externally, and execute multi-step tasks are, by definition, high-risk software. The same properties that make them useful make them dangerous if misconfigured or compromised.
Cybersecurity experts have flagged OpenClaw’s architecture as exhibiting what they call a “lethal trifecta”: persistent access to private organizational data, the ability to communicate with external endpoints, and outputs that could include harmful or manipulated content. Nvidia’s pitch claims NemoClaw addresses these through built-in security and privacy layers. That claim needs scrutiny.
Three specific questions enterprise security teams should demand answers to at GTC and immediately after:
Scope limitation: What mechanisms prevent an agent from accessing data stores beyond its defined scope? Are these enforced at the architecture level or configurable (and therefore breakable)?
Audit logging: Does NemoClaw provide immutable audit trails for every agent action, meeting the evidentiary standards required for SOC 2, ISO 27001, or HIPAA compliance?
External communication controls: How does NemoClaw handle agent-initiated outbound connections? What allowlisting or sandboxing is built in by default?
The Nvidia NeMo platform already includes observability tooling for model monitoring. If NemoClaw extends these to agent-level action logging, that’s a genuine security differentiator. If it doesn’t, the “built-in security” claim is largely positioning.
Until post-GTC technical documentation is published and third-party security researchers have reviewed the codebase, CISOs should treat NemoClaw’s security posture as unverified. That’s not a reason to dismiss the platform; it’s a reason to build evaluation timelines accordingly.
What Enterprise Leaders Should Do Right Now
NemoClaw is pre-announcement. Most decisions can wait for the March 16 keynote and post-GTC documentation. But the strategic questions worth working through now will sharpen your evaluation criteria when the details land.
For CTOs and Engineering Leaders
Map your current AI agent surface area. Which workflows already involve multi-step AI automation? NemoClaw’s relevance depends entirely on whether you’re building in this space or planning to.
Review your NeMo dependency. If your org already runs on NeMo’s model lifecycle tools, NemoClaw integration will likely be low-friction. If not, factor in migration costs.
Define your hardware strategy first. NemoClaw’s hardware-agnostic claim is attractive, but verify it for your specific infrastructure before it influences procurement decisions.
Schedule a security architecture review for Q2 2026 once the codebase is public and external audits begin circulating.
For CISOs
Don’t wait for GTC to start your threat model. Document the data access patterns, external communication requirements, and compliance obligations that any enterprise AI agent platform will need to satisfy for your organization.
Engage your red team to evaluate the “lethal trifecta” risks in your current agent deployments. NemoClaw will inherit these risks unless its architecture explicitly addresses them.
Establish vendor security review criteria now so you can apply them consistently to NemoClaw, OpenClaw derivatives, and proprietary alternatives.
For Founders and Product Leaders
Watch the partnership announcements closely. If Salesforce, Cisco, or CrowdStrike formally integrates with NemoClaw, it signals distribution advantages that could compress your go-to-market timelines in those ecosystems.
Evaluate the open-source community trajectory post-GTC. Platform health in open-source AI frameworks is measurable: GitHub stars, contributor velocity, and corporate sponsorship signal long-term viability better than launch press coverage.
The Timeline to Watch
March 9, 2026: WIRED breaks NemoClaw story; Jensen Huang pitches confirmed to multiple enterprise firms.
March 10, 2026:Engadget and CNBC confirm, noting enterprise focus and five named companies in pitch process.
March 16, 2026:GTC 2026 keynote (San Jose, March 15-19): Expected formal announcement, technical documentation, and potential partner confirmations.
Q2 2026: First enterprise pilots expected; security audits of open-source codebase begin; partnership deal flow becomes visible.
Q3 2026: Earliest credible assessment of adoption metrics, developer community health, and security posture validation.
The Bigger Picture
The pattern emerging from NemoClaw’s pre-announcement is this: the AI agent layer is becoming the new enterprise platform battleground, and every major infrastructure company is now competing for it. Nvidia’s move isn’t surprising in retrospect. What’s notable is the method: open-source, hardware-agnostic, and pitched directly to the enterprise software companies that could otherwise become competitors.
This matters beyond Nvidia’s balance sheet. It signals that the agentic AI market is consolidating around orchestration frameworks faster than most analysts projected twelve months ago. The companies that establish platform relationships now, through integrations, security certifications, and developer toolchains, will shape which agent platforms enterprises standardize on through 2030.
Watch for three developments in the next 90 days: (1) which of the five pitched companies announce formal NemoClaw integrations at or after GTC, (2) whether the open-source codebase draws meaningful external security review or remains primarily Nvidia-controlled, and (3) how Microsoft, Salesforce, and ServiceNow respond with their own agent platform messaging. The organizations that evaluate NemoClaw rigorously now, rather than either dismissing it or adopting it uncritically, will be positioned to make the infrastructure decisions that define their AI roadmap for the next three years.
Editorial note: This article is based on pre-announcement reporting from WIRED (March 9, 2026), Engadget, CNBC, Techloy, and Investing.com. Nvidia had not issued official confirmation of NemoClaw as of publication on March 13, 2026. All technical specifications, partnership details, and security claims are sourced from third-party reporting and should be treated as unverified until Nvidia publishes primary documentation. NeuralWired will update this analysis following the GTC 2026 keynote on March 16.
Meta MTIA Chips: 25x Compute in Under 2 Years | NeuralWired
AnalysisAI InfrastructureMarch 13, 2026
Meta just unveiled four generations of custom silicon in a single announcement. The specs are striking. The strategy behind them is more interesting.
NW
NeuralWired Editorial
AI Infrastructure Analysis
10 min read
25x
Compute gain MTIA 300 to 500 (MX4 FLOPS)
~6mo
Chip generation cadence vs. industry 1 to 2 years
$125B
Meta 2026 capex midpoint for AI buildout
On March 11, 2026, Meta dropped what amounts to a two-year chip roadmap in a single blog post: four generations of its Meta Training and Inference Accelerator, announced together, spanning chips already in production to chips headed for mass production in early 2027. The MTIA 300 is live and running recommendation and ranking workloads right now. The MTIA 500 will deliver 30 petaFLOPS of MX4 compute and 27.6 TB/s of HBM bandwidth when it arrives.
That’s a 25x compute increase over the MTIA 300 across the product line. In under two years.
The announcement raises questions that go well beyond chip specs. Can Meta actually sustain a six-month silicon release cadence? Does this pressure Nvidia in any meaningful way? And what does it mean for the broader enterprise AI market when a consumer tech company starts publishing chip roadmaps that rival semiconductor incumbents? This analysis examines the full picture: what the chips do, who they threaten, where the risks sit, and what decision-makers should do with this information.
The MTIA Roadmap: What Meta Actually Announced
Meta’s MTIA program launched in 2023 with a first-generation inference chip. The March 11 announcement was a different order of magnitude. Meta’s official statement described “four new generations” on a cadence of “every six months or less.” That’s not a product launch. That’s a manufacturing and design philosophy.
Three things jump out. First, the MX4 precision format delivers roughly 6x the throughput of FP16 per clock cycle, which is why the compute numbers look so different between precision tiers. Second, HBM bandwidth grows 4.5x from the MTIA 300 to the 500, tracking the memory wall problem that dominates inference performance. Third, each chip slots into the same Open Compute Project rack standard, enabling data center swaps without infrastructure rebuilds.
The manufacturing stack behind this: TSMC on 3nm process nodes, Broadcom handling compute and I/O chiplet design, CoWoS advanced packaging. This isn’t a skunkworks experiment anymore. Meta is running serious silicon engineering at scale.
Why the Six-Month Cadence Changes the Calculus
The semiconductor industry typically runs on 12-to-24-month product cycles. Nvidia’s H100 to B200 arc took years of engineering. Meta is claiming a six-month generation-over-generation cadence. Whether that’s sustainable long-term is an open question, but the structural reasons it’s possible are worth understanding.
Custom silicon designed for a narrow workload class is far simpler to iterate than a general-purpose GPU. Meta’s chips are inference-first by design. They don’t need to support every CUDA workload, every graphics pipeline, every compute primitive that Nvidia’s customers demand. Narrower scope means faster design cycles, faster tape-out, faster validation.
“We’ve developed a competitive strategy for MTIA by prioritizing rapid, iterative development, an inference-first focus, and frictionless adoption by building natively on industry standards.”
— Meta Platforms, official March 2026 statement
The modularity helps here too. Swapping chiplets within the same rack-scale architecture means Meta doesn’t need to redesign the whole data center each generation. The 72-chip-per-rack MTIA 400 configuration reported by Yahoo Finance gives a sense of the density they’re targeting. New chips drop in. The surrounding infrastructure stays.
Meta is already operating at “hundreds of thousands” of MTIA chips for inference workloads, covering ad ranking, content recommendations, and organic feed algorithms. This isn’t a pilot program. The chips are carrying real production load across billions of daily users. That scale provides a feedback loop that no commercial silicon vendor can match for Meta’s specific workloads.
The Nvidia Rivalry: Competitive or Complementary?
Meta’s announcement landed as a direct competitive shot at Nvidia and AMD. Yahoo Finance coverage noted Meta’s claim that the MTIA 400 is “its inaugural chip that offers both cost efficiency and raw performance that competes with leading commercial products.” That’s a pointed benchmark assertion.
But the full picture is more nuanced. Meta is simultaneously a major Nvidia customer, and Mark Zuckerberg has made no secret of that relationship. The MTIA program isn’t a wholesale replacement strategy. It’s a diversification play targeting specific inference workloads where Meta has enough volume and predictability to engineer a purpose-built solution that beats general-purpose GPUs on cost per operation.
The efficiency claim is significant: analysis from AInvest puts MTIA’s gains at up to 7x for key matrix operations versus general-purpose silicon. For a company running inference at Meta’s scale, that efficiency gap translates directly to billions in infrastructure savings annually.
“The goal is clear: break the AI compute cost curve, aiming for up to 7x gains for key matrix operations.”
— AInvest, Meta MTIA cost analysis, March 2026
For Nvidia, the real concern isn’t Meta. It’s what Meta’s success signals to every other hyperscaler. Google has TPUs. Amazon has Trainium and Inferentia. Apple runs Neural Engines. Microsoft has invested in Maia. Meta’s roadmap is the clearest evidence yet that custom silicon for AI inference is viable at production scale, not just a research exercise. That’s a structural shift in the competitive landscape, even if no single company is abandoning Nvidia GPUs tomorrow.
Technical Architecture: What Makes MTIA Different
MTIA’s inference-first design philosophy produces some specific architectural decisions worth examining for technically-oriented readers.
The MX4 precision format is central to the compute story. MX4 (Microscaling 4-bit) enables roughly 6x the floating-point operations per second versus FP16 at the same clock and power budget. This matters enormously for inference, where you’re running a trained model forward repeatedly at scale, not doing the high-precision arithmetic that training requires. Most inference workloads tolerate the precision reduction. The throughput gains are substantial.
FlashAttention hardware acceleration is built directly into the silicon. For transformer-based models (which now power most of Meta’s AI applications, from content ranking to Llama variants), attention computation is a primary bottleneck. Hardwiring it into the chip rather than implementing it in software on a general-purpose GPU is a meaningful advantage for Meta’s specific workload mix.
The software stack deserves attention. TrendForce reporting confirms native support for PyTorch, vLLM, and Triton, the dominant frameworks in Meta’s (and most of the industry’s) ML toolchain. Teams don’t need to rewrite models or change workflows to run on MTIA. This is the “frictionless adoption” Meta refers to, and it’s not a small detail. The biggest failure mode for custom silicon programs has historically been software ecosystem fragility.
The Data Center Dynamics writeup on the announcement confirms that by 2027, MTIA is targeting full generative AI workloads, not just ranking and recommendation. That’s a significant expansion of scope. Whether the architecture can handle GenAI inference at the scale Meta needs it to remains one of the key unanswered questions.
Risks and Honest Uncertainties
The announcement deserves scrutiny alongside the excitement. Several risk factors are real and worth naming directly.
Where the Skeptics Have a Point
3nm yields are hard. TSMC’s 3nm process is advanced but not without yield challenges. Meta’s cost projections depend on yields at scale that haven’t been publicly validated. TrendForce notes the manufacturing dependency without quantifying the risk.
Development costs are real.Bloomberg reports Meta has spent millions on this program. The ROI case is built on scale that only a handful of companies globally can match.
The six-month cadence is untested at this scope. Claiming it and executing it across four generations while managing yield, packaging, and software integration simultaneously is operationally demanding.
Scope creep risk. Expanding from ranking/recommendation to full GenAI inference means more complex workloads with less predictable access patterns. MTIA’s architecture may face surprises.
No independent benchmarks. All performance comparisons to Nvidia and AMD are Meta’s own assertions. Third-party validation at production scale hasn’t been published.
Meta’s $115 to 135 billion 2026 capex commitment, reported by TrendForce, gives the program a financial buffer that smaller organizations can’t replicate. But it also means the stakes on execution are enormous. A sustained yield problem or software integration failure on MTIA 450 or 500 doesn’t just affect a product line. It affects a quarter of a trillion dollars in planned infrastructure.
A Decision Framework for Enterprise Leaders
Most organizations reading this won’t be designing custom silicon. But this announcement has direct implications for infrastructure decisions being made right now.
Questions to Ask Before Your Next GPU Procurement
What’s your inference-to-training ratio? If you’re running more inference than training (most production AI teams are), the efficiency argument for inference-optimized silicon is directly relevant to your cost model.
Are your workloads predictable enough for custom silicon? MTIA works because Meta’s ranking and recommendation workloads are stable and high-volume. Diverse or experimental workloads still favor general-purpose GPUs.
Do you have the volume to justify it? The economics of custom silicon require scale. For most enterprises, the relevant action is negotiating harder on Nvidia and AMD pricing, not designing chips.
What’s your dependency concentration? If your AI infrastructure is 90%+ Nvidia, this announcement is evidence that diversification is both feasible and strategically important, even if you use commercial alternatives rather than custom silicon.
Can your software stack absorb a hardware swap? Meta’s PyTorch-native approach lowers switching costs dramatically. If your team is framework-agnostic, inference hardware alternatives (Google TPUs, Amazon Inferentia) deserve fresh evaluation against your current Nvidia contracts.
What This Signals for AI Infrastructure Through 2027
The pattern emerging from this announcement isn’t just about Meta MTIA chips. It’s about a fundamental restructuring of how AI compute gets built and procured.
We’re moving from a world where “AI infrastructure” meant “buy Nvidia GPUs” to a world where the compute layer is fragmenting. Custom silicon programs at Google, Amazon, Microsoft, and now Meta are all heading in the same direction: inference workloads, which represent the majority of production AI compute by volume, are increasingly handled by purpose-built accelerators rather than general-purpose GPUs. Training still depends on Nvidia for most organizations, but inference is becoming a contested market.
For investors, the implications for Nvidia’s margins are worth watching. Nvidia’s dominance has historically come from a combination of hardware performance and CUDA ecosystem lock-in. Meta’s PyTorch-native approach for MTIA, and Google’s JAX stack for TPUs, are both evidence that the software moat is more crossable than it looked three years ago. Pressure on inference revenue could emerge as these programs mature.
Watch for three developments in the next 18 months. First, independent benchmarks comparing MTIA 400 to H100 and B200 on real inference workloads. Meta’s internal numbers will eventually face external validation or scrutiny. Second, whether the MTIA 450 and 500 timelines hold, specifically whether the six-month cadence survives the complexity jump to full GenAI workloads. Third, whether any other hyperscalers accelerate their own custom silicon announcements in response.
Meta has published a roadmap. Now comes the harder part: executing it.
Nscale Hits $14.6B Valuation in $2B Series C Round
March 9, 2026 | AI Infrastructure | 8 min read
Nscale Hits $14.6B Valuation in $2B Series C Round
A UK AI infrastructure company founded just two years ago has raised $2 billion in a single round, placing its valuation at $14.6 billion and positioning itself as the most formidable European challenger to US hyperscalers.
Two years. That’s how long it took Nscale to go from founding to a $14.6 billion valuation. On March 8, 2026, the UK-based AI data center operator closed a $2 billion Series C round, bringing its total funding to approximately $4.9 billion in under 24 months. That trajectory doesn’t just turn heads. It rewrites what’s possible for European AI infrastructure companies.
The round attracted a striking investor mix: Norway’s Aker, 8090 Industries, Nvidia, Citadel, Dell, Jane Street, Lenovo, Nokia, and Point72. Customers include Microsoft and OpenAI. The company simultaneously added Sheryl Sandberg, Nick Clegg, and Susan Decker to its board, a signal to public markets that an IPO is not a distant hypothetical.
This analysis examines what drives a $14.6 billion valuation for a company with no public revenue figures, how Nscale’s 1.3GW pipeline and 200,000 contracted Nvidia GPUs compare to rivals like CoreWeave, and what the Series C means for CTOs allocating compute budgets, investors assessing AI infrastructure multiples, and policymakers watching European sovereign AI capacity.
The Funding Trajectory That Shocked the Market
Nscale’s capital raise history reads less like a startup funding story and more like a sovereign infrastructure program accelerated by private capital. Josh Payne founded the company in 2024. By December of that year, Nscale closed a $155 million Series A, which Payne called “one of the largest Series A rounds raised in UK history” at the time.
The pace only accelerated. In September and October 2025, the company raised a $1.1 billion Series B followed immediately by a $433 million pre-Series C SAFE, with Nvidia and Dell among the backers. In February 2026, Reuters reported that Goldman Sachs and JPMorgan had been hired to prepare for a potential IPO, alongside a $1.4 billion GPU-backed delayed draw term loan to fund European cluster builds. The Series C followed weeks later.
That’s $4.9 billion raised in roughly twelve months of active fundraising. For context, CoreWeave, Nscale’s closest US analog, took several years to reach comparable capital scale before its own IPO process.
“The pace with which we have expanded our capacity demonstrates both our readiness and our commitment to efficiency, sustainability and providing our customers with the most advanced technology available,” said Josh Payne, CEO of Nscale, commenting on the company’s Microsoft deal in October 2025.
The Microsoft deal itself was a statement. Nscale secured a contract to deploy 104,000 Nvidia GPUs at a 240MW Texas data center site with the capacity to scale to 1.2GW. That single deployment underpins a significant portion of the valuation narrative and gives investors something concrete to underwrite beyond pipeline projections.
What Justifies the $14.6B Nscale Valuation?
At $14.6 billion, Nscale is being valued on what it can build, not what it has built. No public revenue figures exist. No utilization rates have been disclosed. The valuation rests on three structural arguments that investors appear willing to accept in the current market.
First, the contracted demand is real. Microsoft and OpenAI don’t sign multi-hundred-megawatt compute contracts speculatively. The 104,000 GPU Texas deployment with Microsoft and the ongoing OpenAI relationship represent genuine anchor revenue. These aren’t letters of intent; they’re infrastructure commitments that take years to unwind.
Second, the GPU supply position is a genuine moat. Nscale has 200,000 Nvidia GPUs contracted across its 1.3GW pipeline spanning the UK, Norway, Ohio, and Texas. In a market where hyperscalers are competing for the same Nvidia allocation, holding a contracted supply position at that scale is competitively meaningful. Nvidia’s direct investment in the Series C reinforces this relationship.
Third, the market trajectory makes the multiple defensible. The AI infrastructure market is projected to grow from $32.98 billion in 2025 to $146.37 billion by 2035, an 18% compound annual growth rate. Global AI data center capital expenditure in 2026 alone is estimated at $602 billion, up 36% year over year according to Goldman Sachs. A company holding confirmed capacity in that environment earns a premium.
The honest counterpoint: this is a pipeline valuation. The Next Web noted that the claim of “largest European Series C” deserves scrutiny, and several industry observers have flagged that the gap between contracted capacity and operating capacity remains unbridged. The multiple assumes flawless execution on buildout, grid access, and sustained hyperscaler demand. None of those are guaranteed.
Board Additions Signal IPO Timeline
The Series C announcement came bundled with three board appointments that read like an IPO preparation checklist. Sheryl Sandberg, former Meta COO and one of the most recognized names in technology governance, joins alongside Nick Clegg, the former UK Deputy Prime Minister and most recently Meta’s President of Global Affairs. Susan Decker, former Yahoo President, rounds out the trio.
Each appointment serves a distinct purpose. Sandberg brings institutional investor credibility and US market access. Clegg brings European regulatory fluency and government relations at a moment when UK and EU AI policy is being actively written. Decker’s operational experience with large-scale digital businesses addresses questions about Nscale’s readiness to manage a publicly traded company’s governance demands.
Yahoo Finance noted the board composition signals IPO intent, and Reuters had already reported in February that Goldman Sachs and JPMorgan were engaged. The trajectory points toward a late 2026 public offering, though the company hasn’t confirmed timing publicly.
For investors assessing Nscale’s readiness, The Times reported that the board additions coincided with the funding close, suggesting these weren’t afterthoughts. This level of governance investment at Series C, rather than pre-IPO, reflects how seriously the company’s backers are treating the public market timeline.
The Real Risk: Power, Grid Delays, and Execution
The story Nscale is telling is compelling. The risks embedded in executing it deserve equal attention.
Grid access is the single biggest constraint on AI data center growth globally. Axios reported that approximately 50% of major AI data center projects face risk of postponement due to power infrastructure delays. In Norway, where Nscale has significant planned capacity, Global Data Center Hub flagged that grid queue timelines and renewable energy availability create real execution uncertainty. Cold climates are excellent for cooling; they don’t solve interconnection queues.
Nscale was founded in 2024. It now carries $4.9 billion in obligations. The institutional talent to build, operate, and sell hyperscale AI infrastructure at this speed is genuinely scarce. The company has secured the capital and the contracts, but transforming those into operating megawatts requires execution capacity that takes years to build in most organizations.
The valuation stretch is also real. At $14.6 billion against no disclosed revenue, Nscale’s multiple is priced on future capacity delivery, not current earnings. If one major customer relationship shifts, if GPU delivery schedules slip, or if interest rates affect the economics of its GPU-backed debt facilities, the cushion between pipeline valuation and realized value compresses fast.
What to watch: Track Nscale’s 2026 capacity milestones against announced timelines. The gap between contracted gigawatts and live gigawatts will be the most honest indicator of whether the valuation holds through an IPO.
How CTOs, Investors, and Policymakers Should Read This
Nscale’s Series C isn’t just a funding story. It’s a signal about how the AI compute market is restructuring. Here’s what different decision-makers should take from it.
CTOs and infrastructure teams: Nscale’s model, vertically integrated GPU clusters contracted to hyperscalers, represents a growing alternative to direct cloud provider relationships. For organizations facing compute shortages in 2026, understanding the emerging landscape of AI-native infrastructure providers matters for capacity planning. Long-term GPU contracts with providers that have secured supply will increasingly outperform spot market strategies.
CFOs and investors: The 18% CAGR to $146 billion in AI infrastructure through 2035 justifies aggressive capital allocation to the sector, but the CoreWeave comparison is instructive. Early movers with contracted anchor customers and GPU supply lock-in command premium multiples. Nscale fits that profile. The risk is execution, not demand.
Founders and product leaders: Nscale’s rise illustrates that vertical integration, owning the GPU, the facility, and the software stack, creates stickier customer relationships than reselling hyperscaler capacity. For AI infrastructure startups, the window to carve out sovereign or regional positions before the major players consolidate is narrowing fast.
Policymakers: Nscale is the clearest proof point that European AI infrastructure ambitions can attract institutional capital at scale. The UK now has a hyperscaler-class company. The question is whether grid policy, planning frameworks, and renewable energy commitments can match the pace of private investment.
AI Infrastructure’s Super Cycle and What Comes Next
Nscale’s $14.6 billion valuation doesn’t exist in isolation. It’s a data point in a broader market reordering that’s been building since 2023 and is now reaching a pace that makes individual company announcements feel almost routine.
The $602 billion in AI data center capital expenditure projected for 2026 represents a 36% increase over 2025. Microsoft, Google, Meta, and Amazon have each announced multi-year, multi-billion-dollar infrastructure commitments. The demand signal is unambiguous. What’s less clear is which companies outside the established hyperscaler tier will capture meaningful share of that spending.
CoreWeave, the closest US analog to Nscale, went public and established a template for GPU-native cloud companies. Nscale is building toward that position in Europe and increasingly in the US market, backed by stronger anchor customer relationships at an earlier stage than CoreWeave had at comparable funding levels.
The pattern across this cycle is now consistent: the AI compute super cycle is creating a new class of infrastructure company, one that sits between traditional cloud providers and on-premise deployments, capturing enterprises and AI labs that need dedicated GPU capacity without building their own. Nscale is positioning for that category, and the $4.9 billion it has raised in under two years suggests the market agrees with the thesis.
Watch for three developments in the next twelve months: (1) Nscale’s 2026 capacity coming online against committed timelines, which will determine IPO readiness and public market reception; (2) European grid policy responses to the surge in AI infrastructure demand, which will affect Nscale’s Norway and UK buildout directly; (3) whether Microsoft and OpenAI deepen or diversify their Nscale dependency as their own infrastructure strategies evolve. The organizations that lock in GPU capacity contracts now, at this stage of the cycle, will operate at a structural advantage through 2028 and beyond. The ones still evaluating in twelve months may find both the capacity and the favorable contract terms are gone.
Stargate’s $600M Collapse: Why AI Infrastructure Fails | NeuralWiredNeuralWired
Analysis | AI Infrastructure | March 8, 2026
Infrastructure Analysis
Stargate’s $600M Collapse: Why AI Infrastructure Fails
NeuralWired StaffMarch 8, 20268 min read
Oracle and OpenAI just abandoned a 600 MW expansion of the most-hyped AI campus on earth. The real story isn’t the cancellation. It’s what the Abilene case reveals about the hidden physics of building AI infrastructure at gigawatt scale.
600 MWExpansion cancelled
$150MNvidia’s deposit to Crusoe
4.5 GWStill planned elsewhere
When Donald Trump stood in the White House on January 21, 2025, flanked by Sam Altman, Larry Ellison, and Masayoshi Son, he called the
Stargate AI infrastructure project
“the largest AI infrastructure project, by far, in history.” Less than 14 months later, Oracle and OpenAI quietly abandoned a planned 600 MW expansion of Stargate’s flagship Texas campus, scrapping enough computing capacity to power a mid-sized city’s worth of AI workloads.
This isn’t a story about failure. The core Abilene campus is still being built. Oracle and OpenAI still plan to develop
4.5 GW of capacity
at other sites. But the Stargate data center expansion collapse in Abilene, Texas, reveals something the headlines missed: even a $500 billion project backed by the U.S. president can hit the wall where demand forecasting, financing mechanics, and partner alignment fail to converge.
This analysis breaks down what actually happened, who bears the risk now, and what the Abilene case tells CTOs, CFOs, and infra investors about the physics of building AI at gigawatt scale.
The Anatomy of a Cancelled Expansion
The Abilene Stargate campus is genuinely impressive engineering: roughly 1,100 acres on the outskirts of a mid-sized Texas city, designed to eventually draw 1.2 GW of power, equivalent to supplying around 750,000 homes. Initial deployment hit approximately 200 MW. Ten to twenty “AI factory” halls are planned, each capable of housing tens of thousands of high-density GPU servers. The project’s estimated capex runs to roughly
$3 to $4 billion per GW
of capacity, based on industry benchmarks and partial disclosures.
In September 2025, Oracle and OpenAI announced plans to add another 600 MW adjacent to the flagship campus. By March 6, 2026,
Reuters and Bloomberg reported
that those plans were dead. Two forces killed the expansion: financing negotiations that dragged without resolution, and a shift in OpenAI’s demand forecasts that made the additional capacity harder to justify.
Demand forecasting is the hidden variable in almost every large-scale infra collapse. Changes in model architecture, training efficiency gains, or shifts in deployment strategy can eliminate the need for hundreds of megawatts that looked essential six months earlier. The public reporting doesn’t specify exactly how OpenAI’s requirements changed, whether it was a pivot in training methodology, a reassessment of inference needs, or something else. But the scale of the consequence is clear: 600 MW of planned capacity, representing roughly $2 billion in potential capex at the midpoint estimate, was redirected away from this single site.
Oracle’s stock traded lower after the news emerged. The company has simultaneously been
cutting thousands of jobs
while ramping capital allocation toward AI infrastructure, a rebalancing that signals a painful internal transition even amid an otherwise aggressive buildout strategy. OpenAI, xAI, and Meta are among Oracle’s named AI cloud customers, which means this capacity is being redistributed, not abandoned.
Where the Risk Landed: Nvidia’s $150M Move
Here’s where the story gets structurally interesting. When Oracle and OpenAI walked away from the Abilene expansion, the site didn’t go dark. Crusoe, the data center developer and operator managing the campus, still holds the land, the power commitments, and ambitions to monetize the footprint.
Enter Nvidia. According to
Bloomberg’s reporting,
Nvidia paid a $150 million deposit to Crusoe tied to the expansion site, then actively began recruiting Meta as a replacement tenant. The motive is transparent: Nvidia wants its GPUs filling that facility. If the site sits without a committed buyer, AMD has a window. A $150 million deposit to broker a favorable tenancy arrangement is, from Nvidia’s perspective, an investment in hardware placement, not charity.
Meta is reportedly in discussions to lease the expansion footprint from Crusoe. No lease has been finalized as of this writing, and no MW or term details have been disclosed. But the dynamic illustrates something that will increasingly define AI infrastructure: chip vendors are becoming infrastructure financiers.
“We’re looking for stranded energy, energy that was not being used, to power compute.”
Jamie McGrath, SVP at Crusoe — briefing Abilene city officials, March 4, 2026
This matters beyond this single deal. When a GPU manufacturer puts $150 million into securing placement over a competitor, it signals that the data center real estate game is no longer just about hyperscalers and cloud operators. Nvidia is effectively acting as a demand aggregator, using capital to ensure its hardware stays embedded in new capacity, regardless of which hyperscaler ultimately operates it. For infra developers like Crusoe, that creates a new source of financing and tenant recruitment support. For AMD, it raises the strategic bar for competing in large-scale campus deals.
Crusoe SVP Jamie McGrath told Abilene city officials
on March 4, 2026, just two days before the expansion cancellation became public, that the Abilene campus was built around using under-utilized or curtailed generation capacity from the Texas grid. That strategy didn’t change when Oracle and OpenAI exited. But it underscores how much energy procurement, not just tenant selection, determines whether GW-scale campuses succeed.
The Stargate Data Center Expansion Failure as a Framework
The Abilene case is more than AI industry gossip. It’s a stress test of the decision model every organization building or leasing large-scale compute infrastructure needs to run, and a signal that most current models are broken.
Three failure modes are visible in this story.
Failure Mode 01
Demand Forecasting at Multi-Year Horizons
When OpenAI committed to needing an additional 600 MW adjacent to Abilene, it was forecasting training and inference demand out multiple years based on model roadmaps and utilization assumptions that subsequently shifted. AI architecture is evolving fast enough that 18-month demand projections carry substantial uncertainty. Building 600 MW of shell and power capacity against a single tenant’s forecast creates enormous stranded-asset risk the moment that forecast changes.
Failure Mode 02
Financing Alignment Between Parties With Different Risk Profiles
Oracle, as the cloud operator, needs the build to pencil out against tenant revenue. OpenAI, as the AI tenant, needs flexibility to respond to changing model requirements. Crusoe, as the developer, needs committed capital to build. These interests don’t naturally align. When financing negotiations “dragged,” it likely reflected structurally incompatible assumptions about who bears the risk of utilization falling short. Pre-paid capacity agreements, revenue-share structures, and build-to-suit leases all distribute this risk differently, and the public reporting gives no clarity on what structures were on the table or why they failed.
Failure Mode 03
Multi-Party Misalignment
The Stargate program involves at minimum Oracle, OpenAI, SoftBank, Crusoe, Lancium, Nvidia, and the Trump administration, plus Meta now as a potential tenant. Each party has different time horizons, return requirements, and strategic priorities. Trump’s framing of Stargate as a geopolitical infrastructure project creates pressure to announce and build fast. Crusoe’s incentive is to fill land and power commitments. Nvidia’s incentive is hardware placement. OpenAI’s incentive is flexibility. When these don’t align, projects stall or get cancelled even when macro demand for AI compute remains strong.
The broader Stargate build is
continuing through at least 2028
at the Abilene core site. Oracle and OpenAI are still pursuing 4.5 GW of additional capacity elsewhere. The cancellation is not a sign that AI infrastructure demand has collapsed. It’s a sign that the financing and coordination machinery for GW-scale campuses is still being invented in real time.
What This Means for Infra Decision-Makers
If you’re a CTO, CFO, or infrastructure investor evaluating large-scale AI compute commitments, whether as a tenant, operator, or financier, the Abilene case surfaces four questions worth pressure-testing now.
The Abilene Decision Framework: Four Questions
01 →
What’s your minimum committed-utilization threshold for approving an expansion?
The Abilene cancellation suggests Oracle and OpenAI didn’t have a locked commitment sufficient to justify the financing. Before green-lighting any 200 MW+ build, verify that signed off-take or capacity agreements cover enough utilization to service the debt and hit minimum returns. “We expect to need this” is not a commitment.
02 →
Are your demand forecasts scenario-weighted or point estimates?
Point-estimate forecasting, “we’ll need X exaFLOPs by 2027,” is inadequate for multi-year infrastructure decisions in AI. Scenario-weighted approaches that model architecture shifts, efficiency gains, and competitive dynamics give the decision more credibility and create explicit triggers for pausing or redirecting capacity.
03 →
Is your campus design tenant-agnostic?
Crusoe’s pivot toward Meta was possible because the land, power, and shell infrastructure were separable from the Oracle/OpenAI tenancy. Campuses designed around a single tenant’s specific rack layout, power density, or cooling configuration are harder to re-tenant. Infra developers should build to the most common hyperscale standard, not the specific requirements of one AI lab.
04 →
Who bears the demand risk in your contract structure?
Nvidia’s $150 million deposit to secure GPU placement is a form of demand-risk transfer: the chip vendor is effectively subsidizing tenancy to ensure its hardware gets placed. Developers and cloud operators should assess whether their financing structure accounts for this type of third-party risk subsidy, and whether they can structure equivalent arrangements with other hardware vendors.
The Road Ahead for Stargate
The pattern from Abilene is clear: at gigawatt scale, the gap between announced ambition and executable commitment is large, and it shows up fastest in the expansion phases after the flagship build. This isn’t a reason to dismiss Stargate’s broader goals. It’s a reason to watch the execution methodology more carefully than the headline numbers.
Three things will determine whether the
Stargate data center expansion
program hits anywhere near its 10 GW target: whether demand forecasting gets more rigorous as models and inference architectures stabilize; whether chip vendors like Nvidia continue deepening their role as infra co-financiers; and whether developers like Crusoe build enough tenant-agnostic flexibility into their campuses to absorb anchor-tenant exits without stalling entire sites.
Watch for three near-term signals: a formal Meta-Crusoe lease announcement with disclosed MW figures; Nvidia earnings commentary on pre-payments and partnership structures; and Oracle’s next capex guidance on data center pipeline, which will reveal how much of the
4.5 GW elsewhere
is committed versus aspirational. The organizations that treat those signals as inputs to their own infra planning, rather than just AI industry news, will build more resilient capacity strategies than those still using point-estimate demand forecasts and single-tenant site designs.
Trump called it the largest AI infrastructure project in history. That may still prove true. But the Abilene expansion collapse shows that even the largest projects are subject to the same financing physics as every other capital-intensive bet: ambition is cheap, committed cash flow is not.
The headline numbers are staggering. NVIDIA’s new Rubin GPU delivers 50 petaflops of NVFP4 inference performance, five times the throughput of a Blackwell GB200. Pack 72 of them into a single NVL72 rack, lace them together with NVLink 6 at 3.6 terabytes per second per GPU, and you’re looking at a machine that makes the world’s most powerful AI supercomputers of three years ago look modest.
But here’s what the press releases don’t tell you: the Rubin NVL72 isn’t a GPU upgrade. It’s a facilities project.
Before a single inference token flows through a Rubin rack, your data center needs to deliver 120 kilowatts of liquid-cooled power per rack, route 1.6 terabits per second of external network bandwidth per GPU, and supply 480-volt three-phase AC through four dedicated 30-kilowatt power shelves. The networking optics alone, just the transceivers, can cost between $550,000 and $2.2 million per rack. That’s before you’ve bought a single chip.
Most CIOs discover these constraints about 18 months too late.
This guide is the due-diligence dossier they needed at the start. We’ll walk through the Rubin platform’s architecture, dissect the rack-level engineering reality, quantify the total cost of ownership across multiple deployment scenarios, and give you the decision framework to determine whether, and when, Rubin NVL72 belongs in your infrastructure roadmap.
Section 01
The Six-Chip Architecture Behind the “Rack Is the Computer” Claim
NVIDIA didn’t build Rubin by making a faster GPU. They built a new computing paradigm around six co-designed chips that function as a unified system, and understanding that distinction is essential before you commit a single dollar to planning.
According to NVIDIA’s February 2026 architecture brief, the Vera Rubin platform consists of: the Rubin GPU itself, the Vera CPU, the NVLink 6 switch ASIC, a new networking chip, a DPU, and a next-generation NIC. None of these components is optional. They’re engineered to work as an integrated whole, which is precisely what allows NVIDIA to call the NVL72 rack a single accelerator.
Multiply across 72 GPUs in a single NVL72 rack and you’re looking at 3,600 PFLOPS of inference compute in a single cabinet.
The Vera CPU | More Than a Host Processor
The Vera CPU isn’t just a general-purpose host attached to the GPUs. It’s a purpose-built accelerator for the model management and orchestration work that modern AI inference demands.
Each NVL72 rack pairs 36 Vera CPUs with 72 Rubin GPUs, one CPU for every two GPUs, in a configuration described by SemiAnalysis that also deploys 36 NVLink 6 switch ASICs as the internal fabric spine.
NVLink 6 | The Glue That Makes 72 GPUs Act as One
The most technically consequential component in the Rubin platform isn’t the GPU. It’s NVLink 6.
This architectural choice, treating the rack as a single compute unit rather than a cluster of individual accelerators, drives many of the deployment constraints that follow. To deliver 260 terabytes per second of internal bandwidth at scale, you need to move the NVLink switch complexity inside the rack. That means density. And density means heat. And heat means liquid cooling is no longer optional.
Wheeler’s Network analysis reveals a critical design decision: NVIDIA achieves Rubin’s doubled NVLink bandwidth while maintaining backward compatibility with the Oberon rack backplane introduced with Blackwell. The new NVLink switch tray carries four NVLink ASICs, versus two in the Blackwell NVL72, while reusing 5,184 passive copper cables already embedded in the Oberon spine. This is smart engineering. It protects prior infrastructure investment while doubling internal bandwidth.
The hidden costs, as we’ll see, don’t live in the rack metal. They live in the power distribution, liquid cooling infrastructure, and external optical networking.
Section 02
The Real Power Math | Why 120 Kilowatts Per Rack Changes Everything
Before we get to the Rubin-specific numbers, let’s establish the baseline. Understanding why Rubin-class systems require liquid cooling isn’t optional, it determines whether your current facility can host this hardware at all.
SemiAnalysis established the key thresholds: a general-purpose CPU rack draws around 12 kilowatts. An H100 air-cooled rack manages roughly 40 kilowatts. The GB200 NVL72, Rubin’s immediate predecessor, draws approximately 120 kilowatts per rack. Liquid cooling becomes mandatory once rack density exceeds around 40 kilowatts. The GB200 NVL72 blows past that threshold by a factor of three.
‘The first one is the GB200 NVL72 form factor,’ SemiAnalysis researchers noted in their hardware architecture analysis. ‘This form factor requires approximately 120kW per rack. To put this density into context, a general-purpose CPU rack supports up to 12kW/rack, while the higher-density H100 air-cooled racks typically only support about 40kW/rack. Moving well past 40kW per rack is the primary reason why liquid cooling is required for GB200.’
For GB200 and Rubin NVL72, liquid cooling isn’t an upgrade option. It’s table stakes.