In-depth artificial intelligence analysis: AI agents, LLMs, enterprise deployment, governance, and breakthroughs. Research-backed insights for CTOs, founders, and decision-makers.
Nvidia Hikes AI Server Prices 15%+ as Memory Crisis Bites
AI Infrastructure
Nvidia Hikes AI Server Prices 15%+ as Memory Crisis Bites
By NeuralWired Staff · Published August 23, 2026 · 8 min read
Nvidia just told its biggest customers to expect a bigger bill. Servers built around its flagship Vera Rubin and Grace Blackwell chips are going up more than 15% for systems shipping in early 2027, and the reason has nothing to do with the GPUs themselves. It’s memory, and the shortage behind it is reshaping how every major AI buyer plans its 2027 budget.
The increase, first reported by Bloomberg News on August 22, 2026, lands four days before Nvidia reports Q2 fiscal 2027 earnings, and it confirms something procurement teams have suspected for months: the AI chip price hike isn’t a one-time correction. It’s the visible edge of a memory supercycle that’s already rewritten pricing across the entire semiconductor stack, from data center racks down to the graphics card in a gaming PC.
According to Bloomberg’s reporting, Nvidia has notified some of its largest customers that server prices built around its AI chips will climb more than 15% in many cases, driven by soaring memory chip costs. The increases target systems shipping starting early 2027 and hit both the Vera Rubin and Grace Blackwell platforms, with the exact size depending on chip generation and memory configuration.
The ripple effect is already visible. Contract manufacturers building servers for Microsoft, Alphabet’s Google, and Oracle have started notifying their own customers about the coming increases, according to the report. Nvidia did not respond to requests for comment before publication.
Worth noting: Bloomberg’s sourcing is anonymous, described only as “people familiar with the process.” Reuters said it could not immediately verify the report. Treat this as single-sourced but multiply corroborated, since Fortune, CNBC, and the broader wire picked it up without dispute, and the pattern lines up with independently tracked consumer GPU pricing (more on that below).
Why Memory Chips Are Driving the Increase
Here’s the part that matters for anyone modeling 2027 infrastructure costs: this isn’t Nvidia squeezing more margin out of hyperscalers. It’s a pass-through of a cost shock that started with the companies that make DRAM and HBM, specifically Samsung, SK Hynix, and Micron.
Those suppliers have spent the past year shifting wafer capacity toward high-bandwidth memory for AI accelerators, and that’s left less room for conventional DRAM and NAND used in everything else. The result, according to TrendForce, is one of the sharpest memory price runs on record.
Metric
Figure
Source / Date
DRAM contract price growth, Q1 2026
90% to 95% QoQ
TrendForce, Feb 2026
DRAM contract price forecast, Q3 2026
13% to 18% QoQ
TrendForce, Jul 2026
Combined memory industry revenue, Q1 2026
$97 billion (+81% QoQ)
TrendForce, Jun 2026
Micron fiscal Q3 2026 revenue
$41.46 billion (+346% YoY)
Company earnings, Aug 2026
Memory share of AI system cost (Vera Rubin)
~29%, vs. Nvidia’s 20% target
Wedbush Securities, Jul 2026
That last line is the real story. When memory eats up nearly a third of a system’s cost instead of a fifth, a company with 75% gross margins doesn’t just eat the difference quietly. It redesigns the product and raises the price. Nvidia reportedly cut the capacity of its SOCAMM memory modules in half on the upcoming Vera Rubin platform, a sign the shortage is now shaping hardware decisions, not just invoices.
The Consumer Market Already Felt This
If this all sounds sudden, it isn’t. Enthusiast GPU buyers got hit first. Tom’s Hardware tracked U.S. retail RTX 50-series prices on Newegg and found the median RTX 5070 price jumped 36% between June and August 2026, from $659.99 to $899.99. The RTX 5060 Ti 16GB rose 39% in the same window, and the entry-level RTX 5060 climbed 27%.
Entry-level cards took the biggest hit because they have the least room to absorb a fixed-dollar memory cost increase. Bloomberg’s enterprise-side report is essentially the same story playing out one tier up, on hardware that costs tens of thousands of dollars instead of a few hundred.
Why the Timing Matters: Nvidia’s Q2 Earnings
Nvidia reports Q2 fiscal 2027 earnings on Wednesday, August 26, 2026, after market close, just four days after this pricing story broke. That’s the first moment Jensen Huang and CFO Colette Kress will have to publicly address the increase and what it means for margin.
Nvidia guided Q2 revenue of $91.0 billion, plus or minus 2%, with non-GAAP gross margin around 75.0%. For context, Q1 FY2027 revenue came in at $81.6 billion, up 85% year over year, with the Data Center segment alone hitting $75.2 billion. Analysts on the earnings call will almost certainly push for specifics on how much of the memory cost increase Nvidia is passing through versus absorbing.
What Industry Experts Are Saying
Four voices, four different vantage points on how long this lasts and who’s really driving it.
“Even in 2028, when supply begins to improve gradually, we will see that the demand will continue to be on a robust trajectory as well.”
Sanjay Mehrotra, President and CEO, Micron Technology, earnings call, June 25, 2026 · via NPR
SK Hynix CEO Kwak Noh-jung told Reuters that memory market conditions will get worse in 2027, with demand expected to outstrip supply beyond 2030. Coming from a supplier that benefits directly from tight supply, it’s worth reading as an interested forecast rather than neutral analysis, but it does align with Mehrotra’s timeline.
Matt Bryson, semiconductor analyst at Wedbush Securities, offered the more skeptical read. In a client note reported by Yahoo Finance, Bryson flagged that Nvidia’s decision to halve SOCAMM module capacity on Vera Rubin shows rising memory costs are now shaping product design itself, not just pricing sheets, evidence that the shortage has moved from a supply chain nuisance to an engineering constraint.
James Sanders, an analyst at TechInsights, gave The Register a more measured timeline: DRAM pricing likely won’t peak before 2026, will “settle” somewhat in 2027, then rise again in 2028. He attributed the mismatch to a historically bad three-to-five-year fab buildout cycle colliding head-on with the AI demand surge.
Who This Actually Affects
If you’re a CTO, infrastructure lead, or procurement manager with Vera Rubin or Grace Blackwell orders scheduled for early 2027, this changes your math today, not next quarter.
Re-run your TCO models now. A 15%+ increase on flagship rack pricing is large enough to flip a build-versus-rent decision that looked settled a month ago.
Scale matters more than ever. Hyperscalers with long-term supply agreements can lock in memory allocation. Smaller AI teams without that leverage face both higher prices and lower priority in TrendForce’s documented allocation hierarchy.
This isn’t a 2026 blip. Mehrotra and Kwak, the two executives closest to actual memory supply, both point to tightness lasting through at least 2027 and 2028. Budgeting on 2025-era per-rack costs is no longer defensible.
The Case Against “Shortage Forever”
Not everyone buys the supercycle-forever narrative, and the article would be incomplete without the pushback.
Man Group’s institutional research argues the underlying AI technology is real, but the financial architecture funding it, including circular vendor financing and short-duration assets backed by long-duration debt, is expanding faster than any credible adoption curve justifies. If that thesis is right, today’s scarcity pricing could unwind quickly if hyperscaler capex growth slows.
Morgan Stanley analyst Joseph Moore has also pushed back on how demand is being counted, noting that non-binding “letters of intent” for memory capacity are sometimes conflated with firm orders in market narratives, and that no verified demand destruction has shown up yet despite several macro shocks over the past year.
Our read: this signals Nvidia’s 75% gross margin sits awkwardly next to a “we had no choice” framing. Redesigning Vera Rubin to use fewer memory modules looks a lot like margin protection layered on top of a genuine cost pass-through, not a pure one-to-one transfer of supplier pain.
Frequently Asked Questions
Why is Nvidia raising AI chip prices in 2026?
Nvidia is raising prices on servers containing its Vera Rubin and Grace Blackwell chips by more than 15% because memory chip costs from Samsung, SK Hynix, and Micron have surged amid an AI-driven supply shortage, according to Bloomberg’s August 22, 2026 report.
When will Nvidia’s price increases take effect?
The increases apply to systems shipped starting early 2027 and affect the Vera Rubin and Grace Blackwell platforms, with the exact amount depending on chip generation and memory configuration.
What’s causing the memory shortage behind this?
Memory makers have shifted wafer capacity toward high-bandwidth memory for AI accelerators, squeezing conventional DRAM and NAND supply. TrendForce recorded DRAM contract prices rising up to 95% quarter over quarter in Q1 2026 alone.
Will this affect cloud computing prices?
Likely yes. Contract manufacturers building servers for Microsoft, Google, and Oracle have already notified their own customers of coming increases tied to Nvidia’s price hike, which points toward higher enterprise cloud AI pricing through 2027.
How long will the memory shortage last?
Micron CEO Sanjay Mehrotra expects tight supply into 2028. SK Hynix CEO Kwak Noh-jung has said the market could stay undersupplied beyond 2030. Independent analyst James Sanders of TechInsights projects only a partial “settling” in 2027 before prices climb again in 2028.
What to Watch Next
Three things will tell you whether this pricing shock is temporary or structural. First, watch what Jensen Huang and Colette Kress say about margin on the August 26 earnings call, that’s the clearest signal of how much Nvidia plans to pass through versus absorb. Second, track TrendForce’s Q4 2026 DRAM contract pricing, due out roughly in October, for whether the rate of increase is actually slowing. Third, keep an eye on hyperscaler capex commentary in Q3 earnings; if Amazon, Google, Microsoft, or Meta so much as hint at pulling back 2027 budgets, the entire memory supercycle thesis gets tested in real time.
For now, the numbers are the numbers: a 15%+ jump on flagship AI server pricing, a consumer GPU market that already absorbed increases as high as 39%, and two memory-supplier CEOs both saying this doesn’t ease until at least 2028. Plan accordingly.
Subscribe to The Neural Loop for weekly breakdowns of what’s actually moving AI infrastructure costs, before the headlines catch up.
GitHub Copilot’s Pricing Reset Changes Coding for Beginners
You open GitHub Copilot for your fifth coding session this week and hit a wall you didn’t know existed: “You’ve used your free completions for this month.” That wall didn’t exist six months ago. GitHub quietly rebuilt its entire free tier around a hard cap of 2,000 code completions and 50 chat requests a month, and most of the “best AI coding tools for beginners” lists still circulating online haven’t caught up.
If you’re teaching yourself to code in 2026, that pricing shift is only half the story. The other half is a set of controlled studies, including one published by Anthropic, the company that sells Claude, showing that how you use AI while learning matters more than which tool you pick. Ask AI to hand you finished code and your comprehension can drop by double digits. Ask it to explain, review, and quiz you, and the picture looks very different.
This guide walks through what actually changed, what the research says about learning with AI, and which usage patterns keep you sharp instead of dependent.
On June 1, 2026, GitHub replaced its old “premium request” system with GitHub AI Credits, where one credit equals one cent. The change looks cosmetic on the surface. It isn’t. The old free tier was generous enough that most beginners never thought about limits. The new one hard-caps usage, and once you cross it, the tool simply stops helping until next month or until you upgrade.
Plan
Price
What You Get
Free
$0/mo
2,000 completions + 50 chat requests, Claude Haiku 4.5 and GPT-5 mini access, Copilot CLI
Pro
$10/mo
Unlimited completions, $15/mo in AI Credits, cloud agent, third-party agent access (Claude Code, Codex)
Pro+
$39/mo
$70/mo in credits, access to premium models including Opus
Max
$100/mo
$200/mo in credits, built for sustained agent workflows
Students get a built-in workaround worth knowing about: verified students receive free Copilot Pro access through the GitHub Student Developer Pack. Everyone else needs to budget for hitting that free-tier ceiling faster than expected, likely within a few weeks of daily practice rather than months.
Why this matters right now
Most “best AI tools for beginners” roundups still describe Copilot’s free tier as effectively unlimited. That description stopped being accurate on June 1, 2026. Budget $10 a month into your learning plan from day one instead of discovering the limit mid-project.
The Beginner Tool Landscape in 2026
GitHub Copilot isn’t the only entry point, and it isn’t automatically the right one for every beginner. Replit’s Agent can build a working app from a plain-English description with no prior coding knowledge at all, which makes it the fastest path to “I made something.” Cursor and Windsurf sit closer to Copilot: real code editors with inline AI explanations attached to every suggestion, better suited to someone who wants to actually read and understand the code being written.
None of these tools are mature or settled products sitting still. Mordor Intelligence sizes the AI code tools market at roughly $9.35 to $9.46 billion in 2026, projected to reach $22 to $30 billion by 2030 or 2031, a 26 percent compound annual growth rate. Pricing, free-tier limits, and model access will keep shifting under beginners’ feet for years, not months.
What the Research Says About Learning With AI
Here’s the part most beginner guides skip entirely. In January 2026, Anthropic researchers Judy Hanwen Shen and Alex Tamkin published a randomized controlled trial on exactly this question. Fifty-two mostly junior developers learned an unfamiliar Python library called Trio. One group used AI assistance. One group worked unaided. Both groups then took the same comprehension quiz.
The AI-assisted group scored 50 percent. The unaided group scored 67 percent. A 17-point gap on a same-day test.
“Participants in the AI group scored 17% lower than those who coded by hand, or the equivalent of nearly two letter grades.”
Judy Hanwen Shen & Alex Tamkin, Researchers, Anthropic
Notice what makes this finding unusual: Anthropic sells Claude Code. The company has every commercial incentive to publish research showing AI accelerates learning, not research showing it can undermine it. Anthropic’s own writeup of the study narrows the finding further: comprehension losses concentrated specifically in what the researchers call “AI Delegation,” asking the model to produce finished solutions, rather than in more supervised usage patterns like requesting explanations or reviewing generated code line by line.
Stack Overflow’s 2025 Developer Survey backs this up with adoption numbers. Among the “Learning to Code” segment specifically, 39.5 percent use AI tools daily and 18.7 percent weekly, both lower than the 50.6 percent and 17.4 percent figures for working professionals. Favorability sits lower too: 52.8 percent of learners rate AI tools favorably versus 61.2 percent of professionals, while 26.3 percent of learners report unfavorable views versus 19.7 percent of pros. Learners are, on one narrow measure, more trusting than professionals of AI output (6.1 percent report “high trust” versus 2.7 percent for pros), but that’s still a small minority either way. And 66 percent of all developers surveyed cite “AI solutions that are almost right, but not quite” as their top frustration, with 45.2 percent saying debugging AI-generated code takes longer than writing it themselves.
The speed argument doesn’t hold up well either, even for experienced developers. METR ran a randomized controlled trial in mid-2025 with 16 experienced open-source developers using AI tools, mostly Cursor Pro paired with Claude 3.5 and 3.7 Sonnet. The developers took 19 percent longer to finish real tasks with AI assistance than without it, despite predicting a 24 percent speedup beforehand, and despite believing after the fact that AI had made them 20 percent faster. One important caveat: METR’s own report measured experienced developers on familiar codebases, not beginners, and the organization now labels the result “historical,” tied to early-2025 tool capability. Still, the gap between predicted and measured performance is a useful check against vendor productivity claims.
The Junior Job Market Beginners Are Entering
There’s a labor-market backdrop to all of this that most tool comparisons leave out entirely, and it isn’t speculative. Stanford’s Digital Economy Lab tracks millions of workers through actual ADP payroll data, not surveys or job postings. Their most recent update, dated August 2026, found employment for workers aged 22 to 25 in the most AI-exposed occupations, including software engineering, sitting 19 percent below where it would have landed had it tracked their less-exposed peers. That gap has widened at every update since it was first documented.
Not everyone in the industry agrees on what that means. Erik Brynjolfsson, director of the Stanford Digital Economy Lab, frames it as a diverging-paths story rather than mass job destruction.
“I think it’s fair to say that technology has always been destroying jobs and always been creating jobs.”
Erik Brynjolfsson, Director, Stanford Digital Economy Lab
AWS CEO Matt Garman takes an even more pointed stance against the idea that AI erases the need for junior hires, a position he’s stated publicly on more than one occasion.
“I was like that’s the like one the dumbest thing I’ve ever heard.”
Matt Garman, CEO, Amazon Web Services
He continued: if a company has no talent pipeline and no junior people being mentored up through the code, “at some point that whole thing explodes on itself.” Garman’s comments, first reported in an August 2025 podcast interview and reaffirmed in a December 2025 WIRED interview covered by Fortune, run directly counter to the narrative that junior developer roles are becoming obsolete.
Our read: neither the payroll data nor the executive pushback cancels the other out. The market is genuinely tighter for entry-level, AI-exposed roles right now, and simultaneously, at least one major cloud CEO is on record saying companies that stop training juniors are setting themselves up to fail later. Both things are true at once, and a beginner planning a job search needs to hold both.
How to Actually Use AI Tools Without Skipping the Learning
So what does a beginner actually do with all this? Not “avoid AI.” The Anthropic researchers were careful to isolate which usage pattern caused the comprehension gap, and it wasn’t AI use in general. It was delegation specifically: asking for a finished answer instead of working through the problem first.
Attempt first, then compare. Write your own version of the solution before asking AI for one. Comparing your approach to the AI’s output builds the same kind of retrieval practice that improves comprehension test scores in the Anthropic study.
Ask for explanations, not just code. Prompting for “explain why this works” instead of “write this for me” keeps you in the supervised-usage category the research associates with smaller comprehension losses.
Budget for the free-tier wall. Plan on hitting Copilot’s 2,000-completion cap within weeks of regular use, and decide in advance whether you’ll pay $10 a month or switch tools when you do.
Treat interviews as AI-free zones. Practice explaining and debugging code without assistance regularly. Technical interviews, on-call incidents, and code review are exactly the moments AI assistance is least reliably available.
Build a portfolio that shows your thinking, not just working output. Given the current entry-level hiring gap, projects that demonstrate independent debugging and design decisions carry more weight than a working app you can’t fully explain.
This isn’t a new problem in education. It’s the calculator and spellchecker debate from earlier decades, playing out again with sharper tools and, this time, controlled data instead of just opinions. The framing that holds up best across every source in this piece isn’t “should beginners use AI.” It’s “which usage pattern preserves the learning,” and Anthropic’s own research draws that line clearly.
GitHub Copilot, Replit, Cursor, and Windsurf are the most-recommended entry points in 2026 because each pairs a free tier with plain-language chat rather than requiring memorized syntax. Replit’s Agent can build a working app from a plain-English description with zero prior coding knowledge, while Copilot and Cursor attach explanations to inline code suggestions inside a real code editor.
Is GitHub Copilot free for beginners?
Yes, but with real limits. GitHub Copilot Free includes 2,000 code completions and 50 chat requests per month, no credit card required. That structure took effect after GitHub’s June 1, 2026 shift to usage-based AI Credits billing, replacing a more generous earlier free tier.
Can AI teach me to code from scratch?
AI can meaningfully lower the barrier to writing your first working program, but a January 2026 Anthropic study found learners who leaned on AI to generate code scored 17 percentage points lower on same-day comprehension tests than those who coded by hand, suggesting AI works best as an explainer and reviewer rather than a first-draft generator for beginners.
Will AI replace the need to learn to code?
No major analyst, academic study, or company statement supports that claim. AWS CEO Matt Garman has publicly called the idea of skipping junior-level hiring and training the dumbest thing he’s heard, and Stack Overflow’s 2025 survey shows even the learning-to-code cohort still trusts AI output less than half the time.
Is it harder to get a junior developer job because of AI?
Verified payroll data says yes, directionally. Stanford’s Digital Economy Lab found employment for 22 to 25-year-olds in AI-exposed occupations, including software engineering, sits 19 percent below trend as of mid-2026, a gap that has widened continuously since it was first documented.
Where This Goes Next
You now know something most competing guides still get wrong: Copilot’s free tier isn’t the safety net it used to be, and the “just use AI to learn faster” advice floating around most beginner content isn’t backed by the controlled research that actually exists on the question. Delegation hurts comprehension. Supervised use, where you attempt first and use AI to explain and check, doesn’t show the same drop.
Watch three things over the next 6 to 18 months: whether GitHub’s usage-based billing model spreads to competitors like Cursor and Windsurf, whether Stanford’s entry-level employment gap keeps widening or starts to close as more juniors adapt their AI usage patterns, and whether more AI labs follow Anthropic’s lead in publishing skill-formation research rather than pure productivity claims.
Want the next update on AI coding tools, pricing shifts, and skill-formation research before it hits the mainstream feeds? Subscribe to The Neural Loop, NeuralWired’s newsletter for builders who want the primary sources, not the recycled hot takes.
The short answer: On August 17, 2026, Nvidia filed an SEC 8-K guaranteeing up to $105 billion in lease and power obligations for OpenAI’s new Ohio data center. The guarantee only pays out if OpenAI defaults or goes insolvent, and it covers 4.25 gigawatts of an eventual 8 gigawatt campus built on a former Cold War uranium site.
Jensen Huang spent Sunday on X insisting his company isn’t running a circular financing scheme. That’s not the kind of thing a CEO tweets when nobody’s asking the question. The Nvidia $105 billion OpenAI guarantee, disclosed the same day in a Form 8-K filed with the SEC, is the largest single financial backstop Nvidia has ever put its name on, and it lands squarely on top of a company, OpenAI, that lost $1.22 for every dollar it brought in during the first quarter of 2026.
If you cover semiconductors, AI infrastructure, or anything adjacent to hyperscaler capital spending, this filing is now required reading. Here’s what Nvidia actually signed up for, why the number dropped from an earlier $250 billion figure, and where the real risk sits.
Strip away the SEC language and the structure is fairly simple. SB Energy, a subsidiary of Japan’s SoftBank Group, is building a massive data center campus in Pike County, Ohio, called the PORTS-Pike Technology Campus. SB Energy will own and operate the site. An OpenAI affiliate will lease it for 20 years starting in 2028. Nvidia becomes the exclusive AI compute provider to the campus, with limited exceptions, according to the 8-K filing on SEC EDGAR.
Nvidia’s role is what’s new here. The company has agreed to what its own filing calls “residual value guaranties,” meaning Nvidia will cover the lease and power payments if OpenAI can’t. That obligation is capped at $105 billion, cumulative, across the initial 4.25 gigawatts of IT load. It only becomes a real cash outflow if OpenAI defaults on the lease or becomes insolvent.
Separately, and this distinction matters more than most headlines have made clear, Nvidia is putting $1.5 billion of direct equity into SB Energy itself, described in Nvidia’s release as support for the company’s “evolution into a leading AI infrastructure developer,” per Axios’s reporting. That $1.5 billion is a real, near-term check. The $105 billion is a ceiling that only gets hit if things go wrong.
Why The Guarantee Shrank From $250 Billion To $105 Billion
The Wall Street Journal first reported a proposed backstop of up to $250 billion on August 14, three days before the final filing. Nvidia shares dropped as much as 5% on that report, a clear signal that investors weren’t thrilled about the size of the exposure. By the time the deal was finalized and filed with the SEC on August 17, the number had been cut by more than half, to $105 billion, and scoped down to cover only the campus’s initial phase rather than the full 10 gigawatt buildout planned for the site.
That’s the headline version. The more interesting version is that the cut may be optical rather than structural. CNBC’s same-day reporting noted that Nvidia and OpenAI are separately discussing a financing arrangement of up to $350 billion to fund the actual chip purchases for the site, a deal that has not been confirmed in any SEC filing as of this writing. If that arrangement materializes, Nvidia’s combined exposure to a single customer could end up higher than the original $250 billion figure that spooked the market in the first place. Worth flagging clearly: that $350 billion number is reported, not confirmed.
Inside The Portsmouth Site
The location has its own story. The PORTS-Pike Technology Campus sits on the site of the former Portsmouth Gaseous Diffusion Plant, a decommissioned Cold War uranium enrichment facility roughly 50 miles south of Columbus. Powering an AI campus where the government once enriched uranium for weapons programs is the kind of detail that writes its own headline.
Getting power to the site is its own undertaking. SB Energy and AEP Ohio are jointly investing at least $4.2 billion in transmission infrastructure, including new 765-kV lines and four substations, funded through the project itself rather than passed on to ratepayers. The total site is planned for 10 gigawatts of power draw, including 9.2 gigawatts of new gas-fired generation. OpenAI says the buildout will support 35,000 construction jobs through 2032 and roughly 2,500 permanent operating positions once complete.
The Deal By The Numbers
Figure
What it represents
$105 billion
Cumulative cap on Nvidia’s guaranty, down from an earlier $250 billion figure
4.25 GW
IT load covered in phase one, out of an eventual 8 GW campus
$1.5 billion
Nvidia’s direct equity stake in SB Energy, separate from the guaranty
$4.2 billion
SB Energy and AEP Ohio’s combined transmission infrastructure spend
$81.6 billion
Nvidia’s Q1 FY2027 revenue, up 85% year over year
$852 billion
OpenAI’s post-money valuation as of its March 2026 funding round
-122%
OpenAI’s non-GAAP operating margin in Q1 2026
$63 billion
OpenAI’s projected cash burn for 2027
Put those last two rows next to each other and the reason Nvidia needed to guarantee anything becomes obvious. A tenant with an $852 billion valuation but no investment-grade credit rating and a widening cash burn is exactly the kind of counterparty landlords ask for backstops on.
Is This Circular Financing?
This is the question every analyst note on this deal opens with, and Jensen Huang got ahead of it himself.
“Is this circular financing? No. OpenAI will pay the lease.”
Jensen Huang, Founder & CEO, Nvidia Corporation · posted to X, August 17, 2026
Huang’s argument is that Nvidia is using its balance sheet strength to secure long-lived infrastructure that OpenAI will pay to occupy, not manufacturing demand for its own chips out of thin air. He’s also floated a much bigger number: roughly $600 billion in Nvidia compute opportunity through 2030, tied to OpenAI’s broader buildout plans. That figure is a projection, not a contract, and should be read that way every time it shows up in a headline.
Not everyone is buying the framing. Michael Burry, the investor best known for his short position ahead of the 2008 crash, has been naming this exact deal in his recent writing.
“Circular financing lets capital injected into the AI ecosystem flow back to participants as revenue, while debt makes up a growing share of that capital, which puts the bubble on a clock.”
Michael Burry, Scion Asset Management · Trading Post, Substack, August 13, 2026
Burry has also pointed to roughly $879 billion in hyperscaler commitments that flow back through Nvidia in one form or another, and noted that Nvidia’s credit default swap spread doubled over a two month stretch as bond traders started pricing in this kind of exposure.
Sell-side analysts land somewhere in the middle. Bernstein’s Stacy Rasgon has warned that the sheer size of Nvidia’s guarantees, larger than anything the company has previously disclosed, will “fuel these worries much hotter than what we have seen previously.” CreditSights, a fixed-income research firm, put it more bluntly: the structure is “pro-cyclical,” nearly free to Nvidia while the market is hot, and most dangerous in a downturn, when customers are defaulting at the same time hardware values are falling. Their phrase for it: Nvidia is effectively “writing a put.”
Our read: both things can be true at once. Nvidia probably does get paid the lease under most scenarios. But “most scenarios” isn’t the same as “all scenarios,” and $105 billion is a lot of money to have riding on one customer’s ability to keep growing into an $852 billion valuation it hasn’t earned yet on paper.
The Skeptics’ Case
Set aside the circular financing framing for a moment. There’s a separate, quieter argument building among finance academics and rating agencies that’s less about accusation and more about accounting.
NYU Stern’s Aswath Damodaran, whose valuation work is widely cited across Wall Street, has argued that the big AI hyperscalers have effectively become manufacturing companies dressed in software multiples.
“They now are the equivalent of manufacturing companies. And like all manufacturing companies historically, they’re now going to be judged on whether they can deliver the earnings on this investment.”
Aswath Damodaran, Professor of Finance, NYU Stern School of Business · ProfG Markets, August 7, 2026
That’s a return-on-invested-capital argument, and it applies with more force to OpenAI, the tenant with the cash burn problem, than to Nvidia, the guarantor with the $81.6 billion quarterly revenue base. But it applies to Nvidia too, indirectly: every dollar committed as a guaranty is a dollar of balance sheet capacity that isn’t available for something else.
There’s also a bank-for-central-banks-level warning sitting underneath all of this. The Bank for International Settlements flagged in its June 2026 Annual Report that hyperscaler debt tied to AI buildouts is growing faster than the balance sheets carrying it, a systemic concern rather than a single-company one. And Nvidia’s own filing doesn’t exactly dodge the characterization. The 8-K classifies the guaranty under Item 2.03, “Creation of a Direct Financial Obligation or an Obligation under an Off-Balance Sheet Arrangement,” which is Nvidia’s own language, not a reporter’s spin. Rating agencies have already started treating comparable structures this way. S&P Global has said it will fold Broadcom’s similar residual-value guarantees into its adjusted debt calculations, and there’s no obvious reason Nvidia’s guaranty would be treated differently once the details land in Nvidia’s next 10-Q.
And that’s the honest gap in this story right now: Nvidia hasn’t yet disclosed the guarantee’s trigger conditions, per-lease minimums, or how the $105 billion cap gets allocated across leases. Those details are expected as exhibits to Nvidia’s Form 10-Q for the fiscal quarter ended July 26, 2026. Until that filing lands, a lot of the risk modeling here is still an estimate built on the topline number alone.
What Happens Next
Three things are worth watching over the next 12 to 18 months.
The 10-Q exhibits. Nvidia’s next quarterly filing should finally show the trigger conditions and allocation formula behind the $105 billion cap. That’s when analysts can actually model this instead of estimating around it.
OpenAI’s IPO window. OpenAI confidentially filed a draft S-1 in June 2026, with a possible listing as early as September at a valuation reportedly approaching $1 trillion. A weak public debut would tighten OpenAI’s ability to fund lease payments without leaning on Nvidia’s guaranty.
The $350 billion chip financing talks. If that separate arrangement gets confirmed in a filing, it changes the real size of Nvidia’s total exposure to OpenAI, regardless of what today’s $105 billion headline suggests.
The first phase of the Ohio campus, around 800 megawatts of the initial 4.25 gigawatt commitment, is targeted to come online in 2028. Building gigawatt-scale gas power and a data center shell in two years is an aggressive timeline by utility standards. Nvidia’s “land, power, and shell” approach is designed to decouple the site build from hardware generations, which helps with obsolescence risk, but it doesn’t do anything to change the financing timeline underneath it.
Frequently Asked Questions
What did Nvidia agree to guarantee for OpenAI’s Ohio data center?
On August 17, 2026, Nvidia filed an SEC 8-K disclosing it will guarantee up to $105 billion in lease and power payment obligations for OpenAI’s data center in Pike County, Ohio. The guarantee covers 4.25 gigawatts of an eventual 8-gigawatt campus and pays out only if OpenAI defaults or becomes insolvent.
Is the Nvidia-OpenAI deal circular financing?
Nvidia CEO Jensen Huang has publicly denied it, saying OpenAI will pay the lease itself. Critics including investor Michael Burry and Bernstein analyst Stacy Rasgon argue the structure still lets Nvidia’s capital effectively support demand for its own chips, since Nvidia is guaranteeing debt tied to a facility built to run its hardware exclusively.
Where is OpenAI’s new Ohio data center located?
The PORTS-Pike Technology Campus sits in Pike County, Ohio, on the site of the former Portsmouth Gaseous Diffusion Plant, a decommissioned uranium enrichment facility about 50 miles south of Columbus. SB Energy, a SoftBank subsidiary, will build and operate it under a 20-year lease to OpenAI.
When will OpenAI’s Ohio data center be operational?
The first phase, roughly 800 megawatts of the initial 4.25-gigawatt commitment, is expected online in 2028. The full 8-gigawatt campus would follow in later phases through the early 2030s.
Why did Nvidia’s guarantee shrink from $250 billion to $105 billion?
The Wall Street Journal first reported a proposed $250 billion backstop on August 14, 2026, and Nvidia shares fell as much as 5% on the news. The finalized August 17 SEC filing capped Nvidia’s guaranty at $105 billion, covering only the campus’s initial phase rather than the full 10-gigawatt buildout.
Does Nvidia’s OpenAI guarantee affect its balance sheet or credit rating?
The guarantee is structured as an off-balance-sheet obligation, but Nvidia’s own 8-K classifies it under rules governing direct financial obligations. Rating agencies including S&P Global have said they treat comparable residual-value guarantees, such as Broadcom’s, as debt-like obligations in adjusted debt calculations, which suggests similar scrutiny could apply here.
Where This Leaves You
Here’s what’s actually changed after this filing. Nvidia no longer needs OpenAI to buy more chips to grow. It now needs OpenAI’s Ohio lease payments to keep flowing for the next twenty years, or it needs to be comfortable writing a check as large as $105 billion if they don’t. Those are two different kinds of exposure, and the market has spent the past week trying to figure out which one it’s actually pricing.
Watch the 10-Q exhibits for the real trigger mechanics, watch OpenAI’s IPO timeline for the revenue side of the equation, and watch whether that separate $350 billion chip financing talk turns into an actual filing. Any one of those three could change how this deal reads in six months.
Qwen’s 3 Billion Download Claim vs. the Real Hugging Face Number
Open Source AI · Data Report
Qwen’s 3 Billion Downloads: What Hugging Face Actually Found
By NeuralWired Staff · Published August 16, 2026 · 9 min read
Alibaba says its Qwen models just crossed 3 billion downloads, beating Meta and Google combined. The number making headlines this week comes from a company press statement. The number that came from an independent audit, published one day earlier by Hugging Face, is 2.045 billion. Nobody covering this story has reconciled the two, and the gap tells you more about how AI companies market themselves in 2026 than either figure does on its own.
If you’re a developer, CTO, or ML lead deciding which open model family to build on, the headline number is the least useful part of this story. The methodology behind it, and what Hugging Face’s full report says about where Qwen’s lead actually comes from, matters a lot more.
On August 14, 2026, Hugging Face published its biannual State of Open Models: Summer 2026 Observations report, a survey of Hub activity from January through July authored by staff researchers Adina Yakefu, Apolinário Passos, Irene Solaiman, and roughly 70 contributors. Its number for Qwen: 2,045,000,000 downloads on the Hugging Face Hub, against 418 million for Google and 227 million for Meta over the same window.
One day later, Alibaba sent out an emailed statement, first reported by Bloomberg and syndicated by Business Standard, claiming Qwen had passed 3 billion downloads globally across 460-plus open-sourced models, with 300,000-plus derivative models built on top of them. That figure folds in ModelScope, Alibaba Cloud, and third-party mirrors, none of which Hugging Face’s report can see or verify.
Measurement
Qwen
Google
Meta
Hugging Face Hub (independently logged, Jan-Jul 2026)
The core problem: Every major outlet that covered this story, Fortune, Bloomberg, China Daily, ran the 3 billion figure and the 2.045 billion figure in the same breath, as though they measured the same thing. One is server-side telemetry from a neutral platform. The other is a company’s own count, with no disclosed methodology, covering channels nobody outside Alibaba can audit.
That distinction matters because Hugging Face’s own report contains a direct warning against the interpretation most coverage encouraged. Its methodology notes state plainly that downloads reflect Hub activity, not API usage, private deployments, or distribution through other channels, and should not be read as a proxy for model quality or market share. Almost none of the news coverage repeated that caveat.
What Hugging Face’s Report Actually Measured
Strip away the 3-billion headline and the audited numbers still tell a real story. Qwen’s ecosystem depth, not just its raw download count, is where the report gets interesting.
151,448 Qwen-based derivative models exist on the Hub, roughly 2.6 times Meta’s total derivative count across all its models and 4.7 times Llama’s derivative count specifically. Google’s Gemma family trails with 82,506 derivatives.
New Qwen derivatives are appearing at 180 to 210 repositories per day, sustained through the first seven months of 2026.
Of 28,531 GGUF conversions (the quantized format that lets Qwen run locally on consumer hardware) only 54 came from the Qwen team itself. The rest is unpaid community work.
Qwen pulls 39.6 million GGUF downloads a month for local, on-device inference, nearly double Gemma’s 20.8 million and more than five times Llama’s 7.5 million.
Hugging Face’s own researchers were careful to credit the right party for that lead:
“This position was built largely by the community.”
Hugging Face research team, State of Open Models: Summer 2026 Observations, Aug 14, 2026
Read that sentence again next to Alibaba’s press release. The derivative count, the GGUF conversions, the documentation, most of the infrastructure that makes Qwen usable on a laptop instead of a data center rack, came from developers who don’t work for Alibaba and were never asked to.
The Headline Hides a Small-Model Story
Here’s the number that should reframe the entire “Qwen beat Meta and Google” narrative: 83% of all-time downloads across the entire Hugging Face Hub go to models under 1 billion parameters. And 1.5% of all repositories account for 99.2% of total downloads.
Translation: this isn’t really a story about frontier reasoning models slugging it out for AGI supremacy. It’s a story about which company ships the widest range of small, boring, deployable utility models, the kind that get embedded into a search pipeline or a classification task and never make headlines. Qwen’s flagship 2.4-trillion-parameter Qwen3.8-Max, released July 19, 2026 with 95 billion active parameters per query, is impressive engineering, but it is not what most of those 2.045 billion downloads are for.
If your team is benchmarking frontier capability, Hub download share is close to irrelevant. If your team is trying to figure out where the community troubleshooting, quantized builds, and tooling density will actually be a year from now, it’s the most useful number in the report.
The License Reversal Almost Nobody Is Covering
This is the part of the story that got buried under the download headline, and it’s the part that should worry anyone planning to build a commercial product on the assumption that Qwen stays free forever.
Qwen3.7-Plus, unlike earlier releases in the family, shipped without open weights, a detail first flagged in technical discussion on Hacker News rather than in mainstream coverage. Multiple outlets, citing unnamed sources, now report Alibaba is preparing a revenue-sharing license for Qwen3.8-Max aimed at large commercial users, a structural shift away from the fully permissive Apache 2.0 approach that built the download lead in the first place. It would mirror a move already made by rival Chinese lab Moonshot AI, whose Kimi K3 model requires authorization above $20 million in annual revenue, a licensing detail a Hugging Face Hub community member flagged as inconsistent with the report’s claim that no 20B-plus Chinese release carries non-commercial restrictions. Hugging Face has not issued a correction.
Alibaba’s own researchers, in a January 2026 statement carried by state outlet Xinhua, framed the company’s intent around continued openness:
“…keep pushing the performance frontier of LLMs…”
Unnamed Qwen team researcher, Tongyi Lab, via Xinhua, Jan 13, 2026
Whether that commitment survives contact with a revenue-sharing license for the flagship model is an open question, and one that the “3 billion downloads” framing this week conveniently sidesteps. If Alibaba confirms the shift around its August 20 earnings call, every “Qwen wins open source” piece published this week needs a follow-up within days.
The US-China Framing Problem
Coverage of Chinese open-weight models rarely stays purely technical for long, and this story is happening against a backdrop of congressional scrutiny into Chinese AI generally. Independent technology writer Karl Bode has been one of the more pointed critics of that framing, arguing that national security concerns raised about Chinese open models function to protect incumbent commercial interests more than they reflect a substantiated threat, describing the pattern as built on “a fake concern for national security.”
That’s one side of the argument. It’s not the only one. Anthropic and OpenAI have separately accused Chinese open-weight developers of unauthorized model distillation, a claim distinct from the download-count story but part of the same broader tension over how open the “open” in open-weight Chinese models really is, as detailed in The Conversation’s August 2026 analysis. Readers evaluating the download headline should hold both positions in mind rather than picking whichever confirms an existing view of Alibaba.
What This Means If You’re Building on Qwen
For engineering teams actually shipping product, three things from this report matter more than the topline number.
1. Community tooling really is denser around Qwen
Five times the local-inference download volume of Llama and nearly double Gemma’s isn’t a vanity metric. It means more GGUF builds, more Discord and GitHub troubleshooting threads, and more prebuilt quantizations to pull from when something breaks at 2 a.m.
2. Audit your license before you scale
Apache 2.0 legacy Qwen models are unaffected by anything reported here. But if you’re on a newer flagship variant, or planning to be, check the license terms attached to that specific model version now, not after you’ve built a revenue-generating product around the assumption that it’s free forever.
3. Don’t confuse Hub downloads with frontier capability
With 83% of downloads going to sub-1B models, a high Qwen download count tells you almost nothing about how a 2.4-trillion-parameter Qwen3.8-Max will perform against GPT or Gemini on your specific reasoning task. Benchmark separately.
Frequently Asked Questions
How many downloads does Qwen actually have?
Hugging Face independently measured 2.045 billion Qwen downloads on its Hub for 2026. Alibaba separately claims 3 billion-plus across all distribution channels, including ModelScope and Alibaba Cloud, using a methodology it hasn’t disclosed.
Is Qwen better than Llama?
Qwen leads Llama by a wide margin in Hub downloads and derivative models built on top of it. “Better” still depends on your use case; benchmark against your own task rather than relying on download share as a quality signal.
Is Qwen open source or open weight?
Most Qwen releases use the permissive Apache 2.0 license. Recent exceptions exist: Qwen3.7-Plus shipped without full open weights, and Alibaba is reportedly moving its newest flagship model toward a revenue-sharing license for large commercial users.
Why does Qwen have more downloads than Google or Meta?
Hugging Face credits a wider size range of published models, a faster release cadence, and permissive licensing terms, which together encouraged heavy community-driven derivative and quantization work that Qwen’s own team didn’t have to build itself.
Where This Goes Next
Two numbers came out forty-eight hours apart this week, and only one of them was audited. That doesn’t make Alibaba’s claim false, but it does mean the “Qwen beat Meta and Google” story running across tech media right now is built on a company press release stacked next to an independent report, presented as if they’re interchangeable.
What we now know for certain: Qwen’s community-built infrastructure lead is real and independently verified. What we don’t know: whether the licensing terms that built that lead survive the next flagship release. Watch three things over the next six to eighteen months: Alibaba’s August 20 earnings commentary on AI monetization, whether Qwen3.8-Max ships under the reported revenue-sharing terms, and whether Hugging Face’s next report shows the derivative growth rate holding at 180 to 210 repositories a day or slowing as licensing tightens.
Want stories like this before they hit the feed? Subscribe to The Neural Loop at neuralwired.com/newsletter.
Gemini 3.7 Flash: Half Price Now, Full Price in 2027AI & Enterprise Tech
Gemini 3.7 Flash Is Half Price. Read the Footnote First.
By NeuralWired Staff | Published August 15, 2026
Google DeepMind shipped Gemini 3.7 Flash on August 13, 2026, its fourth Flash-tier model in nine weeks, at half the price of its predecessor. That discount expires December 31, 2026. And the model everyone actually asked for at I/O in May, Gemini 3.5 Pro, still hasn’t shipped.
If you’re choosing a model for coding agents or budgeting inference spend into 2027, both of those facts matter more than the launch headline. Here’s what Google’s own numbers say, what independent testing confirms, and what the pricing footnote is quietly telling you.
Gemini 3.7 Flash went generally available in the Gemini API, Google AI Studio, Vertex AI, and Antigravity, Google’s coding-agent platform, on August 13, 2026, according to Google’s own Gemini API release notes. It also now powers Gemini Spark, Google’s productivity agent.
That’s four Flash-tier releases since Gemini 3.5 Flash debuted at I/O in May: 3.5 Flash, then 3.5 Flash-Lite and 3.6 Flash together on July 21, then 3.7 Flash on August 13. Twenty-three days between the last two. Google says the speed comes from algorithmic improvements, not a bigger base model.
Google AI Studio product lead Logan Kilpatrick framed it as a fast, targeted push rather than a ground-up rebuild.
“A strong intelligence increase, delivered in roughly three weeks through algorithmic work across Google DeepMind teams, focused on making the model feel more usable for real work.”
Logan Kilpatrick, Product Lead, Google AI Studio, Google DeepMind, August 2026 launch announcement
Google’s own positioning line calls it “our most intelligent workhorse model yet for coding and agents.” Positioning aside, the two numbers worth caring about are what it costs and what it can actually do, and the answers to both come with asterisks.
The Pricing Trap: Half Price Until January 1
Gemini 3.7 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens. That’s exactly half of what Gemini 3.6 Flash charges. It’s a genuinely good rate, and it’s live right now.
The part most launch-day coverage skipped:
This pricing runs through December 31, 2026 only. On January 1, 2027, Gemini 3.7 Flash reverts to $1.50 input / $7.50 output per million tokens, the identical permanent rate Gemini 3.6 Flash has charged since its own July launch. The “half price” headline is a five-month promotional window, not a durable cost advantage.
If your team is modeling 2027 inference spend on today’s rate card, that model is wrong by roughly 2x. Build your cost projections around $1.50/$7.50, not $0.75/$3.75, for anything shipping past year-end.
There’s a second cost change buried in the same release: Google removed the “minimal” thinking tier, the cheap setting older Flash models used for high-volume classification work. “Low” is now the floor, and thinking tokens bill at the output rate even though the API only returns a summary of that reasoning. If your pipeline leaned on minimal-tier Flash for bulk, low-stakes calls, re-benchmark it. The effective cost floor just moved up even as the headline price moved down.
Gartner analyst Will Sommer flagged exactly this pattern months before this launch, and it applies directly here.
“Chief Product Officers should not confuse the deflation of commodity tokens with the democratization of frontier reasoning.”
Will Sommer, Senior Director Analyst, Gartner, LLM inference economics forecast, March 2026
Sommer’s point is sharper once you factor in agentic workflows, which is exactly what Google is tuning 3.7 Flash for. Agent loops can multiply token consumption 5 to 30 times per task compared to a single chat completion. Cheaper tokens don’t necessarily mean a cheaper bill when the model is calling itself in a loop.
What Google’s Own Benchmarks Really Show
The capability jump from 3.6 Flash to 3.7 Flash is real, per Google’s own evals methodology page. DeepSWE v1.1 climbed from 49.0% to 65.3%. AutomationBench nearly doubled, from 17.0% to 30.4%.
Independent testing backs up at least the speed claim. Artificial Analysis clocked 3.7 Flash at 340.1 tokens per second in its standardized benchmarking workload, the fastest model the firm currently measures.
But read Google’s own head-to-head comparison table against GPT-5.6 Terra in full, not cherry-picked, and the frontier-leadership framing softens fast.
Notice which four rows it loses: the hardest agentic and terminal-use benchmarks, the exact category Google is marketing this model for. It’s a real value trade-off, not a clean win, and it’s Google’s own chart saying so.
There’s also a small but telling inconsistency worth flagging. Google’s 3.6 Flash model card lists its own DeepSWE v1.1 score as 48.6%. The 3.7 Flash launch blog rounds that same baseline to 49.0%. Minor on its own, but it’s a reminder that even single-vendor self-reported numbers are worth cross-checking against the vendor’s other documents, not just against competitors.
METR, the group that runs independent AI capability evaluations, has warned about a broader version of this problem.
“Benchmarks run without live human interaction can cause models to fail at tasks they could complete with minimal human guidance, making benchmarks unreliable proxies for real capability.”
METR, Experienced Developer Study, July 2025
Translation for anyone building on this: treat DeepSWE and AutomationBench jumps as lab signals worth investigating, not as production-readiness guarantees. Run your own workload against it before you migrate.
The Elephant in the Room: Gemini 3.5 Pro
None of the Flash-tier sprint makes sense without the model that isn’t here. At I/O in May, Sundar Pichai told developers to give Google “until next month” for Gemini 3.5 Pro, implying a June release. It didn’t happen. As of this article’s publication, it still hasn’t.
Bloomberg reported on July 16, citing ten current and former Google employees, that 3.5 Pro was running months behind schedule, largely over coding-capability shortfalls. A late-June training-data update meant to fix that reportedly made results worse, not better. Later reporting sourced to the same chain indicates the problems ran deeper than a bad update: DeepMind concluded the original 3.5 Pro base model had structural failures in recursive tool-calling and SVG generation, scrapped it, and restarted pretraining from a native Gemini 3 foundation. That same reporting says DeepMind has already begun pretraining an entirely new flagship, Gemini 4, mentioned almost in passing in the July 21 announcement.
Pichai himself gave the first public crack in the story, back in May.
“A bit behind on agentic coding.”
Sundar Pichai, CEO, Alphabet/Google, remarks at Google I/O, May 2026
Kilpatrick’s current line on 3.5 Pro is that the team is “testing with partners” and hopes to “land it soon.” That “soon” has now stretched past a second informal window with no date attached.
Wall Street has already priced in the uncertainty. Alphabet shares fell roughly 4.4% the day the Bloomberg delay report landed, an estimated $200 billion in market cap, on top of an earlier ~$225 billion drop in June tied to senior DeepMind researchers leaving for Anthropic and OpenAI. Combined, that’s close to $425 billion in Alphabet market value lost since late June with no change to reported revenue or earnings. Alphabet’s Q1 2026 results were strong (Google Cloud revenue up 63% year over year to $20 billion), which makes the point sharper: this is a narrative problem right now, not yet a fundamentals problem.
Our read: shipping four Flash models in nine weeks while the flagship reasoning tier stalls out looks less like a coincidence and more like a deliberate holding pattern, cover the volume segment on cost and speed while the harder model gets rebuilt underneath it. Google hasn’t confirmed that as strategy though, and it’s worth treating that framing as the most defensible inference from public facts, not as a confirmed internal decision. It could just as easily be ordinary engineering triage under deadline pressure.
What EU and UK Teams Need to Know
Buried in the model card, not the launch announcement, is a jurisdictional exclusion: Gemini 3.7 Flash is not available on the only consumer-facing surface it runs on in the EEA, UK, Switzerland, and Nigeria.
The timing isn’t nothing. The European Commission’s enforcement powers over general-purpose AI providers under the EU AI Act activated on August 2, 2026, penalties up to €15 million or 3% of global annual turnover, whichever is greater. Eleven days later, Google’s newest consumer AI model quietly excludes those exact jurisdictions from that surface. Google hasn’t stated a causal link publicly, but if you’re evaluating this model for an EU-facing product, plan around the exclusion now rather than discovering it in deployment.
The Bottom Line for Engineering Teams
If you’re on Gemini 3.6 Flash today, this is a real upgrade at a genuinely good price, for now. Three things to actually do with that:
Model your 2027 costs at $1.50/$7.50, not $0.75/$3.75. The current rate expires December 31, 2026.
Re-benchmark anything that ran on the old “minimal” thinking tier. It’s gone, and “low” now bills thinking tokens at the output rate.
Don’t lock a roadmap to a Gemini Pro milestone right now. Google has shipped zero Pro-tier models since Gemini 3 Pro in November 2025, despite promising 3.5 Pro for June 2026.
The broader question is whether this efficiency pivot holds. If Gemini 4, reportedly already in early pretraining, also slips, Google will have gone potentially 18 months or more between flagship releases while Anthropic and OpenAI keep a faster cadence. That’s a gap that compounds on reputation even if Flash-tier usage and revenue stay healthy in the meantime. Our related coverage on why enterprise AI inference costs aren’t actually falling and on Google’s agentic AI enterprise adoption gap both dig further into the pieces of this story we didn’t have room for here.
FAQ: Gemini 3.7 Flash and Gemini 3.5 Pro
Is Gemini 3.7 Flash actually cheaper than Gemini 3.6 Flash?
Only through December 31, 2026. Gemini 3.7 Flash launched at $0.75/$3.75 per million tokens, half of 3.6 Flash’s rate, but on January 1, 2027 it reverts to $1.50/$7.50, the same permanent rate 3.6 Flash has charged since July 2026.
When is Gemini 3.5 Pro coming out?
No confirmed date. Google promised it for June 2026 at I/O, but Bloomberg reported in July that coding-performance issues forced a delay, and later reporting indicates Google scrapped the original base model and restarted pretraining. As of August 15, 2026, it remains unreleased.
Is Gemini 3.7 Flash better than GPT-5.6 Terra for coding?
It’s close, not a clear win. On Google’s own 13-row comparison table, 3.7 Flash wins 7 rows but loses on the hardest agentic and terminal-use benchmarks to GPT-5.6 Terra, which costs over three times as much per token.
Why is Gemini 3.7 Flash not available in the EU or UK?
Google’s model card excludes the EEA, UK, Switzerland, and Nigeria from the consumer surface the model runs on. The exclusion lands 11 days after the EU AI Act’s enforcement powers over general-purpose AI providers activated on August 2, 2026.
What happened to Gemini Flash’s “minimal” thinking mode?
Gemini 3.7 Flash removed the “minimal” thinking tier used for cheap, high-volume classification tasks. “Low” is now the cheapest tier, and thinking tokens bill at the output rate even though only a summary is returned, raising the effective cost floor for simple workloads.
Gemini 3.7 Flash is a real, well-priced upgrade for teams already on the Flash tier, for the next four and a half months. What it isn’t is a replacement for the flagship model Google promised in May and still hasn’t shipped. Watch three things over the next 6 to 18 months: whether Gemini 3.5 Pro actually lands, whether the January price reset changes adoption at all, and whether Gemini 4’s pretraining run stays on schedule.
Want the next update the moment Gemini 3.5 Pro ships, or when the pricing resets in January? Subscribe to The Neural Loop at neuralwired.com/newsletter.
Last week, Google quietly flipped a switch that most headlines missed: Gemini Enterprise Agent Platform’s Agent Identity feature went generally available, giving AI agents their own cryptographic identity instead of borrowing a human’s login. It sounds like plumbing. It’s actually the clearest signal yet that agentic AI enterprise adoption in 2026 has quietly crossed a line most CTOs haven’t clocked: agents are no longer experiments sitting in a sandbox. They’re booking meetings, drafting reports, and touching production systems, often with the same shared credentials your interns use.
Here’s the uncomfortable part. Adoption is real. Production maturity mostly isn’t. And the gap between those two numbers is where the next eighteen months of enterprise risk, budget, and board-level scrutiny is going to live.
The number every vendor deck is quoting right now
If you’ve sat through an AI vendor pitch in the last six months, you’ve heard some version of this stat: by the end of 2026, 40% of enterprise applications will have task-specific AI agents built in, up from under 5% in 2025. That’s Gartner’s forecast, and it’s become the shorthand for “agentic AI has arrived.”
Gartner frames the trajectory in five stages: application assistants in 2025, task-specific agents in 2026, collaborative agents within apps by 2027, and cross-app agent ecosystems by 2028, building toward genuine multiagent environments by 2029. In a best case, the firm projects agentic AI could eventually drive close to 30% of enterprise application software revenue by 2035, a market north of $450 billion, up from roughly 2% today.
The deadline Gartner originally attached to that forecast, telling CIOs they had three to six months to define an agent strategy, has now quietly passed. Nobody sent out a memo. The window just closed, and most organizations are still figuring out what “having an agent strategy” even means in practice.
Reality check: Market-size projections vary by $3 billion to $5 billion depending on which analyst firm you ask and what they count as “agentic.” Keyhole Software’s synthesis of more than 20 analyst and vendor reports puts the enterprise agentic AI market at $3.67 billion in 2025, climbing to $24.50 billion by 2030, a 46.2% compound annual growth rate. Treat any single number as a rough directional signal, not a precise figure.
Adoption isn’t the story. Production is.
Here’s where the narrative most CTOs are working from starts to break down. Adoption headlines and production reality are describing two different companies.
McKinsey’s State of AI research found that 88% of organizations now use AI in at least one business function, yet only 23% are scaling agentic AI anywhere across the enterprise. A separate 2026 compilation drawing on Gartner, IDC, McKinsey, Precedence Research, MarketsandMarkets, Capgemini, and PwC found that 79% of companies report some form of AI agent adoption, but only 11% are actually running agents in production. That’s a 68 point gap between “we’re using this” and “this is doing real work.”
Metric
Figure
Source
Orgs using AI in at least one function
88%
McKinsey
Orgs scaling agentic AI enterprise-wide
23%
McKinsey
Orgs reporting some agent adoption
79%
Multi-source 2026 compilation
Orgs with agents actually in production
11%
Multi-source 2026 compilation
Pilots with measurable P&L impact
5%
MIT Project NANDA
CEOs reporting both revenue gain and cost cut from AI
12%
PwC 2026 CEO Survey
The most cited academic data point behind this gap comes from MIT’s Project NANDA. Its July 2025 report, “The GenAI Divide: State of AI in Business 2025,” analyzed 300 public AI deployments and surveyed 153 leaders across 52 organizations. The headline finding: 95% of pilots delivered no measurable profit-and-loss impact, with only 5% of integrated systems creating significant value.
That stat gets misquoted constantly as “95% of AI fails,” and it’s worth being precise here because the nuance matters for anyone making a budget decision. Over 80% of organizations had already explored general-purpose tools like ChatGPT or Copilot, and nearly 40% reported active deployment, with a pilot-to-implementation rate around 83% for those generic tools. The failure MIT documented is concentrated in custom, workflow-embedded agent builds, the expensive, bespoke projects companies commission to automate a specific internal process. Off-the-shelf assistants are doing fine. Custom agentic builds are where the money is disappearing.
Why Google just made identity the real battleground
This is the part of the story that turns an abstract stat into something you can actually act on this quarter.
Google’s Gemini Enterprise Agent Platform, first unveiled at Google Cloud Next in April 2026, reached general availability on its Agent Identity feature in the first week of August. The technical detail matters: each agent now receives its own SPIFFE-formatted cryptographic identifier rather than borrowing a shared human or service account, with an auto-rotating X.509 certificate bound to its access token through mutual TLS. In plain terms, an agent finally gets treated like its own entity, with its own least-privilege permissions and a non-repudiable audit trail, instead of quietly inheriting whatever a human employee happened to have access to.
Why does a hyperscaler shipping an identity feature matter more than another model release? Because identity, not raw capability, is the actual bottleneck standing between “we piloted an agent” and “we trust an agent with production access.” A separate finding from the Cloud Security Alliance, commissioned by Strata Identity, found only 23% of organizations have a formal, enterprise-wide strategy for agent identity management, while 37% are still relying on informal or ad hoc practices. Google is shipping infrastructure for a problem most enterprises haven’t formally acknowledged yet.
This same week, Google’s Gemini Spark agent also demonstrated it can operate the desktop version of Chrome using a user’s logged-in accounts and saved passwords, handling tasks like booking property viewings or preparing flight searches and only returning control for the payment step. It’s a consumer-facing example rather than an enterprise SaaS deployment, but it’s the most concrete real-world illustration yet of what “AI agents can book meetings without you” actually looks like once the identity and permissions layer is solved.
The security blind spot nobody priced in
Adoption without governance has a name in security circles, and it isn’t a flattering one.
Gravitee’s State of AI Agent Security 2026 report, based on a survey of more than 900 executives and technical practitioners, found that 88% of organizations had confirmed or suspected an AI-agent-related security incident in the past year. Only 14.4% required full security approval before an agent went live. A separate survey of over 160 CISOs by NeuralTrust found 72% of organizations had already implemented or were actively scaling AI agents, while just 10% had agents running in full production, a gap that tracks almost exactly with the McKinsey and multi-source figures above.
“Most agentic AI projects right now are early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied. This can blind organizations to the real cost and complexity of deploying AI agents at scale, stalling projects from moving into production.”
Anushree Verma, Senior Director Analyst, Gartner · via RCR Wireless
What makes that quote notable is who said it. Verma works at the same firm that produced the bullish 40 percent adoption forecast driving this entire news cycle. The skepticism isn’t coming from outside Gartner’s narrative. It’s embedded inside it. Gartner’s own June 2025 forecast projects that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls, and the firm has its own term for vendors overselling capability: “agent washing,” the rebranding of existing chatbots or RPA tools as agents without any real autonomous capability behind them.
The case against the hype
Not everyone thinks the adoption curve should be treated as inevitable, and the skepticism doesn’t just come from failure-rate statistics.
Nancy Gohring, Senior Research Director for AI at IDC, points to a more structural problem: vendors have little commercial incentive to make agents interoperable across platforms. “It’s a tech question, as well as a competitive situation,” she told CIO.com, noting that vendors are hesitant to open up interoperability while they’re still figuring out how to monetize the data agents generate and want to keep customers locked inside their own ecosystems. That’s not a capability gap that better prompting or a bigger model fixes. It’s a business-incentive problem, and it means enterprises buying into a single vendor’s agent platform should expect friction the moment they try to connect it to anything outside that vendor’s walls.
Forrester’s own 2026 assessment, titled “Companies Are Chasing, Few Are Catching,” found roughly three-quarters of enterprises adopting agentic AI in some form, but only a small fraction running it in genuine production, with 49% of security decision-makers separately flagging agentic AI as an active security concern in the firm’s 2026 survey.
Gartner’s Hype Cycle placement is arguably the most balanced read available: the firm expects 2026 to be the year agentic AI moves from the “peak of inflated expectations” toward the “trough of disillusionment.” That doesn’t contradict the 40% adoption forecast. It’s the same phenomenon described from two angles: deployment is moving fast, measurable value is not.
Our read: this signals a market where budget approval has gotten easier than governance approval. Getting a pilot funded is no longer the hard part. Getting it certified for production access, with real identity controls and audit trails, is.
What CTOs should actually do this quarter
If you’re evaluating agent vendors right now, the framing question matters more than the feature list. Stop asking “are we using agentic AI.” Start asking whether you have per-agent identity, real-time logging, and defined human-approval thresholds for anything irreversible. Fewer than a quarter of surveyed organizations can currently answer yes to that.
Inventory every agent in use, sanctioned and shadow, the same way you’d inventory unmanaged SaaS accounts.
Map what each agent can actually access, and move off shared API keys and service accounts toward unique per-agent credentials.
Set explicit approval thresholds for which actions an agent can take independently versus which require a human in the loop.
Score vendor pitches against real deployment counts, not roadmap slides. Ask how many customers have agents in production today, not by 2027.
Favor narrow, well-scoped pilots over broad “agentic transformation” programs. MIT’s data says focus, not ambition, is what separates the 5% that work.
CTOs approving new pilots without that governance layer in place are, statistically, more likely to end up inside Gartner’s 40% cancellation cohort by 2027.
Frequently asked questions
What is agentic AI?
Agentic AI refers to systems that independently plan, chain decisions, and execute multi-step tasks with limited ongoing human direction, unlike generative AI, which produces content in response to a single prompt. In 2026, enterprises use it for scheduling, reporting, and workflow management.
How many enterprises are using AI agents in 2026?
McKinsey’s research finds 88% of organizations use AI in at least one business function, but only 23% are scaling agentic AI anywhere across the enterprise, meaning broad experimentation hasn’t translated into widespread production use.
What percentage of AI agent projects fail?
Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls as the primary drivers, not model capability limitations.
What’s the difference between AI agents and AI assistants?
AI assistants respond to prompts and rely on ongoing human input. AI agents are task-specialized systems that can complete complex, multi-step tasks independently, such as monitoring logs and initiating a response without step-by-step direction.
Are AI agents secure?
Adoption is outpacing governance. One 2026 survey of more than 900 practitioners found 88% of organizations had a confirmed or suspected AI-agent security incident in the past year, while only 14.4% required full security approval before agents went live.
How big is the agentic AI market?
Estimates vary by methodology. Keyhole Software’s synthesis of more than 20 analyst reports puts the enterprise agentic AI market at $3.67 billion in 2025, projected to reach $24.50 billion by 2030, a 46.2% compound annual growth rate.
Where this goes next
The story of agentic AI enterprise adoption in 2026 isn’t really about whether agents work. Off-the-shelf assistants clearly do. It’s about the gap between deployment breadth and production trust, and that gap is now the thing being actively engineered around, not just talked about. Google’s identity push is the first major infrastructure response. It won’t be the last.
Watch three things over the next six to eighteen months: whether Gartner’s 40% project-cancellation forecast starts showing up in earnings calls as write-downs, whether other hyperscalers ship their own agent-identity standards or fragment the space further, and whether Forrester’s warning about a publicly disclosed agentic AI breach by the end of 2026 turns out to be right. That last one is a specific, falsifiable prediction worth checking back on.
The adoption curve isn’t the risk. Deploying ahead of your governance is.
Want the next governance-gap story before it breaks?
Google just confirmed the Pixel 11 starts at $899, roughly $100 more than the Pixel 10. The company isn’t blaming inflation or tariffs. It’s blaming a memory chip shortage that’s rewriting phone pricing across the entire industry, and the Pixel 11 is the first flagship to put a hard number on exactly how much that shortage costs.
Key insight: RAM cost Google $2.80 per gigabyte in 2025. In 2026, it costs $12. That sixfold jump, not chip design or R&D, is the single biggest driver of the Pixel 11’s higher price tag.