Category: Machine Learning

Expert machine learning analysis: model architectures, training techniques, MLOps, deployment strategies, and research breakthroughs explained for engineers and technical leaders.

  • Micron Memory Shortage 2026: AI Ate 70% of Chip Supply

    Micron Memory Shortage 2026: AI Ate 70% of Chip Supply

    Memory Chip Shortage 2026: Why Data Centers Are Eating 70% of Global Supply
    Machine Learning • Infrastructure

    Memory Chip Shortage 2026: Data Centers Will Absorb 70% of Global Supply

    The AI training bottleneck nobody’s talking about doesn’t involve a single GPU.

    Your next DRAM order just got 93% more expensive than it was three months ago. That’s not an estimate. It’s what TrendForce recorded in a single quarter of 2026, and it’s the surface symptom of something much bigger: data centers are on track to absorb roughly 70% of all memory chips produced worldwide in 2026, up from just 20% to 30% as recently as 2022.

    If you’re an ML engineer, infrastructure lead, or CTO planning training capacity for next year, this is the memory chip shortage 2026 story you actually need to understand, and it’s not about GPU allocation anymore. It’s about whether there’s enough memory bandwidth on the planet to feed the GPUs you already have.

    What’s Actually Happening to the Memory Market

    Start with the suppliers, because they’re the ones with the clearest view of demand. Samsung’s CFO Park Soon-cheol told investors on the company’s Q1 2026 earnings call that HBM4 “sales volume has already been completely sold out” for the year, with HBM4 expected to make up more than half of Samsung’s total HBM revenue by the third quarter. SK hynix said something almost identical back in October 2025: customers had already claimed the company’s entire 2026 output of both DRAM and NAND.

    Micron’s numbers tell the same story from a different angle. The company’s fiscal Q3 2026 results show HBM4 already in high-volume shipment for its lead customer’s platform, while next-gen HBM4E won’t reach volume production until calendar 2027. Micron guided fiscal Q4 2026 revenue to $50 billion. A year earlier, that number was $9.3 billion.

    None of this is speculation dressed up as forecasting. It’s suppliers describing capacity they’ve already sold, for products they haven’t finished shipping.

    The Numbers Behind the Panic

    Here’s what the reallocation actually looks like in hard figures.

    MetricFigureSource
    DRAM contract price increase, Q1 2026 (QoQ)93% to 98%TrendForce
    36GB HBM3E spot price vs. long-term contract price~$2,100 vs. $300 to $400 (4 to 5x)The Motley Fool
    South Korea DRAM export price, year over year+401%, reaching $92,183/kgChosun Ilbo trade data
    Global memory market forecast, 2026Raised from $551.6B to $889.3BTrendForce
    Global memory market forecast, 2027Over $1.28 trillion (+44% YoY)TrendForce
    Retail 32GB DDR5-6000 kit price, Aug 2026$402, up from $110 to $140 a year earlierTom’s Hardware pricing data
    HBM share of top-3 suppliers’ DRAM wafer input, 2025/2026/202718% / 22% / 30%TrendForce
    Notice the pattern. It isn’t just HBM (the specialized memory stacked directly onto AI accelerators) getting expensive. Ordinary DDR5, the RAM in laptops and servers with no connection to AI training whatsoever, is being dragged up in price because the same fabs, the same wafer starts, and the same clean-room capacity now compete against AI demand for every gigabyte produced.

    Why This Has Nothing to Do With GPUs

    Here’s the part most coverage misses. The GPU shortage that dominated headlines in 2023 and 2024 is largely over. Nvidia, AMD, and their foundry partners have scaled logic production aggressively. What hasn’t scaled at the same rate is the memory that sits next to that logic, and that gap is now the binding constraint on how fast AI models can actually be trained.

    Micron’s HBM Design Architecture Fellow, Raghu Sreeramaneni, put a number on the gap at Hot Chips 2026:

    “Compute scales roughly 3x every two years. HBM bandwidth scales only about 2x every two years. The memory wall persists, and it may be worsening.” Raghu Sreeramaneni, HBM Design Architecture Fellow, Micron Technology — via wccftech, Hot Chips 2026
    That mismatch has a name in chip architecture circles: the memory wall. It means you can add more GPUs to a rack, but if the memory bandwidth feeding those GPUs doesn’t grow at the same pace, the extra compute sits idle waiting for data. Micron’s own materials cite Meta’s Llama 3 training paper, which attributed 17% of unintended training interruptions to HBM issues, a concrete number showing this isn’t a theoretical problem.

    OpenAI’s COO Brad Lightcap confirmed the shift publicly in March 2026, telling reporters the company’s binding constraint had moved: it used to be power availability. Now, in his words, “right now it’s memory.”

    Why this matters for planning: if your infrastructure roadmap is still built around GPU allocation as the scarce resource, you’re solving last year’s problem. The scarce resource in late 2026 is memory bandwidth per accelerator, and that constraint doesn’t get fixed by buying more chips.

    Who’s Feeling the Squeeze

    This stopped being a tech-press story in mid-2026. A coalition representing telecommunications, automotive, medical-device, and retail trade associations formally warned U.S. regulators that expanding AI data centers were consuming an enormous share of available memory chip capacity, according to reporting confirmed by CSIS. That’s four industries with nothing to do with AI, telling Washington the same fabs are now out of reach for them.

    TrendForce analyst Avril Wu, who has tracked the memory sector for close to two decades, doesn’t hedge on how unusual this cycle is:

    “I’ve tracked the memory sector for almost 20 years, and this time really is different. It really is the craziest time ever.” Avril Wu, Analyst, TrendForce — via Tom’s Hardware
    Counterpoint Research’s MS Hwang went further in the same piece, telling buyers to act as if capacity for 2028 is already gone: “you gotta buy a plane ticket and get that allocation from manufacturers right now.”

    That’s not marketing language from a supplier trying to justify a price hike. That’s an independent analyst telling procurement teams the window has already closed for near-term allocation, and the next window (2028 capacity) is closing too.

    When Does This Actually End

    Short answer: not soon, and here’s the specific reason why. New memory fabs take years to build, while GPU compute capacity can effectively double annually. That asymmetry is the whole story.

    SK hynix broke ground on a new HBM fab in Indiana on August 27, 2026, an investment described as “over $4 billion,” with cleanroom completion not scheduled until October 2028, and volume HBM output not expected before 2029. The company’s Korean Yongin fab, part of a separate 54.3 trillion won ($38.3 billion) investment, targets a cleanroom opening in June 2029. Read those dates again. The fabs breaking ground today won’t meaningfully add supply for three years.

    Kushal Fernandes, a partner at Kearney’s product redesign practice, put a specific range on the relief timeline in an interview with Design News:

    “The earliest we see meaningful new capacity is 2028, but that relief will be partial rather than substantial. New fabs largely ramp through 2029, and if AI demand continues at its current pace, we do not anticipate substantial relief before early 2030.” Kushal Fernandes, Partner, Kearney — via Design News
    That’s a wide band (late 2028 to early 2030), and it depends entirely on one variable nobody can currently forecast with confidence: whether AI training demand keeps compounding at its current rate.

    The Skeptic’s Case

    Not everyone accepts that this shortage is a permanent structural feature of the AI economy, and the strongest pushback deserves a real hearing rather than a footnote.

    Ed Zitron, host of the “Better Offline” podcast and a persistent critic of AI infrastructure spending, argues the entire capex cycle underpinning memory demand is itself unsustainable. On his show, he pointed to a gap between announced infrastructure deals and actual revenue: over $178.5 billion in data center deals against less than $1 billion in compute revenue outside the hyperscalers themselves. His warning is blunt: if a major AI lab’s business falters, it “will trigger a brutal collapse of the entire AI bubble,” and memory demand along with it.

    This isn’t just rhetoric. In late June 2026, a sharp tech sell-off saw Samsung and SK hynix shares drop 12% in a single morning, South Korea’s KOSPI fall 10%, and Micron, up nearly 800% over the prior year, plunge 13% on renewed AI-bubble anxiety. Markets themselves aren’t fully convinced this demand is permanent.

    Our read: the memory wall itself (compute scaling 3x against memory bandwidth scaling 2x) is settled engineering fact, confirmed independently by Micron’s own architects. Whether current AI capex is validated by end-market revenue is a genuinely separate, open question, and treating the two as the same debate is where a lot of coverage goes wrong. One is physics. The other is a bet on demand.

    What Engineering Teams Should Do Now

    If you’re planning training or inference capacity into 2027, three things follow directly from the data above.

    • Model memory as its own volatile line item. With HBM3E spot prices running 4 to 5x above contract pricing and DRAM up nearly 100% in a single quarter, any budget built on 2024-era per-gigabyte costs is already wrong. Separate memory pricing risk from GPU pricing risk in your forecasts.
    • Assume allocation now depends on relationships, not budget. Samsung, SK hynix, and Micron have all described 2026 HBM output as effectively sold out. Teams without existing multi-year supply agreements are competing for scraps on the spot market, at multiples of contract price.
    • Treat memory efficiency as a cost-avoidance tool, not a nice-to-have. Roofline analysis (determining whether a workload is memory-bound or compute-bound) can reveal real savings without buying a single new chip. KV-cache compression techniques, better batching, and memory-aware scheduling reduce dependence on scarce HBM allocation directly.
    Teams weighing whether to reduce cloud dependence entirely should also look at how on-device AI is replacing parts of the cloud inference stack in 2026, since edge inference sidesteps data center memory constraints altogether for certain workloads. And if you’re trying to understand how this shortage connects to the broader AI infrastructure financing picture, our coverage of the SB Energy IPO and its OpenAI dependence risk lays out the capex side of the same story.


    Frequently Asked Questions

    What is causing the memory chip shortage in 2026?

    AI data centers are diverting DRAM and HBM production away from consumer electronics to feed GPU-based training and inference. Data centers are forecast to consume roughly 70% of global memory output in 2026, up from 20% to 30% in 2022, according to TechNewsWorld’s reporting on industry-analyst forecasts.

    What is the “memory wall” in AI?

    The memory wall describes the growing gap between how fast AI compute scales versus how fast memory bandwidth can keep up. Micron’s Hot Chips 2026 presentation states compute scales roughly 3x every two years while HBM bandwidth scales only about 2x, leaving processors waiting on data.

    When will the memory chip shortage end?

    No supplier or major analyst firm has confirmed a firm end date. SK hynix’s new fabs in Indiana and Korea don’t target cleanroom completion until 2028 and 2029, and Kearney forecasts meaningful relief is unlikely before early 2030 if AI demand continues at its current pace.

    How much have memory prices risen in 2026?

    Conventional DRAM contract prices rose roughly 93% to 98% quarter over quarter in Q1 2026 alone, the steepest quarterly increase TrendForce has recorded, while some HBM3E spot prices trade 4 to 5 times above long-term contract pricing.

    Is HBM different from regular RAM (DDR5)?

    Yes. HBM stacks multiple DRAM dies vertically, connected through an ultra-wide interface (up to 2,048 bits with HBM4), delivering far higher bandwidth than DDR5. HBM also consumes roughly 3 times the wafer capacity per gigabyte to manufacture, which is why it crowds out conventional DRAM production.

    Which companies make HBM memory for AI chips?

    Samsung Electronics, SK hynix, and Micron Technology are the three merchant suppliers. SK hynix has historically led HBM shipment share, though Samsung’s share has been rising through 2026 as HBM4 output ramps.


    Where This Leaves You

    What’s actually changed since 2024 isn’t that GPUs got scarce again. It’s that the bottleneck moved one layer down the stack, into the memory sitting right next to the compute, and that layer takes years to expand rather than months. The engineering teams that win the next 18 months won’t necessarily be the ones with the biggest GPU order. They’ll be the ones who treated memory bandwidth as the scarce resource it actually is, months before their competitors caught on.

    Three things worth watching over the next six to eighteen months: whether SK hynix and Samsung’s 2028 to 2029 fab timelines hold without slipping further, whether AI training demand shows any sign of the deceleration that would validate the bubble skeptics, and whether memory-efficient training techniques (quantization, KV-cache compression, MoE-aware memory management) become standard practice rather than optimization afterthoughts.

    Want the next development in this story before it hits the front page? Subscribe to The Neural Loop at neuralwired.com/newsletter for weekly briefings on the infrastructure decisions actually shaping AI in 2026.

  • Zillow’s $569M AI Failure: What Is Model Drift? (2026)

    Zillow’s $569M AI Failure: What Is Model Drift? (2026)

    Model Drift: Why Your AI Fails Silently in Production (2026 Guide)
    Machine Learning / AI Infrastructure

    Model Drift: Why Your AI Fails Silently and No One Notices Until the Bill Arrives

  • NVIDIA: Small Language Models Now Beat LLMs in 2026

    NVIDIA: Small Language Models Now Beat LLMs in 2026

    AI Infrastructure

    NVIDIA: Small AI Models Now Beat 70B Giants

    Your AI agent doesn’t need a trillion-parameter brain to check a database field. It needs a fast, cheap, accurate answer, and right now, you’re probably paying frontier-model prices for kindergarten-level work. NVIDIA researchers say small language models now match or beat large language models on narrow, well-defined tasks, at a fraction of the inference cost, and Gartner expects the shift to triple by 2027.

    This isn’t a fringe claim. It’s the thesis of a formal NVIDIA Research position paper, backed by named model benchmarks, a peer-reviewed medical study, and a hard market forecast from one of the industry’s most conservative analyst firms. Here’s what the data actually shows, and where it doesn’t hold up.

    The Paper That Started the Argument

    In June 2025, a team from NVIDIA Research and Georgia Tech, led by Peter Belcak, posted a position paper to arXiv called “Small Language Models are the Future of Agentic AI.” It’s still listed as a preprint under review, not a peer-reviewed benchmark study, and that distinction matters. But the argument inside it has spent over a year working its way through enterprise AI teams, and by 2026, the evidence started catching up to the claim.

    The paper’s definition of “small” is practical, not arbitrary: a model that fits on a common consumer device and runs with latency low enough for single-user agentic work. As of 2025, the authors were comfortable calling most models under 10 billion parameters SLMs.

    Their core complaint: most AI agent systems route 40 to 70 percent of their compute through a generalist LLM, even for tasks that are structurally narrow, things like tool calls, structured extraction, and code-orchestrated steps. That’s the equivalent of hiring a surgeon to change a lightbulb.

    “SLMs are sometimes ‘good enough’ for many nodes in an agent graph, especially tool-calling, structured reasoning, and code-orchestrated steps, sometimes matching or beating larger LLMs for those narrow tasks.”
    Peter Belcak, AI Researcher, NVIDIA Research
    The paper cites named results to back this up. Microsoft’s Phi-2, at 2.7 billion parameters, matches commonsense reasoning and code generation scores of models over ten times its size, while running roughly 15x faster. Phi-3 small, at 7 billion parameters, matches the language understanding of 70-billion-parameter models from the same generation and beats them on code generation. Hugging Face’s SmolLM2 family, some variants under 2 billion parameters, matches the tool-calling performance of 14-billion-parameter contemporaries.

    Two of the more striking claims: DeepSeek-R1-Distill-Qwen-7B reportedly outperforms Claude-3.5-Sonnet and GPT-4o on commonsense reasoning tasks, and Salesforce’s xLAM-2-8B claims state-of-the-art tool-calling accuracy, ahead of both GPT-4o and Claude 3.5, at a fraction of the parameter count.

    The Numbers That Actually Hold Up

    Strip out the vendor blog posts and single-paper claims, and here’s what’s independently verifiable or attributable to a named source:

    Figure Source Date
    0.5B model hits 91.7% accuracy vs. 88.6% for a 72B model on classification Forbes analysis June 2026
    SLMs run 10 to 30x cheaper per token than 70 to 175B LLMs NVIDIA Research paper 2025/2026
    60% of MetaGPT’s LLM queries reliably handleable by SLMs NVIDIA paper, Appendix B.1 2025
    70% of Cradle GUI-agent queries SLM-replaceable NVIDIA paper, Appendix B.3 2025
    Task-specific model usage to triple general LLM usage by 2027 Gartner press release April 2025
    Notice the range in that MetaGPT and Cradle comparison. Sixty percent replaceable for one agent, seventy percent for another. That gap isn’t noise, it’s the real story: how much of your workload an SLM can absorb depends entirely on what your agent is actually doing.

    A Real-World Test: SLMs in Medicine

    Position papers and vendor benchmarks are one thing. A controlled, peer-reviewed comparison is another. In January 2026, researchers from the Bascom Palmer Eye Institute at the University of Miami and the Federal University of São Paulo published a study in JMIR comparing a retrieval-augmented small language model, trained specifically on ophthalmology literature, against GPT-4 on 35 frequently asked glaucoma questions.

    Three independent glaucoma specialists graded the answers on a three-tier accuracy scale, blind to which model produced which response. This is exactly the kind of test the SLM argument needed: narrow domain, real clinical stakes, named institutions, independent graders. It’s a data point the field can build on rather than take on faith.

    Gartner’s 2027 Prediction

    On April 9, 2025, Gartner made it official. The firm predicted that by 2027, organizations will deploy small, task-specific AI models at usage volumes at least three times higher than general-purpose LLMs.

    “The variety of tasks in business workflows and the need for greater accuracy are driving the shift towards specialized models fine-tuned on specific functions or domain data. These smaller, task-specific models provide quicker responses and use less computational power, reducing operational and maintenance costs.”
    Sumit Agarwal, VP Analyst, Gartner
    Read that prediction carefully. It’s a 2027 target, not a claim that the shift has already happened. Most production agent stacks in 2026 are still LLM-first. Gartner is describing a documented trend and a forecast, not the current default state of the industry, and conflating the two is where a lot of the hype gets ahead of the reality.

    The Cost Math Behind the Shift

    This is where the argument stops being academic. Enterprise cost breakdowns put a private SLM endpoint handling 10,000 daily queries at roughly $500 to $2,000 a month. The equivalent workload on frontier LLM APIs runs $5,000 to $50,000 a month, depending on the model and context length. That’s not a marginal saving. At scale, across millions of daily agent invocations, it’s a material line on the P&L.

    Fine-tuning agility compounds the advantage. Parameter-efficient methods like LoRA and DoRA let teams specialize an SLM for a new task in GPU-hours, not the weeks a full LLM fine-tuning cycle typically takes. If your business changes its workflows every quarter, that iteration speed matters as much as the raw inference cost.

    The catch: most of these cost figures trace back to vendor analyses and the NVIDIA paper’s own citations, not independent third-party audits. Treat them as directionally reliable, not laboratory-verified.

    Where the Argument Breaks Down

    To its credit, the NVIDIA paper doesn’t dodge its own weakest points. It preserves the strongest counter-argument verbatim: a substantial body of empirical evidence shows large language models outperform small ones on general language understanding, because LLMs follow scaling laws that reward size with capability. The authors even flag a hypothesized “semantic hub” mechanism, a way larger models may integrate meaning across languages and modalities that smaller architectures structurally can’t replicate.

    There’s also an economics rebuttal the paper admits it can’t fully answer: the per-token savings of a small model can get swallowed by the difficulty of fully utilizing and load-balancing a fleet of specialized SLM endpoints, something a single generalist LLM endpoint doesn’t have to deal with. Add in the MLOps and talent overhead of managing multiple fine-tuned models, and the total cost of ownership gets a lot murkier than the headline per-token numbers suggest.

    Zoom out further and there’s a broader skepticism worth weighing. Gary Marcus, Professor Emeritus at NYU and a longtime critic of scaling-driven AI hype, isn’t commenting on SLMs specifically, but his wider point about the industry is relevant here.

    “A large fraction of what LLMs do is mostly just memorization,” and current systems “still aren’t adding a lot of quantifiable value to the world.”
    Gary Marcus, Professor Emeritus, NYU
    Marcus cites the Remote Labor Index finding that AI could fully complete only about 2.5 percent of remote jobs tested, as reported by the Washington Post. Use his view as a check on compute-versus-capability claims generally, not as a direct rebuttal to the SLM data, which stands on its own narrower footing.

    What This Means for Your Stack

    If you’re an engineering lead running agent workflows on a single frontier-model endpoint, the actionable move isn’t “replace your LLM.” It’s audit first. NVIDIA’s paper actually outlines a six-step conversion process worth stealing: log real usage patterns, curate the resulting data, cluster it by task type, select SLM candidates for the narrow clusters, fine-tune, and iterate.

    Every credible source here, including NVIDIA’s own paper, describes a hybrid architecture, not a replacement. A frontier LLM stays as the planner and orchestrator. SLMs take over the narrow, repetitive, format-constrained work underneath it: classification, extraction, tool calls, structured code steps. Gartner’s own guidance echoes this, recommending small models specifically where an LLM hasn’t met response quality or speed expectations, not as a wholesale swap.

    Our read: the teams that win the next 18 months won’t be the ones who bet everything on either model size. They’ll be the ones who actually measure which of their agent’s tasks are narrow enough to hand to a cheaper, faster model, and which genuinely need the reasoning a frontier LLM provides.

    FAQ

    What is the difference between a small language model and a large language model?

    The core difference is parameter count and what it implies. LLMs, roughly 7 billion to over a trillion parameters, hold broad world knowledge and cross-domain reasoning without task-specific tuning. SLMs typically range from a few million to about 7 billion parameters, trading some generality for speed, low cost, and on-device deployability.

    Can small language models really match LLM accuracy?

    Yes, on narrow, well-defined tasks. One 2026 analysis found a 0.5-billion-parameter model hit 91.7 percent accuracy versus 88.6 percent for a 72-billion-parameter model on simple classification, though LLMs still hold the advantage on broad, open-ended reasoning.

    Are small language models cheaper to run than LLMs?

    Yes. Serving a 7-billion-parameter SLM is estimated at 10 to 30 times cheaper in latency, energy, and compute than a 70 to 175-billion-parameter LLM, according to NVIDIA Research.

    Will small language models replace large language models?

    Not entirely. Gartner predicts organizations will use small, task-specific AI models three times more than general-purpose LLMs by 2027, but researchers and analysts frame this as hybrid adoption, with LLMs still orchestrating and SLMs handling narrow tasks, not a full replacement.


    The Bottom Line

    What you now know that you didn’t before: the “bigger model, better results” assumption doesn’t hold once you narrow the task down to something specific and repeatable. NVIDIA’s research, Gartner’s forecast, and at least one peer-reviewed clinical study all point the same direction, even while the paper behind this movement openly admits where scaling laws and operational reality push back.

    Watch three things over the next 6 to 18 months: whether Gartner’s 2027 usage-volume prediction stays on pace, whether more peer-reviewed domain-specific studies follow the glaucoma model, and whether the MLOps tooling for managing fleets of SLMs matures enough to close the operational gap the NVIDIA paper itself flags as unresolved.

    Small language models aren’t going to replace the model powering your chatbot’s hardest conversations. But if you’re still routing every tool call and classification task through a frontier LLM in 2026, you’re very likely paying trillion-parameter prices for kindergarten-level work.

    Want the next breakdown like this in your inbox?
    Subscribe to The Neural Loop at neuralwired.com/newsletter
  • DeepEval vs RAGAS vs Langfuse: Best LLM Tools 2026

    DeepEval vs RAGAS vs Langfuse: Best LLM Tools 2026

    Best LLM Evaluation Tools 2026: 7 Tested, Ranked, Compared
    ML Tooling / Developer Focus

    Best LLM Evaluation Tools 2026: 7 Tested, Compared

  • vLLM vs SGLang vs TensorRT-LLM: 2026 Benchmark Guide

    vLLM vs SGLang vs TensorRT-LLM: 2026 Benchmark Guide

    Best LLM Inference Optimization Tools 2026: 7 Engines Tested
    Infrastructure / Developer Tools

    vLLM vs SGLang vs TensorRT-LLM: 7 Inference Engines Tested Against MLPerf v6.0

  • Pinecone Says RAG Is Obsolete: Complete 2026 Verdict

    Pinecone Says RAG Is Obsolete: Complete 2026 Verdict

    Pinecone Bets RAG Is Obsolete. The Data Disagrees
    AI Infrastructure

    Pinecone Bets RAG Is Obsolete. The Data Disagrees

    The company that made retrieval-augmented generation a household term just told its own 800,000 developers to stop doing it. Here is what that means if you are choosing between RAG and a 2 million token context window in 2026.

    Your engineering team spent 2024 building a retrieval pipeline. Chunk the docs, embed them, store them in a vector database, retrieve the top matches, stuff them into a prompt. It worked, mostly. Then Gemini shipped a 2 million token context window, Claude and GPT-5.4 hit 1 million, and someone on Slack asked the question everyone is now asking: why not just paste the whole knowledge base in and skip the plumbing?

    That question has a real answer now, and it is not the one either side of the debate wants. A 2 million token context window does not replace retrieval-augmented generation. It changes what retrieval is for. And the company that spent four years teaching the industry how to build RAG vs long context pipelines just told the market, in public, that the pattern it popularized is already the bottleneck.

    The context window race just hit a new ceiling

    By April 2026, five frontier labs had all crossed the same line. Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro, Qwen 3.6 Plus, and Llama 4 Maverick each shipped a 1 million token context window. Meta pushed further with Llama 4 Scout, advertising 10 million tokens, though independent testers found its usable recall breaks down well short of that number. Google’s Gemini line has sat at the 2 million token mark since early 2026, which is why “2 million token context window” is now the phrase enterprise buyers type into Google before they type anything else.

    By June 9, at least 13 models had crossed the 1 million token line, according to a pricing comparison from Morph. What that comparison also revealed is that “1 million tokens” is not one product. It is thirteen different products with wildly different economics.

    ModelCost to fill a 1M-token context window
    DeepSeek V4 Flash$0.14
    Claude Fable 5$10.00
    Source: Morph, June 9, 2026. A 71x spread across the field.

    That 71x spread is the first sign that “just use a bigger window” is not a strategy. It is a pricing decision you have not made yet.

    Context rot: why bigger windows are not always better

    In July 2025, three researchers at the vector database company Chroma published a report that has become the most-cited technical pushback on long-context marketing copy. Kelly Hong, Anton Troynikov, and Jeff Huber tested 18 frontier models, including the GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 families, on tasks specifically designed to hold difficulty constant while varying only input length.

    The finding that should worry anyone planning to dump a full knowledge base into a prompt: every single model got less reliable as the input got longer, even on tasks a human would call trivial. And in a twist that inverts a common assumption among RAG engineers, models performed worse on well-organized, logically coherent source documents than on the same content shuffled into random order.

    “Models do not use their context uniformly.” Kelly Hong, Anton Troynikov, and Jeff Huber, Chroma Research, “Context Rot: How Increasing Input Tokens Impacts LLM Performance,” July 2025
    Chroma is a retrieval infrastructure vendor, so this finding is also commercially convenient for the company publishing it. That is worth disclosing. It does not make the methodology wrong. The 18-model benchmark is open source and independently reproducible, and it lines up with a separate, older finding known as “lost in the middle”: accuracy drops 20 to 30 percentage points when the answer sits in the middle of a long document instead of at the start or end, a pattern first documented by Liu et al. and replicated across model families since.

    Put together, these results point to a rule NVIDIA’s own RULER benchmark backs up: the effective, reliable portion of a context window typically runs at 50 to 65% of the number on the marketing page. Some of Chroma’s own findings suggest the real, safe margin for production workloads is tighter still, closer to a quarter or a third of the advertised maximum.

    The real cost of going long

    Even where accuracy holds up, long-context prompting is not cheap next to modern retrieval. A 2026 arXiv study titled “Long Context vs. RAG for LLMs” ran a direct cost comparison across GPT-5.4-mini and nano on document-grounded question answering. The result: long-context prompting averaged roughly $0.1181 per query, against $0.0045 to $0.0046 for keyword or semantic retrieval. That is a 10x-plus cost gap, and it is the more conservative of the figures floating around; some blog posts cite gaps as high as 1,250x, but those appear to compare different cost baselines and should be treated with skepticism.

    Anthropic’s prompt caching cuts input costs by up to 90% and latency by up to 85% on repeated long prompts, which matters more than raw context size for most production bills. The lesson is not “context is expensive.” It is that caching, batching, and retrieval scope are the real levers, and a bigger window without any of those disciplines is the most expensive way to solve the problem.

    Worth flagging for enterprise architects: access to any single frontier model is not guaranteed to be stable. In June 2026, Anthropic temporarily suspended access to Claude Fable 5 and Mythos 5 to comply with U.S. Department of Commerce export controls, restoring it on July 1 after the controls were lifted (Anthropic’s statement). Whatever architecture you pick, model availability is now a variable you plan around, not an assumption you make.

    Pinecone just bet against the category it built

    On May 4, Pinecone, the vector database that made RAG a standard pattern for roughly 800,000 developers and 9,000 paying customers, launched Nexus, which it calls a “knowledge engine for agents,” alongside KnowQL, a query language built around six primitives: intent, filter, provenance, output shape, confidence, and latency budget.

    Pinecone’s own framing is blunt. It describes retrieval-at-inference, the classic chunk-and-embed pattern the company spent four years teaching the market, as the “ten blue links era of agentic retrieval.” Its argument: agents stuck in retrieve-read-retrieve loops complete only 50 to 60% of tasks and burn 85% of their effort just fetching context, before any actual reasoning happens.

    Instead of retrieving raw chunks at query time, Nexus precompiles source data into structured, cited, task-specific artifacts ahead of time, so an agent queries a compiled answer rather than a pile of documents. Harrison Chase, the CEO of LangChain and the person widely credited with popularizing the term “context engineering,” backed the framing on Pinecone’s own launch post.

    “Building reliable, long-horizon agents is fundamentally a context engineering problem.” Harrison Chase, CEO, LangChain, on Pinecone’s Nexus launch post, May 4, 2026
    Janakiram MSV, the cloud and AI analyst who covers infrastructure shifts for The New Stack, called out just how unusual this is. Most vendors keep selling into a category long after the market has moved past it. Pinecone named the shift itself.

    “Pinecone just declared the RAG era over.” Janakiram MSV, The New Stack, “The company that made RAG mainstream is now betting against it,” May 6, 2026
    Our read: MSV’s framing is closer to right than Pinecone’s own marketing copy. This is not “RAG is dead.” It is RAG’s naive, retrieve-then-hope form getting replaced by something more deliberate, the same shift Anthropic’s Skills and Cursor’s project rules are pushing at the editor and agent-framework layer. The pattern is not new. The vendor saying it out loud is.

    The Subquadratic wildcard: 12 million tokens, unverified

    One day after Pinecone’s launch, Miami-based startup Subquadratic emerged from stealth with $29 million in seed funding and a model called SubQ, built on what it calls a Subquadratic Selective Attention architecture. Founded by CEO Justin Dangel and CTO Alexander Whedon, both veterans of Meta, the company claims SubQ’s research version supports a 12 million token context window, roughly 120 books, while scaling compute linearly rather than quadratically with input length.

    The headline number, as reported by SiliconANGLE: SubQ scored 95% on the RULER 128K benchmark at about $8 in compute, against 94% accuracy and roughly $2,600 for Claude Opus on the same test, a claimed 300x cost reduction. Backers reportedly include Tinder co-founder Justin Mateen and early investors in Anthropic, OpenAI, Stripe, and Brex.

    Treat every one of those numbers as “reported by Subquadratic” until someone outside the company replicates them. As of this writing, no independent benchmarking team has confirmed the 52x attention speedup, the 92.1% needle-in-haystack recall at 12 million tokens, or the roughly 1,000x compute reduction the company claims at full context length. If verified, it would be the largest single jump in usable context the field has seen. If not, it joins a long list of long-context claims that looked revolutionary on launch day and ordinary six months later.

    So is RAG dead? The growth data says no

    Here is the part the “RAG is dead” headlines tend to skip: RAG-adjacent infrastructure spending is still growing fast, and growth data does not lie the way marketing copy can. Market-sizing firms disagree sharply on the exact dollar figures. Grand View Research puts the market at $1.2 billion in 2024, growing to $11 billion by 2030 at a 49.1% compound annual growth rate. Precedence Research estimates $2.76 billion in 2026 climbing to $67.42 billion by 2034. MarketsandMarkets lands in between, at $1.94 billion in 2025 growing to $9.86 billion by 2030. Cite one firm at a time, since the numbers do not reconcile with each other, but the direction across all three is the same: a technology genuinely on its way out does not post 38 to 49% annual growth.

    Production engineers writing on DEV Community made the practical case bluntly: no context window, however large, holds an enterprise knowledge base running to millions of documents. A single 1 million token Claude Sonnet-class prompt runs roughly $3 at list pricing, and that does not scale to production query volumes the way retrieval does. Their position is that RAG’s continued growth is itself the strongest evidence against the “dead technology” framing, not despite the long-context hype but because of what enterprises are actually shipping underneath it.

    What this means for your stack

    Stop treating this as RAG versus long context. Treat it as a context budget you have to manage regardless of which technique you use.

    • Cap your assumptions at 25 to 30% of the advertised window. That is roughly what Chroma’s own findings suggest is the safe, reliable slice of any long-context claim, sticker number aside.
    • Pair retrieval with compaction. For long agent sessions, summarization and compaction loops matter more than raw window size, because irrelevant content is what causes context rot, not length alone.
    • Do not rip out retrieval infrastructure on the assumption long context replaces it. Teams that did this in 2024 and 2025 are the ones now eating the 10x-plus cost premium documented above.
    • Watch where vendor R&D is actually pointed, not where the marketing copy points. Pinecone’s own pivot from raw retrieval toward precompiled, agent-queryable artifacts is a better signal than any single benchmark chart.
    • Evaluate new entrants before migrating production workloads. Subquadratic’s numbers are compelling on paper and unverified in practice. Run your own evals on your own data first.
    One more thing regulated industries should not skip: RAG’s retrieval logs double as an audit trail. Raw long-context prompting does not produce one by default. In finance, healthcare, or legal workflows, that gap is not academic. It is a compliance requirement waiting to surface during an audit, usually at the worst possible time.


    Frequently asked questions

    Does a bigger context window replace RAG?

    Rarely. Long context reduces the need for aggressive retrieval on smaller, bounded corpora, but no window, even 12 million tokens, holds an enterprise knowledge base with millions of documents. Long-context prompting also runs roughly 10x or more expensive per query than modern retrieval in controlled 2026 benchmarks.

    What is “context rot”?

    Context rot is measurable performance degradation as an LLM’s input length grows, even on simple tasks. Chroma Research tested 18 frontier models in 2025 and found every one degraded with length, with logically coherent documents sometimes hurting performance more than shuffled ones.

    What causes the “lost in the middle” problem?

    Models attend most reliably to information at the very start and end of their context window. Liu et al.’s benchmark found accuracy drops 20 to 30 percentage points when the answer sits mid-context, a pattern replicated across GPT, Claude, and other model families since.

    How much does a 1 million token prompt cost?

    It depends heavily on the model. As of June 2026, filling a 1 million token window ranges from about $0.14 on DeepSeek V4 Flash to $10.00 on Claude Fable 5, a 71x spread, before caching discounts are factored in.

    Is RAG still worth building in 2026?

    Yes, for most production systems with large, dynamic, or compliance-sensitive corpora. RAG-related infrastructure spend kept growing at 38 to 49% CAGR across multiple market estimates even as long-context windows expanded, and 2026 is shaping up to be a hybrid-architecture year rather than a winner-take-all contest.


    The bottom line

    Nothing here says long context is a bad bet or that RAG is finished. What the evidence actually supports is narrower and more useful: raw context length is not the same thing as usable context, cost scales against you faster than accuracy does, and the vendor that built the RAG category is now telling the market to build the next layer up, not to abandon retrieval altogether.

    Watch three things over the next 6 to 18 months. First, whether independent labs confirm any of Subquadratic’s numbers, since that would be the first real architectural break from quadratic attention costs. Second, whether Pinecone’s Nexus and KnowQL numbers hold up in production the way they did in Pinecone’s own benchmarks. Third, whether “context engineering,” the discipline of deliberately curating what enters a model’s window regardless of technique, becomes a formal job function the way “prompt engineering” did in 2023.

    The teams that win this cycle will not be the ones who pick a side in the RAG-versus-context debate. They will be the ones who stopped treating context size as a proxy for context quality months before everyone else did.

    Want the next infrastructure shift in your inbox before it hits the front page? Subscribe to The Neural Loop at neuralwired.com/newsletter.

  • Unsloth AI: How ORPO and GaLore Cut Fine-Tuning Cost

    Unsloth AI: How ORPO and GaLore Cut Fine-Tuning Cost

    Machine Learning

    How Unsloth Made LLM Fine-Tuning 2x Faster in 2026