Category: Technology

NeuralWired’s Technology section covers the developments reshaping how the world builds, deploys, and regulates digital innovation. We report daily on the stories driving global conversation in artificial intelligence, big technology companies, startups and venture funding, cybersecurity, consumer gadgets and devices, and blockchain and cryptocurrency.

Our technology coverage goes beyond product announcements. When a major AI model launches, we explain what it can actually do and where its claims are overstated. When a startup raises a large funding round, we look at whether the business behind it can sustain that valuation. When a cybersecurity breach hits the news, we explain who is affected and what comes next, not just what happened. Each article is built from original research into primary sources, including company statements, technical documentation, regulatory filings, and verified data, and is written by our editorial team rather than generated automatically.

Readers come to this section for daily updates on the technology stories that matter globally, from shifts inside major technology companies to emerging tools changing how people work, communicate, and build. Whether you are a founder, an investor, an engineer, or simply someone trying to understand where technology is heading next, NeuralWired’s Technology coverage is built to keep you informed without wasting your time on hype.

  • CrowdStrike AI SOC: The 2% Failure Rate Hiding Cobalt Strike

    CrowdStrike AI SOC: The 2% Failure Rate Hiding Cobalt Strike

    AI SOC Automation: How AI Closed 43% of Alerts Before a Human Saw Them — and What the 2% Failure Rate Actually Cost | NeuralWired
    Security Operations • Enterprise AI

    AI SOC Automation Closed 43% of Alerts Before a Human Saw Them. Here’s What Lived Inside the 2% It Got Wrong.

    Every weekday morning, a real threat is hiding inside a low-severity alert at the average enterprise. The AI already looked at it. The AI already closed it. The analyst never saw it.

    That is not a hypothetical from a vendor white paper. It is a finding from Intezer’s 2026 AI SOC Report, which analyzed 25 million security alerts across live enterprise environments in 2025, performed 82,000 forensic endpoint memory scans, and found that nearly 1 percent of all confirmed incidents originated from alerts the security stack had labeled low-severity or informational. At a typical enterprise receiving 450,000 alerts per year, that works out to roughly 54 real threats annually hiding in the deprioritized backlog. One per week. Every week.

    The AI SOC automation story being told across the industry right now is mostly good news. Platforms are reaching 98 percent triage accuracy. Analysts are getting 40-plus hours of manual work back every week. Breach containment timelines are shrinking by 80 days. All of that is real and documented. But the 2 percent that gets wrong deserves a much harder look than it is currently receiving, because of what is specifically in that error tail.

    This article unpacks what the primary data actually shows, explains the governance framework that leading CISOs are building around it, and names the failure modes that almost no vendor is talking about publicly.


    The Numbers Behind the Headline

    The 43 percent figure in the headline sits comfortably within the documented range of AI triage automation rates across real enterprise deployments. It is a representative midpoint, not a single published statistic. Here is what the primary data actually shows:

    >98% Triage accuracy for CrowdStrike Charlotte AI, measured against Falcon Complete MDR expert decisions
    <2% Of 25 million enterprise alerts escalated to human analysts in Intezer’s 2026 dataset
    61% Reduction in analyst alert queue from AACT academic system across 3.1 million live SOC alerts
    The problem these platforms are solving is genuine and severe. Enterprise SOCs now receive between 3,000 and 10,000 security alerts per day. Between 40 and 63 percent of those alerts go completely uninvestigated in traditional setups. Ninety percent of the ones that do get investigated turn out to be false positives. The global cybersecurity workforce gap sits at 4.8 million unfilled positions, growing at 19 percent year-over-year. Seventy-one percent of SOC analysts report burnout. Sixty-four percent say they are considering leaving within a year.

    The human model of alert triage is structurally broken. AI SOC automation is not an efficiency preference at this point. For most enterprises, it is an operational necessity.

    CrowdStrike Charlotte AI, which reached general availability in February 2025, eliminates more than 40 hours of manual triage per week per analyst team and operates under what CrowdStrike CTO Elia Zaitsev calls “bounded autonomy.” The system does not act unilaterally. Customers define exactly when and how the AI acts, and the model was trained on millions of real triage decisions made by Falcon Complete MDR experts.

    “Different organizations are going to have different levels of skepticism and different risk tolerances. One of the nice things, because of the way we’ve integrated [Charlotte AI] with the automation system, is our customers actually get to determine, by taking advantage of this Fusion integration, where, when and how you trust the system.”

    Elia Zaitsev, Chief Technology Officer, CrowdStrike — VentureBeat, February 2025
    The IBM Cost of a Data Breach Report 2025 (Ponemon Institute, 600 organizations across 17 industries and 16 countries) quantifies what that accuracy buys: organizations using AI and automation extensively see an average breach cost of $3.62 million versus $5.52 million for those with no AI. That is a $1.9 million per-breach saving. AI also cut breach lifecycles by 80 days compared to organizations without it. Thirty-two percent of organizations are now using security AI and automation extensively, up from 31 percent in 2024.

    The efficiency case is not in dispute. The governance case is where things get complicated.


    What Actually Lives Inside the 2% Error Rate

    When an AI SOC system reports 98 percent accuracy, the immediate question any serious CISO should ask is: what is specifically in the 2 percent? Not in aggregate. Not blended with false positives that just wasted analyst time. What threats specifically are being missed?

    Intezer’s forensic data answers this with uncomfortable precision.

    Of the 82,000 endpoints that underwent live forensic memory scans in Intezer’s 2025 dataset, 2,600 had active infections. That alone is significant. But the finding that should change how every enterprise thinks about AI triage closure is this: 51 percent of those confirmed compromised endpoints had already been marked “mitigated” by the source EDR vendor. The machine had been declared clean. It was not clean.

    The malware families found active in memory on those “mitigated” endpoints were not proof-of-concept tools or research artifacts. They were Mimikatz, Cobalt Strike, Meterpreter, and StrelaStealer. These are active criminal and nation-state workhorses. They were sitting in memory, on machines that the security stack had officially declared safe, in environments where the AI was using EDR verdict as an input signal for closure decisions.

    Critical Finding
    1.6 percent of all forensic endpoint scans in Intezer’s 2026 dataset found active compromise despite EDR reporting “mitigated.” The AI did not invent the error. It inherited it from a flawed upstream input. This is the operational gap most AI SOC deployments are not designed to catch.

    This is a layered failure. The EDR declared the machine clean. The AI received that verdict as a trusted data point. The AI closed the alert. No human ever reviewed it. Cobalt Strike stayed in memory.

    Itai Tevet, CEO and co-founder of Intezer and former head of IDF cyber incident response, frames what this finding demands of security leadership:

    “Security teams have normalized the idea that some risk must be accepted because it is impossible to investigate everything. Our research shows that this acceptance is increasingly misaligned with how modern attacks unfold. When genuine threats consistently emerge from alerts we have trained ourselves to ignore, the definition of acceptable risk needs to be reexamined.”

    Itai Tevet, CEO, Intezer — GlobeNewswire, February 3, 2026
    The peer-reviewed academic data reinforces this picture from a different angle. The AACT system (Automated Alert Classification and Triage), deployed in a real managed SOC environment across 3.1 million live alerts over six months, achieved a false negative rate of 1.36 percent. That sounds small. At 3 million alerts, it represents 40,800 real threats that the system incorrectly closed. The precision of that number matters: it came from an independently published, peer-reviewed academic paper using actual production SOC data, not vendor-reported customer telemetry.

    At an enterprise receiving 10,000 alerts per day, a 2 percent blended error rate produces 200 wrong dispositions every single day. The critical question is whether those errors skew toward false positives (wasted time) or false negatives (missed threats). That calibration is not set by the AI vendor. It is a policy decision that the deploying organization must make explicitly, before deployment, based on its own risk tolerance.


    When AI Automation Attacks Its Own Network

    The failure mode that nobody wants to include in their AI SOC pitch deck happened at a real enterprise, and the documentation is on record.

    An enterprise AI-driven SOC response system was programmed to automatically isolate endpoints showing signs of compromise. A software update triggered false positives across hundreds of devices simultaneously, including critical production servers. The AI executed correctly according to its programming. It locked every flagged endpoint. The result was a self-inflicted denial-of-service attack on the organization’s own production infrastructure.

    The root cause investigation found something more troubling than a simple misconfiguration. Over time, the AI had been trained to ignore certain low-level anomalies that had repeatedly proved benign. That created a model drift blind spot. When new attack patterns emerged that resembled previously-benign behavior, the system missed them. The same suppression mechanism that reduced false positives also lowered the detection threshold for real threats that looked familiar.

    This is the automation complacency trap, and it is not unique to security AI. A 2024 peer-reviewed study from ETH Zurich found that human-in-the-loop designs increase uptake of AI recommendations but decrease overall accuracy. Participants were statistically less likely to intervene on the AI’s least accurate recommendations. The implication is that human oversight can create a false sense of verification without actually catching the errors it is supposed to catch.

    Key Insight
    Stale training data is now the leading cause of false positive spikes in AI SOC tools, according to the SANS 2025 SOC Survey. A system tuned to eliminate false positives compensates by raising its detection threshold, which also suppresses low-signal real threats. Attackers learn to look boring. The AI learns to ignore boring. This is not a theoretical concern. It is a documented, measurable attack surface.

    Vectra AI’s 2026 State of Threat Detection research found that 40 to 63 percent of alerts still go uninvestigated at organizations running traditional setups, and that stale model data is the primary driver of false positive inflation in AI-augmented environments. The AI SOC solves the volume problem. Model drift creates a new version of the coverage problem.

    There is no industry standard for AI model drift monitoring in SOC deployments. ISO/IEC 42001, the December 2023 global standard for AI management systems, requires continuous monitoring of AI decisions, but only 21 percent of enterprises have full visibility into their AI agent activities, according to Akto’s 2025 report. The remaining 79 percent are running models of unknown currency against an adversary landscape that evolves continuously.


    The Policy Framework Fixing Both Problems

    The governance answer emerging across serious enterprise deployments is not a binary choice between autonomous AI and human-reviewed everything. It is a tiered autonomy framework that assigns different levels of human oversight to different categories of action based on their risk profile.

    Gartner’s four-mode SOC maturity model, presented by analyst Kevin Schmidt at the Gartner SRM Summit 2025, defines the progression:

    Mode Description AI Role Human Role
    Mode 0 Manual operations None Everything
    Mode 1 Semi-automated (SOAR, playbooks) Predefined playbook execution Approves and monitors
    Mode 2 Augmented (AI copilot) Recommends; enriches context Approves all actions
    Mode 3 Autonomous agents Handles triage, hunting, some remediation Oversees; handles novel/high-stakes cases
    At Gartner’s Security Summit in June 2026, Gartner confirmed Mode 3 as the industry destination while explicitly warning about AI washing in vendor claims. Most enterprises currently operating on Mode 1 or early Mode 2 are being sold Mode 3 outcomes. The gap between those two things is exactly where the 2 percent problem lives.

    The operational architecture that tiered autonomy translates to in practice looks like this. Triage and enrichment run fully autonomous: the volume is high, the risk of an individual wrong decision is relatively low, and this is where the efficiency gains live. Containment actions require human approval: isolating an endpoint, blocking a network segment, or disabling a user account has real operational consequences if wrong. Remediation is human-executed: the blast radius of a wrong remediation action is too high to automate.

    Every AI action at every tier must be logged with an auditable reasoning chain. ISO/IEC 42001 compliance, NIS2 in Europe, and DORA for financial services are making this a regulatory requirement in addition to a governance best practice. CISOs in regulated industries who do not have AI governance documentation in place now are building compliance debt that will become costly to resolve under active regulatory scrutiny.

    Pete Shoard, VP Analyst at Gartner and the credentialed industry voice on this topic, has been consistent on where the line is:

    “If you think you can sack your SOC staff just because you’ve suddenly bought an AI function, I think you’re going to be soundly disappointed. AI won’t replace your security staff, so use it to enhance them and make them better in their jobs.”

    Pete Shoard, VP Analyst, Gartner — Cybersecurity Dive, Gartner SRM Summit, June 2025
    Shoard’s December 2024 Gartner research, “There Will Never Be an Autonomous SOC,” includes a warning that gets too little attention in vendor-led conversations: by 2030, 75 percent of SOC teams will experience erosion of foundational analysis skills due to AI over-dependence. The L1 analyst role, which is the training ground for senior investigators, disappears when AI handles Tier 1 and Tier 2 autonomously. A decade from now, when a truly novel threat requires human expert judgment, the pipeline of experienced analysts who would catch it may not exist.

    The TIAA CISO, Upendra Mardikar, distilled the enterprise buyer position at the same Gartner conference:

    “We don’t want complete autonomy. We have to have a human in the loop.”

    Upendra Mardikar, CISO, TIAA — Cybersecurity Dive, Gartner SRM Summit, June 2025

    Five Things CISOs Must Do Before Expanding AI Autonomy

    The Intezer and AACT data, combined with the Arctiq case study and Gartner’s maturity framework, point to five concrete operational changes that should precede any expansion of AI autonomy in a SOC environment.

    1. Stop Treating EDR “Mitigated” as a Closure Signal

    The finding that 51 percent of confirmed compromised endpoints were already marked “mitigated” by EDR is operationally decisive. Security teams must add a forensic verification layer for cases where AI systems are considering closure. The EDR verdict is one data point. It is not ground truth. Any AI triage architecture that treats EDR “mitigated” as a final state is inheriting the EDR’s error rate on top of its own.

    2. Define Bounded Autonomy Policies Before Deployment

    The “automation gone wrong” case, where an AI-triggered response created a self-inflicted denial-of-service attack, happened because containment policies were not defined before the system went live. The question of which actions AI can execute without approval, which require human sign-off, and which are never automated must be answered in writing, reviewed by legal and compliance, and tested against tabletop scenarios before any autonomous capability is activated in production.

    3. Track Mean Time to Conclusion for All Alerts, Not Just Escalated Ones

    If your AI resolves 98 percent of alerts and your MTTD and MTTR look excellent, but the 2 percent error includes real threats hiding in low-severity backlogs, your dashboard is measuring speed rather than coverage. Mean Time to Conclusion must be tracked across the entire alert population, including the cases the AI autonomously closed. Auditing a statistically significant sample of AI-closed alerts monthly is the minimum viable oversight practice.

    4. Build Model Drift Detection Into Your SOC AI Governance

    Stale training data degrades AI SOC accuracy on a timeline that no vendor will proactively disclose to you. Define a retraining cadence based on your threat landscape velocity. Instrument the system to alert when false positive or false negative rates shift outside defined thresholds. ISO/IEC 42001 requires continuous monitoring of AI decisions. Build that monitoring before you need it, not after a breach investigation reveals the drift window.

    5. Preserve the L1 Analyst Pipeline Deliberately

    If AI handles all Tier 1 triage, the entry-level analyst role that trains the next generation of senior investigators disappears. Organizations running Mode 2 or Mode 3 autonomy need a deliberate career development path that keeps analysts engaged with real investigation work, not just AI oversight. The Gartner prediction that 75 percent of SOC teams will erode foundational analysis skills by 2030 is not a passive forecast. It is a consequence of a specific architectural decision that can be reversed with equally specific policy.


    The Strongest Arguments Against the AI SOC Narrative

    This article would not meet its own standard if it did not engage seriously with the case against the mainstream AI SOC story. Here are the strongest objections, stated plainly.

    98 percent accuracy at scale is still a lot of wrong answers. At 10,000 daily alerts, 98 percent accuracy means 200 wrong triage decisions per day. The industry presents this as a success story. The correct question is: what is the false negative rate specifically, not the blended accuracy, and what types of threats are in that 2 percent? Advanced persistent threats and novel zero-day attacks are disproportionately likely to be in the error tail, because AI is trained on historical patterns and these threats are, by definition, outside historical patterns.

    Vendor accuracy claims have self-serving methodologies. CrowdStrike’s 98 percent accuracy is measured against Falcon Complete expert decisions, meaning it is measured against itself. Intezer’s 98 percent is self-reported from its own customer telemetry. Neither has been validated by an independent third party against ground truth attack data. The only truly independent figure in available primary data is the AACT academic paper, which found a 1.36 percent false negative rate over 3.1 million alerts in a single managed SOC environment with characteristics that may not generalize to every deployment.

    The AI creates a new, harder-to-find blind spot. A system tuned to eliminate false positives compensates by raising its detection threshold, which means it also starts suppressing low-signal real threats. Attackers learn to mimic the patterns the AI has been trained to ignore. This is not a theoretical concern. The SANS 2025 data confirms it is already happening. Stale model data is the leading driver of false positive spikes, which means organizations respond by raising the threshold further, which makes the blind spot larger.

    Our read: the enterprise case for AI SOC automation is sound. The efficiency gains are real, the cost data is credible, and the alternative (a human-only model drowning in 10,000 daily alerts) is not viable. But the governance case has to be built with the same rigor as the technical case, and right now the governance conversation is at least two years behind the deployment conversation.


    Frequently Asked Questions: AI SOC Automation

    What percentage of SOC alerts can AI automatically resolve?
    Real-world AI SOC platforms report autonomous resolution rates ranging from 61 percent (the peer-reviewed AACT system, 3.1 million alerts) to over 98 percent (Intezer, 25 million alerts). The range reflects differences in environment, alert type, and how “resolved” is defined. Enterprise deployments commonly target 40 to 60 percent automation as a conservative, auditable starting point before expanding autonomy.

    What happens when AI gets a SOC alert wrong?
    AI triage errors fall into two categories. False positives, meaning benign alerts incorrectly flagged, waste analyst time. False negatives, meaning real threats incorrectly closed, are the more dangerous failure. Intezer’s 2026 forensic analysis found that 1.6 percent of endpoints the AI cleared still had active Cobalt Strike or Mimikatz infections in memory. The correct policy response is tiered autonomy: AI handles routine closures independently, but containment actions require human approval.

    Will AI replace SOC analysts?
    No. Gartner’s December 2024 research explicitly titled “There Will Never Be an Autonomous SOC” states this is not a realistic outcome. AI automates Tier 1 and Tier 2 triage, eliminating repetitive alert-sorting work. Analysts shift to case validation, threat hunting, and AI oversight. The risk Gartner warns about is the opposite: by 2030, 75 percent of SOC teams may lose foundational analysis skills from over-reliance on automation.

    What is “bounded autonomy” in AI cybersecurity?
    Bounded autonomy means AI operates within customer-defined guardrails. Organizations control which triage actions the AI executes independently and which require human approval. CrowdStrike CTO Elia Zaitsev coined the term for Charlotte AI, launched February 2025. It sits between full autonomy (AI acts without human approval) and copilot mode (AI recommends; human always decides), and it is now the industry consensus model for production SOC AI deployment.

    How much money does AI save in security operations?
    IBM’s 2025 Cost of a Data Breach Report found that organizations using AI and automation extensively saved an average of $1.9 million per breach compared to those with no AI ($3.62 million versus $5.52 million average breach cost). They also cut breach lifecycles by 80 days. Faster detection means shorter dwell time, which directly reduces the total cost of a breach.

    What is the false negative rate of AI SOC systems?
    The most rigorous published figure comes from the peer-reviewed AACT system deployed in a real managed SOC: 1.36 percent false negative rate over 3.1 million alerts. CrowdStrike claims greater than 98 percent accuracy, implying roughly 2 percent combined error. Intezer reports 98 percent verdict accuracy across 25 million alerts. No vendor has published a standalone false negative rate independently verified by a third party.

    What is model drift in cybersecurity AI?
    Model drift occurs when an AI SOC system’s accuracy degrades because the threat landscape has changed but the model has not been retrained. Stale training data is the leading cause of false positive spikes in AI SOC tools, according to SANS 2025. In one documented case, an AI trained to dismiss certain low-level anomalies later failed to detect new attack techniques that resembled previously-benign behavior, creating a breach that a retrained model would have caught.

    What is tiered autonomy in a SOC?
    Tiered autonomy is the governance framework defining which SOC actions AI performs independently versus which require human approval. The consensus model: triage and enrichment are fully automated (high volume, low risk if wrong); containment actions require human sign-off (medium risk); remediation is human-executed (highest impact). Every AI action must be logged with an auditable reasoning chain for compliance with ISO/IEC 42001, NIS2, and DORA.


    What You Now Know That You Didn’t Before

    The AI SOC automation story is not a story about replacing human judgment. It is a story about redirecting it. AI handles the volume that was drowning analysts in noise. Analysts handle the cases that require genuine expertise. The failure is not in the model. The failure is in the governance architecture that surrounds it.

    The specific risk that the Intezer data exposes is not that AI makes mistakes. Every triage system makes mistakes. The risk is that AI mistakes are invisible by default. When a human analyst incorrectly closes an alert, there is a record of the reasoning. When an AI closes it, the reasoning is there too, but nobody is reviewing it. The 51 percent of confirmed compromised endpoints that were already marked “mitigated” by EDR represent exactly this failure: a machine trusted a machine, and Cobalt Strike sat in memory undisturbed.

    In the next 6 to 18 months, watch three things. First, whether the Gartner prediction about 30 percent of SOC leaders failing to integrate GenAI into production (due to hallucinations and governance gaps) materializes at the organizations that deployed most aggressively in 2025 without building the policy layer. Second, whether ISO/IEC 42001 and DORA enforcement creates a meaningful accountability mechanism for AI triage errors in financial services. Third, whether any vendor publishes independently verified false negative rates broken out by threat category, which would finally let buyers compare AI SOC platforms on the metric that actually matters.

    If you are a CISO making a SOC AI decision right now, the question is not whether to deploy. The question is whether you have defined, in writing, what your system is allowed to close on its own. If the answer is “we configured the vendor defaults and moved on,” you have inherited someone else’s risk tolerance on behalf of your organization.

    That is the policy this article is about.

  • CrowdStrike AI SOC Threat Detection 2026

    CrowdStrike AI SOC Threat Detection 2026

    AI Threat Detection Cuts Breach Costs by $1.9M. So Why Are 68% of Enterprise SOCs Still Flying Blind?
    AI Cybersecurity • Enterprise SOC

    AI Threat Detection Cuts Breach Costs by $1.9M. So Why Are 68% of Enterprise SOCs Still Flying Blind?

    Here is the problem, stated as plainly as possible. The average enterprise cybercriminal gains initial network access and begins moving laterally in 29 minutes. The average SOC analyst, working a manual triage queue packed with over 10,000 daily alerts, takes significantly longer than that just to confirm an alert is real.

    That is not a performance failure. That is a structural mismatch between the speed of modern intrusion and the design limits of human-pace security operations. And the data makes the gap measurable: AI-augmented SOC environments have demonstrated a 50% reduction in mean time to detect (MTTD) and a 60% drop in manual triage workload. Non-autonomous AI agents in documented deployments reduced investigation times from 30-plus minutes to under two minutes per incident.

    So the question this article sets out to answer is not whether AI threat detection works. The data on that is clear. The question is why approximately 68% of enterprise security operations centers are still not using it at scale.

    Key Data Point IBM’s 2025 Cost of a Data Breach Report found organizations using AI and automation extensively pay $3.62 million per breach on average. Those without pay $5.52 million. That $1.9 million gap is the largest single-technology cost difference IBM has ever recorded in this study’s history.

    The Detection Gap Nobody Wants to Admit

    Traditional SOC architecture was designed for a threat landscape that no longer exists. In the model that most enterprises still run, Tier 1 analysts review alerts manually, escalate to Tier 2 for investigation, and escalate further to Tier 3 for complex incidents. This model worked when attacks unfolded over hours or days. It doesn’t work when the initial-access-to-lateral-movement window is measured in minutes.

    The alert volume problem compounds this. Modern enterprise SOCs process an average of 10,000 or more alerts per day, with false positive rates hovering around 45%. The SANS Institute’s 2025 survey found 73% of security teams cite false positives as their primary detection challenge, not insufficient tooling, not budget. False positives. The noise is so overwhelming that up to 40% of alerts go uninvestigated entirely.

    Analyst burnout cycles average 18 months before turnover. That number tells you everything about what it means to be a Tier 1 SOC analyst in 2026: you are drowning in 100,000-plus daily alerts where between 1% and 5% are real threats, you cannot distinguish signal from noise fast enough to matter, and the job grinds people down until they leave.

    10,000+ Average daily alerts per enterprise SOC, with a 45% false positive rate
    40% Share of security alerts that go completely uninvestigated
    18 mo. Average analyst burnout cycle before SOC Tier 1 turnover
    50% MTTD reduction demonstrated by AI-augmented SOC operations
    This is the structural problem that AI threat detection is designed to solve. Not by replacing analysts. By absorbing the volume of mechanical triage work that is consuming their capacity and preventing them from doing the judgment-based work only they can do.


    When the Attacker Moves in 29 Minutes

    The CrowdStrike 2026 Global Threat Report, published February 24, 2026, documents something that should recalibrate how every CISO thinks about incident response timelines.

    The average eCrime breakout time in 2025, defined as the elapsed time from initial access to lateral movement, dropped to 29 minutes. That represents a 65% increase in attacker speed from 2024. The fastest observed intrusion moved from access to lateral movement in 27 seconds. In one documented case, data exfiltration began within four minutes of initial compromise.

    “This is an AI arms race. Breakout time is the clearest signal of how intrusion has changed. Adversaries are moving from initial access to lateral movement in minutes. AI is compressing the time between intent and execution while turning enterprise AI systems into targets. Security teams must operate faster than the adversary to win.” Adam Meyers, Head of Counter Adversary Operations, CrowdStrike
    The 29-minute average is an organizational benchmark, not a theoretical worst-case. If your incident response workflow takes longer than 29 minutes from detection to analyst action, you have already ceded the lateral movement window to the attacker. In a significant share of intrusions, the attacker has established persistence and begun moving toward their objective before the alert even surfaces in the SOC queue.

    The attacker speed story gets worse when you consider what those attackers are now equipped with. AI-enabled adversary operations increased by 89% year-over-year in 2025. And 82% of detections in 2025 were malware-free, meaning adversaries used valid credentials and trusted identity flows to move through networks without triggering traditional signature-based detection.

    The Identity Shift Changes Everything When 82% of intrusions use valid credentials rather than malware, traditional endpoint detection loses most of its relevance. The attack surface has shifted to identity and behavior. AI threat detection that correlates behavioral anomalies across identity, endpoint, and network simultaneously is not optional. It is the only architecture that matches this threat model.
    AI-generated phishing reduced attack preparation time from 16 hours to 5 minutes (IBM 2025 data). That’s not an incremental efficiency gain for attackers. It is mass-personalized social engineering at industrial scale. The volume increase this enables on the offensive side directly translates to the alert volume problem on the defensive side.


    The Adoption Paradox: The Advantage Exists. Most Aren’t Using It.

    IBM’s 2025 Cost of a Data Breach Report surveyed 604 organizations across 17 industries and 16 countries. Only 32% report using AI threat detection and automation extensively in their security programs. A separate Anvilogic survey conducted in collaboration with the SANS Institute found 45% of respondents have integrated AI into their threat detection workflows, but “integration” in many cases means a limited deployment in one tool category, not a systematic AI-augmented SOC architecture.

    That leaves a majority of enterprise security operations running detection workflows that are structurally outpaced by the attacker speed documented above.

    What’s behind that gap? The research points to four primary barriers, and they are not the ones most vendors would have you believe.

    Barrier 1: Trust and Explainability

    McKinsey’s March 2026 survey of approximately 500 organizations found nearly two-thirds cite security and risk concerns as the top barrier to fully scaling AI security systems, ahead of regulatory uncertainty and technical limitations. The cost or complexity of AI platforms ranked below trust.

    “AI can discover anomalies faster, but adoption does not automatically create trust. The challenge is that too often, AI produces answers without showing its work. In the SOC, trust has always been built on verifiable evidence that stands up to scrutiny. Analysts move forward when they can see the data, understand the connections, and explain the reasoning behind a decision. AI earns its place in the SOC the same way: by making its insights clear, traceable, and grounded in proof.” Kyle Pearson, Global Solutions Architect, Graylog • Security Boulevard, March 2026
    This is not irrational resistance to change. When an AI system flags a threat and an analyst cannot trace the reasoning path, they face a binary choice: act on an alert they cannot verify, or ignore it. Most analysts default to skepticism, which means the AI detection advantage is wasted at the last mile of the workflow.

    Barrier 2: Alert Volume Gets Worse Before It Gets Better

    Adding AI detection layers without proper tuning can increase alert volume before it decreases it. During transition periods, organizations run legacy rule-based detection alongside the new AI system, generating duplicate alerts and compounding the false positive problem. Most organizations underestimate the tuning timeline and the temporary analyst workload spike that comes with it.

    Barrier 3: Governance Gaps Create New Exposure

    IBM and the Ponemon Institute found that 97% of organizations that experienced an AI-related security incident lacked proper AI access controls. And 63% of organizations have no AI governance policies in place. This is the governance paradox of 2026: organizations know AI is the answer to their SOC capacity problem, but deploying AI without governance infrastructure recreates the same exposure problem at a different layer. The security team’s own AI infrastructure becomes an attack surface.

    Barrier 4: Budget Politics, Not Technology Readiness

    Only 11% of security professionals trust AI completely for mission-critical tasks, per Splunk’s 2025 State of Security survey (n=2,058). That number is worth interrogating carefully. It is not a statement that AI threat detection doesn’t work. It is a statement about organizational trust, procurement cycles, and the difficulty of attributing breach prevention to a tool that works by stopping things before they escalate.

    Our read: the budget and trust barriers are linked. Security teams that cannot demonstrate clear ROI from AI SOC investments face annual budget battles they often lose. IBM’s $1.9 million per-breach savings figure is the most powerful counter-argument available, but it requires a breach to make the case in retrospect.


    The Financial Stakes Are No Longer Theoretical

    IBM’s 2025 Cost of a Data Breach Report provides the clearest financial framework for the AI SOC adoption decision. These numbers are not projections or vendor estimates. They come from an activity-based costing methodology applied to 604 real organizations with documented breaches.

    Organization Type Avg. Breach Cost Detection Timeline
    Extensive AI + automation users $3.62 million 80 days faster than average
    No AI or automation $5.52 million Baseline
    US organizations (average) $10.22 million US record high
    Global average (2025) $4.44 million 241-day mean identify + contain
    The 241-day mean time to identify and contain a breach is actually an improvement: it is the lowest in nine years, driven by faster breach containment powered by AI among organizations that have adopted it. The organizations without AI are dragging that average upward.

    For US enterprises specifically, the $10.22 million average breach cost is a record. Building and maintaining a full in-house 24/7 SOC runs $2 to $2.5 million per year in staffing alone, before SIEM licensing, EDR tools, or management overhead. The AI investment conversation needs to happen inside that cost context, not against it.

    The ROI Calculation CISOs Are Missing The $1.9 million average saving per breach for extensive AI users is not a ceiling. It does not account for reputational damage, regulatory penalty avoidance, or the compounded value of the 80-day reduction in attacker dwell time. Organizations using AI are containing breaches before attackers can maximize damage. Non-users are paying for the full extent of attacker access.

    The Workforce Math Doesn’t Work Without AI

    The global cybersecurity workforce gap stands at approximately 4.8 million unfilled positions. The total workforce needed globally is 10.2 million, against 5.5 million currently employed. The US alone has 750,000 empty cybersecurity roles.

    Those positions are not going to be filled by traditional hiring. The pipeline for trained security professionals cannot be expanded fast enough to close a 4.8 million person gap, and the burnout cycle means that even the analysts you do hire are leaving within 18 months of experiencing the alert volume of a modern SOC.

    Gartner projects that more than 50% of SOC Tier 1 analyst responsibilities will be handled by AI by 2028. That projection is not a threat to analyst careers. It is a necessary architectural shift that frees human analysts from the mechanical work that is burning them out and preventing them from doing the higher-judgment work that actually requires human reasoning.

    “Organizations are already seeing efficiency gains of roughly 40 to 50% for lower-tier SOC tasks, freeing human analysts to focus on more advanced investigations and response activities.” Martin Sordilla, Senior Technology and Security Architect, Accenture • CSO Online, April 2026
    The practical implication: a team of 10 analysts augmented with AI can cover the workload that would previously have required 18 to 20 analysts. In a market where those 8 to 10 additional analysts simply may not be available, AI is not a competitive advantage. It is the only viable operational model.

    This connects directly to how the cybersecurity analyst role is evolving alongside AI tools. The demand for analysts is not disappearing. It is shifting toward the strategic, judgment-based work that AI cannot automate.


    The Honest Counterargument: Why Skepticism Is Legitimate

    The case for AI SOC adoption is strong. But the skeptics are not wrong about everything, and enterprise security teams deserve a version of this argument that doesn’t paper over the real risks.

    The “Seconds” Claim Needs Qualification

    When AI threat detection is described as identifying anomalies in seconds, that framing refers to alert generation, not analyst-confirmed response. An alert that fires in seconds and sits unreviewed in a queue for six hours still represents a six-hour window of attacker opportunity. The metric looks good. The actual detection performance was poor. AI earns MTTD credit when it reduces the time to analyst action, not just the time to alert generation.

    The Same AI Infrastructure Gets Targeted

    CrowdStrike’s 2026 report documents prompt injection attacks against enterprise AI tools across more than 90 organizations. ChatGPT was mentioned in criminal forums 550% more than any other AI model. The AI infrastructure being deployed for defense is actively being targeted by adversaries who have learned to weaponize it. Deploying AI SOC capabilities without AI governance frameworks simultaneously opens a new attack surface. This is not an argument against AI adoption. It is an argument for deploying governance alongside the technology, not after it.

    Implementation Failure Rates Are Real

    A 2026 enterprise AI adoption survey (n=2,400) found 79% of organizations face significant challenges in adopting AI. Only 29% see meaningful ROI from generative AI despite individual productivity gains. Purchasing an AI SOC platform and achieving operational security value from it are very different outcomes separated by months of integration, tuning, and workflow redesign. The 6 to 24 month deployment timeline to operational maturity is not a vendor warning label. It is the realistic planning horizon CISOs need to build into their roadmaps.

    “The first question enterprises ask about AI SOC isn’t ‘how fast is it?’ It’s ‘can we trust it?’ That question deserves a serious answer. Explainability, auditability, and clear escalation paths aren’t nice-to-haves. They’re the difference between AI that improves your SOC and AI that introduces new risk into it. Scale without accountability isn’t efficiency. It’s a different kind of risk.” Enterprise Security Practitioner, cited in Prudent Consulting Cybersecurity Priorities Report, May 2026
    This concern about governance sits at the intersection of the generative AI threats facing enterprise security teams and the AI deployment challenges covered in depth by IBM’s threat research. The responsible AI SOC conversation has to include both the offensive capabilities of AI and the defensive governance structures that keep deployed AI from becoming a liability.


    What a Real AI SOC Actually Looks Like

    An AI SOC is not a product. It is an operational model, and the distinction matters. Organizations that treat it as a product purchase and discover that tuning, integration, and workflow redesign are the actual work are the ones with 79% implementation challenge rates.

    The operational model that practitioners are documenting in 2026 follows a tiered autonomy structure:

    Function Who Handles It Why
    Alert triage, enrichment, correlation AI (autonomous) Volume too high for human triage; pattern matching is AI-native
    Initial investigation and classification AI with human review AI surfaces evidence; analyst confirms before escalation
    Containment decisions Human approval required High-stakes action with potential false-positive consequences
    Complex incident response Human-led, AI-assisted Novel threats, strategic decisions, stakeholder communication
    Post-incident learning and tuning Human-led Requires contextual judgment to reduce future false positives
    The 70%-plus of attacks that occur outside traditional business hours are the clearest argument for AI handling the autonomous triage layer. A human analyst is not reading alerts at 3 a.m. with the same speed and accuracy as a system that never tires, never has a bad night, and applies the same detection logic to every alert regardless of shift timing.

    Vendors with documented case studies in this space include CrowdStrike Falcon, Palo Alto XSIAM, Microsoft Sentinel with Copilot for Security, SentinelOne Singularity, and UnderDefense. The choice of platform matters far less than the design of the autonomy tiers and the governance framework governing escalation paths.

    The regulatory pressure to get this right is accelerating. NIS2 is in active enforcement, with approximately 19,000 companies estimated non-compliant as of March 2026. DORA is in effect for financial services. The EU AI Act moves to full enforcement from August 2026. Organizations that have been deferring AI SOC decisions as a technology question will discover it has become a compliance question. The timeline context connects to the broader regulatory timeline enterprises are navigating on multiple security fronts simultaneously.


    What CISOs Should Do This Quarter

    The argument that “we’re waiting for the technology to mature” is no longer available. AI threat detection platforms exist at commercial maturity, vendor case studies document real deployments, and the regulatory pressure is live. These are the decisions that need to happen now.

    1. Map your SOC workflows against the 29-minute window. If your end-to-end detection-to-analyst-action time exceeds the average eCrime breakout time, every intrusion is potentially a full lateral movement event before your team engages. Identify specifically where AI triage would compress that timeline.
    2. Separate the autonomy decision from the vendor decision. Decide what your AI should be allowed to do autonomously before you evaluate which platform does it. Organizations that buy a platform first and design governance after tend to lock in the wrong architecture.
    3. Treat explainability as a non-negotiable procurement criterion. Evaluate any AI SOC platform on whether analysts can trace the reasoning behind alerts. Black-box AI fails at the last mile regardless of detection accuracy. XAI-integrated platforms that show confidence scores, contributing features, and attribution paths build the analyst trust that sustains adoption.
    4. Build AI governance before you deploy AI detection. The 97% of AI breach victims who lacked proper AI access controls made their AI infrastructure a liability. Governance frameworks for your deployed AI are not a Phase 2 item. They are a prerequisite for Phase 1.
    5. Watch the 2028 Gartner projection as a planning horizon. If 50%+ of Tier 1 responsibilities shift to AI by 2028, your current staffing model, your training pipeline, and your incident response playbooks all need to be redesigned for that operating reality. The planning window is now, not when the transition is already underway.
    The organizations that document clear operational results from AI SOC deployments this year will have 12 to 18 months of institutional learning before the late majority begins their implementations. In the 2026 threat landscape, that compounding advantage in detection speed and analyst capacity is not incremental. It is strategic. The real-world breach consequences for organizations without that advantage are documented and public.


    Frequently Asked Questions

    How does AI detect cyberattacks faster than human analysts?

    AI threat detection processes millions of log events simultaneously, applying behavioral anomaly detection in real time rather than waiting for signature matches or analyst review. AI-augmented SOCs reduce mean time to detect by 50% versus manual operations and correlate cross-domain signals in seconds while human analysts handle triage sequentially, one alert at a time. The speed advantage compounds at high alert volumes where human capacity breaks down entirely.

    What is the average time for a SOC analyst to detect an intrusion without AI?

    Without AI augmentation, mean time to detect (MTTD) ranges from hours to days for sophisticated intrusions. Mandiant’s M-Trends 2025 places median attacker dwell time at 11 days globally. IBM’s 2025 Cost of a Data Breach Report found the average breach takes 241 days to identify and contain. AI-augmented SOCs have reduced investigation times from 30-plus minutes to under two minutes per incident in documented deployments, with AI users detecting and containing 80 days faster on average.

    Why aren’t more enterprises using AI for cybersecurity?

    The top barriers are trust and explainability, not cost or technology readiness. McKinsey’s March 2026 survey of approximately 500 organizations found nearly two-thirds cite security and risk concerns as the primary obstacle to scaling AI security systems. Budget constraints, integration complexity, and governance gaps follow closely. Many organizations also underestimate the 6 to 24 month tuning and integration timeline required to reach operational maturity.

    How fast do cyberattacks move in 2026?

    CrowdStrike’s 2026 Global Threat Report documents the average eCrime breakout time at 29 minutes, a 65% speed increase from 2024. The fastest observed intrusion moved from initial access to lateral movement in 27 seconds. In one documented case, data exfiltration began within four minutes of initial compromise. At scale, 82% of 2025 detections were malware-free: attackers used valid credentials and trusted identity flows, bypassing traditional signature-based detection entirely.

    How much does AI reduce cybersecurity breach costs?

    IBM’s 2025 Cost of a Data Breach Report found organizations using AI and automation extensively incur $3.62 million per breach versus $5.52 million for non-users, a saving of $1.9 million per incident. This is the largest single-technology cost reduction IBM has measured in the study’s history. US organizations face an average breach cost of $10.22 million, a record high, making the AI investment calculation increasingly straightforward for American enterprises.

    Can AI replace SOC analysts?

    No. AI handles triage, enrichment, correlation, and alert classification: the mechanical workload that is currently consuming analyst capacity and accelerating burnout. Analysts handle complex investigation, containment decisions, novel threat response, and stakeholder communication. Gartner projects AI will handle more than 50% of Tier 1 SOC responsibilities by 2028. The consensus operational model is human-supervised AI augmentation, not replacement, and the 4.8 million global workforce gap makes that augmentation structurally necessary.

    What is an AI SOC?

    An AI SOC (Security Operations Center) is an operational architecture where AI handles alert triage, enrichment, and cross-tool correlation at scale, while human analysts supervise critical decisions and execute containment. It is not a single product but a tiered autonomy model that enables 24/7 detection coverage without proportionally scaling headcount. The key design decision is which functions operate autonomously, which require human review, and which require human approval before action.


    What You Now Understand That Changes the Conversation

    The AI SOC adoption gap is real, but it is not a story about technology laggards. It is a story about a legitimate set of governance, trust, and implementation challenges that most vendors have strong incentives to downplay. The organizations that close the gap successfully do so not by buying the fastest AI threat detection platform but by designing the right autonomy tiers, building governance infrastructure before deployment, and investing in explainable AI that earns analyst trust at the last mile of the workflow.

    The next 12 to 18 months will likely define which enterprises have the institutional AI SOC capabilities to operate at attacker speed and which are still designing the framework. By 2030, AI-first SOC operations will be the global standard. The organizations still running human-pace triage workflows against AI-accelerated adversaries will not fail because the technology wasn’t available. The technology is available now.

    Three things to track: the EU AI Act enforcement calendar from August 2026 and how it changes AI governance requirements for deployed security systems; the Gartner 2028 Tier 1 automation projection and whether enterprise procurement cycles are moving fast enough to meet it; and whether XAI (explainable AI) design becomes a competitive differentiator among SOC platform vendors or remains an afterthought. The trust problem Pearson and others describe will not resolve itself without explicit explainability engineering.

  • Arup Deepfake Scam: Inside the $25M CEO Fraud Case

    Arup Deepfake Scam: Inside the $25M CEO Fraud Case

    Deepfake CEO Fraud: Arup’s $25M Wake-Up Call | NeuralWired
    Enterprise Security

    Deepfake CEO Fraud: Arup’s $25M Wake-Up Call

    A finance employee at the global engineering firm Arup joined a video call with five colleagues, including the company’s UK based chief financial officer. Every person on that screen except him was an AI generated fake. Over the following weeks he approved 15 wire transfers totaling HK$200 million, roughly $25 million, to bank accounts the criminals controlled. That is deepfake CEO fraud, and it stopped being a one-off curiosity the moment the FBI started tracking it as its own crime category. The real lesson from Arup has less to do with spotting a fake face on a screen and more to do with who, inside your company, is allowed to approve a transfer in the first place.

    If you sit anywhere near a payment approval chain, in finance, security, or the general counsel’s office, this is the case study worth understanding properly. And the fix is cheaper, and far less exotic, than most detection vendors would like you to believe.

    What Actually Happened at Arup

    The attack didn’t start with a video call. It started with an email, supposedly from Arup’s UK based CFO, requesting a confidential transaction. The employee who received it suspected phishing and didn’t act on it immediately, which is exactly the instinct security teams spend years trying to train into staff.

    Then came the follow up: an invitation to a video conference where the CFO and several other colleagues appeared to be present. They weren’t. Every other participant on that call had been recreated using publicly available video and audio of the real executives, including footage from internal company meetings. Convinced he was speaking with real leadership, the employee proceeded to authorize 15 separate transfers to five Hong Kong bank accounts, totaling HK$200 million.

    Nobody caught it in real time. The fraud only surfaced when the employee later checked in with Arup’s head office. Hong Kong police disclosed the case publicly on February 2, 2024, with senior superintendent Baron Chan Shun-ching giving the on record account. Arup confirmed in May 2024 that it was the company involved, telling press that “fake voices and images” were used and that attacks of this kind had been rising sharply in sophistication. As of the most recent reporting available, no arrests have been announced and the funds haven’t been recovered.

    “Once fraudsters start making money, they fuel their fraud components with that funding.” Matthew Miller, Principal, Cybersecurity Services, KPMG US · via CFO Dive
    Miller’s point, made shortly after Arup went public, is the uncomfortable economic logic underneath all of this: deepfake fraud isn’t a novelty attack run by a handful of specialists. It’s profitable enough now to fund its own expansion.

    Arup Wasn’t First, and It Won’t Be Last

    The Arup case gets the headlines because of its scale and its use of live video, but it sits on a timeline that stretches back further than most coverage admits, and continues well past it.

    DateCaseLossWhat Made It Notable
    March 2019UK energy firm (via German parent company impersonation)€220,000 (~$243K)First widely documented AI voice clone CEO fraud
    January 2024Arup, Hong Kong~$25MFirst major case using a live, multi person video deepfake
    January 2026Entrepreneur in canton Schwyz, Switzerland“Several million” Swiss francsVoice deepfake sustained across a two week call sequence
    April 2026FBI IC3 2025 Annual Report$893M (AI related fraud, all categories)First year “AI related” tracked as a formal crime descriptor
    There’s also a quieter 2020 case, cited in academic research on deepfake detection, where a Hong Kong bank manager authorized $35 million in transfers after a deepfake phone call impersonating a company director he’d actually spoken with before. It got far less press than Arup, but it tells you the same trick worked years before anyone had a name for it.

    The Swiss case in January 2026 matters for a different reason: it confirms this hasn’t tapered off since Arup made headlines. It’s also worth saying plainly what the data does not support: there’s no verified cluster of three additional named, dollar confirmed enterprise deepfake cases within 90 days of any single incident. Several vendor blogs imply otherwise with vague “more cases followed” framing that doesn’t trace back to primary reporting. Be skeptical of round, dramatic numbers in this space that don’t link to a named source.

    The Numbers: How Big Is This, Really

    Individual cases make for a good story. The aggregate numbers are what should actually change how your company approves money.

    StatisticSourceDate
    62% of organizations hit by at least one deepfake incident in 12 monthsGartner survey of 302 security leadersSept. 2025
    $893M in AI related fraud losses reported to the FBIFBI IC3 2025 Annual ReportReleased April 2026
    1,300% surge in deepfake fraud attempts at enterprise contact centersPindrop, analysis of 1.2B+ calls2024 data, June 2025 report
    73% human accuracy detecting AI speech deepfakes by earPeer reviewed listening study, NCBI/PMCPublished study
    87% of finance staff would process a payment if “called” by their CEO or CFOMedius Financial Census, 1,533 respondentsJune 2024
    $20.9B total IC3 reported cybercrime losses (AI fraud is ~4% of that)FBI IC3 2025 Annual Report2025 (reported 2026)
    The 62% figure from Gartner is the single most useful “how common is this” data point in the field right now, because it comes straight from the analyst firm’s own release rather than a secondhand paraphrase.

    “Employees really are on the frontline of trying to spot something unusual.” Akif Khan, Senior Director Analyst, Gartner Research
    Worth flagging a number that often gets misused: the widely cited $1.1 billion figure for total US deepfake fraud losses in 2025 (a tripling from $360 million in 2024, per Surfshark’s analysis) is mostly driven by something different than what happened at Arup. Roughly 80% of that total comes from celebrity and executive impersonation investment scams spread through social platforms like Facebook, WhatsApp, and Telegram, not targeted B2B wire fraud against a single employee. Conflating the two makes for a scarier headline, but it’s the wrong comparison.

    And one honesty check on scale: AI related fraud, at $893 million, is still roughly 4% of the FBI’s total $20.9 billion in 2025 cybercrime losses. The FBI itself flags this as a likely undercount, since most victims don’t identify the AI component when they file a complaint. That cuts both ways: the real number could be higher, but claiming certainty about “how big this is right now” overstates what the data actually shows.

    Why You Can’t Detect Your Way Out of This

    The instinct, understandably, is to fight AI with AI: buy a tool that flags synthetic voices and faces before anyone wires money. The evidence says that’s not where the real advantage sits, at least not yet.

    In a controlled listening study covering both English and Mandarin speech, human listeners correctly identified AI generated speech deepfakes only 73% of the time, even after being shown examples beforehand. That’s well above the unsourced “24.5% detection rate” figure that circulates on vendor blogs without a clear citation trail (treat that number as unverified if you’ve seen it elsewhere), but 73% is still nowhere near reliable enough to bet a wire transfer on.

    Automated detection isn’t meaningfully better yet. Gartner’s own newer research on deepfake heavy social engineering warns that detection remains probabilistic and that benchmarks lag behind how fast generation tools improve. In practical terms: any vendor promising a near perfect detection rate today is selling you a number that won’t hold up against next year’s model.

    A small case study in misinformation, inside a misinformation story Two security blogs published in March 2026 describe the Arup attack as happening “in September 2025.” It didn’t. The verified date, confirmed by Hong Kong police and reported by outlets including CFO Dive, is January 2024. Nobody seems to have made this up maliciously. It’s more likely that someone paraphrased a paraphrase, the date drifted, and search engines rewarded the version that ranked. If a foundational fact like a case’s date can mutate this easily in cybersecurity reporting, it’s worth asking what else has drifted by the time a stat reaches your inbox.
    A more defensible architecture, one NeuralWired has covered separately in our guide to zero trust security, treats every request as unverified by default rather than trying to spot the fake in real time. That principle, “never trust, always verify,” is exactly what the next section is built on.

    The Real Fix: Kill the Trust, Not the Deepfake

    Here’s the uncomfortable part. The Arup fraud worked not because the deepfake was flawless, but because the company’s process let one employee’s belief, however reasonably formed, authorize a $25 million transfer.

    Medius surveyed 1,533 finance professionals across the US and UK and found that 53% had already been targeted by a deepfake scam, and 43% had fallen for one. The number that should worry every CFO most: 87% admitted they would process a payment if “called” by their CEO or CFO, and 57% of finance professionals can authorize transactions independently, without a second approval.

    “Scammers are creating fake audio clips of CEOs and CFOs.” Ahmed Fessi, Chief Transformation & Information Officer, Medius
    Fessi’s broader point is that executives generate their own attack surface just by doing their jobs: earnings calls, conference panels, YouTube interviews, LinkedIn videos. All of it is raw material. You can’t stop a CEO from giving an earnings call. You can stop a single voice, no matter how convincing, from being sufficient authorization to move money.

    A research team at the security publication DeepStrike makes the contrarian case worth sitting with: a basic rule requiring callback verification through a pre-registered phone number before any high value transfer “would have stopped the attack cold,” regardless of how perfect the deepfake was. Their broader argument pushes back on the industry’s heavy spend on detection tooling, suggesting companies stop trying to turn every employee into a forensic audio analyst and instead fix the approval workflow itself.

    The minimum viable version of that fix looks like this:

    • Out-of-band callback verification for any urgent, confidential, or high value transfer request, using a number pulled from an internal directory, never one given during the suspicious call itself.
    • Dual authorization above a fixed dollar threshold, removing any single employee’s ability to independently move large sums, regardless of how senior the request appears to come from.
    • A documented “no exceptions” policy that survives social pressure, including a fake executive expressing urgency or annoyance about the delay.
    • Scenario based simulation, using mock deepfake calls rather than slide deck training, since the exploit here is authority compliance, not unfamiliarity with the concept of deepfakes.

    What This Means for Your Team

    For Finance and Treasury Teams

    If you can independently authorize a wire transfer today, that’s the gap an attacker is counting on, not a convenience worth keeping. Push for mandatory dual sign off and a documented callback policy before your company becomes the next case study, not after.

    For CISOs and Security Leaders

    Annual phishing-style training has shown limited effect on deepfake susceptibility specifically, because the vulnerability is trust in authority, not unfamiliarity with the attack format. Live simulation exercises are cheap relative to a detection tool purchase, and they target the actual failure point. Also worth checking: how your cyber insurance policy classifies this. Deepfake enabled wire fraud is typically bucketed as social engineering fraud, which many standard policies exclude or cap well below data breach coverage.

    For General Counsel and Compliance

    Regulatory exposure is shifting. The FCC’s 2024 ruling that AI generated voices count as “artificial” under the Telephone Consumer Protection Act, the FTC’s 2024 rule banning AI impersonation of businesses, Tennessee’s ELVIS Act, and the FBI naming “AI related” as a formal crime category all point the same direction: a company that suffers a loss without a documented verification protocol will have a harder time in a regulatory or insurance dispute than one that had a tested process, even if that process failed once. Our earlier coverage of the FBI’s IC3 guidance on ransomware prevention walks through how documented controls increasingly shape post-incident outcomes, and the same logic now applies here.

    The Honest Limit of Prevention

    None of this eliminates the underlying problem. Callback verification, dual authorization, zero trust workflows, all of it addresses the moment of the transfer. None of it touches the first stage of these attacks: reconnaissance using an executive’s own public footage.

    CEOs and CFOs can’t realistically stop giving interviews, earnings calls, or conference talks. That means the raw material for cloning a voice or a face will keep accumulating no matter how tight your internal controls get. The honest conclusion isn’t that this is solvable. It’s that companies can meaningfully cut the success rate of these attacks through process design, while accepting that the vulnerability itself, executives having public voices and faces, isn’t going away.

    Frequently Asked Questions

    How does deepfake CEO fraud actually work?

    Attackers gather public audio and video of an executive from earnings calls, interviews, or conference talks, then generate a synthetic voice or video. They contact an employee, often through a spoofed email followed by a live or recorded video call, impersonating the executive to authorize an urgent, confidential wire transfer.

    How much money did Arup lose in its deepfake scam?

    A Hong Kong finance employee at Arup transferred HK$200 million, roughly $25 million, across 15 transactions in January 2024 after a video call where the CFO and several colleagues were entirely AI generated. The case was disclosed by Hong Kong police on February 2, 2024.

    Can you actually detect a deepfake voice or video call?

    Not reliably. A peer reviewed listening study found humans correctly identify speech deepfakes only about 73% of the time, and Gartner’s own research warns that automated detection remains probabilistic, with benchmarks that lag behind how fast deepfake generation tools improve.

    How do companies protect themselves against deepfake CEO fraud?

    The most effective defense is procedural, not technological: requiring independent, out-of-band verification, such as a callback to a pre-registered phone number, for any high value transfer requested by voice or video, no matter how convincing the request sounds or looks.

    Does cyber insurance cover deepfake fraud?

    It depends on the policy. Deepfake enabled wire fraud is typically classified as social engineering fraud, which many standard cyber insurance policies exclude or cap at lower limits than data breach coverage. Companies should review their specific social engineering and funds transfer sub-limits now.

    How many organizations have experienced a deepfake attack?

    62% of organizations reported experiencing at least one deepfake related incident, whether social engineering or exploitation of automated identity verification, in the prior 12 months, according to a Gartner survey of cybersecurity leaders released in September 2025.


    Where This Goes Next

    What Arup actually proves isn’t that AI fakes are unbeatable. It’s that most companies still let a single, well meaning employee’s judgment stand between a convincing phone call and a multi million dollar wire transfer. Fix that approval chain and the sophistication of the deepfake stops mattering nearly as much.

    Three things worth watching over the next 6 to 18 months:

    • H.R. 1734, the Preventing Deep Fake Scams Act, which would stand up a Treasury led task force on AI financial fraud best practices. Its progress is worth tracking as a signal of where federal policy lands.
    • Cyber insurance language. Watch for insurers introducing explicit deepfake or synthetic media sub-limits, separate from general social engineering fraud coverage, as claims data accumulates.
    • The FBI’s 2026 IC3 report, due in early 2027, which will be the first year-over-year comparison for the “AI related” crime category and should clarify whether $893 million was a baseline or an outlier.
    The pattern so far is consistent: the technology keeps getting better, and the fraud keeps working through the same gap in approval process. That gap is the one part of this problem any company can close this quarter, without buying a single piece of detection software.

    Stay Ahead of the Next Wire Fraud Headline

    Get NeuralWired’s analysis on AI security threats, straight to your inbox, before they make headlines.

    Subscribe to The Neural Loop
  • Google’s 2029 Post-Quantum Cryptography Deadline

    Google’s 2029 Post-Quantum Cryptography Deadline

    Google Post-Quantum Cryptography 2029 Deadline Explained
    Enterprise Security · Post-Quantum Cryptography

    Google Just Moved Its Quantum Deadline to 2029

    Your cryptography migration roadmap probably says 2035. Google’s doesn’t anymore. On March 25, 2026, Google set a new post-quantum cryptography 2029 deadline for its own systems, six years ahead of the federal backstop most enterprise security teams have been planning around for the past two years.

    That’s not a marketing decision. It’s a response to math. Six days later, Google Research published the resource estimates behind it, and they’re the kind of numbers that make a CISO reread an email twice (and then forward it straight to the budget committee).

    If you’re responsible for cryptographic risk at your organization, here’s exactly what happened, what the new research means and doesn’t mean, and what NIST’s separate, still-unfinished post-quantum cryptography deadline requires of you in the meantime.

    The Announcement That Moved the Goalposts

    Google’s announcement came from Heather Adkins, VP of Security Engineering, and Sophie Schmieg, Senior Staff Cryptography Engineer, in a post titled “Quantum frontiers may be closer than they appear.”

    “We’re setting a timeline for post-quantum cryptography migration to 2029.” Heather Adkins & Sophie Schmieg, Google Security Engineering
    That’s a specific, internal, engineering-driven deadline, not a regulatory mandate. Google ties it directly to faster than expected progress in quantum hardware, error correction, and updated estimates of what it actually takes to break current encryption standards.

    The first concrete product step: Android 17 is integrating ML-DSA-based digital signature protection, on top of post-quantum support already shipping in Chrome and Google Cloud. Google is treating signature and authentication migration as the more time-sensitive half of the problem, ahead of encryption. The logic is straightforward once you sit with it: forging a signature only requires the attacker to break the math at the moment of the attack, so there’s no advance window. Encrypted data, by contrast, can be captured and stored today, then decrypted years later once the hardware catches up, which is the harvest-now, decrypt-later risk we’ll come back to in a moment. Both problems are urgent. They’re just urgent on different clocks.

    The Math That Convinced Google to Move Early

    Six days after the deadline announcement, on March 31, 2026, Google Research published the paper that explains why. Ryan Babbush, Director of Research for Quantum Algorithms, and Hartmut Neven, VP of Engineering at Google Quantum AI, laid out two optimized quantum circuits for solving the elliptic curve discrete logarithm problem at the 256-bit security level, the math underneath ECDSA, the signature scheme securing most TLS connections, SSH sessions, code signing, and the majority of cryptocurrency wallets, including Bitcoin and Ethereum.

    The numbers: one circuit uses fewer than 1,200 logical qubits and roughly 90 million Toffoli gates. The second uses fewer than 1,450 logical qubits and about 70 million Toffoli gates. Run on a superconducting quantum computer, Google estimates either circuit could complete the attack with fewer than 500,000 physical qubits in a matter of minutes, an approximately 20-fold reduction from prior estimates of what the attack would require.

    It’s not an isolated revision either. Craig Gidney, also at Google Quantum AI, updated his RSA-2048 factoring estimate in 2025 to under 1 million physical qubits and less than a week of runtime, down from his own 2019 estimate of 20 million qubits and roughly eight hours. Same researcher, same category of algorithm, same order-of-magnitude drop. Two separate 20-fold reductions in attack-resource estimates, from the same research group, inside a single decade, is arguably more significant than either individual number. That’s the trend that moved Google’s internal calendar, not one paper.

    One more detail worth knowing: Google didn’t publish the actual attack circuits. It used a zero-knowledge proof, developed in coordination with the U.S. government, that lets outside researchers verify the resource estimate without releasing a usable attack blueprint, an approach modeled on standard coordinated vulnerability disclosure practice.

    NIST’s Slower, Still-Unfinished Deadline

    Google’s 2029 timeline is the freshest news, but the post-quantum cryptography schedule most compliance teams actually have to plan against still comes from NIST. NIST finalized its first three post-quantum standards back in August 2024: FIPS 203 (ML-KEM, for key encapsulation), FIPS 204 (ML-DSA, for digital signatures), and FIPS 205 (SLH-DSA, a hash-based signature scheme). A fourth standard, FIPS 206, based on the Falcon algorithm, is still in draft and isn’t expected to finalize until late 2026 or early 2027. A fifth, HQC, selected as a backup key-encapsulation method in March 2025, won’t see a final standard until 2027 at the earliest.

    The 2030 and 2035 dates you’ve probably seen cited everywhere actually come from a separate document, NIST Internal Report 8547, which proposes deprecating 112-bit-security algorithms like RSA-2048 and ECC P-256 by 2030, and disallowing all quantum-vulnerable public-key algorithms by 2035. Here’s the part that doesn’t get repeated often enough: IR 8547 is still an initial public draft. It was released in November 2024, public comment closed in January 2025, and it has not been finalized as of this writing. Treat the 2030 and 2035 dates as the most likely outcome of an open process, not as settled law, especially if your organization sits outside direct U.S. federal scope.

    National security systems run on a separate, tighter clock. The NSA’s CNSA 2.0 guidance sets preference dates as early as 2025 for some categories, required adoption between 2030 and 2033, and full quantum resistance by 2035, with ML-KEM-1024 and ML-DSA-87 specified as the mandated parameter sets. If you’re a defense contractor or anywhere in that supply chain, this is the schedule that actually governs you, not IR 8547.

    Where Other Governments Stand

    Canada, the EU, and the UK are each running parallel post-quantum cryptography tracks, on slightly different clocks. If your organization operates across any of these jurisdictions, the deadline that matters is whichever one applies to your weakest-governed system, not the most generous one.

    Authority Key Deadline(s) Status
    NIST (U.S. civilian federal) Deprecate by 2030, disallow by 2035 Draft (IR 8547), not finalized
    NSA CNSA 2.0 (U.S. national security systems) Required adoption 2030 to 2033, full by 2035 Active guidance
    Google (internal corporate policy) 2029 Announced March 2026
    Canada (federal departments) Migration plans due April 2026 Active mandate
    European Union (NIS Cooperation Group) Critical infrastructure by end of 2030, medium-risk systems by end of 2035 Endorsed by 18 member states, June 2025
    United Kingdom (NCSC) Map dependencies by 2028, complete migration by 2035 Phased guidance, March 2025

    Not Everyone Is Convinced This Is Urgent

    Worth saying plainly: no quantum computer capable of breaking today’s public-key encryption exists yet. CISA’s own Post-Quantum Cryptography Initiative says so directly, while still flagging harvest-now-decrypt-later as a present-tense risk for long-lived data. Both statements are true at the same time, which is exactly why the expert debate over urgency hasn’t settled.

    Scott Aaronson, the Schlumberger Centennial Chair of Computer Science at the University of Texas at Austin and one of quantum computing’s most consistent public skeptics, wrote in comments reported in early May 2026 that people whose hardware judgment he trusts more than his own now think a fault-tolerant, attack-scale quantum computer “ought to be possible by around 2029.” That’s a notable shift for Aaronson. It’s also worth knowing his own caveat: an earlier prediction of his would technically count as fulfilled even by a trivial demonstration, like factoring 15 into 3 times 5, a calculation a human can do faster by hand than any quantum computer currently running. Most headline coverage drops that part.

    Matthew Green, a cryptography professor at Johns Hopkins University, takes a more grounded stance. In comments reported by CyberScoop in April 2026, Green called the recent research a useful precautionary exercise, but questioned whether quantum computing has enough near-term, lucrative applications to accelerate past foundational research into deployed attacks on the timeline implied by recent coverage. He raises a sharper point too: several of NIST’s own earlier post-quantum candidates turned out to have classical, non-quantum vulnerabilities. SIKE, one of the standardization process’s later-round finalists, was broken in 2022 using ordinary computers. Calling something “post-quantum” doesn’t automatically make it secure against everything else.

    Adam Back, CEO of Blockstream and one of the cypherpunk movement’s earliest figures, pushed back specifically on cryptocurrency alarm following Google’s paper, telling Bloomberg the practical threat to Bitcoin remains decades off.

    “The biggest calculation it’s performed is factoring 21 into 7 times 3.” Adam Back, CEO, Blockstream
    Back still thinks Bitcoin and other chains relying on the same elliptic curve signatures, the kind of dependency that also shows up in cross-chain bridge designs, should start migrating to quantum-resistant signatures now. He just doesn’t think anyone should be panicking about it this month.

    The research group a16z crypto goes further, arguing the field isn’t close to a cryptographically relevant quantum computer by any reasonable reading of public progress data. Even reporting on Google’s own resource-estimate reduction tends to include the same caveat: shrinking the qubit count on paper doesn’t solve the unsolved systems-engineering problem of running hundreds of thousands of physical qubits with real-time error correction at scale. Reducing one bottleneck just exposes the next one.

    What This Actually Means for Your Organization

    Strip out the deadline debate and the actual post-quantum cryptography adoption gap is the unglamorous part. A Propeller Insights survey of 1,042 senior cybersecurity managers, commissioned by DigiCert and published in 2025, found that 69% of respondents recognize the quantum risk to current encryption, but only 5% have actually implemented quantum-safe encryption anywhere in their environment. Nearly half, 46.4%, believe a substantial share of their own encrypted data could eventually be compromised.

    A separate Ponemon Institute study of 1,426 IT and security practitioners across the U.S., EMEA, and Asia-Pacific found 61% don’t expect to be ready, only 30% have allocated budget, and just 52% have even started a cryptographic inventory, the mandatory first step in any migration. A more recent Omdia survey of over 400 senior IT leaders, published in early June 2026, puts the share who’ve fully assessed their systems for cryptographic risk at just 22%.

    None of that requires believing a quantum computer will exist next year. It requires believing harvest-now-decrypt-later is real today. Adversaries can capture encrypted traffic now and simply wait. Anything with a confidentiality requirement stretching into the mid-2030s, financial records, legal archives, government communications, long-lived intellectual property, is exposed under today’s encryption the moment it’s intercepted, regardless of when the decryption key eventually becomes breakable.

    The takeaway isn’t “panic.” It’s “inventory now.” You can’t migrate what you haven’t found. NIST’s own migration methodology treats the inventory phase alone as a six-to-twelve-month project for a complex enterprise, before remediation even starts, which is exactly why most organizations need to begin before they feel ready.

    This is the same regulator setting the clock on zero trust security architecture, so if your team is already mapping NIST-aligned controls for that initiative, cryptographic inventory belongs on the same project plan rather than a separate one.

    A Practical Migration Checklist

    Here’s the post-quantum cryptography migration sequence security teams are actually using, in order:

    1. Run a cryptographic inventory. Find every system, certificate, library, and hardcoded dependency using RSA, ECDSA, ECDH, DSA, or Diffie-Hellman. Most teams underestimate how many places this math is buried.
    2. Prioritize by data lifespan, not system criticality. A low-priority system with 20-year data retention requirements is a higher quantum risk than a high-priority system that only handles short-lived sessions.
    3. Deploy hybrid cryptography first. Pairing a classical algorithm with a post-quantum one means you stay protected even if one half is later broken, which matters given Matthew Green’s point about unproven new candidates.
    4. Treat signatures as time-sensitive on their own clock. Migrate authentication and signing separately from encryption, since forged signatures become possible the moment a capable quantum computer exists, with no advance-harvest grace period.
    5. Wait on unfinished standards. Hold off building production dependencies on FIPS 206 (Falcon) or HQC until they’re finalized, expected sometime between 2026 and 2027.
    6. Budget against the four-year window. Use Gartner’s “unsafe by 2029, fully breakable by 2034” framing as a planning heuristic for budget approval, not a precise countdown.

    The Real Timeline vs. the Headlines

    So which date should actually go on your roadmap? Probably more than one. Google’s 2029 is the most concrete forcing function available right now, especially since it comes from the organization that’s done the most original research on how hard this attack actually is. NIST’s 2030 and 2035 dates remain the closest thing to a regulatory backbone, even in draft form, and they’re the dates auditors, cyber insurers, and procurement teams will most likely reference once IR 8547 finalizes. The skeptics aren’t wrong that no cryptographically relevant quantum computer exists today. They’re just answering a different question than “when should my migration start.”

    Our read: the binding constraint here isn’t compute, it’s organizational inertia. Most enterprises won’t miss the 2029 or 2030 deadlines because the cryptography isn’t ready. They’ll miss it because the inventory phase alone quietly eats two of the four years they thought they had.

    The Y2K comparison shows up constantly in coverage of this story, and it’s useful shorthand with one real flaw: Y2K had a single, fixed, universally agreed date. Q-Day doesn’t. The Global Risk Institute’s seventh annual Quantum Threat Timeline Report, built from structured input from 26 named quantum-computing experts, puts a full-scale cryptographically relevant quantum computer at “quite possible” within 10 years and “likely” within 15. That’s a probability distribution, not a deadline. Plan accordingly.


    Frequently Asked Questions

    When will quantum computers be able to break encryption?

    No cryptographically relevant quantum computer exists yet. The Global Risk Institute’s expert survey puts a full-scale version at “quite possible” within 10 years and “likely” within 15. Google’s own internal deadline targets 2029, six years ahead of NIST’s 2035 federal backstop.

    What is NIST’s deadline for RSA-2048?

    NIST’s draft document IR 8547 proposes deprecating RSA-2048 and ECC P-256 by 2030 and disallowing all quantum-vulnerable public-key algorithms by 2035. The document remains an unfinalized draft as of mid-2026, so treat the dates as a likely outcome, not finished law.

    Why did Google move its quantum deadline to 2029?

    Google cited faster than expected progress in quantum hardware and error correction, plus new research showing elliptic curve cryptography could be broken with roughly 20 times fewer qubits than previously estimated, published alongside its March 2026 deadline announcement.

    What is harvest now, decrypt later?

    It is an attack pattern where adversaries capture encrypted data today and store it, planning to decrypt it once a powerful enough quantum computer exists. It makes long-lived encrypted data vulnerable right now, even though no quantum computer can currently break it.

    Is Bitcoin vulnerable to quantum computers?

    Bitcoin’s signature scheme relies on the same elliptic curve math Google’s research targeted, so it is theoretically exposed long term. Most cryptographers, including Bitcoin advocate Adam Back, consider the practical threat years to decades away, but support migrating gradually now.


    What to Watch Next

    Here’s what you didn’t know a few minutes ago: the urgency in this post-quantum cryptography story isn’t coming from a finished government deadline. It’s coming from a trend line, two separate 20-fold reductions in attack-resource estimates from the same Google research group inside a decade, that’s compressing every other timeline built around it.

    Over the next 6 to 18 months, watch three things. First, whether NIST finally finalizes IR 8547 or pushes the date again, the way an earlier proposed target was already revised once before. Second, whether other major infrastructure operators follow Cloudflare’s reported move to align with Google’s 2029 date, which would turn one company’s internal policy into something closer to an industry standard. Third, whether FIPS 206 and HQC actually land in their projected 2026 to 2027 window, since both are needed before several hybrid deployment strategies can fully mature.

    None of this requires belief in an imminent Q-Day. It requires an honest inventory, a realistic budget conversation, and a migration plan that survives whichever date turns out to be the right one.

    Want the next development before it hits your feed?
    Subscribe to The Neural Loop at neuralwired.com/newsletter

  • Gartner: AI Inference Cost Won’t Drop Your Bill in 2026

    Gartner: AI Inference Cost Won’t Drop Your Bill in 2026

    Reduce LLM Inference Cost: Gartner’s 2026 Enterprise Guide
    Enterprise AI · Cost Architecture

    Gartner Says Your LLM Bill Is About to Double

    In April 2026, Uber’s engineering organization ran out of its entire annual AI coding budget two-thirds of the way through the year. Two months later, Microsoft pulled back most internal developer access to Claude Code over the same problem: cost. Neither company is careless with money. Both got caught by the same math now hitting finance teams across the industry: token prices are falling, but the bill keeps climbing.

    If part of your job in 2026 is figuring out how to reduce LLM inference cost across an enterprise deployment, “wait for prices to drop further” stopped being a strategy the moment it failed at Uber and Microsoft. What works instead is an actual cost architecture, built from seven specific moves. Some you can ship this week. Some take a quarter to stand up properly.

    Quick answer: Per-token LLM prices have fallen as much as 280-fold since 2022, yet enterprise AI bills keep rising. The reason is agentic workflows, which burn 5 to 30 times more tokens per task than a simple chatbot query, according to Gartner, with adoption accelerating faster than unit prices fall. Cutting your real bill takes architecture, not patience: prompt caching, smart model routing, inference-stack tuning, agentic spend controls, sound self-host-versus-API math, FinOps-grade attribution, and pre-deployment cost modeling.

    Why Falling Token Prices Won’t Reduce Your LLM Inference Cost

    Here’s the number every “AI is getting cheaper” headline traces back to. Gartner forecasts that by 2030, running inference on a one-trillion-parameter model will cost providers more than 90% less than it did in 2025, with LLMs overall becoming up to 100 times more cost-efficient than the earliest 2022-era models of similar size, per Gartner’s March 2026 forecast. Stanford’s AI Index already clocked a 280-fold drop in GPT-3.5-equivalent inference pricing between late 2022 and late 2024.

    None of that is showing up as savings on enterprise invoices. Gartner says so explicitly, and so does the analyst who built the forecast.

    “Chief Product Officers (CPOs) should not confuse the deflation of commodity tokens with the democratization of frontier reasoning.” Will Sommer, Senior Director Analyst, Gartner, via Gartner
    Translation: cheap tokens fund better models, not smaller bills. Demand, meanwhile, is exploding underneath the price drop. Agentic models require 5 to 30 times more tokens per task than a standard chatbot exchange, per Gartner, and Goldman Sachs projects global token consumption will climb roughly 24-fold by 2030, reaching something like 120 quadrillion tokens a month. That’s Jevons Paradox playing out in real time: a 160-year-old economic principle stating that efficiency gains tend to increase total consumption rather than reduce it. Cheaper tokens have historically unlocked more AI usage, not lower total spend.

    This has also stopped being a quiet, internal cost-center problem. An AI consultant told Axios that one unnamed enterprise reportedly burned through roughly $500 million in Claude API spend in a single month after failing to set employee usage limits (treat that figure as reported, not officially confirmed; no company has put its name on it). What is confirmed: the Linux Foundation announced a new standards body, the Tokenomics Foundation, in the first week of June 2026, modeled directly on how the FinOps Foundation standardized cloud-cost discipline a decade earlier. When an industry spins up a dedicated standards body to police a cost problem, that problem has stopped being optional to manage.

    The 7-Step Framework to Reduce LLM Inference Cost at Enterprise Scale

    Treat the list below as a stack, not a checklist. The first two steps are mechanical and fast, most teams see results within days. The last three are organizational and slower, taking a quarter to stand up properly, but they’re what stops the bill from doubling again next year.

    StepWhat it fixesTime to first results
    1. Prompt cachingRepeated context reprocessed on every single callDays
    2. Model routing & cascadingFrontier pricing applied to tasks a small model could handle1–2 weeks
    3. Inference-stack tuningBatching and decoding choices mismatched to traffic2–4 weeks
    4. Agentic token-sprawl controlUngoverned agents multiplying spend with no budget ceiling2–4 weeks
    5. Self-host vs. API mathHidden personnel costs erasing on-paper savings4–6 weeks
    6. FinOps-grade attributionNo one can say which team or agent is spending what1–2 quarters
    7. Pre-deployment cost modelingCost surprises discovered after a feature shipsOngoing

    Step 1: Turn On Prompt Caching First

    If you do exactly one thing this week, do this one. Prompt caching stores the computed representation of a prompt’s repeated prefix, things like system instructions, tool definitions, and reference documents, so later calls skip reprocessing it from scratch. Anthropic made it generally available on December 17, 2024, pricing cached input tokens at roughly 10% of standard input cost, with cache writes running 1.25 to 2 times standard pricing depending on how long the cache is held.

    Anthropic’s own figures claim up to 90% cost reduction and 85% latency reduction on long prompts. Vendor claims are easy to discount, except this one has independent backing: a 2026 academic evaluation titled “Don’t Break the Cache” ran the first comprehensive third-party test across three major LLM providers on long-horizon agentic tasks and found real-world savings of 41 to 80%. That’s the single most defensible number in this whole article: a vendor claim and an independent study landing in the same range.

    “We’re excited to use prompt caching to make Notion AI faster and cheaper, all while maintaining state-of-the-art quality.” Simon Last, Co-founder, Notion, via Anthropic
    Implementation here is an engineering task, not a procurement project. Most teams can audit their cache hit-rate and restructure prompts (static content first, variable content last) inside a single sprint.

    Step 2: Route and Cascade, Don’t Just Pay Frontier Prices for Everything

    Routing sends each request to the cheapest model that can actually handle it. Cascading starts cheap and escalates only when a quality check fails. Stanford’s FrugalGPT, the paper that founded this whole technique, showed up to 98% cost reduction while matching the best individual model’s performance. UC Berkeley’s RouteLLM project later trained a classifier that hit 95% of GPT-4-level quality at 85% lower cost on MT-Bench, and comparable quality at 45% lower cost on MMLU.

    Treat those percentages as proof the technique works, not as a number to promise finance. Both figures are tied to one specific model pairing on two specific benchmarks. Your prompts and your traffic will produce a different result, so run the eval on your own workload before it goes into a budget deck.

    We already broke down the small-versus-frontier pricing tables, the self-hosting break-even math, and named enterprise case studies in our companion piece, GPT-5 vs Small Language Models: 2026 Enterprise Cost. If routing is the step you want to go deep on, that’s where the depth lives. Here, the point is narrower: routing is one piece of a seven-piece architecture, not the whole fix.

    Step 3: Tune the Inference Stack to Your Actual Traffic

    This is where “obvious” optimizations get dangerous. Speculative decoding is marketed as a clean win: predict several tokens ahead, verify them in parallel, cut latency and cost together. At small batch sizes (16 or fewer concurrent requests), independent energy-efficiency research found it cuts cost by up to roughly 29%. At large batch sizes (128 concurrent requests), the same technique increased total energy cost by roughly 25.65%.

    Same technique, opposite result, depending entirely on traffic pattern. Batching, hardware allocation, and decoding strategy aren’t settings to copy from a blog post. They’re decisions that need your own load data behind them, and most teams haven’t measured their traffic closely enough to know which side of that line they’re actually on.

    Step 4: Put a Leash on Agentic Token Sprawl

    This is the fastest-growing line item, and the one with the least governance attached to it. EY’s analysis found the cost of a single agentic customer-service interaction rose from roughly $0.04 in 2023 to $1.20 in 2026, a roughly 30-fold increase, driven by orchestrated multi-tool agent workflows replacing simple linear chatbot exchanges.

    The fix is governance, not a model swap: per-agent token budgets, hard ceilings on tool-call chains, and visibility into which agent, team, or feature is generating which slice of the bill. We covered the broader shadow-AI version of this problem, agents spun up without IT’s knowledge, quietly multiplying spend nobody’s tracking, in AI Agent Sprawl: The Shadow AI Crisis Hitting Enterprise. If ungoverned agent proliferation is your specific problem, start there.

    Step 5: Get the Self-Host vs. API Math Right

    Self-hosting looks cheaper on a per-token spreadsheet, and often isn’t, once the team required to run it gets added in. Maintaining a fine-tuned or self-hosted model typically costs $180,000 to $300,000 a year per ML engineer, and round-the-clock coverage adds another $800,000 to $1.2 million annually. Run that math forward and the break-even point lands around 500,000 tokens per day of sustained load. Below that, API-based smaller models beat self-hosting on total cost. Above it, self-hosting can pay off, assuming the team is already in place rather than hired specifically for this.

    (If you’re staffing up purely to bring a model in-house, you’ve likely already lost the math before writing a line of code.)

    Step 6: Build FinOps-Grade Cost Attribution

    The share of FinOps teams managing AI spend rose from 31% to a projected 98% in two years, according to a survey of 1,192 practitioners representing more than $83 billion in annual cloud spend. That’s not a niche trend. That’s an entire discipline getting rebuilt around one new line item.

    “In April and May, I started hearing from companies: ‘Oh my god, we are 3x over our entire 2026 token budget and it’s only April.’ We started hearing existential crises, and the whole conversation shifted from tokenmaxxing and ‘go fast’ to ‘we need guardrails, how do we control this?’” J.R. Storment, Executive Director, FinOps Foundation, via TechCrunch
    The Tokenomics Foundation, announced by the Linux Foundation in the first week of June 2026 and set to formally launch in July, exists to standardize exactly this: canonical metrics like cost-per-intelligence and tokens-per-watt, the same way the FinOps Foundation standardized cloud-cost discipline a decade ago.

    “AI is forcing FinOps to answer a harder question. It’s not ‘what did we spend?’, it’s ‘what did we actually get out of that spend?’ The hardest part of AI isn’t building it, it’s proving it was worth it.” Rajeev Laungani, Head of Product, Virtasant, via Virtasant
    Practically, this means cost attribution by team, feature, and agent, not just one line on a monthly invoice. Build it before finance asks for it. At this point, they will.

    Step 7: Model Cost Before You Ship, Not After

    Engineering platform Jellyfish analyzed production usage data across its customer base and found the developers consuming the most AI tokens were roughly twice as productive, but used ten times the tokens to get there. Per-developer token consumption rose roughly 18.6-fold in nine months.

    Is that a good trade? Nobody actually knows yet, and that’s the real problem with treating cost optimization as a purely engineering exercise.

    “Whether extreme spend pays off comes down to the ultimate business value of shipped code (e.g. revenue), which most companies still can’t measure.” Nicholas Arcolano, Head of Research, Jellyfish, via TechCrunch
    The shift this forces: a cost model needs to exist before a feature ships, not get reconstructed from an invoice afterward. Our deeper reporting on why so many AI deployments never get to prove that value lives in Enterprise AI Failure Rate 2026: MIT Says 95% Miss ROI.


    Where This Framework Breaks Down

    Every one of the seven steps above is real and verifiable. None of them reverses the underlying economics, and an honest article says so.

    Jevons Paradox is the structural problem. Efficiency gains increase total consumption rather than reducing it, and that’s exactly what’s happening with token pricing. Optimization slows the rate at which your bill grows. It does not reliably make the bill go down once agentic adoption is already underway across your organization.

    Gartner’s own analyst makes the sharpest version of this argument, and it’s worth taking seriously precisely because it comes from the firm that published the optimistic 90%-by-2030 number in the first place. Token deflation, in his framing, helps providers fix their own margins long before any of it reaches the enterprise customer.

    “The customer isn’t going to see all of this money.” Will Sommer, Senior Director Analyst, Gartner, via CIO Dive
    Measurement itself is still broken, too, which undercuts any specific savings percentage offered without an attribution system already in place.

    “Even getting clarity on relatively basic metrics, like the number of tokens being used, works differently in different areas. It’s very fragmented across providers and even across services within a single provider.” Jon Thompson, CTO, Virtasant, via Virtasant
    Our read: this framework will slow your bill’s growth rate. It will not reliably shrink the bill itself, not while agentic adoption is still in its land-grab phase across the industry. The honest claim is “stop runaway growth and start measuring,” not “halve your invoice by next quarter.” Anyone promising the second one is selling something.

    What to Watch Over the Next 6 to 18 Months

    • July 2026: The Tokenomics Foundation formally launches. Watch whether its proposed metrics, cost-per-intelligence and tokens-per-watt, actually get adopted by vendors, or stay aspirational.
    • Late 2026 into 2027: Goldman Sachs projects token consumption climbing toward 24 times current levels by 2030. If that holds even roughly, expect more public “3x over budget” stories like Uber’s, not fewer.
    • Ongoing: Whether Anthropic, OpenAI, and Google start publishing standardized, comparable token-accounting data, the same shift cloud computing went through when FinOps forced billing transparency a decade ago.

    Frequently Asked Questions About Enterprise LLM Cost Optimization

    Why are AI inference costs rising if token prices are falling?

    Per-token prices have fallen as much as 280-fold since 2022, but total enterprise spend is rising because agentic workflows use 5 to 30 times more tokens per task than simple chatbot queries, according to Gartner, and adoption is accelerating faster than unit prices decline.

    What is prompt caching and how much does it save?

    Prompt caching stores a prompt’s repeated prefix, such as system instructions, tools, and documents, so later requests skip reprocessing it. Anthropic reports up to 90% cost reduction and 85% latency reduction on long prompts; independent academic research confirms 41 to 80% real-world savings on agentic workloads.

    What is LLM model routing or cascading?

    Routing sends each request to the cheapest capable model; cascading starts small and escalates only if a quality check fails. Stanford’s FrugalGPT framework demonstrated up to 98% cost reduction using this approach while matching the performance of the best individual model.

    Is it cheaper to self-host an LLM or use an API?

    It depends on volume. Below roughly 500,000 tokens a day of sustained load, API-based smaller models typically beat self-hosting once you account for ML engineer salaries and round-the-clock operations coverage. Above that threshold, self-hosting can pay off.

    How much does GPT-5-class inference cost in 2026?

    OpenAI’s frontier model runs roughly $5.00 per million input tokens and $30.00 per million output tokens as of mid-2026, while budget-tier alternatives like Microsoft’s Phi-4 cost roughly $0.065 to $0.14 per million tokens, a difference of more than 70 times for tasks that don’t need frontier-level reasoning.


    The Bottom Line on Reducing Enterprise LLM Inference Cost

    Here’s what changes once you’ve read this far: you stop waiting for token prices to fix your budget, because they won’t, not at the rate agentic adoption is growing. You start treating LLM cost the way mature engineering organizations treat any other infrastructure spend, with caching turned on by default, routing decisions backed by your own evals instead of someone else’s benchmark, agent budgets that exist before an agent ships, and a cost model built before a feature launches instead of reconstructed from an invoice afterward.

    Three things to do this week: pull your prompt-caching hit-rate and see how far it sits from Anthropic’s claimed ceiling. Set a hard per-agent token budget on whatever’s currently ungoverned. And ask finance whether anyone owns AI cost attribution yet, because if the FinOps Foundation’s numbers hold, 98% of FinOps teams will own a piece of it within the year, whether or not engineering looped them in first.

    Reducing enterprise LLM inference cost in 2026 was never going to be about waiting for cheaper tokens. It’s about building the architecture that makes the token price you already have work for your budget instead of against it.

    Want the next breaking development on enterprise AI economics before your competitors see it? Subscribe to The Neural Loop at neuralwired.com/newsletter.

  • Google ML Study: How Often to Retrain Models in 2026

    Google ML Study: How Often to Retrain Models in 2026

    How Often Should You Retrain an ML Model? Google’s Data
    MLOps / Production ML

    How Often Should You Retrain an ML Model? Google’s Data

    A 450,000-model study out of Google and UC Berkeley just answered a question most MLOps teams have been guessing at for years, and the answer has almost nothing to do with drift schedules.

    Somewhere inside Google, an ML pipeline retrained itself nine times before lunch and shipped exactly one of those runs to production. That is not a glitch in someone’s dashboard. It is the production reality behind a question every CTO funding a machine learning team eventually asks: how often should you retrain a machine learning model? A study from Google and UC Berkeley researchers, built on provenance data covering 3,000 production pipelines and more than 450,000 trained models, finally answers it with real numbers instead of conventional wisdom. The answer is stranger, and more useful, than picking a calendar cadence or chasing drift alerts.

    Most retraining advice circulating online treats cadence like a calendar problem: pick weekly, monthly, or quarterly, and move on. The Google data suggests the real bottleneck isn’t how often models get retrained. It’s how rarely those runs actually matter once they’re finished.

    What 450,000 Models Actually Told Researchers

    The study behind these numbers comes from Doris Xin, Hui Miao, Aditya Parameswaran, and Neoklis Polyzotis, who analyzed the full provenance graph of production ML pipelines inside Google over a four-month window: 3,000 pipelines, 450,000-plus trained models, every training run and every push tracked end to end. It’s one of the largest empirical looks anyone has published at what production ML actually does, as opposed to what teams assume it does.

    The numbers don’t match the “quarterly refresh” mental model most engineering orgs still budget around.

    What the data measuredWhat it found
    Average retrains per pipelineRoughly 7 times per day
    Pipelines retraining more than 100 times a day1.12% of all pipelines
    Retrained models that actually get deployedAbout 1 in 4
    Mean model training time168 hours
    Mean gap between deployed modelsRoughly 40 hours
    Source: Xin, Miao, Parameswaran & Polyzotis, “Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities,” arXiv:2103.16007.

    If your team is retraining far less than seven times a day, that’s not necessarily a problem. Google’s pipelines include extremely high-velocity systems (ad ranking, search relevance) that skew the average up. But the second number matters regardless of your industry: only about one in four of those models ever ships. The rest, roughly 80 percent, get trained and quietly discarded.

    The Habit Nobody Budgets For

    Most retraining compute buys nothing

    For every four models retrained inside Google’s pipelines, only one reached production. The other three consumed GPU hours, engineering attention, and CI capacity, then went nowhere. At a mean training time of 168 hours per model, that’s not a rounding error. It’s a budget line most platform teams don’t separate out, because most dashboards only track “models trained,” not “models trained and discarded.”

    This is the number that should reframe the cadence conversation. The question isn’t “are we doing this often enough.” It’s “what happens to the three out of four runs that don’t ship, and why.”

    Why Drift Might Be the Wrong Villain

    The instinctive answer is drift: the model degraded, the data moved, the model didn’t keep up, so it got rejected. The Google researchers tested that instinct directly, and it didn’t hold up.

    They compared models that got pushed to production against models that didn’t, looking specifically at input-data similarity and code-change rates between the two groups. If drift or code changes were driving the decision to deploy, you’d expect a clear gap. There wasn’t one: data similarity scored 0.101 for pushed models versus 0.099 for unpushed ones, and code-match rates came in at 84.6% versus 83.8%. Statistically, that’s noise, not signal.

    So if drift isn’t the reason most runs die quietly, what is? The researchers point to pipeline-level inefficiency and push-rate throttling instead, meaning the bottleneck sits in process and infrastructure, not in the data the model is learning from.

    “Or another aspect is model drift. Things change over time.” Chip Huyen, ML engineer and author, TechTarget
    Huyen’s point is correct and worth holding alongside the Google data rather than against it: drift is real, and it’s common. A 2022 study in Scientific Reports by Vela et al. tested 128 model-dataset combinations across healthcare, weather, airport traffic, and finance, and found measurable temporal degradation in 91% of pairs. That figure is genuine and peer-reviewed (read more in our breakdown of the underlying drift mechanics), but it measures whether degradation happens at all, not whether degradation is what’s killing your discarded retrains specifically. Those are two different claims, and conflating them is how “drift” becomes the catch-all explanation for problems that are really about pipeline design.

    Our read: if your team explains every discarded retrain as “drift,” you’re probably explaining away a pipeline problem, not a data problem. Researchers Shreya Shankar, Rolando Garcia, Joseph Hellerstein, and Aditya Parameswaran reached a related conclusion from a different angle: in 18 interviews with practicing ML engineers at companies running chatbots, autonomous vehicles, and finance systems, they found that engineers consistently treat production behavior as something that can’t be fully known until the model is live, which is precisely why monitoring infrastructure, not retrain frequency, ends up being the deciding factor in whether a model ships.

    There’s a second failure mode hiding in here too: alert fatigue. Statistical drift tests like the Population Stability Index and the Kolmogorov-Smirnov test routinely flag distribution shifts that never translate into a measurable performance drop, a pattern confirmed across multiple independent analyses (arXiv:2003.12808). Teams that don’t tune thresholds to actual business impact eventually start ignoring alerts altogether, real ones included. That’s arguably a bigger operational risk than drift itself, and it’s a pipeline-design problem too, not a data problem.

    This lines up with broader patterns we’ve tracked in MLOps pipeline failures: the infrastructure layer, not the model layer, is where most production ML actually breaks.

    Building a Retrain Schedule That Matches Production, Not a Calendar

    If push-rate, not how often you retrain, is the real lever, your policy should be instrumented around it instead. Four steps, in order:

    1. Baseline your own push rate first

    Before touching your retrain schedule, measure how many of your team’s runs actually reach production today. That number, not your calendar, is your true starting point.

    2. Track “retrained” and “deployed” as separate metrics

    Most teams report retrain count as a proxy for ML activity. Splitting it into trained versus deployed exposes exactly the gap Google’s data found, and tells you where compute is leaking.

    3. Instrument the deploy decision itself

    Log the reason every retrained model did or didn’t ship: performance gate, manual review, throttling, rollback. That log tells you more about your real bottleneck than a drift dashboard ever will.

    4. Use the 40-hour benchmark as a sanity check, not a target

    Google’s pipelines averaged roughly 40 hours between deployed models. If your gap looks wildly different in either direction, investigate that gap before you touch the retrain calendar at all.

    Decisions like these tend to fall on whoever owns ML infrastructure, a role that, per our reporting on the enterprise AI skills gap, many organizations still haven’t clearly assigned.


    When Nobody Catches It in Time

    Process gaps like these aren’t abstract. Zillow’s iBuying arm, Zillow Offers, shut down in November 2021 after its pricing models systematically overvalued homes the company then had to sell at a loss. The numbers, from Zillow’s own Q3 2021 SEC filing, were stark: a $304 million quarterly operating loss, $175 to $230 million in additional impairment costs, and a roughly 25 percent workforce reduction.

    “We were unintentionally purchasing homes at higher prices.” Rich Barton, Co-founder & CEO, Zillow Group, GeekWire
    We’ve covered the full Zillow case study in detail elsewhere, so we won’t retell it here. The relevant point for this article: Zillow’s failure wasn’t primarily a story about retraining too rarely. It was a story about a pricing signal that kept degrading without anyone instrumenting the gap between “the model said X” and “X turned out to be wrong,” which is the exact same blind spot the Google study found at much smaller, less catastrophic scale across thousands of unremarkable pipelines.


    Frequently Asked Questions

    What is model drift in machine learning?

    Model drift is the gradual decline in a deployed model’s predictive accuracy as real-world data or relationships diverge from training conditions. It shows up as data drift, where inputs change, or concept drift, where the relationship between inputs and outputs changes entirely.

    How often should you retrain a machine learning model?

    There is no universal schedule. Production data from a 450,000-model Google study shows pipelines retrain roughly seven times a day on average, but only about one in four of those retrained models ever gets deployed, so cadence matters less than your deploy-decision process.

    What is the difference between data drift and concept drift?

    Data drift means the distribution of input features shifts while the relationship between inputs and outputs stays the same. Concept drift means that relationship itself breaks, so an input that looked normal now warrants a different correct answer, which makes it harder to catch.

    Does model drift cause most ML deployment failures?

    Not necessarily. A Google and UC Berkeley study of 450,000 production models found no meaningful difference in data similarity or code changes between models that got deployed and models that did not, suggesting pipeline inefficiency, not drift, explains most discarded retrains.

    How do you detect model drift?

    Compare live production data against a training baseline using statistical tests like the Population Stability Index or the Kolmogorov-Smirnov test, paired with direct performance tracking against labeled outcomes. Watch for alert fatigue: poorly tuned thresholds flag shifts that never affect real accuracy.


    The Bottom Line

    The number worth carrying out of this article isn’t 91 percent (how often models drift) or even seven times a day (how often Google’s pipelines retrain). It’s one in four: how often a retrain actually earns its compute. Most conversations skip straight from “is our model degrading” to “how often should we retrain,” without ever asking whether that cadence was the bottleneck in the first place.

    Over the next 6 to 18 months, expect this question to get more urgent, not less. Gartner forecasts worldwide AI spending will hit $2.59 trillion in 2026, up 47% year over year, and the same firm predicts that 40% of organizations deploying AI will adopt dedicated observability tooling by 2028. Budget is arriving faster than judgment about where to point it. The teams that benchmark their own push rate now, before the next wave of tooling spend, will be the ones who can tell the difference between buying real visibility and buying a more expensive version of the same blind spot.

    It’s also worth keeping this separate from the broader AI-project failure narrative. The 70 to 95 percent failure-rate figures that get cited from MIT, Gartner, and RAND research on AI ROI are measuring pilot-to-production failure broadly, for reasons that often have nothing to do with this issue specifically. Treating them as the same problem inflates the apparent size of the drift issue and obscures the much narrower, much more fixable pipeline question this study actually answers.

    Three things to watch from here: whether more vendors start publishing push-rate benchmarks the way this Google study did, whether the EU AI Act’s risk-monitoring provisions start requiring documented retrain-versus-deploy decisions rather than just drift scores, and whether the same questions get applied to hosted LLM and agent pipelines, where there’s often no training data to inspect at all. That last one is where this entire conversation is heading next.

    Get analysis like this in your inbox. Subscribe to The Neural Loop at neuralwired.com/newsletter

  • Enterprise AI Failure Rate 2026: MIT Says 95% Miss ROI

    Enterprise AI Failure Rate 2026: MIT Says 95% Miss ROI

    95% of Enterprise AI Pilots Fail to Deliver ROI | NeuralWired
    Enterprise AI • Research Analysis

    95% of Enterprise AI Pilots Fail to Deliver ROI. Four Research Teams Just Confirmed It.

    Your company just spent six months and a million dollars on a generative AI pilot. The vendor demos looked flawless. The internal presentations sparked genuine excitement. Then the results came in, and nothing moved. Not revenue. Not costs. Not customer retention.

    If that sounds familiar, you are in the majority. A large, expensive, embarrassingly well-funded majority.

    Four major research institutions, using four separate methodologies, all landed on the same uncomfortable finding about enterprise AI investment in 2025: somewhere between 60% and 95% of organizations are spending real money on AI and producing nothing measurable in return. The AI ROI crisis is not a pessimist’s talking point anymore. It is the consensus position of the best-sourced data in the field.

    Here is what the research actually shows, why the failures keep happening, and what the small group of winners is doing that everyone else isn’t.


    The Real Numbers: What Four Independent Studies Found

    The most important thing to know before citing any AI failure statistic is that several of the most-shared numbers online are not traceable to real research. A figure that circulated widely in early 2026, claiming “$684 billion invested with 80.3% producing nothing,” appears across dozens of content sites but traces back to no primary dataset, no named methodology, and no actual report. It is not from RAND. It is not from Gartner. It does not exist in the primary literature.

    What does exist is more interesting, and more damning.

    95%
    of GenAI pilots show zero measurable P&L impact
    MIT Project NANDA, July 2025
    39%
    of organizations report any enterprise-wide EBIT impact from AI
    McKinsey State of AI 2025
    42%
    of companies abandoned most of their AI initiatives in 2025, up from 17% in 2024
    S&P Global Market Intelligence, March 2025
    25%
    of AI initiatives delivered expected ROI, per CEO self-report
    IBM Institute for Business Value, May 2025
    MIT’s Project NANDA published the sharpest number. Their GenAI Divide: State of AI in Business 2025 report reviewed over 300 publicly disclosed AI initiatives and conducted structured interviews with representatives from 52 organizations, plus survey responses from 153 senior leaders. The finding: 95% of GenAI pilots delivered no measurable profit-and-loss impact. Only 5% of integrated systems created significant value.

    McKinsey’s State of AI 2025 is the largest survey, with nearly 1,993 respondents across 105 countries. It found that just 39% of organizations report any enterprise-wide EBIT impact from AI. Only about 5.5% to 6% of respondents qualify as true high performers, meaning their organizations attribute more than 5% of EBIT to AI use.

    S&P Global Market Intelligence surveyed more than 1,000 respondents across North America and Europe in March 2025 and found that the share of companies abandoning most of their AI initiatives jumped to 42% in one year, up sharply from 17% the prior year. The average organization scrapped 46% of AI proof-of-concepts before they ever reached production.

    The IBM Institute for Business Value CEO Study surveyed 2,000 CEOs across 33 countries and found that only 25% of AI initiatives had delivered expected ROI over the preceding few years, and only 16% had scaled enterprise-wide.

    Key Context These four studies use different definitions of “failure” and different methodologies. MIT tracks GenAI pilots specifically on P&L impact. McKinsey tracks EBIT at enterprise scale. S&P tracks initiative abandonment rates. IBM tracks CEO self-reported ROI. The fact that all four land in the same territory (a small single-digit percentage of organizations capturing most of the value) is more persuasive than any single number would be on its own.
    For spending context: Stanford HAI’s AI Index 2025 tracked $252.3 billion in corporate AI investment in 2024, with private investment climbing 44.5% year-over-year. Gartner forecasts worldwide AI spending will reach $2.52 trillion in 2026. The money is real. The returns, for most organizations, are not.


    Why Enterprise AI Projects Actually Fail

    The RAND Corporation’s 2024 qualitative study, based on interviews with 65 experienced AI and machine learning practitioners, is explicit: “By some estimates, more than 80% of AI projects fail. That’s twice the failure rate of non-AI IT projects.” RAND framed this as an estimate rather than a hard statistic, which is the intellectually honest position. But the directional claim is consistent with every quantitative study that followed.

    What RAND and the subsequent research agree on is that the failure causes are almost never technical. Models work. APIs work. The infrastructure, mostly, works. The failures are organizational.

    The Integration Gap

    McKinsey’s data reveals something counterintuitive: function-level AI wins (software engineering and IT teams seeing 10% to 20% cost reductions, for example) coexist with near-zero enterprise-wide EBIT impact. Teams are building AI tools. Those tools are producing local efficiencies. But the enterprise-level needle doesn’t move because the tools exist inside departmental silos, disconnected from the workflows that drive revenue and cost at scale.

    The organizations in McKinsey’s high-performer cohort are not running more pilots. They are forcing workflow redesign. There is a meaningful difference between adding AI to an existing process and redesigning the process around AI’s actual capabilities.

    The Data Readiness Problem

    Gartner predicted in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by end of 2025, citing poor data quality, inadequate risk controls, escalating costs, and unclear business value. Gartner’s June 2025 follow-up extended that prediction to agentic AI: over 40% of agentic AI projects are expected to be canceled by end of 2027, for essentially the same reasons.

    The pattern Gartner keeps documenting is that most enterprise AI failures trace upstream to data infrastructure, not model selection. Organizations with fragmented, inconsistent, or poorly governed data are running state-of-the-art models on inputs that guarantee mediocre outputs. The models perform exactly as well as the data allows. The data, in most enterprises, does not allow much.

    The FOMO-Driven Pilot Problem

    IBM’s CEO study names a failure mode that does not appear in the other research but is instantly recognizable to anyone who has sat through an enterprise AI strategy meeting: “FOMO-driven” pilots. Organizations launch AI initiatives because competitors are launching AI initiatives, not because they have identified a specific problem that AI is the right tool to solve. The result is a portfolio of proofs-of-concept that demonstrate capability without establishing business value, and that get quietly shelved when the next technology cycle begins.

    “Some large companies’ pilots and younger startups are really excelling with generative AI. It’s because they pick one pain point, execute well, and partner smartly with companies who use their tools.”

    Aditya Challapally, Lead Author, MIT NANDA GenAI Divide Report (Fortune, August 2025)
    The “one pain point” framing is more radical than it sounds. Most enterprise AI strategies involve multiple simultaneous pilots across multiple functions. The MIT data suggests that approach produces 95% failure rates. The alternative is a level of focus that most organizations, politically and structurally, find difficult to achieve.


    What the 5% of Winners Do Differently

    BCG’s Widening AI Value Gap report (published September 2025, surveying 1,250-plus global firms) found that only 5% of companies are achieving AI value at scale. BCG calls this cohort “future-built” firms. They share specific structural traits, not just better models or more budget.

    What Average Firms Do What Future-Built Firms Do
    Multiple simultaneous pilots across functions Single focused use case with defined P&L ownership
    Measure success by pilot completion Measure success by business metric change within 90 days
    Deploy AI into existing workflows Redesign workflows before and during AI deployment
    Data readiness addressed post-launch Data infrastructure audited and fixed before launch
    AI team isolated in IT or innovation lab AI ownership embedded in business unit P&L
    “Agentic AI isn’t a future concept. It’s already reshaping workflows and redefining roles. Companies should view it as the next step in scaling AI, not as the starting point.”

    Amanda Luther, Managing Director and Senior Partner, Boston Consulting Group, co-author of The Widening AI Value Gap
    BCG’s data also shows that 60% of companies are not achieving material value at all, reporting minimal revenue and cost gains despite substantial investment. The distribution is not a bell curve. It is a winner-take-most dynamic where a small cohort is pulling away from the field. The companies in that cohort are not smarter. They moved earlier on data infrastructure, defined success in business terms before launch, and treated AI deployment as a change-management problem rather than a technology rollout.

    If you’re building an enterprise AI program and you haven’t done a formal audit of your data readiness before approving new spend, Gartner’s prediction applies directly to you.


    The Skeptic’s Case: Is the AI Failure Narrative Overblown?

    The most credentialed critic of the 95% figure is Paul Roetzer, founder and CEO of the Marketing AI Institute. Speaking on The Artificial Intelligence Show in August 2025, Roetzer was direct about the MIT NANDA methodology: “Please don’t put any weight into this study. This is not a viable, statistically valid thing.”

    His critique is specific and worth taking seriously. MIT’s 95% figure tracks GenAI pilots on a narrow, 6-month P&L-only definition of success. That definition excludes efficiency gains, cost reductions, customer churn improvements, and sales pipeline velocity. Roetzer’s argument is that an organization that deploys a GenAI tool and reduces its customer support ticket resolution time by 40% would count as a “failure” under NANDA’s methodology, because that improvement did not show up as a measurable P&L impact within six months.

    “Anytime you see a headline like that, you have to immediately step back and say, okay, that seems unrealistic.”

    Paul Roetzer, Founder and CEO, Marketing AI Institute, speaking on The Artificial Intelligence Show, Episode 164 (August 2025)
    He also notes a potential framing consideration: NANDA’s research mission is building an “Internet of AI Agents,” meaning the report’s implicit argument is that today’s static GenAI tools fail while adaptive agentic systems succeed. That is not a reason to dismiss the report, but it is a reason to hold the 95% figure as directional rather than precise.

    Our read: Roetzer’s methodological critique is valid. The 95% figure almost certainly overstates the failure rate under a broader definition of value. But it probably understates the failure rate under a strict enterprise-ROI definition, because organizations are generally terrible at measuring AI value even when it exists. The honest answer is that somewhere between 60% and 95% of enterprise AI initiatives are producing less value than their sponsors expected, which is damning enough without needing to settle on a single number.


    What Happens Next: The 18-Month Outlook

    Gartner’s “Trough of Disillusionment” framing for GenAI in 2026 fits the historical Hype Cycle pattern and is a reasonable, falsifiable prediction. After peak hype comes a period where the gap between expectation and delivered value becomes impossible to ignore, investment gets more selective, and the organizations that built real infrastructure during the hype phase begin pulling away from those that didn’t.

    Three things are worth watching over the next 12 to 18 months.

    Agentic AI cancellation rates will become the new headline metric. Gartner predicts more than 40% of agentic AI projects will be canceled by end of 2027. Given that agentic AI is currently in an earlier hype phase than GenAI was in 2024, the cancellation rate could be higher. Watch for enterprise announcements of agentic AI programs in Q3 2026, and note whether they include defined success metrics and timelines.

    The winners will start becoming identifiable by name. The BCG “future-built” 5% is currently an anonymous cohort. As the field matures, the firms that built the right infrastructure and redesigned workflows rather than layering AI on top of broken processes will start producing public case studies with real numbers. Those case studies, when they arrive, will be more valuable than any survey data.

    CFO scrutiny will reshape how pilots get approved. IBM found that only 25% of AI initiatives delivered expected ROI. That number is entering boardroom conversations. Finance leaders who previously approved AI spend on the basis of competitive parity (“our competitors are doing this”) are beginning to demand pre-defined success metrics and ROI timelines before sign-off. That shift, if it continues, will produce fewer pilots and better ones.

    Three Actions for Technology Leaders Right Now
    1. Audit your data infrastructure before approving any new AI spend. Gartner’s data consistently shows that data readiness, not model selection, is the primary predictor of AI success.

    2. Define success in business terms, with a timeline and an owner, before a pilot launches. “Measurable reduction in customer support costs by Q3” is a success metric. “Explore AI capabilities” is not.

    3. Consider stopping two current pilots before starting one new one. The evidence suggests that focus produces better outcomes than portfolio diversification when it comes to enterprise AI.


    FAQ: Enterprise AI Failure Rates

    Why do most enterprise AI projects fail?
    Independent research from MIT, RAND, McKinsey, and S&P Global converges on organizational causes rather than technical ones: poor data readiness, unclear success metrics, weak workflow integration, and treating AI deployment as a technology rollout instead of a change-management initiative. The models mostly work. The organizations often don’t.

    What percentage of AI projects fail in 2026?
    Estimates vary by study and definition. MIT found 95% of GenAI pilots show no measurable P&L impact. McKinsey found only 39% of organizations report any enterprise-wide EBIT impact. S&P Global found 42% of companies abandoned most AI initiatives in 2025. No single authoritative percentage exists across all AI project types, but the consistent finding is that fewer than 10% of organizations capture most of the value.

    Is the MIT 95% AI failure statistic accurate?
    The MIT NANDA report is real, published in July 2025, based on 300-plus initiative reviews and 52 organizational interviews. The 95% figure reflects a strict 6-month P&L-only definition of success. Marketing AI Institute’s Paul Roetzer has publicly challenged the methodology for excluding efficiency and productivity gains. The figure is directionally useful but should not be treated as a precise universal failure rate.

    How much are companies investing in AI in 2026?
    Stanford HAI’s AI Index 2025 tracked $252.3 billion in corporate AI investment in 2024, with private investment rising 44.5% year-over-year. Gartner forecasts worldwide AI spending will reach $2.52 trillion in 2026. GenAI-specific pilot investment was estimated at $30 to $40 billion in 2025 per MIT NANDA’s own baseline.

    What do successful enterprise AI programs have in common?
    BCG’s analysis of 1,250-plus global firms found that the 5% achieving AI value at scale share three traits: they identify one specific business pain point rather than launching broad pilot portfolios, they redesign workflows around AI rather than layering AI onto existing processes, and they fix data infrastructure before deployment rather than after. The differentiator is organizational discipline, not model selection.


    The enterprise AI ROI crisis is not a story about artificial intelligence failing. It is a story about organizations failing to create the conditions under which AI can succeed. The technology works. The problem is that most enterprises are deploying it into environments it cannot fix: fragmented data, undefined success metrics, siloed workflows, and approval processes driven by competitive anxiety rather than business logic.

    The 5% that are winning are not using better models. They built better foundations first.

    Over the next 18 months, as Gartner’s predicted “Trough of Disillusionment” plays out and CFO scrutiny tightens, the gap between that leading cohort and the field will widen. The organizations that survive the trough will be the ones that treated their first round of AI failures as diagnostic information rather than sunk costs.

    For more on how enterprises are navigating this gap, read our analysis of why CTOs are falling behind on AI skills and our breakdown of the shadow AI crisis hitting enterprise governance.

    Stay Ahead of the AI Signal

    The Neural Loop delivers the enterprise AI research that actually matters, without the vendor spin. One email, twice a week.

    Subscribe to The Neural Loop