Category: Artificial Intelligence

In-depth artificial intelligence analysis: AI agents, LLMs, enterprise deployment, governance, and breakthroughs. Research-backed insights for CTOs, founders, and decision-makers.

  • Colorado AI Act SB 26-189: What Employers Must Know

    Colorado AI Act SB 26-189: What Employers Must Know

    Colorado AI Act SB 26-189: What Employers Must Do by 2027
    AI Regulation · Employment Law

    Colorado’s AI Law Died Before It Lived. Here’s What’s Next

  • China AI Export Ban 2026: Qwen and DeepSeek at Risk

    China AI Export Ban 2026: Qwen and DeepSeek at Risk

    China May Ban Its Own AI Models: Qwen, DeepSeek at Risk
    Artificial Intelligence / Policy

    China Is Reportedly Weighing Its Own AI Model Export Ban

  • Gartner Multimodal AI 2030 Forecast: Now the Default

    Gartner Multimodal AI 2030 Forecast: Now the Default

    Artificial Intelligence

    Multimodal AI Enterprise Adoption Is Now the Default

    Your next vendor RFP just changed shape. Multimodal AI enterprise adoption is no longer a checkbox feature you evaluate after picking a model, it’s the baseline architecture assumption you build the RFP around. Gartner says 80% of enterprise software will be multimodal by 2030, up from under 10% in 2024. That’s not a slow curve. That’s a rewrite of procurement criteria happening while most teams are still finishing their 2026 roadmap.

    Here’s the tension nobody’s resolving cleanly: the same month frontier labs pushed multimodal models to mass-market default pricing, a peer-reviewed study in Nature Medicine found those same models reasoning incorrectly under adversarial testing, even when they landed on the right answer. Adoption and reliability are moving on different timelines. This piece is about both, because you can’t plan around one without the other.

    The adoption curve, in Gartner’s own numbers

    Gartner has now published two forecasts, a year apart, that both point the same direction. In September 2024, Distinguished VP Analyst Erick Brethenoux told the Gartner IT Symposium that 40% of generative AI solutions would be multimodal by 2027, up from just 1% in 2023. By July 2025, the firm went further: 80% of enterprise software and applications will be multimodal by 2030, up from less than 10% in 2024, according to Senior Director Analyst Roberta Cozza.

    ForecastBaselineTargetSource
    Generative AI solutions, multimodal1% (2023)40% by 2027Gartner, Sept. 2024
    Enterprise software and apps, multimodal<10% (2024)80% by 2030Gartner, July 2025
    Enterprise apps with task-specific AI agents<5% (2025)40% by end of 2026Gartner, Aug. 2025
    Multimodal is a fundamental transformation, letting AI shift from supporting individual productivity to proactive, contextual decision intelligence across healthcare, finance, and manufacturing.
    Roberta Cozza, Senior Director Analyst, Gartner · Gartner press release, July 2025
    Note the small inconsistency across Gartner’s own materials: some releases cite the 2024 baseline as “less than 5%,” others say “less than 10%.” Neither figure changes the shape of the curve, but it’s worth knowing the exact baseline moves depending on which Gartner document you’re reading.

    Real-world numbers back the direction, if not the pace. Two recent frontier releases landed within a day of each other on June 30, 2026: Anthropic’s Claude Sonnet 5 became the default model for every free and paid Claude user starting July 1, and Google shipped two new multimodal image models, Gemini 3.1 Flash Image and Gemini 3 Pro Image, through Google AI Studio. Neither company is treating multimodal as a premium add-on anymore. It’s the base tier.

    Why enterprises are consolidating around multimodal now

    Picture a claims adjuster at a mid-size insurer. Five years ago, that job meant one tool for reading the intake form, another for the damage photos, a third for the call transcript, and a spreadsheet to stitch it all together. Multimodal AI enterprise adoption promises to collapse that into one system that reads the form, looks at the photo, and listens to the call in the same pass. That’s the pitch, and it’s why McKinsey found 88% of organizations now use AI in at least one business function, with generative AI use jumping to 72% from just 33% in 2024.

    But adoption and scale are different claims. The same McKinsey survey found nearly two-thirds of organizations haven’t started scaling AI across the enterprise. Most of what gets counted as “multimodal adoption” in market surveys is still pilots, not production.

    According to Distinguished VP Analyst Erick Brethenoux, the case for native multimodal architecture is structural: real-world data was never single-format to begin with, and stitching together separate vision, audio, and text models introduces latency and accuracy problems that a unified model avoids.

    The Nature Medicine problem: benchmarks lie

    Here’s the part the vendor decks leave out. A peer-reviewed study published in Nature Medicine in June 2026, “Evaluating the robustness and readiness of large frontier models in health AI applications,” stress-tested frontier multimodal models, including GPT-5, Claude 3.5, and Gemini 2.5 Pro, on multimodal medical reasoning tasks. Researchers used adversarial perturbations, removing key details from an image or swapping which modality carried the critical information, and found the models frequently reached the correct answer for the wrong reasons. That means faulty reasoning, inappropriate shortcuts, and outright hallucinations were hiding behind passing benchmark scores.

    Why this matters for your rollout: A model that scores well on a public multimodal benchmark isn’t the same as a model that reasons reliably when the input is messy, adversarial, or simply real. The Nature Medicine authors concluded that popular health benchmarks don’t reliably measure multimodal robustness at all.

    The finding echoes a related pattern documented in Communications Medicine: across 300 doctor-designed clinical vignettes, leading LLMs repeated or built on a single planted fake lab value or diagnosis in up to 83% of cases before any mitigation prompt was applied. Explicit “verify before answering” instructions roughly halved the error rate. They didn’t eliminate it.

    One caveat worth flagging for readers who follow this closely: by the time a peer-reviewed paper like this clears review, the exact models it tested are often a generation behind whatever just shipped. That’s a structural limitation of academic AI evaluation, not evidence the newest models are automatically safer. Treat it as a reason for more testing, not less.

    The contrarian read: Gary Marcus and the ROI gap

    Not everyone is buying the adoption-curve optimism, and it’s worth hearing the strongest version of that case. NYU professor emeritus and longtime AI reliability critic Gary Marcus has argued for months that generative and multimodal systems remain fundamentally unreliable regardless of which lab built them, and that reported enterprise ROI hasn’t come close to matching the capital poured into these systems.

    The industry keeps converging on models with essentially the same class of reasoning flaws, no matter how much scale you throw at them, and the spending-to-revenue gap tells its own story.
    Gary Marcus, cognitive scientist and NYU professor emeritus · Marcus on AI, June 2026
    Marcus has specifically pointed to the Nature Medicine findings as proof that frontier multimodal models “are not ready” for high-stakes reasoning, and he’s not alone in reading McKinsey’s own numbers as a warning sign rather than a victory lap. A companion 2025 McKinsey survey found more than 80% of respondents weren’t yet seeing measurable EBIT impact from generative AI. Adoption curve and value capture are two separate stories, and they get conflated constantly.

    Our read: the skeptics aren’t wrong that governance is lagging. McKinsey’s 2026 AI Trust Maturity Survey put the average Responsible-AI maturity score at just 2.3 out of a possible higher band, up only slightly from 2.0 in 2025, with roughly a third of organizations scoring 3 or above on strategy and agentic-AI governance. Capability is outrunning oversight, and that gap is exactly where the Nature Medicine failures live.

    What this means for your stack

    If you’re the one signing off on the next platform migration, three things follow directly from the research above:

    • Assume multimodal ingestion by default. Document, image, audio, and video inputs should be evaluation criteria from day one of any vendor RFP, not a phase-two add-on.
    • Match the use case to the confidence level. Practitioner reporting from July 2026 converges on the same lesson: multimodal pays off in high-friction, measurable workflows like support tickets with screenshots or full-coverage compliance QA, not in low-stakes novelty pilots.
    • Fund governance at the same pace as capability. If your Responsible-AI maturity score would land near McKinsey’s 2.3 average, that’s your signal to slow autonomous, unsupervised deployment in regulated domains until review processes catch up. NeuralWired’s own reporting on AI code review adoption found a similar pattern: capability scaling faster than the human oversight built to catch its mistakes.
    There’s precedent for how this plays out badly. Gartner has separately warned that more than 40% of agentic AI projects will be abandoned by 2027 over cost, unclear value, or inadequate risk controls, and NeuralWired’s reporting on AI agent deployment failures found roughly 70% of agent projects never reach production. Multimodal rollouts are highly likely to follow the same adoption-curve-versus-production-reality gap.

    Frequently asked questions

    What is multimodal AI in enterprise environments?

    Multimodal AI refers to systems that process and reason across more than one data type, text, images, audio, video, and structured data, within a single unified model rather than separate tools per format. Gartner projects 80% of enterprise software will be multimodal by 2030, up from under 10% in 2024.

    Why are enterprises investing in multimodal AI in 2026?

    Enterprises are consolidating fragmented single-modality tools into unified platforms to cut integration overhead, reduce latency, and enable workflows like reviewing contracts, call recordings, and dashboards together. McKinsey reports 88% of organizations now use AI in at least one business function.

    Is multimodal AI reliable enough for high-stakes decisions?

    Not yet, based on peer-reviewed evidence. A June 2026 Nature Medicine study stress-tested frontier multimodal models on medical reasoning and found faulty logic, inappropriate shortcuts, and hallucinations under adversarial testing, meaning benchmark scores alone don’t prove real-world robustness.

    What’s the difference between multimodal AI and agentic AI?

    Multimodal AI is about perception: processing text, images, audio, and video together. Agentic AI is about action: autonomously executing multi-step tasks. Gartner projects agentic AI capability will reach 40% of enterprise applications by the end of 2026, typically built on multimodal foundations.

    How much of enterprise AI adoption is still just piloting, not production?

    A significant majority. McKinsey found that while 88% of organizations use AI somewhere in the business, nearly two-thirds haven’t begun scaling AI programs across the enterprise, meaning most “adoption” headlines still describe isolated pilots rather than production systems.


    What to watch next

    The honest version of this story has two halves that both hold up under scrutiny. Gartner’s forecasts describe real, well-documented product availability: multimodal is becoming the default architecture, not a premium tier. The Nature Medicine findings describe something different and equally real: benchmark performance and production-grade reliability are not the same claim, and right now the evidence for the second one is thinner than the marketing around the first.

    Over the next 6 to 18 months, watch three things. First, whether McKinsey’s Responsible-AI maturity scores climb faster than the 2.0-to-2.3 pace they’ve shown so far, since that gap is what’s actually gating safe deployment. Second, whether the next generation of academic evaluation catches up to model release cycles, so reliability claims stop lagging capability claims by a full peer-review cycle. Third, whether the 40%+ agentic-AI-project abandonment rate Gartner is forecasting for 2027 repeats itself in multimodal rollouts specifically, or whether the sector learns from the agentic AI stumble first.

    None of that means wait. It means build for the workflows where multimodal already earns its cost, and keep governance funded at the same pace as capability.

    Want the next multimodal AI enterprise adoption story before it breaks? Subscribe to The Neural Loop.

    Subscribe at neuralwired.com/newsletter →
  • Gartner: AI Agent Governance Rules Are Failing 2026

    Gartner: AI Agent Governance Rules Are Failing 2026

    AI Agent Governance 2026: Why ‘One Size’ Rules Fail | NeuralWired Enterprise AI / Governance

    AI Agent Governance 2026: Why ‘One Size’ Rules Fail

    Your AI agent can already read your database, draft an email, and push a config change. The question nobody in the room can answer is who signed off on that, and whether anyone would even notice if it went wrong. That gap has a name now: AI agent governance, and Gartner just told the industry it’s building the wrong kind.

    On May 26, 2026, Gartner published research warning that enterprises applying identical governance rules to every AI agent, regardless of what that agent can actually do, are setting themselves up to fail. The firm’s prediction is blunt: by 2027, 40% of enterprises will demote or decommission autonomous AI agents after governance gaps surface the hard way, in production, after something breaks.

    If you’re a CTO, CISO, or VP of Engineering deciding what your agent fleet is allowed to touch next quarter, this is the framework everyone else is now quoting. Here’s what it actually says, what the data shows is already happening, and what changes on your calendar because of a deadline that isn’t hypothetical: August 2, 2026.

    The binary governance problem

    Most organizations still treat AI agent governance as a light switch: locked down or fully trusted, nothing in between. Shiva Varma, Senior Director Analyst at Gartner and the author of the May 26 research, says that’s exactly the root cause of the failures his team is now tracking.

    “Agents operate at different autonomy levels and across different trust boundaries.” Shiva Varma, Senior Director Analyst, Gartner
    Gartner Newsroom, May 26, 2026
    Apply heavy controls to a document-summarizing agent and you get a bottleneck: delivery slows, and engineers start building unsanctioned workarounds instead of waiting for approval. That’s shadow AI, and it’s a governance failure in its own right. Flip it around and under-restrict a powerful, autonomous agent, and you’ve expanded your attack surface without expanding your ability to see it.

    CIO Dive’s follow-up interview with Varma put it more plainly still: a lot of companies simply don’t have agent-specific governance at all, they have one blanket policy stretched over everything.

    Gartner’s four autonomy tiers, explained

    Gartner’s fix isn’t more governance across the board. It’s proportional governance, matched to what each agent can actually do. The framework splits agents into four tiers by autonomy level, and pairs each with the controls that tier actually needs, not more, not less.

    Tier What the agent does Governance required
    Observe Read-only access, outputs visible only to the requesting user. Document summarization, retrieval, code explanation. Scoped access, authentication, usage logging, basic testing.
    Advise Generates recommendations or drafts; a human reviews and executes manually. Output-quality review, hallucination testing, reliance training.
    Act with approval Writes data, sends communications, or changes configurations, only after explicit human sign-off per action. Security testing, clear approval workflows with audit trails, agent-specific incident response.
    Act autonomously Executes independently within set guardrails; humans review exceptions and aggregated outcomes, not individual decisions. Continuous monitoring, enforced guardrails, rollback mechanisms, circuit breakers.
    The third tier is where Varma’s warning gets sharpest. Human-in-the-loop approval only works as a control if it stays meaningful, and under time pressure, approval fatigue quietly turns a real check into a rubber stamp. And the fourth tier carries its own physics problem: once an agent acts on its own, it operates at a speed no human reviewer can keep pace with in real time. That’s why circuit breakers and rollback mechanisms aren’t optional at that level, they’re the only brake left.

    Gartner adds one more distinction worth sitting with: autonomy and access scope are two separate dials, not one. An agent can be low-autonomy but high-scope (it touches a lot of systems, but a human approves every move), or high-autonomy but narrow-scope. Risk climbs with either dial, independently.

    The data: this is already causing incidents

    None of this is theoretical. The numbers from three separate 2026 surveys point the same direction: deployment is outrunning oversight, and it’s already producing damage.

    The gap, in four numbers:
    • 88.4% of organizations had at least one AI-agent-related security breach in the past 12 months, per AvePoint’s State of AI 2026 report (750 IT leaders surveyed).
    • ~52% average monitoring coverage across deployed agents, meaning roughly 48% run with no meaningful oversight, per Gravitee’s State of AI Agent Security report (750 senior technology leaders, April 2026).
    • 7.2% of organizations have a single named person formally accountable for agent behavior. The rest call it unclear, informally shared, or simply undiscussed. (Gravitee, same survey.)
    • 62% of organizations now name security and risk, not technical limits, as the top barrier to scaling agentic AI, according to Stanford’s 2026 AI Index, cited by Speakeasy.
    Put those together and you get a picture that should worry anyone signing off on an agent rollout: agent fleets roughly doubled in size since December 2025, while monitoring coverage barely moved. The fleet is growing faster than anyone’s ability to watch it.

    Anushree Verma, another Senior Director Analyst at Gartner, offers a useful counterweight here. Much of what gets called “agentic AI” in 2026 is still early and experimental, and treating it as more mature than it is can blind teams to what real production deployment actually costs. That matters: some of the governance panic is running ahead of how much genuinely autonomous work is happening yet. But it doesn’t erase the incident numbers above, and it doesn’t change who’s accountable when the agents that are live go wrong.

    The August 2026 deadline you can’t negotiate

    If your agents touch EU users in employment, credit, insurance, or critical infrastructure decisions, there’s a date on the calendar that matters more than any vendor roadmap. The EU AI Act’s high-risk system obligations reach full enforcement around August 2, 2026, requiring documented human oversight, record-keeping, and audit logging for those systems.

    The penalties aren’t symbolic. Fines scale up to €35 million or 7% of global annual revenue, and they apply regardless of where the company is headquartered, as long as outputs reach EU users. Headquarters in Austin doesn’t buy you an exemption if your hiring agent screens applicants in Berlin.

    Kiteworks’ 2026 forecast puts a sharper edge on why this matters right now: 63% of organizations can’t currently enforce purpose limitations on their AI agents, and 60% can’t terminate a misbehaving one. An agent you cannot stop is, by definition, an agent without governance. That’s not a compliance nuance, that’s the whole ballgame.

    What mature governance actually looks like

    The cloud vendors spent Q2 2026 building governance into the product, not bolting it on after. Microsoft made its Agent 365 SDK generally available at Build 2026, pairing it with an Execution Container SDK and Purview data-loss-prevention for agent prompts. Google built its Gemini Enterprise Agent Platform around an Agent Identity and Agent Registry system, giving every agent a cryptographic identity separate from any human user. AWS took the lighter path, leaning on Bedrock AgentCore to get agents into production fast while still offering identity and tool management.

    The case study everyone in this space keeps citing is Uber’s internal build: an LLM gateway handling PII redaction and audit logging across every model call, an MCP gateway governing every agent-to-tool connection across more than 10,000 internal services, and an agent identity system with cryptographically attested lineage on every action taken.

    Worth saying plainly: that took Uber years and a dedicated platform engineering team whose only job was AI infrastructure. Most companies reading this don’t have that team, and they don’t have that runway either. Uber is proof the model works, not a template you can copy over a weekend.

    The skeptic’s case

    A fair amount of the loudest governance-urgency content in 2026 comes from companies that sell governance software. The underlying statistics are usually real and independently sourced, but the framing tends to land in the same place: buy the platform. Worth reading the data and discounting the pitch separately.

    There’s a sharper irony buried in Gartner’s own research. The firm’s 2026 Hype Cycle for Agentic AI places governance and security tooling on the curve as an early, still-maturing category, not a solved one. Enterprises are being told to urgently adopt governance platforms in a product category Gartner itself flags as immature. That’s not a reason to skip governance. It’s a reason to be honest that the tools for doing it well are still catching up to the sales pitch.

    A more pointed critique comes from outside the analyst world entirely. A recent opinion piece put the capability gap bluntly: in practice, today’s AI agents behave less like autonomous employees and more like “junior staffers who work quickly, confidently and often incorrectly.” That’s commentary, not analyst research, but it’s a useful check on any narrative that assumes agents are already reliable enough that governance is the only thing standing between them and full autonomy.

    What to do this quarter

    You don’t need a platform purchase to make progress before your next planning cycle. Three moves cost nothing but time.

    1. Tier your existing agents. Sort every live agent into Observe, Advise, Act-with-approval, or Act-autonomously. Most teams have never done this classification exercise, and it surfaces mismatches immediately.
    2. Name an owner. Only 7.2% of organizations have done this. It costs nothing and it’s the single most concrete accountability fix available right now.
    3. Check your kill switch. If you can’t answer, in one sentence, how you’d stop a specific agent from acting in the next five minutes, that’s your highest-priority gap, ahead of any new deployment.
    Our read: the enterprises that get hurt in 2027 won’t be the ones that moved slowly on agents. They’ll be the ones that scaled fast without ever doing the tiering exercise above, then discovered their most powerful agent had the governance of their least powerful one.


    Frequently asked questions

    What is AI agent governance?

    AI agent governance is the set of policies, ownership structures, and enforcement controls that determine what AI agents are allowed to do, on whose authority, and under what regulatory constraints, covering identity, permissions, monitoring, and accountability for systems acting on a company’s behalf.

    Why does AI agent governance matter in 2026?

    Gartner found 62% of organizations now cite security and risk, not technical limits, as their top barrier to scaling agentic AI. AvePoint reports 88.4% had at least one agent-related security incident in the past year, and roughly 48% of deployed agents run without adequate monitoring.

    What happens if a company doesn’t govern its AI agents?

    Gartner predicts 40% of enterprises will demote or decommission autonomous AI agents by 2027 after governance gaps surface through real incidents. Ungoverned agents also create direct EU AI Act exposure, with fines reaching €35 million or 7% of global revenue for high-risk systems.

    What are Gartner’s four AI agent autonomy levels?

    Observe (read-only, lightweight controls), Advise (drafts a human reviews and executes), Act with Approval (agent acts only after human sign-off on each action), and Act Autonomously (independent execution within guardrails, monitored through exception review, rollback, and circuit breakers).

    When does the EU AI Act apply to AI agents?

    High-risk obligations under the EU AI Act, covering agents used in employment, credit, insurance, and critical infrastructure, reach full enforcement around August 2, 2026, requiring documented human oversight, audit logging, and conformity assessments regardless of where the company is headquartered.

    Who is responsible for AI agent behavior inside a company?

    Currently, almost no one, formally. Only 7.2% of organizations report having a single named individual with accountability for agent behavior, according to Gravitee’s April 2026 survey of 750 senior technology leaders. Most describe accountability as unclear or undiscussed.


    Where this goes next

    Here’s what you now know that you didn’t ten minutes ago: governance isn’t a checkbox you add after deployment, it’s a dial you set per agent, based on what that agent can actually touch and how fast it can act. Uniform rules break in both directions, over-restricting the harmless agents and under-restricting the dangerous ones.

    Watch three things over the next 6 to 18 months. First, whether Gartner’s 40%-decommission prediction starts showing up as real earnings-call language from enterprises walking back agent rollouts. Second, whether the governance platform market (projected past $1 billion by 2030) actually matures fast enough to catch up with the Hype Cycle placement it currently sits at. Third, how EU regulators enforce the August 2026 deadline in the first few months, since the first fine or the first quiet non-enforcement will set the tone for everyone watching from outside the bloc.

    None of this requires a platform purchase to start. Tiering your agents and naming an owner are free, and they’re the two moves most companies still haven’t made.

    Related coverage: Cursor AI Code Review: 86% Now Skip Human Checks, Klarna, Replit, Zillow: 12 Companies Whose AI Failed, and The $52 Billion Question: Why 70% of AI Agent Deployments Fail.

    Want this kind of breakdown in your inbox? Subscribe to The Neural Loop at neuralwired.com/newsletter for the enterprise AI stories that matter, before they hit everyone else’s feed.
  • EU AI Act August 2026 Deadline: What Really Changes

    EU AI Act August 2026 Deadline: What Really Changes

    EU AI Act’s Real August 2 Deadline: What Actually Changes Regulation / EU Tech Policy

    The EU AI Act’s Real August 2 Deadline: What Actually Changes

  • Cursor SpaceX $60B Deal: AI Code Review Risks 2026

    Cursor SpaceX $60B Deal: AI Code Review Risks 2026

    Cursor, SpaceX, and the End of Human Code Review
    Artificial Intelligence / Software Engineering

    Cursor, SpaceX, and the End of Human Code Review

  • MLflow 3.0: Databricks Merges MLOps and LLMOps in 2026

    MLflow 3.0: Databricks Merges MLOps and LLMOps in 2026

    MLOps vs LLMOps: Why the Split Just Ended in 2026
    Machine Learning · Enterprise AI

    MLOps vs LLMOps: Why the Split Just Ended in 2026

    Databricks, CoreWeave, and Weights & Biases have already merged the tooling. Most enterprise teams have not, and that gap is quietly draining their AI budgets.

    Somewhere inside a mid-size bank right now, one team is watching a fraud model’s accuracy drift on a Tuesday afternoon dashboard. Down the hall, a different team is squinting at a LangSmith trace trying to figure out why the company’s new support chatbot just hallucinated a refund policy. Neither team talks to the other. Neither uses the same registry, the same on-call rotation, or the same vocabulary for “this broke in production.”

    That split is the whole story of MLOps LLMOps convergence in 2026. The platforms that manage classical machine learning and the platforms that manage large language models are merging into a single discipline, driven by real product launches and real acquisitions, not by a marketing buzzword. But the merger is happening at the vendor level far faster than it’s happening inside actual companies. Teams still running two separate stacks are paying for it in duplicate infrastructure, duplicate headcount, and blind spots that show up right when an AI agent goes off the rails in front of a customer.

    This piece breaks down what’s actually converging, what the data says, where the maturity gap still bites, and what to do about it if you’re the person who has to justify the tool budget next quarter.

    What’s Actually Converging (And What Isn’t)

    Start with the clearest evidence: Databricks shipped MLflow 3.0 in June 2025, and it wasn’t a minor version bump. The release was built to bring the same rigor Databricks already applied to classical ML models to generative AI workloads, on one platform, so teams stop juggling separate systems for the two. It added tracing across more than 20 GenAI libraries, LLM-judge style evaluation, and one shared registry for models, prompts, and datasets through Unity Catalog.

    MLflow isn’t a niche tool. The open-source project sits at over 30 million monthly downloads with contributions from more than 850 developers, which makes it the closest thing MLOps has to a standard, and the fact that Databricks pointed that standard directly at LLM workloads is a signal worth taking seriously.

    Then there’s the money. In March 2025, CoreWeave agreed to acquire Weights & Biases, one of the most established names in ML experiment tracking. CoreWeave CEO Michael Intrator didn’t frame the deal as buying an MLOps company or an LLMOps company. He framed it as buying both categories at once, folded into infrastructure CoreWeave already sells.

    “Weights & Biases has built a phenomenal platform to help organizations of any size and across a range of industries to build, deploy and monitor AI training and inference applications.” Michael Intrator, Co-founder & CEO, CoreWeave — CoreWeave official announcement
    Weights & Biases now sells two products under one roof on purpose: W&B Models for the classical MLOps work (training, fine-tuning, deployment) and W&B Weave for LLMOps (tracing, evaluation of non-deterministic outputs). The company’s own positioning is “one platform, one audit trail, from first notebook to production LLM.” That’s not incidental phrasing. It’s the whole pitch.

    W&B CTO Shawn Lewis told VentureBeat that Weave was never meant to stand alone.

    “It’s foundational, so there’s a lot that you can do on top of this.” Shawn Lewis, CTO & Co-founder, Weights & Biases — VentureBeat
    This isn’t only a vendor story. PayPal extended its internal MLOps platform, Cosmos.AI, to natively handle LLM workloads, adding retrieval-augmented generation, semantic caching, and prompt management directly onto infrastructure it already had, rather than standing up a second stack. Uber built a unified “GenAI Gateway” mirroring the OpenAI API spec to serve both external and self-hosted models across more than 60 internal use cases. Neither company treated the LLM layer as a separate discipline requiring a separate org chart.

    Our read: the pattern across every one of these examples is the same. Nobody built a parallel LLMOps stack from scratch and kept it walled off. Every serious player extended what already worked for classical ML and bolted LLM-specific capability on top. If your team is planning a from-scratch LLMOps buildout in 2026, that’s worth questioning before you sign anything.

    The Numbers: How Big Is This, Really

    The market-sizing reports diverge, sometimes by 20 to 40 percent, depending on how each firm scopes “MLOps.” That’s normal for a young category, but it means no single number deserves to be treated as gospel.

    Grand View Research, the most methodologically transparent of the reports reviewed for this piece, puts the MLOps market at roughly $2.19 billion in its 2024 base year, projected to reach $16.6 billion by 2030, a compound annual growth rate above 40 percent. Fortune Business Insights puts 2026 alone at $4.39 billion, heading toward $89.91 billion by 2034. Precedence Research lands closer to $3.33 billion for 2026, reaching $56.6 billion by 2035. Three different firms, three different numbers, one consistent direction: steep, sustained growth concentrated in the platform segment rather than point tools.

    LLMOps, meanwhile, is already nearly its own heavyweight category. Estimates put the LLMOps market at $7.14 billion in 2026, growing to $15.59 billion by 2030. That means LLMOps alone is now roughly the size the entire MLOps market was just two years ago. These aren’t two small categories slowly circling each other. They’re two large, fast-growing budgets on a collision course.

    The adoption pressure behind all of this is agents. Gartner estimates that 40 percent of enterprise applications will feature AI agents by 2026, up from under 5 percent in 2025. Agents need both classical-ML-style evaluation gates and LLM-style prompt and tool governance running at the same time, which is precisely the kind of workload a split toolchain struggles to support.

    And the failure rate underneath all this growth is not small. A widely cited figure puts the share of AI and ML models that never reach production above 85 percent. Separately, S&P Global Market Intelligence found that 42 percent of companies abandoned most of their AI initiatives in 2025, more than double the 17 percent abandonment rate the year before.

    MLOps vs. LLMOps vs. Unified Platforms

    Dimension Classical MLOps LLMOps Unified / xOps (2026)
    Core artifact Trained model weights, features Prompts, RAG pipelines, agent traces Shared registry for models, prompts, datasets
    Evaluation method Deterministic metrics (accuracy, F1, drift) Non-deterministic, LLM-as-judge, human review Combined eval pipelines with both metric types
    Maturity Standardized since roughly 2019 to 2022 3 to 4 years younger, not yet standardized Emerging, led by vendors, not yet universal
    Typical tools MLflow, Kubeflow, DVC LangSmith, Langfuse, Braintrust, Portkey MLflow 3.0, W&B Models + Weave
    Cost profile Predictable, per-prediction Can run roughly 100x the cost per inference Single FinOps layer covering both, still maturing

    The Tax: Why Fragmented Teams Are Paying For This

    Here’s the tension the vendor press releases don’t put in the headline: platform convergence is real, but tool-stack convergence inside most companies is lagging well behind it. Practitioner guides reviewed for this piece describe enterprise LLMOps deployments that still stitch together three to five specialized tools, a tracing tool like LangSmith or Promptflow, an observability layer like Arize AI or Langfuse, a registry like MLflow, an eval pipeline like Braintrust, and a gateway like Portkey or LiteLLM, because no single platform yet covers the whole stack end to end.

    That’s the tax. Every one of those tools needs its own login, its own on-call rotation, its own budget line, and its own translation layer back to whatever the classical ML team is running. ISG’s Jeff Orr put the broader platform-strategy version of this argument plainly.

    “Platform consolidation is no longer an efficiency play. It is now a structural necessity.” Jeff Orr, Director of Research, IT and Technologies, ISG
    Is that overstated? Maybe a little, depending on your company’s size. But the direction is hard to argue with once you look at where budget is actually flowing. Both Grand View Research and Fortune Business Insights show double-digit growth concentrated specifically in the “platform” segment rather than point solutions, meaning the money is already voting for consolidation even where the org chart hasn’t caught up yet.

    The Skeptic’s Case: Governance Is the Real Bottleneck

    Not every analysis buys the clean convergence story, and it’s worth sitting with the pushback. Practitioner research from Atlan argues that LLMOps tooling is structurally three to four years younger than MLOps tooling and simply hasn’t standardized the way MLflow, Kubeflow, and DVC did between 2019 and 2022. Their analysis ties this to a governance deficit rather than a tooling gap: one financial institution’s LLM gateway logs can’t be connected back to its governance platforms at all. Another enterprise, per the same research, still stores its AI model information in PowerPoint.

    That last detail is almost funny until you remember it’s describing companies making real deployment decisions in 2026. Unifying the ops tooling doesn’t retroactively fix an organization’s data lineage practices or its audit trail. A single dashboard sitting on top of a governance mess is still a governance mess, just with a nicer front end.

    Our read: the “platforms have merged” claim is true. The “discipline has merged” claim is not, at least not yet. Treat vendor unification announcements as directionally correct on tooling and meaningfully premature on governance, compliance, and cost attribution. Cost is the sneakiest part of this: a single LLM inference can run roughly 100 times the cost of a traditional ML prediction, so a genuinely unified FinOps layer has to reconcile two wildly different cost profiles under one roof. That’s a much harder systems problem than unifying an experiment tracker, and it’s exactly the part MLflow 3.0 and the W&B deal have not fully solved yet.

    What CTOs Should Actually Do Now

    If you’re the one deciding whether to consolidate, a few things matter more than the vendor slide deck.

    • Verify LLM-specific depth before you consolidate. A unified registry is only as good as its weakest layer. Check tracing coverage, eval rigor, prompt versioning, and guardrail integration against what your current point tools already do, don’t assume feature parity with five-plus years of mature MLOps tooling.
    • Follow the PayPal and Uber model, not a rip-and-replace. Both companies extended existing MLOps infrastructure instead of building a parallel LLMOps org from zero. That’s a lower-risk path than a wholesale platform swap.
    • Fix data lineage before you fix the dashboard. If your model information still lives in spreadsheets or PowerPoint, a unified platform will not solve that for you. Governance work has to happen in parallel with, not after, tooling consolidation.
    • Budget for the cost-attribution problem separately. Don’t assume your FinOps tooling for classical models will cleanly extend to LLM inference costs. It’s a different order of magnitude and needs its own line item.

    Frequently Asked Questions

    What is the difference between MLOps and LLMOps?

    MLOps manages the lifecycle of traditional predictive models: training, versioning, deployment, and drift monitoring. LLMOps manages generative and foundation-model workloads: prompt versioning, retrieval-augmented generation, hallucination monitoring, and evaluation of non-deterministic output. In 2026, unified platforms increasingly handle both under one registry and observability layer.

    Is LLMOps part of MLOps?

    LLMOps functions more as an extension of MLOps than a fully separate discipline. It inherits MLOps’ versioning, CI/CD, and monitoring principles, then adds LLM-specific layers such as prompt pipelines, RAG evaluation, and cost-per-token tracking that classical MLOps tooling was never built to handle.

    Do companies need separate teams for MLOps and LLMOps?

    Not necessarily. PayPal extended its existing Cosmos.AI platform to cover LLM workloads with one team instead of standing up a parallel org. That said, most enterprises in 2026 still run three to five specialized LLM tools alongside their MLOps stack rather than a single unified toolchain.

    What is a unified AI operations platform?

    A unified AI operations platform, sometimes called “xOps,” manages classical ML models and LLM or GenAI applications through the same registry, monitoring, and deployment infrastructure. MLflow 3.0’s shared abstraction layer for both traditional ML artifacts and GenAI traces, prompts, and evaluations is the clearest current example.

    How big is the MLOps market in 2026?

    Estimates vary by research firm. Grand View Research projects the market growing toward roughly $16.6 billion by 2030 from a 2024 base near $2.2 billion. Fortune Business Insights puts 2026 alone at $4.39 billion, heading toward $89.91 billion by 2034. The wide range reflects differing scope definitions across methodologies, not disagreement about the growth trend itself.


    Where This Goes Next

    The vendor-level merger of MLOps and LLMOps is no longer a prediction. MLflow 3.0 shipped it, CoreWeave paid for it, and Weights & Biases built its whole product line around it. What hasn’t merged yet is the actual discipline inside most companies: the governance, the cost attribution, the on-call rotations, and the org charts that still treat classical ML and generative AI as two different jobs.

    Over the next 6 to 18 months, expect three things to matter more than the platform announcements themselves. First, watch whether unified vendors close the governance gap Atlan identified, not just the tracing gap. Second, watch cost-attribution tooling specifically, since that’s the systems problem nobody has solved cleanly yet. Third, watch whether agent adoption, which Gartner expects to hit 40 percent of enterprise applications this year, forces the remaining split-stack teams to consolidate faster than they’d planned, simply because agents don’t respect the old boundary between the two disciplines.

    The teams that treat this as a maturity-catch-up story, and not a symmetrical merger of two equally mature fields, are the ones that will avoid paying the tax twice.

    Want more research like this before it hits the mainstream feed? Subscribe to The Neural Loop at neuralwired.com/newsletter.