Author: Team_Neuralwired

  • FBI Deepfake Fraud Report: $893M Lost by Businesses

    FBI Deepfake Fraud Report: $893M Lost by Businesses

    Cybersecurity

    Nearly 3 in 10 Small Businesses Hit by Deepfake Scams in 2026

    NeuralWired Cybersecurity Desk · Published July 23, 2026

    In February 2024, a finance employee at UK engineering firm Arup joined what looked like a routine video call with the CFO and several colleagues. He wired $25.6 million across 15 transactions before anyone realized every face on that call except his own was AI generated. Two years later, that trick has trickled all the way down to businesses with a dozen employees and no IT department: 29% of small businesses now say they’ve experienced a deepfake scam in the past year, according to a new survey from cybersecurity firm VikingCloud.

    That number, buried inside VikingCloud’s 2026 SMB Threat Landscape Report, is the clearest signal yet that deepfake fraud stopped being an enterprise problem sometime in the last eighteen months. It’s now a Tuesday-afternoon problem for a plumbing company in Ohio or a marketing agency in Manchester. And the FBI, for the first time in its Internet Crime Complaint Center’s roughly 25-year history, agrees the threat is big enough to track on its own.

    1. What the New Numbers Actually Show2. The FBI Just Made It Official3. The Case That Changed Everything: Arup’s $25.6 Million Call4. When the Defense Works: WPP’s Near Miss5. Why Small Businesses Are the Easier Target6. Can You Trust Your Own Eyes?7. The Regulatory Clock Is Ticking8. Reader Beware: Not Every Stat Holds Up9. The One Habit That Beats the Software10. Frequently Asked Questions

    What the New Numbers Actually Show

    VikingCloud surveyed small business owners and operators for its 2026 threat report, and the results reorder what SMBs are worried about. More than a quarter said they’d experienced a deepfake scheme (29%), a customer data breach (27%), a ransomware attack (26%), or a denial of service attack (26%) in the past year. Taken together, 75% of SMB owners now rank cyberattacks as their number one operational threat for 2026, the first time in this survey series that cybersecurity has outranked economic pressure. Forty percent said a cyberattack costing $100,000 or less could put them out of business entirely.

    A note on the source VikingCloud hasn’t published full survey methodology, sample size, or margin of error in its public summary; the underlying data sits behind a lead-gen form. That doesn’t make the 29% figure false, but it means it should be read as “according to a vendor survey of small business owners,” not as census-grade data. Compare it against the FBI figure below, which is independently audited.

    The FBI Just Made It Official

    For 2025, the FBI’s Internet Crime Complaint Center broke out AI-enabled fraud as its own standalone category for the first time. IC3 logged 22,364 complaints with a reported AI nexus, totaling $893,346,472 in adjusted losses, according to the FBI IC3 2025 Annual Report published in April 2026. That figure is the closest thing this space has to a government-audited number, and it’s worth breaking down by category.

    Fraud category (AI referenced)2025 adjusted losses
    Investment fraud$632.0 million
    Business email compromise$30.3 million
    Tech and customer-support scams$19.5 million
    Confidence and romance scams$19.0 million
    Employment scams$12.6 million
    Business email compromise is the line that should matter most to a small business owner. It’s the category built entirely around impersonating someone the victim already trusts, a vendor, a boss, a bank contact, and it’s exactly the mechanism behind the Arup case.

    The Case That Changed Everything: Arup’s $25.6 Million Call

    Arup’s Hong Kong finance team received what appeared to be a standard request from the company’s UK-based CFO: move funds for a confidential transaction. The employee had doubts, so he did what security training tells you to do. He joined a video call to verify. Every other participant on that call, including the person who looked and sounded like the CFO, was an AI-generated deepfake. He made the transfers. Reporting from the Financial Times and CNN in May 2024 confirmed the total loss at $25.6 million across 15 wire transactions, and the case has become the reference point every security vendor cites when explaining why video verification alone is no longer enough.

    Arup is a global engineering firm with sophisticated finance operations, not a small business. That distinction matters, and we’ll come back to it. But the mechanics of the attack, real-time video and voice synthesis convincing enough to fool someone who was actively trying to verify, work exactly the same way against a five-person accounting team as they did against Arup’s.

    When the Defense Works: WPP’s Near Miss

    Not every attempt succeeds, and the counter-example is worth knowing. Scammers targeted WPP CEO Mark Read using a cloned voice and a spoofed Microsoft Teams meeting invite, built around a fake WhatsApp account using his public photo, according to an entry in the OECD.AI Incident Database and reporting from Marketing-Interactive. Staff escalated before any money moved. WPP confirmed zero losses.

    What stopped it wasn’t detection software. It was a human asking a question the scammer couldn’t answer and refusing to proceed until someone verified through a separate channel. That’s a cheap lesson, and it’s the same one at the center of the advice section below.

    Why Small Businesses Are the Easier Target

    Here’s the uncomfortable part for small business owners: being small isn’t protection. It’s the opposite. VikingCloud’s data shows 84% of SMB owners self-manage their own cybersecurity, with no dedicated IT or security staff. That means the same person approving a vendor invoice is also the last line of defense against a fraudulent one, with no gatekeeper, no second sign-off, no layered approval chain to slow things down.

    An enterprise like Arup still has structural weaknesses attackers can exploit, but it also has finance controls, compliance teams, and escalation paths. A twelve-person business usually has one bookkeeper and a Slack channel. Attackers know which door is easier to walk through.

    “We only have like one really good example in the news right now of that organization in Hong Kong that ended up falling for and sending $25 million based on a deepfake audio and video scam, and I think we’re going to see a lot more business email compromise style events because of AI.”Rachel Tobac, CEO, SocialProof Security · 8th Layer Insights podcast, The Cyber Wire, April 9, 2024

    Tobac’s prediction has aged into the current data. The FBI’s BEC-with-AI-nexus figure alone hit $30.3 million in 2025, and that’s before counting the cases that never get formally reported, which fraud researchers generally assume is the majority of them.

    “AI-generated media is not just a future risk, it’s a real business threat. We’re seeing executives impersonated, hiring processes compromised, and financial safeguards bypassed with alarming ease.”Tony Lee, Head of Consulting, Hong Kong & Macau, Trend Micro · Media OutReach Newswire, July 10, 2025

    Worth flagging: Lee’s employer, Trend Micro, sells deepfake detection tools, so treat the quote as an informed but interested voice rather than a neutral one.

    Can You Trust Your Own Eyes?

    Most SMB owners assume they’d notice if something felt off on a call. The data says otherwise. Controlled lab studies compiled by security research firm DeepStrike found human accuracy at spotting high-quality deepfake video sits at just 24.5%, even though roughly 60% of people believe they could identify one. That gap between confidence and competence is arguably the more dangerous number in this whole story.

    The technical barrier to producing convincing fakes keeps dropping too. McAfee’s consumer research found a voice clone with about 85% similarity to the original can now be generated from just three seconds of audio, easily pulled from a podcast clip, a local news interview, or a company’s own marketing video.

    The Regulatory Clock Is Ticking

    Two regulatory shifts land right around this article’s publish date. The EU AI Act’s Article 50 transparency rules, requiring disclosure and labeling of AI-generated content, take effect in August 2026, with penalties reaching €35 million or 7% of global turnover for noncompliance. Meanwhile, roughly 46 to 47 US states have now passed some form of deepfake-specific legislation, spanning election-related disclosure rules, non-consensual imagery protections, and fraud statutes, according to MultiState’s legislative tracking.

    None of this stops a scam call from reaching a small business tomorrow morning. But it does signal that lawmakers on both sides of the Atlantic have stopped treating deepfakes as a novelty problem.

    Reader Beware: Not Every Stat Holds Up

    Scroll through enough 2026 deepfake coverage and you’ll hit percentage increases that sound apocalyptic: 2,137%, 3,892%, four-digit growth claims stacked one after another. A research team at Digital Applied spent its July 2026 audit picking these apart, arguing that the field is crowded with numbers nobody actually verifies, loss figures with no traceable primary source, surge percentages that contradict each other depending on which vendor published them, and forecasts that get recycled as if they were measurements.

    Our read: most of those huge percentage jumps are real in direction but misleading in scale. A fraud category that goes from 0.1% to 6.5% of total fraud attempts, which is roughly what’s happened according to fraud-detection firm Signicat, produces an enormous percentage increase almost automatically, simply because it started near zero. That’s still a genuine and fast-growing threat. It’s just not the same thing as the flat “up 3,892% this year” headline that gets repeated without context.

    It’s also worth being honest about scale. Most of the largest documented deepfake losses, Arup’s $25.6 million among them, hit large enterprises with the kind of finance operations that can move eight figures in a single transfer. A small business physically can’t lose that much in one incident. The realistic SMB exposure looks more like tens of thousands of dollars per event, which is still enough to close a business operating on thin margins, but the “small businesses are next in line for a $25 million loss” framing overstates the individual stakes even while understating how often SMBs get hit.

    The One Habit That Beats the Software

    Security researchers keep landing on the same conclusion, and it isn’t a product pitch. Verizon’s Data Breach Investigations Report, cited across multiple 2026 industry analyses, consistently finds the human element involved in more than 60% of breaches. A basic callback-verification habit defeats a deepfake exactly as well as it defeats a decades-old phone scam, because the fake voice or face is only dangerous if the person on the other end skips the second check.

    • Set a callback rule. Any request to move money, change banking details, or reset credentials gets verified by calling a number pulled from your own records, never one supplied in the suspicious message or call.
    • Agree on a code word. A pre-shared phrase for high-stakes requests costs nothing and a real-time deepfake can’t guess it.
    • Slow down on urgency. Scammers manufacture time pressure because it stops people from verifying. Treat “this has to happen right now” as the red flag it is.
    • Train the one person who approves payments. If your business doesn’t have a finance team, whoever signs off on transfers is your entire defense layer. Make sure they know this playbook exists.
    Gartner had already predicted where this was heading: by 2026, the firm projected that 40% of enterprises would stop trusting standalone identity verification because of deepfakes. That prediction is landing now, and the fix it points to isn’t more software, it’s a second channel that a synthetic voice or face can’t fake its way through.

    Frequently Asked Questions

    What percentage of small businesses have experienced a deepfake scam?

    According to VikingCloud’s 2026 SMB Threat Landscape Report, 29% of small businesses reported experiencing a deepfake scheme in the past 12 months, making it one of the most common cyber incidents SMB owners now report, alongside data breaches and ransomware.

    How much money has been lost to deepfake and AI-enabled fraud in 2025?

    The FBI’s Internet Crime Complaint Center logged $893,346,472 in adjusted losses from 22,364 US complaints referencing AI in 2025, the first year the FBI tracked AI-enabled fraud as its own standalone category.

    How can a small business protect itself from deepfake scams?

    Require a second-channel verification, a callback to an internally stored phone number or a pre-agreed code word, for any request involving wire transfers, banking-detail changes, or credential resets, even ones that arrive by video call. It consistently ranks above detection software as the lowest-cost, most effective defense.

    Why are small businesses targeted by deepfake scammers more than large companies?

    Small businesses often rely on informal, trust-based approval processes with no dedicated IT or security staff. Eighty-four percent of SMB owners self-manage their own cybersecurity, per VikingCloud’s 2026 report, which removes the layered sign-off chain that would otherwise catch a fraudulent request.

    Can humans reliably spot a deepfake video?

    No. Controlled studies find human accuracy at identifying high-quality deepfake videos is only about 24.5%, even though roughly 60% of people believe they could spot one, a gap that itself increases risk by creating false confidence.


    Where This Goes Next

    Two things are converging right now that weren’t true even a year ago. The FBI has an audited number to point to for the first time, and small business owners are, for the first time in this survey series, ranking cyberattacks above the economy as their biggest worry. Neither of those happens without the other. Watch three things over the next six to eighteen months: whether EU AI Act enforcement actually produces fines large enough to change vendor behavior, whether cyber insurers start pricing deepfake-specific BEC into small business premiums, and whether the “29%” figure gets replicated by a source willing to publish full methodology.

    The takeaway for anyone running a small business isn’t to panic about AI. It’s to put a five-minute verification habit in place before you need it. The businesses in the Arup and WPP stories both had smart people on the call. Only one of them had a process that didn’t depend on trusting what they saw.

    Want more coverage like this before it hits the mainstream feeds? Subscribe to The Neural Loop at neuralwired.com/newsletter.

    More cybersecurity coverage: NeuralWired’s Cybersecurity hub. For related coverage on how deepfakes intersect with crypto fraud, see NeuralWired’s Crypto Regulation by Country 2026.

    Nearly 3 in 10 Small Businesses Hit by Deepfake Scams in 2026 Cybersecurity

    Nearly 3 in 10 Small Businesses Hit by Deepfake Scams in 2026

  • Microsoft, Amazon, Cisco Layoffs 2026: AI Capex Math

    Microsoft, Amazon, Cisco Layoffs 2026: AI Capex Math

    Big Tech Layoffs 2026: Why AI Capex Explains It All
    Big Tech

    Big Tech Layoffs 2026: Why AI Capex Explains It All

    Microsoft posted its best quarter ever and cut 4,800 jobs in the same three months. Amazon hit a record 13.1% operating margin and eliminated 30,000 corporate roles. Cisco broke its own revenue record and announced 4,000 layoffs the same week. None of that is a coincidence, and none of it is really about saving money on payroll either. It’s about where the money is actually going.

    Big tech layoffs in 2026 keep landing next to record earnings, and the pattern only makes sense once you put the two numbers side by side: what these companies are cutting from headcount, and what they’re pouring into AI infrastructure. The gap between those numbers is the story.

    The math nobody’s putting in the headline

    Start with Meta, because the comparison is cleanest there and it sets the pattern for everyone else. Meta’s 2026 capital expenditure guidance sits at $125 billion to $145 billion. Its entire human compensation bill, salaries, benefits, equity, all of it, runs around $27 billion. Even if Meta fired every single employee tomorrow, it wouldn’t cover a fifth of what it’s already committed to spend on AI infrastructure. That comparison comes from a Yahoo Finance analysis of company disclosures, and it’s the single most useful number in this entire story.

    Apply the same logic to Microsoft, Amazon, and Cisco and the picture holds. These aren’t companies trimming staff to fund a data center. They’re companies redirecting capital toward compute at a scale where headcount decisions barely register on the balance sheet.

    Why this matters: If layoffs were really about cost savings, the numbers would be close. They’re not. Payroll cuts save these companies low single-digit billions. AI capex commitments run into the hundreds of billions. Two completely different orders of magnitude, decided by two largely separate processes inside the same company.
    Company2026 AI CapexJobs CutSame-Quarter Result
    Microsoft~$190B (guided)4,800 (plus 9,100 prior round)Record $82.89B quarterly revenue
    Amazon~$200B (guided)~30,000 corporate rolesRecord 13.1% operating margin
    Cisco$9B AI orders (raised guidance)Fewer than 4,000 (~5%)Record $15.84B quarterly revenue
    Sources: Microsoft and Amazon Q1/Q3 2026 earnings disclosures; Cisco Q3 FY2026 earnings call.

    Microsoft: record revenue, Xbox gutted

    Microsoft’s fiscal Q3 2026 numbers, reported April 29, were about as strong as a quarter gets: $82.89 billion in revenue, up 18% year over year, with net income jumping to $31.78 billion from $25.82 billion a year earlier. The company also guided full calendar-year 2026 capex to roughly $190 billion, a 61% jump from 2025 and well past what Wall Street had modeled.

    The same quarter, Microsoft cut 4,800 jobs, most of them in the Xbox gaming division, on top of 9,100 roles eliminated about a year earlier. Chief people officer Amy Coleman told staff the cuts weren’t direct AI replacements, even while acknowledging AI is reshaping how the company runs. Microsoft also rolled out its first-ever voluntary buyout program, open to senior director level and below with enough age plus tenure to qualify.

    Here’s the part that undercuts the simplest version of the story: Xbox isn’t where Microsoft’s AI money is going. The division that got hit hardest wasn’t competing for capex dollars with Azure’s AI buildout in any direct sense. It just wasn’t the priority, and priority is what actually decides who keeps their job in 2026, not whether AI can technically do the work.

    Amazon: 30,000 gone, $200 billion committed

    Amazon’s Q1 2026 results, also reported April 29, delivered $181.5 billion in revenue and a record 13.1% operating margin, the highest in the company’s history. AWS grew 28% year over year to $37.6 billion, its fastest growth rate in 15 quarters. Amazon reiterated guidance toward roughly $200 billion in full-year 2026 capex.

    Against that backdrop, Amazon cut around 30,000 corporate jobs across rounds in October 2025 and January 2026, with further cuts hitting Selling Partner Services staff and a temporary Homestead, Florida warehouse closure eliminating 600-plus jobs between July and September.

    Unlike Microsoft and Cisco, Amazon’s leadership hasn’t tried to soften the connection. CEO Andy Jassy told staff in a memo, later reiterated into 2026, that generative AI and agents would reduce the company’s total corporate workforce over time as efficiency gains materialize, alongside creating new roles elsewhere. That memo dates to July 2025, a full year before this round of cuts, which makes it less a same-day justification and more a stated multi-year strategy Amazon is now executing on schedule.

    Cisco: “not a savings-driven restructure”

    Cisco reported record quarterly revenue of $15.84 billion on May 13, up 12% year over year, alongside AI infrastructure orders of $2.1 billion that quarter and $5.3 billion cumulative through three quarters. That pushed Cisco to raise its full-year AI order guidance from $5 billion to $9 billion, roughly 4.5 times fiscal 2025’s total.

    The same week, Cisco began notifying employees that it would cut fewer than 4,000 jobs, about 5% of its global headcount, with restructuring costs running as high as $1 billion, mostly severance. CFO Mark Patterson gave analysts a line worth sitting with:

    “This was really not a savings-driven restructure.” Mark Patterson, CFO, Cisco Systems, Q3 FY2026 earnings call, via Yahoo Finance
    Patterson framed the cuts as a realignment toward silicon, optics, security, and AI rather than a cost play. That’s a notably different posture from Amazon’s Jassy, and it matters: two companies profiled in the same story, cutting staff in the same season, and disagreeing with each other about whether AI is even the reason.

    Is AI actually the reason, or the excuse?

    Not everyone buys the AI-driven narrative, and the skepticism comes from serious places.

    Layoffs are often just standard cost-cutting with an AI label attached. Paraphrased position of Justin Wolfers, Professor of Economics and Public Policy, University of Michigan, via Benzinga/Finviz
    Wolfers argues AI functions as a convenient cover story for restructuring that companies would likely have pursued regardless. JPMorgan’s 2026 economic outlook backs that skepticism with data: the bank’s own labor-market analysis found the AI capex surge hasn’t shown much measurable impact on broader labor dynamics, despite the headlines.

    Wall Street’s bull case sees it differently. Wedbush’s Dan Ives, writing about Meta’s own 8,000-role cut against $135 billion in AI capex, called the layoffs financially minor next to the infrastructure commitment, damaging to morale but not decisive to the balance sheet. Evercore ISI’s Mark Mahaney goes further, noting this pattern of workforce actions followed by 12 to 18 months of margin expansion has repeated roughly every one to two years since 2022. In his read, this isn’t new behavior. It’s a recurring capital-discipline cycle that happens to be colliding with an AI narrative people want to believe.

    Our read: both things are probably true at once. AI capex is real, historically large, and reshaping where investment goes inside these companies. But the specific decision to cut a specific team often has more to do with which function sits outside this year’s priority list than with any AI system actually replacing a job. Xbox wasn’t cut because a model can ship games. It was cut because it wasn’t where the $190 billion was going.

    Worth flagging: Tracking firms don’t agree on the scale of 2026’s layoff wave. Layoffs.fyi-based counts put tech layoffs past 100,000 by early May and over 165,000 by July. SkillSyncer’s broader tracker counts 205,832 people affected across 322 events as of July 22, with 54% of those events explicitly citing AI or automation. The methodologies differ (tech-only versus all-industry, corporate versus contractor), so treat any single total as one tracker’s view, not a consensus figure.

    What this means if you work there

    If you’re an engineer or manager inside one of these companies, or one like them, the record-revenue headline is not protection. Internal budget decisions are increasingly decoupled from how well your specific team is performing. A well-run division can still get cut if it sits outside whatever core the company is funding this year, silicon, optics, security, and AI at Cisco; cloud and AI infrastructure at Microsoft and Amazon.

    The more useful signal than “is the company doing well” is “where is the capex actually going.” Read the earnings call transcript, not just the headline. That’s where you find out whether your function is this year’s priority or this year’s line item.

    There’s a real hedge here too. Industry-wide, roughly 275,000 AI-related roles are sitting open while laid-off tech workers largely can’t cross the skills gap to fill them, according to an Invezz analysis of labor-market data. Internal mobility toward AI or infrastructure teams is a legitimate near-term move, but it requires demonstrable fluency, not just tenure at the company.

    If you’re evaluating these companies as a vendor rather than an employer, the same logic applies from the other side. Support and account-management staff, the exact functions cut at Amazon’s Selling Partner Services, may thin even as infrastructure capacity grows. That’s worth a line in any vendor risk review.


    Frequently asked questions

    Why are Microsoft, Amazon, and Cisco laying off workers if their revenue is at record highs?
    Revenue and layoffs aren’t directly linked. Each company is redirecting tens of billions toward AI infrastructure capex, roughly $190 billion at Microsoft and $200 billion at Amazon for 2026, while separately restructuring specific divisions that sit outside their AI and cloud growth priorities.

    How much is Big Tech spending on AI infrastructure in 2026?
    Google, Amazon, Meta, and Microsoft combined are projected to spend roughly $725 billion on AI capital expenditure in 2026, up about 77% year over year, according to aggregated company guidance.

    Is AI actually causing the 2026 tech layoffs?
    It’s contested. Economists including Justin Wolfers argue AI often serves as a convenient explanation for ordinary cost-cutting. JPMorgan’s own analysis found little measurable labor-market impact from the AI capex surge, even as companies cite AI in layoff announcements.

    Did Cisco lay off workers despite good earnings?
    Yes. Cisco posted record Q3 FY2026 revenue of $15.84 billion, up 12% year over year, and in the same week announced plans to cut nearly 4,000 jobs as part of a restructuring its CFO described as not savings-driven.

    How many tech layoffs have there been in 2026?
    Estimates vary by tracker. Layoffs.fyi-based counts show over 165,000 tech layoffs by July 2026, while SkillSyncer’s broader tracker counts 205,832 people affected across 322 events as of July 22, with 54% of events citing AI as a factor.


    Where this goes next

    What’s changed by walking through all three companies together is this: the “AI is taking jobs” framing is too simple, and so is “it’s just normal cost-cutting.” What’s actually happening is a capital reallocation on a scale large enough that headcount decisions have become almost a separate conversation from infrastructure decisions, loosely connected at best, openly denied at Cisco, openly claimed at Amazon.

    Watch three things over the next 6 to 18 months. First, whether Microsoft’s own admission that it will remain capacity-constrained through 2026 even after this spending turns into visible AI revenue, or into a monetization lag that makes the capex look premature. Second, whether more executives start talking like Jassy (AI explicitly reducing headcount) instead of like Patterson (AI reorganizing headcount). Third, whether policymakers, following California’s move this June to build a state tracking tool for AI’s workforce impact, start requiring the kind of capex-versus-headcount disclosure that would make stories like this one unnecessary.

    None of these companies are lying when they post record revenue. None of them are lying when they cite AI in a restructuring memo either. They’re just optimizing for two different things at once, and reading the earnings call is currently the only way to tell which one is driving a specific decision.

    Want the next capex disclosure and layoff filing broken down like this one? Subscribe to The Neural Loop at neuralwired.com/newsletter.

  • Claude Opus 4.8 vs GPT-5.6: Best AI Coding Model 2026

    Claude Opus 4.8 vs GPT-5.6: Best AI Coding Model 2026

    Best AI Models for Agentic Coding Tasks 2026: 6 Tested on Cost-Per-Task
    Developer Focus

    Best AI Models for Agentic Coding Tasks in 2026

    Six frontier and workhorse models, tested against real cost-per-task benchmarks, not just leaderboard bragging rights.

    Your engineering team just spent $4.82 running Claude Opus 4.8 on a routine bug fix that a $0.07 model would have solved just as well. That’s not a hypothetical. It’s the real spread Artificial Analysis measured on its Coding Agent Index this year, and it’s the single most important fact in the best AI models for agentic coding tasks 2026 conversation right now. Model choice used to be about which one scored highest. In 2026, it’s about which one earns its price on the specific task in front of you.

    That shift didn’t happen quietly. Six weeks ago, one of the most capable coding models on the market vanished overnight because of a U.S. export control order, then came back three weeks later. Vendors quietly stopped reporting the benchmark everyone used to trust. And developers, according to a JetBrains-backed survey, now spend more hours reviewing AI-written code than writing it themselves. This piece walks through what’s actually true, what’s marketing, and which model belongs on which job.

    Why the old benchmarks stopped telling the truth

    For most of 2025, SWE-bench Verified was the number everyone quoted. Scores climbed from single digits to the high 80s and low 90s in under two years, a curve that looked like genuine progress until you asked the obvious question: how do models keep getting smarter at solving GitHub issues that were published years before their training cutoff?

    In February 2026, OpenAI’s own Frontier Evals team answered that question by walking away from the benchmark entirely. Their reasoning was blunt: model training had absorbed enough of the dataset that the score stopped measuring skill on unseen code and started measuring memorization. An independent audit of the top 30 leaderboard entries found that roughly 19.78% of cases labeled “solved” were passing unit tests by coincidence or by gaming the evaluation harness rather than by producing correct code.

    That’s why serious 2026 comparisons have moved to two newer references: SWE-bench Pro, built on private, professional repositories that no model has seen in training, and Terminal-Bench 2.1, which scores the model and its coding harness together as they complete a real terminal-driven task from start to finish. If a vendor is still leading its marketing with a SWE-bench Verified score above 90%, read it the way you’d read a car’s mileage sticker before the EPA got involved.

    The 2026 lineup, ranked

    Here’s where the six models actually land once you strip out the marketing and look at SWE-bench Pro and Terminal-Bench 2.1, the two benchmarks least contaminated by memorization.

    Model Vendor SWE-bench Pro Terminal-Bench 2.1 Pricing (input/output per MTok)
    GPT-5.6 “Sol”OpenAINot separately reported88.8% (highest recorded)Not disclosed at review time
    Claude Fable 5Anthropic80.3% (leader)83.1% (Claude Code)$10 / $50
    Claude Opus 4.8Anthropic69.2%78.9% (Claude Code)$5 / $25
    GPT-5.5OpenAI58.6%83.4% (with Codex)Not disclosed at review time
    Gemini 3.5 FlashGoogle DeepMind55.1%76.2%Not disclosed at review time
    Grok 4.5xAINot separately reportedNot separately reported$2 / $6
    Two things jump out. First, Claude Fable 5 leads the harder, contamination-resistant benchmark by a wide margin, 11 points ahead of Anthropic’s own Opus 4.8. Second, GPT-5.6 Sol leads the benchmark that best reflects how a coding agent behaves in an actual terminal, doing real multi-step work rather than generating a single patch. Neither model is the “best” one. They’re the best at different jobs.

    “But the improvement I keep coming back to is honesty.” Rahul Patil, CTO, Anthropic, on Claude Opus 4.8’s jump on SWE-bench Pro, via EdTech Innovation Hub
    Patil described the target workload for Opus 4.8 as the kind of job that “used to take a quarter and a working group,” meaning codebase-scale migrations and bug fixes spread across hundreds of files. That framing matters. It’s a tacit admission that raw benchmark points matter less than whether the model can survive a genuinely large, messy, real-world job without losing the thread.

    Where the open-weight tier fits in

    Not every team needs frontier pricing. GLM-5.2 from Z.ai, released under an MIT license, scores 62.1 on SWE-bench Pro and 81.0 on Terminal-Bench 2.1 at $1.40/$4.40 per million tokens, the strongest published open-weight coding numbers available right now. Kimi K2.6 Code lands an 80.2% SWE-bench Verified score at open-model pricing, close to Opus-class accuracy for a fraction of the bill. Neither will win a head-to-head against Fable 5 on the hardest tasks. Both will handle the routine 80% of your ticket queue for pennies.

    What each task actually costs you

    Here’s the number that should reshape how your team budgets for AI coding tools: on Artificial Analysis’s Coding Agent Index, an identical fixed task costs about $0.07 to run through Cursor’s Composer 2.5 and $4.82 to run through GPT-5.5, a roughly 60x spread for only three to four points of quality difference on the index.

    The mechanism most teams miss Coding agents burn most of their budget on reading, not writing. PointFive’s July 2026 index found that a single realistic task, reading a handful of files, reasoning about them, and revising a diff, can pull in up to 200,000 input tokens against a diff of roughly 30,000 tokens written back. That’s why input pricing, and whether a model caches repeated reads at a discount (often around 10% of the standard rate), swings your real bill far more than the headline output price per token.
    Zoom out and the frontier-to-workhorse spread gets even wider. Claude Fable 5 charges $10/$50 per million tokens. DeepSeek V4 Flash charges roughly $0.14. That’s close to a 180x difference in raw token pricing between the most expensive and cheapest options a team might reasonably put in production this year.

    None of this means cheap wins by default. GLM-4.6 costs about $0.059 per task but is statistically tied on accuracy with pricier open options, which means the math sometimes favors a marginally more expensive model like DeepSeek or Qwen instead. The lesson isn’t “buy the cheapest model.” It’s “stop assuming the most expensive model is the safest default,” and start routing tasks by difficulty: cheap model first, escalate to a frontier model only when the cheap one fails.

    The Fable 5 warning every team should have caught

    Claude Fable 5 launched June 9, 2026, and immediately topped the SWE-bench Pro leaderboard. Three days later, on June 12, 2026, Anthropic suspended it worldwide to comply with a U.S. Department of Commerce export control order. Access came back on July 1, 2026, after the controls were lifted, and Anthropic confirmed the restoration directly.

    Three weeks of downtime for a model teams were actively shipping in production. If your pipeline depended entirely on Fable 5 during that window, you didn’t have a benchmark problem. You had a supply chain problem, and most engineering leaders still aren’t tracking it as one.

    There’s a second wrinkle independent evaluators caught after Fable 5 came back online: Artificial Analysis and Vals AI both measured Fable 5 refusing roughly 8 to 9% of test prompts, quietly falling back to Opus 4.8 for those cases. That means the headline SWE-bench Pro score doesn’t fully describe what a production deployment experiences. A meaningful slice of real traffic never actually touches the model you thought you were paying for.

    Best practice going forward: never build a single-model dependency into a critical pipeline. Keep at least one fallback model configured, and treat a vendor’s top model the way you’d treat a single-region cloud deployment. It works great, right up until it doesn’t.

    The bottleneck nobody’s marketing deck mentions

    Every vendor above is racing to add benchmark points. Almost none of them are talking about the actual reason enterprise AI coding adoption stalls, and that’s reliability, not raw capability.

    “It unpacks different factors that I see tangled together in almost every eval I’ve ever seen.” Bryan Silverthorn, Director of AGI Autonomy, Amazon, at VB Transform 2026
    Silverthorn, who joined Amazon through its Adept AI acquisition, argues that “reliability” isn’t one thing. He breaks it into four separate dimensions, borrowing a framework from Princeton research: consistency, robustness, predictability, and safety. He described a customer whose agent performed a serial number extraction task flawlessly for two months, then quietly started misreading numbers with no warning and no obvious trigger. No benchmark on this list would have caught that failure mode before it hit production.

    Paul Gauthier, creator of the open source pair-programming tool Aider, has built a reputation on the opposite end of the spectrum: refusing to rank his own tool against agents that won’t publish their evaluation methodology. If a vendor won’t show its work, that’s a signal worth weighing as heavily as the score itself.

    There’s a human cost showing up in the data too. A developer survey compiling adoption research found that engineers using AI coding tools now spend 11.4 hours a week reviewing AI-generated code, against 9.8 hours writing new code themselves, a reversal from the pattern two years ago. The “10x productivity” pitch quietly assumes review time is free. It isn’t.

    How to actually choose, task by task

    Stop asking which model is smartest. Ask which model fits the task type sitting in your queue right now.

    • CLI-heavy DevOps and multi-step terminal work: GPT-5.6 Sol currently leads Terminal-Bench 2.1 at 88.8%, the strongest publicly reported score for real terminal-agent workflows.
    • The hardest multi-file repository repairs: Claude Fable 5 leads SWE-bench Pro, provided you’ve built in a fallback for its 8 to 9% refusal rate and you’re comfortable with the export control volatility above.
    • Large-scale migrations and refactors across hundreds of files: Claude Opus 4.8, purpose-built by Anthropic for exactly this workload, with its Dynamic Workflows feature fanning work out to parallel subagents.
    • Tool-orchestration-heavy agent work: Gemini 3.5 Flash leads MCP Atlas at 83.6% even though it trails on raw SWE-bench numbers, making it a genuine specialist pick for agentic tool-calling.
    • The routine 80% of your ticket queue: An open-weight model like GLM-5.2 or Kimi K2.6, or a workhorse like Cursor’s Composer 2.5, saves 10 to 60x on cost for a 3 to 4 point accuracy trade-off most teams won’t even notice.
    Our read: the real 2026 skill isn’t picking a single model and standardizing on it. It’s building a routing layer that sends each task to the cheapest model likely to solve it, and escalates only on failure. Teams still budgeting per seat instead of per completed task are leaving real money on the table, and the PointFive and Artificial Analysis data above shows exactly how much.

    Frequently asked questions

    What is the best AI model for coding in 2026?
    There’s no single winner. Claude Fable 5 leads the hardest contamination-resistant benchmark, SWE-bench Pro. GPT-5.6 Sol leads real terminal-agent work, scoring 88.8% on Terminal-Bench 2.1. The right choice depends on task type and budget, and open-weight models like GLM-5.2 close most of the gap at a fraction of the cost.

    How much does an AI coding agent cost per task?
    Cost per completed coding task ranges from roughly $0.07 to $4.82 depending on the model, according to Artificial Analysis and PointFive benchmark data. Workhorse models like Cursor’s Composer 2.5 cost around $0.07 per task, while frontier models like GPT-5.5 or Claude Opus can run $4 or more for only a few extra benchmark points.

    Why did OpenAI stop reporting SWE-bench Verified scores?
    OpenAI’s Frontier Evals team announced in February 2026 that it would stop reporting SWE-bench Verified results because training data contamination had inflated scores past the point where they reflected real coding ability on unseen code. SWE-bench Pro, built on private repositories, is now the more trusted reference.

    Is Claude Fable 5 still available?
    Yes. Claude Fable 5 launched June 9, 2026, was suspended worldwide on June 12, 2026 under a U.S. Department of Commerce export control order, and access was restored on July 1, 2026 after the controls were lifted. Teams building on it should keep a fallback model plan in place given that volatility.


    Where this goes next

    The benchmark story of 2026 is really a trust story. Vendors spent two years optimizing for a number that eventually stopped meaning anything, and the market is only now rebuilding around harder, more honest measures like SWE-bench Pro and Terminal-Bench 2.1. Cost-per-task, not leaderboard rank, is fast becoming the metric that actually determines what ships to production.

    Three things worth watching over the next six to eighteen months: whether Anthropic can keep Fable 5 and Mythos 5 available without another export control disruption, whether the 60x cost gap between frontier and workhorse models narrows as competition in the open-weight tier intensifies, and whether reliability metrics like Bryan Silverthorn’s four-part framework get standardized into a benchmark of their own. Gartner’s projection that 40% of new enterprise production software will involve vibe coding by 2028 is a forecast, not a fact on the ground today, and it deserves the same skepticism this piece just applied to SWE-bench Verified.

    Want the next model launch, export control ruling, and cost benchmark broken down the same way? Subscribe to The Neural Loop at neuralwired.com/newsletter.

  • vLLM vs SGLang vs TensorRT-LLM: 2026 Benchmark Guide

    vLLM vs SGLang vs TensorRT-LLM: 2026 Benchmark Guide

    Best LLM Inference Optimization Tools 2026: 7 Engines Tested
    Infrastructure / Developer Tools

    vLLM vs SGLang vs TensorRT-LLM: 7 Inference Engines Tested Against MLPerf v6.0

  • FATF: Crypto Crime Hits Record as Bitcoin Crashes 2026

    FATF: Crypto Crime Hits Record as Bitcoin Crashes 2026

    Crypto Crime Hits Record as Market Crashes 52%
    Cybersecurity & Regulation

    Crypto Crime Hits Record as Market Crashes 52%

  • JPMorgan Kinexys Hits $4 Trillion: Blockchain 2026

    JPMorgan Kinexys Hits $4 Trillion: Blockchain 2026

    64% of Institutions Are Now Tokenizing Assets: Inside the 2026 Enterprise Blockchain Market
    Enterprise Blockchain · Market Analysis

    64% of Institutions Are Now Tokenizing Assets: Inside the 2026 Enterprise Blockchain Market

    JPMorgan’s Kinexys platform just crossed $4 trillion in cumulative volume. The Federal Reserve says tokenized assets doubled in a year. But the “market size” numbers everyone’s citing don’t agree with each other, and one popular stat about institutional adoption is being misquoted across the web.

  • Tether Controls 58% of the $300B Stablecoin Market 2026

    Tether Controls 58% of the $300B Stablecoin Market 2026

    Stablecoins Hit $300B — Tether Controls 58%. Who’s Fighting for the Rest?
    Crypto & Markets

    Stablecoins Hit $300B, Tether Owns 58%. Who’s Fighting for the Rest?

  • Pinecone Says RAG Is Obsolete: Complete 2026 Verdict

    Pinecone Says RAG Is Obsolete: Complete 2026 Verdict

    Pinecone Bets RAG Is Obsolete. The Data Disagrees
    AI Infrastructure

    Pinecone Bets RAG Is Obsolete. The Data Disagrees

    The company that made retrieval-augmented generation a household term just told its own 800,000 developers to stop doing it. Here is what that means if you are choosing between RAG and a 2 million token context window in 2026.

    Your engineering team spent 2024 building a retrieval pipeline. Chunk the docs, embed them, store them in a vector database, retrieve the top matches, stuff them into a prompt. It worked, mostly. Then Gemini shipped a 2 million token context window, Claude and GPT-5.4 hit 1 million, and someone on Slack asked the question everyone is now asking: why not just paste the whole knowledge base in and skip the plumbing?

    That question has a real answer now, and it is not the one either side of the debate wants. A 2 million token context window does not replace retrieval-augmented generation. It changes what retrieval is for. And the company that spent four years teaching the industry how to build RAG vs long context pipelines just told the market, in public, that the pattern it popularized is already the bottleneck.

    The context window race just hit a new ceiling

    By April 2026, five frontier labs had all crossed the same line. Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro, Qwen 3.6 Plus, and Llama 4 Maverick each shipped a 1 million token context window. Meta pushed further with Llama 4 Scout, advertising 10 million tokens, though independent testers found its usable recall breaks down well short of that number. Google’s Gemini line has sat at the 2 million token mark since early 2026, which is why “2 million token context window” is now the phrase enterprise buyers type into Google before they type anything else.

    By June 9, at least 13 models had crossed the 1 million token line, according to a pricing comparison from Morph. What that comparison also revealed is that “1 million tokens” is not one product. It is thirteen different products with wildly different economics.

    ModelCost to fill a 1M-token context window
    DeepSeek V4 Flash$0.14
    Claude Fable 5$10.00
    Source: Morph, June 9, 2026. A 71x spread across the field.

    That 71x spread is the first sign that “just use a bigger window” is not a strategy. It is a pricing decision you have not made yet.

    Context rot: why bigger windows are not always better

    In July 2025, three researchers at the vector database company Chroma published a report that has become the most-cited technical pushback on long-context marketing copy. Kelly Hong, Anton Troynikov, and Jeff Huber tested 18 frontier models, including the GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 families, on tasks specifically designed to hold difficulty constant while varying only input length.

    The finding that should worry anyone planning to dump a full knowledge base into a prompt: every single model got less reliable as the input got longer, even on tasks a human would call trivial. And in a twist that inverts a common assumption among RAG engineers, models performed worse on well-organized, logically coherent source documents than on the same content shuffled into random order.

    “Models do not use their context uniformly.” Kelly Hong, Anton Troynikov, and Jeff Huber, Chroma Research, “Context Rot: How Increasing Input Tokens Impacts LLM Performance,” July 2025
    Chroma is a retrieval infrastructure vendor, so this finding is also commercially convenient for the company publishing it. That is worth disclosing. It does not make the methodology wrong. The 18-model benchmark is open source and independently reproducible, and it lines up with a separate, older finding known as “lost in the middle”: accuracy drops 20 to 30 percentage points when the answer sits in the middle of a long document instead of at the start or end, a pattern first documented by Liu et al. and replicated across model families since.

    Put together, these results point to a rule NVIDIA’s own RULER benchmark backs up: the effective, reliable portion of a context window typically runs at 50 to 65% of the number on the marketing page. Some of Chroma’s own findings suggest the real, safe margin for production workloads is tighter still, closer to a quarter or a third of the advertised maximum.

    The real cost of going long

    Even where accuracy holds up, long-context prompting is not cheap next to modern retrieval. A 2026 arXiv study titled “Long Context vs. RAG for LLMs” ran a direct cost comparison across GPT-5.4-mini and nano on document-grounded question answering. The result: long-context prompting averaged roughly $0.1181 per query, against $0.0045 to $0.0046 for keyword or semantic retrieval. That is a 10x-plus cost gap, and it is the more conservative of the figures floating around; some blog posts cite gaps as high as 1,250x, but those appear to compare different cost baselines and should be treated with skepticism.

    Anthropic’s prompt caching cuts input costs by up to 90% and latency by up to 85% on repeated long prompts, which matters more than raw context size for most production bills. The lesson is not “context is expensive.” It is that caching, batching, and retrieval scope are the real levers, and a bigger window without any of those disciplines is the most expensive way to solve the problem.

    Worth flagging for enterprise architects: access to any single frontier model is not guaranteed to be stable. In June 2026, Anthropic temporarily suspended access to Claude Fable 5 and Mythos 5 to comply with U.S. Department of Commerce export controls, restoring it on July 1 after the controls were lifted (Anthropic’s statement). Whatever architecture you pick, model availability is now a variable you plan around, not an assumption you make.

    Pinecone just bet against the category it built

    On May 4, Pinecone, the vector database that made RAG a standard pattern for roughly 800,000 developers and 9,000 paying customers, launched Nexus, which it calls a “knowledge engine for agents,” alongside KnowQL, a query language built around six primitives: intent, filter, provenance, output shape, confidence, and latency budget.

    Pinecone’s own framing is blunt. It describes retrieval-at-inference, the classic chunk-and-embed pattern the company spent four years teaching the market, as the “ten blue links era of agentic retrieval.” Its argument: agents stuck in retrieve-read-retrieve loops complete only 50 to 60% of tasks and burn 85% of their effort just fetching context, before any actual reasoning happens.

    Instead of retrieving raw chunks at query time, Nexus precompiles source data into structured, cited, task-specific artifacts ahead of time, so an agent queries a compiled answer rather than a pile of documents. Harrison Chase, the CEO of LangChain and the person widely credited with popularizing the term “context engineering,” backed the framing on Pinecone’s own launch post.

    “Building reliable, long-horizon agents is fundamentally a context engineering problem.” Harrison Chase, CEO, LangChain, on Pinecone’s Nexus launch post, May 4, 2026
    Janakiram MSV, the cloud and AI analyst who covers infrastructure shifts for The New Stack, called out just how unusual this is. Most vendors keep selling into a category long after the market has moved past it. Pinecone named the shift itself.

    “Pinecone just declared the RAG era over.” Janakiram MSV, The New Stack, “The company that made RAG mainstream is now betting against it,” May 6, 2026
    Our read: MSV’s framing is closer to right than Pinecone’s own marketing copy. This is not “RAG is dead.” It is RAG’s naive, retrieve-then-hope form getting replaced by something more deliberate, the same shift Anthropic’s Skills and Cursor’s project rules are pushing at the editor and agent-framework layer. The pattern is not new. The vendor saying it out loud is.

    The Subquadratic wildcard: 12 million tokens, unverified

    One day after Pinecone’s launch, Miami-based startup Subquadratic emerged from stealth with $29 million in seed funding and a model called SubQ, built on what it calls a Subquadratic Selective Attention architecture. Founded by CEO Justin Dangel and CTO Alexander Whedon, both veterans of Meta, the company claims SubQ’s research version supports a 12 million token context window, roughly 120 books, while scaling compute linearly rather than quadratically with input length.

    The headline number, as reported by SiliconANGLE: SubQ scored 95% on the RULER 128K benchmark at about $8 in compute, against 94% accuracy and roughly $2,600 for Claude Opus on the same test, a claimed 300x cost reduction. Backers reportedly include Tinder co-founder Justin Mateen and early investors in Anthropic, OpenAI, Stripe, and Brex.

    Treat every one of those numbers as “reported by Subquadratic” until someone outside the company replicates them. As of this writing, no independent benchmarking team has confirmed the 52x attention speedup, the 92.1% needle-in-haystack recall at 12 million tokens, or the roughly 1,000x compute reduction the company claims at full context length. If verified, it would be the largest single jump in usable context the field has seen. If not, it joins a long list of long-context claims that looked revolutionary on launch day and ordinary six months later.

    So is RAG dead? The growth data says no

    Here is the part the “RAG is dead” headlines tend to skip: RAG-adjacent infrastructure spending is still growing fast, and growth data does not lie the way marketing copy can. Market-sizing firms disagree sharply on the exact dollar figures. Grand View Research puts the market at $1.2 billion in 2024, growing to $11 billion by 2030 at a 49.1% compound annual growth rate. Precedence Research estimates $2.76 billion in 2026 climbing to $67.42 billion by 2034. MarketsandMarkets lands in between, at $1.94 billion in 2025 growing to $9.86 billion by 2030. Cite one firm at a time, since the numbers do not reconcile with each other, but the direction across all three is the same: a technology genuinely on its way out does not post 38 to 49% annual growth.

    Production engineers writing on DEV Community made the practical case bluntly: no context window, however large, holds an enterprise knowledge base running to millions of documents. A single 1 million token Claude Sonnet-class prompt runs roughly $3 at list pricing, and that does not scale to production query volumes the way retrieval does. Their position is that RAG’s continued growth is itself the strongest evidence against the “dead technology” framing, not despite the long-context hype but because of what enterprises are actually shipping underneath it.

    What this means for your stack

    Stop treating this as RAG versus long context. Treat it as a context budget you have to manage regardless of which technique you use.

    • Cap your assumptions at 25 to 30% of the advertised window. That is roughly what Chroma’s own findings suggest is the safe, reliable slice of any long-context claim, sticker number aside.
    • Pair retrieval with compaction. For long agent sessions, summarization and compaction loops matter more than raw window size, because irrelevant content is what causes context rot, not length alone.
    • Do not rip out retrieval infrastructure on the assumption long context replaces it. Teams that did this in 2024 and 2025 are the ones now eating the 10x-plus cost premium documented above.
    • Watch where vendor R&D is actually pointed, not where the marketing copy points. Pinecone’s own pivot from raw retrieval toward precompiled, agent-queryable artifacts is a better signal than any single benchmark chart.
    • Evaluate new entrants before migrating production workloads. Subquadratic’s numbers are compelling on paper and unverified in practice. Run your own evals on your own data first.
    One more thing regulated industries should not skip: RAG’s retrieval logs double as an audit trail. Raw long-context prompting does not produce one by default. In finance, healthcare, or legal workflows, that gap is not academic. It is a compliance requirement waiting to surface during an audit, usually at the worst possible time.


    Frequently asked questions

    Does a bigger context window replace RAG?

    Rarely. Long context reduces the need for aggressive retrieval on smaller, bounded corpora, but no window, even 12 million tokens, holds an enterprise knowledge base with millions of documents. Long-context prompting also runs roughly 10x or more expensive per query than modern retrieval in controlled 2026 benchmarks.

    What is “context rot”?

    Context rot is measurable performance degradation as an LLM’s input length grows, even on simple tasks. Chroma Research tested 18 frontier models in 2025 and found every one degraded with length, with logically coherent documents sometimes hurting performance more than shuffled ones.

    What causes the “lost in the middle” problem?

    Models attend most reliably to information at the very start and end of their context window. Liu et al.’s benchmark found accuracy drops 20 to 30 percentage points when the answer sits mid-context, a pattern replicated across GPT, Claude, and other model families since.

    How much does a 1 million token prompt cost?

    It depends heavily on the model. As of June 2026, filling a 1 million token window ranges from about $0.14 on DeepSeek V4 Flash to $10.00 on Claude Fable 5, a 71x spread, before caching discounts are factored in.

    Is RAG still worth building in 2026?

    Yes, for most production systems with large, dynamic, or compliance-sensitive corpora. RAG-related infrastructure spend kept growing at 38 to 49% CAGR across multiple market estimates even as long-context windows expanded, and 2026 is shaping up to be a hybrid-architecture year rather than a winner-take-all contest.


    The bottom line

    Nothing here says long context is a bad bet or that RAG is finished. What the evidence actually supports is narrower and more useful: raw context length is not the same thing as usable context, cost scales against you faster than accuracy does, and the vendor that built the RAG category is now telling the market to build the next layer up, not to abandon retrieval altogether.

    Watch three things over the next 6 to 18 months. First, whether independent labs confirm any of Subquadratic’s numbers, since that would be the first real architectural break from quadratic attention costs. Second, whether Pinecone’s Nexus and KnowQL numbers hold up in production the way they did in Pinecone’s own benchmarks. Third, whether “context engineering,” the discipline of deliberately curating what enters a model’s window regardless of technique, becomes a formal job function the way “prompt engineering” did in 2023.

    The teams that win this cycle will not be the ones who pick a side in the RAG-versus-context debate. They will be the ones who stopped treating context size as a proxy for context quality months before everyone else did.

    Want the next infrastructure shift in your inbox before it hits the front page? Subscribe to The Neural Loop at neuralwired.com/newsletter.

  • Google & EU AI Act: New Ad Disclosure Rules for 2026

    Google & EU AI Act: New Ad Disclosure Rules for 2026

    AI Generated Content Disclosure Rules 2026: The August 2 Deadline Marketers Can’t Miss
    Policies

    AI Ad Disclosure Rules 2026: The August 2 Deadline That Hits Meta, Google, the EU, California and New York at Once

  • Colorado AI Act SB 26-189: What Employers Must Know

    Colorado AI Act SB 26-189: What Employers Must Know

    Colorado AI Act SB 26-189: What Employers Must Do by 2027
    AI Regulation · Employment Law

    Colorado’s AI Law Died Before It Lived. Here’s What’s Next