Gemini 3.7 Flash: Half Price Now, Full Price in 2027AI & Enterprise Tech
Gemini 3.7 Flash Is Half Price. Read the Footnote First.
By NeuralWired Staff | Published August 15, 2026
Google DeepMind shipped Gemini 3.7 Flash on August 13, 2026, its fourth Flash-tier model in nine weeks, at half the price of its predecessor. That discount expires December 31, 2026. And the model everyone actually asked for at I/O in May, Gemini 3.5 Pro, still hasn’t shipped.
If you’re choosing a model for coding agents or budgeting inference spend into 2027, both of those facts matter more than the launch headline. Here’s what Google’s own numbers say, what independent testing confirms, and what the pricing footnote is quietly telling you.
Gemini 3.7 Flash went generally available in the Gemini API, Google AI Studio, Vertex AI, and Antigravity, Google’s coding-agent platform, on August 13, 2026, according to Google’s own Gemini API release notes. It also now powers Gemini Spark, Google’s productivity agent.
That’s four Flash-tier releases since Gemini 3.5 Flash debuted at I/O in May: 3.5 Flash, then 3.5 Flash-Lite and 3.6 Flash together on July 21, then 3.7 Flash on August 13. Twenty-three days between the last two. Google says the speed comes from algorithmic improvements, not a bigger base model.
Google AI Studio product lead Logan Kilpatrick framed it as a fast, targeted push rather than a ground-up rebuild.
“A strong intelligence increase, delivered in roughly three weeks through algorithmic work across Google DeepMind teams, focused on making the model feel more usable for real work.”
Logan Kilpatrick, Product Lead, Google AI Studio, Google DeepMind, August 2026 launch announcement
Google’s own positioning line calls it “our most intelligent workhorse model yet for coding and agents.” Positioning aside, the two numbers worth caring about are what it costs and what it can actually do, and the answers to both come with asterisks.
The Pricing Trap: Half Price Until January 1
Gemini 3.7 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens. That’s exactly half of what Gemini 3.6 Flash charges. It’s a genuinely good rate, and it’s live right now.
The part most launch-day coverage skipped:
This pricing runs through December 31, 2026 only. On January 1, 2027, Gemini 3.7 Flash reverts to $1.50 input / $7.50 output per million tokens, the identical permanent rate Gemini 3.6 Flash has charged since its own July launch. The “half price” headline is a five-month promotional window, not a durable cost advantage.
If your team is modeling 2027 inference spend on today’s rate card, that model is wrong by roughly 2x. Build your cost projections around $1.50/$7.50, not $0.75/$3.75, for anything shipping past year-end.
There’s a second cost change buried in the same release: Google removed the “minimal” thinking tier, the cheap setting older Flash models used for high-volume classification work. “Low” is now the floor, and thinking tokens bill at the output rate even though the API only returns a summary of that reasoning. If your pipeline leaned on minimal-tier Flash for bulk, low-stakes calls, re-benchmark it. The effective cost floor just moved up even as the headline price moved down.
Gartner analyst Will Sommer flagged exactly this pattern months before this launch, and it applies directly here.
“Chief Product Officers should not confuse the deflation of commodity tokens with the democratization of frontier reasoning.”
Will Sommer, Senior Director Analyst, Gartner, LLM inference economics forecast, March 2026
Sommer’s point is sharper once you factor in agentic workflows, which is exactly what Google is tuning 3.7 Flash for. Agent loops can multiply token consumption 5 to 30 times per task compared to a single chat completion. Cheaper tokens don’t necessarily mean a cheaper bill when the model is calling itself in a loop.
What Google’s Own Benchmarks Really Show
The capability jump from 3.6 Flash to 3.7 Flash is real, per Google’s own evals methodology page. DeepSWE v1.1 climbed from 49.0% to 65.3%. AutomationBench nearly doubled, from 17.0% to 30.4%.
Independent testing backs up at least the speed claim. Artificial Analysis clocked 3.7 Flash at 340.1 tokens per second in its standardized benchmarking workload, the fastest model the firm currently measures.
But read Google’s own head-to-head comparison table against GPT-5.6 Terra in full, not cherry-picked, and the frontier-leadership framing softens fast.
Notice which four rows it loses: the hardest agentic and terminal-use benchmarks, the exact category Google is marketing this model for. It’s a real value trade-off, not a clean win, and it’s Google’s own chart saying so.
There’s also a small but telling inconsistency worth flagging. Google’s 3.6 Flash model card lists its own DeepSWE v1.1 score as 48.6%. The 3.7 Flash launch blog rounds that same baseline to 49.0%. Minor on its own, but it’s a reminder that even single-vendor self-reported numbers are worth cross-checking against the vendor’s other documents, not just against competitors.
METR, the group that runs independent AI capability evaluations, has warned about a broader version of this problem.
“Benchmarks run without live human interaction can cause models to fail at tasks they could complete with minimal human guidance, making benchmarks unreliable proxies for real capability.”
METR, Experienced Developer Study, July 2025
Translation for anyone building on this: treat DeepSWE and AutomationBench jumps as lab signals worth investigating, not as production-readiness guarantees. Run your own workload against it before you migrate.
The Elephant in the Room: Gemini 3.5 Pro
None of the Flash-tier sprint makes sense without the model that isn’t here. At I/O in May, Sundar Pichai told developers to give Google “until next month” for Gemini 3.5 Pro, implying a June release. It didn’t happen. As of this article’s publication, it still hasn’t.
Bloomberg reported on July 16, citing ten current and former Google employees, that 3.5 Pro was running months behind schedule, largely over coding-capability shortfalls. A late-June training-data update meant to fix that reportedly made results worse, not better. Later reporting sourced to the same chain indicates the problems ran deeper than a bad update: DeepMind concluded the original 3.5 Pro base model had structural failures in recursive tool-calling and SVG generation, scrapped it, and restarted pretraining from a native Gemini 3 foundation. That same reporting says DeepMind has already begun pretraining an entirely new flagship, Gemini 4, mentioned almost in passing in the July 21 announcement.
Pichai himself gave the first public crack in the story, back in May.
“A bit behind on agentic coding.”
Sundar Pichai, CEO, Alphabet/Google, remarks at Google I/O, May 2026
Kilpatrick’s current line on 3.5 Pro is that the team is “testing with partners” and hopes to “land it soon.” That “soon” has now stretched past a second informal window with no date attached.
Wall Street has already priced in the uncertainty. Alphabet shares fell roughly 4.4% the day the Bloomberg delay report landed, an estimated $200 billion in market cap, on top of an earlier ~$225 billion drop in June tied to senior DeepMind researchers leaving for Anthropic and OpenAI. Combined, that’s close to $425 billion in Alphabet market value lost since late June with no change to reported revenue or earnings. Alphabet’s Q1 2026 results were strong (Google Cloud revenue up 63% year over year to $20 billion), which makes the point sharper: this is a narrative problem right now, not yet a fundamentals problem.
Our read: shipping four Flash models in nine weeks while the flagship reasoning tier stalls out looks less like a coincidence and more like a deliberate holding pattern, cover the volume segment on cost and speed while the harder model gets rebuilt underneath it. Google hasn’t confirmed that as strategy though, and it’s worth treating that framing as the most defensible inference from public facts, not as a confirmed internal decision. It could just as easily be ordinary engineering triage under deadline pressure.
What EU and UK Teams Need to Know
Buried in the model card, not the launch announcement, is a jurisdictional exclusion: Gemini 3.7 Flash is not available on the only consumer-facing surface it runs on in the EEA, UK, Switzerland, and Nigeria.
The timing isn’t nothing. The European Commission’s enforcement powers over general-purpose AI providers under the EU AI Act activated on August 2, 2026, penalties up to €15 million or 3% of global annual turnover, whichever is greater. Eleven days later, Google’s newest consumer AI model quietly excludes those exact jurisdictions from that surface. Google hasn’t stated a causal link publicly, but if you’re evaluating this model for an EU-facing product, plan around the exclusion now rather than discovering it in deployment.
The Bottom Line for Engineering Teams
If you’re on Gemini 3.6 Flash today, this is a real upgrade at a genuinely good price, for now. Three things to actually do with that:
Model your 2027 costs at $1.50/$7.50, not $0.75/$3.75. The current rate expires December 31, 2026.
Re-benchmark anything that ran on the old “minimal” thinking tier. It’s gone, and “low” now bills thinking tokens at the output rate.
Don’t lock a roadmap to a Gemini Pro milestone right now. Google has shipped zero Pro-tier models since Gemini 3 Pro in November 2025, despite promising 3.5 Pro for June 2026.
The broader question is whether this efficiency pivot holds. If Gemini 4, reportedly already in early pretraining, also slips, Google will have gone potentially 18 months or more between flagship releases while Anthropic and OpenAI keep a faster cadence. That’s a gap that compounds on reputation even if Flash-tier usage and revenue stay healthy in the meantime. Our related coverage on why enterprise AI inference costs aren’t actually falling and on Google’s agentic AI enterprise adoption gap both dig further into the pieces of this story we didn’t have room for here.
FAQ: Gemini 3.7 Flash and Gemini 3.5 Pro
Is Gemini 3.7 Flash actually cheaper than Gemini 3.6 Flash?
Only through December 31, 2026. Gemini 3.7 Flash launched at $0.75/$3.75 per million tokens, half of 3.6 Flash’s rate, but on January 1, 2027 it reverts to $1.50/$7.50, the same permanent rate 3.6 Flash has charged since July 2026.
When is Gemini 3.5 Pro coming out?
No confirmed date. Google promised it for June 2026 at I/O, but Bloomberg reported in July that coding-performance issues forced a delay, and later reporting indicates Google scrapped the original base model and restarted pretraining. As of August 15, 2026, it remains unreleased.
Is Gemini 3.7 Flash better than GPT-5.6 Terra for coding?
It’s close, not a clear win. On Google’s own 13-row comparison table, 3.7 Flash wins 7 rows but loses on the hardest agentic and terminal-use benchmarks to GPT-5.6 Terra, which costs over three times as much per token.
Why is Gemini 3.7 Flash not available in the EU or UK?
Google’s model card excludes the EEA, UK, Switzerland, and Nigeria from the consumer surface the model runs on. The exclusion lands 11 days after the EU AI Act’s enforcement powers over general-purpose AI providers activated on August 2, 2026.
What happened to Gemini Flash’s “minimal” thinking mode?
Gemini 3.7 Flash removed the “minimal” thinking tier used for cheap, high-volume classification tasks. “Low” is now the cheapest tier, and thinking tokens bill at the output rate even though only a summary is returned, raising the effective cost floor for simple workloads.
Gemini 3.7 Flash is a real, well-priced upgrade for teams already on the Flash tier, for the next four and a half months. What it isn’t is a replacement for the flagship model Google promised in May and still hasn’t shipped. Watch three things over the next 6 to 18 months: whether Gemini 3.5 Pro actually lands, whether the January price reset changes adoption at all, and whether Gemini 4’s pretraining run stays on schedule.
Want the next update the moment Gemini 3.5 Pro ships, or when the pricing resets in January? Subscribe to The Neural Loop at neuralwired.com/newsletter.
LiteLLM Breach 2026: Why Your SDLC Checklist Failed
Cybersecurity
LiteLLM Breach 2026: Why Your SDLC Checklist Failed
Published August 14, 2026 | NeuralWired Cybersecurity Desk
One credential from February didn’t get rotated. Five months later, that single oversight had cascaded through a vulnerability scanner, a code analysis tool, and an AI gateway used by thousands of companies, exposing an estimated 2,500 organizations and roughly 434,000 CI/CD pipelines. If your team runs LiteLLM, Trivy, or Checkmarx KICS anywhere in its build process, this story isn’t background reading. It’s an open incident.
Two threat intelligence firms independently confirmed the scale of the damage this week. On August 11, 2026, CloudSEK published its exposure dataset. Two days later, Hudson Rock corroborated it from a completely separate 153GB archive. Neither firm was working from the other’s data. That’s what makes this LiteLLM breach different from the usual single-source security scare: the numbers hold up.
LiteLLM is a popular open-source gateway that lets developers call dozens of large language model APIs through one unified interface. It sits in front of, or alongside, a huge number of production AI workloads. That’s exactly why the FBI’s Internet Crime Complaint Center formally named the threat group behind this campaign: TeamPCP, in a July 2, 2026 advisory that confirmed Trivy, Checkmarx KICS, LiteLLM, and the Telnyx Python SDK as compromised links in one escalating campaign.
The breach itself happened back in March. The public reckoning is happening now, in real time, which is why this is the story to understand this week rather than next month.
The Attack Chain: One Credential, Three Tools, Thousands of Companies
Strip away the acronyms and the sequence is almost mundane, which is what makes it unsettling.
A credential from a late-February 2026 breach never got fully rotated. TeamPCP used it to hijack the service account behind Aqua Security’s Trivy vulnerability scanner.
March 19, 2026: the group force-pushed malicious code across 76 of the 77 version tags in the aquasecurity/trivy-action GitHub repository.
Two days later: Checkmarx’s KICS scanner was compromised using stolen GitHub tokens, extending the campaign to a second widely used security tool.
LiteLLM’s own CI pipeline auto-installed the compromised Trivy version, and two malicious LiteLLM releases, versions 1.82.7 and 1.82.8, went live on PyPI.
Forty minutes doesn’t sound like much until you understand what version 1.82.8 actually shipped: a file called litellm_init.pth that executes automatically the moment Python starts up. Teams that thought running --ignore-scripts protected them were wrong. That flag blocks install-time scripts. It does nothing against a file designed to fire on interpreter startup, which is the detail that should worry anyone who assumed a single defensive habit was sufficient.
“Trivy, then the build system, then the release: one unrotated token, three tools deep. That chain is what turns a single credential leak into ecosystem-wide exposure.”
CloudSEK, via SecurityWeek, August 12, 2026
By the Numbers: Third-Party Breaches Are Accelerating
The LiteLLM breach isn’t a one-off. It’s the loudest recent data point in a trend that’s been building for two years. Here’s what the most credible sources actually say, since the headline stats floating around social media don’t all agree.
Source
Figure
What it measures
Verizon 2025 DBIR
30% of breaches, double the 15% a year earlier
Confirmed breaches with third-party involvement, across 12,195 incidents globally
SecurityScorecard / HIPAA Journal
35.5% in 2024, up from 29% in 2023
Breaches that originated from a third-party compromise
IBM Cost of a Data Breach 2025
30%, described as doubling year over year
Corroborates Verizon’s directional finding
SecurityScorecard / Secureframe
75% of third-party breaches
Specifically hit the software and technology supply chain
Which number should you actually cite?
A widely repeated “29% of breaches start with a third party” figure is outdated. It’s SecurityScorecard’s 2023 baseline, and it climbed to 35.5% by 2024. If you need one number to anchor a board conversation or a budget request, use Verizon’s 30%, doubled from 15% the prior year, drawn from the largest DBIR dataset on record. It’s the most methodologically transparent figure in the industry right now.
Sonatype’s 2026 State of the Software Supply Chain report adds scale to the picture: 1.233 million malicious open source packages have now been identified, with open source malware up 75% year over year and 454,648 new malicious packages found in the past twelve months alone, based on analysis of more than 10 trillion downloads across Maven Central, PyPI, npm, and NuGet. And 86% of Maven Central traffic in 2025 came from cloud service providers rather than humans, which tells you something important: the attack surface has moved from developers clicking “install” to automated build systems pulling dependencies at machine speed, unsupervised, thousands of times a day.
This Isn’t Isolated: The Shai-Hulud npm Worm Wave
If LiteLLM feels like an isolated AI-ecosystem incident, it isn’t. It’s the PyPI chapter of a story that’s been unfolding in npm for almost a year.
September 2025: “Shai-Hulud,” the first documented self-replicating npm worm, compromised more than 500 packages, according to a CISA advisory.
November 24, 2025: “Shai-Hulud 2.0” backdoored 796 unique npm packages representing over 20 million weekly downloads, per Datadog Security Labs. It self-replicates without needing a command-and-control connection back to the attacker.
March 2026: a related campaign, tracked by StepSecurity and CloudSEK, exfiltrated 78,330 secrets from CI/CD pipelines across 2,186 organizations in five days.
April 2026: a “Shai-Hulud: The Third Coming” variant compromised the official @bitwarden/cli package, which had more than 250,000 monthly downloads, through a malicious preinstall hook.
Between August 2025 and May 2026, npm went from occasionally hosting malware to becoming one of the most actively exploited software supply chains anywhere. A maintainer-phishing wave briefly poisoned a combined 2.6 billion weekly downloads across the chalk and debug packages alone. The pattern connecting npm’s worm wave to the LiteLLM breach is the same: attackers no longer need to compromise your code. They just need to compromise something your code trusts.
Why Your Secure SDLC Checklist Didn’t Catch This
Here’s the uncomfortable part. LiteLLM’s own development practices weren’t the failure point. The breach succeeded because of one unrotated credential, several hops upstream, inside a security scanner that most engineering teams never think to audit as an attack surface in the first place. A checklist that only covers your own code and your direct dependencies would not have caught this. The failure happened inside the tooling that exists specifically to provide security assurance.
Not everyone agrees this is an AI story at all, and that disagreement matters.
Ordinary DevOps hygiene failures under pressure to ship AI features quickly, not novel AI risk, is how independent researcher Kevin Beaumont frames the root cause.
Reported via Help Net Security, August 13, 2026
Beaumont’s contribution goes beyond commentary. He personally tested a major tech company’s public claim that it had rotated every exposed credential, and found working credentials still active months after the company said the issue was closed. That’s arguably the single most concrete finding to come out of this story: a “we already fixed it” statement from March may still be false in August.
Alon Gal, Co-Founder and CTO of Hudson Rock, described the scale of the credential archive as demanding a genuinely different tier of industry response than incidents like this have typically drawn.
Help Net Security, August 13, 2026
There’s a counterpoint worth holding onto, though, because it complicates the “the industry is failing” narrative that’s easy to reach for. GitHub’s Octoverse 2025 report found that average fix time for critical severity vulnerabilities improved 30%, dropping from 37 days to 26 days, and that 26% fewer repositories received critical security alerts over the same window. Dependabot adoption climbed to more than 2.6 million projects. Automation is working, where teams actually use it.
Our read: this isn’t a uniform industry failure. It’s a bifurcation. Teams running automated software composition analysis and enforced dependency gates are getting measurably safer. Teams without that tooling remain exposed to worm-class threats that spread faster than a human reviewer can react. The gap between those two groups is widening, not narrowing.
One counterweight worth flagging in the other direction: Broken Access Control overtook Injection as the most common CodeQL security alert in 2025, appearing in more than 151,000 repositories, a 172% year-over-year jump that GitHub’s own engineers link partly to misconfigured CI/CD permissions and AI-generated code scaffolds that skip authorization checks by default.
NIST, CISA, and the EU’s SBOM Mandate
Institutional responses exist, and they’re maturing, but nobody serious is calling them sufficient yet.
NIST SP 800-218, the Secure Software Development Framework, remains the most-referenced U.S. framework, required for FedRAMP and federal vendors. CISA’s Secure by Design pledge now has 68 signatory manufacturers, including AWS, Cisco, GitHub, GitLab, and Microsoft, all committing to specific security-by-default practices. And the EU’s Cyber Resilience Act is pushing Software Bills of Materials from a nice-to-have into a legal requirement for anyone selling software into the EU.
Saša Zdjelar, Chief Trust Officer at ReversingLabs, has credited CISA’s Secure by Design work with maturing the industry conversation on software security, while noting that current guidelines don’t yet fully address the complexity of the modern software supply chain.
ReversingLabs, “CISA’s Secure by Design Pledge”
Read between the lines and the honest assessment is this: these frameworks were largely built before ecosystem-scale, self-replicating worm attacks were a realized threat rather than a theoretical one. They’re catching up, not leading.
One caution flag before you cite this story elsewhere
A widely circulating quote calling the LiteLLM incident “the AI era’s SolarWinds moment” traces back to an April 2026 press release from a competing AI-gateway vendor promoting its own product, not to CloudSEK, Hudson Rock, Unit 42, or the FBI. A “36% of all cloud environments” statistic attached to that same quote appears in none of the independent datasets. Treat it as marketing, not research.
What Engineering and Security Teams Should Do Now
If your organization touches LiteLLM, Trivy, or Checkmarx KICS anywhere in a build pipeline, here’s the practical checklist, drawn directly from the FBI’s own recommended mitigation in FLASH-20260702-01.
Pin to commit hashes, not version tags. Floating tags are exactly what let TeamPCP force-push malicious code across 76 of 77 Trivy release tags in one move.
Audit your security tooling as an attack surface, not just your application code. The scanner meant to protect you is now a documented entry point.
Don’t trust a “credentials rotated” announcement at face value. Beaumont’s test proved a major company’s public claim was false months after the fact. Verify independently.
Check whether your org appears in the CloudSEK or Hudson Rock datasets. Inclusion means exposure evidence was found, not confirmed compromise. Treat it as an investigation trigger, not a panic button, and not a dismissal either.
If you’re not already running automated SCA scanning and dependency pinning enforcement, this incident is the concrete, current justification to get budget approved. GitHub’s own data shows it works.
FAQ
What percentage of data breaches involve third parties?
Verizon’s 2025 Data Breach Investigations Report found third-party involvement in 30% of breaches, double the 15% reported the prior year, based on 12,195 breaches, the largest dataset in the report’s history.
What happened in the LiteLLM supply chain attack?
In March 2026, threat group TeamPCP compromised the Trivy security scanner through an unrotated credential, which cascaded into LiteLLM’s build pipeline. Two malicious LiteLLM versions sat live on PyPI for roughly 40 minutes, later linked to over 2,500 exposed organizations.
What is a Secure Software Development Lifecycle?
An SSDLC builds security activities, like threat modeling, automated scanning, and code review, into every development phase instead of treating security as a final gate before release. NIST SP 800-218 is the most widely referenced U.S. framework for this.
How many npm packages did the Shai-Hulud worm compromise?
Shai-Hulud 2.0, identified in November 2025, backdoored 796 unique npm packages representing more than 20 million combined weekly downloads, and it self-replicates without needing a command-and-control connection.
Does pinning dependencies to a version number protect against this kind of attack?
No. TeamPCP force-pushed malicious code across 76 of 77 version tags in one Trivy repository. Pinning to an immutable commit hash, not a floating version tag, is the mitigation the FBI explicitly recommends.
Where This Goes Next
What’s changed after this week isn’t just the exposure count. It’s the assumption that “we fixed it in March” means anything in August. TeamPCP’s campaign proved that a compromise several tools upstream, in software meant to secure you, can sit undetected for months while credentials stay valid and reusable. That’s a longer blast radius than most incident response plans are built for.
Watch three things over the next six to eighteen months: whether the EU’s Cyber Resilience Act SBOM requirement actually forces vendors to disclose dependency provenance in a way that would have caught this earlier, whether the gap between automated and manual security teams keeps widening the way GitHub’s Octoverse data suggests, and whether more organizations quietly confirm they’re still exposed the way Beaumont’s test did. Five months of silence between compromise and disclosure was too long. The next one probably won’t be different unless the incentives change.
LiteLLM Breach: CloudSEK and Hudson Rock Diverge on Scale
Cybersecurity / Supply Chain
LiteLLM Breach: CloudSEK and Hudson Rock Diverge on Scale
Updated August 13, 2026 · 9 min read
Headline options considered:
★ LiteLLM Breach: CloudSEK, Hudson Rock Diverge on Scale
LiteLLM Hack: 2,500+ Firms Named, Credentials Still Live
Inside the LiteLLM Supply Chain Breach, Five Months Later
Two threat intelligence firms just published two different datasets about the same breach, and the numbers almost, but don’t quite, agree. If your organization runs the LiteLLM AI gateway anywhere in your CI/CD pipeline, that gap between “almost” and “exactly” is the part you need to understand today.
On August 11, 2026, CloudSEK published a victim-exposure dataset claiming more than 2,500 organizations and roughly 434,000 CI/CD pipelines were potentially exposed by the LiteLLM supply chain breach. Two days later, Hudson Rock came back with its own numbers, pulled from a separately obtained archive, and they landed close but not identical. That’s not a rounding error. It’s two independent forensic teams checking each other’s work in public, and the small gaps between their findings are as newsworthy as the breach itself.
Strip away the marketing language from both firms’ blog posts and the underlying story is a fairly classic dependency-chain compromise, just with an AI gateway as the final target instead of a bank or a pipeline operator.
The threat group known as TeamPCP first got in through Aqua Security’s Trivy vulnerability scanner, not LiteLLM itself. According to Unit 42’s research, the group used a credential from an earlier late-February breach that hadn’t been fully rotated to hijack Trivy’s service account, then force-pushed malicious code across 76 of 77 version tags in the widely used aquasecurity/trivy-action repository on March 19. Two days later, the same group used stolen GitHub tokens to do the same thing to Checkmarx’s KICS scanner.
LiteLLM was next. Its own incident report confirms that two malicious versions, 1.82.7 and 1.82.8, sat live on PyPI for roughly 40 minutes on March 24, between 10:39 and 11:19 UTC, before being pulled. Version 1.82.8 shipped a file called litellm_init.pth, a mechanism that runs automatically the moment Python starts, regardless of whether a team thought --ignore-scripts was protecting them at install time.
Then came the long silence. The FBI didn’t issue its public advisory, FLASH-20260702-01, until July 2, more than three months after the initial compromise. It formally named TeamPCP and confirmed Trivy, KICS, LiteLLM, and the Telnyx Python SDK as compromised vectors in a single, escalating campaign.
It took another five weeks after that for named-victim-level detail to surface publicly, first from CloudSEK on August 11, then from Hudson Rock and independent researcher Kevin Beaumont on August 13. Nearly five months passed between the original compromise and the point where affected companies could actually see whether they were on a list.
The numbers: CloudSEK vs. Hudson Rock
Here’s where the story gets more interesting than a standard “breach affects X companies” writeup. Both firms say they worked from separately obtained data, and their headline totals are close enough to corroborate each other, but different enough that neither should be treated as the final word.
Metric
CloudSEK
Hudson Rock
Organizations identified
2,500+
2,488 corporate domains
Underlying scope
~434,000 CI/CD pipelines
118,829 CI runner dumps
Source archive
Confidential intelligence sources
153GB archive, 433,909 files
Published
August 11, 2026
August 13, 2026
Notice that CloudSEK’s “434,000” figure and Hudson Rock’s “433,909 files” figure are suspiciously close in raw count, even though one is labeled pipelines and the other is labeled files. Cyber Kendra flagged this directly in its own analysis, arguing the widely repeated pipeline figure may actually describe leaked files or records rather than distinct pipelines, a distinction that changes how any single company should estimate its own exposure.
Why this matters for your risk math
Appearing in either dataset means information tied to your organization was identified in leaked material, not that an intrusion into your systems was confirmed. CloudSEK, Hudson Rock, and the FBI advisory all draw that line explicitly. Treat inclusion as a trigger for investigation, not a verdict.
What security researchers are saying
Hudson Rock co-founder and CTO Alon Gal framed the scale of the finding in blunt terms, telling Help Net Security that the size of this dataset demands a genuinely different tier of response from the security industry than incidents of this kind have typically drawn.
Independent researcher Kevin Beaumont pushed back on the framing that AI itself is the villain of this story. In the same Help Net Security piece, he argued the real failure is ordinary DevOps hygiene under pressure to ship AI features fast, not some novel danger inherent to AI systems. That reframe matters because it shifts the accountability conversation from “AI is scary” toward “your credential rotation policy is the problem,” which is a much more actionable takeaway for engineering leadership.
“The LiteLLM supply chain attack is the AI era’s SolarWinds or NotPetya moment.”
Craig Alberino, CEO and Co-Founder, APERION · BusinessWire, April 2, 2026
One caveat worth stating plainly: Alberino’s comparison came from a press release tied to the launch of his own company’s competing AI-gateway product, and it included a “36% of all cloud environments” claim that appears in no independent dataset from CloudSEK, Hudson Rock, Unit 42, or the FBI. Read that quote as vendor commentary, not as verified research.
Beaumont also did something more concrete than commentary. He personally tested a major tech company’s public assurance that it had already rotated all exposed credentials, and per the same reporting, found working credentials still active despite that claim. That’s the single most damning, verifiable data point in the entire story, and it’s a preview of what security teams should expect when they audit their own “we already fixed this” statements from March.
What to check in your own pipeline right now
If you’re a DevOps engineer, platform lead, or CISO reading this, the patch-and-move-on instinct doesn’t apply here. Confirming you’re not currently running LiteLLM 1.82.7 or 1.82.8 tells you nothing about whether credentials exposed in March are still sitting active somewhere five months later.
Audit CI/CD logs and Docker build history for any install of LiteLLM 1.82.7 or 1.82.8 on March 24, 2026, specifically between 10:39 and 16:00 UTC.
Search your site-packages directory for a leftover litellm_init.pth file, which persists even after the package itself is upgraded.
Search your GitHub organization for unexpected repositories named tpcp-docs or docs-tpcp, a known artifact of the malware’s fallback exfiltration path.
Rotate everything that touched an affected build: cloud IAM keys, SSH keys, Kubernetes service-account tokens, package-publishing tokens, and any AI-provider API keys.
Pin GitHub Actions and dependencies to verified commit hashes instead of floating version tags, per the FBI’s own recommended mitigation in FLASH-20260702-01.
Confirmed safe by LiteLLM’s own postmortem: anyone on the official LiteLLM Proxy Docker image, LiteLLM Cloud, or any self-hosted version at 1.82.6 or earlier who didn’t upgrade during the 40-minute window. Versions 1.78.0 through 1.82.6 and 1.83.0 onward were independently SHA-256 verified against Git commits, with Google Mandiant assisting the forensics.
What’s still unresolved
Reporting on this cleanly means resisting the urge to present a single tidy mechanism, because the primary sources themselves don’t agree on one. CloudSEK, LiteLLM’s own team, and Unit 42 each describe the path from the Trivy compromise to the malicious PyPI publish slightly differently, whether it was the poisoned build process itself producing the release, a direct unauthorized upload that bypassed CI/CD entirely, or stolen publishing tokens harvested during the Trivy breach. When asked about the discrepancy, CloudSEK told The Hacker News these were different stages of one attack chain rather than competing explanations, which is a reasonable answer but not the same as a confirmed one.
There’s also an unreconciled figure floating around secondary coverage: some outlets have cited an archive size as large as 195TB for what may be the same or a related dataset, against Hudson Rock’s own stated 153GB. Until one of the primary sources clarifies that gap, treat it as unverified.
Our read: this signals that “responsible disclosure” as a framing is starting to strain under its own timeline. Five months between compromise and named-victim transparency is a long runway for stolen credentials to be resold, reused, or simply forgotten by the teams who should have rotated them.
Frequently asked questions
What is the LiteLLM supply chain attack?
A March 2026 breach in which the threat group TeamPCP compromised the Trivy security scanner used in LiteLLM’s build pipeline, then published two malicious LiteLLM versions, 1.82.7 and 1.82.8, to PyPI for about 40 minutes before they were pulled.
How many companies were affected by the LiteLLM breach?
CloudSEK’s dataset lists more than 2,500 organizations as potentially exposed. Hudson Rock, working from a separately obtained archive, independently identified 2,488 corporate domains. Both figures describe potential exposure, not confirmed intrusion.
Is LiteLLM safe to use now?
Yes, if you’re on version 1.82.6 or earlier, or 1.83.0 and later, all of which BerriAI verified via SHA-256 hashing against Git commits with Google Mandiant’s help. Versions 1.82.7 and 1.82.8 remain compromised and were removed from PyPI.
How do I check if my organization was affected?
Audit CI/CD logs for installs of the two malicious versions on March 24, 2026, check for a leftover litellm_init.pth file, and search your GitHub organization for repositories named tpcp-docs or docs-tpcp.
What is TeamPCP?
A financially motivated group active since at least September 2025 that shifted in 2026 from ransomware toward compromising trusted developer and security tools, including Trivy, Checkmarx KICS, LiteLLM, and the Telnyx Python SDK.
Where this goes next
What’s clear five months in: this wasn’t a LiteLLM problem so much as a trust problem in the tools that sit upstream of nearly every AI deployment pipeline. TeamPCP didn’t need to break LiteLLM’s own defenses. It needed one unrotated credential in a scanner most teams never think about.
Three things worth watching over the next six to eighteen months: whether CloudSEK and Hudson Rock ever formally reconcile their overlapping-but-different datasets, whether regulated industries named in either list face disclosure obligations tied to insurance or compliance frameworks, and whether “pin to commit hash, not floating tag” becomes a default CI/CD posture industry-wide rather than a lesson learned twice.
If your team touched LiteLLM, Trivy, or KICS anywhere in a build process this year, the credential rotation audit isn’t optional. Five months of exposure is a long time for a stolen key to sit around waiting to be used.
Editorial note: An earlier NeuralWired piece referenced a “4TB exposed in under 3 hours” figure tied to LiteLLM and Mercor. That figure does not appear in CloudSEK’s, Hudson Rock’s, Unit 42’s, the FBI’s, or LiteLLM’s own reporting, and this article’s numbers should be treated as the current, sourced account of the breach’s scale.
Big Tech’s $760B AI Bet: Who’s Cashing In, Who Isn’t
Meta’s free cash flow just fell to $784 million. SpaceX’s AI capex grew sixfold in a single quarter. Amazon’s cloud arm is finally showing the receipts. Q2 2026 earnings season didn’t answer whether AI spending is a bubble. It answered something more useful: which companies can prove it, and which ones are still asking investors to trust them.