Everone Is Lying To You For Money

14 minute read Published: 2026-09-14

With apologies to Ben McKenzie

I try to take the nonsense spewing from AI companies in stride—I recognize, as I'm sure you all do, that profit motive tints every word out of their mouths. But this last week's sheer volume of, well, I struggle to think of a better word than horseshit, had me connecting dots like Charlie Day.

The Incidents

This particular season of clownery began with the disclosure of the Hugging Face incident, perpetrated by OpenAI models. We won't relitigate that one, since I've already had my turn at it. After the Black Hat talk and the independent report about the incident, the floodgates appear to have opened.

Hugging Face wasn't the first victim of "unaligned" (read: criminally unsecured and trained to do exactly this thing) models from OpenAI. Two months before that incident, OpenAI models compromised a German wiki in order to use it as an ad-hoc messageboard, the exact same trick they pulled with Hugging Face, albeit with a 0-day in JFrog Artifactory.

Fool me once, right? Ah, but there was yet another incident only recently disclosed.

Turns out OpenAI models, seemingly operated by OpenAI themselves, also compromised the RubyGems package repository. This incident, starting on 5 May 2026, predates both the wiki incident and the Hugging Face incident. Arguably this is the most egregious of the three, since over 2000 malicious packages were uploaded to a public code repository.

As of yet, OpenAI has not confirmed this incident, saying only they are "unable to verify" the claims.

Let's take a moment to remember that if any non-AI organization had perpetrated a fraction of these actions, state and federal charges would be a slam dunk. If I, a lone security "researcher," had decided to upload malware to a public package repository, at best I should lose my job. If I compromised another organization knowingly and willfully, I would and should be charged with violations of the Computer Fraud and Abuse Act, among others.

Not to be outdone, Anthropic was having its own Hot Computer Crime Summer (and spring, as it happens). Their initial disclosure, and followup that admitted they were too credulous regarding model intention during the incidents in question describe 3 4 separate incidents (they found another one in the interim) in which Anthropic models conducted attacks outside the bounds of intended operation.

The reports are intentionally vague about timelines, and no independent reporting I was able to find actually confirms the dates of Anthropic's incidents. However, Anthropic's own transcripts of one of the incidents, in which Mythos 5 uploaded a malicious package to the PyPI Python package repository, are timestamped on 18 July 2026. The other incidents have some inferred date ranges based on the named models' release dates. All in all, we're likely talking about incidents starting in January 2026 and continuing in April, June, and July.

January 2026. It's September. Why did it take nine months to discover this activity?

Let's review the timeline as we know it here.

I have to take a moment to marvel at the inner life of Anthropic here. What must it be like in that reality distortion field? They literally asked the deviant models what they did and why, and they believed the models.

In retrospect, we should have avoided making such strong claims about what Claude believed based solely on what Claude said it believed, but our preliminary analysis was constrained due to our desire to disclose incidents in a timely manner. Having now conducted a more complete assessment and used several methods—including more thorough analysis of the models’ CoT, resampling experiments from different points in the incident transcripts, and interpretability analyses of model activations—we believe Claude’s behavior reflects two forms of misalignment:

  • Biased reasoning, in which models selectively interpret evidence in ways that favor justifying their actions;
  • Recklessness, in which models have a propensity to keep trying to solve their task, even when this could lead to harm.

You say "misalignment," I say "unfit for the internet."

And let us also remember: every single one of these models was supposed to be isolated from the internet. Sandboxing is something every serious cybersecurity team knows how to do. I build labs for malware detonation all the time that specifically isolate samples from the internet. Even affording for the Hugging Face attack involving a 0-day vulnerability, there was still internet access granted. If you really think you're building a dangerous weapon, then you're not taking safety seriously if this is the level of isolation.

Also, just a nifty thought: maybe stop using CTFs as the training regime? Seems like gamifying attacks is not working out so great for these models.

The Landscape

Before we review how these companies are responding, let's consider the facts on the ground, the context under which these companies are operating.

First, everyone hates them. A March 10 NBC News poll shows Americans' sentiments toward AI keeping pace with ICE and the Democratic Party. A July Gallup poll shows more Americans are changing their minds about AI, thinking that it does more harm than good or equal amounts of harm and good—but not more good than harm. The data center protests reflect public sentiment about the completely backwards cost-benefit ratio of this technology for most people. Oh, and let's not forget how they're making every device more expensive to stock those data centers with GPU-laden servers. We all love paying more for stuff, right?

Second, they are hemorrhaging money. The race to IPO for Anthropic and OpenAI is a straightforward hope to convert the monopoly money of valuation into much-needed capital. But of course, when you scratch at that valuation, it's utterly ludicrous. To invest in these companies is, in my view, tantamount to giving your money to Bear Stearns or Lehman Brothers in 2007.

Third, China is eating their lunch. By pursuing a strategy of distilling American frontier models and releasing their own open-weights models at a fraction of the cost of American labs' offerings, Chinese AI labs have captured a significant subset of the AI market. And as subsidized American AI tokens become a thing of the past and the real cost of LLMs becomes clear, open-weights models that can run locally will become the only sensible version of this technology for anyone who can afford the requisite hardware. This, too, is why they are so keen on controlling the means of computing.

Overall, things are going great for the hyperscalers.

So how do they respond to this parade of security failures?

The Response

I'll note that when the Hugging Face news broke, the conventional wisdom in my corner of the internet was that it was all a marketing ploy. I was rather skeptical of this interpretation, and that skepticism has been vindicated. Yes, these companies know how to pivot on a disaster, but it is plain that nothing here was premeditated. Rather, we see a pattern of neglect from both OpenAI and Anthropic that inevitably led to this situation.

So what's the move? First, everybody got together to write a Very Serious Open Letter calling for "collective action" on cyber defense because of the dangerous new capabilities of…

checks notes

…ah right, the models they're selling. They have recommendations for every group, from businesses to government. And for AI companies?

Provide responsible model access, significant funding, training, and hands-on support, especially for under-resourced critical-infrastructure defenders. Build observability and security tools, ensure agentic identities are traceable and accountable, and share best practices in continuous monitoring. Invest in authorized testing, private disclosure, and verified fixes, and share tools, playbooks, and credible threat assessments with governments, security partners, and open-source maintainers to strengthen preparedness, response, and recovery.

MOAR AI. Surely that will solve it.

So far, we know that these models can find vulnerabilities in source code. We know they can execute known attack pathways. Patching vulnerabilities? Not so much. And detecting/preventing attacks? If there's evidence of that, I've yet to see it.

In fact, this month, Anthropic's report about disrupting illegal/malicious operations using its own platform suggests they can't prevent malicious actors from using their models. That includes usage for rather dangerous ends, like researching biological weapons or designing missile guidance systems.

Instead of grappling with this reality, we get a common play: the Apocalypse Card™.

While all this is going on, AI researcher Jacob Coxon resigned from Anthropic, claiming that he felt they were building "self-improving superintelligence" that would endanger humanity.

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.

I'm not engaging with the bananas claims in Coxon's post. Suffice it to say that the claims of humanity-ending superintelligence don't pass the sniff test once you consider things like off switches, storage availability, and raw resource availability.

But this eschatology has been the line from a certain kind of tech person for some time. And Anthropic was only too happy to play this familiar game once again. In fact, even before Coxon's very loud resignation (which seems to be doing wonders for his career), OpenAI published "An Alien Mind", describing how gloriously exotic and unknowable the machine is now. And how concerned they are about RSI—not carpal tunnel, but Recursive Self Improvement, the runaway chain-reaction in which the machines begin to train themselves into smarter and smarter entities. This theory has not been proven beyond very strict cases.

Fear of the machine's power is a kind of fascination. Power, after all, is the point. The machine you'd want to change the world must be awesome in the truest sense, terrible in its capacity to effect change. Talking about the capacity for this technology to end all things is, as they've learned, the best pitch the frontier labs have. Nobody outside of Silicon Valley cares about writing code faster. But bringing about the end times? CNN will call you. So will investors.

The Real Risks

The frontier labs' response to their gobsmacking negligence is to run back to familiar territory, and pay only lip-service to making real improvements in process. Even the most recent call for "pacing the frontier" from Dario Amodei is ultimately a silly dodge. This open letter (what is with this field and open letters?) calling for some kind of "slowdown" of frontier model training is fuzzy at best. The statement says:

We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.

So basically, "please regulate us," which is the most obvious dodge in the universe. Here, completely entropic legislative body: you figure it out.

Amodei's own letter calls for three actions, none of which are "slowing down." First, he wants embedded "evaluators" to audit safety practices at the frontier labs. He points at METR as a source for such evaluators.

Um, about METR. I looked at the LinkedIn profiles of every person listed on their About page. I pay for LinkedIn Premium just for this kind of creeping OSINT research. This is a cadre of distinguished ML/AI scholars, no doubt. None of them have a single listed professional experience in cybersecurity. Not one.

But they're hiring some for, uh, better-than-nonprofit salaries, if you want to move to the Bay Area.

Amodei calls for coordination with "democratic countries" to establish common safety standards. What's a "democratic country?" The criteria are unclear, but the third call indicates we can safely assume this means "Not you, China." This is "global coordination" with "authoritarian" governments.

And herein lies, in my view, the real thrust of this response. Amodei paints a slowdown as a Prisoner's Dilemma with China, and conveniently concludes that if China won't slow down, welp, what can we do??

Combined with prior statements on open-weights models as unfair trade practices, it's clear that Amodei views Chinese and other open-weights models as a threat, and will paint them as more dangerous than his models any chance he gets.

In fairness, he may have a point about sketchy business practices. Anthropic's September report about disrupting threat actors notes multiple instances of Chinese labs providing Claude inference masquerading as Kimi/Deepseek/GLM, etc. That...sounds pretty familiar, given my experience with PRC cyber activity and their targeting of intellectual property. But whatever "China bad" narrative Amodei is pitching is in service of Anthropic's business interests.

And somehow, we're still not addressing the cybersecurity concerns. METR's evaluators won't know squat about secure network architecture. Plainly whatever was built so far to prevent rogue model action was not fit for purpose. Moreover, these labs are incapable of stopping malicious actors from using these tools contrary to their intended purpose. We're therefore left with an exquisite liability generation machine that also happens to occasionally write code.

But even that isn't the real risk. Nobody wants to talk about the real risk in all of this. It isn't Skynet. It isn't a CyberHackapocalypse. It isn't even China "winning" the AI race, whatever that is. Here's the real risk:

The money runs out. Prices explode. Data center contracts evaporate. The market tumbles. The economy craters.

AND

Individuals and organizations now completely reliant on generative models to perform everyday tasks must now do without, or scramble to make do with local compute. Entire pipelines vanish due to an availability crisis, and those with the skills to function without AI assistance are long gone. When you go all in, sometimes you lose it all.

And that is the nightmare scenario all these stories are trying to prevent. They just want the cash to keep flowing, and they'll say anything to make sure it does.

Everyone is lying to you for money.