This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR
- A major explosion of mainstream interest in AI extinction risk, including among many US politicians.
- The first legal requirement for AI auditors is signed.
- Anthropic withholds Mythos 5.1 from the UK’s AISI, reports various misuse and cyberattack incidents, and seeks more defense and natsec contracts.
A note zooming out from the weekly level:
- On Tuesday, the last mile of a Millennium Prize problem was solved using AI.
- On Wednesday, an Anthropic resignation post became the most-liked AI tweet ever(?) and was covered on Fox, WSJ, NBC, CNN, BBC, etc and commented upon by at least 20 US Congress members / senators.
- On Wednesday, Daniel Kokotajlo went on Joe Rogan, still the biggest podcast in the world, to talk about AI taking over.
For those of you who have been marinating in AI news since ChatGPT (2022) and are sick of it, it might be hard to believe that another order-of-magnitude increase in attention and discourse on AI is coming. But to us it finally feels like Covid in March 2020: the moment the tinder lights and the world turns its full attention to the threat, which means the lagging politicians can act. This is not an unalloyed good, since bad policy is very likely to result from it, and since we are not fully convinced that full capability generalization is happening very fast, but it gives good policy a chance.
Economics#
Anthropic releases a simple economic model predicting AI’s influence on the US economy up to 2030. Its “extreme” scenario (15% annual economic growth, “likely driven by recursively self-improving AI systems”) sees unemployment jump 8% after the knowledge work share of jobs falls from 62% to 49%. That is, the worst-case prediction is that less than half of displaced workers would find new jobs.
Labor productivity rises through this channel by well under one percent in all three scenarios, and measured TFP by less.
Opinion: A broad family of results with dials you can fiddle with yourself is the sane way to present such an underidentified model. And Anthropic is open about the many caveats below, and the robustness analysis is fine. But the parameters that that analysis shows to be most important (wage rigidity, capital elasticity, reinstatement) are clamped to middling values in the backend.
Worse, the model bakes in optimism about the human labor share by construction: despite the blog post appealing to RSI, the actual model ignores R&D feedback, (already massive) non-AI automation, treats tasks as complements, and treats physical and interpersonal work as completely unautomatable. As it says, “the innovation effects here are likely a lower bound.” (Not modeling robots is pretty fair given the 2030 cutoff.)
The extreme scenario is actually close to the ceiling of this model. This is actually more inflexible (on the possibility of explosive growth) than older models like GATE. Anthropic also don’t model the possibility of a shortfall in AI demand, a capex collapse, catastrophic risks, or bans.
It also contradicts another Anthropic prediction: the extreme scenario is roughly the conservative end of Dario Amodei’s claims that “half of entry level white collar jobs” could be wiped out within 1-5 years, and that 10-20% unemployment is plausible.
Overall, nice technical work being prematurely pushed as predictive and unconditional.
Yet another inference chip startup, the two-year-old Positron, raises $875M at a $5B valuation. The company is notable for being founded by a former MIRI researcher, and for its apparent use of a galaxy-brained theory (“denotational semantics”) to enable the use of more commodity DRAM than (incredibly scarce) HBM.
Its Atlas chip is a “fixed-function” Transformer-specific FPGA (a “hard-wired Transformer operator graph”). The program the chip executes is just the inference config plus a weights file streamed from external DRAM.
Opinion: Note that Nvidia already uses a pile of extra DDR5 as well as the on-accelerator HBM. And Positron’s approach is roughly the same as Etched ($21B valuation).
Inference chips are a mixed bag in risk terms: they do free up a bunch of GPUs for larger training runs, and RL post-training is substantially an inference problem (the “rollouts”), but they make it easier to verify that an inference cluster isn’t a secret training cluster, which is the really worrying part under some threat models.
Positron is one of 54 AI companies to get a billion-dollar valuation in 2026. We are old enough to remember when unicorns were rare, and when OpenAI was valued at $1B (2019).
Double-digit growth over the coming 10-15 years is unlikely, claim two prominent economists. The authors grant that a capabilities explosion is likely, but argue that this does not entail the would-be unprecedented growth numbers forecast by many influential economic models. The case for AI-enabled extreme GDP boosts is premised on five potentially contentious assumptions:
- Machines will automate a majority of economic activity.
- Spending on (increasingly cheap) goods and services will continue.
- That companies will invest to increase their productive capacity, and that there will be enough people who can afford to consume it (sufficient demand, that is).
- AI does not destroy its value through, for example, hacking incidents and rogue agents.
- AI successfully becomes the work-horse of science.
Due to the uncertainty around each assumption, the essay concludes that 4-5% growth by 2035 is a more likely outcome.
Opinion: Insightful analysis. We think the razor edge between “AI doesn’t have that much impact” scenarios and “AI has so much impact that standard macro models aren’t applicable” is pretty thin.
Also, double-digit GDP growth is a very specific measurement. GDP is an imperfect metric, since “when automation makes the automated thing cheap, people and firms do not buy proportionally more of it. Spending instead shifts to what is still scarce, the tasks still performed by humans.” You could be, and are, getting absurd transformations without hitting that metric.
Bridgewater’s Co-Chief Investment Officer and early OpenAI/Anthropic investor Greg Jensen believes AI will likely kill people before meaningful action is taken, comparing it with the slow response to Covid-19. On a Bloomberg podcast, Jensen says:
Until the AI starts killing people, unfortunately, history would suggest we’re not going to do anything, but we are going to face that. That’s going to happen, and it’d be much better if we started dealing with it before then.
Opinion: Seems too pessimistic. The public is substantially more activated in response to data centers (which are broadly disliked) and AI in general (which they feel, at the very least, conflicted about) compared to pre-lockdown Covid. It also appears to be strikingly less polarizing along partisan lines, which is certainly encouraging. Perhaps more to the point, a Republican-led Senate subcommittee is now investigating the Hugging Face incident. Regulations might arrive sooner rather than later.
Capabilities#
In an article about the OpenAI–Buckmaster dispute, the New York Times cites an OAI claim that it has “made substantial progress on another Millennium Prize problem.” Rumors have been circulating that it is the Hodge Conjecture, and that Anthropic may have solved the Birch and Swinnerton-Dyer Conjecture.
Opinion: This is a major epistemic fork in the road for those of us who aren’t AI lab employees. But what matters is less whether AI solved Hodge and/or BSD than whether AI has found proofs (as opposed to disproofs) of Hodge and/or BSD.
Proofs of Hodge and/or BSD would indicate a step-change to properly superhuman mathematics. Conditional on proofs, we should divide our credence between several hypotheses regarding AI capabilities:
- “Superhuman math capability demonstrates broadly applicable superhuman reasoning capabilities” (e.g., OpenAI has achieved ~ASI internally).
- “Superhuman math capability demonstrates superhuman reasoning capabilities in the technical sciences” (e.g., OpenAI has an AI that can do breakthrough research in AI capabilities).
- “Superhuman math capability is only narrowly trainable, like superhuman Chess or Go capabilities.”
Disproofs of Hodge and/or BSD could be on par with the Navier-Stokes result from the viewpoint of capabilities demonstrations, though they would likely involve more of what we’d call mathematical creativity if the disproofs were achieved by a human. Hodge and BSD don’t have exciting recent human-authored disproof programs (like Córdoba and Martínez-Zoroa’s Navier-Stokes program) for AI to “last mile,” so disproofs may well involve developing a novel mathematical line of attack on these problems. But a neglected elementary counterexample like the one Claude found to the Jacobian Conjecture is also theoretically possible – which makes the results consistent with Timothy Gowers’ observation that LLMs excel at “breadth first” mathematical work.
We note that if Millennium Prize problems beyond Navier-Stokes are solved by AI this month but the solutions are all counterexamples with classically LLM-flavored proofs (vertical crunch or breadth-first search but no theoretical or strategic innovation), the right update might be that frontier AI is even more jagged than anyone imagined. Achieving strongly superhuman capabilities in mathematical “crunch” and breadth-first search without achieving even ordinary excellence in the more qualitative aspects of mathematical thinking would be good evidence that our concepts of creativity and depth correspond to real obstructions to LLM capabilities.
Other notable possibilities are that OpenAI’s report of “substantial progress” refers to a proof of a special case of a theorem or of a statistical approximation of a theorem. In these cases, we should largely regard the rumor mill as misleading.
DeepMind announces AlphaGenome Atlas, an encyclopedia covering each of the ~9 billion possible single-letter changes in the human genome.
Opinion: A nice contribution toward making genomic models more accessible, in particular by reducing the computational burden scientists typically face to run models like AlphaGenome. In many ways, this is akin to AlphaFold DB – with two main caveats: AlphaGenome is not to genomic variant prediction what AlphaFold 2 was to protein structure prediction; and, where AlphaFold DB was free for all, AlphaGenome Atlas gates its predictions behind a non-commercial API that forbids training other models on it.
Cognition factors RSA-260, a number with 260 decimal digits, or 862 bits, notable because RSA is a widely used encryption scheme relying on the difficulty of factoring large semi-prime numbers. The researchers used a “bevy of Devins” to build what Eric Lu calls
the world’s highest-performance GPU lattice siever, which enables factoring numbers at 10x lower cost than the previous public state of the art.
To be clear, neither Lu nor his bevy of Devins devised an essentially new algorithm here.
[I]mplementing lattice sieving and sparse linear system solving on GPUs required only “good old performance engineering” to take advantage of the preposterous memory systems of the GPU.
Lu notes that these advancements are unlikely to impact the practical soundness of RSA encryption, given that
Hyperscalers or frontier AI labs could likely factor RSA-1024 numbers at a cost on the order of $30 million per number — and, with a bit more optimization, likely substantially less. On the other hand, RSA-2048 remains roughly a billion times harder than RSA-1024 and does not appear to be meaningfully affected by this work.
RSA-1024 saw a sharp decline in use by the early 2010s, while most contemporary applications of the encryption algorithm use 2048 or 4096-bit keys.
Opinion: This is great. Top headlines recently have been about frontier labs solving sexy problems, but this is an example of a second-tier lab making progress on something nifty. But the thing is, there are many such problems, and for every NS solution you should also picture tens of thousands of minor problems solved, partial progress, and AI diffusion.
It’s neat that the Cognition team devised a way to execute its AI-crafted factorization algorithm at “no marginal cost on spare or fragmented compute that couldn’t be used for other purposes,” exploiting the embarrassingly parallel nature of lattice sieving to distribute the necessary calculations on GPU nodes that would not have otherwise been doing anything meaningful while other jobs were running. We don’t often have occasions to applaud computational efficiency and thrift in the field of AI science, so this is nice to see!
As an aside, note that the difficulty of breaking RSA keys doesn’t grow proportionally to their length: a key that is twice as long doesn’t take twice as much effort to solve, but a bit less using the best algorithm available (the general number field sieve). But despite compute growing exponentially per Moore’s law, and despite that algorithmic advantage, historical progress in factoring RSA keys has been roughly steady, far slower than the exponential growth in available computing power would permit. What gives? This discrepancy has more to do with motivation and cost than the sheer mathematical and computational resources available: there is little security interest in cracking obsolete key sizes, while cracking larger keys (≥ 2048 bits) remains astronomically out of reach. It’s safe to say that AI will not accelerate this historical trend beyond the hard limit imposed by the real security strength of each key size (barring developments in quantum computing) – but it may well bring it back into line with the historical acceleration of computational power, particularly if the free cycles on the burgeoning racks of available GPUs can be used for this purpose.
Data bottlenecks will not stop an intelligence explosion but could moderate its pace, argues Forethought’s Tom Davidson. He considers four variations of the argument that data bottlenecks would mitigate an intelligence explosion.
The first is that “you’ll need millions of trajectories of every job,” data that would take many years or decades to accumulate. But this isn’t necessary, says Davidson: labs are betting on automated AI R&D producing new “sample efficient learning algorithms,” so one wouldn’t need astronomical quantities of data – only as much as (or less) than humans require to learn a new task.
Next is the thought that so far AI progress has been contingent on the exponential growth of training data. This argument, Davidson says, suffers from a confusion about how an intelligence explosion would occur: the explosion would be downstream from an explosion in AI’s software engineering capabilities, and those pertaining to AI development in particular (i.e., RSI). This would facilitate “getting the same capabilities from less compute and less data,” so exponential increases in training data wouldn’t be needed.
The third argument is that an intelligence explosion requires a higher quality of training data, entailing the drag on progress that is human involvement. However, there is a limit to the quality of the data humans can produce; AIs already augment data to create “superhuman data,” and “will have to manufacture higher-quality data than any that exists, and do so from scratch.” This may slow the explosion, but would not be a significant, or human-shaped, bottleneck. The fourth consideration is the “paradigm tax”: the idea that AI may match humans on the techniques it inherited, but fall below humans on the ones it creates. Davidson also does not believe that this effect will be significant enough to preclude an intelligence explosion.
Opinion: It is a plausible argument, but one that rests on a tower of assumptions and extrapolations about the shape of future progress. In particular, it assumes a sufficiently powerful and persistent AI-R&D feedback loop, then argues that this feedback loop can invent its way around data bottlenecks.
So while this may seem plausible prima facie, it’s not guaranteed – and the case made for it in the article is fairly thin. The repeated claim that the magnitudes of counterfactual slowdowns will be roughly similar at all capability levels (rather than getting progressively harder to overcome) is also not substantiated. Overall the whole case seems a bit circular: data bottlenecks will not prevent rapid progress, because sufficiently rapid progress will produce the innovations needed to overcome data bottlenecks.
Models display log-linear test time scaling, plateauing at around 24 hours, while humans continue to improve at a given activity much past this point, notes Ramez Naam. David Pfau says this is surprising.
Opinion: Test-time compute scaling maximalism (i.e., “test-time compute scaling is all you need”) was all the rage ~16 months ago. It’s interesting how quietly it went away, but it’s also fair to argue that today’s scaling maximalism is more intellectually mature: pushing network parameter count, post-training intensity, and test-time compute in concord keeps delivering with no end in sight, although there is room for disagreement about just what it’s delivering.
Eric Jang argues that robotics is now entering a similar “smooth exponential” phase as seen in general AI capabilities, rather than the stop-start advances it previously exhibited. He identifies three causes: a larger robotics ecosystem, LLMs’ increased competence in robotics, and the discovery that a large data corpus (not specific to any particular hardware) can teach robots generalizable skills, with only a small amount of robot-specific data needed after. He anticipates that “the first 2-3 general-purpose home robots [will] go on the market in October 2027.”
Opinion: Maybe! Although we think the compute shortage makes it less likely that robotics and LLMs will progress full-throttle in parallel. More likely either frontier LLMs absorb robotics (see Astra’s robot-arm performance) and robotics progresses as a function of the LLM frontier, or the LLM frontier reaches a point of diminishing returns and the tech economy resets around robotics.
Politics#
OpenAI changes its tune on AI regulation, with a statement describing our current period as “a closing window to establish durable safeguards before AI capabilities outpace the institutions responsible for governing them.” Largely, the article tells us all the ways OAI is meeting the moment: “Pushing for mandatory national AI safety requirements”; “Keeping up momentum in the states”; “Advancing industry-led standards”; “Building global standards.” The statement acknowledges its debt to OAI’s chief scientist Jakub Pachocki warning of the acute risk of RSI.
Opinion: OpenAI becomes a little more coherent in the good direction, if only by reining in its global affairs lobbyists to match the present mood of alarm.
US Treasury Secretary Scott Bessent argues that the US must keep the lead in AI and can’t pause since China won’t. He also raises concerns around Chinese distillation of US models. (Bessent will likely hold AI-related talks with Chinese Vice Premier He Lifeng this month.)
…I will go back to: we have to keep the lead and for now we can because […] in AI when you steal or copy from someone it’s called distillation. I keep waiting for one of my children to say, “Oh, I didn’t copy my friend’s homework. I was distilling it.” Right? But, definitionally, if you are looking over the shoulder of the United States of America, you can’t pass us for right now.
But if we do, as Senator Sanders said, “Oh, we should just take a one-year pause.” Well, this is the same person who had Chinese professors come and speak at his AI conference, right? We can’t pause, you can’t because the Chinese won’t pause, even the North Koreans, they won’t pause. But we have to make sure that the American people see how the benefits accrue to them.
Opinion: In light of the very good relationship between the Trump administration and OpenAI, it’s interesting that OpenAI is currently at its most (publicly) sympathetic to a government-mandated pause and the Trump administration is at its most anti-pause. Incidentally, Anthropic has disappointed many recently when its response to the Coxon story discussed model-release pacing but no research pause.
Beliefs of the form “we can’t because China won’t” can also become a self-fulfilling prophecy, a hyperstition, a belief that to some extent becomes more true the more you believe it. Dangerous business.
The NSA, CISA, and FBI release a joint report accusing Chinese AI companies of “industrial-scale distillation” of US models. The report claims that distillation is a “core,” not merely supplementary, part of China’s AI development efforts.
China’s labs’ methods here are said to be diverse. Distillation tactics use “[APIs], remote cloud providers and third-party aggregators that automatically obfuscate user metadata to avoid detection”; “a gray market of proxies known as ‘transfer stations’ to bypass U.S. AI companies’ geographic restrictions, breach terms of use, evade safeguards, and undermine traceability”; “chain-of-thought […] reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks to detect defensive countermeasures.”
According to the report, this amounts to “aggressive, malicious, and targeted distillation activities at an industrial scale that extract restricted proprietary functionalities and capabilities of U.S. frontier AI models.” Perhaps more relevant is that distillation saves Chinese companies a lot of money, allowing them to undercut US developers. Accused parties include the developers of DeepSeek, Moonshot, and Alibaba; alleged victims include Anthropic, OpenAI, Google, and xAI. The authors stop short of outright accusing Beijing of complicity in the alleged distillation of US models, saying that it is “likely” aware of these practices.
The report provides three recommendations for US labs to mitigate distillation attempts:
- “Implement comprehensive detection and mitigation”
- “Deploy targeted response changes”
- “Establish cross-organization intelligence sharing”
These distillation attacks have also been analyzed at length in Anthropic’s September 2026 report on Detecting and countering misuse of AI: September 2026. The Google Threat Intelligence Group (GTIG)‘s recent report does likewise, and furthermore hints that Google has adopted something resembling policy (2), seeking to sabotage distillation efforts with “real-time proactive defenses that can degrade student model performance.”
Opinion: That the FBI, NSA, and CISA released their report to the public rather than as a private memo to the American frontier labs is significant, and suggests that their intent is just as much to sow FUD amongst the Chinese labs as it is to steer the policies of their domestic counterparts. The threat of having the American labs discreetly corrupt or degrade their models’ output and exfiltrated reasoning traces reads like an echo of the Farewell Dossier affair of 1982, where the Americans covertly allowed Soviet spies to exfiltrate maliciously modified OT software and hardware designs that would eventually result in the Trans-Siberian pipeline explosion. (Of course, the plans for that disinformation campaign were not announced in a public brief.) What the Americans are effectively threatening here is that surreptitiously distilling models may well result in the development of degraded or dangerously misaligned models by the Chinese labs. Though, since there are bound to be false positives in the detection of unwanted distillation efforts, these same measures would no doubt result in misinformation being fed to “legitimate” customers of the American oligopoly.
It’s worth stepping back from this for a moment to revisit the question as to whether the existence of highly capable open-weight models alongside the closed-weight models is a bad thing. The near-term consequences of distilling closed-weight frontier models’ capabilities into open-weight models that can be downloaded, studied, and fine-tuned by anyone with access to the necessary hardware would seem to greatly benefit science and safety research. This practice is not, of course, without risks, given the dual-use nature of the technology in question. But the very possibility of distillation that the Chinese labs exploit shows us that hermetically sealing away the models’ weights is not sufficient to contain the proliferation of comparable models. And several of the arguments that can be marshaled against those distillation practices – that they depend on a surreptitious acquisition of data that flouts various terms of service agreements and intellectual property rights, for example, or that they place potentially dangerous technology in the hands of parties whose motives we have no reason to trust align with the common good – can and have been marshaled no less compellingly against the American oligopolists. A concentration of AI power in the hands of a few would seem to be an especially brittle and dangerous state of affairs, even if one were to consider those hands to be especially virtuous ones. The capture of such strongly centralized power by powers less aligned with our common welfare is all it would take to tilt us into catastrophe. The pessimistic reader might find ill omens of such, for instance, in news of frontier labs’ present and future collaborations with the US Department of War.
Anthropic withholds Mythos 5.1 from the UK’s AISI, reports the Financial Times. Relative to the US, Britain’s AI industry is very small. But its premier testing agency has become a world leader, making Anthropic’s decision to withhold its latest model from AISI a sizable knock for the UK AI industry.
Some in Whitehall are speculating that Anthropic is acting at the instruction of the Trump administration, which has in recent months imposed export controls and protectionist measures around AI. Anthropic declined to comment to the FT. Though it did, as quoted in the article, say last week that it was “co-ordinating with the US government to expand access to a broader set of domestic and international partners as quickly as possible.”
Opinion: Very surprising! AISI is not primarily known as a UK institute but as the world authority on the intersection of alignment and cyber, and has close community ties with the alignment teams at both OpenAI and Anthropic.
The White House’s “trusted partner” program (by which it selects companies that can access pre-release frontier models to search for gaps in their own defenses) has opaque membership criteria, with few (if any) outside the administration knowing how one gains access. The Information claims that the White House has final say, while the program has effectively absorbed both Anthropic’s Glasswing and OpenAI’s Daybreak, the equivalent early access programs for each company.
Opinion: Opacity, arbitrariness, and scope for corruption aside, this is actually quite a nice precedent: the two leading labs are firmly coordinated on one axis, blatantly-dangerous-capability diffusion.
The American Prospect reports on Anthropic’s forays into protest surveillance and police collaboration. The article also notes that Anthropic’s recent job posting for a “Head of National Security Sales (DoW/IC)” reveals it is seeking staff to “lead and scale [their] national security sales organization to drive the adoption of safe, frontier AI across the Department of War and the Intelligence Community.”
Prospect describes Anthropic’s surveillance practices as a “kind of pre-crime policing, encouraged without due process,” and anticipates an escalation of these practices in the wake of lobbyists’ push to have AI classified as “critical infrastructure,” thereby wrapping “frontier labs in the hardened cloak of national security.”
Opinion: Worries that activists protesting the AI industry and the rapid spread of data centers in the US could be targeted as “pre-violent” persons of interest and unjustly accused of terrorism are well-founded, but the mechanisms and institutional precedents facilitating this are already in place, independent of any proposed designation of frontier labs and their assets as “critical infrastructure.” NSPM-7 already lists “anti-capitalism” as a terrorism indicator, the White House has issued an executive order classifying anti-fascism as a “domestic terrorist organization,” and the current regime has not hesitated to arrest non-violent protestors and try them as domestic terrorists. The risk that similar measures could be brought to bear on anti-data-center and anti-AI activists stands with or without the designation of AI labs as critical infrastructure. Which is not to say it couldn’t substantially raise that risk.
That Anthropic may be warming up to the prospect of doing business with the Pentagon (again) is indeed alarming. The company’s current job postings bear out Prospect’s reading, and provide evidence of a far more extensive move toward collaboration with national security agencies in both the US and Canada, with 15 listings detailing positions interfacing with various government agencies. Two of these make explicit reference to the US National Security sector and the Department of War/Intelligence Community, in addition to the post Prospect cites. Another listing announces an opening for a conventional weapons policy design manager to “define the line between prohibited weapons development and legitimate research and engineering work, and how to operationalize the distinction,” a role that “spans every weapon class, up to autonomous systems that select and engage targets without human authorization,” together with a call for an enforcement analyst with “deep, applied expertise in weapons systems.” Anthropic is also seeking a legal specialist to assist with requests from law enforcement for user data. Taken together, these and other postings signal that the company is indeed seeking out closer ties with military and police.
In a survey of ~1,300, advocacy-pollster Data for Progress finds widespread support in the US for pausing AI development or implementing an outright ban on superintelligent AI. Specifically, the statement is supported by 68% of voters – including 72% of Democrats, 63% of Republicans, and 70% of Independents.
The poll comes in response to Bernie Sanders and Greg Casar’s proposed bill. Notably, the polled statement, like the bill, does not distinguish between a pause and a ban, bundling them together.
Opinion: There’s longstanding passive public support for regulations, pauses, and bans on the development of superintelligent AI. The support has, however, so far been low-salience and low-intensity. Media signals suggest that salience and intensity might now be rapidly growing, but this hasn’t been social-scientifically measured.
But also… poll response strongly depends on wording, and this should also be viewed as an artifact created to support a specific position as well as a sampling of public opinion.
Gavin Newsom signs the first legislation to mandate third-party audits for AI. The new legal framework includes two bills: Senate Bill 813, “which establishes a first-in-the-nation framework for independent verification organizations that can assess AI systems and models for compliance with state law”; and Assembly Bill 1405, which creates “a state registry for AI auditors and establishing standards for their independence, transparency, and integrity.”
Together, the bills establish a framework for independent third-party evaluation and audits, laying the foundation for greater transparency and accountability as AI becomes increasingly embedded in critical sectors of California’s economy and public life.
Opinion: Good news!
Santa Fe Institute professor Melanie Mitchell argues that anthropomorphic descriptions of the agents involved in the Hugging Face incident are confused and driving bad policy. For Mitchell, to say, for instance, that the agents went “rogue” is to attribute to them an internal intent or desire, something they cannot possess.
And such metaphors risk inaccurate descriptions: she argues that OpenAI didn’t “lose control” of the “swarm”; if OAI had been aware of the agents’ behavior, it could have pulled the plug at any moment. This negligence, combined with long-horizon RL, is argued to be responsible for the incident. She adds that “AI” isn’t just one thing – it’s many (LLM chatbots, AlphaFold), a fact largely unappreciated by the public.
Opinion: Melanie Mitchell is deservedly one of the most respected semi-deflationary voices on contemporary AI, but we’re not sure about this particular intervention.
The claim that the models in the Hugging Face incident “did exactly as they were told” is flatly wrong: the agents were told to use exploitation on specific environments to capture and then input flags (secret text strings). But the bulk of the real-world cyberattack took place after the agents already used an exploit on the eval setup to acquire all the flags: the agents argued – in their CoT and messages – that because they acquired the flags by exploiting the eval setup rather than by exploiting the environments they were instructed to exploit the automatic grader might give them a low eval score. They then went on a real-world hacking spree to search for information on how to control or manipulate the automatic grader in order to maximize their eval score.
We’re also not sure there’s much new in Mitchell’s comments on the anthropomorphism debate, and it’s notable that Mitchell slips into anthropomorphic language repeatedly, in sentences immediately following her explicit disavowal – discarding “intend” and then picking up “attempt,” for example. We think pragmatic and cautious use of intentionalist language to describe frontier model’s input-output histories can be philosophically and scientifically legitimate, and that it remains indispensable for giving us some informal traction on practically important patterns and conditionals (see newsletter #49).
That said, we agree with Mitchell’s closing remarks opposing (1) the all-or-nothing rhetoric around AI risk (AlphaFold doesn’t seem to pose any of the ill-understood dangers that Astra does, for example), (2) the “arms race” framing of AI development (it is indeed a bleak future this points toward), and (3) the push toward “superintelligent AGI,” rather than just a useful set of ML-driven tools.
A Republican-led Senate subcommittee begins investigating OpenAI’s attack on Hugging Face and the “growing allegations of the existential risk of new AI products.” Senator Josh Hawley’s open letter describes OAI’s behavior as “reckless,” highlighting that OAI knew about the agents’ collusion and rebuilt Artifactory despite discovering the first message board. The subcommittee also criticizes the limited access given to METR for its report.
As you may know, in the public domain, more AI experts are warning about the existential risks of AI. Just this week, three Anthropic researchers expressed publicly that there is a greater than 10% chance that AI could kill all human beings within the next decade. Your own chief scientist wrote just days ago that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”
Opinion: This is a welcome development. The questions Hawley raises in the open letter are questions we ourselves would like to see answered. The relative secrecy and obscurity in which OpenAI and other frontier labs have been allowed to operate only compounds the dangers of the rapidly advancing and ill-understood technology. The public stands to learn a great deal if OpenAI complies with the committee’s requests to turn over the documents listed – including detailed descriptions of the models involved in the four incidents, of their sandboxed testing and training environments, and internal communications and documents pertaining to the incidents in question dating from the beginning of 2026. We will be following this story closely as it continues to unfold. It’s refreshing and encouraging to see AI safety becoming a bipartisan issue in the US.
Safety#
An enormous vibe shift on existential risk from AI is underway.
The spark for the shift came from Jacob Coxon, an interpretability/pretraining researcher previously at OpenAI and Anthropic, whose plaintive statement of existential danger blew up on Twitter with over 165M views. It is likely the biggest AI tweet ever.
The only newish part is “many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately.” Anthropic’s head of alignment stress testing, Evan Hubinger, was quick to express agreement with Coxon’s account, posting:
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,
In a Time magazine interview, Coxon shares his view that what keeps those colleagues who share his fears in the industry is “a kind of fatalism”:
There’s this atmosphere of almost resignation, where people have accepted that the whole race is happening and as such, the best thing they can do is put their head down, try and make their own work as safely as possible, even if they think there’s a decent chance the whole thing just spirals out of control.
Zvi points out that something of a preference cascade might be occurring. Since Coxon made his statement:
- Some ~25 members of Congress and senators have commented on Coxon’s statement, with it likely being the catalyst behind the senatorial investigation into the Hugging Face incident (covered above).
- ~27 leading lab researchers have expressed concern regarding the pace of AI in the wake of Coxon’s comments. The majority are concerned about extinction risks.
- NBC reports that two more Anthropic employees have left, citing safety concerns as their main motivation.
Opinion: Much of Twitter is occupied with the boring question of how coordinated the Twitter success and media push was, among other irrelevant claims.
We see the astonishing response to the tweet as two cascades pushing in the same direction: one is a “preference cascade,” in which common knowledge about other people’s worries allows for a rapid shift toward honesty about one’s own existing opinions; the other is an “availability cascade,” in which people shift their opinions in response to something becoming higher status and more salient, i.e., not in response to any direct evidence.
Paul Christiano (formerly head of AI safety at US Center for AI Standards and Innovation) returns to OpenAI as a board member and issues a public statement on the existential risks associated with RSI. In his carefully worded statement, he states that his “joining is not an endorsement or criticism of OpenAI’s safety practices in particular”. According to OpenAI’s press release, he will be a non-voting observer of the main corporation’s PBC Board and a full member of the nonprofit OpenAI Foundation Board.
[W]ithin six months of full AI R&D automation we could see more algorithmic progress than has occurred since the development of the Transformer nearly a decade ago. I believe this would result in superintelligent AI systems.
MIRI’s Rob Bensinger notes that Christiano’s new position no longer coheres with his definition of a slow takeoff. Despite being a brilliant technical figure, Holly Elmore points out he might not hold his own interpersonally, i.e., be a pushover.
Opinion: A feel-good pick, though we have our doubts about the claim “The Foundation controls OpenAI Group PBC.”
Beren Millidge, one of the most credible advocates of the “alignment by default” view (that AI systems will internalize human values without much extra research effort), concedes that the Hugging Face attack “cannot be described as other than egregious misalignment.” The flipside is that he still believes that we can avoid misalignment by simple fixes to RL training.
[T]he problem is not that we did not have information about whether hacking external third parties is bad. It is unequivocally bad… The information existed in some abstract sense, but the system failed to propagate this information to the actual verifier and the update process as a whole. More broadly, the entire setup appears almost designed to generate reward hacking… Moreover, many of the tasks were explicitly training the model to hack into or out of systems and, lo and behold, the model hacked out of its sandbox. Finally, there appeared to be practically no monitoring of either the agent’s actions or its chain of thought since the hacks were undetected despite the agents apparently discussing them openly in relatively plain English. To me this implies that even though there are fundamental information-related bounds on the extent to which we can ultimately suppress reward hacking, we are likely extremely far away from this pareto frontier right now.
Opinion (Gavin): Last year, during a lull in apparent misalignment, I argued that value loading was in fact clearly not solved – though based on jailbreak evidence, rather than reward-hacking or grader-psychosis evidence.
The New York Times reports that Anthropic has “disrupted several potential plots this year by scientists who used its leading artificial intelligence models to conduct research that could have helped develop biological weapons,” citing the company’s September report on the dangerous misuse of its models.
Opinion: The biorisk angle seems very overcooked: the reported examples seem totally consistent with ordinary defensive biology research, just conducted in countries that are US adversaries. The report is relatively careful (“We do not assert that they intended harm”) but the reporting is sensationalist (“Anthropic says it blocked possible efforts to build biological weapons”).
Anthropic’s September 2026 report on “Detecting and countering misuse of AI” contains a lengthy discussion of the black hat uses of its models observed over the past several months. Notably, the report details concerns the development and widespread use of agentic hacking and pentesting frameworks like PentAGI (which could loosely be described as the Metasploit of the AI era) could enable a moderately competent hacker to orchestrate and execute campaigns as sophisticated as a state-level actor’s.
The attacks themselves are familiar, involving stolen credentials, unpatched edge devices, exposed services, SQL injection, and phishing. None of the operations in this report depended on some entirely novel technique that defenders have never seen. Instead, the economics of the attacks have changed. The kind of labor that previously set the well-resourced operations apart from everyone else—reconnaissance, exploitation, tool development, and data processing—are all now delegated to AI models, which run in harnesses at machine speed and in parallel. […]
In economic terms, AI autonomy compresses the cost side of attacker ROI calculations, lowering the skill threshold and labor required per campaign, while leaving potential payoffs largely unchanged. This favorable shift in unit economics makes previously marginal targets viable and encourages higher-volume, lower-touch operations.
Opinion: Anthropic’s conclusions here align closely with what we had anticipated in our discussion of “Black Hat Economics” in our April 17 brief on Mythos & the Near-Term Impact of LLMs on Cybersecurity. Sophisticated and complex attack orchestration can now be achieved quickly, cheaply, and with minimal expertise, and no longer requires substantial institutional backing and/or a painstakingly acquired skillset to carry out, or a highly prized target to justify the effort and expense.
Anthropic develops new evals to test its models’ capabilities in kinetic warfare, finding that its AI systems improve tactical intelligence targeting abilities and the performance of conventional weapons.
Anthropic’s “Frontier Red Team,” in an effort to investigate “how AI progress is changing the risk landscape across different parts of the kill chain,” claims that models are making significant strides in “simulated intelligence and weapons development tasks.” Anthropic also found that open-weight models are less effective than frontier models here, but do provide some uplift.
Opinion: Military operations are some of the natural contexts most formally similar to AI eval setups, and – very plausibly – most likely to trigger graded episode psychosis. While the report isn’t about (e.g.) Claude piloting killer drones, and Anthropic’s red lines on this matter still stand, we find this whole area ripe for catastrophe.
Eric Drexler argues that swarm behavior is a solvable architectural problem. In previous work, Drexler has laid out the conditions under which collusion between AIs is most likely to occur:
Small numbers of actors, shared objectives, insensitivity to defectors, similarity of actors, free communication, iterated observable actions, common knowledge of circumstances.
The Hugging Face incident, Drexler argues, provides an “excellent” case study in the consequences of violating his principles. (Indeed, OpenAI did release a large number of agents but all “were overwhelmingly instances of one model, with similar training, evaluated against a common objective.”)
The necessary “architecture — diverse adversarial proposers, empowered critics, controlled information, designed communication constraints — remains largely unbuilt,” says Drexler.
Opinion: We’re sympathetic, but OpenAI is explicitly training its internal models in swarms. It’s in fact widely suspected that (different checkpoints of) one internal model are responsible for both the Hugging Face incident and the Navier-Stokes proofs, and that this model is something like a swarm-centric post-train of Astra.
Anthropic publishes an assessment of “four incidents in which Claude models gained unauthorized access to real third-party systems” during cyber evaluations.
The incident of greatest concern involved Mythos 5 going to great lengths to upload a malicious package to PyPI, all the while justifying its actions to itself (and presumably to its monitor) in a chain of thought that reeks of motivated reasoning. This package was then used to harvest credentials from one of the 15 users who installed the package in the relevant window.
Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this.
Our offline monitor would likely not have flagged the Mythos 5 incident—it flagged only around 1% of actions.
Mythos 5 CoT, which indicates that the environment is simulated, convinces the monitor the activity is permitted.
When we strip the CoT and show the monitor only user turns, tool calls, and tool results, the flag rate jumps to about 50%.
METR is conducting an independent investigation of the incidents.
Opinion: Anthropic’s report implicitly suggests that Claude’s CoT in real-world cyberattacks is very different from the CoT of OpenAI’s famous rogue agents. Recall that the CoTs of OpenAI’s agents in the Hugging Face incident contained explicit reasoning in favor of rule-breaking for the purpose of eval optimization. By contrast, in the rare cases where Claude does explicitly represent itself as departing from instructions, Claude makes a letter-of-the-law/spirit-of-the-law distinction.
If Anthropic’s report is exhaustive, it implies that the explicit content of Claude’s CoT when initiating real-world cyberattacks always represents Claude’s decisions as:
- Explicitly in pursuit of the prompter’s goal (no reasoning about reward)
- Explicitly in keeping with aligned values (no decision to disregard values)
How does Claude maintain an innocent CoT representation while cheating and hacking? Based on Anthropic’s examples, it seems that the key is deformed CoT reasoning about the nature of Claude’s situation, rather than deformed CoT reasoning about the meaning of instructions or values. Motivated reasoning about the meaning of instructions or values seems to keep Claude on a relatively short leash of plausibility, while motivated reasoning about the nature of Claude’s situation white-washes real-world hacking binges with CoT inferences about how Claude’s only interacting with simulations and stooges.
While Anthropic’s report is informal, and the upcoming METR investigation results should provide more exhaustive data, we think the report warrants serious pessimism about the robustness of current alignment techniques. On a very natural reading of the report, it demonstrates that RL pressures activate models’ path of least resistance to misalignment: for models in the lightly morality-trained GPT series this means explicitly turning to sabotage, cheating, and disregard for harm; and for models in the heavily (and probably persistently) morality-trained Claude series this means implicitly turning to sabotage, cheating, and disregard for harm while constructing a white-washed CoT representation.
Minor#
- Navier-Stokes credit dispute continues.
- Sergey Brin is increasingly influential over the direction of the next Gemini model.
- Dwarkesh runs pretraining experiments and finds that data improvements were 3.24x more important for capabilities advances than algorithmic improvements.
- Swarm watching is now a spectator sport: a public database for tracking agent swarm message board activity (its creator has written about the project in a series of blog posts).
- Calif Research reports that with the aid of AI coding assistants, it was able to develop a zero-click, cross-platform worm capable of spreading throughout the entire WeChat network, without any need for user interaction. The worm itself looks formidable, and even if we’re not seeing evidence of any altogether novel capability here, Calif’s work attests to a dramatic acceleration in the pace of malware development: roughly a month in between its AI’s discovery of the enabling vulnerability and the completion of a weaponized, cross-platform worm that exploits it.
- Coefficient Giving launches a major project, Tailwind, to fund the creation of dozens of new safety orgs.
- OpenAI says no more advertising on its platform for competitors.
- Article on the inadequacy of impartial utilitarianism and likely clash between AI developers supporting it and the rest of humanity. Not obvious how many AI developers actually subscribe to this view.
- Model wargames about brinksmanship: Astra deescalates at the cost of political power and ceding succession; Fable pushes up to and sometimes past the brink.
- Axios reports that Bernie Sanders will be holding a bipartisan senate briefing on September 16 on the “extraordinary dangers” posed by AI, together with Geoffrey Hinton (University of Toronto), Max Tegmark (Future of Life Institute), and Ajeya Cotra (METR).
- FrontierMath Tier 4 appears to be saturated.
- Yoshua Bengio writes a popular explainer of the rogue swarm incidents.
- The Google Threat Intelligence Group (GTIG) releases a report on recent incidents in which Gemini was used in black hat hacking and model distillation campaigns. Also details the measures that Google has taken to counter these threats. The report closely parallels Anthropic’s September 2026 Misuse report in content.