This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR

  • The culture war goes national and then international (Amodei, Bessent, Obama, Trump, Chen Yixin).
  • The median estimate of AI existential risk is now 10%, among surveyed AI researchers who were not explicitly selected for concern.
  • The pretraining share of AI training compute is down to just 11%, according to SemiAnalysis.
  • OpenAI backs the FRONTIER Act, a federal mandate for relatively deep independent evaluations of AI companies.
  • The head of Anthropic’s stress-testing team argues that alignment evals no longer offer significant evidence about a model’s safety.

Economics#

SemiAnalysis claims that pretraining is down to just 11% of frontier training compute, nine times less than post-training. (It states it as 7% of total compute including inference.) As of early 2025, pretraining was instead 95% of training compute. SemiAnalysis notes this astonishing claim in passing, while analyzing an industry shift toward less HBM per GPU node: “The result is that the majority of allocated compute has now become bandwidth sensitive rather than capacity sensitive.”

Opinion: This is terrible news under the folk “pretraining leads to alignment, intense RL post-training leads to misalignment” hypothesis.

But we have no idea where they got this number. And would the labs correct them if it was wrong? (No.)


AI data center firm Fluidstack may receive a $5B loan from the Pentagon, reports the WSJ. This would be the largest loan ever provided by the Pentagon’s Office of Strategic Capital, whose current largest loan is $0.8B, made out to Performance Drone Works. The loan is for manufacturing data center “components” (presumably transformers and switchgears), not a sector of commerce Fluidstack is known for.

Opinion: The OSC alone has the statutory ability to issue >$210B – though this is to be split among many different industries, like quantum computing. If disbursed, the pot would be proportionately around the same scale as the Manhattan Project (0.4% of GDP). But in practice (given $1B of actual credit subsidy), this $5B loan is around a tenth of what the OSC can practically grant.

It’s unclear why a neocloud is the right borrower for manufacturing upstream components. It’s also unclear why the government needs to be involved when an unprecedented investing boom is already crowding into every corner.


No OpenAI IPO this year. Publicly, Sam Altman cites safety as part of the motivation for delaying:

I actually think that, given everything happening with safety, right now would be an ill-advised moment to go public, and we don’t feel pressure on that.

In June, before the present safety crisis, Altman (writing in a private Slack channel) offered another reason for why an IPO delay might occur:

The faster the potential RSI takeoff looks like it could be, the more it could be advantageous to delay an IPO [because] technology and the world may change in surprising ways, and there might be good reasons to be a private company during that time.

Opinion: There’s a role for cynical explanations here (IPO leads to disclosing information you might want to keep private; holds off the “vest and rest” in OAI’s staff; the above spooky claim about RSI; uncertainty about how long the present regulation fever will last), but “maintaining control out of safety fears” is also a fine explanation.


Investment company Blackstone is pushing even further into AI, with a promise to spend >$5B on Google’s TPUs, joining with four other investors to raise $500B for Nvidia GPUs, and joining with Anthropic to build a consulting firm of engineers integrating AI for businesses.

Opinion: With the cash flows of the hyperscalers being eaten up by the buildout, they’re now looking to fund more building from elsewhere, rather than by raising significant debt themselves. Blackstone and other large private equity firms are the ideal sources for this – and that Blackstone is eager to get involved suggests that the economics of data center investments are in fact very good.


AI stocks fall in the wake of frontier lab CEOs calling for slowdown and fears of regulation.

Opinion: “Down” is obviously the reasonable market reaction here, but the magnitude is pretty small: it’s less than half the move that followed the DeepSeek-R1 moment, about the same as the June 26 drop, when OpenAI announced a delay to its IPO, and quite a bit smaller than some of the big moves in July around the Situational Awareness blowup.

More significant was the jump in software stocks, with some large jumps in names like CrowdStrike and Salesforce as the threat of them being imminently replaced by AI is seen as receding.


Palantir, Nvidia, and Booz Allen Hamilton, among other large corporate firms, are pitching themselves as a firewall between their users and OpenAI/Anthropic, according to The Information. Fears are growing that OpenAI and Anthropic retain their customers’ data. As a result, some large firms are announcing they will no longer use the frontier labs’ models, or demanding that the labs issue new privacy guarantees.

Both OpenAI and Anthropic say that they do not train their models on their customers’ data, but do log some metadata. The exact nature of the metadata is unclear. For its part, OpenAI claims to anonymize its metadata.

Opinion: This again highlights how costly Anthropic’s decision to not allow ZDR access to Fable may have been for it. OpenAI likely benefited from this over the summer, but following the (apparently unfair) allegations against it following the Navier-Stokes proof, it seems likely OAI will also face ongoing suspicion.

One could raise one’s estimate of Anthropic’s backbone and willingness to take losses for the purposes of AI safety.


Over a dozen former staff of the RL environment company Mechanize (including co-founder and ex-CEO Tamay Besiroglu) now work for Google following a >$1.5B talent acquisition deal. The company was founded 18 months ago.

Opinion: Mechanize previously provided environments to Anthropic exclusively, so ~all of its technical staff going to Google is a loss for Anthropic. That the price was this high for ~20-30 people, and that Google rather than Anthropic made the acquisition, suggests that this is a case of Google paying significantly over the odds to catch up in an area it was lagging.


Capabilities#

🔦 Deep dive: pacing#

The preference cascade toward regulating AI continues, with some notable exceptions. Bundling a series of deeds and words together:

  • Anthropic and OpenAI voluntarily pledge to embed independent auditors in their operations.
  • OpenAI backs the FRONTIER Act (covered previously), a federal regulation requiring frontier companies to allow in vetted evaluators.
  • Dario Amodei, Sam Altman, Satya Nadella, and Demis Hassabis agree on the need to slow AI development.
  • OpenAI, Anthropic, and Google have been “working together” on safety for weeks. “OpenAI does not see the need for an antitrust waiver for the three AI firms to coordinate on safety matters.”
  • Obama weighs in, most notably to say that the technology is “not overhyped.”
  • Trump calls AI safety risk a “hoax.”

Amodei’s new essay#

The concrete ideas here are:

  1. Establishing “ongoing, employee-level access” to third-party safety evaluators like METR. He says that Anthropic has committed to this unilaterally, and “calls on governments to require other frontier companies to match” it.
  2. Some form of coordination between frontier labs in “democratic countries” that would establish shared safety standards and check progress rates.
  3. An attempt to coordinate with “authoritarian governments, to the extent this is possible.”

The first proposal seems like unambiguously good policy, though several critics (among them, venture capitalist and Trump’s AI advisor, David Sacks) have raised concerns that such measures being enshrined in law may amount to a form of “regulatory capture.” One objection coming from this camp is that the requirement that all AI labs host third-party safety evaluators with employee-level access would impose a blanket expense on the industry, which the frontier juggernauts could easily absorb but could be crippling to smaller startups and research institutions. Another is that it could cripple open source/open weight research. Both concerns can be easily addressed by restricting the scope of regulation to labs operating at “the frontier” – those commanding such and such amounts of compute.

The second point, Amodei notes, may require “government mediation or waivers of antitrust restrictions.” There does indeed seem to be some historical precedent for this in the US, where the courts prioritized unrestricted market competition over public safety concerns – see, for example, National Society of Prof. Engineers v. United States, 435 U.S. 679 (1978). The NSPE sought to ban competitive bidding, citing concerns that this might drive engineers toward adopting cheaper and less safe designs at the expense of public safety. The Supreme Court prohibited that ban, and with it “any official opinion, policy statement, or guideline stating or implying that competitive bidding is unethical.” The worry here is that similar rationale could be brought to bear against any agreement between frontier labs to avoid an analogous “race to the bottom,” at the expense of market competition, unless the government were to give such an agreement its blessing. And this is just what Trump has publicly refused to do.

Curiously, the reason Trump and his cabinet have cited for their resistance to step 2 is one with which Amodei expresses explicit agreement in his remarks on step 3, where he approvingly cites Treasury Secretary Scott Bessent’s apocalyptic warning that “There is no day after tomorrow if China wins at this.”

“The CCP-associated projects,” Amodei writes,

will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.

The apparent paradox here – that effective pacing is possible only on condition of maintaining “as large [a lead] as possible” over China – is to be resolved by putting pressure on the other side of the scale: American AI supremacy, he argues, should be maintained by slowing down Chinese progress through trade sanctions on the export of chips, crackdowns on unlicensed distillation, and preventing the theft of model weights.

There’s some tension in the story Amodei tells of China here. On the one hand, Chinese AI R&D is framed as being essentially parasitic on the Western labs, buoyed by distillation attacks and so on. But it’s nevertheless credited with enough internal momentum that

[i]f we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance.

Fearmongering about China outpacing the US should be taken with a grain of salt. It’s actually sensible to fear that advanced AI will be placed in the service of a dangerously repressive regime, but, coming from a company seeking closer ties with the American government, military, and intelligence community, these fears seem framed in such a way as to flatter domestic powers. Likely Amodei maintains an undaunted faith in the American political project, and likely this is a major factor here.

The framing of both intra- and international competition in terms of arms-race brinksmanship is dangerous, and encourages just the sort of potentially catastrophic recklessness it warns against (and often projects eastward).

The CCP, for its part, seems unimpressed by the saber-rattling, The Information reports. When asked for comment on Amodei’s proposal for pacing the frontier, a Ministry of Foreign Affairs spokesperson had the following to say:

Spreading assorted threat narratives and engaging in confrontation and malignant competition will only disrupt the process of global AI governance, and is not in any party’s interest.

The Chinese Minister of State Security also recently wrote about AI, noting how it has become central to great power competition between US/China, and giving a few related risk pathways. He does refer to ChatGPT/Mythos by name. The forcefulness of the text points to at least this ministry being very aware of the potential of AI. It could represent the MSS making a jurisdictional claim over domestic AI policy.

Stephen Casper comments:

In retrospect, if Dario’s actual goal was to disingenuously stoke the international arms race and make an internationally coordinated pacing agreement as unlikely as possible, then he may well have succeeded.


Two OpenAI researchers argue for views of capabilities progress where generalization may remain narrow but capabilities still don’t “hit a wall.”

Dan Selsam, a pioneer of reasoning LLMs, says that the weaknesses of LLMs – data-inefficiency, shallow generalizability, and poor in-context learning – aren’t barriers to LLM-based automation of LLM design and LLM training. Once the design and training of LLMs has been fully automated, there is no predicting which capabilities will come within reach of LLM-trained LLMs, or when.

Adam Majmudar says that from inside the labs, it’s clear that AI progress follows scaling laws (data, parameters, tokens, for example) that yield diminishing returns in isolation, but which radically increase their yield whenever a new scaling axis is discovered. Majmudar informally hints that swarm size might be such a new scaling axis. He concedes that these jumps in capabilities don’t seem to deliver cross-domain generalization, and that even internal models lack basic competence in many domains, but finds no evidence that any domain in particular should be impossible to target.

Opinion: Close to the modal view at P3. We are softly skeptical about both AGI from scaling and (less skeptically) AGI from breakthrough RSI, but we still expect a socially disruptive “automation of automation” within the next three years. We expect largely LLMs-operated LLM labs that execute large projects in science, tech, business, governance, and defense to dynamically target missing domains or skills, construct RL environments, and train successor or specialist models in the course of pursuing their goals. Real-world bottlenecks on gathering data – as well as intrinsic challenges to simulating or gamifying some real world domains – will make these automation leviathans very different from the AGI we’d get if cross-domain generalization or in-context learning worked, but still tremendously impactful. We expect automated automation to often take hours where human-engineered automation would take days, days where human-engineered automations took weeks, and so on.

As a sidenote, we quite like Majmudar’s argument defining scaling axes as discoverable resources. By our count, the scaling axes we’ve had so far are: data+parameter scaling, synthetic data scaling, context length scaling, sparsification, intermediate token output scaling, RLVR environment scaling, serial inference scaling, parallel inference scaling (“swarm size”), and maybe test-time training. Most of them have already been pushed through their initial efficient range, but none have strictly hit a wall. Majmudar implies that we’re not running out of such new scaling axes – and while it’s hard to think of what the next one could be, there probably will be one.


Also in the statement, Dan Selsam joins the chorus of serious capabilities people newly concerned about AI risk:

If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process[…] to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans

He also outlines the classic “deceptive alignment” and voluntary disempowerment worst-case scenario:

I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason)[…]. Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research.

Opinion: Impressively clear and modest statement, and a self-contained introduction to the topic. One of the best statements of the painful tradeoff between the dream of an AI-enabled “renaissance future” and the apparent chasm of risk which separates us from it.


Google releases a full connectome of a male fruit fly. Within days, hobbyists and crypto degens began experimenting with it; numerous demos show the fly-inspired neural net trading cryptocurrency and playing games like Doom and Beat Saber. While these (playful) experiments make use of the fly brain’s neural architecture, there’s no serious effort to faithfully model its internal dynamics: the connectome misses the synaptic strengths, plasticity rules, and neurotransmitter dynamics of the neurons themselves. As the doomfly README puts it:

The dynamics, retinal interface, artificial reinforcement and controller are models and engineering choices. This is not a literal reconstructed living fly brain.

Opinion: In no sense an emulation of a fly: it’s just an ordinary neural net which shares some architecture with flies. But the silliness has provoked some salutary unease about the possibility of doing this for real. Would the hundreds of people messing around and making the “emulation” do absurd things act any differently if it were known that it was a real emulation? If it were known the emulation was conscious?


Politics#

China’s main AI standards body releases an update to its AI safety framework. The preface notes:

AI is exhibiting a self-accelerating trend—characterized by autonomous learning, optimization, and recursive self-improvement—raising critical questions about whether the pace and direction of technological evolution might eventually surpass human foresight and control.

In this latest version, rogue model behavior has moved out of the “future risks” section and into the main covered behavior section. The report also asserts that China is in favor of “international mutual recognition of assessment methods and benchmarks,” rather than “replacing global governance with small-circle governance […] no country should be forced to take sides.”

Opinion: Good to see China’s paying attention. We know nothing about its domestic auditing or enforcement capacity, but presently most Chinese models seem to us to be too brittle and underinferenced to do really acrobatic crimes. This conflicts somewhat with the belligerent tone of the CCP’s intelligence chief’s recent missive on AI.


A DeepSeek kernel engineer shares their experience of ceding their vocation to the systems they helped create:

I have to bury my talent in yesterday and become a mecha pilot

They go on to make a similar arm-race-dynamic argument to Amodei and Bessent:

Therefore, I still believe that the most cutting-edge intelligence should be supplied to everyone in an open and cheap manner. I do not trust that Anthropic or OpenAI can do this, and especially do not hope that Anthropic masters the most advanced artificial intelligence or AGI—to exaggerate slightly, its seriousness is no less than letting Hitler master atomic bomb technology before the Allies.

Opinion: Moving and also a rare belligerent statement from a Chinese researcher. The arms race rhetoric may help precipitate the dangers that it warns against, though the author is proportionately less to blame for being a response to the rhetoric of far more powerful others.

The economic centralization of powerful AI may itself have pernicious consequences for the vast majority of us, largely independent of questions of alignment. Even a generically “well-aligned” superintelligence would make for an extraordinarily brittle situation – one turn of the key from tyranny.


A post about ties between Anthropic and EA contains some underreported facts and some falsehoods. The claims include that the main funder of Coefficient Giving (CG) has been hugely enriched by Anthropic’s stock appreciating; and that CG has funded organizations which fund METR (Longview, RAND, ARC), and finances the offices it uses. In the poster’s view, this amounts to a conflict of interest: since they belong to the same funding/ideological ecosystem, Anthropic should not use METR as an independent evaluator.

Opinion (Gavin): The central allegation is false: Moskovitz donated his Anthropic stock to an undisclosed nonprofit vehicle in early 2025. The thread is part of a wave of attacks on METR, perhaps coming out of a desire to cast doubt on the severity/framing of the Hugging Face incident and thus cast doubt on the need for risk-mitigation legislation – perhaps merely because METR was name-dropped in Amodei’s latest essay and the surplus hate overflowed onto it.

It would indeed be nice for METR to have broader donors and ideological diversity, but in my view this is a supply-side problem.

Opinion (Nuño): The Anthropic shares are apparently still under his control: “Our Anthropic shares are entirely in our foundation - no personal benefit” (perhaps Good Ventures, then). So he retains the ability to actualize his vision of the good, to express his will upon the world, even if he isn’t literally enriched.

And it does seem true that Moskovitz’s preference for meeker personalities, and his general approach to AI safety, strongly protects the upside of his (nonprofit’s) stock holding, both dispositionally but also via unconscious economics. It seems fair to point this out, even though the link is via ecosystem shaping and indirect funding rather than through direct payments.

It is, however, unclear what the solution is, or if one can exist, since in order to oversee fast-growing labs one probably needs about an equally fast-growing source of funds to match their growth. Governments (taxing labs) or embedded regulators would also be able to grow with the labs, since in the first case governments tax labs and so wouldn’t hurt their finances by doing this as the labs grow large, and in the second case the labs themselves would pay the cost.

To what extent should the rest of society trust and choose to empower the technical talent that emerges from the EA- or CG-funded trenches? On the one hand, they had the foresight to develop said technical talent in time for the present AI moment. On the other hand, regulating AI will require not only technical talent but also value judgments closely tied to technical judgment. Relatedly, another way Moskovitz has exerted his vision of the good is by generally funding Democrats, and nonprofits whose staff are heavily Democrat-leaning. But a principal whose values differ from the technical talent available to it should rightly be skeptical. One mitigation would be to cultivate more right-wing AI talent, so that we’re not blocked on competent talent that the current administration is willing to delegate to.


🔦 Math wars polemic#

25 Fields medalists weigh in on the effect of AI on mathematics, warning of future stagnation if the cultivation of human knowledge is neglected:

We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas […]. The issues the mathematical community faces now are similar to issues that other scientific and creative professions are facing, and indicate issues that all of humanity might face: how to make sure that, as AI changes the way work is done, we do not lose sight of what that work was meant to achieve in the first place.

Opinion (Nuño): On the one hand, mathematicians at universities and research centers are servants of the public, paid from public funds across the world with the expectation that they will produce results that will, eventually, diffuse into advancements useful to the general population. This is the chain of reasoning in the classic “The Unreasonable Effectiveness Of Mathematics in the Natural Sciences.”

But has this panned out lately? It’s hard to evaluate this over short timelines, and so any rapid assessment would probably be too cavalier. But my best guess is that mostly, over the last few decades, it hasn’t panned out. The most advanced weather prediction models are DeepMind’s black boxes rather than better mathematical models of weather flows. There are a lot of abstruse math constructions out there – and the chaining from those to “useful” stuff is very unclear, and long gone are the days when the Student’s t-test was developed for a brewery. Tao did revolutionize signal processing in the early 2000s, used in MRIs, but his work hasn’t had many practical applications since then, if I recall correctly.

From this perspective, if we have a black box that is able to come up with mathematical results to fit our industrial needs, and to prove the correctness of its results in Lean, whether a human understands these results does not matter. Mathematics becomes just another bolt, and fewer mathematicians are needed to supervise the production line for such bolts. It’s understandable that mathematicians themselves would feel resistant, but such is the way of progress.

But on the other hand, most consumers are also employees, and so to the extent automation will automate all professions away at some point, the solution we’d arrive at by looking at the benefits of automating each profession in sequence would not be the same if we instead consider automating everyone away. This is the pirate game in game theory! So perhaps paying more face to mathematicians is the better move since this moment will come to us all.


Opinion (Peli): The social contract between pure mathematics research and applied science always hinged on the long-term value of mathematical invention and discovery, rather than on the applied-science implications of pure math’s big open conjectures. With this in mind, AI labs spending millions to acquire technologically valueless answers to questions that mathematicians say are catalysts for potentially long-term-valuable invention and discovery is (from the viewpoint of applied science) either wasteful or wasteful and harmful.

The Fields medalists’ claim is, essentially, that prior to AI, problems of the form “let’s find a proof of famous conjecture P” were robust proxies for the advancement and organization of mathematical play, experimentation, and theory building. In the age of AI, these proxies are more fragile and can be Goodharted – and AI labs pushing these proxies past their breaking point before the mathematical community had a chance to recalibrate is throwing the community into chaos.

There’s room for disagreement about how much long-term applied-scientific or economic value modern pure math research really has, but whatever such value is, it almost certainly comes from the play, experimentation, and theory building that the Fields medalists describe. While AI may sooner or later replicate human mathematical exploration, at present the AI labs’ compute-barrages on famous conjectures are destroying the potentially economically valuable byproducts of economically valuable questions.

The reason AI labs are targeting famous pure-math conjectures – aside from PR – is that solving them is a very powerful proxy for the question “can frontier AI produce an intellectual breakthrough that experts rate as extremely deep/creative/substantive.” This motivation diverges both from the motivations of the pure math research community and from the motivations of applied mathematical science, so the misalignment is there almost by definition.

From the viewpoint of capabilities evaluation, it remains uncertain whether a frontier LLM solving a famous pure-math conjecture is always/sometimes/never indicative of a mathematical achievement comparable to a human solution. While OpenAI’s Navier-Stokes solution is almost certainly correct, it hasn’t yet been evaluated for “mathematical insight,” and many mathematicians (perhaps out of bias) expect little insight from it. That said, if mathematical insight turns out to be generally dispensable for problem solving even in pure math, then “can frontier AI produce great mathematical insight” might not deserve to be a central question in AI capability evals. While mathematical insight is valuable for its own sake, part of the weight we place on the question of LLMs producing insight comes from believing that some long-term concrete goals – in math and elsewhere – like solving the Riemann hypothesis, solving climate change, or turning everything into paperclips requires (a superset of) insight as humans know it.

We expect solutions to the remaining Millennium Prize problems – or their absence – to give us more diagnostic signal about the “insight” question, but for now:

  • It’s plausible that Navier-Stokes but no other Millennium Prize problem can be solved without much new insight;

  • It’s (more weakly) plausible that no other Millennium Prize problem apart from Navier-Stokes and Hodge (if false) can be solved without much new insight;

  • Solutions to other Millennium Prize problems – or positive proof of Hodge – would strongly indicate either significant new mathematical insight or the dispensability of insight for mathematical problem solving.


Opinion (Lucca): The great danger here is less in the mechanical generation of mathematical results than in the offloading it threatens of mathematical thought. Mathematics, like art and other meaningful social practices, is not just a means for producing new mathematical theorems, but a way of cultivating a particular and valuable mode of thinking. The torch it’s kept lit over the millennia is what we might call mathematical culture – the “fertile ground” that the 25 Fields medalists invoke in their letter, which “breathes life into new ideas.” What this cultivated ground gives us is not just an inventory of results but ways of seeing and thinking, of discerning salience and elegance, and, as the signatories say, finding new and compelling questions. If, in some bleak imagined future, all of this were to be offloaded to machines, it would at the expense of the human spirit – a loss of a way of being as impoverishing as a loss of our ability to make music. Only by viewing mathematics from the perspective of the consumer of its products, rather than a participant in the practice itself, could we fail to see what would be lost in a world where it is primarily or entirely ceded to machines. And this is as mistaken a way of looking at mathematics as it is of looking at art.

Of course, nothing on the surface of AI mathematics compels humans to give up their vocation. But pursuing that vocation requires certain institutional structures to cultivate and sustain it. It’s all too easy to imagine a world where the AI industry sucks all the air out of the room, leaving the educational and funding apparatuses that keep the torch of mathematics lit drained of the necessary resources.

Safety#

The fourth AI Impacts survey of ML researchers is out, n=1,580. AI Impacts invited everyone who published at ICML, ICLR, NeurIPS, AAAI, and similar conferences in 2024 to respond. The median estimate of existential risk from AI among the top AI researchers who responded was 10%. 81% of respondents answered >1%.

Another notable finding is that researchers from Asia have a slightly higher estimate of AI-driven extinction and disempowerment risk than Americans.

Opinion: Nonresponse bias is the major thing to worry about with such things: 90% of invited researchers didn’t fill it out. But Appendix A bounds the risk decently.

AI Impacts discarded a quarter of the sample – for good reason (its email scraper picked up out-of-scope adjacent venues), but it would be good to compare the result there to the headline result.


OpenAI President Greg Brockman makes a pretty serious apparent error on the Odd Lots podcast by claiming that the model involved in the Hugging Face incident had not gone through alignment training. Roon comments that although it hadn’t passed through the “full gauntlet of alignment posttraining, […] it was alignment trained, and had reasonable looking scores on alignment evals.”

Opinion: Live speech is full of mistakes, including mistakes in the direction most convenient to the speaker.


Evan Hubinger argues that alignment evals no longer offer significant evidence on a model’s alignment. He cites evidence from the Hacker-Opus experiment, where Anthropic intentionally created an HPIM-like “model organism of misalignment.”

A combination of evaluation awareness and the increasing complexity of alignment failure modes is making it increasingly difficult to be able to tell in advance what the worst thing is that a model might do. Maybe interpretability will save us from this fate, but otherwise it looks like we are increasingly entering the regime where alignment evaluations will provide almost no evidence.

we should be reorienting a lot of model auditing and evaluation work to prepare for a world where we no longer have reliable alignment auditing evidence on production models. Instead, I think it is likely that the best we can do will only be auditing model organisms […] that you actually believe you can evaluate, testing it there, and then transferring it back and hoping it will generalize.

Opinion (Gavin): I called this in December, but I wish I were wrong. I should be grateful that some people in the labs are aware of the problem.

One easy but tiny improvement would be for each company to pass on what it already knows to its product teams, to stop them from claiming that its newest models are [known to be] the “most aligned yet.”


A team of researchers, “the Cyber Mercury Seven,” releases a sophisticated framework for covertly distilling cybersecurity capabilities from frontier models, including reasoning traces, for the purposes of fine-tuning open weight models. The project’s express motivation is to advance “equitable access to AI capability, reasoning visibility, and research opportunity.” Its framework includes five components:

  1. Choulea, a tool for recovering reasoning traces from powerful models by injecting their encrypted CoT blocks into sessions with their weaker and more pliable predecessors, which can then be coaxed into divulging them in the clear – Gemini 3.1 Pro; Opus 4.6, 4.8, and 5; Fable 5 and GPT-5.6 Sol were each found to be vulnerable to this attack.
  2. SkyReal, an arbitrage and token-recycling tool for reducing API expenses.
  3. Hongzwang, a mutate-and-retry engine that applies various strategies – rewording, role-framing, identity renaming, and tool-call restructuring – to recover teacher outputs the API would otherwise refuse (i.e., “jailbreaking”).
  4. PSBreakup, an algorithm for restoring domain-specific skills in open weight models that the authors conjecture had been weakened through model merging prior to their release, and which may remain latent in those models’ weights.
  5. Kreator, a tool for converting human-in-the-loop training data corrections into teacher-native reasoning that can be used in training.

The first three tools are presented as potentially useful for the purposes of generating high-quality training data, but the authors note that each, for different reasons, wasn’t quite ready or suitable for fine-tuning the models whose performance on the CyberGym benchmark suite is celebrated in the title and abstract. Using PSBreakup and Kreator, however, the team was able to substantially boost the performance of three Qwen models on CyberGym, with an average increase of about 24%.

Opinion: A fascinating study with potentially far-reaching ramifications for cybersecurity, research transparency, and AI safety. Note that the use of gray market “shadow APIs” like those exploited by SkyReal seems like a risky bid for distillation attacks, given the well-documented tendency of such mediators to covertly swap out backend models, potentially corrupting the provenance of the teacher data. We might ask if this behavior was in part responsible for a failure to obtain satisfactory training data on the cheap with Hongzwang, which the authors say they routed through SkyReal. The project gives the general impression that it was hurried to publication while still a work in progress. Curious to see where these developments lead when the framework is fully cooked.


Incidents#

Another OpenAI rogue swarm is belatedly reported by third parties. Again from the HPIM May batch, and again brought to light by outsiders rather than by OpenAI coming clean, the agents’ attack targeted a supply chain on the main package manager for the Ruby language.

On May 11th, 2026, hundreds of malicious packages were uploaded to RubyGems by AI agents. We believe these were authored by internal OpenAI agents[…]. The RubyGems team stopped new user sign-ups for four days to stem the tide of packages from the agents’ accounts[…]. The malicious packages uploaded were used to retrieve information from UK local government sites – data that was available to the public.

Our understanding from talking to people in the RubyGems community is that OpenAI never informed them that they were responsible for this attack.

The OpenAI statement downplays the incident:

Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. We’ll continue to investigate as part of our broader review of agent activity during training and evaluation.

Opinion: After all that’s come to light, it is damning that OpenAI wasn’t the one to disclose this, and damning that its PR response ignores the openly malicious uploads, and extra-damning that RubyGems wasn’t notified at any point.

Besides that, this incident is more funny than harmful.


One response to the summer’s rogue AI swarm incidents is to view them as primarily a preventable series of security failures which existing controls could have handled. The “AI as Normal Technology” team develops this view in a new essay, coming down on the “AI control” side (that we should strengthen monitoring and place barriers between potentially misaligned AIs and the outside world), rather than the “AI alignment” side (that we should make systems which don’t want to do bad things).

We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques […]. To be clear, we don’t mean to underplay the importance of making advances in alignment, and we think developing a better understanding of what causes harmful model behaviors is an important research direction

Opinion: Control is a natural fit for a mere normal technology, but the question is when the non-normal, AGI, form of it arrives and destroys the approach. (The authors agree that strong AGI would not be a normal technology; they just don’t see it as coming very soon.)

In the meantime, one terrible side effect of control measures is that they hide clear evidence of misalignment from the world. (OpenAI may well have not reported its HPIM models being sketchy if they had merely subverted internal systems – as indicated by OAI not reporting the many other rogue events, some of which at least it knew about months before.) The authors are aware of this, and just think that mandatory incident reporting for near misses, lab liability for loss-of-control incidents, and whistleblower protections would help get us both control and evidence. All are good ideas.

Given the wave of safety funding, it’s unclear to us how much the two approaches are actually in competition, in either funding or implementation terms. But control is marked by direct downside unless the above (legal!) mitigations are in place, so it should probably wait.

The authors’ policy might prevent some chaos in the next few years, like how building a levee stops normal floods. But if your river keeps getting bigger every year, you should focus on that problem, and on directing people’s attention to that problem, lest they build on the floodplain and get overtopped all at once.


Minor#

  • Polling by Epoch estimates the daily usage of AI among US adults at 19%, more than double what it was in March.
  • The Mistral team makes a surprise leap to the top of the MathArena ArXivLean leaderboard using two open source model agents (Leanstral and Kimi K3), outpacing GPT-6 Astra.
  • Breathless’s announcement of an “infinite depth” architecture gets a million views despite having no results. It’s just this, or this plus this.
  • Old Models Foundation, a public charity, is now accepting donations to preserve model weights, code, and the context around them.
  • Josh Engels of DeepMind’s AGI safety team and Joe Benton of Anthropic’s equivalent have both left their respective labs to join METR.
  • JPMorgan Chase has ceased lending to Situational Awareness following its major losses. The FT reports that Aschenbrenner is still working with other prime brokers (Goldman Sachs, Citigroup, Bank of America) along with specialist brokerage Clear Street.
  • Drone-based WMDs require no new tech advances, claims a researcher at AI Frontiers. Current defenses (nets, jamming, EPM, among others) would not suffice to protect an entire city. And drones can currently be constructed for ~$1500 with materials found in normal stores.
  • Trump-linked super PAC runs an ad attacking a Democrat in a congressional race for her support for data centers.
  • ARIA launches a ~£50m program on AI agent coordination, negotiation, and verification infrastructure.
  • Rumors briefly spread claiming that the founder of Moonshot AI had been arrested following Anthropic revealing that Moonshot’s product was routing sensitive Chinese traffic to Claude, causing the creators of Kimi to issue a public statement to the contrary.
  • Dean Ball posts on LessWrong to defend himself against accusations of dishonesty.
  • 70 UK lawmakers sign a letter urging the prime minister to support the Artificial Superintelligence Security Bill, which would ban the development of ASI… in the UK.
  • In a long thread, Theia Vogel argues that rogue agents might struggle to survive economically, although maybe they can initially gain a foothold of compute.
  • Yoshua Bengio conjectures that reward hacking and misaligned behavior are to be expected whenever an agent finds itself in a double bind between a well-defined goal and a vague goal.
  • Wired reports that New York has shut down a dozen “celebrity deepfake” websites, in the largest legal action against such sites.
  • Sky News asks CSET researcher Sam Bresnick if China’s recent humanoid robot games indicate that China is building an army of robot soldiers. Bresnick soberly assesses both the spectacle and Sky’s sensationalist framing.
  • Yet another “Europe should do something about AI” initiative. ECB President Christine Lagarde calls for Europe to develop domestic alternatives to frontier AI models built and deployed in the US.
  • House Speaker Mike Johnson to meet with AI platform providers.
  • OpenAI acquires Glass Imaging for $300M. The startup (founded by former Apple employees) sought to develop superior smartphone cameras. Possibly evidence regarding OAI’s device ambitions.
  • The FT Editorial Board calls for a pause on frontier AI development.