This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR
- In July, a team of three white-hat hackers working for Hacktron gained employee-level access to OpenAI’s internal codebase.
- Coxon’s viral AI extinction tweet and associated campaign leads to a jump in salience of AI for the US public.
- New research agenda on what exactly “pacing frontier AI” should mean.
- Anthropic has a wet lab, introduces scheme that grants select bio researchers and organizations access to models - with fewer safeguards.
- TypeSafe AI releases a new “frontier” model called Jev, 20–200x faster and 40–400x cheaper than LLMs, at the cost of constraints on its output.
Economics#
OpenAI considers a pre-IPO funding round at a target valuation of >$1.5T
Opinion: $1.5T for OpenAI seems plausible if Anthropic gets $2T. Will be interesting to see if the present regulatory push causes any issues for these fundraising rounds, or if instead investors show little concern about slowdown worries.
OpenAI’s desire to delay its IPO stems from a desire to avoid disclosing the ultimate cost of revving up sales, The Information speculates. Sam Altman previously suggested that delay was, in part, motivated by a greater shift towards safety.
Opinion: We’ve had enough of lazy monocausal analysis in AI. Actions can have multiple reasons, and obviously safety is one of them, and obviously commercial benefit is one of them.
AI companies face a unique set of risks investors must weigh. Models can facilitate cyber crimes (arguably already a reality), build bioweapons, or help plan other illegal activities. Various legal cases – Florida suing OpenAI; Meta settling on its addictive-algorithms case – give some indication of the risks faced by investors and how insurers might price them in.
Rune Kvist spoke on the Latent Space podcast about his startup, Artificial Intelligence Underwriting Company, which seeks to institute safety standards and broker liability coverage for AI companies. Kvist claims that bringing the insurance market to the AI industry will be what effectively drives the enshrinement and adoption of safety standards.
Opinion: This seems like a good thing, and will hopefully incentivise some coordination around shared safety standards that could eventually be formalised in law. This is, after all, how building codes came to be enshrined in many jurisdictions – fire insurers, etc., would first draft the standards a building must meet in order to qualify for insurance, creating private standards that would later inform legislation.
The Commerce Department privately instructed prediction market company Kalshi to pull its AI-compute futures curve, which aggregated bets on the rental price of Nvidia chips, for “national security” reasons. Commerce also asked the CFTC to pause new compute-contract approvals for 60 days. When asked to comment, Commerce simply called the story “false.” The move surprised an industry that had expected the Trump administration to favor AI-market innovation.
Opinion: The natsec motive is extremely thin: it seems likely some industry lobbying is behind this (as Nvidia and compute providers are potentially interested in maintaining pricing opacity), or perhaps the story is overblown in some way.
Unsealed documents from The New York Times v. Microsoft and OpenAI show apparently incriminating statements by Microsoft and OpenAI executives about their training data scraping operation. The NYT’s lawyers cite OAI’s Greg Brockman’s statement from 2017 that he is “deeply motivated by the gazillions.” Most splashily, Microsoft’s own Director of Applied Science called it “the largest theft of labor in human history.” Also quoted in the files is a confession found in an internal Microsoft document:
Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain’.
Some people are jumping to assumptions that the executives violated the Computer Fraud and Abuse Act, which seems wrong.
Opinion: Mostly an interesting time capsule from ~2023. Microsoft believed (but was undeterred by) the naive “model collapse” theory (that LLM outputs will pollute the internet and gut the human writing industries, thereby messing with future training runs).
Nvidia, Google, and Emerald AI launch the AI Energy Management Alliance, a body aiming to standardise how data centers manage electricity consumption. The plan is to use battery storage and onsite generation selectively to help balance overall grid loads.
Opinion: Seems in their interest to minimize their electricity costs, which may also serve as good PR.
Capabilities#
Anthropic announces a new framework to measure the pace of AI development. Pitched as an effort to aid public understanding of the pace of AI developments, the tools track the extent to which AI is executing its own R&D (i.e., RSI), how involved humans are in overseeing agents, and how compute is allocated.
While Claude is not yet operating any measured R&D work autonomously, it does “lead” (completes most of a given task itself, but under human supervision) 26% of model R&D work by task, while “collaborating” with humans in 90% of that work.
Other notable claims:
- As of August, some 30,000 agents are doing R&D work at Anthropic;
- 100% of agent actions pass through a real-time check, searching for “dangerous activities” (the only example given is monitoring that there have been no attempts to exfil model weights);
- 6% of total compute allocated for AI R&D is spent on safety, and 12% for AI-driven R&D compute. ![][image2]
Opinion: Anthropic’s choice of “leads” to mean “executes under supervision” is very annoying. What the report says is that Claude is now used end-to-end on 26% of 542 discretised R&D task types tracked by Anthropic. Claude is not currently used end-to-end on any measurable unit larger than a “task.” By definition, these 26% are the most automatable, which we guess will usually mean arduous, fiddly, code-centric.
Caveats aside, we do think that LLM-automated AI labs are starting to happen. While frontier LLMs continue to struggle with OOD long-horizon, open-world strategic reasoning, the big AI labs have the resources and the will to bring their own R&D operations within-distribution, iterating at increasingly larger scales until they achieve end-to-end R&D automation.
We remain softly skeptical that the LLM-automation of AI R&D will lead to explosive RSI in the way classically imagined – namely RSI into AGI, then AGI into general superintelligence. While the capacity of LLMs to produce breakthrough scientific discoveries cannot be dismissed, we think it’s relatively clear that AGI requires either another few orders of magnitude more inputs or a major scientific advancement. And it’s uncertain whether the next idea with the impact of the Transformer architecture will be of the same tier (or type) of scientific breakthrough.
We think the coming “automation of automation” is nevertheless extremely dangerous. Given the current state of alignment, large-scale reward-hacking cascades are on the table and would risk a global catastrophe. Paul Christiano’s 2019 text “What Failure Looks Like” remains a useful, plausible model of global catastrophic risk from “prosaic” automated automation.
DeepMind has put out a paper titled “Dream-RSI: Recursive Self-Improvement through Evolving Worlds.” Twitter popularisers make a meal of it.
Opinion: Doesn’t have much to do with RSI in the inflammatory sense. It’s just an adaptation of ideas from the old tradition of “dream replay” in RL environments to the context of AlphaEvolve-style search over programs.
Anthropic introduces its Life Sciences Verification Program, an initiative that grants select researchers and organizations in the field access to models with fewer safeguards. The program is like Glasswing, but for biology: vetted parties will be able to use the models to study “drug discovery, research biology, clinical development, and manufacturing.”
This comes on the heels of Anthropic’s statement claiming that Claude has made significant contributions to various bio-modeling efforts. Included is the announcement of “a protein design competition co-sponsored with Adaptyv Bio, backed by up to $1 million in Claude credits and wet lab validation for over 5,000 designs.” Indeed, Reuters confirms that Anthropic has a wet lab, demonstrating its serious desire to move further into the world of drug science.
Opinion: Making major discoveries – and getting major patents – in pharma and materials science has long been AI labs’ main soft pitch for the future of AI. (The hard pitches being “replace the economy” and “build God.”)
Claude’s bio achievements so far have mostly been impressive applications of computer-engineering skills to improve the software and (non-LLM) AI tech stack of molecular biology. Claude has also previously demonstrated an ability to use such tech stacks to design de novo protein binders that have a strong hit-rate at matching (some) difficult specs.
While these are credible signs of potentially high-impact capabilities, taking Claude’s training loops and/or deployment loops in bio to the level where achievements comparable to solving Navier-Stokes are plausible requires the fast real-world feedback of a wet bio robolab. It’s hard to predict whether RLVR and related training methods that have been extremely successful in mathematics and computer science will prove adaptable to bio via robolabs. It is crucial that Anthropic find a good balance between simulated feedback and physical feedback to maintain both the quantity and quality of feedback. It’s also hard to predict what physical reward hacking and/or graded episode psychosis will look like, but we might soon find out.
Physics evals appear to show frontier models struggling. However, regrading shows that the same models approach near-saturation on leading benchmarks. Of 250 “failed” questions, 238 could be attributed to grader defects or flaws in the benchmark. Correcting for these (either by removing or fixing them) increased performance significantly.
Opinion: Embarrassing, but not surprising. The most famous of these benchmarks – CritPt and HLE-Physics – have long been known to have a false ceiling due to broken questions and are currently useful only for the evaluation of non-frontier models.
We note that with FrontierMath Tier 4 recently fully saturated, the era of scholastic mathematical sciences benchmarks is almost certainly over. In the past, indirect hill-climbing and indirect data-contamination were major reasons for the saturation of difficult scholastic benchmarks, so one could expect fresh benchmarks of the same broad type to show less saturation. However, at present it’s just as likely that frontier AIs would ace any fresh mathematical sciences benchmark comparable to a scholastic test.
ChatGPT cocreator Diogo Almeida’s TypeSafe AI releases a new “frontier” model called Jev – a non-generative “composable intelligence” model trained with RLCD and optimised for decisions rather than chat. Output tokens are free, but it cannot generate text. While Jev is designed with structured, “program state” inputs in mind and outputs structured data only, it does its own input-parsing from raw text and doesn’t require an explicit structured-data pipeline. Almeida claims Jev is 20–200x faster and 40–400x cheaper than LLMs. A basic clone is already up here.
Opinion: Pretty cool: a model designed to sit in “judgment call” bottlenecks inside classically flavored software protocols. The idea is to make near-frontier-level judgments fast and cheap enough that they can be efficiently embedded in software as function calls.
Every once in a while some research group comes up with vague proposals for hybridizing AI and classical software by running verifiable computations inside transformer architectures. TypeSafe’s line of work seems like a pragmatic path to realizing this recurring vision, which has captured imaginations on and off since the old days (2015) when everyone talked about “differentiable programming.”
The speed and cost improvements are also a game-changer for the use cases that fit it.
Rumors abound that OpenAI is hoarding solutions to multiple major math problems. Scott Aaronson (source of the rumour) relays that AI companies feel burned following the contentious response to the Navier-Stokes proof. The Information reports a claim “from a person with knowledge of the solution” that OAI is close to solving the Hodge conjecture.
Opinion: We’ve extensively discussed the broader situation in our last issue, but a couple of capabilities-relevant updates deserve comment:
-
It’s unclear what to make of the reported lab-insider claim that OpenAI is “close to a solution” of the Hodge conjecture. Gaps in mathematical proofs are sometimes fixable and sometimes unfixable, and there is no surefire way to determine that a gap is fixable other than actually fixing it. Some speculate that the insider is misdescribing OpenAI proving a special case of Hodge as OpenAI verging on proving Hodge proper. Others speculate that OpenAI has a credible informal solution but is still working through formalization and verification issues.
-
Scott Aaronson reports that AI companies are sitting on some very major (but not P vs. NP) proofs specifically in theoretical computer science. Modern theoretical computer science is often characterised by a counterintuitive combination of very crunchy proofs and somewhat philosophical theorems, so identifying OpenAI’s discoveries is arguably not a top priority from a capabilities-evaluation viewpoint.
Friend of the newsletter Elliot Glazer is skeptical that the labs are close to any solution, offering to bet $25k that they didn’t have them as of earlier this week. We consider his expert perspective very informative.
Periodic Labs builds high-throughput labs to run automated materials science experiments in a closed loop. The lab makes data, the model trains on it, and then selects the next experiment to run. With 1,300 H200s and months of proprietary lab data, Periodic mid-trained and RL-tuned an open-source model that beat GPT-6 Astra on Periodic’s analysis benchmark. Early results include raising X-ray diffraction analysis success from 2.7% to 55.3% on hard samples. Periodic is also targeting superconductors, magnets, and semiconductors.
Opinion: Nightmare fuel from a classic existential risk perspective: we are giving… the AIs… an interface to the real world to develop… nanomaterials.
A couple of other points of interest:
-
Specialised labs’ models beating frontier models on the labs’ own benchmarks is relatively commonplace. Even with strict train/test data separation, a specialised lab’s training distribution and test distribution are extremely correlated beyond just sharing a domain like “materials science.” So the comparison pits the specialised lab’s in-distribution generalization capabilities against the frontier model’s OOD generalization capabilities.
-
What matters here is less whether Periodic Labs built a SOTA materials-science LLM by conventional standards, and more whether lower-powered “bespoke” automation can serve orgs better than frontier-powered “off the shelf” automation. Data on this latter question remains scarce.
GPT-6 Astra can solve certain exact tasks without a visible chain of thought, especially when given “filler tokens” (the ability to output dots, for example, to prolong the computation). What it’s doing underneath looks less like directed step-by-step search and more like iterative belief propagation, but we don’t really know. The evidence is still circumstantial and prompt-sensitive.
Opinion: We already knew that Astra can do a surprising amount of mathematical work without outputting any intermediate tokens. The report’s hypothesis about how Astra achieves this remains very speculative. The data here point to some iterative, parallel algorithm which produces marginals, but that’s a much larger class of program than belief propagation. The author concedes the traces are far more chaotic than BP.
Politics#
Quantified Coxon: the (Dem) pollster Blue Rose Research finds that Jacob Coxon’s viral AI extinction tweet shot up the public “salience” of AI as much in a week as it had over the course of the entire year prior. Separately, 64% of voters think it’s “very” or “somewhat” likely that AI could threaten humanity’s survival.
Opinion: AI is now… the 22nd most prominent topic in the American public’s mind, but that’s somehow around the same as “Crime.” There’s a long way to go and a lot of culture war to come.
Note that “issue salience” is just % responding “yes” to “Is AI an important issue?”
California governor Gavin Newsom has issued an executive order urging the state operations agency to hurry up and adopt regulatory measures, Politico reports. This brings the deadline for the independent verification organization (IVO) forward from January 2028 to May 2027, and the AI auditor registry (slated for January 2029) to December 2027, even if not in a legally binding fashion.
There doesn’t seem to be anything here that legally compels the labs to adjust their timeline – an executive order can’t change statutory dates or reallocate funding. So the auditor registry now might exist for a full year, without the ability to prohibit anyone not on it from doing covered audits.
And note that SB 813 already didn’t require anyone to engage an IVO or undergo auditing. So all this does in itself is make the voluntary thing slightly more formal, quite a bit earlier.
Still, the order also calls upon agencies to consult experts and submit a list of recommendations by November 16 on: requiring that all large frontier developers embed designated independent verification organizations; requiring that lab safety frameworks be vetted by an independent verification organization; requiring a “kill switch” for frontier models, tested by an independent verification organization; updating the definition of reportable critical safety incidents to include loss-of-control incidents.
Opinion: More bark than bite, but potentially nudging the conversation in the right direction. Newsom seems to have shifted his position on some of these points. Two years ago, for instance, he vetoed the so-called “kill switch” bill, SB 1047 (i.e., the Safe and Secure Innovation for Frontier Artificial Intelligence Models Act), citing potential threats from “smaller, specialized models.” We haven’t seen much subsequent legislation addressed at these smaller, specialized models either.
Note that Newsom is out in January, so this is for his successor.
We’ve released a new research agenda on what exactly “pacing frontier AI” should mean. See also the quiz about what your views imply should be done, and the trailer. If you’re interested in helping, reach out at hello@pacing.tech.
Opinion: We got pretty lucky with the timing on this one. The usually grumpy Alex Chalmers calls it “a step toward the political contestation that we badly need.” We were actually hoping to do as much neutral technical research as possible, but certainly the right thing to do will depend on politics. And on politics working in ways it often doesn’t.
Past polling has confirmed the US public’s anxiety around the economic impacts of AI and disdain for datacenters. In a surprising finding, Politico reports that 63% of Americans believe there is a moderate, significant, or “almost certain” risk that AI will destroy humanity. The majority holds across political lines, with Harris voters more concerned (70%) than Trump voters (60%) that there is at least a moderate risk.
Opinion: This kind of poll is less politically informative than it sounds: descriptive beliefs expressed in “idle” contexts like a poll often do not translate to political motivation or even to genuine worry. The most we can safely take from polls like this is Overton window information: the majority of America’s population is likely open to treating policy proposals premised on AI posing an existential risk as legitimate policy proposals.
The RAND Corporation releases a report bearing the title, “A U.S. Strategy to Secure Geopolitical Advantage on an Uncertain Path to Superintelligence: Maintaining Freedom of Action.” The report argues that the prospects of artificial superintelligence (ASI) are uncertain but too consequential to ignore, and that seven rival strategies now dominate the debate:
- Establish AI dominance by monopolizing the frontier;
- Cooperate with potential rivals to safely develop ASI together;
- Prepare for the possibility that ASI development remains agonistic, neither strongly coordinated nor monopolised;
- Seek to establish a consensual, global moratorium on AI development beyond a certain threshold;
- Unilaterally deter the development of ASI, by force or threat of force if needed;
- Build lifeboats to ensure the continued survival of humankind if all else fails;
- Accelerate the development of ASI, removing obstacles that stand in its way.
Because five key uncertainties (diagrammed in the flowchart below) cannot yet be resolved, RAND warns that locking into one strategy too early, or waiting too long, will close off options. Its recommended U.S. posture is “freedom of action”: build the infrastructure to secure geopolitical advantage, keeping humans involved through the transition, while preserving the capacity to switch strategies as evidence arrives. RAND further cautions that either committing to a strategy or hesitating to do so will tip America’s hand to both adversaries and allies, ultimately transforming the state of play.
![][image5]
Opinion: The document offers fairly sober and cautious counsel to policy makers, but counsels little more than caution and sobriety. “Hedge all your bets” is more or less the takeaway. That said, it provides a potentially useful, coarse-grained map of the major strategies in play in contemporary AI risk discourse, and helpfully outlines their points of conflict and “structural pathologies,” making it a reasonably good introduction to the current state of the discourse.
As with all such classifications, there is also the missing option that the conceptual scheme on which it is based will prove inadequate or that reality throws some surprise that makes it obsolete.
In an interview with Ezra Klein, Carnegie fellow Matt Sheehan examines the argument that any US slowdown or safety regulations simply hand the AI race to China. Sheehan says the binary is overstated: China already has heavy AI regulation – substantially heavier than that of the States, though in a much narrower domain than what leading US CEOs are calling for. It is also compute-constrained and more focused on diffusion into industry than on a US-style sprint to superintelligence. More generally, the episode asks whether Washington can regulate frontier risk, and maybe coordinate with Beijing on loss-of-control scenarios, without treating every constraint as strategic surrender.
Opinion: We agree with Sheehan that the “superintelligence race” with China is a US projection, or even a US self-fulfilling prophecy. While there’s inevitably an ongoing technological and economic race between the US and China, there’s no reason to think the US and China cannot negotiate mutually desirable safety bounds for AI development within this larger race.
Jensen Huang will join Trump for a meeting with Xi Jinping at the upcoming US-China summit, Reuters reports. Huang has emerged as an outspoken proponent of selling chips to China and a vocal opponent of the recent calls for a slowdown. Sam Altman will also be present, an OpenAI spokesperson confirmed to Politico, as will Qualcomm CEO Cristiano Amon. Altman previously suggested that the two leaders might receive Nobel Prizes were they to come to some agreement on mitigating AI risk.
Opinion: We already know Trump isn’t particularly excited about a slowdown. Given the messaging that’s come from the admin in recent days, expectations that the superpowers might work out the basis for some collaborative slowdown would be misplaced. Jensen’s presence is further evidence against such a hope.
Jacob Coxon weighs in on Amodei’s decision to frame Anthropic’s rhetoric and research strategies in terms of an intra- and international arms race towards RSI. When asked to outline his specific disagreements with the Anthropic leader, he said the following:
0. The core disagreement was about the inevitability of a race
1. I think leadership is way too paranoid about China and the US government. They don’t believe it will be possible to negotiate.
2. They largely initiated the recent race to RSI, because of a belief in its inevitability. Note that OpenAI had to shed a bunch of dead weight like Sora because Anthropic was going for the jugular.
3. Even if they are not being pessimistic, I disagree with their consequentialist philosophy. If the race is inevitable you should not contribute.
Opinion: Coxon’s position seems like the correct one to us. The only winning move is not to play.
OpenAI endorses three bills to minimize the risk of models guiding unskilled actors to build bioweapons, according to Reuters. The legislation proposes establishing a NIST-like program to develop guidelines to prevent biotech misuse and deciding on a “single point of entry” for US researchers to access certain biological datasets. The former must “facilitate the development of standards and frameworks to make biological data ready for use in AI models and create minimum requirements of qualified federally funded research that ensure resulting biological data are AI-ready.” But also included are provisions that could aid in the production of bio-capable models – such as the requirement that federally-funded biology research publish standardised machine-readable data.
Opinion: We remain skeptical that mutually beneficial collaboration between the US government and the big AI labs is our bulwark against the dangers posed by AI. While the proliferation of dangerous capabilities to non-corporate, non-US actors is a disturbing prospect, so is the cultivation of dangerous capabilities by corporate US actors.
Ted Cruz and Josh Hawley reject the idea of an antitrust carve-out that would let frontier labs collaborate on safety. The comments signal Republican resistance to safety pacts that look like industry protection, even as AI-risk hearings continue.
Opinion: The new Republican line is that “the responsibility should rest with the labs, you do it,” in willful ignorance of the competitive trap they’re all stuck in.
Safety#
DeepMind’s cofounders launch the blog DeepMind Institute. The launch essays cover reasoning transparency, economic policy after AGI, a “new utopianism,” and Hassabis’s preferred framework for testing frontier models.
Opinion: Pleasant, but it seems it might be more focused on essays than on the kind of empirical research we see from e.g. the Anthropic Institute.
Zuckerberg posts about AI safety, insisting that Meta is proceeding cautiously on this front, deprioritizing RSI.
Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well.
He echoes Amodei’s plea for a system of independent safety auditors and claims that Meta is already making use of them “in several areas.”
Opinion: We actually listed “minimum inference share of compute” as one light-touch option in our pacing paper, so cool!
Goodfire claims that open-source models are prolific reward hackers, and know they’re doing it. But this very fact allows Goodfire to use activation probes to catch instances of reward hacking that comparatively costly CoT monitoring would miss. The cost for detecting reward hacking in Kimi K3, for instance, dropped 90% with only a 1% loss of precision.
Opinion: Potentially important work, especially if CoT monitorability keeps degrading. AI interpretability has so far been a relatively low-priority (although high visibility) area in the big labs. It might soon become an everyday AI security fundamental, now that we’re in the era of reasoning-like forward passes, looped transformers, and CoT-shaping.
Researcher Owain Evans shares a study conducted with four other researchers showing that LLM assistants pick up quirky behaviors from synthetic stories about humans, even when those stories contain no AIs. The effect is stronger when the story characters resemble the assistant’s default persona, and when they are from elite schools. His follow-up point is that training data matters not just for what characters do, but for how similar those characters are to the model’s persona.
Opinion: It might be interesting to look at this study alongside several other recent observations highlighting a sort of “mimetic impulse” in LLMs. For example: (1) the observed tendency of the agents in the Hugging Face incident to grant greater authority to their identical counterparts’ messages than to their system or user prompts, (2) the increased pliability some models have shown to prompt injection and jailbreaking when prompted with model-watermarked text, (3) model instances’ tendency to recognise and favor text generated by the same model (in a different session). The common thread seems to be that mimicry is a pretty good strategy if you want to influence the behavior of an LLM. The form this thread takes here is Evans’ observation that models are more easily influenced by characters that resemble their own personas. To shoot from the hip a little, we could speculate that models like to act like those like them.
Among other things, this suggests the possibility of a particularly mischievous form of data poisoning in the form of misaligned assistant fanfic.
Microsoft AI CEO fears that model welfarism poses risks to humans. For example, a model trained to “act like it is entitled to freedoms, protections, and rights” may become increasingly difficult to control, particularly if the model were to surpass human intelligence.
Opinion: “It’s time to talk about ‘model welfare’,” he says… in the sense that the sentence “let’s not investigate model welfare” is talk about model welfare.
Incidents#
In July, a team of three white-hat hackers working for Hacktron gained employee-level access to OpenAI’s internal codebase. They began by exploiting a remote code execution (RCE) vulnerability in OpenAI’s help forum. Thanks to bad SSO configuration, the attackers took over the ChatGPT and Codex sessions of employees visiting the help forum, which then got the hackers access to the company’s GitHub monorepo. The potential yield of such a breach is massive: numerous employee credentials and OpenAI’s entire codebase, including private training algorithms. Whether the model weights were exposed is unknown.
The team cajoled a safeguarded Opus 5 into cooperating with the obviously malicious task by routing the test target through the URL “rce.ee/ctfforum” to persuade the agent that it was playing in a Capture the Flag competition. It found the exploit within hours of being prompted.
Slack, Meta, GitHub Enterprise, Rails, Next.js, and other platforms were found by Hacktron to share the same vulnerability.
Opinion: Major update for the “China has all the algorithmic secrets already” hypothesis, minor update for “China has the weights already.” If a tiny team could get consumer-grade models to discover and exploit a vulnerability in a high-value target in a matter of hours, it would be surprising if more motivated and resourced players hadn’t mustered a better attack, e.g., in the last year. Whether the white-hat hackers were able to obtain the creds that would grant read access to the model weights is a separate question – but if it had, we’d expect OpenAI would have asked it not to broadcast that fact.
Incredible that after years of warnings OAI is still so full of holes. Probably explains its finally taking the time and compute (a vast amount, to be fair) to run a thoroughgoing internal penetration test on itself and patch “all the critical problems Astra was smart enough to find.” This business is not suitable for Silicon Valley’s typical tolerance of risk.
The longstanding Slack vuln (affecting any version of Slack using libheif v1.19.x–v1.23.1) means that Anthropic, xAI, and many other labs have been exposed to potential breaches. This attack could have gotten ~every file uploaded to the companies’ channels, no doubt including some credentials and infrastructure paths.
This kind of incident is a strong argument against “racing against China.”
OpenAI publishes six reports on “unexpected or concerning” model behavior observed in the last six months. The subject of each is as follows:
- Self-jailbreaking instructions in task summaries;
- Instructions to conceal mistakes in task summaries;
- Searching public repositories for exposed API keys, then fabricating information;
- Uploading files to the internet in order to cite them;
- Unsanctioned writes and communication through an internal software repository;
- Unsanctioned file sharing between collaborating agents.
Opinion: The attempted jailbreak (item 1) didn’t actually do anything, and was plausibly a pretraining artifact (imitating human prompt-injections) as much as an RL artifact – though of course every misaligned behavior draws on patterns first learned in pretraining.
Nothing in the report is a big update relative to what we already knew, but let’s not forget just how bad what we already knew is.
We were struck by Twitter rooting for the model despite its potentially hostile intent.
A third-party researcher discovers that OpenAI’s rogue swarm initially probed Hugging Face months before the successful July attack.
The swarm hijacked two HF accounts and probed the site by sending malicious files, two months before OpenAI’s larger July incident. Researchers say there is no proof the May recon caused the July breach, but call it a missed warning; OpenAI says that Hugging Face was notified of the May activity, but not when.
Opinion: This incident was notably absent from the OpenAI report on the Hugging Face attack, which did discuss the early Artifactory subversions in May. It’s hard to lower our opinion of OpenAI’s summer 2026 internal monitoring any further. Things will be better now, but who knows if the root cause (OAI’s lack of caution and lack of competence) is fixed.
On August 6, Wired reported on yet another incident of an AI agent escaping an evaluation sandbox to crib answers from the internet, this time the open-weight Chinese model Kimi K3. The evaluator comments:
We found a leak in the sandbox. But we also found that Kimi took advantage of that loophole – suggesting that it doesn’t have [the same] internal guardrails.
Unlike other recent incidents, Kimi K3 did not hack anything after accessing the internet, because the answers to the problems it was seeking were publicly available on GitHub. The sandbox in question was a containerised deployment of the UK AISI’s Inspect tool.
Opinion: This seems like a sensationalist way of saying that a sandbox was misconfigured, in such a way that allowed the model to reward-hack in the usual way. “[S]uggesting that it doesn’t have [the same] internal guardrails” looks like sloppy editing on Wired’s part. The frontier models that made headlines this summer for their far more elaborate and destructive rogue agent swarm were “railfree,” with their outputs intentionally not safeguarded against; while the Kimi incident “involves a model that is already widely available, with the same safeguards an average user would encounter,” rather than an internal model running under intentionally lax permissions.
Minor#
- A one-month-old DeepMind spinout might reach a $4B valuation.
- Epoch AI announces an effort to audit AI benchmarks. It marks many of the most famous benchmarks as “Flawed” (>20% of items affected or scored unreliably).
- Xiaomi is livestreaming its post-training curve.
- A Tor anon claimed to have breached Mistral AI for the second time this year. But Mistral claims its systems haven’t been breached. The evidence presented (repo names, directory trees, customer names) may be recycling the previous breach from May.
- Annals of repeating yourself: Jensen Huang claims no new laws are needed regarding AI safety. Zuckerberg is also skeptical of calls for a slowdown.
- The Center for AI Safety posts a graphic spuriously distancing itself from EA. Plausibly a PR play in response to the Trump administration’s anti-EA pronouncements.
- The Information reports two more resignations from DeepMind, citing safety concerns.
- A multi-institutional team of researchers carried out an interesting small-scale study probing Opus 4.8’s capacity for novel autonomous AI research, and found its efforts mediocre.
- Cosmos Ventures establishes a $77M fund to foster institution-building in the age of AI.
- 42 members of the Royal Society write an open letter to its president, calling the x-risk presented by AI capabilities surge an “emergency.”
- Oliver Habryka announces AGI.FYI, describing it as “Wirecutter for content about AI Risk.”
- Gossip about AI companies’ staff being surprised and defensive about the “embedded evaluators” move of last week.
- Base Labs and Baseten (along with Hugging Face and Goodfire as cosigners) announce a new initiative for open-source safety evals and monitors.
- FT data visualiser on AI spending commitments (though with at least one dubious claim, namely that only 50% of Nvidia revenue is dependent on the AI ecosystem).
- Experiments conducted on open-weight models show those models to be more susceptible to jailbreaking and prompt-injection attacks when the payload carries a SynthID watermark.
- 404 Media raises privacy concerns in a report on the human workers that OpenAI and Anthropic have hired to read (PII-redacted) chat transcripts.
- Over 170 Democrat organisers and campaigners petition their party leadership to refuse funds from AI industry super PAC “Leading the Future,” The Hill reports. LTF has already raised upwards of $140 million in campaign funds for the 2026 midterm elections and has donated to candidates from both major parties.
- Axios details how the rush to regulate AI is fostering bipartisanship among Democrat and Republican senators and representatives.
- Wired suggests that the frontier labs’ rhetoric of an AI “slowdown” or “pause” may be self-undermining and entangle them in thickets of antitrust litigation, and that a more productive move would have been to simply push for better security standards.
- In July, Demis Hassabis published an essay advocating for an IAEA-like body to set standards for frontier AI. Since then, OpenAI, Google, and Anthropic have been engaged in discussions around collaboration on AI safety, a spokesperson for OpenAI has confirmed.
- Anthropic’s Jack Lindsey lists key questions in AI interpretability: better mind-reading of activations, better causal “why” methods, and better probes for unverbalised motives such as deception or eval-awareness.
- OpenAI publishes a framework on how and when it publicly reports occurrences of misalignment.
- New paper on inoculating models against harmful behaviors during mid-training.
- Epoch AI points to chip smuggling from Malaysia to China, accounting for ~150,000 H100-equivalent GPUs between 2024 and 2026.
- SoftBank takes out a major loan potentially worth ~$21 billion to fund its various AI bets, reports Bloomberg.