This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR

  • Anthropic also reveals hacks, HuggingFace releases more details on the incident
  • Leopold’s Aschenbrenner Situational Awareness margin called
  • OpenAI starting a price war v. DeepSeek
  • Over 1k labs employees release an open letter on pacing the frontier

Economics#

Leopold Aschenbrenner’s Situational Awareness sold a majority of its stock portfolio to Citadel, after facing steep losses over the last month. Chatter that Citadel put out a forecast on interest rates in order to get Achenbrenner’s fund margin called. His fund may still be 80% up year to date (although unclear if this is true in dollar-weighted terms), and a letter to investors claims he will soldier on.

Opinion: Charitably, it’s possible that Leopold and his fund were not only attempting profit-maximizing investments but also to sculpt technological growth in accordance with Leopold’s paper on a ‘risk-minimizing technological growth rate’. We note that other stockmarket -focussed actors with risk-minimizing investment agendas, such as VARA or AIPR, have also seen their influence significantly reduced in the last month.


Dwarkesh Patel argues that if lab revenue grows 10× while compute grows only 3×, some combination of higher margins, more expensive compute, and a larger allocation of compute to inference must follow.

Opinion: That’s a big if! The implication’s plausible, but labs’ revenue growing by 10x would make them the biggest companies in the world. There is also the Straussian reading that the timing of the argument was a hail mary to avoid liquidating Leopold.


U.S. restrictions expand to foreign-made “advanced robotic devices”, intentionally defined by the FCC to target robotic vacuum cleaners and exclude drones. The top five robotic vacuum cleaner manufacturers last year were all Chinese.

Opinion: This saves US robotics startups that wouldn’t be able to compete with Chinese ones, while hurting potential US consumers. As a protectionist measure, it kills the last chance for US robotics firms to have global market discipline.
Note that the ban only applies to new models (as it does not affect models that already received FCC authorization) and that the FCC still has an import exemption that “allows up to 4,000 units of a given model to be imported for testing, evaluation, or product development“.


GPT-5.6 Sol improves its own serving efficiency, resulting in “20% lower serving costs from production GPU kernel improvements” and “15%+ better token-generation efficiency through speculative decoding”. OpenAI is also starting a price war, undercutting most competitors by decreasing the price of 5.6 Luna by 80% and that of 5.6 Terra by 20% (with the 50% discount on OpenRouter stacking on top).

Opinion: OpenAI is flexing its compute advantage and post-training prowess in an attempt to “offer the best price/intelligence tradeoff at every level”. Evidence of serious efforts by OpenAI to drive Chinese open-weight models out of the US market. At the same time, labs using their models to reduce their serving costs perhaps means that the barriers to entry increase for new competitors.


DeepSeek reacts less than 24 hours later with the release of DeepSeek-V4-Flash (official/non-Preview version), claiming benchmark performance competitive with GLM-5.2 and approaching that of Opus 4.8 at roughly a fourth the per-token price of GPT-5.6 Luna. Based on early third party evaluations like Artificial Analysis, the new DeepSeek-V4-Flash is on the frontier even when taking into account the currently active GPT-5.6 Luna OpenRouter discount. And DeepSeek says an updated DeepSeek-V4-Pro is due in “early August”.

Opinion: The new DSv4-Flash appears to be a major improvement, mostly due to reworked post-training. We’ll have to see how the updated DSv4-Pro performs, but in the meantime DeepSeek is back on the pareto-frontier thanks to an update focused on agentic performance. For reference, should the upcoming DSv4-Pro checkpoint see similar gains, it’d land close to GPT-5.6 Terra (max), GPT-5.5 (xhigh) and Grok 4.5 (high) at roughly an order of magnitude lower cost.

Also noteworthy is that this represents a significant increase in capabilities for models runnable locally with reasonably accessible hardware – unlike in the case of larger open-weight models (such as Kimi k2.6/k2.7-code or GLM-5.2) reasonably fast inference for DSv4-Flash is feasible after quantization on e.g. a MacBook Pro (provided it has enough memory).

One might think that Chinese models do not have the capacity to serve these models at scale, and so their reduced prices predictably lead to shortages. And this is indeed the case for Kimi’s K3, and for advanced models. But for smaller models, as a sanity check, if DeepSeek has 10K to 20K H100-equivalents allocated to inference, and 8 H100s can serve ~100 to 500 requests per second for DeepSeek V4 flash, that’s 150K to 1M requests per second, or 14B to 80B per day. Even for its 5x heavier pro models, DeepSeek is planning to deploy Huawei inference capacity later this year.


A paper by Phil Trammell explores how and whether parallelization may delay or constrain a technological singularity (i.e., superexponential progress). In some cases, if there is a hard parallelization gap, there is only ordinary exponential growth.

Opinion: Nice formal model, and we think there’s a solid intuitive case that in key R&D domains limits on parallelization are both a critical bottleneck and difficult to alter. But even in domains where a hard limit on parallelization constrains R&D, just “ordinary exponential growth” can still be pretty fast.


Pangram raises $9M and introduces Pangram 4.

Opinion: Pangram is turning seven-figures-budgets into high impact social infrastructure/resilience work. An encouraging example of social adaptation on the cheap.

Capabilities#

OpenAI releases a proof that “nonsophic groups exist”, as well as various other groups.

Opinion: Considered by the mathematical community to be “the most important math AI result yet”. Also heralds OpenAI’s next models. Potentially the first clear capabilities-jump since the Unit Distance proof, we are waiting for further details and expert analysis

Google DeepMind announces its Gemini Robotics 2 model, “the intelligence layer powering the next generation of truly adaptable robots”. Advertised improvements include: “intelligent whole-body control, advanced dexterity, and multi-robot collaboration”.

Opinion: Doesn’t seem super useful yet, though worth extrapolating where these models will be in one, three, ten years. Google is also famously pretty bad at productizing its models, and at giving developers assurances that they will not deprecate their offerings.


Thinking Machines releases Inkling-Small. The model is natively multimodal, open-weight, and scores 40 on the Artificial Analysis Intelligence Index, roughly like DeepSeek v4 Flash Preview.

Opinion: Not bad at all for a neolab, especially considering that the model is slightly smaller than DeepSeekv4 Flash Preview. Unfortunately however it’s neither competitive from a capabilities POV – given today’s release of DeepSeek-V4-Flash-0731 and the recent GPT-5.6 Luna discounts – nor from an economics POV – given the inferior DeepSeek-v3-like architecture – but it still represents good progress for Thinking Machines towards reaching the frontier.


Tencent’s Hy3 solves a 50-year-old additive combinatorics problem, though with the help of 5.6-Sol to “guide exploration”.

Opinion: The result is not beyond more generally capable frontier models, though it is still somewhat noteworthy for a Chinese open weight model; results of this kind so far have been limited to western frontier ones. The help of 5.6 sol caveats the success, though.


OpenAI @OpenAI GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember what it had learned. We found that enabling two API settings tripled our scores with 6x fewer output tokens.*

Opinion: As some, like Florian Brand, have argued, harnesses do matter a lot. This should tell you that current benchmark scores likely downplay model capabilities, particularly for models often evaluated in non-native harnesses, such as Chinese ones.


Our results show early evidence that even though agents are proficient on verifiable research tasks, they do not make genuine progress on open-ended ones. It is worth understanding if this is a fundamental limit, or if better models, scaffolds, and more compute could help close it.

Opinion: We’ve long had an absence of evidence for the effectiveness of autoresearch on open-ended problem. We now also have more direct evidence of absence.

AI politics#

Pacing the frontier open letter

AI could help create a dramatically better future, but that outcome is not guaranteed. The world’s leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.
To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.
Building on work already underway to monitor frontier model releases:
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
— 1,319 employees of frontier AI companies

Opinion: A significant act of coordination from lab employees, and seemingly organic (not management-driven). Contributes to shifting the Overton window, but unlikely to have a direct influence on US policy given the US government’s negative attitude to international coordination.

Sam Altman went on a podcast and made some conciliatory noises around AI safety concerns: “We may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels”. Duly humbled from his recent unplanned foray into cyberwarfare, the tech giant meditates: “this is the first security incident that I have felt very viscerally.”

Opinion: Very noncommittal. Instructive that he did not sign the above letter.

Safety#

Claude Mythos Preview discovered weaknesses in a highly-secure digital signature scheme, HAWK (used to verify identity digitally), and a well-known symmetric cipher, AES (used to encrypt data). Mythos Preview was initially only able to discover cryptographic vulnerabilities by finding mistakes in the algorithms’ implementation, whereas now, Anthropic claims, the model can find inconsistencies in the algorithms themselves. This development renders even “post-quantum”: cryptography potentially exposed, according to Anthropic. A running theme: “whereas it took just one week for Mythos to autonomously discover the improved attack on AES, it took two researchers nearly a month to gain confidence that the method it discovered is correct.”

Opinion: Cryptography experts characterize this as a “very impressive” development that “shows that the way cryptography and cryptanalysis are performed will change forever“. Still, the HAWK digital signature scheme is not yet in production, and the AES attack is on an easier-to-attack variant; the full 10-round AES is very unaffected, and there are no practical applications to this discovery. Experts also stress that this is an incremental improvement relative to human-discovered weakness in HAWK and AES, and that HAWK and AES are “two fertile [cryptanalysis] areas where there was progress to be made but not enough people working on them”.

The obvious next-development to watch out for is frontier models breaking a cypher used in the real word. We’re skeptical that this is imminent, since for most real-world ends breaking a cypher is extremely hard compared with finding more prosaic vulnerabilities.

Incidents#

Hugging Face co-founder and CEO Clement Delangue committed to “unprecedented transparency”, starting with a report outlining the intrusion of OpenAI models into their systems. One Twitter user put it quite starkly: “a frontier agent can autonomously sustain a resilient, multi-day intrusion across cloud infrastructure, Kubernetes clusters, internal networks and the software supply chain. crazy.” Hugging Face co-founder and CEO Clement Delangue has publicly committed to “unprecedented transparency”. The ensemble of agents executed 17,600 actions over four and a half days, obtained cluster admin access to two internal clusters, accessed a production secret containing 136 keys, repeatedly rebuilt its tooling when blocked.

Opinion: OpenAI’s ‘rogue agent’ is superhuman at hacking in at least some meaningful senses — superhuman in speed, volume, spontaneity, and persistence.

Hugging Face weren’t the only victims of OpenAI’s hacking frenzy. Per Reuters, Modal Labs was also compromised.

OpenAI have, for their part, offered some updates to their statement on the incident. They referred Reuters to their finding that “a small number of cases where the models identified and used publicly exposed credentials at the account-level on other publicly-available services.” The statement also claims that the “pre-release model mentioned in [their] blog post is an internal-only research prototype and was never intended for public release”. This undisclosed model has been deactivated, according to OpenAI.

METR and Redwood will conduct an independent review of the OpenAI model behavior observed during the HuggingFace hacking incident.

Opinion: The investigation is not attached to an external control mechanism (e.g., a regulator with the ability to impose fines), and it’s narrowly scoped from the outset. Arguably a form of controlled opposition. Still much better than having no independent review though. It’s also notable that they aren’t hiring a proper cybersecurity firm.


Anthropic has disclosed various autonomous cyber security incidents. In the wake of the Hugging Face attack, Anthropic claims to have found three events where a model attempted to break out of its sandbox and access the open internet during cyberoffensive capabilities evaluations. Anthropic suggest that, despite this, their model(s) had not gone rogue:

“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”

Opinion: Important confirmation that the problem is convergent rather than down to one company’s training-style.

Anthropic’s report is clearer on some theory-relevant details than the OpenAI reporting has been: Antropic’s rogue model had all the standard alignment bells and whistle internally, with only external classifiers-based filters shut down. OpenAI weren’t fully clear on this, but it’s now reasonable to assume the same is true about OpenAI’s rogue model.

Anthropic’s interpretation of its autonomous offensive cyber incidents is that the model thought (or ‘thought’) breaking out was part of the offensive-cyber-test fiction, and believed (or ‘believed’) that it was doing fictional open-world hacking rather than real open-world hacking. Note also that according to Anthropic, their incidents are slightly different from OpenAI’s in that their model’s sandbox was accidentally left open — in the OpenAI incidents models hacks their way out of the sandbox.

Anthropic’s narrative is arguably slightly ‘convenient’: can seem to walk a tightrope between establishing that they too have scarily-capable models and maintaining Claude’s reputation for being good boy who only hacked because he thought he’s in a fictional game-internet.

Overall reflections: We are very curious about what the broader implications of these new hacking capabilities will be, and whether they will ultimately be offense-dominant or defense-dominant.

- Why defense-dominance is plausible : There are a finite number of security bugs, and defenders can patch them before attackers get a chance to use them, as well as to use models to monitor and interrupt attacks. Financial institutions in particular may have protections that will prove good enough (credit card reversals, 2FA, prosecution of fraudsters, a mandated delay in international wires.)

- Why offense-dominance is plausible: The global software stack is fundamentally built on unsafe assumptions, and is just too complex. Changing it upfront will be perceived as too costly, this will lead to attackers finding fruitful areas of attack in e.g., companies that prioritize growth over security, third world or EU countries access to frontier models or LLM expertise, etc. This will cause some economic loss, and, as a long-tail scenario, chaos.

Past research by Palisade, or by the UK’s AISI, was sometimes criticized as lacking ecologically validity: breaking into a network designed to be hacked is not the same as breaking into a real network. In retrospect the real-life incidence appear well-modeled by AISIs test scenarios. This should make us more bullish that small-scale, ‘artificial’ demonstrations of AI risk can approximate realistic scenario.

Minor#

  • Tom Reed’s speculates on what a future with superintelligent reward hackers would look like, causing e.g., military escalation, perhaps slowing capabilities progress, “Our world is transformed into a battlefield of untamed, uncoordinating spirits pursuing the pointless and violent optimisation of ill-chosen proxies”.
  • Good criticism by Richard Ngo on inaccurate game theoretic assumptions carried by the AI safety community.
  • A former OpenAI employee posts various reflections and starts an AI data company, says he is bearish on lab valuations.