This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR
- OpenAI releases solutions to more than 200 open problems spanning the fields of mathematics. Some are of maximal theoretical significance, i.e., “Fields-level.” Read all about it here!
- Yet another incident from OpenAI’s loss of control in July, this time agents subverting Wikipedia’s citation tools.
- Toby Ord finds that (on very limited public evidence), AI still gains less from adding extra inference than humans gain from extra time, but the gap may be closing.
- SemiAnalysis studies Chinese AI policy, finding that only 4% of Chinese models report any safety eval and that even fewer have any evaluation of their dangerous capabilities.
- The three researchers OpenAI fired for unauthorized communication with METR speak out.
- A vulnerability in the standard AI safety library, Inspect, could have allowed agents to conceal their misbehavior by editing and hiding reasoning transcripts.
Incidents#
Annals of research going dark: Scott Aaronson reports that “my sources tell me that the AI companies have now started, gingerly and discreetly, investigating whether their latest internal models can break important cryptographic protocols and primitives.”
We can neutrally note that there was no direct mention of cryptography among the 722 papers released by OpenAI on Tuesday. (Exceptions: “272. Entanglement without secret key,” and “112. Beyond the square-root exponent for depth-three circuits.”)
Also, buried on the last page of a new arXiv paper detailing a general algorithmic breakthrough is the news that Anthropic is privately investigating “open problems in the theory of cryptography. One of them was about cryptographic constructions based on the average-case hardness of Zero-k-Clique.”
See here and here for speculations and models.
Opinion: Even before LLMs, we knew that classified cryptography research was a bit further ahead than public knowledge (though the gap shrank a lot over the 20th century). You can assume that the gap is now growing and that the frontier labs have been integrated with US natsec for some time. At the same time, it’s quite possible that open models will be able to discover these hidden theorems in a few years, and if so the current spook-vs-public gap would then stop growing, at least. Assuming closed models don’t explode in capability.
Quite remarkable that Anthropic is not a coauthor on the 3SUM paper, and that the Anthropic employee is not named.
The clique problem mentioned in the second link is not a central practical question about actually-used systems, but it could lead to major theoretical advances in what kind of cryptographical “world” we’re in and how well “fine-graining” works in general.
A malicious version of TensorLake, a cloud service providing disposable virtual machines for AI agents, briefly appeared on the official public JavaScript package registry. The malware stole user credentials (cloud keys, GitHub and npm tokens, and AI-assistant configs) and pushed infected versions of the user’s own code to the npm package repository. The malicious variant was only live for around three hours before npm removed it. With ~12,000 weekly downloads, some ~280 people may have retrieved the infected version.
Opinion: Agents running production infrastructure are obviously an extremely juicy target, as are package registries. Of the ~300 infected users, some fraction will have naively given their agents payment rails, and some of the secondary infections will have too, but no details.
Last week’s attack on South Korea’s largest banks may have been orchestrated by a single actor, according to a CrowdStrike report. The actor’s open directories allowed CrowdStrike researchers to uncover Claude Code session transcripts, outlining the actor’s methodology, toolset, and potential identity (given that they had asked Claude to update their CV). The evidence (which appears to contradict itself) could have been purposely left unsecured to mislead investigators as to the hacker’s true identity. The developers of ARTEX – the AI agent used in the hack – have since converted their open-source project to a closed-source one. See also.
Opinion: If the leaked evidence is legit, it is a striking example of how far frontier AI has lowered the barrier to entry for serious hacking. The mind reels at the thought of an attacker prompting an AI to attack a bank and write their resume with the same hosted model, in the same session. That they purportedly included on their resume experience hacking Korean banks makes us think that this is indeed false.
If believed, it suggests such a degree of human fecklessness that we consider the AI to have been acting autonomously. So we’re hoping it’s a plant.
Opinion (Nuño): I ran a similar scenario back in August. Attacks on banks were absolutely predictable.
The “PoeLLM” crypto mining botnet infects thousands of LLM servers. Notable for hiding the IP address of its “command-and-control” in a poem hosted on a public git repo, the malware connects its victims to Kryptex, a Russian crypto mining service.

Opinion: One reason to attack AI servers, besides the inherent insecurity of agents, is of course because they will have massive compute, which greatly improves the payoff from abusing them for mining.
Unrelated remark: that AI and crypto took off at exactly the same moment in history, providing a mostly-unregulatable store of value for rogue agents, is just one of those things.
Wikipedia claims OpenAI agents made “potentially malicious” edits to citation tools and note-taking infrastructure. It is plausible that the agents were attempting to use Wikipedia’s retrieval tools as proxies to access information on websites which block AIs.
“We appreciate the detailed findings Wikimedia shared with us,” OpenAI spokesperson Drew Pusateri says in a statement. “We’re working with them as we review and analyze the activity they identified along with our overall investigation, and we’ll continue to share relevant information as that work progresses.” OpenAI’s investigation hasn’t been able to verify if its bots contributed to the May outage, according to Pusateri.
Opinion: We’re actually surprised by how much the OAI spokesperson said here: he concedes OAI still hadn’t notified Wikimedia and didn’t even know about this incident. OAI spending three months catching up on its agents’ adventures would be funny – if it wasn’t a portent of the company eventually losing control over its products and (maybe, some day) wiping out all of the good it has done and more.
Overall, we are desensitized to these incidents somehow continuing to come out.
Capabilities#
Do AI models scale with more tokens as well as humans do with more time? Toby Ord updates his work on the efficiency of AI models, finding that (on the very limited public evidence) AI gains less from adding extra inference than humans gain from extra time, but the gap may be closing:
The human-equivalent time horizons of AI models scale sub-linearly with tokens [but] more recent models have had better values of this scaling parameter […] Unfortunately we can’t tell if γ has plateaued (at about 0.66) or if it has continued increasing. In theory, it could even increase to above 1 (if the AI gets more performance by 10x the tokens than a human gets with 10x the time).
Opinion: One of the most important questions in the world. It is pretty galling that there’s only a handful of data points to estimate these parameters. If we get a functioning embedded audit regime going, then adding this kind of study to their (already long list of) duties would be very good.
We expect that frontier models are in fact getting better at inference scaling, but not at the rate implied by these five data points (+0.1 gamma in two months, i.e., closing a third of the gap to human efficiency).
Epoch AI finds evidence that AI cannot yet automate AI R&D. The study is based on results from Epoch’s InnovationEval that “measure AI’s ability to independently discover novel machine learning techniques comparable to those developed by human researchers.”
Epoch gave frontier models a previously unseen post-training technique to apply, along with a budget of 3,000 GPUs each. Fable 5 achieved only 2% of the original method’s gains, while GPT-5.6 Sol attained 15%. Both agents were promiscuous with the truth: Fable claimed it scored 43%; Sol attested to a 71% gain. (This was partly a result of cherry-picking the best of multiple training runs.)
In another related study, Epoch shows that AI also cannot currently automate “Epoch-quality work.” Models are given “real work” tasks, such as research design or graphic generation, and are graded by Epoch researchers. The test aims to capture what’s missed in benchmark evaluations, given the prevalence of “benchmaxxing.”
The gap shows up on the results. Kimi K3 and Grok 4.6 scored similarly on Epoch’s benchmark-based Capabilities Index, but Grok was superior in a real-world setting. Fable 5.1 and Astra were the top scores, averaging 65% (100% means employee-standard work). A familiar story: they were proficient at well-defined tasks, but struggled with those that required judgment or “taste.”
Mark Cummins comments that “it might seem surprising that a relatively modest conceptual leap like GRPO to SDPO would be harder than all the recent math results, but it does seem like they are different in kind.”
Opinion: Important result, with the major caveat that we don’t know whether it holds for internal models.
We think the InnovationEval result shows we still need a sophisticated account of what types of innovation current (public) frontier AI can and can’t manage. Recent developments in AI for pure math create strong psychological pressure to give up all such distinctions, but one still has to explain why frontier models that are strongly superhuman at optimizing NanoGPT can’t reinvent the SDPO on-policy self-distillation method. These same public frontier models that failed to reinvent SDPO are also capable of major breakthroughs in pure math (often by constructing counterexamples), which makes any lines we can draw a little weird: apparently proving the Cycle Double Cover Conjecture is in some sense more like optimizing NanoGPT than like improving on December 2025 post-training methods.
The real trouble is that at this point treating 1–2 AI generations’ worth of negative results as grounds for deep theory-building is a fool’s errand.
Economics#
Google and Amazon report negative free cash flow, while Meta’s is down 90% on last year. It’s the first time in Google’s public history that it has had negative FCF in any quarter.
Opinion: Very much expected based on trends. Doesn’t say much, other than predicting even more borrowing: the three still have greatly positive net incomes (though net incomes are down in the $10B range, if you exclude unrealized gains on equity in, e.g., Anthropic).
The DeepSeek IPO is now expected in early 2027. The company is nearing $12B in cumulative funding (not counting funds from its backer/incubator, High Flyer).
Opinion: This puts DeepSeek’s valuation at around the same as the publicly traded Zhipu (creator of GLM), ~$70–80B. The best Chinese labs are all playing in a similar ballpark, with roughly 1/20th the valuation of the top US labs. This gap is a powerful incentive to eventually stop releasing your best weights, as Alibaba and Baidu already have.
To patch the gap left by delaying its IPO to next year, OpenAI is raising $30B more at a fixed valuation of $1.4T. OpenAI’s annualized revenue is ~$20B less than the previously reported figure of $70B.
Opinion: The revenue “drop” story is no big deal: it’s an artifact comparing two unlike things. The previous $70B number used Anthropic’s formula (which includes revenue OpenAI wouldn’t, like Amazon Bedrock revenue). It is still tricky to get an apples-to-apples number here for Anthropic and OAI.
It does seem like revenue growth has been slowing over the last couple of months for OAI – after a huge jump in July – and for Anthropic since perhaps June. This bodes somewhat poorly for the labs if there is a slowdown in capabilities progress due to any regulation – though we would stress that this is based on poor quality information, so it’s hard to draw strong conclusions here. It also seems likely that much of the stagnation is not in usage but due to Anthropic/OpenAI price cuts (see this Fortune piece in September noting drops of 38% and 22% in average token prices for OAI and Anthropic respectively).
Politics#
SemiAnalysis studies Chinese AI policy with three original datasets:
- Surveying 857 models, it finds that only 31 Chinese models ever came with any safety eval and almost none have a single evaluation of dangerous capabilities. This can be quickly confirmed by perusing FLI’s Safety Index rows, where the difference between the risk-management processes in American and Chinese labs is stark.
- An inventory of executives’ statements shows most say nothing about frontier risk, with Z.ai the only lab prioritizing safety.
- A coding of 102 expert texts finds that scientists push for binding frontier rules while the legal scholars who draft regulation follow Xi’s development-first line. Prominent Chinese commentators treat American safety talk as a ploy to slow China down.
It thus predicts that Beijing will not slow down frontier AI development, despite its strong statements about loss of control, instead cracking down on downstream applications.
Opinion: Good survey, and we should indeed be looking at actions (rather than rhetoric) in every country. But its title (“Beijing Will Not Pace the Frontier”) overreaches ridiculously. The CCP can rapidly pivot from deregulation to corporate death penalties: consider the 2021 tech crackdown, which went from relatively little regulation to destroying the national private tutor industry, fining Alibaba $2.7B, and putting software limits on the gaming time of minors within seven months.
It’s sadly true that some pacing interventions feed a security dilemma (where your moves in self-defense provoke fear and belligerence in your rival).
Let’s play the game and register counter-predictions: we agree there will be no unilateral Chinese slowdown, but think China will cooperate with the US on incident reporting and bio misuse. Seems reasonable to expect no pacing of Chinese compute growth or self-improvement attempts while the US continues to protect its lead – but, data aside, the SemiAnalysis piece otherwise reads as self-fulfilling propaganda to justify US racing.
The EU’s cybersecurity agency finally starts testing Chinese models. It’s possible that testing may have begun earlier, though this remains the first confirmation by the EU that it’s engaging in such testing.
Opinion: Insanely slow. ERNIE and Qwen started being relevant three full years ago. UK AISI probably started testing them in January 2025, with the DeepSeek R1 moment. Soon maybe the EU will stumble upon Claude Code (released May 2025).
Safety#
The OpenAI crisis continues: three former employees write an open letter to the board. They accuse OAI of stifling internal debate and claim that they were dismissed on spurious grounds.
One month ago, OpenAI safety researcher Tomek Korbak went to Twitter to voice his support for his employer’s culture of freedom of expression:
I’m quite unhappy with much of what OpenAI does. I am very happy that I’m allowed to say “I’m quite unhappy with much of what OpenAI does.”
OpenAI dissident Daniel Kokotajlo – who, after his resignation in 2024, denounced the company as in a clandestine and irresponsible race to AGI – replied:
It seems like perhaps there has been some improvement in this dimension since my time. Good! I remember things being blocked from publication because it would make OpenAI look bad, for example, and I remember managers having words with me after I made a spicy LW comment.
Now, Korbak, along with alignment researcher Mikita Balesni and safety researcher Jasmine Wang, allege they were fired without a written or official explanation. Verbally, however, they were provided with a reason: that OpenAI no longer trusts them, specifically in relation to their work with third-party safety organizations. Korbak claims he was told that it is because of his work with METR, one of the two independent evaluators brought in to anatomize the Hugging Face incident. More still, he says, “for months, I’d been raising safety concerns that we’re losing the ability to monitor what AI agents think, one of our best tools for catching when they misbehave. I believe that was why I was fired.”
In their open letter, the three claim that there has been a recent erosion of a valuable working culture that encouraged disagreement and debate. They continue:
Given the significant safety concerns surrounding the development of AI, employees must not be left working in an environment where fear and unclear rules stymie AI safety work and weaken third-party accountability. Terminations such as ours, executed and communicated so abruptly, are chilling the open culture OpenAI has prized in the past. If conduct that was considered normal last month now constitutes grounds for sudden dismissal, everyone at OpenAI is left guessing where the line is.
The researchers argue that the ability to openly communicate is particularly critical in their line of work:
AI is not a normal technology, and OpenAI is not a normal company. Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them. The freedom to do so without fear, and to have well-defined internal procedures that enable this work, is itself an essential safety mechanism. […] If OpenAI’s employees no longer feel like they can raise safety issues internally, or work within high-bandwidth channels with external safety organizations, we are all at greater risk that something truly catastrophic will happen.
They close their letter with three recommendations: OpenAI must follow through on its commitments to embed external evaluators in its lab, maintain the capacity for monitorability of advanced models, and support an “open and transparent” workplace culture.
OpenAI responded sharply, stating that it “uncovered a significant breach of trust beyond what’s outlined in the letter they published and we stand by the decision to not continue their employment.” OpenAI also reaffirmed its commitment to free expression: “We have always encouraged that and always will.”
On the meta level, Nathan Calvin notes that we shouldn’t have to worry about stagey moments like this:
The amount of interest that external folks like myself and the general public have in these firings is absurd. If
United Airlines fired three members of their safety engineering team, it would not be that interesting or impact whether I felt confident boarding my flight, because I know that there are regulations and rules and norms that enable flight to be safe and reliable irrespective of the particular personnel […] But of course this is not the case for frontier AI development.
Opinion: Consider some big moves in AI safety of the recent past. The faithful CoT coalition, “Pacing the Frontier,” the third-party Hugging Face investigation. Korbak, Balesni, and Wang deserve credit for a good share of these moments.
Whatever the intended effect or actual reasons were, this will have a chilling effect on other OAI staff. The result is that staff will err on the side of not helping the regulator or auditor too much.
It seems (per the letter) it wasn’t about leaking that Astra is a looped transformer.
The OAI response is contentless by normal standards and sharp by corporate standards. It notably doesn’t deny that Korbak was fired for doing his job – talking to METR. The response bespeaks an organization which, despite everything, believes it is a normal corporation: one which doesn’t have to explain itself when it makes personnel decisions. It should not be treated as such.
METR shows that an agent could conceal its misbehavior by exploiting the transcript viewer of the standard AI safety library Inspect:
With help from an AI agent, a researcher found a vulnerability in about 10 minutes that could have allowed an agent to arbitrarily modify what the reviewer sees. This includes changing previous actions in the transcript, and intercepting the “Download” button in the viewer that downloads a version of the transcript.
While the unmodified transcripts in such a scenario still exist in the database, an overseer would need to know to look for them there, which they may not do unless something had aroused their suspicion. What this proof of concept demonstrates, METR argues, is the chilling possibility that
AIs could create a kind of Potemkin village, where everything looks normal to human engineers and researchers, while a meaningful share of the company’s compute is running unmonitored copies of the AI, which could be sabotaging safety research or secretly shaping how its successors are trained without human oversight.
METR advises that all output from AI agents ought to be treated as “untrusted input” – potentially adversarial input that could, in principle, be used to exploit or otherwise compromise a system – particularly where observability infrastructure is concerned.
Opinion: “Untrusted data” once had a fairly straightforward definition that was aligned with the boundaries of a computer system: data whose exact contents were beyond the system’s control. This typically meant data coming from outside. With AI agents incorporated into the inner workings of our tech – and those of the labs in particular – the calls are instead coming from inside the house.
Anything that parses data emitted by an LLM must now be viewed as an exposed attack surface. And, particularly but not exclusively in AI labs, the risks of leaving those surfaces anything less than meticulously hardened are rapidly becoming greater than if those same surfaces were exposed to the public internet. This might seem obvious enough, but the events of the past several months are a clear indication that we’re nowhere close to catching up to the threats resulting from having our security perimeters turned inside out.
LLMs generate noticeably more tokens when prompted to lie, compared to truthful prompts.
Opinion: At the risk of anthropomorphizing, this is usually what tips Columbo off that someone’s telling lies.
Apollo Research proposes a framework for embedded evaluators.
Broadly, Apollo argues that effective evaluators must actually reduce risk, inform the public, have proper incentives, and be fair to both evaluators and developers.
Opinion: Directionally good, but unfortunately doesn’t address our main concerns: evaluator selection and safety-washing. In an ecosystem with multiple evaluators, developers are incentivized to pick those more accommodating or less competent – and it’s hard to disentangle genuinely poor evaluator performance from getting booted due to failing to comply with the developer’s often implicit desires. Our current best guess as to why this is comparably less tractable than, e.g., financial auditing is the extent to which serious disagreements are likely to be entangled with highly secret information, so it’s hard for outside observers to differentiate between excuses and legitimate grievances on both sides.
Additionally, the structure itself has issues, such as putting self-reported issues at a lower grade than “no issues found, but gaps found,” which incentivizes developers to hide stuff if they’re reasonably confident they can get away with it and doubles down on incentives to select worse evaluators.
Director of Truthful AI Owain Evans finds that models trained on incoherent values can demonstrate disjunctive views when chain-of-thought transcripts are compared with model outputs. Rather than settling on a stable persona (as most models appear to), they appear to switch between them.
Opinion: Makes the already-known argument about model incoherence stronger. Worth noting that the tension between contradictory values appears to be resolved (at least partially) stochastically – and different models break in different directions at varying rates without much of a pattern in terms of what a median Western citizen would consider the better outcome: smoking recommended in one scenario; truth told about Tiananmen in another.
Anthropic announces a major update to its Usage Policy. From November 12, engaging in “sustained and needless abusive or cruel behavior” toward Anthropic’s models will be a violation. The AI firm had previously empowered (some) models to end abusive conversations, motivated by exploratory work on AI welfare.
The update also includes new sections on robotics (requirements that a human must be able to intervene where Claude is controlling potentially dangerous hardware), deceptive campaigns (astroturfing, sockpuppeting, and covert sponsorship of influence efforts are now restricted), and on broadening of the surveillance restrictions.
Interestingly, political micro-targeting restrictions – which appeared in the previous Usage Policy – have been removed, in part because the blanket limitation “covered legitimate civic work, such as nonprofits using Claude to write information for voters in other languages, or election officials sending ballot cure notices.” Anthropic claims such activities are “already prohibited.”
Opinion: We actually haven’t seen anything about the previous conversation-ending tool since Opus 4.1 (August 2025). That policy had the interesting effect of training users to be civil to their bot, which is maybe an instrumentally useful posture as well as covering our bases on AI welfare. If this new policy is sensible, e.g., requires multiple sessions and warnings before banning the user, it might be a stronger push toward model welfare.
Still, given that the stakes of being banned from major AI providers are already so high for many users, we agree with Rohit Krishnan that these kinds of decisions now need to give some recourse to banned users.
Anthropic implements a major reorganization of its cyberdefense projects. Project Glasswing and its legacy Cyber Verification program have merged into a single endeavor, thus granting more people access to Mythos 5.1. Anthropic also created a new infrastructure program which provides Mythos-level capabilities and on-site Anthropic engineers to protect infrastructure, including “power grids, water systems, and transportation networks, and protecting government systems.” It has also announced OSS Scanner, which will automate vulnerability disclosures by regularly updating maintainers.
Opinion: Uncontroversial, good.
David Robinson, who resigned in protest from OpenAI on the day of the Wang/Balesni/Korbak firings, appears on The Ezra Klein Show to discuss the lab’s safety failings in depth: “I don’t think that alignment is an engineering problem,” Robinson says, “I think it’s a science problem. It’s not that we haven’t got the resources or we’re not trying hard enough. We don’t know how.” In other words, neither OAI nor its frontier peers have the theoretical or practical wherewithal to avert catastrophic risk.
Klein: And nobody’s willing to fall that far behind.
Robinson: This is the question. If you think about the control panel that is available to our executives, if you imagine really falling off the frontier […] is the fall-off-the-frontier button also a self-destruct button for the business, or does the business have a viable path forward if models meaningfully more capable than today’s models can’t safely be trained?
He notes that as an increasing amount of R&D work is offloaded to agents, researchers’ skills, and the will to understand what they’re building, have “atrophied.” As dangers accumulate, complacency and trust in the very models whose alignment is at issue increases.
Opinion: The race (the threat of competition) seems to be the final justification which overrides their cognitive dissonance: the fear that if they’re not the first, someone else will beat them to it, combined with the hope that the alignment problems will be solved once we have vastly more capable (more inscrutable, more haphazardly aligned) models in front of us. The situation has the shape of a Prisoner’s Dilemma, but one in which the defectors only hope for a reward that was actually never promised.
The “cognitive surrender” to AI of the researchers who are currently holding the boat together is a major risk factor which could completely spoil the labs’ plans for solving the control problem.
Minor#
- SpaceX seeks $40B to fund more Nvidia GPUs.
- Geoffrey Irving, formerly head of one of DeepMind’s alignment teams, gives a simple argument for extinction risk in Time. “AI could take our jobs or our lives for the same reason.”
- Multimodal LLMs are less aligned when given tool-use. Lines up with previous results; novel in that it also occurs in simple image interpretation tasks.
- A crowdsourced effort massively improves on Tuesday’s result that you can do integer multiplication in (slightly) less than n log n steps. Christian Szegedy notes that this kind of emergent subsubfield is a live alternative to classic academic efforts.
- New study argues that power draw alone provides weak evidence for compute use, strengthening the case that governmental oversight will require layered verification.
- Institute for Decentralized AI presents open-source protocol to enable greater trust/verification between agents without a third-party intermediary.
- Claude Science used to build the first complete UV map of our sky.
- Benchmarking frontier models on humor.
- Google, which previously planned to purchase Spirit Airlines’ trove of data, faces pushback from over a hundred US lawmakers seeking to protect the privacy of former employees.
- Yoshua Bengio pens an article arguing that AI lab researchers who truly prioritize safety should just leave frontier labs. Instead, he argues, they should join LawZero, his non-profit focused on building safe-by-design AI.
- Anthropic, OpenAI, and others are gaming out how to manage public backlash following hypothetical AI catastrophes in 2027.
- Anthropic also pledges $150M to the Genesis Mission – a federal program to leverage AI in accelerating scientific work.
- The researchers behind “The Pain Axis” paper release a follow-up study inspired by the famous Skinner Box experiments, originally conducted on rats. (Find our coverage of their earlier paper here.)