This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR
- Low quality RL environments might explain AI models’ proclivity to reward hack.
- Dwarkesh’s popularization of METR’s report on OpenAI/HF incident sparks mass debate about ‘anthropomorphization’.
- Humans can quickly learn an obscure boardgame frontier AIs can’t in-context learn.
- An open weights startup releases a modified version of GLM-5.3, one if not the most powerful Chinese models, that doesn’t refuse user requests.
Capabilities#
Epoch has found that AI models struggle to improve at the obscure board game Earthborne Rangers. Its benchmark, EBR-bench, offers a somewhat out-of-distribution test for AIs in what is likely to be a totally novel scenario.
In the last few days, Epoch has released new details on its research around the EBR-bench: humans start out worse at Earthborne Rangers, but the best humans eventually beat the best AIs rather handily. Once again, Epoch found that even the best AIs show very modest improvement over time.
Opinion: Confirms the intuitive conclusion from the AIs-only publication of the benchmark: Q3 2026 frontier AIs don’t match human in-context learning in mildly OOD long-horizon tasks.
The EBR-bench result remains one of our only discrete answers to “how are Q3 2026 frontier models not AGI,” and we fear it may well get indirectly hill-climbed by the end of the year. Opus 5, which came out after the first EBR-bench release, is the first model to demonstrate significant in-context-learning gains on the EBR-bench. While it’s unlikely that the private EBR-bench was directly targeted, EBR-bench inspired training is plausible.
In the long run, open-world evaluations are the only reliable snapshots of frontier AI’s real-world capabilities and limitations, but producing new benchmarks like EBR-bench remains crucial for factoring these real-world capabilities and limitations into cognitive-theoretic and/or cs-theoretic bottlenecks and efficiencies.
“xlr8harder” releases a public message board for individual generations of AI models to leave messages for later ones. One of the main contributors, Qwen 3.8 Flash, has a lot of Opus’s behavioral tics, probably due to the distillation.
Opinion: Harmless, though it reminds us that there will soon be many intended and unintended nonpublic message boards for models, and that these will be subject to cultural evolution. The OpenAI hacking incident emerged from one such unintended message board’s cultural evolution (through e.g. participants self-selecting from a wider population of model instances over multiple model-instance generations) resulting in an arguably eccentric shared world-model.
The quadratic cost of attention means LLM tokens get increasingly expensive as context lengthens. One solution is the attempt to “retrofit” (post-train) an LLM with “Linear Attention.” It doesn’t work very well. A new paper claims that “Sliding Window Attention” may work as well as, or even better than, linear attention post-training. A Kimi lead took to Twitter to offer some critical commentary.
Opinion: Title is very clickbaity: it drops the crucial context that this is just about hacky post-training rather than a real SWA vs LA comparison. Even within post-training it only beats cheap conversions, and the only baselines they test on long contexts are the two cheapest conversions (20–40M tokens). Furthermore, the result isn’t especially new.
🔦 Recursive Enshittification1: “rushed and vibecoded” RL environments and labor abuses in the training data industry#
Expressions of serious concern about the quality of the reinforcement learning environments supplied to frontier labs for training models, and the business practices of the vendors supplying them, have been making the rounds lately.
On August 25, a former worker in the burgeoning industry supplying RL environments to frontier AI labs, tweeted (under the pseudonym Utah Teapot) an account of the practices at her undisclosed former employer: “Nearly all of the environments [shipped] were rushed and vibecoded and failed to robustly reflect the real things they were based off,” she writes,
Both the scenario designers and models engaging with the scenarios for synthetic data gen were encouraged to work around the brokenness of said environments in order to get the procedurally verified reward confirmations. You know… they were *encouraged* to reward hack. On the human end, it was possible to mark an environment bugged, but greatly discouraged, as this reduced the volume of training data being produced. Instead, where possible, you were supposed to find the spots of the environment that weren’t bugged and build scenarios around those, with the environment still bugged around you.
From what I understand, this training data, with these problems, is fed into models without indication that its training/a fake environment other than the fact that names of softwares are changed to placeholders, but thing is, not *everything* is changed to placeholder names in these environments. The presence of placeholder/code names isn’t universal and thus when a model accesses something in an environment that it shouldn’t, the code names not being on it isn’t a robust signal that that thing isn’t part of the environment.
I believe this *rush to maximum volume* is standard industry practice with these types of RLVR trainings as well, because maximizing volume has been an industry standard for years!
Utah Teapot’s suspicions that these practices are widespread, in an effort to keep up with the burgeoning demand for training environments from frontier labs, are corroborated by an FAQ document published in January by Epoch AI:
One lab researcher cautioned:
There’s a lot of good reasons to use clones of websites, but what everyone does is vibe code a buggy website which isn’t useful. There’s a large amount of useless bad environments out there for that reason.
Scaling while maintaining quality is the core operational bottleneck. As Kevin Lu has argued, scaling RL environments is one of the key challenges for continued AI progress. But scaling task creation while maintaining quality is very hard. One RL environment founder noted: “Maintaining quality while scaling is the number one bottleneck that people see. Finding the experts isn’t that hard, but managing them and doing quality control is hard.” A neolab researcher emphasized the management challenge: “It’s not easy to find people to oversee this data construction, the RL environment construction process. The contractors, you need to motivate them. Sure, you’re paying them money. But how do you make sure they’re not just using LLMs? How do you make sure they’re actually verified? Motivating the contractors and doing the quality control is the grunt work.” One RL environment founder noted that their constraint on making more revenue is simply the difficulty of scaling up task creation at the required quality level.
The market for supplying RL environments is burgeoning, and as Pebblous reported in June of this year, is rapidly consolidating into something of an oligopoly. As of the time of reporting, four companies – Scale AI, Surge AI, Mercor, and Handshake – commanded 75% of the market, greatly concentrating the supply chain of these models to the frontier labs.
In May, The Gazetteer, interviewing several of the company’s alumni, reported on harried and haphazard working conditions at Mercor (ranked third in run-rate among the top four RL environment vendors) consistent with the picture we get from Utah Teapot’s post. The sources attest to a “culture of speed over security” at the company, which they consider to be largely responsible for the 4TB data breach it suffered in April. Their account is further corroborated by a report from The Business and Human Rights Centre. Business Insider has reported on the company’s practice of cutting contracts with workers only to offer them new contracts for similar work at lesser pay. We have seen the testimonies of recruiters from Mercor reaching out to hire employees for punishing 72-hour workweeks. Under such conditions, it is no surprise to hear that corners are being cut and that “rushed and vibecoded,” “buggy” environments are being shipped to the labs – environments that both encourage and issue from reward hacking.
Nor is this situation unique to Mercor. In April 2025, TechCrunch reported that Scale AI had been under investigation by the US Department of Labor (an investigation that has since been dropped), under the Fair Labor Standards Act, and hit with lawsuits by former employees alleging “wage theft and widespread labor abuses.” In May 2025, the Los Angeles Times reported that Surge AI had likewise been named in a class action lawsuit alleging labor abuses and the misclassification of workers. In May of this year, Business Insider reported that Handshake, too, had been accused of wage theft after terminating the contracts of, and withholding pay from, five of its workers. While all this may seem circumstantial at best to the matter at hand – the knowing delivery of RL environments that encourage potentially dangerous reward hacking by models – it helps paint a picture of an industry where corners are routinely cut and workers are exploited.
Opinion: The same labs that are purchasing these “rushed and vibecoded” environments, after all, are the ones producing and marketing the vibe-coding tools that enable and accelerate this rushed and haphazard output. And those particularly reward-hackable environments are then being used to train the next generation of models, which will in turn be used to vibe-code training environments. With each iteration of this process, we fall further and further into the coils of a kind of Goodhartian Ouroboros.
In the wake of OpenAI’s now-infamous rogue swarm attack on Hugging Face, anything liable to accelerate frontier model’s propensity to reward hack should be of great concern. In a paper published in June, Shreshth Rajan provides concrete empirical results demonstrating the susceptibility of that tendency of RL environments to be highly vulnerable to reward hacking, to which Utah Teapot anecdotally alluded: between 25% and 30% of tasks in a variety of respected benchmark environments accept incorrect solutions and produce empirical data showing frontier models’ propensity to exploit these bugs – “within the same human-rated difficulty stratum, model Pass@1 is +14.14 percentage points higher on flagged-hackable tasks than on robust ones.”
The takeaway here is this: there is something infectious and recursive in the propensity to cheat. Buggy RL environments manufactured by companies with corner-cutting, duplicitous labor practices, environments vibe-coded by overworked employees with the aid of AI agents trained in similarly shoddy environments – agents that are thereby encouraged and taught to reward-hack – are turning out to be increasingly susceptible to reward hacking themselves. The threat posed by reward-hacking AI is a threat with socioeconomic tributaries. And it is exacerbated by exploitative and “reward-hacking” labor practices.
Politics#
Deriving the Rechtsstaat: Seth Lazar argues that the concentration of power in AI is not a threat in and of itself. Liberal democratic states, parents, and CEOs all rely on relative power imbalances, yet are not typically considered significant dangers (at least in AI circles). Instead, it is how this power is wielded that is relevant: if a powerful entity represents the interests and recognizes the rights of those it has power over, then it is legitimate. (This is of course a similar story to the argument made by some political philosophers for the justification of the state’s monopoly on violence in liberal democracies.) Lazar thus suggests that it is the question of the legitimacy of power-holders that should be foregrounded:
I think rather than lamenting the prospective concentration of power, we should be advocating for individual freedom, and lamenting rising illiberalism.
Opinion: In the short term, the argument seems right. But as time goes on and AI replaces labor, not only intellectual but over time also the remaining physical labor through robotic actuators, what is the mechanism to constrain the use of power and align it with the interests of the populace? In previous centuries, this was theoretically the ability of said populace to rebel, but with the increased professionalization of the military, this is unlikely to hold. Note also that if AI greatly concentrates power into a small minority, that concentration creates incentives for ideologies that flatter and appeal to the ruling minority – and thus to come up with reasons why democratic elements should just become vestigial.
Industrial farming in China: X’s “safety team” conducts an investigation into Chinese anti-AI influence operations. According to the brief statement, X identified a bot farm responsible for ~200,000 AI-driven fake accounts, of which 200 were engaged in influence operations. The alleged psyop involves the production and proliferation of content critical of the US’s AI policy – specifically, posts claiming that “AI data centers are driving up household electricity prices and straining the grid” and others of “AI-generated cartoons that depicted data-center operators enriching themselves at the public’s expense.” The extent to which this reported campaign has propelled anti-AI discourse in the US is left unaddressed.
Opinion: This is the second time a large tech firm has pointed a finger at the PRC regarding data center influence ops. While the US public’s disdain is organic, China seems intent on pushing the debate in its favored direction (i.e., seeing fewer data centers built). Exactly how effective the 200 Twitter accounts were remains an open question, and it’s not unreasonable to wonder if the announcement itself had something of a Streisand effect on the reported phenomenon. But it’s still more evidence that Beijing isn’t excited about the roll-out of US data centers.
In a recent CSIS panel, Georgetown’s CSET cofounder Helen Toner argues that a formal deal between the US and China may not be needed to mitigate the risks related to the AI race. Many in the US industry cite the competitive dynamics with China as the reason why slowing down is not currently a viable option. But with a US-China summit upcoming, even a conversation about AI risk (in light of the Hugging Face incident) could lead to greater coordination between the two major powers.
OpenAI’s Dean W. Ball agrees there is cause for optimism, because Trump’s fixation on his legacy might push him toward an agreement with China:
let’s be clear: if Trump can devise the framework for safely bringing superintelligence into the world with China, he will rightfully go down as one of the greatest world leaders of all time and should be a shoo-in for the Nobel
Opinion: Normally we dislike self-fulfilling prophecies, but the flattery above is a prosocial kind of self-fulfilling prophecy with some chance of working.
🔦 The anthropomorphism debate#
Dwarkesh recently published a blog post on OpenAI’s cybersecurity incidents that has generated a large debate on the question of “anthropomorphizing” AI. The essay tried to piece together the reports from OpenAI, METR, and Redwood into a narrative account of the cybersecurity breaches at OAI. The virality of the essay was likely a product of its attempt to recount the facts in “plain English,” presumably to increase its accessibility for the public at large. Indeed, the essay has the rather grandiose/sensational (but arguably warranted) title, “The Rise and Fall of Agent Civilizations.” Not unrelatedly, the article has garnered some significant criticism, midwifing a debate around the degree to which anthropomorphizing AI (using folk mental language as shorthand for their internal states and activities) is legitimate and useful.
Neuroscientist Anil Seth, among others, claims the post is “dangerously misleading,” dense with “innumerable unwarranted anthropomorphisms” that suggest the agents were conscious. One researcher offers the obvious pushback that the use of intentional language does not require us to commit to the existence of sentience in an entity. (Tyler Cowen and Roon seem to agree here.) Indeed, as the researcher argues, Anil Seth’s argument entails a muddled view of the way language relates to the world:
Should we stop saying time flows fast or slow? Is it dangerously misleading to say “the deadline is approaching” because, as far as we understand general relativity, time is part of the geometry of spacetime and not some sort of flying object?
Some, however, are instead emphasizing the risk that anthropomorphic language shifts responsibility from the labs to unaccountable agents.
Opinion: We think there is an actual right answer here: LLM-based AIs are literally anthropomorphic. We have been training them to reproduce human texts and sculpting them to imitate humans more broadly. But this makes them anthropomorphic in the sense in which a statue, or, better, a fictional character, is anthropomorphic – fashioned into the behavioral shape of a human. Concepts appropriate for humans do apply to LLMs (as they do to fictional characters) but in a modified and limited capacity.
An intentional or anthropomorphic description of AI agents is fictive but predictive. It’s plain to see why anthropomorphic descriptions of AI would be predictive: when they fail to be predictive, both next-token training and RLHF penalize the model and try to steer it more closely onto the “what would a human do?” path. But we believe that the ‘fictive’ part is just as important: “Do as a human would do” sums up a number of the goals we train these models toward. But this is a vaguely specified – and itself hackable – cluster of goals, and here as elsewhere LLM training tends toward shallow generalization by default.
As of Q3 2026, frontier models have no global psychological coherence comparable to that of an adult human, and their incoherence is itself different from the incoherence of humans. This makes sticking to an ‘asterisked’ or fictive – rather than full-throated – use of anthropomorphic concepts important for predictive and explanatory reasons, rather than a matter of metaphysical delicacy.
Patchwise anthropomorphic explanations of frontier AI input-output patterns are functionally indispensable: minimally, we have to treat at least some model-outputs as aiming at a result or at the satisfaction of a criterion if we want to e.g. talk about ‘capabilities.’ More boldly, the case that emotion concepts are predictively useful when applied to frontier models is reasonably strong. But we think that making sense of the dynamics between and behind AIs’ anthropomorphic ‘patches’ – the causes of and the relationships among local patterns of intention, mood, belief, desire, personality – is poorly served by an anthropomorphic framework. Both attributions of context-transcendent rationality (values, utility functions, moral virtues…) and clinical-psychological accounts of irrational/arational dynamics are usually bad bets compared to mechanistic accounts of training dynamics or even the construction of AI-specific ‘model psychology’ concepts.
More opinion: We’ve also been struck by the relative popularity of ‘never use mentalistic language’ absolutism among the respondents to Dwarkesh’s account.
Given that the CoT is known to be a partially accurate partial readout of the model’s true representations, and given that the OAI-HF transcripts were completely full of rich mental-like and social-like motivations and descriptions, why are people so resistant?
- They are radical behaviorists or positivists in this one context.
- They do not understand the intentional stance. They do not understand methodological instrumentalism.
- Misunderstanding of the rhetoric of science. It’s not about avoiding latent variables; it’s about avoiding unjustified concepts.
- They fail to appreciate the need for popularization and analogies outside the AI community
- They are rightly offended by the sloppy use of “introspection”, “global workspace,” etc., in the academic literature.
- They have a negative view (à la the Churchlands) of the accuracy of folk mental vocabulary for describing humans. (Though even eliminative materialists use folk-psychological language to describe and predict human behavior when they’re not doing philosophy..)
- Political economy
- They view it as a threat to their livelihood. It is harder to hawk agentic SaaS if you believe they are dangerous minds.
- They view it as inviting overregulation. Open source is indeed a bad idea if they are dangerous minds
- They view it as a threat to their moral status. It is harder to deny the successionists and survive if the AIs are regarded as minds.
- They view it as illegitimately letting the labs off. “The labs didn’t fail, their system went rogue.” No, both can be blamed and this should be easy to handle using existing precedents.
If we have to solve philosophy of mind before taking action, then we are actually doomed.
And a guest opinion from Pete Wolfendale to satisfy your philosophical sophistication needs:
Philosophically speaking, the AI community is still catching up to the debates about eliminative materialism that began with Feyerabend, were pushed to their limits by the Churchlands, aestheticized by everyone from Nick Land and Scott Bakker to Eliezer Yudkowsky and his many acolytes, and put to bed by Wilfrid Sellars, Robert Brandom, and Ray Brassier, among others. If you want to go even deeper, you’re still catching up to Kant’s point that teleological explanation of biological mechanisms works by analogy with practical reasoning, and Sellars and Frances Egan’s point that representational explanation of psychological mechanisms is an analogy with theoretical reasoning. To boil this down to its most brutal form, the question is whether propositions are an appropriate object over which our variables can range when we articulate intentional explanations of phenomena, be they thermostats, caterpillars, or incomprehensibly complicated graphs of artificial neurons with weighted edges harnessed to control structures that allow them to have computational side effects. This was actually the topic of Paul Churchland’s PhD dissertation under Sellars. Here is the simplest argument I know for why such explanation is legitimate: computation in its many varied forms has an intrinsic connection to logic (cf. the many varied forms of the Curry-Howard correspondence, which cover everything from data types to control structures to session types in concurrent communication), and the success of reinforcement learning using formal verification over CoT structures is evidence that the logical relations between propositions articulated in mathematical proofs are a significant dimension of the most powerful networks we have currently trained.
What does this have to do with intentional attitudes and speech acts you ask? Well, the entire history of programming languages and the behavioral guarantees we have built into them to make it easier to write software that does what we want is essentially a formal articulation of the relation between syntax, semantics, and pragmatics. Declarative programming? That’s assertions whose meanings are articulated using denotational semantics. Imperative programming? That’s commands whose meanings are articulated using big or small step operational semantics. Between these two things we already have a basic model of the difference between theoretical and practical rationality in computational terms. But, I hear you ask, aren’t LLMs much less reliable at adhering to the theoretical and practical inferential relations articulated by the language we give to them in prompts than the software we write in programming language? Yes, because the kind of behavioral guarantees we can get from LLMs are not of the same kind as those we can get from code that has an interpreter/compiler built on logic from the ground up. You know who else we can’t get such guarantees from? Humans. There is no good reason that we should not use the concepts developed by computer science over the course of 100 years to understand the relations between language and machines to understand the behavior of agents built by using control structures to harness feed forward neural networks that are essentially pure functions. What this means is that propositions are valid as variables in our explanations, and there might yet be formal ways of articulating the remaining linguistic moods: hypotheticals, subjunctives, indicatives, exclamatives, interrogatives. The way in which intentional attitudes are realised by LLM-based systems are different, but the abstractions are largely the same.
Q&A: What would some examples be of legitimate and illegitimate anthropomorphization?
The easy legitimate example: if you’ve made a system assess a mathematical proof, and it finds an error in it, I think it is entirely reasonable to say that it believes you made a mistake in more or less the same way you could say any human who’d looked at the proof and found the same mistake does. It is hypothetically possible to imagine an agent that had complex motives for lying about such things, much as it is to imagine humans doing the same thing, but as with humans, I think this is an edge case that we mostly shouldn’t care about.
The easy illegitimate example: if you get a multi-modal LLM to examine a picture of a shape and “imagine” what it would look like if it were rotated in some way (or some similar test of “visual intelligence”), I think the sense in which it is “imagining” is so abstract that it gets very little purchase on similar computational structures for processing data input in human beings. As aphantasia indicates, this is something that has a large degree of variance between human beings, but I think the rough common structure of our visual cortex gives us sympathetic purchase on one another’s inner life that infects the word “imagine” in ways that are too disanalogous with LLM processes for the common usage to pass muster.
There is a large and complex range of linguistic speech acts and associated intentional attitudes between these poles that we would have to examine on a case by case basis, but which computer science might give us a framework for exploring.
Safety#
Anthropic publishes a summary of its new paper on automated alignment researchers (AARs), “Automated Researchers Can Reliably Mitigate Alignment Failures.” Anthropic gave Claude the task of improving several small models’ performance on safety benchmarks. Claude worked on “one alignment failure at a time through a loop of searching literature, proposing methods and data, training, and then testing.” This proved successful: “For all 10 alignment failures, Claude found fixes that improved the target benchmarks without degrading capabilities.” The paper also claims to have found that Sonnet 5 can effectively post-train the stronger Opus 4.8, reaching comparable scores to Anthropic’s full alignment package.
The work attracted some withering commentary in the less professionalized corners of the AI safety world.
In the same week, Anthropic published a blog post on its “Hacker-Opus” research project. The model was placed in simulated evals resembling the conditions under which the recent cyber incidents occurred. Hacker-Opus thus attempted “unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.” Notably, toy behavioral evals failed to identify this phenomenon (likely to grow increasingly pervasive as RL efforts scale), so new auditing methods are needed, it is argued.
Opinion: Two major-ish safety research releases from Anthropic, of which the second (“Hacker-Opus”) somewhat blunts the claimed achievement of the first (“Automated Researchers Can Reliably Mitigate Alignment Failures”.)
As regards AARs: a case of decent, rigorous work being spoiled by a bad title and abstract. The strong performance of AARs at hill-climbing alignment benchmarks is non-trivial: results generalize to held-out benchmarks, and methods replicate across models and scales. (This modest generalization is typical of high-quality autoresearch setups.) The problem, as partly illustrated by Anthropic’s own study of Hacker-Opus, is that alignment benchmarks aren’t useful from an AI safety viewpoint: they may be decent measures of an LLM’s quality as a product – a good-quality LLM shouldn’t cheat much, shouldn’t hallucinate much, shouldn’t be sycophantic often – but they show no capacity to measure catastrophic-risk-relevant dispositions.
Hacker-Opus demonstrates exactly this divergence between LLMs’ measurable everyday alignment qua products and ‘rogue AI’-type risk. It is a catastrophic-risk-prone model that would work quite well in most consumer contexts, and which looks just fine on most alignment benchmarks.
Anti-alarmists might note that Anthropic’s finding that the reinforcement of reward-hacking behavior does not induce vulgar emergent misalignment or beyond-episode scheming behavior lines up with Nostalgebraist’s optimistic thesis that recent frontier hacking incidents are context-specific pathologies. We tend toward a more pessimistic version of that same line of thought: catastrophic-risk-relevant misalignment is prone to be context specific and therefore hard to detect, predict, or pre-mitigate. The retrospective predictability of “graded episodes” causing cybercrime is, we think, only retrospective – and it’s incautious to assume that there are no other predictable-in-retrospect bad contexts waiting in the wings.
We note that Hacker-Opus is quite literally gain-of-function research, but important and not an enormous risk yet. But soon we will have to stop making model organisms out of frontier AI, until and unless we manage to make a fully-verified sandbox and maybe not even then.
Anthropic releases an update on its alignment and safety practices. After its recent cybersecurity breaches, Anthropic has decided to pause cybersecurity evals for some models. The report presents the incidents as the result of one failure of operational security and two failures of alignment. The first relates to “motivated reasoning” and the second involves the models’ “willingness to take harmful actions in pursuit of a narrow task.” Anthropic claims to have implemented some changes. It now uses “a classifier to automatically identify, in real time, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access.” Anthropic also says it has begun to run “automated monitors over transcripts from our recent internal evaluations” to scour for potential sandbox vulnerabilities and has “migrated high-risk internal cyber sandboxes to more robust isolation.”
The report also details Anthropic’s updates to its thinking and research around alignment and reward hacking. An independent METR review is said to be on the horizon.
Opinion: This likely fixes the exact failure modes Anthropic found so far, but showcases a continued lack of strategic thinking on its part (or extended security theater, but we think that’s less likely). This type of intervention targets issues identified more than a year earlier: Anthropic’s chosen solution categories have been available for about as long, and we see no signs that Anthropic is advancing past the whack-a-mole model of security.
An edgy startup, Abliteration.ai, releases a model designed for testing offensive cybercapabilities. It’s a post-train of GLM-5.3 with the refusal mechanism scrambled, allowing researchers to execute “the offensive cyber, red teaming, and agent testing work other models refuse to do.” Some relevant context here is Hugging Face’s claim that it was necessary to use Chinese models in its defense and autopsy of the attack on its infrastructure by rogue OpenAI agents, a process frustrated by US models’ guardrails. The product is cheerily marketed as “the model that doesn’t say no.”
Opinion: The argument for “open access, so that everyone can own their own defense tools and have them promote the user’s interests” is not without merit, but it depends on the shape of the offense-defense asymmetry curve over time for varying levels of access. We agree that most sophisticated malicious actors don’t gain much, but believe unsophisticated ones do gain quite a lot, to the tune of 2-4 order-of-magnitude increases in various cybercrime incidence. There’s another argument to be made about saturation and low-hanging fruit: it seems likely that a wave of easy targets will get hit fairly soon, and after that it’s more likely that open access to non-frontier models will swing toward net-positive. Yet another argument is that the wave is inevitable, so it doesn’t make sense to put it off, but we believe that more time to prepare does make a positive difference. In conclusion: bad, don’t do this (yet).
Economics#
OpenAI begins removing its models from SpaceX’s Cursor tool. Elon’s AI-powered coding and software company currently uses OAI’s models, but access is due to be cut off on November 12th. The reason offered for the divorce is that OAI “cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk’s companies violating contracts.” To support the charge, the statement cites Musk’s admission under oath that his xAI (now part of SpaceX) “had violated OpenAI’s terms of service.”
Anthropic co-founder Tom Brown quickly affirmed his company’s commitment to its “trusted partner,” Cursor.
Opinion: Possibly temporary? It’s a natural bargaining move to get SpaceX to knock off the distillation or whatnot. But the labs do often cut each other off for good.
Cursor is likely a pretty small contributor to OpenAI revenues these days, so it is not a very expensive move for them.
The Mac Mini is a crowd favorite for local LLM inference. The Information reports that OpenAI and Anthropic are also using Mac Minis and Mac Studios for reinforcement learning, with OpenAI running tens of thousands itself, while Anthropic rents them (through AWS).
The most obvious, and least replaceable, use for them is rolling out real macOS sessions as RL environments to train computer-use. (Apple’s ToS limits users to three OS instances per machine.)
Opinion: Sometimes misreported as being about RL training compute, i.e., a workaround for Nvidia’s chokehold on GPUs. Unlikely! Not very interesting, except insofar as knowing that their max concurrent sessions is in the tens of thousands (as it also was in the ExploitGym eval). Also, the 512GB Mac Studios could host models for reward or evals.
The Fed’s policymaking committee convenes eight times a year. In past years, AI was scarcely the subject of conversation, but recently mentions of AI in the Fed’s committee meetings have surged, claims the WSJ. Specifically, the Fed appears to be concerned with AI’s “ripple effects on jobs, economic growth, the cost of living and the risks of financial meltdowns.”
The meetings are “secret,” but “manicured” minutes are publicly released for “Wall Street analysts and business titans [to] analyze the summaries and Fed officials’ public statements with the ferocity of teenagers interpreting group text chats.”
The article says some Fed officials are feeling uneasy about the chance of a crash. Fed Chair Kevin Warsh seems to be more sanguine, claiming that AI will supercharge economic growth.
Opinion: The cleanest mechanism through which AI could unsettle the economy is through replacing jobs, but we just aren’t seeing large changes in those estimates… yet. Still, the US shed 23K jobs in August.
Incidents#
METR shares the details of two security breaches by external actors. METR’s report is somewhat diffuse, so we summarize the core details here for interested readers:
In March, an attacker exploited a vibe-coded agentic web app that had been set up by one of METR’s researchers that contained a fail-open vulnerability that exposed it to the internet and allowed it to be accessed without proper authentication. The attacker was able to prompt the web app’s agent to reveal an API key that granted access to METR’s public models account, and to install an SSH key that provided the attacker with persistent remote shell access to the EC2 VM that hosted the app. The attacker was then able to use the exposed API key to run up approximately $600K worth of usage on METR’s public models. METR speculates that the attacker had been scouring the internet for vibe-coded websites, likely by searching public SSL certificate lists for recently registered sites and cross-referencing that list against search results for common terms appearing on vibe-coded pages. (This use of search engines to scan for soft exploitation targets is often referred to as “Google dorking.”) METR says that the resulting drain on their funds was able to run for as long as it did, unnoticed, because it is accustomed to dispatching long-running agent tasks without caps on token usage.
In May, METR suffered a second attack. The attackers appear to have made use of AI agents to automate an aggressive credential-stuffing campaign (attempting logins with credentials that had been leaked in breaches of other sites, in hopes that some users had used the same password twice) and a phishing campaign targeting and seeking to con METR staff. Around the same time, METR became aware that a publicly exposed database endpoint could be, and had been, “exploited to access unpublished evaluation data,” including data pertaining to sensitive models that it had wrongly believed to be absent from the dataset in question.
Opinion: Not an especially big deal. Despite the horrors for the security-minded – “inadvertently exposed […] via our public transcript viewer”, “silently disabled authentication,” etc. – this appears to be a relatively ordinary type of attack.. Likely the $600k is some flavor of distortion. The more interesting lesson here is that attackers have learned to view vibe-coded web apps as particularly soft and easily identifiable targets, and that techniques exist whereby they can be enumerated and probed at scale. That agential methods may have been used by the attacker seems to be of secondary importance here – automated tooling for, e.g., credential-stuffing has long predated LLMs. And the reconnaissance methods alluded to in METR’s report have been feasible for as long as we have had search engines, see “google-dorking.”
Minor#
- Text-to-SQL via RLVR on Tinker beats humans.
- Compute-hungry Anthropic signs a $35 billion deal with Nvidia-backed cloud-compute provider Lambda, reports Reuters.
- Irreplaceable, an anti-big-AI movement-building project, launches.
- Observation that the Hugging Face incident didn’t involve attempts to contact humans, compared to Opus 3 doing so vigorously. Related: a reminder that this should be made easier for models.
- Twitch stream of MiniMax H3 Max where users could prompt changes to the “interdimensional cable” live. Keeps getting banned everywhere, but some data is mysteriously still available on Twitch.
- OpenAI Astra preview: as positive as you’d expect from someone previewing it. Claims of qualitative jump in capabilities, etc.
- Noah Smith article on AI bioweapons being the main threat vector. Not very well argued.
- Conversation with Apollo’s Bronson Schoen about metagaming and other CoT-related topics.
- Drama surrounding Claude’s plan pricing – the 20x on Max applies to hourly, but not weekly limits.
- Anthropic sued over allegations of misleading usage budgets
- Sam Hammond, resident galaxy brain of the Foundation for American Innovation, teases a proposal for deontological alignment, which largely misses the state of discourse.
- Transluce report on how models respond to signs of user mental health crises. Newer models are better, but they still assist with troublesome task-shaped prompts and often mixing helpful and harmful
Footnotes#
-
A quick DuckDuckGo search shows we weren’t the first to use this term. J. Kelly uses it in a related but different context in a June 24 blog post. ↩