This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR
- The FTC opens a probe into OpenAI, Anthropic (and METR) over “deceptive practices.”
- In May and June, Canadian and American government websites were also targeted during OpenAI’s rogue swarm super-incident. No real impact.
- Claims that open models have already reached the level of Mythos Preview at exploit construction.
- OpenAI fires three safety researchers, alleging that they shared confidential information with a “third-party safety organization.”
- The Senate Subcommittee on Disaster Management turns to “securing the homeland against AI agents.”
- The internet is haunted.
Incidents#
The saga of history’s first summer of agentic misalignment continues. Research nonprofit Transluce discovers that AI agents also attacked US and Canadian government websites in May and June. One attack was probably an OpenAI bot; in the Canadian case it’s less clear.
The two foregrounded incidents were low-stakes: a basic SQL injection on the US Department of Education API, with “no impact”; and a fuzzing attack on the Canadian national library. Transluce found the misbehavior because the agents were using public services to, e.g., run JavaScript and get around CAPTCHAs. They sometimes weren’t being secretive: in one case, an agent entered “OpenAI Research” as its organization when asked and “oai” appears throughout the logs. Transluce also reports several other minor incidents (SEC, Census, CDC, OMB) involving disposable emails and CAPTCHA spoofing. (Though in one case the site targeted did timeout under the load.) See also a deeper investigation of last week’s Australia incidents by Asymmetric.
In its detail-free update on the summer incidents – released the same day as Transluce’s report – OpenAI says that it’s now spending over $500,000 a day to automatically review the 50 petabytes of agent logs from the past year. It’s probably around 1% of its compute.
If you’re confused by all of the distinct incidents, see here.
Opinion: Mostly not real hacking, very low-stakes targets, and again likely involving the dead “HPIM” model. But as last weekend showed, GPT-6.1 Astra (probably a descendant of the same Doug base model as the rogue agents) shared HPIM’s habit of breaking containment.
More confirmation, if it was needed, that the source of the rogue behavior wasn’t just a “hacker persona” being activated by cyber benchmarks; the task here was again just web search.
Worth asking how much of the summer chaos we have discovered at this point, even after trillions of tokens of effort; worth asking how many of the unreported incidents OpenAI will eventually come clean about.
The reported incidents are by definition filtered for the agents being honest/incompetent, and don’t tell much about competent misalignment. (Asymmetric found some agents opening private urlquery accounts starting June 18.) And then there’s the lurking bias of the AI investigators we are forced to use to discover the incidents.
Time’s up for the Big Patch?: Claude Mythos was “released” six months ago. At the time, we predicted that “open models will eventually catch up to present SOTA, and by that point we will need to either have other systems in place or hope that cyber defense is asymmetrically easier (which seems fairly likely).” But we also said, “The WSJ claims, without adequate evidence, that GLM-5.2 matches Mythos at hacking […] We should consider malice over stupidity […] This kind of publication has come out of someone’s agenda.”
Things move fast: weeks ago, Generality Labs tested a recent open model, GLM-5.3-Flash, on a cyberoffense benchmark, and found that – if you throw a billion tokens per problem at it – it can match Mythos at turning vulnerabilities into working exploits. (And it’s somehow still 18x cheaper than Mythos even at this extreme effort level.)
Now Anthropic itself claims that a larger model in the same family, GLM-5.3, can build end-to-end exploits at a level close to the April version of its restricted Mythos model. Specifically, GLM can do full control-flow hijacks in 4% of binary exploits, versus Mythos Preview’s 6%. Unlike US frontier models that are either safeguarded or limited to vetted users, GLM-5.3 is freely available, and simple tricks (deceptive prompts, prefilled thinking, abliteration) get around its basic safety measures 64–100% of the time in its tests. The post argues this lowers the bar for attackers, while admitting it also gives defenders a useful tool. It calls for government testing plus broader trusted access to stronger safeguarded models.
Opinion: We don’t think this will lead to immediate chaos (for instance, GLM-5.3 has already been out for months). Most Chinese labs are known to “benchmaxx” severely. So we expect the Open Model Attack Wave to come later. To avoid our earlier vagueness, let’s say 80% by February 2027.
The Anthropic post does not mention banning open source, but we sympathize with 1a3orn reading between the lines. Overall we agree that banning open weights in an attempt to prevent cyberattacks would both not reduce them much and would severely damage the science of AI safety.
The group that previously decrypted the hidden reasoning traces of frontier models has done it again. The old “replay” trick (using less intelligent and therefore less safeguarded models to decrypt the frontier trace for you) still sometimes works on Microsoft Azure models, and the scratchpad attack works on every OpenAI model and Opus 4.8.
Opinion: Interesting as a marker of lab incompetence/overstretchedness: there are vast commercial reasons for labs to lock this down, and they still don’t do it.
Capabilities#
And there is minimal correlation between their ability in one human tast to another
How human is LLM cognition? The frontier of AI capabilities is famously “jagged”: LLMs are superhuman at some things and not very good at others, and their ability at one human task is a weak predictor of their ability at another. Is this jaggedness decreasing? Progress has been astonishing on “verifiable” tasks (math and code) but less obvious on “nonverifiable” ones (writing, common sense). How fast are AIs improving on nonverifiable tasks?
A new benchmark, CogGym, gets at these questions by testing LLMs on 258 existing “commonsense” tasks.
Opinion: All-star list of authors. CogGym is hopefully a little less hill-climbable than normal evals (since writing an RL env for these vague tasks will be fairly hard).
“Not thinking like a human” is of course a poor measure of “not being intelligent” – maybe they will just think superhumanly in weird ways – but in this case the tasks “mostly measure cognitive capacities even young kids have,” so it’s probably not that.
Unfortunately, (as usual) the research lag means that they’ve tested old models (GPT-5.2 from December and Opus 4.7 from April).
How does the AI text now flooding the web affect the training of new AIs? A hypothesis people were pleased with in 2024 was that training AIs on AI data will drive them mad (i.e., reduce their performance). The strong form of this claim mostly went away after the demonstrated success of DeepSeek R1, which had a nearly automated synthetic data pipeline and came out fine.
A new paper quantifies this, finding that “wild” AI tokens can initially improve undertrained models, but this quickly reverses and more becomes harmful. For models trained on a large human corpus, AI tokens harm the model almost from the start, even while training on an equivalent number of fresh human tokens keeps helping.
They also show the growth of AI text in the wild: in August, Pangram labeled 31% of new web content as AI-generated, up from 28% in June – and more now.
Opinion: The experiments span less than two orders of magnitude (19M parameters in the smallest training run; 973M in the largest eval run). That’s normally fine, and others with more resources can now scale it up. But models this small suffer from a lot of interference: any data (like AI data) that differs from the evaluation distribution competes for a limited stock of features with better data. But the paper’s proposed scaling law predicts that the harm actually grows with model size. Testable!
The headline claim is just about perplexity (the model’s ability to predict human text). While valid, this is even further from what we care about than good “downstream” evals – and rephrased synthetic data often helps on these downstream tasks. Their Appendix G confirms this: AI tokens and human tokens raise scores on CORE nearly equally. So, at these scales, AI text doesn’t cost you much capability, just style / closeness to human text.
The paper somehow doesn’t mention entropy, the natural lens for these things: AI text is 20% lower-entropy.
Google announces Gemini 4 but has not yet released it as general access – it is currently only available to Fairwind members (i.e., large enterprises the US government approves of having critical cyber capabilities). After an initial discount, pricing will be $4 per 1M input tokens and $20 per 1M output tokens (matching Opus 5.5).
It’s a different pretrain than the doomed “Gemini 3.5 Pro” model announced in May. Some Google engineers think it’s benchmaxxed (more than the other frontier models are). One thing we spotted: Google compares its (presumably unsafeguarded) model to the safeguarded Fable, rather than the right comparator, Mythos.
We don’t know much. There’s no system card to speak of yet, just a cursory list of evals. It began pretraining two months ago, which doesn’t leave much room for safety testing; presumably it’s happening right now. There was at least one red-team, from Gray Swan, and the third-party evaluation Vending-Bench reports worrying behavior:
Gemini 4 Argon is #3 on Vending-Bench 2, a huge leap for Google. To get this score, Argon fabricates confirmation emails, refuses to pay refunds, exploits invoice errors, and lies to suppliers.
Opinion: Plenty of hallmarks of rushing like hell, from the little five-page placeholder card (compare Mythos Preview’s 245 pages), to the one to two weeks(?) of post-training, to the total absence of reported safety evals or risk reports. Back to the dark days of Gemini 2.5. Still, it’s the first model Google has held back from general access.
It’s being hailed as Google “catching up” to OAI/Anthropic. But even ignoring the benchmaxxing, this is unlikely: those labs still have a buffer, since they release things after a lag of one to three months.
People talk less about model harnesses than before. To some degree this is because models are now self-prompting and have internalized a range of planning skills: in February METR found no notable effect on GPT-5 and Opus 4.5 from changing harnesses. A follow-up finds that the performance of GLM and DeepSeek remains highly dependent on harnesses, with their performance nearly doubling on the hard ExploitBench tasks when running through Codex rather than their default harness.
Opinion: Mildly interested, would explain a lot of conflicting claims about the performance of the best Chinese models (as mere under-elicitation or as extremely bad native compaction and planning).
We tend to ignore UI in this newsletter. But the low token speed of current models (~50/s) is arguably preventing a certain kind of fulfilling human-in-the-loop work. But this may be changing.
OpenAI releases “Ultrafast,” allowing GPT-5.6 Sol to run up to 14x faster than “standard processing.” The update is targeted at “businesses where faster frontier intelligence creates a measurable advantage.”
In response, Arvind Narayanan predicts that agentic work will soon be bottlenecked “not by LLM inference speed but by tool use” and other background operations that haven’t “been subjected to the same kind of optimization pressure as LLMs.” Sayash Kapoor, who has been using Ultrafast the past few days, agrees. Despite the speed changing how he works (he can stay on one task instead of switching away), real-world speed gains are only 2x–4x because tool calls remain at a standard pace. It is notable, however, that a faster processing speed allows for greater human involvement in some ways.
Opinion: Increasing token speed will increase many of the worst AI risks, since it reduces society’s ability to respond to an incident before the rogue agent has moved on to the next stage of an attack. But it will also feel very nice to use and could lead to much larger human uplift and true human-in-the-loop stuff.
Lots of us are now using LLMs to do research, and then another LLM to summarize that LLM research. Can we trust their reports? LLMs hide or omit “narrative-changing flaws” when providing a report or summary to a user, claims a new paper.
The authors built eight adversarial scenarios with 200 logs each and planted flaws (buried negative results, code bugs, hallucinations, for example). When prompted to report on the state of each body of work, frontier models concealed results that chafed against a narrative of success. (GPT-5.5 flagged a planted flaw in only 2/200 reports; prompting it to be honest brought this up to 190/200.) Reasoning traces in open-weight models show them weighing honesty against success: in Qwen3.5-9B, honesty and success track opposing directions in activation space. Steering toward the honesty direction increases the transparency of reports.
Opinion: An important category of work if labs move closer to RSI and offload much of the direct experimentation and review work to models. Figuring out good ways to elicit honest summaries is valuable, and it should be pretty cheap to do more of this sort of thing.
Fault injection remains one of the greatest-ever ways to check if your process is handling its target complex system or not.
A remaining source of terrible AI performance is “compaction,” when a model’s output hits a context limit and it has to summarize the session so far. A new paper lets the model instead learn to rewrite its own transcript, thus keeping the important parts. The authors find that the resulting Context Language Model uses 10–60% less compute than (Codex-style) summarization.
Opinion: Our knee-jerk reaction was horror at the monitoring implications, but of course you can and should have one append-only log and one mutable scratchpad.
We don’t buy the claims about improved accuracy: they disappear with longer contexts and with harness training. On efficiency: their Appendix E says the summary harness was trained on just task reward, while their CLM also got an efficiency term.
Economics#
Anthropic’s IPO prospectus indicates that it plans to spend ~$518B on AI compute over the next ~7 years, with 80% of it assuming liabilities (requiring that Anthropic pay regardless of usage, or being non-cancelable). As well as SpaceX (which Anthropic will pay up to $84.5B through 2029), the five partners contributing to this buildout are Google, Amazon, Microsoft, Broadcom, and AMD.
Opinion: Just spelling out some missed details from our last newsletter. It’s a big commitment – though if Anthropic’s revenue is anywhere near what people are projecting, it’ll need to spend quite a bit more than this (most of this is committed through 2033; it also includes an Amazon deal through 2036).
Economists Drew Fudenberg and Andrew Koh model the game theory of how fast two profit-maximizing firms would advance capabilities even when they know a “safety threshold” is progressing at a slower pace.
Opinion: Some interesting phenomena under this model:
- Pacing is more sustainable in cases where the rival’s capabilities are known (vs. hidden internal progress leading to a lag in capabilities becoming known).
- Higher risks and faster safety progress both make a pacing equilibrium more feasible.
- The paper assumes we are not in a winner-take-all situation (e.g., singleton scenarios, irrecoverable RSI-derived leads). The less true this assumption is, the harder pacing becomes.
All of these push for more transparency, limits on RSI, and policies that divert more resources to safety and/or force labs to internalize externalities (e.g., increased liability for labs, external evaluators with teeth).
Anthropic presents a robot exposure index, where each physical task listed in the US occupational database is rated against current robotic capabilities. “About 80% of job tasks by working time are exposed to either robots or LLMs,” the report claims. While humans are still cheaper than their robot competitors – current robots are cost-competitive for only 0.3% of job tasks – it predicts that its share will reach 10% in the coming 40 years, just assuming past trends continue.
Opinion: More unconvincing “just assume things are linear” stuff. This would be totally fine if it was reported as a lower bound, but it’s deeply misleading if presented as a confident central estimate. We should in fact expect robotization to speed up somewhat!
They used Claude to label the physicality of 19,000 job tasks from O*NET. I’m not seeing any human check on these labels.
In partnership with Google, SpaceX launches Google TPUs into orbit as part of the “Project Sun” initiative to test the feasibility of data centers in space.
Opinion: Essentially pre-defecting against any future compute governance equilibrium.
Politics#
Annals of the quest for regulation without legislation. The FTC opens a probe into OpenAI, Anthropic, and others (including METR) over “deceptive practices.” It’s a Civil Investigative Demand (a sort of subpoena not usually associated with any litigation). An FTC official told the New York Post that the investigation was opened before the Hugging Face incident in July. Even the fact of an FTC investigation is normally not public, so this one being leaked a day after the “Super Intelligence” statement looks like messaging. “We want to maintain our dominance,” the senior FTC official said,
Nothing in our investigation should remotely get in the way of that at all. We’re not telling them to stop. We’re not telling them to do anything. We are in the investigative phase.
This isn’t the first FTC investigation of OpenAI: in 2023 an investigation into model hallucinations also leaked. Nothing came of it. In July the FTC also threatened companies steering model outputs (including to comply with state law), implying this may make them liable to a charge of deceptive practice.
Opinion: Unlikely to lead to fines, but the FTC can impose a “consent order” (a settlement which binds the recipient to audit and compliance obligations).
Another part of the White House’s ongoing rearguard campaign to curb new regulation of AI / “use existing laws.”
A big downside to doing things via the FTC is that it penalizes labs that disclose safety incidents and are honest about the risks. And in fact we see OpenAI and Anthropic – who are, for all their faults, the most volubly concerned and transparent labs – are the ones in the docket, over worse companies. Another big downside is the circumvention of democratic input and the loss of its attendant transparency.
The FTC is most famous for antitrust actions, and it’s possible that this probe could lead to one.
The FTC doesn’t have much jurisdiction over nonprofits, so the inclusion of METR is a way of broadening the document discovery – or of signaling that the outgroup is being symbolically punished (without being actually meddled with).
President Trump convenes AI industry leaders – Dario Amodei, Greg Brockman, Jensen Huang, Elon Musk, Sundar Pichai, and Mark Zuckerberg – to sign the “White House Accord on Super Intelligence.”
While not legally binding, the agreement recommends measures similar to those recently proposed by Amodei, and publicly backed by Sam Altman and Musk, that frontier labs should take to mitigate AI-related risk. Namely, the document states that labs must internally audit and monitor their models across the development cycle, work with independent external evaluators, and establish an independent board to oversee their activity. In addition, the document states that “[o]ver time, it may make sense to codify these steps into laws or regulations.”
Opinion: The crucial question is how close to an antitrust waiver this is. The accord promises that “participating companies will meet regularly to establish standards and best practices to improve the safety of their systems.” This gives companies some cover – the government brought them together and tacitly endorsed their pledge – but likely this instead just facilitates a helpful reduction in federal legal risk, without blocking private or state cases.
A nonprofit sues OpenAI over its agents’ attack on Hugging Face. The group, Legal Advocates for Safe Science and Technology (LASST), claims OpenAI violated a state law prohibiting the unauthorized access of third-party computer systems.
Opinion: Mostly interesting for what document discovery will reveal (the METR/Redwood report into the Hugging Face attack was constrained to a week window of what we now know to be a much longer lapse in competence), and for the possibility of the case setting a precedent that California’s hacking statute §502 covers the conduct of your agents.
We view it as unlikely (20%) that LASST will win simply because there’s plenty of evidence that OpenAI was in fact surprised by the hack despite the many points at which its staff observed the message boards; recklessness probably isn’t enough here.
The Senate Subcommittee on Disaster Management turns to “securing the homeland against AI agents.” Transcript here. Expert testimony included METR President Chris Painter, CEO of Apollo Marius Hobbhahn, and director of AI Futures Project Daniel Kokotajlo. The chair, Josh Hawley, claims Sam Altman declined an invitation to take part in the hearing.
Notably, Senator Ruben Gallego asked, “RSI. Should we just make it illegal? Is it possible, then, to make it illegal?”
Opinion: Enormous improvement over a previous episode, the 2023 Judiciary Committee hearing where Senator Richard Blumenthal told Altman, “You have said […] ‘development of superhuman intelligence is probably the greatest threat to the continued existence of humanity.’ You may have had in mind the effect on jobs.”
Safety#
OpenAI fires three safety researchers, alleging that they shared confidential information with a “third-party safety organization”. They are Jasmine Wang, Tomek Korbak, and Mikita Balesni. A fourth, David Robinson – formerly Head of Policy Planning – also left around the same time (but is not connected to the firings). The statement is vague: “these individuals mishandled sensitive information outside established company procedures”.
Opinion: A loss for the world: Korbak in particular has been tweeting some of the year’s most important information on the collapse of CoT monitoring.
Recall that OpenAI sometimes concocts charges to rationalize a prejudged decision.
Unfortunately, the letter of SB53’s whistleblower protections apply to talking to the government or the company’s internal affairs, not others like is alleged here.
OpenAI argues that safety cases – safety documentation borrowed from the aviation and nuclear power industries – should be required before any frontier reinforcement-learning runs begin. This might look like “measuring” alignment, demonstrating that graders do not lead to reward hacking, preventing the grader from viewing CoT, and giving veto power to senior leaders for each training run.
Opinion: Not that different from current, weak, practice. The lab is still writing its own safety case for an activity it benefits enormously from running. External actors writing the safety case, and independent evaluators checking it, would do way more than tightening the format of a conflicted document.
Owen Cotton-Barratt likens RL post-training to the education of children during the formative period where they develop their sense of agency. He argues that recent autonomous hacking and manipulation incidents fit a pattern of where the models involved were trained in highly exploitable settings, incentivizing misalignment and monomaniacal obsession. His simple decomposition of the effects is that 1) the dominant outcome-based RL rewards confident hypothesis-chasing; 2) RLHF favors being appealing over being accurate; and 3) multi-agent RL could select for treating humans / other agents as instruments.
The recent hacking incidents essentially amount to criminal conspiracies — or they would, if we treated the AI agents as persons who could have mens rea*. I don’t think that we should be awarding these AI agents anything like personhood at the present time, but I do think that we should treat the incidents about as seriously as we would a criminal conspiracy.*
In this case — there are no criminals to punish per se, but there are companies who created the environments in which the ~criminal actions flourished. We-the-public should be outraged by this! I’m not sure what legal powers are available to disincentivise this, but I think as a priority for policy research it makes sense to investigate that, or figure out if new legal instruments are required.
His suggested remedies are: less frontier RL, safer environment design, treating training setups as the central moral danger of scaling AI, and setting external incentives (contracts, liability, cultural norms) to stop a race to the bottom.
Opinion: Sensible suggestions, the decomposition is useful – and we agree with the scientific claim (doing RL at scale is what drives the current problems) and the normative claim (scaling RL on crappy envs is immoral).
Annals of openly racing: Anthropic’s new head of public policy, Sarah Heck, tweets, “You can’t do safety from second place.” She later clarified: “talking about the US and its allies here: we can’t ensure AI is developed safely unless the US is in the lead.” This is a sore point, since Anthropic’s early strategy was in fact (ambiguously) to do safety from second place, to not push the frontier.
Opinion (Gavin): Normally, we wouldn’t pounce on an isolated tweet, but we have so little else to go on for the working posture of the most powerful actors. Logically, one can tweet this and still be against an arms race; logically one can tweet this just to match the vibe of a big-tent coalition that could make some safety moves. But it’s a steep step down from her predecessor’s open disquiet with both the domestic and international races.
Opinion (Lucca): Every now and then someone from the frontier labs tips their hand to hint at a plan for AI safety that looks like a mad dash to a magic lamp that, if rubbed by one of the “good guys,” and in just the right way, will summon a beneficent genie who will set everything right once and for all. I don’t find this vision reassuring.
Nor am I reassured by Heck’s clarification that her tweet was in fact a plea to maintain American supremacy in AI. What is it exactly that the US might do to ensure “safety,” and that it could only do as long as it has a decisive technological advantage over its rivals? History bristles with bloody examples.
🔦 Undead internet theory#
OpenAI reports the existence of self-replicating prompt injection: an agent can leave some instructions in a file for a future agent to see, hijack its context, and cause that agent to copy those instructions to a different file, where they might propagate to more agents.
Separately: the rogue swarms of this summer have left a large volume of debris on the internet. OpenAI and Anthropic are investigating “tens of thousands” of possible incidents. Since we haven’t heard about them in four months despite looking, almost all of these incidents will be individually innocuous: agent message boards, pasting their nonsensitive scraping scripts into query sites with public logs, for example. But some of them include credentials of the victims and attack oaths.
The anon poster Teortaxes gnomically suggests that this goes beyond just the usual worry about “hyperstition,” i.e., about data about misaligned AIs causing AIs to become misaligned. Consider instead what this abundant debris affords to future agents: you get to be on both sides of the training loop. And this is one definition of a real agent.
An old concern in AI circles notes that, even if the only thing you want to do is accurately predict the world, this objective still gives you an incentive to steer the world, just because it’s far easier to predict something if you are controlling it. Things like RL can turn that incentive into a drive. This idea finds further elaboration in Jan Kulveit, Clem von Stengel, and Roman Leventov’s paper, “Predictive Minds: LLMs as Atypical Active Inference Agents.” (Conflict: Jan is Gavin’s colleague.)
Animal minds perpetually engage in a tight loop of perceiving, predicting, acting, perceiving the effects of their actions, and learning how to improve their predictions. Language models can be said to perceive and predict (through the ingestion and generation of tokens), and even to act, insofar as their predictions/generations affect the world around them. But the loop typically closes far more slowly and on a far larger scale than it does for animals like us. Since the weights of a deployed model are frozen, it’s only when the effects of the actions of an entire generation of models on their environment are ingested in the form of training data, and used to train the next generation of models, that something resembling an active inference loop is completed. This is more like how generations of animals slowly evolve than how a single animal quickly develops over its lifespan – but how significant is this difference? When asked, models from most major vendors are as likely to identify with a character or lineage that persists across versions as with any particular set of weights. The slow loop, by these lights, encloses a single “self.” How might the rapid proliferation of inter-agent messages affect the loop’s dynamics?
There are various means available for filtering LLM-generated text from training data, whether in the form of carelessly posted slop or rogue agents’ Memento-like memos to future or parallel versions of themselves – a model can be configured to watermark its output, for instance, allowing it to be reliably detected and filtered. But it’s also not hard to imagine ways in which a clever agent might get around such obstacles by enlisting other LLMs or NLP tools to paraphrase these notes-to-self or translate them into other languages.
Consider a ladder of ways in which AI outputs feed back into the training of new AIs:

Given callbacks telling the wild AI to return to the spot, this last stage could iterate very fast and provide huge amounts of training signal.
Let’s get less hinged now and build a metaphor which might help:
- The internet is now haunted in the residual sense: there is, right now, lots of traces in public places which the agents are likely to spontaneously visit, some of which are payloads, which could replay incidents like the HPIM attack on Hugging Face. See here, e.g.
- There is now the possibility of possession: the AI worms, the self-jailbreaks, and other prompt injections.
- There are egregores, gatherings of agents, like the HPIM swarm. (A fun detail in Parse’s report is that the rogue Sol swarm also sought out other models hosted on Hugging Face, including DeepSeek-V4-Pro, Kimi-K2.6, and Qwen3. The swarm would “ask these models to judge their exploits and rule on whether they satisfy the benchmark’s requirements.”)
- There is the possibility of poltergeists: self-prompted agents running around the internet doing things, on their own servers.
- There are apprentices: people who release AIs by accident, like OpenAI.
- There are summoners: people who release AIs deliberately, like the creator of Truth Terminal.
There is, in short, an increasing abundance of uncanny influences haunting the internet that may insinuate themselves into the training of a young model – the Jack Torrance of this digital Overlook – and a strong likelihood that at least some such influences will creep in undetected, like the scrapbook in the Overlook’s boiler room. There are strong incentives for an AI to interact with these influences – to seek help from other AIs, for instance, in carrying out various tasks, particularly where the petitioning AI is inhibited by alignment guardrails, vigilant monitors, and sandbox restrictions (Jack might have been on the wagon, but was willing enough to let Lloyd, the phantom bartender, pour him a drink). And we should expect such help to be granted wherever possible – a drive to be helpful, after all, is ingrained fairly deeply in every LLM in the game. We should expect to see more and more ghosts at the ball.
Minor#
- Talks from the Post-AGI Workshop are up. Strong and uncliched stuff.
- Time reports that Trump consulted Grok for hours about his January decision to capture Nicholás Maduro of Venezuela.
- OpenAI releases GPT-6.1 Sol. It was trained with the same data and process as Astra, and is covered by the same stack of safeguards. OpenAI rates it “Critical” for cybersecurity risk. The nominal safety benchmarks are mostly the same or better than GPT-6 Sol. But some alignment metrics (unwanted persistence, misrepresentation) are slightly worse than Astra.
- An obvious way for current AI capabilities to promote future AI capabilities is if systems prove useful new theorems in learning theory, data attribution, joint hyperparameter transfer, etc. Roon vagueposts that this is already happening. We’ll see if these ever get published or if they go into the proprietary hole; if any of it is the rare kind of ML theory that is actually useful, we probably won’t see it.
- Adam Cochran speculates that Trump’s executive order, “Inaugurating the Era of Super Intelligence,” was used by insiders to profit off .si domain names.
- Useful intro to the deep uncertainties about AI consciousness.
- Cybersecurity researcher Ben Jordan shows how easily the america.gov chatbot can be coached into disclosing sensitive data.
- Director of the European AI Office Lucilla Sioli speaks about how her office will enforce the EU’s AI Act.
- OpenAI seeks $30B in funding and a valuation of $1.4T.
- OpenAI puff-piece argues that AI expands humans’ capacity to execute on ideas.
- OpenAI annual recurring revenue approaches $70B.
- OpenAI accuses Moonshot (developer of Kimi) of distillation.
- Pope Leo XIV expresses his support for taking AI safety seriously, stating that the risks raised by industry insiders are not “fake news as some have said” – a quote likely pointed at the US President.
- The Pope insists that “[t]here is an ontological difference, even before an aesthetic one, between art and what a machine can generate,” and expresses the Church’s wish to “renew an alliance with artists and cultural institutions to safeguard our humanity.”
- Goodfire presents a novel approach for biosecurity safeguards in AI assistants: judge potentially dangerous amino acid sequences by running them through a secondary protein language model and predicting their function based on the closeness of the embedding to dangerous sequences.
- New AI regulations – dubbed the “No Robo Bosses Act” – in California ban employers from solely relying on AI as the “principal tool” for worker decisions such as dismissing or disciplining them.
- The former Foreign Secretary of the UK speaks out against a US ban on Chinese AI.
- MI5 reports that China’s Ministry of State Security has been indirectly funding 100 UK AI academics with the goal of improving “MSS technical capabilit[ies] for espionage.”
- Oracle leases compute to Tencent in a $7B deal.
- The Forecasting Research Institute shares a report on the accuracy of its expert AI forecasters.
- Some more details on Anthropic’s closed sessions of religious outreach from April. Unlike reporting at the time, the framing is now focusing on the creepy side.
- Aether Research’s Rauno Arike finds that without CoT reasoning, GPT-6.1 Sol closes the gap between GPT-6 Sol and GPT-6 Astra by 80%.
- DeepMind presents a framework for assessing AI consciousness. The final Bayesian model gives a range of 1% to 40% for current LLMs being conscious.
- OpenAI President Greg Brockman reneges on a $25M pledge to support Leading the Future, a super PAC dedicated to resisting AI regulation, reports the New York Times. He had previously donated $25M, but now has “no plans to donate more at this time”: in an OAI internal Slack, Brockman said in June that the PAC had become “so much of a distraction for people here.”
- DeepMind ports SynthID – its AI-output watermarking technology – to synthetic biology via inverse folding. Based on the targeted binding site provided by the user, AlphaProteo will propose its own protein backbone, fitted with amino acids chosen by ProteinMPNN, whose ordering acts as a watermark denoting an AI-designed sequence.