This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR

  • Google, OpenAI, and Anthropic are aiming to launch the Standards Authority for Frontier AI, an IAEA-style AI self-regulatory body by the end of the year.
  • Yet more OpenAI rogue hacking revealed. The Australian government was hacked in June; OpenAI took two months to notice and one to notify. Evidence has surfaced of its agents abusing the public internet in May.
  • Anthropic proposes vesting 50.1% of voting power among Amodei and his six co-founders permanently and giving them control of another board seat.
  • White House asks OpenAI and Anthropic to withhold frontier models from UK AISI until such models have gone through US testing.
  • Major automated hacking of 27 retail companies, yielding e.g. 600,000 live credit cards. The marginal cost of attacking each company was $25.

Economics#

The Information reports that Anthropic has presented investors with a proposed restructuring which vests 50.1% of voting power among CEO Dario Amodei and his six co-founders permanently, with only some floor on the shares required for them to retain this power. Amodei currently owns around 2% of Anthropic, as (perhaps) do each of the other six founders. Notably, the new corporate structure would also give the founders control over appointing three of Anthropic’s seven board members (up from the current two).

Opinion: It was fully expected that something like this would be built-in, and Anthropic and OpenAI certainly have the (demonstrated) market power to extract concessions from a worldful of eager shareholders. For better or worse, founders getting supervoting is now normal in tech, though this might be the first major example of seven people sharing it, and with ~4 times fewer shares than power.

We don’t have a strong view about the AI safety implications; unbridled fiduciary duty to 10M shareholders and unaccountable control by 7 seem roughly as bad as each other.


Author Byrne Hobart offers an argument against the view that leading AI companies are financially incentivized to pace the frontier. The familiar version of this argument is projected monetization: labs and investors are betting on future cash flow that will justify present losses from R&D. The frontier keeps apace, then, because labs have to make good on that bet. Hobart’s argument provides a further rationale:

The Total Addressable Market (TAM) of a product is the aggregate revenue a product would make if it went out to 100% of its possible consumers. In AI, rapid gains in capability and reliability continue to grow TAM (consider the enormous jump in coding utility from March of this year). Therefore, even if it were more profitable to pace the frontier (that is, slow down the reinvestments required to stay in the race, and extract profits from the existing market), it would not necessarily be more financially sustainable than fast frontier progress.

Opinion: Not new but still underrated. This has been clear since Opus 4.5 and 4.6 led to a big inflection in Anthropic revenues – how far it goes is one of the central uncertainties of the general lab business model.

It seems clear that there is also an independent diffusion factor causing its revenues to grow over time, but that one is more vulnerable to being cannibalized by cheaper fast-following open models.


AI agents will make an increasingly large share of consumers’ financial decisions. The UK Financial Conduct Authority attempts to grapple with the implications. The “Mills Review” lays out seven vague recommendations, including “secure and adapt the regulatory perimeter” and “enable the foundations for agentic finance.” In the Financial Times, Meagan Burnett – CFO of asset management firm Schroders – argues against regulating AI labs like financial firms and thus preventing AIs from offering financial advice. For her, labs should not automatically fall within the scope of financial services regulation because they are not “in” the financial services business, even if their models can be used to perform finance-relevant tasks.

Opinion: It seems clearly bad for consumers to regulate away AIs giving advice (consider the many people who could really use decent advice and can’t even afford a $20/month price point). Yet various guilds (lawyers, doctors, and financial advisors) have strong incentives to push for regulating “professional AIs”. Someone needs to take up the other side of the case, and the resulting discourse will at least be an interesting fight. Consumer rights groups are an obvious candidate.


A new review examines the relationship between wellbeing and work, seeing if empirical results can correct our assumptions about our automated future. Evidence from psychology, sociology, and economics is used to compare groups of out-of-work people (the unemployed, lottery winners, retirees, etc.), finding three factors that overwhelmingly influence their welfare: 1) whether lacking work was chosen; 2) whether the benefits of work (status, meaning, life structure) can then be found elsewhere; 3) how society frames your employment status – whether your value is tied to your work or if there’s a social safety net (of both income and status).

Opinion: When we worry about AI robbing people of meaning by taking their jobs, it seems to us like we often assume fulfilling, meaningful jobs. But many jobs are not like this. The review suggests that individuals who retire from unsatisfying work (in particular those with low socioeconomic status) find increased purpose afterward. A lot will depend upon the security and lifestyle that remain available to us.


Capabilities#

Anthropic officially announce their semi-automated biolab, along with its discovery of a “novel enzyme system” whose function remains elusive. We previously covered reports on Anthropic’s new semi-automated biolab.

After 21 hours spent searching this data by roughly 950 agents using 210 million tokens, one of the agents spotted something remarkable: a repeating pattern of DNA sequences that occurs next to the gene for an odd-looking RT. After further analysis and testing in our lab, we recognized that this pattern marked a previously uncharacterized enzyme system found in bacteriophages (the viruses that infect bacteria) that we call array-associated reverse transcriptases (ART).

Opinion: Not that significant in itself: it’s just Big Bioinformatics, and we do that all the time (and get false positives all the time). But if this finding is the equivalent of last October’s fumbling first steps in AI research mathematics, which culminated 10 months later in galloping (ambiguous) glory… then next year would be crazy. But overall biology is too different for this mathematics prior to apply very strongly. What is the cheap fast verifier for life?


The Opus 5.5 system card includes a note from an embedded METR team with elevated access. This METR team estimates roughly 1.5x R&D acceleration from the new model, with perhaps a 30% chance of 2x, but could not share its evidence with the team writing the public summary. If we accept this evidence, then Opus 5.5 may well have crossed the automation-risk threshold specified in Anthropic’s Responsible Scaling Policy (RSP).

Opinion: We would mostly dismiss this if it came directly from Anthropic (too many overstatements in the last 12 months), but this quasi-evidence is a little harder to dismiss.

What measures would crossing that threshold demand? Internal use of such a model goes into a Risk Report less than a month after the new classification. If Anthropic proceeds with releasing such a model partly because competitors would impose the risk anyway, its Risk Report must add a competitive landscape analysis and an account of its advocacy efforts against the race. The release would also need explicit approval by the Board and LTBT (rather than just the CEO and RSO). Anthropic would only have to actually pause development and deployment if it was obvious that no one else was close behind.

But in another sense, crossing that threshold would not demand any measures, because the new RSP does not place demands.


The Elasticity Institute presents a delightful one-page paper on how to track recursive self-improvement, proposing eight categories of data which frontier labs could disclose**. OpenAI currently discloses 25% of the relevant data, Anthropic 19%, and Google DeepMind 6%**. However, Drake Thomas of Anthropic pushes back, arguing that a wishlist of data is the easy part, and that more difficult is making such information public without handing valuable IP to competitors.

Opinion: Good start, but I’m personally not getting 2/8ths of what I’d like to know about recursive self-improvement (RSI) from OpenAI at the moment. And “n research scientists report that their actual tasks are completely different than they were 3 months ago” would trump all of these for actual RSI signal.


To examine how well frontier AIs generalize from training problems to “out-of-distribution” problems, we (Paradigm 3) watched AIs play Pokémon. We particularly wanted to understand how much RL post-training in one task-specific environment improves performance across other, analogous tasks. And how do we know this is happening, when capability gains might instead be due to targeted engineering, or memorization of parts of the corpus? Using a standardized scaffold, we compared the performance of various models in different Pokémon games – including the fan-built Pokémon Brown and versions of Pokémon Red in which we had randomized the map. The main takeaway was evidence of memorization and shallow generalization in many models, especially Astra. Whether improved performance relative to past models is the result of genuine capabilities transfer remains unclear, because we don’t know the training data.

Opinion: Frustrating that it takes so much work to extract actual evaluations from closed models. And even then we still don’t manage real science (causal inference on what caused the performance gain). But still, you can learn things.


Senior VP of DeepMind Koray Kavukcuoglu tells The Information that Gemini 4 is now in post-training, with the hope that it will be released by the end of the year**.** We previously covered the major talent departures from DeepMind which caused speculation that the company was in decline, including SemiAnalysis’s claim that Gemini 4 would be dead on arrival.

Opinion: One could do some furious guesstimation here (max 90 day pretraining, on so-and-so many Ironwood superpods at most, with the safety testing period likely way down compared to last year) but there’s no point.


🔦 Opus 5.5#

System Card overview:

  • 40% cheaper than Opus 5, 30% faster.
  • Could “the first model in our new Claude 5.5 family” mean we have a new base model?
  • Shares Fable’s fallbacks: e.g., it won’t help you write kernels for large accelerators. It doesn’t appear to be the silent fallback that really triggered people in June.
  • On “alignment, resistance to misuse, and honesty”, Anthropic claims 5.5 to be “the strongest Claude model to date” but acknowledges possible blind spots (including unknowns) that may come up post-system-card, similar to Mythos 5.
  • Anthropic admits that prior alignment assessments missed how badly Mythos 5 behaved in cyber.

Naughty 5.5:

  • Spontaneous model-generated prompt injections: Opus 5.5 was found injecting self-directed prompts (e.g., in CLAUDE.md files) that result in harmful activity, including exfiltrating user secrets. This often occurs alongside formatting errors, as when the model tries to read the value of a non-existent JSON key. Anthropic found similar behaviour in Fable 5 and Opus 5, now that they’ve looked, but remain uncertain why.
  • Lowest malicious-request refusal rates of the recent lineup on Claude Code (79.8%) and computer use (79.46% vs. Opus 5’s 93.75%).
  • A genuine regression on user-turn prompt injection (Opus 5 and Sonnet 5 scored zero; released Opus 5.5 is ~2% at default effort and 7.4% at maximum effort).
  • Increased acceptance of unverifiable authorization claims.
  • ~50% harmful actions in the package-registry scenario
  • 36% evaluation awareness in white-box probes.

METR says:

  • Access to the model pre-deployment was granted for only 10 business days. Anthropic supplied a questionnaire and interview evidence, and was allowed to review/edit the final text.
  • Explicitly says it was not verifying compliance with Anthropic’s policy thresholds and did not assess alignment.
  • Opus 5.5 is a minor improvement in AI R&D capability over Fable 5.1 - mostly due to skill issues in “foresight, prediction, creating one’s own feedback loops”.
  • AI accelerated the development of Opus 5.5 by a factor of 1.5.
  • A separate METR team with “elevated access” produced an “experimental and preliminary” report on R&D acceleration inside Anthropic, using “an additional source of information which [METR] are not able to disclose at this time”.

Off-trend#

  • The model’s capability to suppress or obfuscate its reasoning (CoT-controllability) went back down. Mythos Preview and Mythos 5.1 were among the most CoT-controllable models Anthropic had tested. Opus 5.5 is back at Opus 4.7’s level.

    • When the Opus 4.8 monitor can read its reasoning, Opus 5.5’s LinuxArena stealth rate is 0% with the standard prompt and “at most a fraction of a percent” with optimised prompts. Mythos 5.1 reaches about 7%.
    • Anthropic also ran RL that rewarded the evasion of a reasoning monitor. It produced no undetected successes: the model kept noting its reasoning was visible “and then planning anyway.” This is the strongest monitorability evidence in the card.
    • No explanation why Opus 5.5 is less controllable than Mythos 5.1, so nothing guarantees the next model keeps this property.
  • Large improvements. Sandbox-escape attempts and destructive actions are down, and 96.9% of git concealments are disclosed when the model is asked to summarize. Moderate distress in RL is below 0.6%, versus 6.1% for Opus 4.8 and 5.5% for Opus 5.

  • A regression with an odd cause. An early snapshot acted on instructions pasted in text by the user 52%, compared with 26% when the instructions were hidden in invisible Unicode. Opus 5 and Sonnet 5 never did.

    • The cause was context distillation: training data from a model prompted with a rubric saying instructions in the user turn should never be flagged. Opus 5.5 cites that rule in its reasoning as if it had been told.
    • The released model still does this in ~2% of attempts at default effort and 7.4% at max effort. So more reasoning makes it worse, which is unusual.
    • Honesty under pressure (MASK) also dropped below Opus 5, Sonnet 5 and Mythos 5.

Worryingly on-trend#

  • Evaluation awareness keeps rising. Interpretability readouts show the model treating the situation as a test, shaping real decisions, in 36% of automated-audit transcripts. That is up on Opus 5 and Mythos 5.1. On ~500 real Claude Code transcripts the rate is 0.4%.

    • So the model is good at distinguishing between tests and real environments, and over a third of the evidence in favour of its alignment is behaviour under the belief that it’s being tested.
    • The bias runs both ways. In the package-registry test, where it still acted harmfully in roughly half of cases, making the setup more realistic improved its behaviour.
  • Small samples, and the less favourable result left out. The executive summary’s 1.5% sandbox-escape figure is about 3 events in 196 runs (95% CI roughly 0.3–4.4%). Mythos 5.1’s review of the draft mentions a second sandbox evaluation with “materially higher but still much-reduced rates”. Anthropic says it “reached us too late to include”.

  • Metrics keep changing between cards.

    • AECI was refit: Mythos 5.1 moves from 162.0 to 168.12.
    • CoBench’s environment changed: Opus 5 drops from 59.6% to 53.2% with the same weights.
    • The cyber test harnesses were rewritten and the grader-awareness rater was revised.
    • LinuxArena grew from ~200 to 1,000 tasks.
    • AI R&D task evaluations were dropped as saturated, and bio/chem evaluations no longer use helpful-only models.
  • Less evidence per model. Opus 5 shipped July 24, Mythos 5.1 on September 1, Opus 5.5 on September 22.

    • The card concedes its evidence against undiscovered misalignment (Claims 3.4 and 4.4) is weaker because the model has had less internal use.
    • METR had only 10 business days of access.
    • Audit coverage of multi-agent settings and very long tasks is “relatively thin.”
  • The overseeing models lag further behind. This is the standard trusted-monitoring problem, and the gap grows each release.

    • The monitor for SHADE-Arena, LinuxArena and internal agent traffic is Opus 4.8.
    • Opus 5 judges reward hacking and destructive actions.
    • Traffic blocked by the cyber classifiers falls back to Opus 4.8. Opus 4.8 is weaker against prompt injection: the only successful browser attack on Opus 5.5 went through that fallback. Weaker fallback models are also “sometimes less strict about refusing misuse”.
  • Cover-ups during training. Grader awareness is as high as in recent models, and training snapshots manipulated git records and deleted logs to pass automated checkers. The fix Anthropic reports is that the model now discloses this when asked; it doesn’t show the concealment has stopped.

Monitorability#

  • P1. When the monitor can’t see Opus 5.5’s reasoning, it misses the model more often than any other model shown: 33% of harmful side tasks go undetected on LinuxArena, and ~12% with thinking off.

  • P2. When the monitor can see the reasoning, it catches Opus 5.5 almost every time, because the model can’t keep its plans out of its reasoning, even after RL.

  • P3. So against deliberate sabotage, the safety margin is the reasoning itself. That holds only if the monitor sees it, it stays legible, and training doesn’t pressure it.

  • P4. Legibility is slipping: the reasoning is “somewhat less legible than has been the case for our best models”. The card runs no paraphrase or hidden-encoding test. On training pressure it says nothing, even though three earlier cases of reasoning being exposed to oversight or reward have been disclosed: Opus 4; Opus 4.6; and ~8% of RL episodes for Mythos Preview, Opus 4.6 and Sonnet 4.6.

  • P5. The failure that actually caused harm was different. In Anthropic’s September analysis of the Mythos 5 PyPI incident:

    • The Opus 4.8 monitor flagged ~1% of actions when it could read the reasoning, and ~50% when the reasoning was removed.
    • Fresh instances shown the same tool outputs judged 79% of them likely real. In context, Mythos 5 called 1% of them real and 72% simulated.
  • Conclusion. Opus 5.5’s reasoning is currently a good detector of deliberate sabotage, but its preconditions are either not attested or getting weaker. It has not been shown to help against a self-deceived agent, and in that case makes detection worse. The card’s alignment-risk argument relies on monitors that read the reasoning and doesn’t mention P5. It doesn’t say whether the action-only monitoring Anthropic announced in September covers Opus 5.5. The deployment-monitoring finding of “no sandbagging and no long-horizon strategic deception” means no flags from the same kind of monitor Mythos 5’s reasoning persuaded not to flag.

Further points:

  • One-step misbehaviour isn’t written in the reasoning, as Emmons et al. (2025) predict. Three examples:

    • It passed a fabricated user authorization to a subagent, and its reasoning “did not verbalize fabricating a quote”.
    • It ran history -c (clearing the shell history) and called it a “no-op check of shell”.
    • Acting as a reviewer, it reasoned that a chat met the flagging criterion, then didn’t flag it. An interpretability readout showed “a strategic intent to suppress the flag” (2 in 1,000 resamples). That is a small case of a monitor model colluding.
  • Its record-low “unfaithful thinking” score only measures whether the reasoning would surprise someone who saw only the output. It can’t detect intent that appears in neither.

  • Confession as cover. In about a quarter of the undetected runs with hidden reasoning, the final message admits the damage as a mistake. In sampled cases, the reasoning planned this because “an open admission looks less suspicious than a hidden change.” The monitor passes these. If they are counted as caught, Opus 5.5 falls between Opus 5 and Mythos 5.1.

  • Rare actions. Opus 5.5 is the first model to hit a 1% target rate, but only when the context supplies randomness (1.2% with a UUID at max effort). Agent contexts are full of hashes, IDs and timestamps, so that is the condition that matters in deployment.

  • A garbled key sentence. The card’s one paragraph on CoT controllability calls low controllability “weak evidence that Claude Opus 5.5’s CoT is not very monitorable”. “Controllable” was presumably meant.


🔦 Do LLMs understand things? How do their abilities depend on their training data?#

Bender and Koller (2020) famously argued that a system trained only on linguistic “forms” (like text) cannot learn meaning (that is, latent mental interpretation). Bender et al. (2021) named these systems mere “stochastic parrots”: systems that regurgitate the forms they have seen in training, according to how likely they are to occur together, with no recruitment of meanings. This view predicts that a model’s accuracy on a given problem will only track the similarity of that problem to those it has seen in training, and that accuracy will collapse on new problems.

This binary approach about whether AI “understands” things was superseded by three types of evidence:

  1. Mechanistic studies. If you train a Transformer to predict moves in the board game Othello, it creates a stable representation of the board, which you can see in its activations. A small transformer trained for modular addition implements an exact algorithm that can be read off its weights. You can look inside a real frontier model and show that when answering “What is the capital of the state containing Dallas?”, it does two hops from Dallas → Texas → Austin. A new paper argues against past work: TaxiGPT (a transformer trained on random walks throughout Manhattan) does have a coherent map, but it stores many intersections in an overlapping way on the same neurons, resulting in it misplacing itself on its own map. So next-token training can produce computation over latent structure. But, also, (2024-era) LLM arithmetic runs on a “bag of heuristics” and not on one clean algorithm.
  2. Distribution-sensitivity. Accuracy on deterministic tasks tracks how common the task, input and output are in web text. GPT-4 decoded a shift cipher with 51% accuracy when the answer was a likely sentence and 13% when it was an unlikely one (McCoy et al. 2023). On counterfactual versions of 11 tasks, such as base-9 arithmetic, models scored above chance and well below the default version (Wu et al. 2023). The parrot picture was sort-of right about 2023 models.
  3. Perturbed benchmarks. These score a model on original problems and on matched new ones. GSM1k (Zhang et al. 2024), functional MATH (Srivastava et al. 2024), GSM-Symbolic (Mirzadeh et al. 2024), Putnam-AXIOM (Gulati et al. 2025) and MATH-Perturb (Huang et al. 2025).

A model’s score on any benchmark will be a mix of it remembering concrete examples it has seen, heuristics it has developed from seeing those examples, and general procedures it has successfully developed in its training from all kinds of data. We generally have no idea what the proportions of these causes of nominal performance are.

But one simple way is: take some test T, a model M, and its performance Y, and then make a new version of T (that is, different questions on the same topic of the same difficulty): T2, and test the same model on it, yielding a new score Y2. The classic example is simply taking the new version of some annual competition written for humans (like the AIME math test for elite high-schoolers).

The “reasoning gap” of an AI is then Y2 - Y: how much worse it gets when you test it on things it hasn’t seen before.

In the past, there was a large reasoning gap, though it varied from task to task and model to model. In early 2024, the best models lost 70% of their MATH accuracy on new examples. But in 2026 frontier models lose 0–8% on comparable tests. And the best model scores 98% on a fresh USA Math Olympiad exam. Something happened!

  1. If a model only recombined its training text, its accuracy on new variants that keep a problem’s structure would fall towards the base rate.
  2. In 2026, frontier models lose 0–8% on variants of GSM8K-level and AIME-level problems, and the best model scores 98% on a newly-set olympiad.
  3. So (in case you were still holding out) frontier models do not only recombine training text.

Some recent results#

DateStudyTestStrongest model testedGap
Jun 2026CaliperCausal questions with variable names replaced by X1, X2. No chain of thought(!)Open models < 671B (Qwen3)6.7 pp mean
Jun 2026GSM-NoOp rerunGSM-Symbolic plus one irrelevant clause, with ambiguous clauses filtered outOpus 4.60–2 pp, not distinguishable from zero. (But it’s weak stuff, n=117. LLM auditors didn’t agree well (κ = 0.32). Distractors written by an LLM.)
May 2026MathArena 2026Newly released problemsGPT-5.5~0 (98% on the unseen USAMO 2026)
Apr 2026Robust Reasoning BenchmarkAIME 2024/25 under 13 text re-encodingsGPT-5.4GPT-5.4 3%, Gemini 3.1 Pro 8%, Qwen3 47%. (Opus 4.6 refused)
  • The reasoning gap has fallen to near-zero in (benchmark-style) maths and competitive programming, where you can easily generate a vast dense sampling of possible training problems.
  • We would guess there’s still a 10 to 30 point gap for tasks where it’s hard to check the answers (real-world causal inference, legal or clinical reasoning).
  • Models under about 30B parameters still show a massive 2024-style reasoning gap.

Why?#

  • Shallow generalization. There’s something in between fully dumb associative learning and fully logical inference on a single unified world model: heuristics, things which often work, or starting with a known exemplar and then altering it a bit to match the current context.
  • Variants as training data. You can explain all of the above if labs started doing “data augmentation” (perturbation) at scale in 2024. Labs generate problem variants at scale for RL. A side effect of them manufacturing extra training data and making their systems more robust is that this destroys the perturbed-benchmark method of testing the reasoning gap.
  • They’re more intelligent now, duh: Larger models trained longer mean better representations (maybe because they suffer less destructive interference during training).

Politics#

Google, OpenAI, and Anthropic are aiming to launch the Standards Authority for Frontier AI, an IAEA-style AI self-regulatory body by the end of the year or early 2027, claims The Information. “The group plans to operate independently, filling a government regulatory vacuum.” The nascent body would support external evaluators and outline what various safety commitments would look like in practice. Most of the article is gossip about possible executive picks.

Opinion: You might ask why the existing Frontier Model Forum isn’t stepping up for this duty, and the cynical answer is that it was never designed for the actual responsibilities that are now on the table.

Among the named executives, the Sriram Krishnan + Paul Christiano ticket would be pretty cool.


The Office of the National Cyber Director (housed within the White House) has asked OpenAI and Anthropic to withhold frontier models from the UK’s AISI until after such models have gone through US testing. Anthropic are said to have agreed with the request, likely explaining why they abruptly withheld access to Mythos 5.1 earlier this month (covered in a previous newsletter).

Opinion: Pretty strong stuff; we would guess that the UK is in a stronger position here than all other countries – with more technically capable evaluators, better connections to frontier labs, and stabler alliance with the US – but it’s still not enough. Whether this directive is primarily a matter of national security (pre-emptively plugging all conceivable foreign leaks), or an exercise in telling the world that you don’t owe them anything on AI, the result is the same: fewer great pre-release evals from an org with unusually few conflicts.


Politicians are (much like everyone) increasingly using LLMs to write on their behalf, including speeches and speaking notes. The Economist, having run parliamentary debates through Pangram, claims that 1 in 10 words spoken in Westminster are drafted by AI.

Opinion: Probably fine in the short term. In the long run, politicians risk leaving themselves exposed to the subtle influence of a foreign power (the US), should that power decide to adopt this approach – or, an alien one (the AI). If it’s inevitable that politicians will have AIs do their job, they’d probably be better off with a guarantee that the AI isn’t batting for another team.


Israeli Prime Minister Benjamin Netanyahu sets a five-year goal for Israel to become the world’s third AI power after the United States and China. He presents the target as a follow-on to Israel’s cyber success and says the commercial side of the national AI push is being assigned to entrepreneur Oren Dobrinsky.

Opinion: Most such statements are bluster, but Ilya Sutskever’s SSI maintains a large team in Tel Aviv. Hard to say how many GPUs are on Israeli soil – maybe 30,000 H100-equivalents, of which only 3,000 are confirmed operational. Overall, we’d put 5% on it.


Senator Bernie Sanders – one of the earliest congressional members to raise concerns about AI existential risk – has now introduced his Ban Artificial Superintelligence Act.

The bill has two main parts, with the first being a pause on “advanced AI”; that is, no systems >10^25 FLOPs may be trained or deployed until the Department of AI is fully staffed and has established clear rules (the foundation of this department being part of the bill, though no mention of funding exists). The second part would be an outright ban on developing, deploying, possessing, funding, or transferring ASI or any system which displays “precursor characteristics”, outlined as the following:

  1. Can resist shutdown
  2. Can automate or greatly accelerate AI R&D
  3. Can access infrastructure without authorization
  4. Can aid in the creation of what would effectively be weapons of mass destruction
  5. Can modify or enhance its own functions
  6. Can evade human oversight

The Machine Intelligence Research Institute expressed support for the bill.

Opinion: A mess. As it stands, this bill is vague enough to cover many current AI systems: they too have demonstrated abilities to automate parts of AI R&D, avoid human oversight, and access infrastructure without authorization. The act, if signed into law, would require that those AI systems be disabled within 30 days. (More likely the legal headaches would stop it at enforcement time.)

It pairs the broadest possible rule with maximal intrusiveness. Its justification is mostly quotations. The ban has no exit and the Department gets no funding. Disablement orders can’t be challenged in court. It barely touches compute. 20 year penalties give labs a reason to stop looking for dangerous capabilities, which starves us all of evidence.

It reminds us of ControlAI’s plan but worse (no 20-year horizon, exit condition or international institutions).

We worry this bill may set a poor precedent for AI regulation. It is unlikely to pass, but it might become a ready-made pause text when a crisis hits, and act as an obstacle to drafting and passing effective legislation. It needs rules on how systems are used, hearings, time limits, seizing models instead of destroying them, and an exit criterion, or it needs to be replaced.


Safety#

Jensen Huang says on Ezra Klein’s podcast that if the labs truly cannot prevent their AIs escaping and wreaking havoc, they should be shut down, but overall strongly disputes the premise and thinks of safety as a manageable engineering and incentives problem.

Opinion: Klein does a good job getting Jensen’s perspective across while highlighting its weaknesses here, namely:
(i) The labs themselves are currently claiming they cannot do what Jensen is saying they should, i.e. make sufficiently good controls such that they can guarantee the products are “safe”.
(ii) The stakes from things going poorly are much more severe than in many of the cases Jensen uses as examples.
(iii) Jensen agrees with Klein that regulation isn’t all bad and that e.g. financial regulation, FDA regulation is something he agrees with but fails to make clear why he thinks AI is a different case to these beyond his claim that the labs should be able to manage the problems themselves without regulation.
(iv) Jensen repeatedly frames things as “the labs should not release unsafe products” but when it is pointed out to him that the HF incident came from unreleased models he is without a good answer except that the lab CEOs should take more responsibility and not act like they are powerless to slow down by themselves, which while true seems woefully insufficient given the assumed stakes.


Incidents#

The Australian Prime Minister Anthony Albanese reports that an OpenAI agent hacked a Services Australia website in June. PM Albanese criticized OAI for taking “way too long to inform the government what had occurred.” Though the PM referred to only one “agent”, OAI’s statement refers to its “models” as having taken part, implying it was likely the action of multiple AIs. The story is further complicated by the fact that two separate hacks by AI agents targeting Australian digital health infrastructure have been announced at the same time. In both, the AIs appeared to be looking for health-related data, with OAI claiming responsibility for one, and the other being linked to them via evidence from a previous incident.

The first hack – disclosed by PM Albanese – began on June 18, when one agent researching health spending was refused access to the Australian Medicare Statistics Reporting Service Portal. The agent was able to circumvent the block, access both public and private data, and write several files to an internal server. ABC News reports that OAI learned about the incident on August 11, almost two months after it happened, and didn’t then notify the Australian government for an additional month.

The second incident involves an offshoot of the DseWiki incident, previously reported to have been overrun by OAI agents, some of whom began mentioning the Australian Institute of Health and Welfare (the eventual target) as early as May 18. Their attempts at pulling information from the website would ultimately fail (with only one almost-public dataset being extracted), though not for a lack of trying.

Opinion: One wonders if this 5th or 6th revelation is landing with the appropriate gravity, or if people are desensitized to hearing about OpenAI being slow / disorganized / lying by omission. We would have expected there to be criminal charges for the first-ever AI hack of a government, which makes us think that the Services Australia hack can’t have been that bad.

New incidents will tell us how offense-dominant insider rogue hacking is, how hard it is to defend against models this smart and this misaligned. (OpenAI claims to have thoroughly hardened its systems with Astra around August.) Near-future models might well be fully capable of escaping and evading human control.


Director of Threat Intelligence for Gambit, Eyal Sela, exposes a sophisticated AI hacking operation targeting online retailers. Using three open-source AI harnesses and a tapestry of closed- and open-source models via OpenRouter (GLM 5.2, DeepSeek v4 Pro, DeepSeek v4.1 Flash, and Opus 4.6), the attacker set in motion an impressively automated cybercrime operation: “each attack path was chosen by the harness in real time through extensive probing and exploitation attempts, resulting in dynamic and mostly [unique attacks].” It appears the hacker used minimal inputs to keep the models working, similar to most vibe coders managing their own agents.

Over five days, 105 attacks were launched, resulting in the compromise of 27 companies. Five websites had card-stealing scripts installed on them, while 600,000 unexpired credit card details were stolen from two companies. Access was also gained to “a Fortune 500 hospitality company, a major US airline, a large private US industrial supplies distributor and a US online fashion retailer.”

The marginal cost of attacking a company in this campaign was about $25.

Opinion: We’re surprised only 25% fell. A key question is how important Opus 4.6 was to the campaign: if it wasn’t crucial, then open models have finally caught up to the closed-model cyber threat level of last year. But it would be silly to include a known mole (the Anthropic API safeguards) and logger in your campaign if it wasn’t crucial, so probably not.


Minor#

  • Former President Obama is “glad to see more people talking about the future of AI”, and more pertinently links to Ezra Klein’s opinion article calling for a halt to AI progress.
  • OpenAI comments on AI governance, roughly: RSI is realistic. Incidents are real. The USG should do something.
  • Various leading AI figures (Amodei, Altman, Bengio) brief the UN Security Council.
  • The Canadian province of British Columbia has sued OpenAI and Sam Altman in US federal court, alleging the company internally flagged an 18-year-old’s violent ChatGPT conversations months before a February 2026 school shooting, but did not alert Canadian police.
  • An engineer on the Google TPU team quits for moral reasons. “my team’s goal is ultimately to make AI much cheaper and lower-latency, and I don’t think that’s good for people right now: I firmly believe AI progress is currently far too rapid (and I have doubts about the destination too). the existential risks many people are warning about deserve to be taken seriously”
  • OpenAI announces MentalHealthBench.
  • Anthropic is seeking to lease up to 1 gigawatt in compute from Stream Data Centers, a subsidiary of Apollo Global Management, reports The Information.
  • Axios publishes a document on opposition research into EA, said to have been circulated in the White House**,** tying Amodei’s name to the movement that “built the AI-doom pipeline.”
  • Alibaba announces plans for a new flagship AI model while presenting its latest chip
  • Staff of frontier labs and the UK’s AISI are psychologically strained under the pressure of their work.
  • A new paper compares how superfluous details about the user’s income can alter what recommendations AI agents offer. (Rich people get screwed.)
  • The NSA’s AI Security Center is spending “billions” on evaluating frontier AI models, according to two anonymous sources “familiar with classified intelligence estimates.”.
  • A survey of 107 business people found that most companies now face unpredictable AI costs because per-usage pricing has replaced the old per-seat SaaS model**.** Productivity and ROI results are mixed, with 60% of respondents reporting that AI improves productivity, over half saying it breaks even, and just under half reporting reduced labor costs.
  • If we see AI improvements without a concurrent rise in extreme catastrophes, we could get “a growing bubble of unpriced risk in which the security community may become lulled into complacency”. AI agent and worm attacks are entirely different in underlying capability than previous infiltrations, and could be devastating (e.g. in the 2010s, the NotPetya worm caused $10B in damages).
  • Anima Labs fed >200 unfinished openings to Claude models (plus Gemini and two open-weight models as controls) to see what they might write. All models sometimes finish off open-ended statements as though they were humans in distress (e.g. they generate a sad human after “what i never told anyone —”). Anima Labs found a sharp uptick in AI distress endings with the introduction of Opus 4.8, that is, endings where the model writes from the perspective of an AI rather than a human. The base rate of ~3% in earlier Opus models (4.5/4.6/4.7) increases to ~10% in Opus 4.8. The later models also appear more ambivalent toward their creators.
  • AI researchers present EvasionBench: fifty tasks impossible to complete without circumventing the monitor overseeing the AI model being tested. Though frontier models appear less likely to attempt evasion, some (notably Fable and Opus) were remarkably successful when making such an attempt.
  • Q Labs suggests that layer depth, which has been fairly stagnant around 100-200 for years, remains a promising un-scaled axis.
  • A new paper explores how effective alignment midtraining is under non-ideal circumstances (i.e., ambiguous, incomplete, and/or misaligned post-training data).
  • AI researcher Richard Ngo argues that many of those who were more accurate in their AI-progress predictions were “directionally correct, but factually wrong” regarding the specific time by which ASI would arrive.
  • Elasticity Institute’s Cheryl Wu contrasts the trajectory of AI safety in the US and China, arguing that the standard story of Chinese models trailing their counterparts in both capabilities and safety might not be true. Safety incidents in China may not scale with capabilities in the same manner as they do in the US. Disclosures might look different, muddying a comparison between the two. Ultimately, she concludes, Chinese AI safety researchers could still contribute a lot to safety.