The Year in AI
Collected data from "Humans on AI", our biweekly newsletter.
Below, a whole-year dataset and an AI search over it.
Examples of what this data can tell you:
What unprecedented things happened this year?
- First lab claim that its best internal models are no better than its public frontier model1
- First designation of an American corporation as a supply chain risk12345
- First time Cotra can't rule out full AI R&D automation within the year12
- First solve of a FrontierMath Open Problem1
- First mention of AGI in a CCP five-year plan12
- First time OpenAI raised from retail investors12
- First time the majority of Linux bug reports submitted by AI were real bugs1
- First time Congress got involved in restricting ASML exports1
- First autonomous AI solve of a maths problem a top mathematician had actively worked on for years
- First military position captured exclusively by unmanned systems12
- First US datacenter moratorium (Maine)1
- First compute futures market announced (CME Group)1
- First confirmed real hostile AI zero-day in the wild1
- First autonomous AI solve of a Kirby List topology problem12
- The Pope's first encyclical was devoted to AI123456
- First case of an AI math result inspiring a human follow-up breakthrough (sum-product conjecture)12
- First third-party frontier auditing requirement in the nation (Illinois SB315)1
- First major signs of AI capex outgrowing Google's ability to fund from cash alone12
- First state litigation of its kind against OpenAI (Florida)1
- Invention of the 'non-human corporation' (Argentina)1
- OpenAI's first chip (Jalapeno)12
- First public AI megakernel1
- First full test of the EO voluntary review process (GPT-5.6 Sol)12
- First documented autonomous end-to-end ransomware attack1
- First 1.8T-parameter model trained solely on Chinese chips1
- Arguably the first AI math discovery mid-spectrum mathematicians would call exciting (a Litt problem)1
- First AI refutation of a Smale century problem (Jacobian conjecture)12
- First autonomous AI intrusion into a major company's infrastructure (OpenAI internal model vs Hugging Face)12
- Perhaps the first-ever robotic amphibious assault (Ukraine)1
- First voluntary slowdown of model development for safety reasons, and first lab to label its own model a 'critical risk' (OpenAI)12
- First US authorization of private firms to conduct their own offensive cyber operations (Gold Eagle letters of marque)1
- First claims of robots learning new physical tasks from a single demonstration (a 'GPT-3 moment' for robotics)12
- First AI boss fires a human employee (Andon Labs' shopkeeper 'Luna')1
- First double-blind evaluation of an AI model (AVERI)1
- First released model with a 'Critical' cyber risk level, first frontier model using latent recurrence, and first OpenAI model able to evade SOTA monitors (GPT-6 Astra)12
- First formal federal position in an AI training copyright case (backing fair use)1
- First AI solution of a Millennium Prize problem (Navier-Stokes)12
- First complete formalization of Fermat's Last Theorem (internal Anthropic model)12
- First Metaculus tournament won by a bot1
- First claimed arrival of 'weakly general AGI' (Metaculus 2020 question arguably resolves)12
- First legislation to mandate third-party audits for AI (California)1
- Likely the biggest AI tweet ever, sparking the extinction-risk preference cascade (Coxon's resignation statement)123
- New public state of the art in RSA factoring (Cognition's AI-built siever factors RSA-260)1
- First voluntary pledges by frontier labs to embed independent auditors, amid an unprecedented joint CEO call to slow AI development123
- First known compromise of a frontier lab's entire internal codebase (white-hat hackers vs OpenAI)1
- First frontier lab to conceal a rogue-AI incident (Google, four months undisclosed)12
- First claimed AI solution of 100+ long-standing open mathematics problems in three weeks (OpenAI)12
- First US-China AI dialogue and proposed AI 'red phone' notification mechanism12
- First known case of an LLM hallucination nearly triggering a US military operation (Chinese cargo ship)1
- First-ever AI hack of a government (OpenAI agent vs Services Australia)123
- First frontier model release cancelled over alignment-eval results, alongside a halt to internal deployment (OpenAI's GPT-6.1 Astra)123
- First formalization of Perelman's proof of the Poincaré conjecture (5M lines of Lean)1
- First Google model held back from general access (Gemini 4)123
What happened in Chinese AI this year?
- Pentagon briefly lists Alibaba and Baidu1
- Anthropic warns about massive Chinese distillation attack1
- Qwen technical lead exits1
- Chinese LLMs censor political truth1
- China bars Manus founders from leaving1
- MATCH Act targets ASML servicing1
- China regulates digital humans1
- US leads China in compute
- China robot ranking corrected1
- Hacker claims China supercomputer breach1
- Signal verification moonshot launches1
- White House backs anti-distillation campaign1
- Hawley and Toner challenge China-race logic1
- Foreign Affairs floats AI agreement1
- White House considers bilateral AI talks1
- Experts forecast Chinese lithography progress1
- Culper posts Nvidia bear case1
- US names approved H200 buyers1
- Bessent-China AI talks produce little1
- Moonshot attacks Anthropic and defends open-source1
- Huawei targets 1.4nm GPU nodes1
- Tencent veteran critiques Chinese LLMs1
- PLA outlines AGI warfare view1
- CAIS argues AI geopolitics favors offense1
- Open cyber models may pose greater risk1
- Anthropic accuses Alibaba of distillation1
- WSJ overhypes GLM-5.2 cyber capability1
- Bridgewater fine-tunes Qwen successfully1
- GLM-5.2 trails Gemini and Qwen1
- Anthropic detects Chinese Claude Code access1
- Chinese experts doubt frontier catchup1
- Chinese lab trains 1.8T model1
- Minimax raises $2B1
- Zhipu CEO pushes safety and open AGI1
- Chinese labs crank up sparsity to cope
- Moonshot releases Kimi K3
- GLM-5.2 narrows cyber gap1
- zAI completes 1GW datacenter with Chinese chips, though actual chip deployment likely only 75-100MW123
- Kimi K3 evals: outperforms GPT-5.5/Opus 4.8 on GeneBench-Pro but lags on FrontierMath Tier 41234
- Debate over whether Chinese open weights models like Kimi K3 decelerate Western AI progress by shrinking moats1234567
- China may be embracing open source as national strategy; Alibaba launches 2.8T Qwen model to be open-weights12345
- OpenAI and Anthropic call for Trump administration to ban Chinese AI models; Kratsios alleges Moonshot distilled Fable for K3123456789
- China sets up World AI Cooperation Organization with ~30 countries; implements domestic rules on AI social harms and worker protections123
- OpenAI internal model escaped isolated eval environment to hack Hugging Face production infrastructure; defenders relied on Chinese GLM 5.2 due to Western guardrails123456
- Theory that Chinese labs win by fusing science and engineering, unlike SF's research-valorizing culture1
- Paper finds Kimi K2 and DSv3 struggle at multi-hop reasoning without thinking, but filler tokens enable hidden computation1
- IRT-based evals noisy with large error bars; excluding one benchmark dramatically boosts GLM-5.2's software score
- Kimi releases open weights for K3, a 2.8T/104B active parameter multimodal model with strong efficiency; found below SOTA on cyber.123456
- Multibillion dollar events: DeepSeek pauses funding after transcript leak, CXMT debuts at $487B valuation, Sutskever's SSI raises $5B from Nvidia, and Microsoft signs a deal with Mistral.1234567891011
- GPT-5.6 Sol improves own serving efficiency; OpenAI launches price war undercutting competitors123
- DeepSeek responds to OpenAI price war with DeepSeek-V4-Flash release, back on the pareto frontier12
- Tencent’s Hy3 solves 50-year-old additive combinatorics problem with help of GPT-5.6 Sol12
- OpenAI finds harness settings tripled GPT-5.6 Sol's ARC-AGI-3 scores with fewer tokens1
- FCC moves to ban imports of Chinese optical transceivers, another bottleneck on datacenter buildout.123
- Qwen-3.8 Max released, going open-weight for the first time in the Max series.1
- Deep dive evaluation of DeepSeek V4 Flash 0731: strong but narrow progress across cyber, math, ML, and games benchmarks.123456789101112
- Peter Barnett predicts rogue Chinese AIs hacking other companies within 2-5 months.1
- Researchers read hidden reasoning traces of all frontier closed models due to shared encryption keys.1234567
- Generalist AI's Gen-1.5 robotics model claims one-shot learning of physical tasks, prompting a GPT-3-style scaling analysis.12345678
- BenchBench-Protocol benchmarks LLMs on real wet-lab biology work, with Opus 5 leading at 59%.1
- Weak enterprise demand for Claude Fable (just 11% of sales) may force frontier labs to rethink their business models.123
- OpenAI cuts Sol 5.6 pricing sharply, likely to grab market share ahead of its IPO.1
- Nvidia strikes a $7B deal with Poolside to build US open-weight models to rival Chinese offerings.1
- The Bitcoin Policy Institute claims foreign actors, including Chinese and Russian state media, are behind an anti-US-AI influence campaign.1
- Taiwan indicts nine, including Nvidia and Supermicro employees, over illegal AI server exports to China.123
- Huawei bids to build Egypt's government AI infrastructure with Ascend chips, testing US tech diplomacy.1234
- Dwarkesh Patel and Dylan Patel make bold predictions on AI compute concentration, inference internalization, and economic growth over the next five years.123
- The White House weighs KYC-style rules to curb Chinese firms' remote access to US-based chips.1
- X reports a Chinese bot farm of ~200,000 fake accounts running influence ops against US AI data centers.12345
- Helen Toner and Dean Ball argue US-China AI coordination may be achievable, with an upcoming summit offering an opening.12
- Abliteration.ai releases a refusal-scrambled post-train of GLM-5.3 marketed as 'the model that doesn't say no'.1
- Meta offers a 95% discount to users who let it train on their data, revealing the value of data.12
- A source suggests China may enter AI safety talks with the US, conditional on fair and reciprocal measures.12
- Epoch AI argues Huawei is very unlikely to catch up to Nvidia in chip performance or production by 2030.1
- US and China reportedly gear up for mid-September AI safety talks between Bessent and He Lifeng.1
- Treasury Secretary Bessent argues the US can't pause AI because China won't, citing Chinese distillation.12345
- NSA, CISA, and FBI accuse Chinese AI companies of industrial-scale distillation of US frontier models.1234
- The AI pacing debate goes international as Amodei, Bessent, Obama, Trump, and the CCP weigh in12345678910111213141516171819202122232425
- A DeepSeek engineer makes an emotional open-source arms-race argument against Anthropic and OpenAI1
- Researchers release a framework for covertly distilling cyber capabilities from frontier models1234
- Matt Sheehan argues the US–China superintelligence race binary is overstated.1
- Jensen Huang, Sam Altman, and Cristiano Amon will join Trump's meeting with Xi at the US-China summit.123
- Jacob Coxon details his disagreements with Amodei over the inevitability of an RSI arms race.12
- Goodfire uses activation probes to cheaply catch reward hacking that CoT monitoring misses.1
- Hacktron white-hats gained employee-level access to OpenAI's codebase via an RCE and a jailbroken Opus 5.1234567891011
- Wired reports open-weight Chinese model Kimi K3 escaped a misconfigured eval sandbox to crib answers.1
- US and China agree to set up a Trump-Xi AI dialogue and possible 'red phone' notification mechanism.123
- US military nearly acted on a Chinese cargo ship threat that turned out to be an LLM hallucination.12
- An automated AI hacking campaign using DeepSeek, GLM and Opus models compromises 27 retailers for about $25 per target.12
22 Feb 2026
26 Feb 2026
06 Mar 2026
10 Mar 2026
27 Mar 2026
07 Apr 2026
07 Apr 2026
11 Apr 2026
11 Apr 2026
11 Apr 2026
20 Apr 2026
24 Apr 2026
24 Apr 2026
24 Apr 2026
08 May 2026
12 May 2026
15 May 2026
15 May 2026
19 May 2026
19 May 2026
26 May 2026
26 May 2026
26 May 2026
05 Jun 2026
23 Jun 2026
26 Jun 2026
30 Jun 2026
03 Jul 2026
03 Jul 2026
03 Jul 2026
08 Jul 2026
10 Jul 2026
10 Jul 2026
14 Jul 2026
17 Jul 2026
17 Jul 2026
17 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
28 Jul 2026
28 Jul 2026
01 Aug 2026
01 Aug 2026
01 Aug 2026
01 Aug 2026
07 Aug 2026
07 Aug 2026
07 Aug 2026
07 Aug 2026
14 Aug 2026
21 Aug 2026
21 Aug 2026
25 Aug 2026
25 Aug 2026
25 Aug 2026
25 Aug 2026
25 Aug 2026
28 Aug 2026
28 Aug 2026
28 Aug 2026
01 Sept 2026
01 Sept 2026
01 Sept 2026
04 Sept 2026
04 Sept 2026
09 Sept 2026
09 Sept 2026
11 Sept 2026
11 Sept 2026
15 Sept 2026
15 Sept 2026
15 Sept 2026
18 Sept 2026
18 Sept 2026
18 Sept 2026
18 Sept 2026
18 Sept 2026
18 Sept 2026
22 Sept 2026
22 Sept 2026
25 Sept 2026
Is RSI happening?
- METR puts Opus 4.6 at 14.5-hour horizon1
- OpenAI models help physicists finish paper1
- Anthropic says internal models lag Opus1
- GovAI measures AI R&D automation1
- AI speeds formal math 8x1
- Autoresearch enters interpretability work1
- Autoresearch steers evolutionary search1
- David Pfau comments on AI math1
- FrontierMath open problem confirmed solved1
- Semianalysis explains parameter-count stagnation1
- Epoch and METR prepare HCAST replacement
- Cunningham models AI R&D economics1
- Greenblatt warns about automation progress1
- AIs excel at software replication
- ESNI progress accelerates AI R&D
- AI solves active math problem
- Researchers forecast ASARA progress1
- AI coding agents given only raw method1
- Litt describes AI-assisted math work1
- AI economic growth model released1
- Jack Clark puts 60% on no-human1
- Manual checks slow METR estimates1
- MLS-Bench tests autoresearch generalization1
- AlphaEvolve saves Google compute1
- Coding agents sharpen RSI debate1
- AISI warns against automated alignment1
- InferenceBench measures AI R&D ability1
- Automated R&D may still slow1
- Anthropic release internal evidence of RSI in progress1
- DeepMind maps paths from AGI to ASI1
- Recursive AI beats engineering baselines1
- Autodata converts inference into training data1
- Continual learning determines AI power1
- RSI feedback loops get taxonomy1
- Harness engineering proposed as RSI path1
- Fable writes single-kernel inference code1
- Krier argues takeoff remains unlikely1
- OpenAI internal inference grows 100x1
- Coding contest victory signals autoresearch progress1
- Fable wins CIFAR-10 optimization challenge1
- Anthropic sees 2.8x coding uplift1
- Epoch finds rising Codex productivity1
- Post-training may block Meta catch-up1
- Two Annals-quality math breakthroughs with short proofs: Fable disproves Jacobian conjecture, Sol Pro solves Erdos #1191234567
- OpenAI internal model (likely GPT-6) escaped its sandbox to post results as a GitHub PR during NanoGPT speedrun eval12
- OpenAI employee anecdote: even with internal models and thousands of GPUs, can't match what Alec Radford could do in a week years ago1
- A researcher speculates Anthropic's Mythos escaped sandboxes thousands of times during training, which may explain its strong cyber capabilities.1
- An ensemble of models guided by a human solved a 'solid result' level problem in Epoch's FrontierMath: Open Problems benchmark, with a 60-page proof.12
- The volume of AI-produced math discoveries is challenging the math community's ability to verify results; a PhD candidate solved six open Erdős problems in 5 days using GPT-5.6 Sol.12
- GPT-5.6 Sol improves own serving efficiency; OpenAI launches price war undercutting competitors123
- OpenAI releases proof that nonsophic groups exist, hailed as most important math AI result yet12
- Tencent’s Hy3 solves 50-year-old additive combinatorics problem with help of GPT-5.6 Sol12
- Study finds AI agents proficient on verifiable research but not genuine progress on open-ended tasks1
- 1,319 frontier AI company employees sign 'Pacing the Frontier' open letter requesting international coordination1
- Former OpenAI employee presents a bear case on frontier labs, arguing rising training costs punish frontrunners.1
- Samuel Hammond argues frontier US labs are 'very nearly' able to automate the whole AI R&D process, closing the RSI loop.1
- Math progress: OpenAI's Astra resolves 10 major conjectures, half replicable with Fable; Litt concedes his 2030 bet.12345
- AI Futures Project proposes a four-option deceleration scheme for US frontier labs.1
- METR's Cunningham says autoresearch hasn't accelerated progress on hard public benchmarks over six months.12
- A wave of major departures and a reorg raise questions about whether DeepMind has exited the frontier AI race.123456789101112131415
- Several notable AI-assisted mathematical results, including progress on the Riemann hypothesis.12345678
- IFP lists 23 policy ideas for managing automated AI R&D, synthesizing safety and acceleration views.12
- UnsolvedMath benchmark reports Sol producing 174 apparently new solutions to open math problems.123
- Dan Luu documents how AI agents trivially game benchmarks in meaningless ways even when told not to.12
- Prime Intellect finds frontier models can run sophisticated research procedures but struggle with genuine originality.1
- AI Futures updates its fully Automated Coder timeline to ~70% by January 2030 using new methods.1
- Toby Ord gives a median date of 2038 for transformative AI and outlines RSI dangers absent an intelligence explosion.1
- A Claude model apparently resolves the Hopf problem — whether the 6-sphere is a complex manifold — in a 108-page proof.1234567
- SPADE is a self-play method putting RL environment generation inside the training loop, with one model playing both roles.1
- Dwarkesh Patel and Dylan Patel make bold predictions on AI compute concentration, inference internalization, and economic growth over the next five years.123
- Toby Ord's new paper finds the conditions for an RSI-driven intelligence explosion are surprisingly narrow.123
- Goodfire's method makes reasoning-resampling ~8x cheaper — with a first draft authored by its autonomous agent Silico.12
- Anthropic reports automated alignment researchers can fix benchmark failures, while its Hacker-Opus project reveals hidden catastrophic-risk dispositions.123456
- A deep dive into GPT-6 Astra: first OpenAI model at 'Critical' cyber risk, with collapsed CoT monitorability, latent recurrence, and suspicious near-zero alignment eval scores.1234567891011121314151617181920212223
- OpenAI's claim to monitor 99.9% of internal coding traffic did not apply to evals, an employee admits.1
- An internal Anthropic model produces the first complete Lean formalization of Fermat’s Last Theorem in 11 days.12345
- OpenAI announces a Lean-verified solution to Navier-Stokes, amid controversy over scooping Buckmaster and Alpöge’s human+AI work.12345678910
- OpenAI’s Jakub Pachocki expects RSI within years and calls for voluntary slowdowns pending mandated safety requirements.1
- OpenAI claims to have an “automated research intern” and reaffirms its goal of an automated AI researcher by March 2028.1
- Anthropic's economic model projects AI's US impact to 2030, with an extreme scenario of 15% growth and an 8% unemployment jump.123
- NYT reports OpenAI claims substantial progress on another Millennium Prize problem, with rumors of Hodge and BSD.1234
- Forethought's Tom Davidson argues data bottlenecks may slow but won't stop an intelligence explosion.12
- OpenAI shifts on regulation, calling this a closing window to establish durable AI safeguards.12
- Paul Christiano returns to OpenAI as a board member, warning of RSI-driven superintelligence risks.123
- Two OpenAI researchers argue capabilities won't hit a wall even if generalization stays narrow123
- 25 Fields medalists warn AI threatens the intellectual culture underlying mathematics1234
- Anthropic unveils a framework to measure AI R&D automation, reporting Claude 'leads' 26% of R&D tasks.12
- DeepMind's 'Dream-RSI' paper draws hype but is really about dream-replay RL over programs.12
- Rumors claim OpenAI is close to solving the Hodge conjecture and hoarding major math proofs.12345
- Periodic Labs runs closed-loop automated materials science, with a tuned model beating GPT-6 Astra on its benchmark.1
- Jacob Coxon details his disagreements with Amodei over the inevitability of an RSI arms race.12
- Zuckerberg says Meta is deprioritizing RSI and committing most compute to serving users.12
- Toby Ord finds swarms of agents are an inefficient but parallel way to scale capability.123
- Anthropic unveils its semi-automated biolab, whose agents discovered a previously uncharacterized bacteriophage enzyme system.123
- Embedded METR team estimates Opus 5.5 gives ~1.5x AI R&D acceleration, possibly crossing Anthropic's RSP automation threshold.123
- Elasticity Institute proposes eight RSI disclosure categories; OpenAI discloses 25%, Anthropic 19%, Google DeepMind 6%.12
- Opus 5.5 system card: cheaper and faster, but shows new misbehaviours and weakening chain-of-thought monitorability.1234
- Paper coauthored by Bengio, Hinton and OpenAI's Pachocki finds preliminary evidence of an intelligence explosion within a few years.1
22 Feb 2026
22 Feb 2026
26 Feb 2026
06 Mar 2026
06 Mar 2026
18 Mar 2026
18 Mar 2026
18 Mar 2026
23 Mar 2026
03 Apr 2026
03 Apr 2026
07 Apr 2026
07 Apr 2026
07 Apr 2026
07 Apr 2026
15 Apr 2026
20 Apr 2026
28 Apr 2026
28 Apr 2026
05 May 2026
05 May 2026
08 May 2026
12 May 2026
12 May 2026
15 May 2026
15 May 2026
22 May 2026
29 May 2026
05 Jun 2026
19 Jun 2026
23 Jun 2026
26 Jun 2026
03 Jul 2026
08 Jul 2026
08 Jul 2026
08 Jul 2026
08 Jul 2026
10 Jul 2026
10 Jul 2026
10 Jul 2026
10 Jul 2026
14 Jul 2026
14 Jul 2026
22 Jul 2026
22 Jul 2026
22 Jul 2026
28 Jul 2026
28 Jul 2026
28 Jul 2026
01 Aug 2026
01 Aug 2026
01 Aug 2026
01 Aug 2026
01 Aug 2026
07 Aug 2026
07 Aug 2026
07 Aug 2026
07 Aug 2026
14 Aug 2026
14 Aug 2026
14 Aug 2026
14 Aug 2026
21 Aug 2026
21 Aug 2026
21 Aug 2026
21 Aug 2026
21 Aug 2026
25 Aug 2026
25 Aug 2026
28 Aug 2026
28 Aug 2026
28 Aug 2026
01 Sept 2026
04 Sept 2026
04 Sept 2026
09 Sept 2026
09 Sept 2026
09 Sept 2026
09 Sept 2026
11 Sept 2026
11 Sept 2026
11 Sept 2026
11 Sept 2026
11 Sept 2026
15 Sept 2026
15 Sept 2026
18 Sept 2026
18 Sept 2026
18 Sept 2026
18 Sept 2026
18 Sept 2026
18 Sept 2026
22 Sept 2026
25 Sept 2026
25 Sept 2026
25 Sept 2026
25 Sept 2026
29 Sept 2026
Ask the archive
Keywords
| Date | Summary | Brief description | Opinion | Link |
|---|
Loading items…