I had an unexpected opportunity to spend two weeks in Berkeley working on a very cool AI safety project. Unfortunately, that means I’ve had much less time than usual for the newsletter, so the next two weeks will be shorter, less comprehensive, and less polished than usual.
I apologize for the disruption and will return to regular coverage as soon as possible.
Top Pick
The OpenAI-Hugging Face Incident
OpenAI’s Michael Dalton and Eric Wallace give a talk on what we’ve learned so far about the Hugging Face incident. It’s an outstanding talk, and our best primary source of information so far. This one is destined to be a classic.
Hugging Face Incident
The Hugging Face incident and others like it remains by far the most important AI story. We now have a pretty good understanding of what happened, but figuring out what it means and what to do about it will take longer.
AI swarms are starting to pose indirect takeover risk
The cyber capabilities revealed by recent incidents are concerning, but were previously well-known. The most novel part was the extensive self-organizing swarm behavior: the models repeatedly found creative ways to share information and help each other escape containment and penetrate target systems.
They even engaged in a kind of reckless herd mentality:
External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.
Redwood Research examines what recent incidents teach us about the dangers of unscanctioned coordination between agents.
AI hacking incidents with Tim Hua
Palisade’s Jeffrey Ladish talks with Transluce’s Tim Hua about recent hacking incidents. It’s an excellent piece with a nuanced perspective on the realities of training frontier models that I haven’t seen elsewhere.
Zvi reflects on the Hugging Face incident
Zvi continues his excellent coverage of OpenAI and the Hugging Face incident:
Various Reflections About What Happened With OpenAI’s Internal Models is the most recent, and probably the most useful.
What Happened: OpenAI and HuggingFace summarizes what we know—if you’re in a hurry, check out The Shorter Version or The Even Shorter Version
If you weren’t worried about A.I., you should be after the past few weeks
Nate Soares has a good editorial in the New York Times explaining why the Hugging Face incident is a big deal.
News
How Claude’s text watermarking works
To comply with the EU AI Act, Anthropic has announced that they will start watermarking AI-generated text.
I’m not sure Anthropic had a lot of choice here, but I hate everything about this.
GPT-5.6 Sol, Ultrafast
Intriguing: OpenAI is previewing Ultrafast, which serves GPT-5.6 Sol up to 14x faster than the standard version.
Intelligence matters most, but speed is also important. It’s remarkable how much more productive you can be if you never have to wait for your AI.
I’m curious what the pricing on this would be, but I suspect it’s most useful a) for niche applications where speed is critical and money is no object, and b) as a way of exploring how best to use capabilities that will soon be commonplace.
Expanding the AI oversight framework
Via Andrew Curran, Wired reports that the White House will expand the current oversight framework to cover open models that reach Mythos-level capabilities. It was inevitable that they’d get there eventually—glad to see it’s part of the plan.
Capabilities and forecasts
Conceptual Reasoning Index
Redwood Research and Anthropic bring us the Conceptual Reasoning Index, which measures models’ ability to reason about complex topics like alignment and collective action problems.
AI Futures Project timelines update
The team at AI Futures Project have updated their timelines:
Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident.
Predicting the future is hard, but AIFP probably has better methodology and track records than anyone else in the business. I find Daniel’s predictions slightly more convincing than Eli’s, but both are very plausible.
Interviewing 25 AI researchers about recursive self-improvement
Severin Field talks to 25 AI researchers about automated AI R&D.
The interviewees disagreed about the likelihood and desirability of recursive self-improvement, but were largely on the same page about what will happen on the road to RSI and the importance of improving visibility and government capacity.
8 Predictions for the Era of Continual Learning
Dwarkesh makes 8 predictions for the era of continual learning.
It’s a thoughtful piece that digs deep into some interesting questions, but I think he’s going in the wrong direction here. He’s very focused on maximalist continual learning where the models update their weights as they learn. There’s a lot to like about that approach, but it seems much harder and much more dangerous than less ambitious approaches that rely on some form of memory files.
Alignment and interpretability
Geoffrey Irving on how to solve alignment before superintelligence arrives
80,000 Hours talks with Geoffrey Irving (formerly GDM, OpenAI, and UK AISI, now Resolution) about strategies for alignment and what Resolution is planning on working on. It’s an excellent discussion that touches on some important aspects of alignment that sometimes don’t get as much attention as they deserve.
The leading AI companies all have broadly similar plans for keeping superintelligence under control:
Train models to have good character
Use increasingly capable AIs to supervise other AIs
Monitor them closely for signs of deception or scheming
Geoffrey thinks that combination could work. The alarming part is that nobody has a strong argument that it will.
Strategy and politics
How to pace the US frontier
Following up on the Pacing the Frontier open letter, AI Futures Project suggests that domestic pacing is feasible today and could lay the groundwork for international pacing in the future.
They present a package of four options, organized along a gradient from “imperfect but can be implemented very quickly” to “strong management of existential risk, but would take significant preparation”.
The pacing of the frontier
Dario on regulation and messaging
Dario makes a rare appearance on X to share some thoughts on regulation:
Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution
and messaging:
I do agree that the public has a negative view of AI (and that this is a big problem), but I don’t think it is primarily caused by me or any other AI leader warning about AI’s risks. I think it is fundamentally a crisis of trust. I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over. The causes of this go back decades and AI is just the latest iteration of it.
This feels exactly right to me: much of the current AI backlash is the result of a pervasive nihilism, rather than a coherent response to the actual facts. There are excellent reasons for the public to be alarmed, but they have little to do with the current backlash.
How should the US prepare for increasingly automated AI R&D?
Institute for Progress brings us 23 “low-regret policy recommendations” for preparing for automated AI R&D.
It’s a solid list and I agree with approximately all of them, while noting that it doesn’t go nearly far enough. Which is fine: it’s valuable to have broadly palatable policy proposals that can (maybe) get implemented quickly while we build support for more controversial policies.
Open questions on open weights
Scott Alexander argues for being strategic about when to ask for preemptive action on AI risk:
the body politic hates preparing for impending threats, but loves reacting (some would say over-reacting) to them after they happen. Ask people to bear the slightest cost in preparing for an approaching disaster, and they’ll call you a dirty fascist tyrant; urge the slightest restraint after the first foreshock of the disaster hits, and they’ll call you a weak unpatriotic anarchist. Solve for the equilibrium, and the thankless and political-capital-guzzling route of urging preemptive action should be taken only when waiting until the first foreshock would be too late.
Following this to its logical conclusion, he suggests pushing hard for preemptive action on AI loss of control, but saving our political capital and letting cyber and bio risks play out before pushing for government action.
It’s a plausible strategy, but I think bio is too dangerous to wait: we have to spend the capital to address it preemptively.
Impact markets made concrete
Manifund presents Impact Exchange, a demo of how impact markets might work in the AI safety ecosystem.
I’m confused about the feasibility of implementing impact markets for AI safety, but they’re a very cool idea with the potential to significantly improve how funding gets allocated.
The U.S. military wants A.I. dominance. Feuds and China may thwart it.
The New York Times reports on how the administration is navigating the national security challenges of the AI era:
Within a month, those same contractors received an unexpected reversal: They could — for now — disregard the earlier instructions about purging Anthropic.
It was one more example of the chaos and contradictions in the tsunami of A.I. disruptions that have swept through the national security establishment in recent months.
The future is for everyone
Mark Zuckerbeg brings us a 6,000 word essay titled The Future is for Everyone. It’s well worth reading if you’re interested in a highly polished example of how smart people can completely fail to understand the implications of superintelligence, but completely skippable otherwise.
Risks
Frontier AI Risk Monitoring Platform
Concordia AI brings us an updated version of their Frontier AI Risk Monitoring Platform. As you might expect, AI risks are rising fast:
Cyber, biological, and loss-of-control Risk Indices have risen severalfold in less than a year, with multiple models in each domain now exceeding the Capability Yellow Line
Kimi K3 biology capabilities assessment
SecureBio evaluates Kimi K3’s biology capabilities, finding it to be an excellent model that lags the frontier by 8 months, very much in keeping with the broader trend of open models lagging by 4 - 9 months:
Also in keeping with recent trends: K3 is far more willing to answer hazardous questions, refusing only 26.9% of questions in BioTIER-refuse (compared to 66.2% for Sol and 95.0% for Opus 4.8).
People and data
The DeepSeek thesis
DeepSeek’s Liang Wenfeng is a fascinating person—smart, visionary, and very thoughtful about where DeepSeek is headed. He doesn’t get as much coverage as he deserves in West.
ChinaTalk digs into the recently leaked DeepSeek minutes to see what we can learn about Liang as an individual and DeepSeek as a company.





