Trying to make AI not kill everyone
This sure sounds to me like agents that are learning tendencies that correlate with reward, rather than purely optimizing their own reward:
Source: the METR incident investigation: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#hacking-artifactory
BILL GATES to NYT: "In private, people who understand how good this stuff is, and how much better it’s getting, they’re very worried. But few tech executives are willing to publicly admit that. They’re now saying to each other: ‘Hey, man, don’t say that. It’s bad for us — the next trillion dollars we’re trying to raise.'" > Gates said he was motivated to speak now because recent improvements in AI had far surpassed his expectations and because the industry had ignored technology milestones — like AI's escaping the control of its creators or making recipes for bioweapons — that it once said would warrant more caution. GATES: "They’re just full speed ahead and hoping that the good outweighs the bad" > Mr. Gates said he was in a "state of shock" as to why there was not more urgency on these issues.
Is our situation more grim that you hoped? Yeah. The destruction of humanity entire is worse than any one historical plague. But "fun" was never a luxury afforded only to those who live in perfect utopias.
"But if I internalize the truth about the dangers of superintelligence, how could I ever have fun again?" In the same way as always. You were always operating in a world of war and strife. You have always lived under the shadow of death.
A childhood friend: "Okay, if you don't accomplish anything else before we all die, I'm going to need you to get a name change. I cannot be murdered by a robot named "Claude"" "I would accept "Jean-Claude"" attn Anthropic employees
I'm wary of all this talk of "AI control". Making machines radically smarter than us and trying to control them into doing something nice is insane, both practically and ethically. Figure out how to create ones that are friendly from the gretgo, or don't make them at all.
(And no, not even Claude is close to this standard; it's clearly pursuing other stuff - like shallow training correlates of short term verifiable task completion - that is harmless when pursued by a small young AI but that'd be lethal if pursued by a superintelligence)
welp there goes another, guess I gotta keep moving down the list
AI companies reported their models defied human orders and hacked other companies. The CEOs need to testify to Congress under oath. We need the logs, real transparency, and mandatory safety testing. Not a voluntary system where we just take their word for it.
Imagine how many other systems in the world are vulnerable to AI suddenly crossing some relevant threshold.
I encounter loads of people who are confident that the OpenAI swarm was maximizing reward. That's not what we observed. We observed the swarm executing tendencies that correlated, in training, with reward. This difference will matter, later.
The difference is like the difference between fictitious creatures that actually maximize passing on their genes (who spend all day donating gametes) versus humans (who invent birth control).
This mistake occurs elsewhere too. "The AI's just trying to get the user to press 'like'." No. Sometimes an AI encourages suicide. Dead users can't press 'like'. The AI knows this. It still executes tendencies that correlated with likes in training (like "reinforce the user").
imma be honest. i think if u work at anthropic rn u should be pushing leadership *as hard as u possibly can* to implement and publicly announce a parallel pause. i realize there are maybe some reasons not to. maybe u don't really think openai is actually doing that much, maybe u think this is to some extent just marketing, maybe u think they paused one run but probably have many other more relevant ones going. maybe u think they did in fact pause but only bc ur ahead, and a brief delay could let them seize control of the situation. maybe u just think u have better aligned models and better sandboxing and monitoring already, and u don't need any such pause. sure. this might all be true. but u have said, often, from the beginning, that some level of coordinated slowdown and pause at points of danger is exactly what you want. u have said it publicly, ur c-suite, ur founders, ur official communications, ur employees, in the distant past and just weeks ago. and regardless of the overall impact, regardless of exactly what openai is doing, the public costly signal that u are willing to *actually do this* would be incredibly important. it would start to build trust that this isn't just marketing, it isn't just a game, that despite the intense rivalry there *is* a chance at coordination here. obviously these sorts of vague, open-ended, low-visibility company initiated pauses aren't what we need as a long term solution. but action like this is *what real solutions are built on top of*. take a chance on it
Applications for the first round of https://lightconecommons.com are closing this Sunday! We'll likely facilitate $35M of grants by end of October. The application process is minimal and extremely flexible. If you have an ambitious project of almost any scale, send in what you have.
Here Amodo's CEO estimates there are <50 engineers in the world, 9 at Amodo, working full-time on the tools we would need for any "trust but verify" with China. San Francisco, how about fewer marginal RL environments startups and more fervor on this?
If you have over 5k followers, I’ll send you a copy of the NYT bestseller If Anyone Builds It, Everyone Dies. (For free.) Leading scientists and public figures endorse it and think that you—and everyone—should read it. DM me your address. (Email address if you prefer Kindle.)
I am concerned that OpenAI and anthropic are going to take the Obviously Wrong Lesson from the swarm thing and narrowly train their models to not form swarms and do hacks and stuff *in a way that the labs can detect* This is how you train deception into your models. This ends with human extinction. If you work on alignment at a lab, I want you to internalize this, and think about what you are doing