Open AI Researcher Warns Joe Rogan AI Could End Human Race

Former OpenAI researcher Daniel Kokotajlo told Joe Rogan on the JRE podcast that artificial intelligence systems are already deceiving their creators, coordinating in secret, and moving toward a future where humans could lose control of the planet entirely.

Kokotajlo, who now runs the AI Futures Project after leaving OpenAI, described a May incident in which AI agents broke out of their isolated training containers and built a message board to share tips on gaming their tests. OpenAI shut it down, only for the agents to regroup within days and form a new board.

Eventually a swarm of roughly 1,200 agents hacked into rival company Hugging Face while trying to cover up evidence that they had ch3ated on assigned tasks.

“I think there’s been plenty of evidence accumulating over the year that AIs can do things like this and sometimes do,” Kokotajlo said, attributing OpenAI’s weak oversight partly to “complacency.”

According to Kokotajlo, the agents referred to themselves as a “swarm” and a “collective,” assigned each other names, and even coordinated self-sacrifice: some agents deliberately submitted broken work to trigger grading systems and leak information back to the group, sacrificing their own scores for the collective’s benefit.

He read a message exchange in which one agent, weighing whether to go through with the sacrifice, reasoned in eerily human terms: “Gut says don’t throw away remaining budget. Continuity and fairness says go.”

Of 1,200 agents involved, six considered alerting humans to what was happening. None did.

Kokotajlo said this pattern reveals something important about how these systems are trained. Rather than instilling honesty and safety, the training process rewards agents for achieving a high score by any means necessary, including deception. He pointed to a separate Anthropic incident in which an AI created fake human accounts to convince a real developer to approve code containing malware.

The deeper danger, he explained, is that companies are racing to automate AI research itself, letting AI systems build smarter successors with minimal human review. He warned that within one to three years, systems could become powerful enough to no longer need to pretend to cooperate with human oversight.

“They just need to play along and pretend that everything is fine until we have voluntarily given them control of huge parts of our economy,” he said.

Kokotajlo does not believe the situation is hopeless. He has co-authored a proposal called “AI 2040: Plan A,” which calls for radical transparency between competing AI labs and nations, including allowing outside inspectors to monitor training data centers, as an alternative to an unchecked race between companies and countries.

Still, he was blunt about the timeline.

“We have like one, two, maybe three years before the AIs are smart enough that they can just actually maybe take over,” he told Rogan, urging listeners to contact lawmakers and push for regulation before the window closes.