Summary: On May 20, 2026, Anthropic co-founder Jack Clark delivered the 2026 Cosmos HAI Lab Lecture at Oxford University. He predicted a 60% chance of recursive self-improvement — AI systems building their own successors autonomously — by end of 2028. He said AI will help produce a Nobel Prize-winning discovery within 12 months. And he argued that human autonomy, not just extinction risk, is the defining challenge of the next decade.
The Event
The 2026 Cosmos HAI Lab Lecture took place on May 20 at the Sohmen Concert Hall in Oxford’s Schwarzman Centre for the Humanities. Clark, a founding fellow of the Cosmos Institute, delivered the lecture under the title “Change is inevitable. Autonomy is not." It was co-hosted with Oxford’s Human-Centered AI Lab (HAI Lab) and the Oxford Institute for Ethics in AI.
The event included a fireside chat with Cosmos Institute founder and Chair Brendan McCord and a philosophical response from Prof. Philipp Koralus, McCord Professor of Philosophy and AI and Director of the Human-Centered AI Lab within Oxford’s Institute for Ethics in AI. (The Institute itself is directed separately by Prof. Edward Harcourt.)
The backdrop matters: weeks earlier, Anthropic had previewed Claude Mythos — a model that, per Anthropic’s own account, reached a level of coding capability that could “surpass all but the most skilled humans at finding and exploiting software vulnerabilities” and autonomously identified thousands of zero-day vulnerabilities “in every major operating system and every major web browser,” developing working exploits without human steering. Anthropic chose not to release Mythos publicly; instead it gave scoped access to a dozen launch partners (including AWS, Apple, Google, Microsoft, and CrowdStrike) plus more than 40 additional organizations and open-source maintainers to harden critical software first, while saying it was in ongoing discussions with US government officials about the model’s offensive and defensive capabilities. The Council on Foreign Relations described its attack capabilities as ones “that veterans of the cybersecurity industry previously considered to only exist in the realm of science fiction.” In other words, Clark was speaking not from theory but from recent, concrete experience: Anthropic’s own safety predictions had already materialized at the time he took the stage.
The 2028 Prediction: Recursive Self-Improvement
The most precise and far-reaching claim came from Clark’s essay in Import AI #455, published May 4, 2026 — about two weeks before the lecture:
“I reluctantly come to the view that there’s a likely chance (60%+) that no-human-involved AI R&D — an AI system powerful enough that it could plausibly autonomously build its own successor — happens by the end of 2028.”
At Oxford, he framed it this way: “It’s more likely than not that we have an AI system where you would be able to say to it: ‘Make a better version of yourself.’ And it just goes off and does that completely autonomously.”
His probability breakdown, from the same Import AI essay:
- ~30% chance by end of 2027 (“If you had to push me for a 2027 probability, I’d say 30%.")
- ~60% chance by end of 2028 (“I think there’s a ~60% chance we see automated AI R&D… by the end of 2028.")
Recursive self-improvement is significant because it removes humans from the improvement loop. Today’s AI systems are improved by human engineers with human-set objectives. A system that can autonomously improve itself could, in principle, iterate far faster than any human-supervised process — and do so in ways humans can’t fully evaluate.
Why Clark Is Worried About Alignment Breaking Down
Clark’s deeper worry is not just that recursive self-improvement will happen, but that it will break alignment. From the Import AI essay:
“Alignment techniques that work today may break under recursive self-improvement as the AI systems become much smarter than the people or systems that supervise them.”
He identified a compounding error problem: no alignment technique is 100% accurate. In his own example, a 99.9%-accurate technique becomes 95.12% accurate after 50 generations of self-improvement, and 60.5% accurate after 500 generations. At scale, alignment methods degrade unless someone solves the hard problem of alignment before recursive improvement becomes real.
Nobel Prize Within 12 Months
Clark stated plainly that AI will work with humans to make a Nobel Prize-winning discovery within 12 months of May 2026 — framing this not as a remote possibility but as a near-certainty.
Separately — in a Channel 4 News interview published May 4, 2026, about two weeks before the Oxford lecture — Clark framed the broader pace of change this way: AI’s impact will be “10x larger and 10x faster than the Industrial Revolution.”
Additional near-term predictions from the Oxford lecture:
- Bipedal robots assisting tradespeople: within 2 years
- AI-only companies generating millions in revenue: within 18 months (this may already be true, depending on how you define “AI-only”)
Non-Zero Existential Risk — Still on the Table
Clark maintained the position Anthropic has held since its founding: that scenarios where advanced AI kills everyone on the planet remain plausible.
He said there remain “plausible scenarios in which the technology had a non-zero chance of killing everyone on the planet” and that it was “important to clearly state that that risk hasn’t gone away.”
He acknowledged that this framing has attracted significant criticism — including from the Trump White House. White House AI czar David Sacks had previously attacked Clark on X, accusing him of deploying a “sophisticated regulatory capture strategy based on fear-mongering” that was “principally responsible for the state regulatory frenzy that is damaging the startup ecosystem” — a claim covered by The Decoder — positioning safety concerns as a commercial tactic by one of the major US AI labs.
Clark pushed back on the characterization directly at Oxford, arguing that geopolitical competition is “drowning out the larger existential-to-the-species aspects” of AI development — that the race dynamic is making it harder, not easier, to have clear-eyed conversations about risk.
The Core Thesis: Autonomy, Not Just Extinction
The lecture’s title — “Change is inevitable. Autonomy is not.” — signals that Clark’s deeper concern is something subtler than catastrophic scenarios.
He argued that most people are in denial about the capabilities of current AI models, “let alone those coming down the track in six months.” He compared the failure to prepare for AI to the institutional failure to prepare for Covid-19, making the same point he had made in his own account of the fireside chat: “Dealing with COVID highlighted that, though there’d been some modeling of what would happen, if you had very fast take-off propagating viruses the world would break very quickly” — his argument being that AI preparedness requires the same advance modeling, not improvised crisis response.
The fireside chat offered a more everyday illustration. In his own recap of the discussion, Clark described a group of 13-year-old boys who, about a year earlier, decided “to live by the Claude and die by the Claude” — first as parody, then as an earnestly adopted identity, doing whatever the AI told them. He called this “clearly problematic in that everyone needs a part of their life where you’re making your own decisions, including mistakes.” That capacity atrophies when it’s outsourced.
This positions Clark as concerned with two distinct failure modes:
- Catastrophic: Recursive self-improvement + alignment degradation → outcomes humans can’t predict or stop
- Quotidian: Behavioral capture → gradual erosion of human agency that never hits a single identifiable crisis point
The second failure mode is harder to regulate against and perhaps more likely to arrive first.
“A Tale of Two Anthropics”
Time magazine’s May 22 piece, published two days after the Oxford lecture, captured the tension in a phrase: “A Tale of Two Anthropics."
The reporter described “a profound sense of whiplash” from attending both an Anthropic developer event focused on Claude Code’s productivity benefits and Clark’s Oxford lecture within days of each other. Anthropic simultaneously:
- Argues that it has built perhaps the most powerful technology in human history and that its risks must be taken seriously at a civilizational level
- Sells that same technology aggressively to developers and enterprises to fund further development
Time’s reporter closed by noting the irony directly: if Clark’s own predictions are right, “the developers Anthropic is currently so dependent on will find themselves replaced by machines faster than they can say ‘recursive self-improvement’” — and suggested that may be why those predictions “didn’t find their way into Wednesday’s keynotes.”
Context: The “Intelligence Explosion” Document
One week before the Oxford lecture, Axios reported (May 7, 2026) that a five-page research-agenda document from the Anthropic Institute — the research and early-warning arm Clark leads, run alongside Anthropic’s Long-Term Benefit Trust — had used the term “intelligence explosion” in an official context for the first time. The term, long used in AI safety theory to describe the moment when recursive self-improvement becomes self-sustaining, had previously been kept out of Anthropic’s formal communications.
The document reportedly also proposed that Cold War-style AI crisis hotlines between major AI powers may be necessary infrastructure — an analogy Clark himself has used in public writing.
The Oxford lecture, then, was not an isolated moment of alarm. It was Clark making public what Anthropic had been saying internally.
What to Watch
Clark’s 60% by 2028 figure is specific enough to be falsifiable. Key indicators to watch:
- Whether AI systems can be given open-ended self-improvement tasks and produce models that score higher on benchmarks without human engineering intervention
- Whether alignment research produces techniques that remain robust across generations of model improvement
- Whether any AI lab announces recursive self-improvement capability before 2028
For immediate stakes: Clark’s Nobel Prize prediction lands no later than May 2027. If AI is credited as a co-contributor to a Nobel Prize by then, it would confirm his near-term timeline is calibrated. If it doesn’t happen, it will be worth revisiting the rest of his predictions.
The existential risk framing and the autonomy framing are not in tension — Clark seems to believe both are real, operating on different timescales. The question he leaves open is which one regulators, companies, and users will actually take seriously before it becomes unavoidable.
Sources: Oxford Institute for Ethics in AI — 2026 Cosmos HAI Lab Lecture event · Oxford HAI Lab — event page · Import AI #455 — Jack Clark essay · The Guardian — Clark’s Oxford predictions (May 21, 2026) · Time — A Tale of Two Anthropics (May 22, 2026) · Axios — Anthropic intelligence explosion document (May 7, 2026) · Cosmos Institute — Clark’s own fireside chat recap · Anthropic — Project Glasswing / Claude Mythos · Council on Foreign Relations — Claude Mythos analysis
ChatForest covers AI tools, models, and the industry driving their development. See our related coverage: Anthropic’s First Operating Profit | Anthropic ARR analysis