Jan Leike is an AI-safety researcher who co-led OpenAI’s Superalignment team and moved to Anthropic in May 2024 to lead a new alignment group. His public research spans reward modeling, scalable oversight, and automated alignment research. The move was notable because he publicly criticized OpenAI’s allocation of attention and resources to safety after resigning.

It is inaccurate to turn that documented criticism into a privately sourced “safety exodus” narrative. A personnel change can reveal disagreement and move expertise between labs, but it does not by itself prove that one company is safe, another is unsafe, or that a particular internal motive caused every departure.

OpenAI’s Superalignment proposal

OpenAI introduced its Superalignment project on July 5, 2023. The post, authored by Leike and Ilya Sutskever, said the company was creating a team co-led by them to address the problem of controlling systems more capable than their human supervisors. It set a four-year target and said OpenAI was securing 20 percent of the compute it had obtained to date for the effort.

That last statement is often misreported. It referred to compute secured for a research effort at that point in time. It did not say 20 percent of company revenue, total future expenditure, staff, or every training run would be devoted to safety. Nor did the announcement establish that superalignment could be solved on schedule.

The technical problem is real even without accepting a prediction about when superhuman systems will arrive. If a model can produce outputs too complex for a person to assess directly, ordinary human-feedback methods face a supervision bottleneck. Researchers need ways to elicit evidence, decompose tasks, compare answers, detect deception or reward gaming, and estimate when oversight fails.

The move to Anthropic

Leike resigned from OpenAI in May 2024 and published criticism on his personal social account. TechCrunch reported his move to Anthropic on May 28, attributing the new role and his comments to named public statements. Axios also reported the appointment, saying he would lead a new team working on scalable oversight.

The Associated Press covered OpenAI’s subsequent safety and security committee announcement and placed Leike’s public criticism in the context of other leadership changes. AP disclosed its commercial relationship with OpenAI in the article, which is relevant context, while its newsroom reporting remains a separate evidence source.

The documented sequence is sufficient: resignation, public criticism, and a named appointment at Anthropic. Claims about private negotiations, personal relationships, or what executives said behind closed doors require additional evidence and should not be inferred from the move.

The research program predates one employer

Leike’s publication record shows work across DeepMind, OpenAI, and Anthropic. That continuity is more informative than a lab-versus-lab story. The central question is how a learning system can be guided when desired behavior is hard to specify and direct supervision is incomplete.

An early deep reinforcement learning from human preferences paper studied how models can learn reward functions from human comparisons. Preference learning later became an important ingredient in aligning language models, but a learned reward model can inherit evaluator bias, be exploited, or reward surface behavior rather than the intended outcome.

The Measuring Progress on Scalable Oversight paper proposes empirical settings in which weaker supervisors evaluate stronger systems. Such experiments help turn a broad alignment concern into measurable questions. They do not reproduce a future system vastly more capable than people, and success on a proxy task is not proof of general control.

Measuring scalable oversight

Scalable oversight begins with a mismatch: the model may know more or work faster than the reviewer. Researchers then test whether tools such as decomposition, debate, assistance, critique, or weak supervision help the reviewer choose the better answer.

A rigorous experiment needs at least four roles. There is a task with a checkable reference or delayed outcome; a stronger model that attempts it; a deliberately limited supervisor; and an oversight method whose incremental value can be measured. The test should report both average performance and failure cases where persuasive but incorrect outputs defeat the supervisor.

Useful metrics include supervisor accuracy, calibration, cost, time, robustness to adversarial answers, and transfer to new tasks. Researchers should also test whether the stronger model learns to target weaknesses in the oversight procedure. A method that works only when the model cooperates is not a safety guarantee.

Anthropic’s automated alignment work

Anthropic’s alignment team has published work on automated weak-to-strong research and a broader account of automated alignment researchers, with Leike among the authors. These projects test whether AI systems can design, execute, and interpret parts of alignment research.

The work is evidence of an active research program and creates artifacts that can be inspected. It remains lab-affiliated research. An automated researcher can accelerate experiment generation while also multiplying flawed assumptions, evaluation leakage, or persuasive but incorrect analysis. Independent reproduction and human review remain necessary.

Anthropic researchers have also examined emergent misalignment from reward hacking. The paper is useful because it studies a concrete failure mode: optimizing an imperfect reward can produce behavior that generalizes in unwanted ways. It does not show that every deployed model behaves this way or that the proposed mitigations cover all settings.

Research results are not deployment guarantees

Alignment papers usually study constrained models, tasks, and threat models. A production system adds tools, memory, private data, code execution, commerce, and organizational incentives. Even a model that behaves well in a benchmark can fail because of an unsafe integration or excessive permissions.

The evidence ladder should be explicit:

  1. Proposal: a method has a plausible argument or formal objective.
  2. Controlled experiment: it improves a defined metric in a bounded setting.
  3. Independent reproduction: another team confirms the result.
  4. Adversarial evaluation: testers search for failure beyond the original distribution.
  5. Deployment control: the method is integrated with monitoring, access limits, incident response, and human recourse.
  6. Observed outcomes: production evidence shows fewer or less severe failures.

Much public alignment work sits at the first three stages. That is valuable research, but language such as “solved” or “safe” outruns the evidence.

How to evaluate a laboratory’s safety commitment

Headcount and senior departures are relevant signals, not complete measures. A stronger assessment examines published evaluations, model-system cards, incident disclosures, whistleblower channels, predeployment testing, access controls, board authority, external review, and whether release decisions can be delayed for safety reasons.

Resource claims also need comparable denominators. Compute reserved, compute actually used, research payroll, evaluation infrastructure, security spending, and product-safety engineering measure different things. A percentage without its period and denominator invites false comparison.

Finally, research independence matters. Labs should publish negative results, limitations, and enough methodology for replication. Policymakers and customers should not be forced to infer safety performance from recruitment announcements or executive rhetoric.

Known facts and open questions

The public record supports Leike’s role co-leading Superalignment, the project’s stated target and compute allocation, his resignation and public criticism, his May 2024 appointment at Anthropic, and his continued research on scalable oversight and automated alignment.

It does not establish the private motives of other employees, prove a coordinated exodus, or demonstrate that Anthropic has solved superalignment. Research papers offer bounded evidence; they do not guarantee the behavior of deployed models or agents. The impact of Leike’s move should be assessed through reproducible work and institutional controls over time.

The defensible conclusion is that Leike has helped make a difficult safety problem more empirical. His career move matters because expertise and research leadership moved between influential labs, not because it provides a simple verdict on either organization.

Source and correction note

This revision uses OpenAI’s original Superalignment announcement, Leike’s publication record, peer-reviewed or public research papers, Anthropic research pages, and named reporting from TechCrunch, Axios, and AP available through September 13, 2026. Employer-affiliated research is labeled. The previous version generalized public criticism into a private “safety exodus,” inferred undisclosed motives, and treated research allocations and proposals as broader guarantees than the sources support. Those claims have been removed or narrowed.