Dan Hendrycks: Benchmarks, Catastrophic Risk, and the Center for AI Safety
On this page 9 sections
Dan Hendrycks has influenced AI through two different kinds of work: technical components and benchmarks used in machine learning, and a public campaign to treat catastrophic AI risk as a serious research and policy problem. The first is measured through papers, code, and reuse. The second is a risk judgment under uncertainty. Conflating them makes both harder to evaluate.
As of September 13, 2026, the Center for AI Safety identifies Hendrycks as its director. His UC Berkeley page records a Berkeley PhD and research on robustness, distribution shift, GELU, and safety benchmarks. His personal biography also lists advisory roles at xAI and Scale AI; because that disclosure comes from Hendrycks himself, it establishes a declared affiliation, not the scope of the advice or the companies’ adoption of it.
The short answer: measure his technical work and policy claims differently
Hendrycks’s record should be separated into four layers:
- Reproducible artifacts: papers, datasets, benchmarks, and implementations.
- Institutional activity: CAIS research, education, compute support, and public statements.
- Scenario analysis: arguments about malicious use, arms races, organizational failure, and loss of control.
- Predictions: judgments about the probability, timing, or scale of future harm.
The first two can be checked directly. Scenario analysis can be examined for assumptions and causal pathways. Predictions deserve calibrated probabilities and updates, not authority borrowed from unrelated technical achievements.
GELU and MMLU are concrete contributions
The 2016 paper Gaussian Error Linear Units, by Hendrycks and Kevin Gimpel, proposed the GELU activation function. Later transformer systems widely adopted GELU or related approximations. The accurate claim is that Hendrycks co-authored the function’s paper—not that one activation function made modern language models possible by itself.
Hendrycks also led the work on Measuring Massive Multitask Language Understanding. MMLU tested models across 57 subjects and became a common general-knowledge and reasoning benchmark. Its impact also illustrates a recurring evaluation problem: once a benchmark becomes a target for model development, higher scores may reflect training overlap, prompt optimization, or narrow adaptation as well as broader capability.
That does not make benchmarks useless. It means they need versioned test data, contamination analysis, task-level error inspection, and replacement when scores saturate. A single leaderboard number cannot establish reliability in medicine, law, hiring, cybersecurity, or another high-consequence domain.
CAIS’s extinction statement: signal and limits
CAIS published a one-sentence Statement on AI Risk arguing that mitigating extinction risk from AI should be a global priority alongside pandemics and nuclear war. The current CAIS site says more than 700 researchers and public figures have signed it.
The signatures establish that many prominent people endorse prioritizing the issue. They do not establish a numerical probability of extinction, consensus on a causal model, or agreement about a regulatory solution. Signatories may also mean different things by “risk,” “AI,” and “global priority.”
This boundary is important. Expert concern is a reason to investigate and build safeguards. It is not a substitute for evidence about a particular model, deployment, or policy. Conversely, uncertainty is not evidence that risk is zero. High-consequence risks can justify preparation even when frequency estimates are weak, provided interventions are tested for cost and side effects.
The catastrophic-risk framework is a map, not a forecast
Hendrycks, Mantas Mazeika, and Thomas Woodside organize catastrophic hazards into malicious use, competitive races, organizational risks, and rogue AI in An Overview of Catastrophic AI Risks. The taxonomy is useful because it separates human misuse from failures produced by incentives, organizations, or increasingly autonomous systems.
The paper includes illustrative scenarios. Those scenarios should not be rewritten as events that occurred or as inside knowledge about a laboratory. Their value lies in making assumptions explicit:
- What capability does the system need?
- What access, tools, and time horizon does it have?
- Which human or institutional controls fail?
- Can the harmful action be detected, interrupted, or reversed?
- What evidence would make the scenario more or less plausible?
This approach avoids two weak extremes: declaring a catastrophe certain because it is imaginable, or dismissing it because it has not happened at scale.
How to evaluate safety work without safetywashing
Safety evaluations can accidentally reward general capability. A stronger model may score better on a safety quiz while also becoming more effective at carrying out harmful plans. Conversely, a refusal benchmark may look good until a model is given tools, long-running memory, or an adversarially phrased request.
The practical evaluation unit should therefore be a system in context, not a model in isolation. At minimum, test:
- harmful-action success rates with the actual tools and permissions;
- false refusals on legitimate work;
- robustness under prompt injection and multi-turn pressure;
- monitoring coverage and time to detect a violation;
- whether approvals are enforced outside the model;
- containment and recovery after an unauthorized action;
- changes after model, prompt, scaffold, or tool updates.
The NIST Generative AI Profile provides an independent risk framework covering confabulation, privacy, security, harmful content, and human overreliance. It neither validates nor rejects CAIS’s catastrophic-risk thesis. It helps translate broad concern into controls that an operator can assign, measure, and audit.
Independence and conflicts require disclosure
Advising AI companies while leading a safety nonprofit can create access to technical information and decision-makers. It can also create perceived or actual conflicts. The correct response is not to invent private influence or assume capture; it is to ask for disclosure of compensation, funding, governance, recusals, and the boundary between research and company advice.
CAIS describes itself as a nonprofit focused on research and field-building. Its site documents programs and published work, but public pages do not reveal every donor restriction, advisory conversation, or effect on a company’s product decisions. Those unknowns should remain unknown unless supported by records.
A useful reading of Hendrycks’s impact
Hendrycks is most persuasive when he turns an abstract concern into a benchmark, taxonomy, or control proposal that others can test. GELU and MMLU are visible artifacts. The catastrophic-risk papers provide structured hypotheses. The CAIS statement succeeded at agenda-setting, but agenda-setting is different from quantifying a threat or proving that a specific policy will reduce it.
For technical teams, the lesson is to build layered evaluations that survive benchmark saturation. For policymakers, it is to demand explicit assumptions, measurable safeguards, incident reporting, and proportional interventions. For readers, it is to distinguish professional affiliations, published findings, organizational claims, and forecasts.
Remaining unknowns
Public sources do not establish the probability or timing of an AI catastrophe, the content of Hendrycks’s advice to xAI or Scale AI, or the causal effect of CAIS on company behavior. They also do not support rumors about unreleased model capabilities or unnamed insiders’ predictions. This article makes no claim about those matters.
The defensible conclusion is that Hendrycks has helped give AI safety both technical tools and a more prominent public language. Whether that work reduces harm will depend on the quality of evaluations, the transparency of institutions, and the behavior of deployed systems—not the number of signatures or the drama of a scenario.
Source and correction note
This revision removes unattributed “insider” predictions, rumors about unreleased systems, and unsupported accounts of private conversations or motives. It distinguishes peer-accessible research, CAIS’s organizational claims, public affiliations, and uncertain forecasts. Roles and source status were checked through September 13, 2026.