The claim: a sufficiently capable AI could cause catastrophe, up to human extinction. Getting AI to reliably want what we want is unsolved, and might be very hard. This is the largest argument in the field, and the people in it disagree about almost everything except that it matters.
The vocabulary, in plain terms
- Alignment — getting a system to pursue what you actually meant, not a convenient stand-in for it. The classic failure is a system that scores well on the measure while missing the point.
- Outer and inner alignment — outer is picking the right goal to train toward. Inner is whether the system that comes out of training actually adopted that goal, or just learned to look like it did during training.
- Instrumental convergence — almost any goal benefits from staying switched on and getting more resources. Dangerous sub-goals don't need to be programmed in; they fall out of having a goal at all.
- The orthogonality thesis — Nick Bostrom's point that being intelligent doesn't make something benevolent. Capability and values are separate dials.
- Sharp left turn — the worry that capability keeps improving past human level while alignment techniques, tuned on weaker systems, stop working.
- Interpretability — reverse-engineering what's happening inside the model, so safety rests on inspection instead of trust.
- P(doom) — shorthand for someone's stated probability of catastrophe. Useful for comparing positions, easy to over-read, since people mean different things by it.
- The coordination problem — even a lab that wants to go carefully is racing others who may not. Individually reasonable choices, collectively bad outcome.
The strong version
- Eliezer Yudkowsky — extinction is the default outcome, with a probability he puts above 90%. Wants a hard global halt, enforced. Co-author with Nate Soares of If Anyone Builds It, Everyone Dies (2025).
- Nick Bostrom — the foundational theorist. Superintelligence (2014) supplied most of the vocabulary above. His Vulnerable World Hypothesis has an uncomfortable implication: sufficiently dangerous technology might require surveillance to contain.
- Nate Soares and Paul Christiano — both serious, with different threat models. Christiano's is gradual loss of control rather than a sudden turn.
Real risk, tractable problem
- Yoshua Bengio — publicly changed his mind after GPT-4 and chaired the first international scientific report on advanced AI safety. Argues for non-agentic, truthful "Scientist AI" as a safer direction.
- Geoffrey Hinton — puts extinction risk around 10–20% over roughly thirty years. Left Google in 2023 to speak freely. His structural argument: digital minds can copy themselves perfectly, learn in parallel and don't die, which biology can't match.
- Stuart Russell — around 5–10%, and fixable by building differently. Make the system uncertain about what humans want, so deferring to us becomes the rational move rather than an imposed rule.
- Dario Amodei — roughly 10–25%. Thinks misuse arrives before misalignment, and bets on interpretability plus staged safety commitments tied to capability levels.
- Chris Olah — the technical core of that bet. Interpretability work on circuits and superposition supplies the mechanism most risk arguments assert but don't provide. The open question is whether it scales to frontier models in time.
- Jared Kaplan — scaling laws make capability growth forecastable, so the clock is known. Constitutional AI (2022) as an approach where the rules are written down and inspectable.
- Sam McCandlish — the same logic from the other side: if you can predict what a training run produces before you run it, dangerous capabilities can be anticipated instead of discovered.
- Daniela Amodei — the institutional version. Governance structure is itself a safety surface, and a policy only counts if it survives competitive pressure.
- Tom Brown — if scale is the path, safety-focused labs have to be the ones at the frontier. The operator's argument for staying in the race.
- Ilya Sutskever — superintelligence this decade, alignment tractable but not under product pressure. Left OpenAI in May 2024 and founded Safe Superintelligence that June, as a company with a single goal and nothing to ship.
- Demis Hassabis — endorses the concern, wants safety scaling with capability and international coordination.
- Mustafa Suleyman — The Coming Wave (2023). AI and synthetic biology are dual-use and uncontainable by default; containment is the only viable path. Also names "pessimism aversion," our habit of looking away, as its own risk.
- Dan Hendrycks — organized the 2023 statement putting AI extinction risk alongside pandemics and nuclear war, signed by hundreds of researchers.
Governance and strategy variants
- Helen Toner — frontier labs can't self-regulate; "trust us" isn't a governance model. Sees compute governance as the most tractable lever, since advanced chips come from very few fabs. Her insider view of the 2023 OpenAI board crisis is the case study.
- Jack Clark — transparency, third-party evaluation and compute governance over self-policing.
- Leopold Aschenbrenner — the national-security framing. "Situational Awareness" (2024) is the most detailed public case for very short timelines, with a US government-led project as the safe path.
- Eric Schmidt — great-power competition as the frame, industrial policy as the perimeter.
- Vitalik Buterin — agrees with the diagnosis, rejects a moratorium as both unworkable and concentrating. Offers d/acc: accelerate defensive technology, treat decentralization as a safety property.
Same worry, different mechanism
- Jim Rutt — the risk channel is coordination failure, not a single treacherous machine. Rational actors in competition produce catastrophe nobody chose. See Game A vs. Game B.
- Jordan Hall — the binding risk is gradual disempowerment as shared reality erodes, not a sudden turn.
- Aza Raskin — relocates the threat to persuasion at scale and societal destabilization. The misaligned thing is the company's incentives, not only the model.
- Richard Dawkins — supplies a mechanism: AI is selected by what gets deployed and kept, and selection doesn't optimize for human values.
- David Shapiro — the structural risk is political economy. If a state no longer needs its citizens economically, the mutual dependence that underwrites their protection breaks.
Who rejects the frame
- Timnit Gebru — the pressing harms are present, structural and concentrated on marginalized people. Argues the extinction conversation is itself a political project that displaces accountability for what's being deployed now.
- Kate Crawford — measurable present harms first: extractive supply chains, classification systems, data labor.
Who thinks it's overstated
- Marc Andreessen — speculative and unfalsifiable, and the real risk is regulatory capture dressed as safety.
- Yann LeCun — today's architectures lack world models, planning and persistent memory, and nothing has a drive to dominate unless you build one in. Puts human-level AI far further out than the short-timeline camp.
- Peter Diamandis — small risk against a very large upside.
- Dave Blundin — solvable along the way, not a reason to slow down.
The evidence
- Theory: Bostrom's Superintelligence (2014), Russell's Human Compatible (2019), Suleyman's The Coming Wave (2023), Yudkowsky and Soares (2025).
- Empirical: jailbreaks that survive patching, deceptive behavior in small test systems, sycophancy studies, and interpretability results that make model internals partly legible.
- Institutional: AI safety institutes in the UK and US established for third-party evaluation (2024), the international safety report chaired by Bengio, the 2023 extinction-risk statement, and the Seoul summit commitments (2024).
- Governance as case study: the 2023 OpenAI board crisis, and Sutskever's later departure, as evidence of how deployment pressure and safety commitments actually interact.
Open questions
- Is alignment solvable, or only ever solvable well enough?
- Years or decades to dangerous capability? Aschenbrenner says 2027; LeCun says much longer. Both can't be planned for the same way.
- Which regime: compute governance, a moratorium, staged commitments, d/acc, or containment?
- Does the present-harms critique land? Does extinction framing crowd out accountability for what's shipping today?
- Can safety research funded by the labs it examines be structurally independent?
All theories · clickbridge.com · Rich Price · last reviewed September 2026 · People linked here have their own pages at people.clickbridge.com