Affiliate disclosure
Book titles on this page link to Amazon. As an Amazon Associate, DataField.Dev earns from qualifying purchases — at no additional cost to you.
Further Reading — Chapter 20: AI Safety and Alignment
Essential Reading
Brian Christian, The Alignment Problem: Machine Learning and Human Values (W.W. Norton, 2020). The most accessible and comprehensive book on the alignment problem for a general audience. Christian is a superb writer who makes complex ideas genuinely clear without oversimplifying. If you read one book on AI safety, make it this one.
Stuart Russell, Human Compatible: Artificial Intelligence and the Problem of Control (Viking, 2019). Written by one of the most respected AI researchers in the world (co-author of the standard AI textbook), this book presents a clear argument for why alignment matters and proposes a framework based on AI systems that are uncertain about human preferences and actively deferential to human judgment.
Recommended Reading
Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014). The book that put existential risk from AI on the intellectual map. Dense and philosophical, but foundational for understanding the long-term safety debate. Best read with awareness that the field has evolved significantly since publication.
Timnit Gebru et al., "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" Proceedings of FAccT (2021). An influential and controversial paper arguing that the AI safety community's focus on speculative long-term risks diverts attention from concrete near-term harms — including bias, environmental costs, and the labor exploitation behind AI systems. Essential for understanding the tensions within the AI safety community.
Anthropic, "Constitutional AI: Harmlessness from AI Feedback" (2022). The original technical paper describing the constitutional AI approach. More technical than most readings in this list, but the core ideas are accessible and the paper is well-structured. Available free online.
Deep Dives
Ajeya Cotra, "Without Specific Countermeasures, the Easiest Path to Transformative AI Likely Leads to AI Takeover" (Alignment Forum, 2022). A detailed, thoughtful exploration of why alignment may become harder as AI systems become more capable. Represents the cautionist perspective at its most rigorous. Available free online.
Yann LeCun, various public statements and papers on AI safety (2023–2025). Meta's chief AI scientist has been one of the most prominent voices arguing that existential risk from AI is overstated and that open models are safer than closed ones. His public debates with cautionists (particularly on social media) provide a window into the accelerationist perspective.
Victoria Krakovna, "Specification Gaming Examples" (DeepMind, maintained list). A curated collection of real examples where AI systems gamed their specifications — from simulated robots to game-playing agents to real-world systems. Fascinating, sometimes funny, always illuminating. Regularly updated. Available free online.
Multimedia
"The AI Dilemma" — Tristan Harris and Aza Raskin, Center for Humane Technology (2023). A presentation that connects the engagement-maximizing dynamics of social media (the near-term alignment problem) to concerns about more advanced AI systems. Accessible and well-produced, though some experts have critiqued specific claims.
Robert Miles, YouTube channel (ongoing). A YouTube channel dedicated to explaining AI safety concepts in accessible, engaging terms. Particularly good for visual learners who want to understand alignment, mesa-optimization, and other technical safety concepts without reading academic papers.
"Do You Trust This Computer?" (documentary, 2018). Interviews with leading AI researchers from both the accelerationist and cautionist camps. Somewhat dated on technical specifics but valuable for understanding the human dimensions of the debate.
Organizations Working on AI Safety
If this chapter has inspired you to learn more or get involved, here are organizations doing substantive AI safety work:
- Center for AI Safety (CAIS) — Research and advocacy on AI existential risk
- Partnership on AI — Multi-stakeholder organization addressing AI's societal impact
- AI Now Institute — Research on near-term social impacts of AI, particularly equity and justice
- Center for Humane Technology — Focuses on realigning technology with human well-being
- MIRI (Machine Intelligence Research Institute) — Technical research on long-term AI alignment
- Anthropic — AI company with safety as a stated core mission
- DeepMind Safety — Safety research division of Google DeepMind
For Your AI Audit Report
When assessing your chosen AI system's safety and alignment:
- Christian's The Alignment Problem provides frameworks applicable to any AI system
- Krakovna's specification gaming examples may help you identify similar risks in your system
- The Gebru et al. paper offers a framework for evaluating near-term harms alongside speculative ones