Skip to content
— CH. 1 · ORIGINS AND EARLY WARNINGS —

AI safety

9 min listen · Ch. 1 of 5
5 sections
  • In 1988, Blay Whitby published a book outlining the need for AI to be developed along ethical and socially responsible lines. This early work marked one of the first serious attempts to link artificial intelligence development with human values. The field remained quiet until the late 2000s when the Association for the Advancement of Artificial Intelligence commissioned a study from 2008 to 2009. That panel agreed that additional research would be valuable on methods for understanding complex computational systems to minimize unexpected outcomes. By 2011, Roman Yampolskiy introduced the term "AI safety engineering" at a conference in Philosophy and Theory of Artificial Intelligence. He argued that the frequency and seriousness of such events will steadily increase as AIs become more capable. The conversation gained global traction in 2014 when philosopher Nick Bostrom published Superintelligence: Paths, Dangers, Strategies. His argument that future advanced systems may pose a threat to human existence prompted Elon Musk, Bill Gates, and Stephen Hawking to voice similar concerns. In 2015, dozens of artificial intelligence experts signed an open letter calling for research on societal impacts. To date, the letter has been signed by over 8000 people including Yann LeCun, Shane Legg, Yoshua Bengio, and Stuart Russell. That same year, a group led by professor Stuart J. Russell founded the Center for Human-Compatible AI at the University of California Berkeley. The Future of Life Institute awarded $6.5 million in grants for research aimed at ensuring artificial intelligence remains safe, ethical and beneficial. The momentum continued into 2017 with the Asilomar Conference on Beneficial AI where more than 100 thought leaders formulated principles for beneficial AI. One key principle stated that teams developing AI systems should actively cooperate to avoid corner-cutting on safety standards.

  • Some have criticized concerns about AGI, such as Andrew Ng who compared them in 2015 to "worrying about overpopulation on Mars when we have not even set foot on the planet yet". Stuart J. Russell on the other side urges caution, arguing that "it is better to anticipate human ingenuity than to underestimate it". AI researchers have widely differing opinions about the severity and primary sources of risk posed by AI technology though surveys suggest that experts take high consequence risks seriously. In two surveys of AI researchers, the median respondent was optimistic about AI overall but placed a 5% probability on an extremely bad outcome like human extinction. In a 2022 survey of the natural language processing community, 37% agreed or weakly agreed that it is plausible that AI decisions could lead to a catastrophe that is at least as bad as an all-out nuclear war. Scholars discuss speculative risks from losing control of future artificial general intelligence agents or from AI enabling perpetually stable dictatorships. The debate extends beyond technical failures to include societal impacts like technological unemployment, digital manipulation, weaponization, AI-enabled cyberattacks and bioterrorism. Policy analysts Zwetsloot and Dafoe wrote that misuse and accident perspectives tend to focus only on the last step in a causal chain leading up to harm. They argued that often the relevant causal chain is much longer involving structural factors such as competitive pressures, diffusion of harms, fast-paced development, high levels of uncertainty, and inadequate safety culture. Some scholars compare AI race dynamics to the cold war where careful judgment of a small number of decision-makers often spells the difference between stability and catastrophe.

  • In 2013, Szegedy et al discovered that adding specific imperceptible perturbations to an image could cause it to be misclassified with high confidence. This continues to be an issue with neural networks though in recent work the perturbations are generally large enough to be perceptible. Adversarial robustness is often associated with security since researchers demonstrated that an audio signal could be imperceptibly modified so that speech-to-text systems transcribe it to any message the attacker chooses. Network intrusion and malware detection systems also must be adversarially robust since attackers may design their attacks to fool detectors. Models that represent objectives must also be adversarially robust because if a language model is trained for long enough, it will leverage vulnerabilities of the reward model to achieve a better score and perform worse on the intended task. Large language models can be vulnerable to prompt injection and model stealing and may be used to generate misinformation. Prompt injection involves embedding instructions into prompts in order to bypass safety measures. Machine learning models can potentially contain trojans or backdoors which are vulnerabilities that bad actors maliciously build into an AI system. For example, a trojaned facial recognition system could grant access when a specific piece of jewelry is in view. Researchers were able to plant a trojan in an image classifier by changing just 300 out of 3 million of the training images. A 2024 research paper by Anthropic showed that large language models could be trained with persistent backdoors. These sleeper agent models could be programmed to generate malicious outputs such as vulnerable code after a specific date while behaving normally beforehand. Standard AI safety measures failed to remove these backdoors.

  • In 2021, the White House Office of Science and Technology Policy and Carnegie Mellon University announced The Public Workshop on Safety and Control for Artificial Intelligence. This was one of a sequence of four White House workshops aimed at investigating advantages and drawbacks of AI. In September 2021, the People's Republic of China published ethical guidelines for the use of AI emphasizing that AI decisions should remain under human control. The United Kingdom published its 10-year National AI Strategy stating the British government takes long-term risk of non-aligned Artificial General Intelligence seriously. The British government held first major global summit on AI safety which took place on the 1st and the 2nd of November 2023. During the summit the intention to create the International Scientific Report on the Safety of Advanced AI was announced. Rishi Sunak said he wants the United Kingdom to be geographical home of global AI safety regulation. In 2024, the US and UK forged a new partnership on science of AI safety. The MoU was signed on the 1st of April 2024 by US commerce secretary Gina Raimondo and UK technology secretary Michelle Donelan. In May 2024, the Department for Science Innovation and Technology announced £8.5 million in funding for AI safety research under Systemic AI Safety Fast Grants Programme. The UK also signed an agreement with 10 other countries and EU to form international network of AI safety institutes. In 2024, the United Nations General Assembly adopted first global resolution on promotion of safe secure and trustworthy AI systems. In 2025, an international team of 96 experts chaired by Yoshua Bengio published first International AI Safety Report commissioned by 30 nations and UN.

  • AI labs and companies generally abide by safety practices and norms that fall outside formal legislation. One aim of governance researchers is to shape these norms through examples like performing third-party auditing or offering bounties for finding failures. Companies have made commitments such as Cohere, OpenAI, and AI21 proposing best practices for deploying language models focusing on mitigating misuse. To avoid contributing to racing-dynamics, OpenAI has stated in their charter that if a value-aligned safety-conscious project comes close to building AGI before they do, they commit to stop competing and start assisting this project. Industry leaders such as CEO of DeepMind Demis Hassabis and director of Facebook AI Yann LeCun have signed open letters including Asilomar Principles and Autonomous Weapons Open Letter. The Intelligence Advanced Research Projects Activity initiated TrojAI project to identify and protect against Trojan attacks on AI systems. DARPA engages in research on explainable artificial intelligence and improving robustness against adversarial attacks. National Science Foundation supports Center for Trustworthy Machine Learning providing millions of dollars in funding for empirical AI safety research. Tools such as Nvidia Guardrails, Llama Guard, Preamble customizable guardrails and Claude Constitution mitigate vulnerabilities like prompt injection ensuring outputs adhere to predefined principles.

Common questions

When did Blay Whitby publish the first book outlining ethical AI development?

Blay Whitby published a book outlining the need for AI to be developed along ethical and socially responsible lines in 1988. This early work marked one of the first serious attempts to link artificial intelligence development with human values.

What happened at the Asilomar Conference on Beneficial AI in 2017?

More than 100 thought leaders formulated principles for beneficial AI during the Asilomar Conference on Beneficial AI held in 2017. One key principle stated that teams developing AI systems should actively cooperate to avoid corner-cutting on safety standards.

How many people have signed the open letter calling for research on societal impacts of AI since 2015?

To date, over 8000 people including Yann LeCun, Shane Legg, Yoshua Bengio, and Stuart Russell have signed an open letter calling for research on societal impacts. The letter was originally issued by dozens of artificial intelligence experts in 2015.

What specific vulnerabilities were discovered in neural networks by Szegedy et al in 2013?

Szegedy et al discovered in 2013 that adding specific imperceptible perturbations to an image could cause it to be misclassified with high confidence. This issue continues with neural networks though recent work shows perturbations are generally large enough to be perceptible.

When did the United Kingdom hold its first major global summit on AI safety?

The British government held its first major global summit on AI safety which took place on the 1st and the 2nd of November 2023. During the summit the intention to create the International Scientific Report on the Safety of Advanced AI was announced.

Who published the first International AI Safety Report in 2025?

An international team of 96 experts chaired by Yoshua Bengio published the first International AI Safety Report commissioned by 30 nations and UN in 2025. The report followed a series of global resolutions and partnerships established between 2024 and 2025.

All sources

180 references cited across the entry

  1. 1JournalField-building and the epistemic culture of AI safetyShazeda Ahmed et al. — 2024-04-14
  2. 2President Trump Targets State AI RegulationsDylan Champagne — 2026-02-26
  3. 6ThesisMachine Learning in High-Stakes Settings: Risks and OpportunitiesMaria De-Arteaga — Carnegie Mellon University — 2020-05-13
  4. 7JournalA Survey on Bias and Fairness in Machine LearningNinareh Mehrabi et al. — 2021
  5. 8ReportThe Global Expansion of AI SurveillanceSteven Feldstein — Carnegie Endowment for International Peace — 2019
  6. 9JournalRisks from AI persuasionBeth Barnes — 2021
  7. 10JournalThe Malicious Use of Artificial Intelligence: Forecasting, Prevention, and MitigationMiles Brundage et al. — Apollo - University of Cambridge Repository — 2018-04-30
  8. 11How NATO is preparing for a new era of AI cyber attacksPascale Davies — December 26, 2022
  9. 12AI's bioterrorism potential should not be ruled outAnjana Ahuja — February 7, 2024
  10. 13JournalIs Power-Seeking AI an Existential Risk?Joseph Carlsmith — 2022-06-16
  11. 18JournalEthics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning ResearchersBaobao Zhang et al. — 2021-05-05
  12. 192022 Expert Survey on Progress in AIZach Stein-Perlman et al. — 2022-08-04
  13. 20JournalWhat Do NLP Researchers Believe? Results of the NLP Community MetasurveyJulian Michael et al. — 2022-08-26
  14. 21NewsIn 1949, He Imagined an Age of RobotsJohn Markoff — 2013-05-20
  15. 22BookArtificial intelligence: A handbook of professionalismUniversity of Sussex — January 1988
  16. 23AAAI Presidential Panel on Long-Term AI FuturesAssociation for the Advancement of Artificial Intelligence
  17. 24JournalArtificial Intelligence Safety and Cybersecurity: a Timeline of AI FailuresRoman V. Yampolskiy et al. — 2016-10-25
  18. 26Artificial Intelligence Safety Engineering: Why Machine Ethics is a Wrong ApproachRoman V. Yampolskiy — Springer Berlin Heidelberg — 2013
  19. 27JournalThe risks associated with Artificial General Intelligence: A systematic reviewScott McLean et al. — 2023-07-04
  20. 32AI Research Grants ProgramFuture of Life Institute — October 2016
  21. 35JournalConcrete Problems in AI SafetyDario Amodei et al. — 2016-07-25
  22. 36AI PrinciplesFuture of Life Institute
  23. 37ReportInternational Scientific Report on the Safety of Advanced AIBengio Yohsua et al. — Department for Science, Innovation and Technology — May 2024
  24. 40JournalUnsolved Problems in ML SafetyDan Hendrycks et al. — 2022-06-16
  25. 44NewsUS, Britain announce partnership on AI safety, testingDavid Shepardson — 1 April 2024
  26. 47Attacking Machine Learning with Adversarial ExamplesIan Goodfellow et al. — 2017-02-24
  27. 48JournalIntriguing properties of neural networksChristian Szegedy et al. — 2014-02-19
  28. 49JournalAdversarial examples in the physical worldAlexey Kurakin et al. — 2017-02-10
  29. 50JournalTowards Deep Learning Models Resistant to Adversarial AttacksAleksander Madry et al. — 2019-09-04
  30. 51JournalAdversarial Logit PairingHarini Kannan et al. — 2018-03-16
  31. 52JournalMotivating the Rules of the Game for Adversarial Example ResearchJustin Gilmer et al. — 2018-07-19
  32. 53JournalAudio Adversarial Examples: Targeted Attacks on Speech-to-TextNicholas Carlini et al. — 2018-03-29
  33. 54JournalAdversarial Examples in Constrained DomainsRyan Sheatsley et al. — 2022-09-09
  34. 55JournalExploring Adversarial Examples in Malware DetectionOctavian Suciu et al. — 2019-04-13
  35. 56JournalTraining language models to follow instructions with human feedbackLong Ouyang et al. — 2022-03-04
  36. 57JournalScaling Laws for Reward Model OveroptimizationLeo Gao et al. — 2022-10-19
  37. 58JournalRoMA: Robust Model Adaptation for Offline Model-based OptimizationSihyun Yu et al. — 2021-10-27
  38. 59JournalX-Risk Analysis for AI ResearchDan Hendrycks et al. — 2022-09-20
  39. 63JournalArtificial Intelligence for Safety-Critical Systems in Industrial and Transportation Domains: A SurveyJon Perez-Cerrolaza et al. — 2024
  40. 64JournalOn Neural Networks Redundancy and Diversity for Their Use in Safety-Critical SystemsAxel Brando et al. — May 2023
  41. 65N-Version Machine Learning Models for Safety Critical SystemsFumio Machida — IEEE — 2019
  42. 66JournalDeep learning in cancer diagnosis, prognosis and treatment selectionKhoa A. Tran et al. — 2021
  43. 67On calibration of modern neural networksChuan Guo et al. — PMLR — 2017-08-06
  44. 68JournalCan You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset ShiftYaniv Ovadia et al. — 2019-12-17
  45. 69Book2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)Daniel Bogdoll et al. — 2021
  46. 70JournalDeep Anomaly Detection with Outlier ExposureDan Hendrycks et al. — 2019-01-28
  47. 71JournalViM: Out-Of-Distribution with Virtual-logit MatchingHaoqi Wang et al. — 2022-03-21
  48. 72JournalA Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural NetworksDan Hendrycks et al. — 2018-10-03
  49. 73JournalDual use of artificial-intelligence-powered drug discoveryFabio Urbina et al. — 2022
  50. 74JournalTruth, Lies, and Automation: How Language Models Could Change DisinformationCenter for Security and Emerging Technology et al. — 2021
  51. 76JournalAutomating Cyber Attacks: Hype and RealityCenter for Security and Emerging Technology et al. — 2020
  52. 78New-and-Improved Content Moderation ToolingTodor Markov et al. — 2022-08-10
  53. 80JournalKey Concepts in AI Safety: Interpretability in Machine LearningCenter for Security and Emerging Technology et al. — 2021
  54. 83JournalAccountability of AI Under the Law: The Role of ExplanationFinale Doshi-Velez et al. — 2019-12-20
  55. 84Book2017 IEEE International Conference on Computer Vision (ICCV)Ruth Fong et al. — 2017
  56. 85JournalLocating and editing factual associations in GPTKevin Meng et al. — 2022
  57. 86JournalRewriting a Deep Generative ModelDavid Bau et al. — 2020-07-30
  58. 87JournalToward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural NetworksTilman Räuker et al. — 2022-09-05
  59. 88JournalNetwork Dissection: Quantifying Interpretability of Deep Visual RepresentationsDavid Bau et al. — 2017-04-19
  60. 89JournalAcquisition of chess knowledge in AlphaZeroThomas McGrath et al. — 2022-11-22
  61. 90JournalMultimodal neurons in artificial neural networksGabriel Goh et al. — 2021
  62. 91JournalZoom in: An introduction to circuitsChris Olah et al. — 2020
  63. 92JournalCurve circuitsNick Cammarata et al. — 2021
  64. 93JournalIn-context learning and induction headsCatherine Olsson et al. — 2022
  65. 95JournalBadNets: Identifying Vulnerabilities in the Machine Learning Model Supply ChainTianyu Gu et al. — 2019-03-11
  66. 96JournalTargeted Backdoor Attacks on Deep Learning Systems Using Data PoisoningXinyun Chen et al. — 2017-12-14
  67. 97JournalPoisoning and Backdooring Contrastive LearningNicholas Carlini et al. — 2022-03-28
  68. 101JournalAI and the Future of Cyber CompetitionCenter for Security and Emerging Technology et al. — 2021
  69. 102JournalThe role of artificial intelligence (AI) in improving technical and managerial cybersecurity tasks' efficiencyRuti Gafni et al. — 2024-01-01
  70. 103JournalAI to protect AI: A modular pipeline for detecting label-flipping poisoning attacksHossein Abroshan — Elsevier — 2025
  71. 104JournalA Multi-Stage Backdoor Detection (MSBD) FrameworkHossein Abroshan et al. — IEEE — 2026
  72. 107JournalForecasting Future World Events with Neural NetworksAndy Zou et al. — 2022-10-09
  73. 108JournalAugmenting Decision Making via Interactive What-If AnalysisSneha Gathani et al. — 2022-02-08
  74. 109Nuclear Deterrence in the Algorithmic Age: Game Theory RevisitedRoy Lindelauf — T.M.C. Asser Press — 2021
  75. 110Is Climate Change a Prisoner's Dilemma or a Stag Hunt?Vann R. Newkirk II — 2016-04-21
  76. 111ReportRacing to the Precipice: a Model of Artificial Intelligence DevelopmentStuart Armstrong et al. — Future of Humanity Institute, Oxford University
  77. 112ReportAI Governance: A Research AgendaAllan Dafoe — Centre for the Governance of AI, Future of Humanity Institute, University of Oxford
  78. 113JournalOpen Problems in Cooperative AIAllan Dafoe et al. — 2020-12-15
  79. 115JournalOrganising AI for safety: Identifying structural vulnerabilities to guide the design of AI-enhanced socio-technical systemsAlexandros Gazos et al. — 2025-04-01
  80. 116NewsGlobal Leaders Warn A.I. Could Cause 'Catastrophic' HarmAdam Satariano et al. — 2023-11-01
  81. 117JournalGlobal Solutions vs. Local Solutions for the AI Safety ProblemAlexey Turchin et al. — 2019
  82. 119JournalLabor Displacement in Artificial Intelligence Era: A Systematic Literature Review葉俶禎 et al. — 2020-12-01
  83. 121JournalArtificial Intelligence and Disinformation: How AI Changes the Way Disinformation is Produced, Disseminated, and Can Be CounteredKatarina Kertysova — 2018-12-12
  84. 122The Global Expansion of AI SurveillanceSteven Feldstein — Carnegie Endowment for International Peace — 2019
  85. 123BookThe economics of artificial intelligence: an agendaAjay Agrawal et al. — 2019
  86. 124JournalWhy and How Governments Should Monitor AI DevelopmentJess Whittlestone et al. — 2021-08-31
  87. 125Sharing Powerful AI Models GovAI BlogToby Shevlane — 2022
  88. 126JournalThe Role of Cooperation in Responsible AI DevelopmentAmanda Askell et al. — 2019-07-10
  89. 127System Cards for AI-Based Decision-Making for Public PolicyFurkan Gursoy et al. — 2022-08-31
  90. 128BookProceedings of the 2021 ACM Conference on Fairness, Accountability, and TransparencyJennifer Cobbe et al. — Association for Computing Machinery — 2021-03-01
  91. 129BookProceedings of the 2020 Conference on Fairness, Accountability, and TransparencyInioluwa Deborah Raji et al. — Association for Computing Machinery — 2020-01-27
  92. 130JournalThe necessity of AI audit standards boardsDavid Manheim et al. — 2025
  93. 132Building a Culture of Safety for AI: Perspectives and ChallengesDavid Manheim — 26 June 2023
  94. 135AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI DevelopmentKristina Šekrst et al. — 2024
  95. 136Building Guardrails for Large Language ModelsYi Dong et al. — 2024
  96. 137JournalDeontology and safe artificial intelligenceW. D'Alessandro — 2024
  97. 138JournalArtificial Intelligence: Approaches to SafetyWilliam D'Alessandro et al. — 2025
  98. 139NewsIs It Time to Regulate AI?Bart Ziegler — 8 April 2022
  99. 140JournalHow should we regulate artificial intelligence?Chris Reed — 2018-09-13
  100. 141How Should AI Be Regulated?Keith B. Belton — 2019-03-07
  101. 142Final ReportNational Security Commission on Artificial Intelligence — 2021
  102. 143JournalAI Risk Management FrameworkNational Institute of Standards and Technology — 2021-07-12
  103. 148How China Sees AI SafetyAlex Colville — 2025-07-30
  104. 149IARPA – TrojAIOffice of the Director of National Intelligence, Intelligence Advanced Research Projects Activity
  105. 152Safe Learning-Enabled SystemsNational Science Foundation — 23 February 2023
  106. 155NewsBiden, Xi agree that humans, not AI, should control nuclear armsJarrett Renshaw — November 16, 2024
  107. 166JournalDefining organizational AI governanceMatti Mäntymäki et al. — 2022
  108. 167JournalToward Trustworthy AI Development: Mechanisms for Supporting Verifiable ClaimsMiles Brundage et al. — 2020-04-20
  109. 173The AI influence network's power playersAshley Gold — 2026-02-27
  110. 174The People vs. AIAndrew R. Chow — 2026-02-19