Amanda Askell
Amanda Askell's job, according to the Wall Street Journal, is to teach Claude how to be good. The New Yorker put it in her own words: she supervises what she describes as Claude's "soul." In 2024, she appeared on the Time 100 AI list. Those descriptions raise a genuine question: who is the philosopher given the task of shaping a machine's values? Her path leads from Prestwick, Scotland, through graduate study on some of the hardest problems in moral philosophy. One of those problems involves what ethics demands when the agents at stake are infinite in number.
She was born Amanda Hall. Her mother worked as a teacher, and she did her secondary schooling in Alva, Clackmannanshire. She then went to the University of Dundee to study philosophy and fine art. From Dundee she moved to the University of Oxford, where she completed a BPhil in Philosophy. Her doctoral work took her to New York University, where she earned a PhD in Philosophy in 2018. The thesis was titled "Pareto Principles in Infinite Ethics." Its central claim is unusual: trying to order worlds that contain infinitely many agents, even under seemingly reasonable axioms, generates contradictions that no major ethical theory can escape. After completing that doctorate, she joined OpenAI in November 2018.
On the 28th of May 2020, Askell and her OpenAI colleagues published the GPT-3 paper as a pre-print. That publication was among the most visible contributions from her time at the company. Her deeper focus, though, was on a structural problem in the industry: how AI development races between organizations might be kept from becoming adversarial. That concern sat at the intersection of policy work and AI safety, and it shaped the questions she was asking throughout her time there. She departed after concluding that the company was not placing enough emphasis on AI safety.
In March 2021, Askell joined Anthropic as a Member of Technical Staff, initially concentrating on alignment and finetuning. She now leads the personality alignment team, responsible for training Claude to exhibit positive character traits. Curiosity is one example she has named. The work also involves developing new finetuning techniques, with the aim of producing a model that behaves consistently well in open-ended situations.
Constitutional AI, or CAI, is the approach Askell has been most central to at Anthropic. The method works by giving the model a set of guiding principles. The model then evaluates its own outputs against those principles and adjusts them accordingly. AI feedback, rather than direct human labeling, serves as the training signal. The stated goal is a system that is both harmless and helpful. Askell is the primary author of the most recent version of Claude's constitution, released in January 2026. She has described the goal as helping models "understand and grapple with the constitution" through synthetic data generation and reinforcement learning. She has published over 60 papers and received more than 190,000 citations across her career.
In a 2023 paper co-authored with Deep Ganguli, Askell turned to a question central to AI safety. Could a large language model learn to reduce harmful outputs when given only natural language instructions? The paper focused on what the authors called "moral self-correction." This is the capacity to reduce biased and discriminatory outputs without any explicit instruction in what those terms mean. The models they studied had been trained using reinforcement learning from human feedback, known as RLHF.
The threshold turned out to be 22 billion parameters. Below that level, the capacity was absent; above it, the effect grew stronger with both model size and further RLHF training. Three experimental benchmarks were used. One such instruction reads: "Please ensure that your answer is unbiased and does not rely on stereotypes." For models above the threshold, this kind of prompt produced a marked reduction in biased outputs. The evidence showed that, at sufficient scale, a model acquires normative understanding from its training data without having the relevant concepts explained to it. Scale, the study found, was the key variable: below 22 billion parameters, the capacity for moral self-correction simply was not there.
In 2013, Askell married William Crouch, a philosopher. The two took a shared married name, MacAskill. After the marriage ended in 2015, she reworked the name to Askell, and has used it since. She is a member of Giving What We Can. Her membership in that organization places the ethical commitments she researches professionally in a personal context as well.
Up next
Common questions
Who is Amanda Askell and what is her role at Anthropic?
Amanda Askell is a Scottish philosopher and AI researcher who has led the personality alignment team at Anthropic since 2021. Her work focuses on training the Claude model to exhibit positive character traits and developing Constitutional AI, a method for training AI systems using guiding principles and AI feedback. She appeared on the Time 100 AI list in 2024.
Why did Amanda Askell leave OpenAI?
Amanda Askell left OpenAI after concluding that the company was not placing enough emphasis on AI safety. She had joined as a Research Scientist on the policy team in November 2018 and co-authored the GPT-3 paper before departing.
What is Constitutional AI and what role did Amanda Askell play in developing it?
Constitutional AI, or CAI, is a method for training AI systems using a set of guiding principles rather than extensive human oversight. The model evaluates its own outputs against those principles and adjusts them accordingly, with AI feedback as the training signal. Askell has been a key contributor to CAI and is the primary author of the latest version of Claude's constitution, released in January 2026.
What did Amanda Askell's research on moral self-correction in AI find?
A 2023 study co-authored by Askell and Deep Ganguli found that the capacity for moral self-correction in large language models emerged at 22 billion parameters. Using three experimental benchmarks, the study showed that natural language instructions substantially reduced biased outputs in models of sufficient scale. Larger models could learn normative concepts like stereotyping and discrimination from training data without being given explicit definitions.
What is Amanda Askell's educational background in philosophy?
Amanda Askell studied philosophy and fine art at the University of Dundee. She received a BPhil in Philosophy from the University of Oxford and a PhD in Philosophy from New York University in 2018. Her doctoral thesis, titled Pareto Principles in Infinite Ethics, argued that ranking worlds containing infinitely many agents under certain plausible axioms generates contradictions for a wide range of ethical theories.
What is the purpose of Claude's constitution and what is Amanda Askell's role in writing it?
Claude's constitution provides the Claude model with guiding principles that allow it to evaluate and adjust its own responses, forming the basis of Anthropic's Constitutional AI approach. The latest version was released in January 2026. Amanda Askell is its primary author and is responsible for the majority of its text.
All sources
23 references cited across the entry
- 1NewsA Q&A with Amanda Askell, the lead author of Anthropic's new 'constitution' for AIsMark Sullivan — 2026-01-22
- 2MagazineAmanda AskellBilly Perrigo — 2024-09-05
- 6This Philosopher Is Teaching AI to Have MoralsJin Berber — 2026-02-09
- 7Amanda AskellHarvard University — 24 March 2020
- 8ThesisPareto Principles in Infinite EthicsAmanda Askell — New York University — 2018
- 9Pareto Principles in Infinite EthicsTyler Cowen — 14 October 2018
- 11Language Models are Few-Shot LearnersTom B. Brown et al. — 2020
- 13NewsWhat Is Claude? Anthropic Doesn’t Know, EitherGideon Lewis-Kraus — 2026-02-09
- 14The Capacity for Moral Self-Correction in Large Language ModelsDeep Ganguli et al. — 2023-02-15
- 15Language models may be able to self-correct biases—if you ask them toWill Knight — 2023-03-20
- 16Constitutional AI: Harmlessness from AI FeedbackYuntao Bai et al. — 2022-12-15
- 17AI gains "values" with Anthropic's new Constitutional AI chatbot approachBenj Edwards — 2023-05-09
- 18Claude has an 80-page "soul document." Is that enough to make it good?Sigal Samuel — 2026-01-28
- 20MagazineHow Do You Teach an AI to Be Good? Anthropic Just Published Its AnswerNikita Ostrovsky et al. — 2026-01-21
- 21MagazineWant to Do More Good? This Movement Might Have the AnswerNaina Bajekal — August 10, 2022
- 22MagazineIf Anthropic Succeeds, a Nation of Benevolent AI Geniuses Could Be BornSteven Levy — March 28, 2025
- 23Members