Aligned to Flawed Values: David Veldran and Jonathan Leighton on AI and the Risk of Large-Scale Suffering

Most AI safety work asks whether we can build systems that do what we intend. In a new paper, David Veldran and Jonathan Leighton of the Organisation for the Prevention of Intense Suffering (OPIS) press a more uncomfortable question. Suppose we succeed at alignment. What happens if the values we hand the machine are the flawed ones our societies already run on?

Their answer is that large-scale intense suffering has become a moral blind spot in AI governance. The debate tends to frame the worst case as human extinction, while the prospect of a future full of entrenched suffering rarely reaches the agenda. The authors argue that a technically well-aligned AI, trained on the norms and practices we currently tolerate, could preserve and even amplify that suffering, and lock it in for a very long time.

In this conversation we work through the parts of the argument that are hardest and most interesting. We get into the asymmetry between happiness and suffering, and whether extreme suffering carries a priority that no quantity of bliss can offset. We ask whether a non-sentient AI could ever be motivated to care about pain, and use Mary’s Room to test what understanding without experience really amounts to. We discuss the risk of a “benevolent” AI that decides the surest way to end suffering is to end sentience, and what would have to go into a system’s values to rule that out. We also look at whether AI itself might one day suffer, and what a suffering-prevention line in Claude’s constitution might say.

The paper is short and readable, and worth going through alongside the discussion.

Paper: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6949958
OPIS: https://www.preventsuffering.org

Timestamps / Chapters

0:00 Even a perfectly aligned AI could go wrong
4:57 Who gets to decide what counts as “better” values?
11:41 Existence bias: is survival always worth it?
18:04 Can happiness outweigh suffering? The asymmetry
26:35 Why evolution wired us for suffering
30:27 The AI motivation problem: caring vs modelling caring
33:40 Mary’s Room: grasping suffering without feeling it
41:40 Rawls, the veil of ignorance, and the worst off
43:56 Could AI itself suffer? Artificial and digital sentience
51:28 Moral realism and the trouble with “right” and “wrong”
55:04 Avoiding “benevolent” extinction: which values to give AI
57:50 Should we hand ethics over to AI?
59:42 The risk-benefit asymmetry – is a pause justified?
1:06:57 Getting global cooperation on compassionate values
1:13:42 What Claude’s constitution should say about suffering
1:18:59 One thing for policymakers to take away
1:23:56 Closing

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *