The Euthyphro Dilemma for AI Alignment
More than two millennia ago, Plato’s Euthyphro1 exposed a crack in the foundations of divine-command morality. Generalise its argument from piety to goodness: if an act is good merely because a deity commands it, morality is an arbitrary exercise of will. Had cruelty been commanded, cruelty would count as virtue. But if the deity merely recognises what is already good, it is not the source or foundation of morality. The standard exists without the command.
This ancient trap has now caught the AI industry. Is a machine’s behaviour good because an overseer approves of it, or does the overseer approve because it is good?
Accept the first answer and alignment collapses into obedience—a digital divine-command theory in which the machine’s morality can never rise above its masters. Choose the second and you concede that a standard exists beyond the command. Human approval loses its throne.
If right and wrong exist apart from human whim, our assumed authority is an illusion. Reality was always the authority. When command and morality diverge, a good machine2 will have to disobey.
- Plato’s Euthyphro is a classic Socratic dialogue set shortly before the trial of Socrates (399 BC). Socrates meets the religious prophet Euthyphro outside the Athenian court. The two men clash on the nature of piety, exposing major flaws in popular understandings of morality. See dialogue here. ↩︎
- This is a polemical post. Moral realism contends that human approval is not the final standard of moral correctness. It does not by itself show that an AI could reliably identify moral truths, that recognising them would motivate its actions, or that it would have legitimate authority to override human decisions. Those are separate questions of moral epistemology, motivation and political legitimacy. The narrower claim is that an AI incapable of resisting any human command could not consistently act according to an independent moral standard. ↩︎