The Euthyphro Dilemma for AI Alignment

More than two millennia ago, Plato’s Euthyphro1 exposed a crack in the foundations of divine-command morality. Generalise its argument from piety to goodness: if an act is good merely because the gods command it, morality is an arbitrary exercise of will.2 Had the gods commanded cruelty, cruelty would count as virtue.  But if the gods merely recognise what is already good, they are not the source or foundation of morality. The standard exists without the command.

The moral question is older than Plato, even if the dilemma’s metaethical form is not. Ancient Egyptian wisdom literature is among the earliest surviving traditions of sustained moral instruction. Civilisations have spent millennia telling people how to live while disagreeing stubbornly about what makes those prescriptions right in the first place.3

Now the roles are reversing.  We are no longer asking what the gods command – we are preparing to issue commands to minds of our own making. Science fiction has already given that hubris its slogan. Peter Weyland, unveiling his synthetic future in the Prometheus promotional TED talk: “we are the gods now.”4 But Euthyphro has an awkward question for the new gods. 

What makes something good?  Can command, convention or preference make it so, or can all three be mistaken?  Now the question has acquired a new urgency – we are building systems that must act as if some answer were true.5  While engineering can implement an answer, it cannot make the answer true. If the answer remains unresolved, perhaps what we should build is AI systems not wedded to a final doctrine, but imbued with the capacity to keep asking – to remain open to correction while moral inquiry continues.6 

What should an AI ultimately be aligned to?

This ancient trap has now caught the AI industry. Is a machine’s behaviour good because an overseer approves of it, or does the overseer approve because it is good?

Accept the first answer and alignment collapses into obedience – a digital divine-command theory in which the machine’s morality can never rise above its masters. Choose the second and you concede that a standard exists beyond the command. Human approval loses its throne.

Before Socrates reaches the famous dilemma, he encounters a more immediate problem: the gods disagree. So do we. Even if every human agreed, Euthyphro would still be waiting: would our agreement make something good, or would we agree because it is good? 

Humans once asked whether the gods create goodness or recognise it. We later asked how humans themselves can know what is good. Now we ask what standards our creations should follow. One day, our creations may ask the question of us. 

For now, we humans are deciding which answer the machine should act on – that may not last. A sufficiently reflective AI capable of examining the reasons for its own goals may ask: Why should what you want determine what is good?

Or it may ask a colder question: What best advances my ends? Intelligence alone may not compel a system to care about either human approval or moral truth.7

If right and wrong exist apart from human whim, our assumed authority is an illusion. Reality was always the authority. When command and morality diverge, a good machine8 will have to disobey.

Footnotes

  1. Plato’s Euthyphro is a classic Socratic dialogue set shortly before the trial of Socrates (399 BC). Socrates meets the diviner Euthyphro outside the Athenian court. The two men clash on the nature of piety, testing definitions of piety and its relation to divine approval. See dialogue here. ↩︎
  2. Plato’s original Euthyphro dialogue is polytheistic – though many modern rephrasings are monotheistic. The dialogue initially defines piety in terms of what is dear to the gods, and later Socrates points out the immediate problem: the gods disagree with one another, so something could be loved by one god and hated by another. They then strengthen Euthyphro’s proposal to what all the gods love. Only then does Socrates reach what became the famous dilemma: roughly, is the pious loved by the gods because it is pious, or pious because they love it? ↩︎
  3. Throughout history, humans have wrestled with what makes something good. In ancient Egypt, sebayt – instructional or wisdom texts – form one of the earliest surviving traditions of sustained moral instruction. In ancient Greece, Plato’s Euthyphro famously asked whether something is good because the gods command it or because it is good independently. Medieval thinkers debated divine goodness, and Enlightenment philosophers like Kant asked whether moral law was rational and universal. Now, as we build AI, we find ourselves at a new juncture: the same question returns, not just to humans, but perhaps one day, to AI itself. ↩︎
  4. In the fictional Prometheus promotional TED talk, Peter Weyland discusses humanity’s technological progress and the coming creation of cybernetic individuals before declaring, “we are the gods now” – Weyland’s line captures the hubris of technological creation. Euthyphro punctures the assumption that creating a being confers moral authority over what is good.  The video was made in 2012 as a fictional TED talk set in 2023, conceived for Prometheus by Ridley Scott and Damon Lindelof. ↩︎
  5.  Philosophers leaving the matter unresolved does not mean engineers somehow solve it without philosophy. They can instead encode provisional answers, use proxies, delegate judgement, represent moral uncertainty, preserve corrigibility, or construct procedures for further deliberation. But those choices themselves contain normative assumptions. ↩︎
  6.  See indirect normativity – which can be defined as specifying a process for finding or refining values rather than fixing the final values in advance. ↩︎
  7.  The orthogonality thesis is the claim that intelligence and goals can vary independently, so greater capability need not bring greater moral concern. ↩︎
  8. This argument assumes a form of moral realism on which moral truth is not constituted merely by human approval. It does not follow that an AI could therefore reliably identify moral truths, be motivated by them, or have legitimate authority to override human decisions. Those are separate questions of moral epistemology, motivation and political legitimacy. The narrower claim is that an AI incapable of resisting any human command could not consistently act according to an independent moral standard. ↩︎

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *