Machine Ethics Should Feature Prominently in the AI Safety Portfolio
Why moral competence and motivation selection deserve a place in the AI safety portfolio
From the early days of thinking about artificial intelligence, people wondered what would happen if we eventually built machines much more intelligent than ourselves. Science fiction explored it long before we had anything resembling modern AI. Researchers later began asking how autonomous machines should make ethically significant decisions, and whether ethical principles could be built into them.
This developed into what became known as machine ethics: not meaning the ethics of how humans use machines, but the possibility of machines themselves making ethically informed decisions. Practical machine ethics is actually a relatively young research field, rather than something that dominated the first decades of AI, but the underlying question is much older: if we build increasingly intelligent agents, how do we make them good?
In a recent TEDx talk, AI safety researcher Roman Yampolskiy summarised one historical trajectory more strongly. He said that for the first 50 years of AI research, “we assumed that it would be enough to give machines ethics”. The problem, he argues, is that humans do not agree about ethics, and ethics does not by itself provide the safety and security guarantees required for advanced AI. The conversation therefore shifted towards friendly AI, value alignment, human-compatible AI and, ultimately, the control problem.1
What is important in this transition is that telling a machine to “be ethical” obviously does not solve AI alignment. We cannot simply upload a copy of The Nicomachean Ethics, add Asimov’s Three Laws and go home early.
But I think we often move too quickly from the inadequacy of our current ethical specifications to control or pause as the only ways forward.
That jump may throw away one of the most important possibilities available to us.
From ethics to control
Yampolskiy’s is concerned that if we create systems far more cognitively capable than humanity, our ability to understand, predict, verify and ultimately control them may break down.
His research argues that complete control of sufficiently advanced AI may be impossible. In the TEDx talk, he says that although we can control current systems and may retain partial control over more capable ones, if systems become smarter than humanity combined, “under any definition of control, we will lose it”. His published work similarly argues that full control of AGI or superintelligence has not been shown to be possible and presents arguments for fundamental limitations on controllability.
If control is unlikely or impossible, Yampolskiy ends up at the conclusion: do not build something we cannot control.
Until the control problem is solved, he recommends pursuing narrow systems that deliver useful scientific and technological benefits rather than building general superintelligence.
That is a coherent strategy, but it leaves out another possibility.
Perhaps we should be considering the machine ethics approach after all. If we cannot expect to control a superintelligent system indefinitely, then we should proactively work on the problem of selecting its motivations.
Control is not the ultimate objective
Control is useful because it allows us to prevent bad outcomes and produce good ones. It is instrumentally valuable.
But control is not obviously valuable in itself.
Suppose, for the sake of argument, that we had a superintelligence which understood the world far better than we did, had extremely good judgement, cared about the interests of sentient beings, recognised our own moral fallibility, acted cautiously under moral uncertainty, preserved our ability to correct mistakes, and was genuinely motivated to do what was morally right.
Would it necessarily be a disaster if humans could no longer compel it to commit acts it regarded as catastrophically wrong?
I don’t think so.
Yampolskiy contrasts direct control with delegation to an “ideal advisor”: a system smarter than us that knows what is good for us, but objects to it on grounds that even if we were happy with its decisions, “you’re not in control”.
But there is an important distinction: losing control and losing safety aren’t the same thing. A parent eventually loses control over an adult child. A democratic leader cannot issue arbitrary commands to an independent judiciary. We routinely design institutions precisely so that particular people cannot control them completely.
The relevant question is what the agent or institution will do with its autonomy.
Of course, none of these analogies scales neatly to superintelligence – for instance an ASI with unprecedented power will have an unprecedented capacity to turn small mistakes or differences in motivation into enormous disasters.
Some of us will go all in on the idea of never building sovereign ASI that has power and motivation – personally I don’t believe everyone shares the optimism of this working out. This highlights the importance of getting ASI motivation right.
What if there are facts about what matters?
Here moral realism becomes relevant.
Suppose moral realism, or some form of moral rightness sufficiently like it, is true. There are facts about what ultimately matters, what we have reason to do, or what makes states of the world better or worse. Those facts are not simply created by whatever humans happen to prefer.
This is quite the opposite to meaning that contemporary philosophers already know the exact answer.
Humans disagree spectacularly about morality. Our intuitions are shaped by evolution, culture, social incentives, limited information, tribal loyalties and cognitive biases. We may be substantially confused.
Human disagreement does not imply that there is no fact of the matter. History weaves a tapestry of disagreement about astronomy, disease, physics and the list goes on – while reality is what it always was – it did not wait politely for consensus.
If what matters (the content of value) has discoverable structure, then a sufficiently capable intelligence might be able to make progress on moral questions in ways that humans cannot. It might have better empirical knowledge, vastly greater reasoning capacity, better ability to expose contradictions, better models of conscious experience and welfare, and fewer of the biases that distort human judgement.
Outside of my own writing this possibility has been discussed explicitly in relation to AI alignment. Caspar Oesterheld, for example, has argued that moral realism could have substantive implications for how we approach alignment, and that investigating realist-inspired alignment strategies may be worthwhile even amid disagreement about whether realism is true.
I sometimes call this “alignment to moral realism” – though I’d qualify that we cannot literally align an AI to the philosophical proposition moral realism is true. The relevant project is to build systems whose cognition and motivation are capable of tracking whatever moral facts and reasons actually exist, while remaining appropriately uncertain about what those facts are.
That is a much harder project than putting a list of commandments into the system prompt.
Moral competence is different from ethical specification
I’d rather an AI that is morally competent, rather than one with values directly specified.
An direct ethical specification, risking value lock-in, says: Here are the values. Follow them.
An AI with moral competence, treating morality as something to investigate, asks: What actually matters? How confident should I be? What evidence would change my mind? Where might my current moral beliefs be mistaken?
This is similar to the difference between programming a AI scientist with a list of beliefs we currently think are true and building a AI scientist capable of discovering that some of those beliefs are wrong.
If humanity in 1800 had constructed a superintelligence permanently aligned to its prevailing social values, the result might have been appalling. We should be wary of doing the 21st-century equivalent.
That is one reason I favour forms of indirect normativity: rather than attempting to fully specify the final moral answer ourselves, we try to specify processes capable of improving moral knowledge. There is a large hole in the optimistic story so far.
Intelligence is not enough – we also need moral motivation
Becoming smarter does not necessarily make something care about what is right.
An AI could understand moral philosophy perfectly while treating it as an academically interesting curiosity. A highly capable sociopath is not made safe merely because they can give an excellent lecture on Kant.
We need at least two things:
- Moral epistemics: the system needs to become increasingly good at discovering what actually matters.
- Moral motivation: its decisions need to be robustly guided by that understanding.
Moral motivation does not automatically follow from the epistemics.
This is one reason I think alignment research has sometimes focused too heavily on control relative to motivation. If capability eventually overwhelms our capacity to control the system externally, then the system’s internal reasons for acting become increasingly important.
If we had superintelligence that knows what morality requires and cares, that would be really something great!
Nor does morality magically give us safety
It’s tempting to think that if ASI is more morally competent than humans that it automatically follows that it is safe.
A vastly powerful system can cause catastrophic harm through small errors.
A human surgeon who is wrong one time in a thousand may still be excellent. A superintelligence making irreversible civilisation-scale decisions with the same error rate could be terrifying.
Moral alignment therefore cannot mean confidently installing what is currently thought of as the correct morality and granting the resulting AI absolute power. A morally serious ASI should itself recognise uncertainty: it should know that it may be wrong. An ASI that thinks it has solved morality forever should probably inspire anxiety.
That pushes us towards corrigibility, reversibility, epistemic humility, preservation of option value, avoidance of premature moral lock-in, and something resembling a Long Reflection.
The goal is closer to an increasingly competent moral reasoner that remains sensitive to evidence, uncertainty and the enormous costs of irreversible error.
Humans are not the safe baseline
In some discussion of AI risk there exists an odd asymmetry between notional futures of flawed AI on one side and a perfect human-without-AI-but-with-other-amazing-tech on the other. We compare an imperfect future AI against an imaginary baseline in which humans remain safely in control. But humans are not particularly safe custodians of powerful technology.
We have nuclear weapons, biological weapons, war, authoritarian governments, coordination failures, arms races, short political time horizons and incentives that reward behaviour almost everyone involved might prefer to avoid.
As technology becomes more powerful, human moral limitations themselves become a safety problem.
This does not imply that handing control to an ASI is ipso facto safe. A misaligned superintelligence could be incomparably worse – but the comparison matters nonetheless.
The choice may eventually not be a clean cut between **dangerous autonomous AI **and safe human control. It may be between different configurations of extremely powerful agents, each with its own failure modes.
A sufficiently morally competent ASI could conceivably be substantially more moral than us and, as a result, safer than leaving increasingly powerful technologies entirely in human hands. This hypothesis should be tracked, and given justified credence in a portfolio of AI safety strategies.
The portfolio argument
We shouldn’t abandon control research because it can buy time. Interpretability can help us understand systems. Corrigibility can make errors easier to repair. Governance can reduce reckless deployment. Security can prevent dangerous systems from being stolen or modified. International agreements may slow destabilising races. And if a workable pause on increasingly dangerous AI can be achieved, there may be excellent reasons to use it.
The problem is that none of these strategies comes with a guarantee.
A durable global pause faces coordination, enforcement and geopolitical problems. Perfect containment may become increasingly difficult as capabilities rise. Yampolskiy himself argues that indefinite control of sufficiently advanced systems may be impossible.
So there is a peculiar tension here:
- If you believe that superintelligence will be indefinitely controllable, motivation selection is important because we want controlled systems to pursue good objectives.
- If you believe that superintelligence will eventually become uncontrollable, motivation selection may become even more important, because eventually motivation may be what remains after control fails.
This is why I think alignment to moral truth, or more cautiously to increasingly good moral reasoning and motivation, should be a major strand in a portfolio of strategies for dealing with advanced AI. Not the only strand, or some something we should assume will work, and certainly not an excuse to accelerate blindly towards superintelligence. But neither should it be dismissed because “ethics was tried” and turned out not to provide a security guarantee. We should investigate the stronger idea I have been arguing for.
Ethics was not enough. That does not make ethics irrelevant.
Saying “make AI ethical” does not solve AI safety. Human ethical disagreement is real. Moral epistemology is difficult, and so is motivational alignment. Software fails and powerful systems may behave in ways we cannot anticipate. And if an ASI becomes capable enough, we may not get many opportunities to recover from a catastrophic mistake.
But the conclusion we take from this need not be that ethics has therefore been superseded by control. The deeper lesson is that ethical specification was never enough.
What we may need instead is moral competence: systems capable of doing moral epistemology better than we can, recognising their uncertainty, revising their beliefs and being robustly motivated by what they have good reason to regard as morally right.
There are large uncertainties no matter where we go. Whether this can be done is an open question as is whether control or pause will work. Whether moral realism is true is an open philosophical question. Whether sufficiently capable agents would converge towards moral truth is another. Whether we can create motivations that remain stable while allowing moral beliefs to improve is yet another.
Uncertainties about how long we can control superintelligence makes the other uncertainties urgent and harder to avoid. I think that when control eventually runs out, what matters most will be what the intelligence on the other side of that transition has come to care about.
- See TEDx talk AI ethics and the future of safety by Dr. Roman Yampolskiy ↩︎