Is Superintelligence Necessarily Moral? Discussion with Leonard Dung
Intelligence, Motivation and the Orthogonality Thesis
What I found most interesting in the discussion was the gap between the literal orthogonality thesis and an empirical question we should care about. Bostrom’s thesis says that high intelligence is, in principle, compatible with a wide range of final goals. That does not by itself tell us how values will likely develop in increasingly capable AI systems.
Leonard’s paper gives some empirical support to orthogonality by pointing to reinforcement learning: systems can become increasingly competent while the reward function is specified independently of that competence, and reward hacking shows that greater capability need not produce the designer’s intended objective. But this raises an important further question: is a reward function really analogous to a mature artificial agent’s final goal, or is it better understood as a training pressure from which more complex motivations, representations and values may emerge?
On one view, intelligence is largely an optimisation engine to which many different goals can be attached. On another, intelligence, concepts, reflection, social understanding and values develop together. If the latter is closer to how advanced AI develops, then moral cognition and moral motivation may be less independent than stronger less modal versions of orthogonality suggest.
Moral realism adds another layer. If there are moral facts, a sufficiently capable AI may become increasingly good at discovering them. But moral knowledge is not yet moral motivation. The central alignment problem may therefore be neither simply “teach the AI our values” nor “hope intelligence discovers morality”, but to understand what kinds of architectures allow recognised moral reasons to exert genuine motivational force.
That leaves an empirical question I think deserves much more attention: As AI systems become more capable, are they just getting better at producing wise-sounding outputs, or are they developing increasingly coherent, generalisable and robust moral dispositions? The answer matters enormously for how seriously we should take both pessimistic and optimistic readings of orthogonality.
The empirical question we should care about: Will increasingly capable artificial minds likely develop in ways that keep intelligence and moral motivation apart?
Dr. Leonard Dung is a Postdoctoral Researcher at the Institute of Philosophy II, Ruhr-University Bochum
Topics covered:
- What is the core argument of “Is Superintelligence Necessarily Moral?”
- Can moral knowledge motivate artificial agents?
- How does this argument relate to Bostrom’s orthogonality thesis?
- Does moral realism change the alignment problem?
- What do current AI systems suggest about the relationship between learning, goals, and motivation?
See paper ‘Is superintelligence necessarily moral?’: https://philarchive.org/rec/DUNISN
Abstract : Numerous authors have expressed concern that advanced artificial intelligence (AI) poses an existential risk to humanity. These authors argue that we might build AI which is vastly intellectually superior to humans (a ‘superintelligence’), and which optimizes for goals that strike us as morally bad, or even irrational. Thus this argument assumes that a superintelligence might have morally bad goals. However, according to some views, a superintelligence necessarily has morally adequate goals. This might be the case either because abilities for moral reasoning and intelligence mutually depend on each other, or because moral realism and moral internalism are true. I argue that the former argument misconstrues the view that intelligence and goals are independent, and that the latter argument misunderstands the implications of moral internalism. Moreover, the current state of AI research provides additional reasons to think that a superintelligence could have bad goals.
Also see paper ‘Saving Artificial Minds Understanding and Preventing AI Suffering’: https://www.routledge.com/Saving-Artificial-Minds-Understanding-and-Preventing-AI-Suffering/Dung/p/book/9781041144663