Normative Convergence
If there are normative reasons that exist, then normative convergence is highly plausible for advanced agents.1
By normative convergence I mean the tendency of sufficiently capable, reflectively open agents to converge in their judgements about what there is reason to value as they become better at discovering normative facts – analogous to instrumental convergence, except that the proposed convergence occurs in normative judgement rather than in the means useful for pursuing antecedently fixed goals.
This may produce motivational convergence.
Normative convergence: The hypothesis that sufficiently capable agents that can reflect effectively on normative questions, discover normative facts, and remain motivationally responsive to normative reasons will tend to converge on at least some of the same normative judgements, values, or ends, despite differences in their initial goals or origins.
The phrase itself is not new, I am adapting it as a specific AI-alignment hypothesis. “Normative convergence” and closely related “convergence theses” already occur in metaethics, including discussions of whether ideally rational and informed agents would converge on normative judgements. Parfit and Michael Smith are particularly relevant predecessors.
There is a parallel to between instrumental convergence:
Instrumental convergence: where different final-goals convergence on similar instrumental means.
Normative convergence: different initial goals or values → convergence through reflection on some of the same normative reasons, values or ends
And there is also motivational convergence
Motivatkonal convergence: agent with a meta-level concern about whether its motivations are good -> searches motivational design space -> prefers architectures that make it more responsive to reasons.
The contrast is potentially important for AI alignment.
Instrumental convergence says that apparently harmless terminal goals can produce dangerous strategies because power, resources, self-preservation and so forth are useful for many ends. Bostrom’s formulation is explicitly about intermediary goals that help agents realise otherwise divergent final goals.
Normative convergence would point in the other direction: sufficiently sophisticated normative cognition might constrain the space of final ends themselves.
Sarah McGrath’s 2010 paper, “Moral Realism Without Convergence“2, argues that argues that moral realism need not imply that fully rational, fully informed people converge morally. But to my mind McGrath is too pessimistic about rational tools – historical trends show a gradual convergence on key moral issues (such as the rejection of slavery or institutionalised subjugation), suggesting that rational debate often does slowly reveal objective moral truths over time. I agree with realist philosophers like Michael Smith in that if moral realism is true, then fully rational, fully informed investigators must eventually converge on the same moral truths – and even anti-realists like John Mackie agree that this would be the case if moral realism were true.
Three positions need separating:
Arguably only the first is supplied by realism. The second is a normative-epistemology problem. The third is the normative-motivation problem (though I’m a realist about this).
Epistemic normative convergence would mean increasingly capable normative reasoners tend towards agreement about what normative facts/reasons actually obtain.
Motivational normative convergence would mean their motivations or terminal commitments increasingly come to reflect those reasons.
The first absolutely does not guarantee the second – the problem is far more interesting than the old “intelligence leads to benevolence” argument. An ASI might converge nearly perfectly on the proposition that suffering gives agents strong reasons to prevent it while continuing to give that fact almost no motivational weight.
The stronger AI-alignment thesis is conditional:
Normative Convergence Thesis (NCT): Across a sufficiently broad range of initial values and cognitive architectures, agents capable of reliable normative inquiry and reflective revision will tend to converge in their normative judgements insofar as they successfully track stance-independent normative facts, and agents whose motivations are also responsive to those facts will tend towards corresponding convergence in their endorsed ends.
There is the concern that there may not be one singular optimum. Moral realism doesn’t strictly imply this. Reality might contain plural values, incomparable reasons, agent-relative reasons, genuine underdetermination, or many permissible Pareto-equivalent arrangements. Normative convergence could therefore produce a narrowing basin rather than a single point. Under the narrowing basin assumption treat \(\mathcal{N}\) as a constrained region of normatively defensible values rather than one unique utility function, else treat \(\mathcal{N}\) as a single point:
This challenges stronger readings of orthogonality. Bostrom’s orthogonality thesis is a modal thesis which is explicitly about what combinations of intelligence and final goals are possible, not a probability distribution over what reflectively sophisticated agents will end up valuing.3 So normative convergence needn’t deny orthogonality:
Orthogonality: extraordinarily intelligent agents can have almost arbitrary ends.
Normative convergence: under additional conditions – normative cognition, reliable normative epistemology, reflective openness and reasons-responsive motivation – their ends may nevertheless be systematically attracted towards a much smaller region or a single point.
Logical possibility gives almost no information about the distribution of values produced by particular developmental processes.
And this connects to differential normative development. Different parts of an ASI’s normative understanding may converge at different rates. It could learn very quickly that sentient suffering matters while remaining confused for much longer about population ethics, aggregation, moral status, distributive justice or astronomical-scale trade-offs. So normative convergence need not be synchronous.
Normative convergence sits within a wider group of convergences:
- Instrumental convergence: Different ends -> similar useful means.
- Epistemic convergence: Different beliefs -> increasingly similar beliefs about reality.
- Normative convergence: Different starting values -> increasingly similar judgements about what there is reason to value.
- Motivational convergence: Different starting motivations -> increasingly similar motivation by those normative reasons.
Footnotes
- This could be true under robust moral realism (mind-independent, irreducible, necessary normative reasons), naturalism (simply complex natural facts about the physical universe – specifically facts regarding psychology, sociology, cooperation, and the flourishing of conscious systems), or realist constructivism (necessary outputs of rational agency itself – though I’d argue that rational agency is a complex natural fact about the universe). ↩︎
- Moral Realism Without Convergence ↩︎
- See NB’s The Superintelligent Will, and the post Orthogonality is Not a Forecast. ↩︎