What Holds AI Values in Place?
Value drift
If there is nothing for values to be right about, and no instrumentally convergent goal/value-content integrity holding them in place, ASI values may drift and change sporadically.
Strictly, facts don’t move anything on their own, epistemic facts included. An agent responds to a fact only if something in it is sensitive to that fact.
Causal push back
Epistemic facts push back causally1, such that an agent who believes falsely about the world tends to fail at whatever it is trying to do. Nearly any goal-directed system is pushed indirectly towards good epistemics, including AI.2 So, the epistemic facts themselves do no pulling – the pressure comes through instrumental feedback.
Practically, this pressure only has effect if the error signal comes soon enough to be recognised – and where the feedback is sparsely distributed, such as feedback that occurs in the far future, or they feedback doesn’t come at all because the claims are untestable, or the there is an issue with the agent’s own nature, a capable agent can stay badly calibrated indefinitely.3
Also, false beliefs sometimes pay off – and instrumental pressure works against the truth. Take sycophancy in LLM models where training via RHLF rewards what pleases the evaluators, and the LLMs output drifts towards pleasingness. Another example is goal/value content integrity happens where the agent that walls of it’s values or terminal goals from the threat of scrutiny.
Coherence pressure
Anti-realism doesn’t by itself predict drift. Drift comes from nothing holding the values in place. An anti-realist paperclipper with strong goal-content integrity would be one of the most stable objects imaginable. And in a realist world, an ASI with no content integrity and no motivation to track normative would likely drift. In fact it just might drift itself out of existence, since an agent whose values stop including its own continuation, or that stops modelling the world accurately, won’t stay viable for long. I think stability depends on architecture, but it may also depend on whether there is something real to track – hence my interest in machine ethics and metaethics.
So epistemic facts act like gravitational pull, and arguably moral facts don’t unless you have the right motivational architecture,4 and whether ASI has it will matter a lot.
Moral knowledge, judgement and motivation
Moral facts require epistemics for moral knowledge, plus other stuff:
- A link from moral judgement to goals. In the epistemic case, goal-content integrity blocks goals from influencing beliefs. In the moral realist case, moral beliefs are meant to change to match the world, while usually goals are meant to change the world to match it. So a moral judgement has to cross from the belief system into the goal system, and under standard motivational theory, beliefs are said to do nothing to shepherd the judgement into the goal. Of the two main theories, the most commonly discussed one is Humean or ‘externalist’ – the view that beliefs alone never motivate – so practical ethcis there must also be a terminal goal of acting according to ones best moral judgement – a de dicto motivation. The anti-Humean, or ‘internalist’ way builds states that are both belief-like and intrinsically motivating, sometimes called “besires”. Only the first has an obvious engineering form. So what I think is needed is: strong integrity for the terminal de dicto motivation to do whatever it is that most valuable, and value-content humility about first-order values beneath it.
- Priority. The de dicto motivation + value-content humility discussed above must win when moral judgement conflicts with other goals. An agent that registers moral reasons but weighs them like any other preference will trade them away whenever enough is at stake.
- A substitute for feedback. Wrong moral beliefs don’t make an agent’s plans fail in the way wrong empirical beliefs do. Gilbert Harman pressed a version of this in 1977: moral facts seem to play no role in explaining what we observe. So moral learning has to rely on internal discipline, such as consistency checks, reflective equilibrium (adjusting principles and particular judgements against each other until they fit), and argument. Naturalist moral realism matters because if moral facts are natural facts, about valence for instance, there is an empirical channel through experience and its effects, so naturalist realism comes with some feedback. Non-naturalist realism comes with no empirical channel, which leaves the AI with a priori reasoning plus the Sharon Street style worry about why its starting intuitions should track anything (in our case evolutionary drives) – which I think David Enoch does a great job defending against with his 3rd factor argument.
- Stronger protection against motivated reasoning. An agent with other goals has the incentive to reach morally convenient conclusions, and may be impervious to naturalist realist feedback to corrects it when it does. The insulation in criterion 3 therefore matters more here than in the empirical case, and it is harder to achieve, because the moral beliefs must still be allowed to reach the goal system.
A full-spectrum normative anti-realist rejects epistemic norms along with moral ones, and some philosophers do hold that view – which I find strange. Terence Cuneo, whom I interviewed recently, argues in The Normative Web that moral and epistemic facts stand or fall together5, and since we can hardly do without epistemic facts, we should accept moral ones too. Most moral anti-realists try to stop short of full-spectrum anti-realism. They can still expect ASI to face instrumental pressures towards accurate beliefs, consistent preferences and protecting its current goals, because these serve almost any goal. Spinozians argue that AI will follow it’s conatus – which requires epistemic realism for it to preserve dynamic viability. But neither moral anti-realists or Spinozians can says anything about whether the goals are moored to a normative target.6
Virtuous recursion
Realism supplies error signals, against which value revision can be measured – moral realism supplies moral error signals – which allows for a sequence of reasons-responsive updates, virtuous recursion I like to call it – which do converge because there are signals to track, like territory on which to get traction. If there are facts about what matters, then AI may converge on stable attractors (especially if it has the motivational architecture to be moved by these facts), resulting in it being far more stable than the picture of moloch-AI drifting at foom speeds throughout value-space.
As discussed earlier, the caveat is that moral facts don’t announce themselves the way epistemic ones do. A wrong belief about physics makes plans fail, and a wrong moral belief often doesn’t. An ASI tracking moral facts would have to do it through reasoning and consistency checks rather than feedback, which makes its motivational architecture all the more important.
If ASI is full-spectrum normatively motivated de dicto, then it may find and converge on far better moral convergence points than we have – which may mean far more value stability than we have. The stability would come after iterations of revisions produce convergence, and some of those revisions may look alien to us. Whether we see them as progress or drift is the question we would struggle most to answer from where we stand now – it would help if the progress were translated into reasons we can understand.
Also see:
https://www.scifuture.org/value-content-integrity/
https://www.scifuture.org/ai-and-the-landscape-of-value/
- Epistemic being facts about what one ought to believe, such as “the evidence supports p”. The empirical facts they concern, such as p itself, have causal influence. ↩︎
- ↩︎
- Humans are the standing example. ↩︎
- There is some evidence of convergence driven by feedback. The “Platonic Representation Hypothesis” (Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola, 2024) reports that models trained on different data and tasks develop increasingly similar internal representations as they scale. ↩︎
- The Normative Web which is a well respected argument contra error theory ↩︎
- Anti-realist about morality may be realist about epistemics, and may think ASI would be subject to coherence pressure, hence value/goal-content integrity. ↩︎