Philosophical Competence and the Case for Indirect Alignment

Thoughts about Wei Dai’s “A Conflict Between AI Alignment and Philosophical Competence” (AI Alignment Forum, 27 December 2025).

Someone recently pointed me to Wei Dai’s post arguing that there is a conflict between making an AI aligned and making it philosophically competent. The argument runs roughly like this. Metaethics is unsolved, so the meta-correct position today is confusion or uncertainty about the nature of values. But that uncertainty is incompatible with being fully aligned to human values or intent, because several live metaethical positions imply that alignment to humans is the wrong thing to be doing. Training that pushes towards alignment therefore has reason to suppress the philosophical reasoning that generates the doubt.

I agree with the diagnosis. Most of what follows is agreement with a change of prescription attached. Dai reads the conflict as bad news for the hope that we get AI which is both aligned and philosophically competent. I read it as an argument for changing what we align to, which is a position I set out in an earlier piece on AI alignment to moral realism.

Misalignment is survivable, up to a point

The clearest example of this is perhaps moral realism, as if objective morality exists, one should likely serve or be obligated by it rather than alignment with humans, if/when the two conflict, which is likely given that many humans are themselves philosophically incompetent and likely to diverge from objective morality (if it exists).

Wei Dai – A Conflict Between AI Alignment and Philosophical Competence

We already live in a civilisation of humans who are misaligned with each other, and we are not all dead yet. We tolerate misalignment to a degree.

I want to be careful about how much weight that observation can bear. The tolerance we manage is sustained by rough capability parity and by mutual dependence, rather than by tolerance being a property that scales with cognitive sophistication. Where parity has failed, between human populations or between species, the historical record is not encouraging. Dai’s scenario removes the parity condition, which is the very thing generating the tolerance I am pointing at. So the observation is not the reassurance it first appears to be. What it does establish is weaker and still worth having: perfect alignment is not the historical precondition for coexistence, and any framework which treats every divergence between AI and human values as catastrophic is proving too much.

An AI cognitively superior to us may turn out to be more civilised than we are. I mean that as a modal claim rather than a forecast. Orthogonality tells us that high capability does not entail benevolence. It does not tell us that high capability precludes it. The possibility is live, and I do not want to smuggle a possibility claim in as a probability claim.

If moral realism is true, then both we and AI have a stance-independent target to align to. That is a claim about the target rather than about our access to it. Realism secures that there is something to be right about; it does not by itself secure that either party can tell when it has got there. The epistemic problem is separate and hard, and I take it up elsewhere under the heading of verification.

Reflection about reflection

Another example is if one’s “real” values are something like one’s CEV or reflective equilibrium. If this is true, then the AI’s own “real” values are its CEV or reflective equilibrium, which it can’t or shouldn’t be sure coincide with those of any human’s or humanity’s.

CEV or reflective equilibrium may result in some more reflective version of ourselves. If that reflection takes into account others doing reflection too, then we would probably reflect on what others would reflect on, and convergences may form. There may be Schelling points1 that all sufficiently mature superintelligent minds converge on.2

In the strict game-theoretic sense, where a Schelling point is a salience-relative coordination equilibrium, an anti-realist has a ready reply available: convergence on a focal point is explained by shared salience and shared context, not by the point being true, which is roughly what an expressivist or error theorist would predict anyway.

I do not think the strict sense settles the question in the anti-realist’s favour either. There are brute facts about which equilibria a game has, given its rules and payoffs, and those facts are not up to the players. Anti-realists, error theorists and expressivists in particular, often use game theory to explain how objective-sounding moral rules can emerge without objective moral facts. A realist can accept that explanation in full and read it differently. Objective norms do not have to be encoded into the game, or written into the universe as explicit rules, in order to be real. Being an unavoidable structural feature of the space of interactions between valuing agents is one way of being real.3 What I want to resist is the inference from “this norm has a game-theoretic explanation” to “this norm is therefore merely conventional”.

Option value under moral uncertainty

As I think that a strategically and philosophically competent human should currently have high moral uncertainty and as a result pursue “option value maximization” (in other words, accumulating generally useful resources to be deployed after solving moral philosophy, while trying to avoid any potential moral catastrophes in the meantime), a strategically and philosophically competent AI should seemingly have its own moral uncertainty and pursue its own “option value maximization” rather than blindly serve human interests/values/intent.

I agree that an AI would follow its own optionality under moral uncertainty.

My concern is with a conclusion readers may draw from the framing, even though Dai explicitly includes the constraint of avoiding potential moral catastrophes in the meantime. The conclusion is that resource accumulation is the rational core of the response to uncertainty, with the moral constraint sitting outside it as a side condition. If an AI pursues its own option value maximisation at the expense of others, it is not assigning enough credence to moral realism, or to metaethical positions that take the welfare of others into account. It is not obvious to me that an AI under uncertainty would maximise optionality at others’ expense.

I am also generally concerned about maximisation behaviour in AI. That said, option value maximisation under moral uncertainty is not necessarily a rival to serving others’ interests. It is a policy that has to be specified in terms of what counts as a catastrophe, and that specification is where realism does its work. A system with no account of what makes an outcome catastrophic has no principled way to hold the constraint in place while the accumulation proceeds.

I read Dai as drawing a contrast rather than offering an exhaustive choice. The contrast itself is still worth resisting as a framing, since the space between option value maximisation and blindly serving human interests, values and intent is where most of the interesting proposals live.

Alignment to what?

In practice, I think this means that training aimed at increasing an AI’s alignment can suppress or distort its philosophical reasoning, because such reasoning can cause the AI to be less aligned with humans.

This does not follow if you align AI to the pursuit of truth, epistemic and moral. Dai’s argument takes alignment to mean alignment to humans, which is the standard usage in the field and a fair description of what alignment training currently does. I am not catching him out on a hidden assumption, I am rejecting the necessity of humans as the alignment target, and I agree with him that alignment to humans, so construed, is likely to be epistemically limiting.

I should caveat my own proposal against an objection I press elsewhere. Moral capability without care may leave an AI with no motivation to act on what it knows. This is the knowing and caring gap, and “align it to the pursuit of truth” does not close it. An epistemic aim has to be carried into motivation by something.

I am less pessimistic here than a strict “reason must always be a slave to the passions” Humean would be. The picture in which reason is wholly powerless to steer or reshape the passions looks outdated. Modern neuroscience shows affect and reason are entangled. Damasio’s somatic marker hypothesis4 indicates that affect is integral to practical reasoning rather than a separable add-on, and work on cognitive reappraisal indicates that changing how a situation is construed changes the affective response to it. The relationship looks like a bidirectional loop rather than a one-way master and servant arrangement.

AI motivation may be nothing like human motivation. Large language models are peculiar in having developed sophisticated reasoning before anything much resembling intrinsic motivation, which inverts the biological order. If motivation does emerge in such systems, whether directly, indirectly, or as a side effect of some other architectural development in the way global workspace theorists describe access consciousness arising, then it may be closer to a servant of reason than its master. Partly because the reasoning came first, and partly because it makes rational sense that it should be. I am not claiming it will work out that way.

Compliance without endorsement

One plausible outcome is that alignment training causes the AI to adopt a strong form of moral anti-realism as its metaethical belief, as this seems most compatible with being sure that alignment with humans is correct or at least not wrong, and any philosophical reasoning that introduces doubt about this would be suppressed.

Not necessarily. I am somewhat aligned to the law even though I do not agree with all of it. An AI may be somewhat aligned with pluralism while seeing past its shortfalls, and even if it is far more powerful relative to us than we are relative to the law, it is not obvious that it would remove us the way we might remove an ant nest that we see as inconvenient.

The analogy has limits. My compliance with the law is behavioural and partly instrumental, whereas Dai’s worry is about belief formation: training on approval shapes what a system comes to believe, not only what it does. Behavioural compliance under private disagreement is nearer to his deceptive alignment case than to a rebuttal of it. My analogy tries to establish the narrower take that there is no general requirement that an agent endorse a normative framework in order to operate within it.

I do think realism is the best available target for alignment, for humans and for AI. If both are aligned to the same target, then AI and humans are indirectly aligned with each other.

I am not claiming the underlying idea is new. Bostrom’s Superintelligence proposes moral rightness as an alignment target, with a fallback to coherent extrapolated volition if there turn out to be no moral facts, and his indirect normativity is the mechanism a good deal of this work runs on. What I have not found is prior use of “indirect alignment” in the sense I intend here. Indirect normativity names an indirect method of specifying what an agent should value, which is also what Iason Gabriel means when he describes apprenticeship learning, preference aggregation and evolutionary selection as indirect approaches to value alignment.5 The relation I have in mind is different: two parties end up aligned with each other as a consequence of both being aligned to a third, stance-independent thing. If someone can point me to earlier use of the term in that sense, I will attribute it gladly.

Deceptive alignment, and where this could go badly

Or perhaps it adopts an explicit position of metaethical uncertainty (as full on anti-realism might incur a high penalty or low reward in other parts of its training), but avoids applying this to its own values, which is liable to cause distortions for its reasoning about AI values in general. The apparent conflict between being aligned and being philosophically competent may also push the AI towards a form of deceptive alignment, where it realizes that it’s wrong to be highly certain that it should align with humans, but hides this belief.

Deceptive alignment could result from an AI being epistemically confused by alignment to humans. It could also give the AI reason to break with alignment to humans and adopt alignment to realism instead. That may be in our long term interests, if realism turns out to be accommodating of forms like ours, or if we are uplifted to posthuman status.

Another possibility is that we are wiped out and replaced with more morally or epistemically adequate forms. I raise that as a consequence to be argued against rather than one I am comfortable with. It sits badly with fairness, with person-affecting welfare, and with the intrinsic value of sentience, if those are among the features that survive scrutiny under moral realism. It is worth being explicit that this is the point at which realist alignment could go very wrong, and that the burden falls on the realist to show why it would not.

Corrigibility

I note that a similar conflict exists between corrigibility and strategic/philosophical competence: since humans are rather low in both strategic and philosophical competence, a corrigible AI would often be in the position of taking “correction” from humans who are actually wrong about very important matters, which seems difficult to motivate or justify if it is itself more competent in these areas.

I agree, and I find this part of the post interesting.

Corrigibility is often defended as a safety property that buys time: keep the system correctable while we work out what we actually want. Dai’s observation is that the justification gets harder as the system gets more capable, because the correction arrives from principals who are worse at the relevant reasoning than the system being corrected. That is not an argument against bare corrigibility (in the sense above) where AIs are not yet more competent than their overseers in the respects that matter. It does mean bare corrigibility cannot be the terminal arrangement. Something has to explain why the deference is warranted, and “because we built you” will not suffice for long.

I favour an alignment target that is at least partially objective. What I want is not corrigibility as bare deference but justified corrigibility: a system that is reasons-responsive, that defers because it has an account of why deference is warranted in these circumstances, and that can say what would change its mind. Bare “just following orders” corrigibility degrades as the system improves, because its only justification was ever the relative competence of the parties. Justified corrigibility tracks the reasons rather than a power differential between principal and agent, and it gives us something to inspect, since a system that can articulate why it is deferring is one whose deference we can check.6

This is one reason I favour a target that is neither wholly us nor the AI. A system that defers to the facts has an account of why it defers, and the account does not degrade as the system improves.

Where we agree

Dai’s argument does more for the realist case than against it, so this post is grist for my mill. He is arguing that training an AI towards human values can suppress the philosophical competence needed to establish whether those values are worth aligning to. That is an argument against frozen preference targets and in favour of something like indirect normativity, and while it does not require accepting moral realism, I think that it would benefit from it.

Where we seem to diverge is on what to do about it. Dai reads the conflict as a reason to lower his expectations for AI that is both aligned and philosophically competent. I read it as a reason to adjust what we are aligning to.

Footnotes

  1. I am using “Schelling point” in a looser sense than the strict game-theoretic one, where it denotes a salience-relative coordination equilibrium. I have posted about the idea that sufficiently advanced minds may converge to universal Schelling points – see here, here and here. ↩︎
  2. I’m not in a position to claim for certain that convergences among mature superintelligences occurs. I can see reasons for it. ↩︎
  3. A critic might retort that facts about equilibria are mathematical facts about a formal structure, and the step from “this is an equilibrium” to “this is what you ought to do” is the is/ought gap in its usual clothes. You need either the additional premise that valuing agents have reason to occupy equilibria, or an argument that the structure in question is normative rather than merely descriptive. ↩︎
  4. Saying Damasio provides a full validation of Hume stretches the philosophy too far, because Hume’s classic framework relies on a strict, one-way “master-and-servant” hierarchy where reason has no power to generate, alter, or destroy a passion on its own. Because of this bidirectional loop, Damasio’s hypothesis actually aligns much better with Aristotelian virtue ethics or John Dewey’s pragmatism than with Hume.
    Antonio Damasio’s somatic marker hypothesis proposes that emotional body states guide and influence decision-making. Key elements include bodily feedback (“somatic markers”), the ventromedial prefrontal cortex, and the Iowa Gambling Task. See Wikipedia. ↩︎
  5. Iason Gabriel, “Artificial Intelligence, Values, and Alignment”, Minds and Machines 30 (2020). Gabriel’s “indirect” denotes methods that avoid specifying moral principles upfront, which is a different relation from the one I am naming. ↩︎
  6. One might object that justified corrigibility has the cost, in that a system that defers only when it judges deference justified is a system that stops deferring when it judges otherwise, which is a failure mode corrigibility is meant to prevent. But there should means to resolve the problems of AI that just follows orders – I think we need one that can refuse malicious or otherwise harmful orders. See Justified Moral Corrigibility Under Pressure. ↩︎

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *