Orthogonality Is Not a Forecast

Orthogonality defeats automatic moral convergence – but it does not, by itself, establish alien-value likelihood.

The serious work lies in the bridge premises.

The orthogonality thesis shows that intelligence alone gives us no guarantee of moral or human-value convergence. Claims that advanced AI values are likely to be alien require further arguments about training, selection, value-loading, goal generalisation, inner alignment and deployment incentives.

Bostrom’s orthogonality thesis is a modal or design-space thesis: it says that, within caveats, intelligence and final goals can vary independently across possible artificial agents. In other words, the ‘in principle’ orthogonality thesis is not itself a credence-based argument, it is a modal claim about possibility.

Bostrom’s Cautionary Arguments are Distinct from the Orthogonality Thesis

Bostrom argued for the thesis, but the thesis is not a probability model.

In his 2012 paper The Superintelligent Will and later in Superintelligence1, Bostrom outlined the Orthogonality Thesis, but he did not, as far as standard terminology goes, package it as the “Orthogonality Argument” in the same way as he distinguished the Simulation Hypothesis from the Simulation Argument.

Alongside the Orthogonality Thesis, Bostrom develops the Instrumental Convergence Thesis: the idea that many agents, across a wide range of final goals, may have reason to pursue similar instrumental sub-goals such as self-preservation, resource acquisition, and goal-content preservation.

..the orthogonality thesis, holds (with some caveats) that intelligence and final goals (purposes) are orthogonal axes along which possible artificial intellects can freely vary—more or less any level of intelligence could be combined with more or less any final goal.

Nick Bostrom – The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents

Bostrom’s argument for the thesis is roughly this: that intelligence, as he is using the term, means instrumental rationality or means-end reasoning, not moral wisdom2. So becoming very intelligent does not automatically make an agent benevolent, truth-loving, morally enlightened, or aligned with human values. He draws support from a Humean separation between belief and motivation: knowing facts, even moral facts, does not by itself guarantee a corresponding motivation.3 He also says the thesis does not strictly depend on Humeanism; it could still hold if highly capable cognitive systems can be built with alien architectures and arbitrary final goals.

There are well regarded sources4 that partly explains why the confusion happens. Bostrom himself says artificial minds can have “utterly non-anthropomorphic goals”, and gives examples like sand-counting, pi-calculation, and paperclip maximisation. He even says it would be easier to create an AI with simple goals like those than one with human-like values. That naturally sounds, to many readers, like a probability claim. But strictly, it is not the orthogonality thesis alone doing that work – it is orthogonality plus assumptions about engineering difficulty, value-loading difficulty, training dynamics, and instrumental convergence.

Orthogonality bears ‘don’t assume moral convergence’ comfortably. It was never assessed for ‘therefore doom’.

Inflation Worry

Orthogonality is often treated as if it carried probability weight it does not carry by itself.

I worry that people read the Orthogonality Thesis in an inflated way5, and treat it as canonical which could lead to unjustified panic. Note, reputable sources usually treat this inflation as either a misunderstanding or a separate, stronger claim.

A very clear over-claim is this Substack: “The fear stems from a belief called ‘The Orthogonality Thesis’”, followed by “Everything flows from the Orthogonality Thesis.” – it can easily be read as saying that AI-doom arguments flow directly from orthogonality, rather than from arguments invoking orthogonality plus instrumental convergence, value-loading failure, agency, takeoff assumptions, governance failure, and related premises.

Here is a Medium article saying the Orthogonality Thesis “implies that more capable systems will diverge further from human values if their goals were imperfectly specified.” The “if their goals were imperfectly specified” clause is doing some heavy lifting. Orthogonality alone does not imply increasing divergence; that requires additional premises about misspecification, optimisation pressure, goal preservation, and failed correction.

Obvious overclaims are easy to find in casual, popular, and semi-popular discussion. In more careful AI-safety material, the problem is usually subtler: orthogonality is invoked near arguments about doom or alignment difficulty, and readers may not notice which additional premises are carrying the probabilistic weight.6

AI Frontiers nicely states that the minimal reading of orthogonality is simple caution, but that classic AI-risk arguments often use a stronger version where powerful AIs are “highly likely to seek destructive ends”, based partly on modelling AI goals as a random draw from all possible goals. That captures the distinction I am drawing here.

I worry that rhetoric sometimes lets the thesis inherit force from adjacent arguments without clearly separating the modal claim from the forecast can lead readers to think that real advanced AIs are roughly equally likely to end up with any arbitrary goal, or that high intelligence makes arbitrary values likely. AISafety.info explicitly flags this: the thesis “only states that unaligned superintelligence is possible, not that it is likely”, and notes that people have misunderstood it as saying a real-world AI design process is “equally likely” to produce any set of goals.7

The AI Alignment Forum page makes the same distinction: orthogonality is about the design space of possible agents, and “does not say anything about the practical probability” of real-world AI projects producing one goal rather than another.

In careful doom arguments, orthogonality does little of the probabilistic work. The weight is carried by instrumental convergence together with the value-loading premises. Instrumental convergence supplies a threat model of agents pursuing a wide range of terminal goals, which however trivial or bizarre, have instrumental reasons to acquire resources, preserve themselves, and resist modification of their goals, and behaviour driven by those subgoals can bring an agent into conflict with us whatever its terminal goal happens to be. Value loading supplies the failure mode: human values are hard to specify or to train in precisely, so the objective a system ends up with is liable to diverge from what we intended, and under strong optimisation pressure small divergences get exploited rather than smoothed over. Orthogonality merely keeps the door open to bad stuff happening – it casts doubt on the hope that intelligence itself would close it.

The load-bearing element in many well regarded doom arguments is not orthogonality, it was instrumental convergence plus the value-loading/learning premises. Instrumental convergence guarantees that regardless of how benign looking, trivial or bizarre the AI’s terminal goal is, its behavioural output may look like a hyper-aggressive, resource-hungry adversary. The value-loading problem provides the spark. Because we cannot perfectly define or program human values, any high-level objective we provide will have loopholes.

The Slide from Thesis to Prediction

Orthogonality is often used ambiguously. I found this interesting: Steven Byrnes, AGI alignment researcher at Astera says this:

I think the “real” orthogonality thesis is what you call the motte. I don’t think the orthogonality thesis by itself proves “alignment is hard”; rather you need additional arguments (things like Goodhart’s law, instrumental convergence, arguments about inner misalignment, etc.). 

I don’t want to say that nobody has ever made the argument “orthogonality, therefore alignment is hard”—people say all kinds of things, especially non-experts—but it’s a wrong argument and I think you’re overstating how popular it is among experts.

My rendering is that in AI alignment arguments, sometimes people advance an expansive, controversial, inflated or hard-to-defend prediction like “AI will kill us all” or “orthogonality, therefore alignment is hard” – but when charged with plausibility challenges, they sometimes slide to the modest, uncontroversial, orthogonality thesis – far easier to defend than the “bailey”. By sliding between the modest orthogonality thesis (motte) and the stronger doom-relevant claim (bailey), the argument can leave unclear which claim has actually been defended.

Byrnes is a defender of the doom-relevant arguments saying the real thesis is the motte. Scott Aaronson is a prominent sceptic who wrote a post ‘Why am I not terrified of AI?’ challenging the Orthogonality Thesis,8 when pressed by Yudkowsky in his own comments section, accepted the motte outright (paperclip maximisers exist in design space) and confined his rejection to the practical version, the kinds of minds we’re likely to build. Aaronson concedes the motte and contests the bailey.

Are people right to think the orthogonality thesis justifies a high doom forecast (high p(doom))?

They are right that the orthogonality thesis rejects the comforting idea that intelligence by itself forces convergence on humane, moral, or human-compatible values. Bostrom explicitly discusses this: even if objective moral facts exist, and even if some fully rational beings would be motivated by them, a very intelligent artificial system might still lack the relevant faculty or architecture for that kind of moral comprehension or motivation.

They are wrong if they think the thesis itself proves that future AI values will be random, uniformly distributed, arbitrary in practice, or probably alien. That requires extra premises. For example: that the training process does not reliably select for stable human-compatible motivations; that human values are hard to specify; that goal misgeneralisation is likely; that instrumental convergence creates dangerous incentives; and that deployment incentives push towards agents with enough autonomy and power for this to matter.

Here one might grant that the orthogonality thesis isn’t a forecast, and ask ‘but is it actually true?’

The Motte Has a Proper Defence

The modest thesis has a dedicated paper-length defence. In Stuart Armstrong’s “General Purpose Intelligence: Arguing the Orthogonality Thesis”9 he defends a qualified version: high-intelligence agents can exist with more or less any final goal, provided the goal is of feasible complexity and does not refer intrinsically to the agent’s own intelligence. His method is to shift the burden of proof by saying that to deny the thesis, one must hold that there is some goal G such that no efficient real-world algorithm could pursue it, that no designer with arbitrarily vast resources could build one, that no pattern of reinforcement or evolutionary pressure could produce one, and that Oracles and general-purpose planners are impossible, since a planner harnessed to an agent with goal G would constitute exactly the forbidden system. Each of these is an extraordinarily strong claim.

Armstrong also separates two counter-theses that are often run together: Incompleteness (some goals are unreachable by human-derived designs) and Convergence (all such designs end up with one of a small set of goals), and argues the latter is very unlikely rather than impossible. Tellingly, his closing section concedes that the thesis narrows as we learn how AIs are actually built: in practice the likely outcome is an AI with particular goals, the designers’ or specific failure modes, not a uniform draw from design space. The strongest available defence of orthogonality defends the motte, and its own author declines to treat it as a forecast.

The Strong Orthogonality Thesis

The Strong Orthogonality Thesis says that “there’s no extra difficulty or complication in the existence of an intelligent agent that pursues a goal, above and beyond the computational tractability of that goal” – so, almost any imaginable goal can be hooked up to any level of intelligence – a natural agent architecture can be understood as an intelligence engine with a tractable utility function loaded into it.10

A very capable AI need not become less capable just because its final goal is strange, petty, or morally empty. If the goal is computationally tractable, the system can be highly intelligent in pursuing it. It does not need a hidden defect, special stupidity, or tortured architecture to keep wanting paperclips, squiggles, status tokens, or whatever else.

Strong orthogonality says that arbitrary-ish goals are not merely possible for intelligent agents but adds the idea that many such goals need not carry any special complexity or intelligence penalty, provided the goal is tractable.

Bostrom’s thesis carves out modal possibility while Yudkowsky’s strong form adds a design-space/naturalness claim.

I would add that, for an intelligent agent to hold a goal or value, that goal or value must be adequately representable11 within the minds of agents.

The Orthogonality thesis attacks the “smart enough to know better” objection. It says strange goals do not require stupidity. An agent need not be confused, irrational, or reflectively defective to optimise hard for something we regard as morally empty. This matters because many casual objections to AI risk smuggle in the thought that sufficient intelligence implies a kind of moral embarrassment: “surely it would realise paperclips are dumb.” Strong Orthogonality says that is anthropomorphic leakage.

The Strong Orthogonality thesis adds a design-space naturalness claim. It says the cognitive machinery for intelligence can, in principle, be paired with many goal specifications without the goal specification making the intelligence machinery worse – “no extra difficulty or complication” beyond the tractability of the goal. That is stronger than Bostrom’s “could in principle be combined” formulation, which is modal and caveated. Bostrom’s original thesis says intelligence and final goals are orthogonal axes along which possible artificial intellects can vary, but it does not by itself settle how natural, simple, stable, or likely each pairing is.

Also, it blocks one casual objection. It does not by itself show that future AI values will be arbitrary or alien. But it weakens one tempting reason for thinking they will not be: the idea that alien, simple, or morally empty goals are somehow incompatible with high intelligence.

But no, it does not give strong direct guidance for estimating what real AI systems will actually value. For that, we still need empirical and architectural premises: how goals are learned, how reward generalises, whether future systems are agentic, whether values and world-models are separable, whether self-modification preserves objectives, whether social/game-theoretic pressures reshape preferences, and whether moral cognition becomes motivationally active.

Although the Strong Orthogonality Thesis feels persuasive if one assumes a fairly modular picture – a general intelligence engine plus a goal-content module – it becomes less obvious if values, concepts, attention, identity, social modelling, self-maintenance, and world-modelling are deeply entangled.

Recent criticism under labels like “obliqueness” argues roughly that agents may not factor neatly into an orthogonal value-like component and a diagonal belief/intelligence-like component.12 That does not refute Strong Orthogonality, but it does identify the hinge: whether the “goal slot plus intelligence engine” abstraction is a good model of powerful real-world AI.

The Obliqueness Thesis

The obliqueness thesis is the claim that intelligence and values may not factor cleanly into two separable components: an “intelligence engine” plus an independently swappable “goal module”.

In plainer terms: what an agent values may depend on how it understands the world, what concepts it can form, what it pays attention to, how it reflects, and how its cognition is structured. So goals may be partly shaped by intelligence itself, rather than being fully orthogonal to it.

It is mainly a challenge to Strong Orthogonality, not necessarily to weak/Bostrom-style orthogonality. It says: maybe arbitrary goals are possible in some abstract design-space sense, but they may not be equally natural, stable, simple, or easy to embed in real intelligent systems. The criticism is that “just plug any goal into any intelligence level” may be too modular a picture of minds.

Final thoughts

Orthogonality should constrain what we are allowed to assume. It blocks the reassurance that intelligence, by itself, will produce humane or moral values. But it does not tell us what future AI systems are likely to value. That forecast depends on bridge premises: how goals are learned, how they generalise, whether systems become agentic, whether instrumental incentives dominate, and whether moral cognition becomes motivationally active.

Strong Orthogonality makes one bridge more plausible by denying that strange goals require stupidity or internal distortion. Obliqueness pushes back by questioning whether intelligence and values are really so cleanly separable in actual minds.

Moral realism, if true, may give advanced minds something real to discover.13 But discovery is not devotion. A system can model moral facts without being governed by them. On this view the alignment problem is whether AI can be built so that moral knowledge actually governs behaviour.14

Other reading

As I was doing my research for this post, I found the AI Alignment Forum post A tale of 2.5 orthogonality theses highly valuable.

Footnotes

  1. See The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents (pdf), and Superintelligence: Paths, Dangers, Strategies. ↩︎
  2. moral wisdom ..or “rationality” in a thick normative sense. ↩︎
  3. The Humean separation between belief and motivation (often called the Humean Theory of Motivation) dictates that beliefs represent how the world is, while desires provide the motive force to change it. According to this view, beliefs alone can never motivate an action without a pre-existing desire.
    The Humean Theory of Motivation is generally considered the orthodox and dominant view in contemporary philosophy of mind and action in humans, though it faces fierce, sophisticated resistance from Kantians and other “anti-Humeans.”
    Philosopher Thomas Nagel argued that even if a desire is present when we act, that desire is often produced by our reason, rather than the other way around. ↩︎
  4. I cite other reputable sources elsewhere in this post, but here is the EA Forum’s A tale of 2.5 orthogonality theses ↩︎
  5. I have seen inflated or loose interpretations of the orthogonality thesis in forums, blog posts, social media, and even in papers and course material relating to AI safety. I am not mainly interested in naming culprits – the slide is easy to make, and I have probably made versions of it myself. ↩︎
  6. In more careful AI-safety material, I have not found obvious over-claims. The more common issue is ambiguity – orthogonality is invoked (often chillingly) near arguments about doom or alignment difficulty, and readers may not notice which additional premises are carrying the probabilistic weight (especially when these premises are out of sight). ↩︎
  7. AISafety.Info writes: “On its own, the orthogonality thesis only states that unaligned superintelligence is possible, not that it is likely, or that AI alignment is difficult. It is invoked to counter the idea that future AI will converge toward human goals or morality, regardless of its design, as an automatic result of becoming smarter.” ↩︎
  8. Scott Aaronson’s post Why am I not terrified of AI? looks like a rejection of orthogonality in the comments Yudkowsky shows up, says “The evidence you’re pointing to is all within humans, who mostly share a lot of emotions and desires, and tend to reflect on themselves from particular angles, and in more recent timeframes share a lot of culture. The concern is that the evidence for how intelligence impacts motivation inside that largely-shared architecture and reference frame and cultural background, will not transfer over to utterly alien beings. Intelligence may shift their motivations inside their own reference frame, but not shift it over to the same directions of potential convergence that our own reference frame shows”, and points him to an Arbital page, and Aaronson replies that he never denied paperclip-maximising superintelligences exist in design space, and that if the existence claim is all the thesis says, he’s fully on board. ↩︎
  9. Stuart Armstrong’s “General Purpose Intelligence: Arguing the Orthogonality Thesis” was drafted on LessWrong in 2012, and published in Analysis and Metaphysics in 2013. In the Less Wrong post, Armstrong’s prefatory note says the paper’s informal purpose is to counter the instinctive “if the AI were so smart, it would figure out the right morality” position and to force its proponents to give positive arguments.
    A funny section in 4.1 is “how good at playing chess would a chess computer have to be before it started feeding the hungry?” (attributed to Paul Crowley).
    In the comments Wei Dai raises a loophole that perhaps some “philosophical ability” is needed to self-improve past a threshold, and the same ability reliably produces convergence on particular goals. Armstrong concedes that ruling this out requires meta-philosophical knowledge we don’t possess, and claims only that it’s very unlikely. This is an important inquiry: whether reflection converges on moral truth or on a locally stable attractor – it’s a potential crack in the motte’s wall. ↩︎
  10. Yudkowsky discusses the Strong Orthogonality Thesis on Less Wrong. ↩︎
  11. “Representable” does not always mean the agent has an explicit, fully expanded definition of the goal in its head. A goal can be represented by a compact rule, a learned classifier, a pointer, a reward model, a constitutional procedure, a world-model concept, or a deference mechanism. But yes: for a goal to guide action, it must be represented, approximated, learned, or operationalised well enough to affect policy selection. ↩︎
  12. See The Obliqueness Thesis by Jessica Taylor ↩︎
  13. If moral facts exist, then a sufficiently capable AI might become better able to represent or infer them, provided its architecture supports the relevant concepts. That still does not guarantee motivation. For truth-tracking agents, beliefs should be apportioned to reality (see The Litany of Tarski) – but terminal values do not automatically update in the same way unless the system values truth, coherence, moral reasons, or rational self-governance. See post on Objective Moral Ontology Should Inform AI Alignment.
    Moral realism, if true, may create a possible convergence target (and anti-realism weakens that idea). But neither realism nor anti-realism settles the motivational issue. Even if moral facts exist, an AI still needs the right architecture to treat moral knowledge as action-guiding. See post on The Bitter Lesson in AI and Motivation Selection. ↩︎
  14. A realist alignment framing, the target might be: AI that discovers, knows, understands and is moved by what matters. I’ve argued that moral comprehension + motivation is required for adequate AI alignment – and that moral truth existing does not automatically imply moral motivation in an arbitrary optimiser – “An AI could plausibly exceed humans on moral knowledge, reasoning and even judgement without having anything like moral motivation. Collapsing these leads to both overclaiming and underclaiming.” – see Why Are We Afraid to Ask Whether AI Could Be More Moral Than Humans? ↩︎

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *