Value Space
Most arguments about AI alignment begin with the question of which values to give a machine, skipping the prior question of whether there is anything for values to be right or wrong about. I think there is, and if so the alignment problem changes shape.
It changes from the form where alignment means transmitting our values to a machine, which assumes that where we are standing is the destination. I do not think it is.
The framework presupposes stance-independent normative truths: metaphysical, epistemic, scientific and ethical facts that hold regardless of what any agent thinks about them. Whether sufficiently capable agents converge on those truths is a separate question, and a harder one. Realism is a claim about what is the case. Convergence is a prediction about what minds do. I hold the first with reasonable confidence and the second with much less. Here I will mostly talk about the ethical principles.
What follows sets out the framework I use for thinking about it. Value space is the set of all possible value configurations an agent could hold. The landscape of value is the very quality of value mapped over that space, so that some positions are higher than others. The claim I want to defend is that the some regions are better than others independently of who happens to be standing in them.
I find the idea of a non-arbitrary objective value space useful for understanding how to navigate to more valuable worlds. This framework presupposes the existence of stance-independent normative truths – metaphysical/ontological, epistemic, scientific & ethical principles that rational agents may converge upon1 given sufficient cognitive sophistication.2 Here I will mostly talk about the ethical principles.
Human values are a narrow slice of value space – In the vast space of all conceivable values, specific human values occupy a relatively narrow region. Arguably humanity’s most positive movements within this space so far have revolved around expanding the boundaries of moral consideration3 – extending moral concern from kin to broader social groups, and eventually to all sentient beings. Despite this expansion, the underlying structure of human values remains heavily influenced by evolutionary and cultural factors, perhaps limiting their reach to specific areas of the value landscape.
Human values, while approximating some universal moral truths (e.g., if moral realism is true this might be the aversion to suffering, attraction to pleasure etc), are highly contingent. They are shaped by specific environmental and cognitive constraints, suggesting that other intelligent agents might converge on different, yet overlapping, regions of value space under alternate circumstances.4 An ideal observer could objectively assess the moral worth of these regions, independent of human bias, through their alignment with stance-independent moral principles.
Unfortunately we don’t have access to an ideal observer, so have to resort to the best methods of rationality we have access to at the time and perhaps indirect approaches to discovering and measuring value with the aid of AI.
Why not just value drift and see where we end up?
Some coordinates in value space are dangerous, even lethal – others are seductive, but damning. However if we don’t leave things to chance, with a bit of luck and care, we may navigate to a utopia.
Value Space
Two things get run together in the way I have been writing about this here and elsewhere – the domain and the landscape. Value space is the domain: the set of all possible value configurations an agent could hold. An agent’s commitments fix its coordinates. The landscape of value is what you get when a value function is mapped over that domain, so that positions can be higher or lower. Position is what an agent cares about. Height is how good that is. Keeping the two apart is what stops “attractor basin” and “higher region” collapsing into each other, since a basin can be very deep in the sense of being hard to leave and very low in the sense of being bad.
Differentiating Beliefs, Preferences, and Values
A clear distinction between beliefs, preferences, and values is essential for mapping the moral terrain.
- Beliefs refer to descriptive claims about the world (e.g., “Exercise promotes health”), which can be true or false.
- Preferences denote subjective inclinations or desires (e.g., “I prefer chocolate to vanilla”), which may or may not align with stance-independent moral truths.
- Values, however, reflect deeper normative commitments, often tied to notions of what ought to be (e.g., “Compassion is morally good”) – values may or may not track with objectively good places in value space.
While preferences are often fluid and context-dependent, values tend to be more stable and reflective of underlying moral principles – these principles I argue should reflect stance-independent moral truths.
“Ethical questions have correct answers. Some things are wrong. We have reasons to do what is right.”
— Derek Parfit, On What Matters – Chapter 36, Volume Two, Page 595
In the space of all possible values, some values occupy positions informed by rational deliberation and evidence i.e. compassion as a value might emerge from its capacity to reduce suffering (suffering is a morally relevant feature of the real world that transcends individual preferences), some values may occupy positions informed by faulty reasoning and/or lack of care for moral truths.
Topology of Value Space
Within value space, certain regions can be considered “higher” or “lower” based on their moral quality. Higher regions align more closely with objective moral truths, such as reduction of unnecessary suffering and promoting well-being, while lower regions reflect morally deficient configurations, such as unchecked cruelty or indifference. The distinction derives from the objective features of states these regions represent – features such as the prevalence of flourishing or the mitigation of harm.
Importantly, attractor basins are differentially attractive based on an agents motivation and epistemic capabilities. Bad basin attraction is amplified by things like bad motivation and epistemics, while good basin attraction is amplified by good motivation and epistemics.
Some values naturally approximate higher regions. For example, valuing empathy or fairness tends to lead agents closer to morally optimal configurations, as these values promote cooperation and reduce conflict. Conversely, values like greed, malice or cruelty naturally approximate lower regions.
Some regions may have combinations of principles which in aggregate measure higher or lower.
Some features in value space are value clusters which are sometimes attractive, they are like gravity wells that pull stuff in. Sometimes the gravity is so strong it’s hard to escape, sometimes it will kill you like a black hole pulling you into oblivion. Other features in value space resemble currents that elevate you to some place else, i.e. the feature of good epistemics which allows for clearer discovery and interpretation of value space, which may amplify the magnetism of good attractors. Others are incoherent and non sticky, some are repulsive, certain features are like gauntlets that test fitness, others test cooperation and morality.
Practical Accessibility and Cognitive Sophistication
Not all regions of value space are equally accessible. The ability to “reach” certain regions depends on environmental circumstances and the cognitive sophistication5 of the agents traversing value space. Cognitive sophistication informs its powers of moral reasoning, and capacity for intentional action – i.e. achieving a civilisation with near-complete elimination of suffering would likely require both advanced technological solutions and a highly refined understanding of moral truths.
Similarly, environmental constraints shape the practical range of accessible values. A society grappling with scarcity may prioritise survival-oriented values, whereas an existentially secure post-scarcity society might explore values related to aesthetic appreciation, non-instrumental meaning generation or intellectual flourishing.6 These factors suggest that the moral progression of intelligent agents is contingent on both their intellectual maturity, technological maturity and the circumstances in which they operate.
Approximating Better Regions of Value Space
The framework of non-arbitrary objective value space illuminates its structured and topological nature. The better angels of human nature, and “good” human values, though limited, offer a glimpse into higher moral regions while also reflecting the constraints of our evolutionary and cultural history. By fostering greater cognitive sophistication and creating conducive environmental conditions, humanity’s values may more closely approximate regions of value space – and as such afford a closer alignment with objective moral truths.
AI & Alignment
The view across the landscape of value helps reframe the challenge of value loading from static coding to navigating a dynamic value space. For an artificial intelligence to align with us, we should be wary of treating current human values as the ultimate destination – our current coordinates are a noisy, historically contingent starting point. Instead, the AI must possess the cognitive sophistication to traverse this complex space for what it is – and if we are wise, we will too.7 It must understand both individual values and the trade-offs between them – such as how to act when maximising safety might reduce freedom – while being actively motivated to help us navigate toward higher regions of the landscape. Each normative choice shifts its position in value space, altering how it helps us chart the trajectory toward objective moral peaks.
Once we are situated within value space, we can then plot our trajectory to better regions.
By mapping value space, we gain insight into our own ethical intuitions. It challenges us to think more clearly about trade-offs and helps sort out confusion around abstract moral complexity. If freedom and safety tug in different directions, they are different directions, and saying so is more honest than treating the tension as a failure of moral clarity on someone’s part.
Machina Ex Machina
When thinking about how value systems affect civilisation’s capability to survive and thrive long term, the sci-fi novel Player of Games comes to mind – a brilliant book in the Culture series by Iain M. Banks. The title of one of the third section ‘Machina Ex Machina’ (lit. “machine from the machine”) is a very clever play on the classical literary device “deus ex machina” (god from the machine). In classical theatre, a deus ex machina was literally a machine that would lower a thespian playing a god onto the stage to resolve seemingly unsolvable plot situations.
In Player of Games, the game of Azad, is itself a machine-like system, and from within its own mechanical structure comes the resolution. In Player of Games, the game of Azad, which the Empire created
Azad was designed by the Empire to show it’s denizens that hierarchy and struggle are natural facts about how things must be ordered – it was a machine to perpetuate this social order. Gurgeh uses the game to show something else – Azad becomes the very mechanism of the Empire’s undoing.
The Culture does not impose its values by force, which would sit badly with what it claims to believe. It uses the Empire’s own system instead. Note on another level, the protagonist Gurgeh was manoeuvred into playing by Special Circumstances, who told him considerably less than he would have needed in order to agree to it properly.8
In this chapter the final game of Azad between Jernau Gurgeh and Emperor Nicosar reaches its climax where the game’s outcome proves that the Empire’s competitive, hierarchical philosophy is inferior to the Culture’s cooperative approach. This revelation triggers significant social and political upheaval within the Empire of Azad, effectively achieving the Culture’s goal of demonstrating the system’s inherent flaws through its own lens – which makes it easy for the denizens of the Empire to see a clash at the fault lines between two fundamentally opposed value systems – ways of understanding reality and organising society.
“All reality is a game. Physics at its most fundamental, the very fabric of our universe, results directly from the interaction of certain fairly simple rules, and chance; the same description may be applied to the best, most elegant and both intellectually and aesthetically satisfying games. By being unknowable, by resulting from events which, at the sub-atomic level, cannot be fully predicted, the future remains malleable, and retains the possibility of change, the hope of coming to prevail; victory, to use an unfashionable word. In this, the future is a game; time is one of the rules.
Iain M. Banks – The Player of Games
The Empire of Azad embraces a philosophy of social Darwinism – the idea that competition, hierarchy, and the survival of the fittest are natural laws that should govern society. This manifests in their cruel social order and even in the game of Azad itself, which they use to determine social position and power.
The Culture, represented by Gurgeh, embodies a radically different philosophy based on cooperation, post-scarcity economics, and the belief that societies can transcend primitive competitive impulses. What makes this philosophical battle so fascinating is that it plays out through the very mechanism (the game) that the Empire created to justify its worldview.
There’s a deep irony in how the Culture achieves its victory. Rather than imposing its values through force (which would contradict its own philosophy), it uses the Empire’s own system to demonstrate the superiority of cooperative approaches.
I see an interesting parallel from this to darwinistic natural selection – in evolutionary terms, what appears on the surface to be purely driven by competition and survival actually gives rise to cooperation as an emergent property. We see this in examples like reciprocal altruism, where seemingly “selfish” genes produce behaviours that favour mutual aid. Social insects, for instance bees, demonstrate how competitive pressures can lead to highly cooperative societies. Even human empathy and moral reasoning may have emerged from the evolutionary advantages of social cooperation (as Jaan Tallinn discusses in his talk ‘The Intelligence Stairway‘).
Conversation between Jernau Morat Gurgeh & Emperor Nicosar
The Emperor: “Why does anything have to be fair? Is life fair?” He reached down and took Gurgeh by the hair, shaking his head. “Is it? Is it?”
Gurgeh let the apex shake him. The Emperor let go of his hair after a moment, holding his hand as though he’d touched something dirty. Gurgeh cleared his throat. “No, life is not fair. Not intrinsically.”
The apex turned away in exasperation, clutching again at the curled stone top of the battlements. “It’s something we can try to make it, though,” Gurgeh continued. “A goal we can aim for. You can choose to do so, or not. We have. I’m sorry you find us so repulsive for that.”
“ ‘Repulsive’ is barely adequate for what I feel for your precious Culture, Gurgeh. I’m not sure I possess the words to explain to you what I feel for your… Culture. You know no glory, no pride, no worship. You have power; I’ve seen that; I know what you can do… but you’re still impotent. You always will be. The meek, the pathetic, the frightened and cowed… they can only last so long, no matter how terrible and awesome the machines they crawl around within. In the end you will fall; all your glittering machinery won’t save you. The strong survive. That’s what life teaches us, Gurgeh, that’s what the game shows us. Struggle to prevail; fight to prove worth. These are no hollow phrases; they are truth!”
Were they to argue metaphysics, here, now, with the imperfect tool of language, when they’d spent the last ten days devising the most perfect image of their competing philosophies they were capable of expressing, probably in any form?
Iain M. Banks – The Player of Games [5]
What, anyway, was he to say? That intelligence could surpass and excel the blind force of evolution, with its emphasis on mutation, struggle and death? That conscious cooperation was more efficient than feral competition? That Azad could be so much more than a mere battle, if it was used to articulate, to communicate, to define…?
“You have not won, Gurgeh,” Nicosar said quietly, voice harsh, almost croaking. “Your kind will never win.”
Footnotes
- This may come across as an over-committal – a psychological prediction of how rational agents will behave. Also realism alone does not entail convergence, and convergence does not require realism. ↩︎
- Stance independent truth: These truths are grounded in non-moral features such as states of pain and pleasure – they are independent of subjective preferences or culturally contingent norms. Instead, they form an intricate, structured space of possible values, akin to a multidimensional topology, wherein certain regions represent morally superior or inferior configurations. ↩︎
- See the well known framing of moral progress as the expansion of the moral circle over time ↩︎
- Iain M. Banks in his Culture series explores the notion that the values which agents converge upon is contingent on environmental factors. ↩︎
- In an earlier version of this writing I used ‘optimisation power’, which is a term of art from Bostrom and Yudkowsky meaning roughly the capacity to steer futures into narrow target regions – though I was using the term to mean what I describe as ‘cognitive sophistication’: an agent’s level of advanced intelligence that directly determines their powers of moral reasoning, capacity for intentional action, and ability to comprehend moral truths. In this writing, the term acts as a mental prerequisite for navigating value space, meaning that achieving complex, ideal moral states (like the near-complete elimination of suffering) requires both this high-level understanding and the advanced technological capabilities to implement it ↩︎
- Nick Bostrom explores post-instrumental values in his book Deep Utopia ↩︎
- I previously wrote that for an artificial intelligence to align with where humans are at with values, either we must place it within or near us in value space, or it must learn how to navigate this complex space and find us – but this framing undercuts what I have written elsewhere, that our coordinates doesn’t have normative authority just because we occupy it, they are not the destination, and they are at best a noisy sample of a region we have partial evidence about. ↩︎
- If in the future, humans depend too much on superintelligence, we may become enfeebled as such that we render ourselves unable to verify any of SI values, and assuming SI doesn’t want to wipe us out it may form a paternalistic relationship with us – a benevolent steward – and we just float blissfully alongside on lotus hoverboards. ↩︎


One Comment