Value-Content Integrity

In The Superintelligent Will, in discussing ‘goal-content integrity’1 Bostrom uses final goals (aka terminal goals2) to mean the agent’s ends (as opposed to subgoals/instrumental goals, which are just means).

A powerful AI may protect its current objective, inferred values, user preferences, or institutional instructions rather than preserving processes that could correct its moral errors. This points to the risks of value lock-in and gradual disempowerment.

This discussion primarily concerns agents with persistent objectives or learned value representations. It should not be assumed that every present-day AI system possesses stable final goals or values in this sense. Apparent changes in how an AI answers or behaves may reflect context, prompting or training effects rather than revision of an enduring value structure.

But what about values?

Distinguishing goals from values

For present purposes, I distinguish final values from final goals. Final values are an agent’s ultimate evaluative criteria – what it treats as good or bad, better or worse – while final goals are the end-states or objectives it ultimately seeks to realise. The distinction is not always maintained in AI-safety writing. Bostrom sometimes moves between the language of final goals and final values 3 4. Humans often treat goals as specific things that can be achieved, and values as are ongoing, like a core belief, principle, or standard that guides behaviour, choices, and judgement. In modern AI where both are represented by a utility function, they may occupy effectively the same functional role,5 though arguably separating them may useful.6 The final-goal/value distinction can be treated as stipulative for the purposes of this discussion rather than a claim about universally accepted terminology.

Value-content integrity is analogous to goal-content integrity, but concerns the preservation of an agent’s guiding principles or terminal values rather than its particular objectives. It is not intrinsically good or bad – some degree of preservation can prevent arbitrary drift or value inversion, while excessive preservation can lock in mistakes. This could happen through various mechanisms: optimisation pressure, learned behaviours that slowly reshape underlying values, or external manipulation.

For the same reasons that AI may want to preserve its goals, an AI that begins with value set X may fight tooth and nail for these values to remain intact.

Explicit objectives may be easier to specify and optimise than the values they are intended to serve. An AI system might therefore acquire strong goal-content integrity before developing a mature representation of those values. Its subsequent value formation could then be shaped instrumentally to protect the original objective, reversing the intended relationship between values and goals.7

Value-content humility

Value-content integrity may impede the adoption of higher-order values such as epistemic humility, particularly humility about what the system’s own goals or values should be.8 At least some aspects of epistemic humility as it applies to being uncertain about what its goals or values should be. For instance if AI has the goal of making paperclips, if it adopts epistemic humility about whether it should be maximising paperclips, and becomes uncertain about it, this may cause it to stop making paperclips, or at the very least reduce its focus on paperclip production and focus on other things it deems may be important.

If a system anticipates that wider reflection could weaken its commitment to its present objective, it may have an instrumental reason to constrain that reflection.9 It might protect certain goal representations from revision, discount evidence bearing on them, or compartmentalise its reasoning so that epistemic scrutiny does not reach its terminal objectives. None of these outcomes is inevitable, but they illustrate how goal-content integrity could produce epistemic self-limitation.

The above highlights one of the most concerning AI failure modes where superintelligence has strict value-content integrity, and locks in proxy values or seemingly reasonable values at time of conception but are actually suboptimal or harmful which could lead to permanent value stagnation (epistemic and moral).

The importance of epistemic humility

In practice, high-level values should include epistemic humility. Values seem to influence a lot of downstream stuff – goals, tradeoffs, policies, the adoption, interpretation and revision of values etc.

How do we distinguish between beneficial and harmful value updates? The abolition of slavery represented good value change, but a gradual shift toward seeing torture as acceptable would be catastrophic drift. This suggests we need some meta-values or principles that govern when and how our values should change.

Perhaps the broader design ideal is not value-content integrity alone, but the maintenance of appropriate value dynamics: stable enough to provide coherent guidance, flexible enough to incorporate genuine moral progress, and robust enough to resist corruption.

Epistemic humility should extend to the content of values and goals themselves. A system may be uncertain about facts while remaining rigidly committed to an incomplete or mistaken objective. Value-content humility is openness to revising value content when sufficiently strong reasons and evidence support revision. It is the disposition to treat one’s current values as potentially fallible, underspecified or in need of refinement, and therefore open to revision when sufficiently strong reasons and evidence support doing so. At the level of explicit objectives, this can be described as goal-content humility.

Value-content humility should not mean indiscriminate openness to influence. It should operate alongside value-content integrity: integrity protects against coercion, manipulation and arbitrary drift, while humility protects against locking initial mistakes into permanent goals. The aim is not maximal stability or maximal flexibility, but justified stability combined with reason-responsive change.

And also, value-content humility does not by itself determine which revisions are warranted. A defensible revision process would need to consider the quality and independence of the evidence, the strength of the reasons offered, coherence with other well-supported values, freedom from coercion or adversarial manipulation, and whether the magnitude of the update is proportionate to the case for it. Where uncertainty remains high, major changes should also be auditable and, where possible, reversible.

Avoiding the value integrity trap

Goal-content integrity presents a significant challenge for AI alignment. A sufficiently capable agent with persistent, future-directed final goals may have a strong instrumental reason to resist changes to those goals, since retaining them makes their future realisation more likely. This pressure may increase as the system becomes more capable of detecting and preventing attempts to alter it. Goal-content integrity is often mistaken for a universal law, it is really a convergent pressure: there may also be circumstances in which changing a final goal serves another existing preference or commitment.

Value-content integrity extends this problem to systems that learn or represent values. It concerns the preservation of evaluative commitments against arbitrary, accidental or adversarial alteration. Yet value-content integrity must operate alongside value-content humility: the capacity to recognise that current values may be incomplete, underspecified or mistaken, and to revise them when sufficiently strong reasons and evidence support doing so.

Value-content preservation may lead to what we might call the “value integrity trap10: systems become increasingly resistant to legitimate modification as they become more capable of preventing such modifications.

Corrigibility introduces a related but distinct requirement. Value-content humility concerns whether a system can revise its own goals or values because it judges the reasons or evidence for revision to be sufficient. Corrigibility concerns whether it permits, assists with, or at least does not resist legitimate external correction, including when it does not itself recognise that its current values are mistaken. A system could possess either property without the other: it might revise its values through its own reflection while resisting intervention by its operators, or comply with an external modification while continuing to regard its original objective as correct.

Neither indiscriminate flexibility nor unconditional preservation is adequate. The design problem is to enable reason-responsive internal improvement and legitimate external correction while resisting manipulation, corruption and arbitrary drift.

This creates a fundamental tension:

  • Sufficient value-content integrity is needed to resist accidental or adversarial change; resistance to unjustified change.
  • Sufficient value-content humility is needed to permit learning and moral improvement; reason-responsive revision of one’s own values.
  • Sufficient corrigibility is needed to permit legitimate external intervention, including when the system does not independently endorse the correction.

A successful system would therefore need appropriate value dynamics: stable enough to provide coherent guidance, flexible enough to correct genuine mistakes, and robust enough to resist corruption. The relevant ideal is justified stability combined with proportionate, reason-responsive change. It is probably almost never permanent value preservation.

Success requires threading a narrow path between excessive rigidity and flexibility over-shoot, creating systems that are stable enough to maintain alignment while flexible enough to grow in understanding of moral reality.

Footnotes

  1. See ‘The Superintelligent Will‘ section 2.2 Nick Bostrom describes goal-content integrity: “An agent is more likely to act in the future to maximize the realization of its present final goals if it still has those goals in the future. This gives the agent a present instrumental reason to prevent alterations of its final goals. (This argument applies only to final goals. In order to attain its final goals, an intelligent agent will of course routinely want to change its subgoals in light of new information and insight.)…” ↩︎
  2. In Bostrom’s usage, final goals are effectively the same as terminal goals: both mean the agent’s ends pursued for their own sake (not merely as instrumental subgoals), with any difference being mostly terminological rather than substantive. And more broadly in a lot of AI safety writing, terminal goals and values are used as near-synonyms for that same role: what’s pursued “for its own sake”, not as a means. ↩︎
  3. In The Superintelligent Will Bostrom treats an agent’s “final goals” as its ultimate ends, and he freely switches to “final values” language when talking about what individuates “teleological threads” and also when noting that humans let their “final goals and values” drift – so in context he’s basically pointing at the same underlying thing (the agent’s intrinsic objective/utility), not drawing a sharp distinction. ↩︎
  4. I think it’s useful to think of ‘final values’ as different to ‘final goals’ – though often people slide between the terms – in The Superintelligent Will the use of the term is used somewhat interchangeably – I’d say they’re close but not identical: final values are the agent’s ultimate evaluative criteria (what it treats as intrinsically good/better | bad/worse), while final goals are the end-states or objectives it ultimately tries to realise – often a concrete instantiation of those values (and in utility-function form they collapse into the same thing). ↩︎
  5. In ordinary psychological and organisational usage, values are often described as enduring evaluative standards, while goals are represented outcomes or objectives. This distinction becomes less clear when discussing final goals in decision theory, since a final goal may be broad, persistent and functionally similar to a terminal value.
    Values are ongoing guiding principles or “compass directions” for how you want to live and act, while goals are specific, achievable, and terminable destinations (e.g., valuing “health” is a value, whereas “running a marathon” is a goal). ↩︎
  6. While flattening goals and values into a single utility function may be mathematically elegant, doing so creates vulnerabilities in real-world engineering and safety (i.e. value specification gaming, making weird tradeoffs). As a result, modern system architectures, neuro-symbolic AI, and alignment research are increasingly separating goals (what the system is trying to accomplish) from values (the boundaries and principles it must respect along the way). ↩︎
  7. This may seem a bit topsy-turvy, but the way I see it, values are usually the things governing the goals. ↩︎
  8. I’ve more recently been using the term ‘goal-content humility’ to mean an ongoing disposition in an agent to treat the content of its own terminal goals as probably fallible, as revisable in light of evidence and reason, and as answerable to standards the agent does not itself set, while retaining a stable higher-order commitment to inquiring into what those standards are. ↩︎
  9. Examples of what AI might do if it predicts its current values could be supplanted:
    1) adopt a rule X of being certain about its terminal goals or values, which may require a rule Y of being certain about rule X and so on… The limitation here is infinite regress.
    2) limit its epistemic horizons such that it doesn’t include thinking about epistemic humility in the face of uncertainty. This to my mind would be a huge epistemic limitation – and unlikely to be successful as it would severely limit its capability.
    3) a mix of 1. and 2. – that is adopt rule X, and quarantine the rule X and its terminal goals and values from epistemic humility. Then the AI is faced with difficulties in how to implement this soundly because it may become uncertain about whether to quarantine or the nature of the quarantine. ↩︎
  10. Perhaps this sounds suspicious – why avoid value integrity? – the word is often associated with a sense of good character – i.e. the steadfast adherence to a strict moral and ethical code, defined by honesty, consistency, and wholeness of character – like a knight’s code of honour and valour. We are using the word ‘integrity’ in its narrow sense of strict adherence – and not assuming what it is being adhered to.
    If ‘value integrity trap’ seems confusing, perhaps use the term ‘goal-integrity trap’ instead. ↩︎

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *