Value-Content Integrity
In The Superintelligent Will, in discussing ‘goal-content integrity’ Bostrom uses final goals (aka terminal goals) to mean the agent’s ends (as opposed to subgoals/instrumental goals, which are just means). This discussion primarily concerns agents with persistent objectives or learned value representations. It should not be assumed that every present-day AI system possesses stable final goals…