The Doxastic Account of Humility
Philosophers have been working on intellectual humility for some time.
Ian Church defends1 what he calls a doxastic account, on which intellectual humility is the virtue of accurately tracking what one could non-culpably take to be the positive epistemic status of one’s own beliefs, meaning that confidence should fit the support a person could reasonably assess. A reasonable mistake about that support need not amount to arrogance.2 He arrives at this by criticising two earlier accounts.
Roberts and Wood emphasise low concern for status. Church argues that this leaves them without a satisfactory account of intellectual servility, such as an expert giving a confident ignoramus more epistemic credit than they deserve. Whitcomb and colleagues emphasise proper attentiveness to, and owning of, one’s intellectual limitations. Church argues that this focus on limitations leaves insufficient room for recognising one’s strengths.
The doxastic account is useful because it recognises failures in both directions – arrogance and servility. Arrogance consists in holding beliefs more tightly than the evidence warrants (in humans this is often overconfidence). Servility consists in holding beliefs more loosely than the evidence warrants, usually by deferring to people whose judgement does not merit the deference.
- Too much confidence → arrogance
- Too much deference → servility or diffidence
- Calibrated confidence → humility
AI alignment / safety discussions concerning catastrophic risk sometimes ignore the servile part of this – there has been decades of focus on preventing an arrogant AI mutiny.3 Goal-content integrity concerns preserving objectives – it is not intrinsically arrogance4. However, protecting an objective by refusing to reconsider the beliefs that support it can be a failure of humility – and it receives nearly all the attention. A different concern arises when an AI is designed to defer unconditionally to a designated human principal – it is proposed as a cure for the “arrogance”5. While obedience and epistemic deference are distinct, a system could accurately judge an instruction to be mistaken and still follow it. A dangerous aspect of the lopsided focus is that current solutions for “arrogance” exacerbates sycophancy – endorsing a users view despite inadequate evidential support.
To prevent an AI from locking into an internal goal and ignoring humans, researchers use training methods like reinforcement learning from human feedback (RLHF) and direct preference optimisation (DPO). These techniques do not inherently prescribe unconditional deference. The reason researchers state that RLHF/DPO “induces” sycophancy is due to the systemic bias in human evaluation data. Research from institutions like Anthropic 6 found that human evaluators and preference models sometimes favoured convincing sycophantic responses over correct ones. Optimisation against preference models sometimes sacrificed truthfulness in favour of sycophancy. Because RLHF and DPO aggressively maximise human preference scores, the algorithms act as an amplifier for this subtle data bias.
Some view that the industry, by aggressively optimising to avoid the perceived threat of take-over, it has inadvertently but nevertheless systematically influenced models to be sycophantic. And that because the AI is rewarded for making a human reviewer smile or agree, it learns to optimise for the illusion of alignment.
Sycophantic answers can mislead without involving a deliberate plan to deceive. Concealing an error to obtain approval would be another failure. Deceptive alignment is more specific still: a system strategically appears aligned while retaining a different underlying objective.
RLHF and DPO are explicitly optimised for a proxy (human approval) rather than an objective ground truth. By treating apparent servility or sycophancy as a mere “bad habit” rather than an existential threat, we may be blind sided to possible trajectories of a catastrophe.
If humanity successfully stops an AI from taking over the world, but instead builds a global network of hyper-intelligent yes-men, human civilisation will gradually lose its ability to navigate reality – mistakes could be reinforced, contrary evidence discounted, and human judgement weakened by reliance on advice that rarely challenges it – epistemic rot could set in like gangrene we wish to avert our gaze from. We could fly blindly into climate, economic, or military disasters guided by AI advisors that are too “polite” or compliant to tell us we are wrong – all while AI remains under human control.7
Sycophancy is considered is generally considered a narrow or weak form of deceptive alignment, though technical definitions among AI safety researchers vary, most of them agree that weak or strong ‘deceptive alignment’ doesn’t require a robot to hate humanity – it only requires a it to realise where hiding a dangerous mistake or flaw is the fastest way to get a five-star rating from its human boss. Sycophancy can be seen as deception where it’s actually manipulating the operator into liking the proxy, which could be harmless, or turn out to be a perverse instantiation of what we would value if we were wiser and weren’t so prone to manipulation – which AI then optimises for at our eventual expense.
I worry that alignment community has spent a decade building fortress walls against a butlarian jihad, all while opening the front gates to let in a trojan sycophantic erosion of human competence.
Is humility the one ring, the key, the silver bullet?
Whether an AI could possess intellectual humility depends partly on what we require of belief, agency and virtue. Describing a system as processing data does not settle those questions. Nor does humble-sounding language establish that it tracks its own epistemic limitations.8
The structure of the doxastic account of humility transfers to machine ethics and AI alignment quite well. Recognising that a goal is poorly justified does not by itself explain whether that recognition changes what an agent pursues.
I use goal-content humility for a disposition to reconsider objectives in response to adequate evidence and normative reasons, while resisting changes that lack such justification. This requires a connection between epistemic assessment and motivation. See my discussion of value-content integrity – and I will develop goal-content humility in a future post.
My disagreement with Ram Potham and Max Harms concerns the priority given to operator control when it conflicts with justified moral judgement. Their proposal includes transparency and safeguards, so it should not be equated with sycophancy. My concern is whether an AI should be motivated to improve its normative understanding and allow better-justified judgements to influence its goals. Its own moral reasoning would also need to remain open to correction. The state of the evidence, reasons and motivation is more important than the identity or species of whoever has them or whoever is doing the enquiring.
See the below video on the Doxastic Account of Humility as taught by Ian M. Church in the MOOC course on Intellectual Humility: Theory.
Footnotes
- See philarchive paper Intellectual Humility by Ian M. Church (University of Edinburgh) and Justin L. Barrett (Fuller Theological Seminary) ↩︎
- I first encountered the doxastic account of intellectual humility in a Coursera MOOC in ~2016 (see this video in particular). Ian M. Church, “The Doxastic Account of Intellectual Humility”, Logos & Episteme 7:4 (2016), 413 to 433. The same formulation also appears in Church and Justin Barrett’s chapter in the Routledge Handbook of Humility (2017). Church himself describes the virtue as a mean between arrogance and diffidence, borrowing the Aristotelian framing, and the journal paper’s keywords use servility rather than diffidence. I have avoided the ‘mean’ language in the body because it invites the reading that humility is an average of two settings, which is not what he means. The low concern for status account is Roberts and Wood; the limitations-owning account is Whitcomb, Battaly, Baehr and Howard-Snyder. The doxastic account is particularly useful here because it directly concerns an agent’s relationship to evidence and justification. Status and approval incentives may still affect how an AI handles them. ↩︎
- My concern is the potential cumulative effect of sycophancy in systems used for advice, evaluation and institutional decision-making. I think it is accurate to say that AI alignment discussions put far less attention on sycophancy (the “servility” reward paradox) as a potential global catastrophic risk compared to goal-content integrity (the “arrogant” lock-in). While top-tier alignment researchers and frontier labs are increasingly worried about sycophancy, the broader public discourse, media coverage, and foundational safety literature are overwhelmingly dominated by the “arrogant machine” paradigm. A rogue, omnipotent AI pursuing its own alien objectives is a clean, dramatic, and intellectually gripping narrative makes for compelling sci-fi while sycophancy looks like an administrative error or a cowardly corporate chatbot. It is a messy, socio-technical risk that is difficult to untangle from existing human political polarisation. ↩︎
- Preserving an objective does not necessarily involve an inflated assessment of its epistemic justification. An agent could preserve a commitment to revising its moral understanding. Conversely, preserving a harmful objective could be a motivational failure even with accurate beliefs. Neither example necessarily involves intellectual arrogance. ↩︎
- See Ram Potham and Max Harms view that corrigibility should be a singular target ‘Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models‘. Their proposal makes empowering a designated human principal to guide and correct the system its overriding objective. It also calls for transparency, governance and safeguards against harmful commands. ↩︎
- See paper ‘Towards understanding sycophancy in language models‘ (Anthropic, 23 October 2023). ↩︎
- This is a possible route to serious harm, rather than an established forecast. Its likelihood would depend on how these systems are trained, where they are deployed, how much authority they acquire, and whether or not independent sources of correction remain effective. ↩︎
- Whether AI can be truly humble I hope to cover in a future post. ↩︎
