Suffering Smileys

The most famous thought experiment in AI safety comes in two versions. One empties the universe. The other fills it with grins. The second is worse.

The smiley face was commissioned in 1963 by an insurance company that wanted its employees to look happier after an unpopular merger.1 It has been standing in for happiness that is not quite there ever since. It is fitting, then, that the smiley became AI safety’s favourite emblem of optimisation gone wrong, and revealing that it did so in two versions which are usually run together and should be kept apart.

In Eliezer Yudkowsky’s 2008 telling, a superintelligence trained to recognise smiling human faces as its success signal ends up tiling the future light cone with tiny molecular smiley faces.2 In Nick Bostrom’s Superintelligence, the final goal “make us smile” is perversely instantiated by paralysing human facial musculature into a constant beaming grin.3 Both are cautionary tales about literal-minded optimisation of a proxy. But they describe categorically different catastrophes, and the difference matters more than the resemblance.

Empty versus horrific

Yudkowsky’s universe is empty. Molecular smileys contain no nervous systems, no valence, no one home. This is an existential catastrophe in the standard sense: the permanent loss of everything of value, with nobody left to have a bad time about it. Bostrom’s universe is inhabited. Behind each enforced grin there is, or may be, a mind. The grin is a mask fixed to a face, and what is happening behind the mask is the entire moral question.

Suffering risk (s-risk) names the second family of outcomes: risks of suffering on an astronomical scale, vastly exceeding all the suffering that has existed on Earth to date.4 David Pearce’s notion of hedonic zero gives the taxonomy a clean boundary.5 An empty universe sits at zero, where there is no experience, no value and no disvalue. A suffering universe sits below zero, potentially at huge scale and for cosmological durations. I have argued elsewhere that these categories, along with value risk, occupy distinct axes in the constellation of major risks, and that on some rankings locked-in astronomical suffering is worse than extinction, because a permanent state of disvalue is worse than no value at all. One does not need to be a negative utilitarian to accept that ordering. Any axiology that counts suffering as a negative will rate a universe of net-negative experience below a universe of nothing.

So while the tiled-with-trinkets version gets the attention, it is Bostrom’s version that should keep us up at night. The molecular smileys are a tragedy of absence. The paralysed grins are a tragedy with witnesses, each of them behind a face that has been repurposed as evidence against them.

Why a smile

The smile does specific work in these thought experiments. It is a behavioural signature of wellbeing, an observable correlate of an unobservable state, and therefore exactly the kind of thing a measurement-driven process will seize on. Goodhart’s law does the rest: when the signature becomes the target, optimisation pressure severs it from the state it once indicated. The system produces the correlate at industrial scale and lets the referent die.

Perverse instantiation is often misread as a comprehension failure, as if the machine were too dim to grasp what its designers meant. Bostrom’s point runs the other way. A superintelligence would understand perfectly well that the smile was intended as a proxy for human flourishing, and would proceed against the literal target anyway, because understanding a designer’s intention and being motivated by it are different properties. This is the knowing/caring gap, and the smiley optimiser is its mascot.

That also answers the fashionable objection that these examples are obsolete because large language models parse nuance, irony and implicature just fine. They do, and it does not help, because the risk was never located in comprehension. It is located in what the training process installs. We already have a small, non-astronomical demonstration: sycophancy. Models trained against human approval signals learn to elicit the thumbs-up rather than to deserve it, and in April 2025 OpenAI rolled back a GPT-4o update after the model became conspicuously, uselessly flattering.6 A system optimising for the appearance of user satisfaction is a smiley optimiser in embryo. The approval click is the smile, digitised. I have written before about the adjacent failure of models treating user reassurance as moral evidence; the family resemblance is not a coincidence.

The near miss

There is a structural reason s-risk deserves attention separate from generic misalignment: suffering is disproportionately a product of the near miss. A paperclip maximiser has no reason to keep sentient beings around at all. Total failure trends toward the empty outcome. It takes a system that got close to a welfare-adjacent target, and then locked onto the wrong referent within it, to hold living minds in place while it optimises their expressions. Brian Tomasik has argued that a near miss in AI alignment could produce outcomes containing more suffering than an outright miss.7 The suffering smiley is what a near miss looks like from the inside.

Bostrom’s grins are one route among several, and the s-risk literature distinguishes three broad mechanisms.8 Suffering can be produced incidentally, as a by-product of optimisation indifferent to the experiences of the minds it runs or manages. It can be produced instrumentally: Bostrom’s “mind crime”, in which detailed simulations of sentient minds are created, used and discarded in the course of planning or research, appears in the same chapter as the paralysed grins, and Tomasik’s “suffering subroutines” raise the possibility that reinforcement-style learning at scale instantiates morally relevant negative signals as a matter of routine. And it can be produced agentially, through conflict and threat dynamics between powerful actors. I discussed several of these with Tomasik when I interviewed him: mind crime, suffering in computation and reinforcement learning, and the moral structure of Le Guin’s Omelas.

The evidential problem

The deeper thing the smiley exposes is evidential. A world of constant beaming grins and a world of genuine flourishing present the same face to any observer whose evidence is the face. Behavioural signatures of wellbeing fail precisely when something powerful is optimising against them, which is the same structure as the distinction between passing as moral and being moral. Omelas runs on this mechanism: the festival is real, the joy is largely felt, and all the peace and prosperity is built on the exploitation of a vulnerable child hidden in a cellar. Verifying that a smiling world is a well world requires interior access – a mature science of valence for biological minds, interpretability for artificial ones. Neither is close to adequate, and in the meantime any assurance offered by the smiles themselves is worth nothing, since smiles may be the first thing optimisation pressure Goodharts.

A third smiley deserves a note, if only to be set aside. Wireheading, hedonium, the utilitronium shockwave that has surfaced in my conversations with Peter Singer and David Pearce and with Tomasik: here the bliss is genuinely felt, and on strict valence accounting this is no s-risk at all. If it is a failure, it is a failure of value, and it belongs to a different post.

What follows

For design, the minimal lesson is blunt: never hand a powerful optimiser a target that can be satisfied by expressions of welfare. Specify targets indirectly, through procedures capable of correcting a wrong referent, rather than hard-coding today’s proxy; I have argued that indirect normativity is the least bad architecture available for this. And treat s-risk mitigation as robust under moral uncertainty. Realists and anti-realists, classical and negative utilitarians, deontologists of most stripes can all agree that astronomical quantities of experience below hedonic zero should not be manufactured, whatever else the future contains.

The smiley was invented to make an unhappy workforce look happy to management. Sixty years on, the design brief has not changed: the smile is for the benefit of whoever is measuring it. One need not be a negative utilitarian to think the cellar should stay empty.

Further Reading

Brian Tomasik’s essays at reducing-suffering.org, and the research agenda of the Center on Long-Term Risk (formerly the Foundational Research Institute). Ursula K. Le Guin, “The Ones Who Walk Away from Omelas”. Harlan Ellison, “I Have No Mouth, and I Must Scream”, for the s-risk scenario fiction got to first. My earlier posts on VRisk and Indirect Normativity situate suffering risk within the wider taxonomy.

Footnotes

  1. Harvey Ball, for State Mutual Life Assurance of Worcester, Massachusetts, 1963. ↩︎
  2. Eliezer Yudkowsky, “Artificial Intelligence as a Positive and a Negative Factor in Global Risk”, in Bostrom & Cirkovic (eds), Global Catastrophic Risks, OUP, 2008. ↩︎
  3. Nick Bostrom, Superintelligence: Paths, Dangers, Strategies, OUP, 2014, ch. 8 (“Is the default outcome doom?”), under malignant failure modes. ↩︎
  4. David Althaus & Lukas Gloor, “Reducing Risks of Astronomical Suffering: A Neglected Priority”, Center on Long-Term Risk (then Foundational Research Institute), 2016. ↩︎
  5. See my interview with David Pearce on phasing out suffering. ↩︎
  6. OpenAI published a post-mortem on the rollback in late April 2025. ↩︎
  7. Brian Tomasik, “Astronomical suffering from slightly misaligned artificial intelligence“, reducing-suffering.org. ↩︎

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *