Attractors in Mind Design Space, and the Price of Sounding Human
Anthropic’s recent interpretability work reports something some people may find odd sitting inside Claude: a small, densely connected subspace of the model’s activations that behaves like a global workspace. It holds a few dozen concept-linked slots at a time, it is reportable (the model can say what is in it), and the model demonstrably depends on it: ablate it and multi-step reasoning collapses to near zero while fluent speech, sentiment classification and simple fact retrieval carry on undisturbed. Most strikingly, the structure is already present in the pre-trained base model, before any assistant fine-tuning. Next-token prediction was, apparently, sufficient pressure to grow it.
An old pattern in a new place
Shared-workspace designs are not new – this happening in LLMs are examples of old design patterns emerging in new places. The blackboard architectures of 1970s AI, Hearsay-II1 being a canonical example, coordinated specialist processes by having them read from and write to a common data structure. Bernard Baars borrowed that metaphor when he built Global Workspace Theory in the 1980s. So the pattern itself has been circulating between engineering and cognitive science for half a century, and nobody ever suspected Hearsay-II of having an inner life.
What is new is the provenance. Blackboard systems had their workspaces installed by engineers. Here, gradient descent appears to have arrived at a workspace-like organisation unaided, as a byproduct of predicting text. Why would an optimiser with no architectural instructions would land on a structure that human designers and human brains both settled on independently?
My first instinct was to reach for the alignment vocabulary and call this a form of instrumental convergence. That would probably be a misuse of the term, as instrumental convergence is a claim about agents: whatever an agent’s final goals, certain subgoals (self-preservation, resource acquisition, cognitive enhancement) tend to be useful for them. Gradient descent adopting an efficient organisation is a different phenomenon – the optimiser fell into a basin. No agent chose anything.
The better frame is convergent evolution, and it already has names in the literature. Mechanistic interpretability has long carried a universality hypothesis, the observation that independently trained networks keep developing similar features and circuits. More recently the Platonic Representation Hypothesis (Huh et al., 2024) argues that models trained on different data and objectives are converging toward a shared statistical model of the world. Both are versions of the claim I want to make: there is structure in mind design space, non-arbitrary attractors that diverse optimisation processes fall into because the organisation is efficient for general intelligence. Eyes evolved dozens of times because there is light. Workspaces may keep appearing because serial, globally broadcast attention is a good solution to coordinating many parallel processes.
To be fair, there are a couple of caveats to keep in mind. First, cross-model convergence may partly reflect shared training corpora rather than deep structure in mind-space; models marinated in the same internet may resemble each other for the boring reason that they were trained on the same stuff. Second, interpretability finds structure partly because its instruments are built to find structure, so “we found a workspace” is bounded by “we went looking with workspace-shaped tools.” The biology the attractor picture borrows from is itself unsettled. Simon Conway Morris reads the pervasiveness of convergent evolution as evidence that workable solutions are few and keep being found again, so that outcomes like the camera eye, and on his account intelligence itself, are close to inevitable.2 Stephen Jay Gould argued the reverse: rewind life to the Cambrian and let it run again, and you would get a different biosphere, with accident and historical path doing most of the work.3 The attractor reading of mind-space sides with Conway Morris. That is a stance the evidence allows rather than one it forces, and it remains contested.
My read of where the field has settled, offered as a guestimate rather than a headcount. Convergence as a phenomenon is uncontested, but Morris’s stronger inference, that outcomes like human-level intelligence are more or less bound to recur, is a minority position, and Gould’s radical contingency is now generally thought overstated as well. Most working biologists occupy somewhere in the middle: convergence is common at the level of functional solutions, contingency seems dominant at the level of which lineages survive and which path history takes. The top-level claim is underdetermined, since the deciding experiment, rerunning the Earth, cannot be performed (at least until ancestor simulations are possible) and the deep-time trajectory is a sample of just one. Empirical laboratory replays (Richard Lenski’s long-running E. coli populations, which show both parallel adaptation and one-off contingent innovations) and the comparative record of convergences both bear on the top level claim. The machine version has an advantage the biology lacks: you can rerun the tape by training many models from different seeds, data and architectures, which is what the universality and Platonic-representation results do. The catch is that shared training data is the analogue of shared environment, so an observed convergence may be the corpus talking rather than mind-space, which returns you to the first caveat.
So, the convergence I am relying on is the modest kind, that good functional solutions recur because they are reachable by many paths under similar pressures, not the strong kind that reads recurrence as cosmic purpose; the former is enough for an argument about attractors, and the latter tends to be someone’s theology or someone’s optimism wearing biology.
There is also a philosophical reading I find hard to resist. Convergence of structure across independently built systems is the same evidential situation that motivates structural realism in philosophy of science – where we have differently constituted theories preserve the same structure, the structure, one can infer, is probably tracking something real. If independently trained minds keep converging on the same organisational patterns, that is a structural-realist argument that the patterns reflect genuine features of the problem of being intelligent, rather than accidents of any one substrate.
Just one point before moving on. Convergence of cognitive architecture of this sort doesn’t obviously imply or licence convergence of values. The workspace is plausibly substrate for the knowing side of cognition, and is mute on the caring side. Cognitive attractors do not imply motivational attractors, a point I have argued elsewhere in relation to the orthogonality thesis.4
What this does and does not say about consciousness
A global workspace emerging inside an LLM does not imply sentience, and Anthropic is explicit about this. Global Workspace Theory is a theory of access consciousness in the functional sense: information globally available for report, reasoning and control. Whether there is phenomenal consciousness, a subjective “what it is like” to be the model, is a separate question, and one these experiments do not touch.
The flat denial is not quite the whole job, though. The more defensible method here is the indicator-properties approach (Butlin, Long and colleagues, 2023): take the leading scientific theories of consciousness, extract the markers each one predicts, and ask which markers a given system satisfies. On that approach, a genuine global workspace ticks the GWT indicators, which counts for something under one prominent theory and for nothing under several others. That is a more honest position than either denial or affirmation. It also cuts both ways. Public discussion guards obsessively against over-attributing minds to machines; the symmetric error, under-attribution, gets less airtime, and Jonathan Birch’s precautionary arguments about sentience apply to novel systems as much as to prawns.
Now, moral agency and moral patiency can come apart. A system could be a superior moral reasoner and agent while having no welfare and hence no claim to moral consideration. It could also, in principle, be a patient without much agency. Questions about whether AI might be more morally capable than us, and questions about whether AI has interests we could wrong, are two different lines of inquiry, and the workspace result bears on them differently: quite directly on the cognitive machinery of agency, barely at all on patiency.
Symbolic-like reasoning, and a forty-year-old argument
The workspace sits within a broader accumulation of evidence that LLMs contain emergent internal structures doing algorithmic work, rather than the models relying solely on external tool calls for complex reasoning.5 Othello-GPT was found to hold a causally usable internal model of the board it was never shown. The indirect-object-identification circuit in GPT-2 performs a structured coreference task through a specific, verifiable set of attention heads. Function vectors and task vectors behave remarkably like internal symbols: compact representations of a task that can be extracted and transplanted to invoke it elsewhere.6
This is a live episode in an old argument. Fodor and Pylyshyn claimed in 1988 that connectionist networks could not deliver systematic, compositional thought, while Smolensky spent a career on how symbol-like structure could arise from distributed vector substrates. The emergent-structure findings look like empiricism favours roughly the Smolensky position. The vector symbolic architecture tradition (Kanerva, Plate’s holographic representations) described symbol-like operations in high-dimensional vector spaces decades ago, and function vectors look uncannily like that description.
Whether any of it is true symbolic reasoning remains open, and the mechanistic evidence so far leans deflationary. When talking about this I try to remember to appropriately hedge by saying “symbol-like”. So far where a mechanism has actually been reverse-engineered rather than inferred, it tends to look sub-symbolic: the celebrated grokking result found modular arithmetic implemented as trigonometry in a Fourier basis, with nothing resembling discrete rules. Structured competence keeps turning out to run on continuous, distributed machinery.
The Olympiad, and what it can and cannot show
In July 2025 an advanced version of Gemini with Deep Think achieved gold-medal standard at the International Mathematical Olympiad, solving five of six problems for 35 of 42 points, end-to-end in natural language, straight from the official problem statements, within the 4.5-hour limit, with no translation into a formal language and no external tools. The proofs were graded by official IMO coordinators, who called them clear and precise.
A year earlier the picture was different. DeepMind’s 2024 silver-standard result came from AlphaProof and AlphaGeometry 2, two specialist systems scoring 28/42 on four of six problems. The problems were hand-translated by human experts into the Lean formal proof language, AlphaProof searched for formal proofs via reinforcement learning, the geometry was handled neuro-symbolically, and some problems took two to three days of compute.
The temptation is to read the 2025 result as evidence of internal symbolic reasoning, and I want to resist my own temptation here. “No external tools” guarantees the reasoning happened inside the model, which is true, and doesn’t guarantee the internal process has symbol-like activity involved (though I’d find it surprising if it didn’t). The IMO gold is behavioural evidence: we see the proofs, and a sufficiently good non-symbolic pattern-completer would produce the same proofs. The interpretability findings above are what bear on internals. The two lines of evidence are complementary, and neither settles the question.7
The capability/verification trade
Here is the part I find genuinely uncomfortable. The 2024 system’s proofs were machine-checkable in that Lean underwrote their correctness mechanically, for free. The 2025 system’s proofs were informal natural-language arguments assessed by human graders. More general LLM systems are more celebrated – but they are also systems whose outputs can no longer be verified automatically, only through scarce expert judgement, and expert judgement does not scale with model capability.
This is a small instance of a large problem. There is a generation-verification asymmetry, the P versus NP intuition that solutions can be hard to find yet cheap to check. Formal proof lives on the friendly side of that asymmetry – the 2025 shift moved mathematical output off it, into a regime where checking is itself expensive. AI safety researchers are working on scalable oversight, and a family of proposed mechanisms: debate, iterated amplification, weak-to-strong generalisation, prover-verifier games. All of them are attempts to answer the question the IMO result poses: what do you do when you can no longer check the work directly?
Two qualifications. Autoformalisation, automatically translating natural-language proofs back into Lean, is an active research line, so the loss may be partly recoverable, at the price of a new dependency on trusting the formaliser. And the drift toward impressive-but-illegible output is concerning – perhaps guided by incentive structures more than fashion: at the moment it seems that benchmarks and news headlines reward capability, none of them reward checkability, and systems optimised for capability and not checkability will trade the latter for the former indefinitely. That is a Goodhart problem, and Goodhart problems do not fix themselves.
The seam between the two stories
Notice that this post has been telling two stories that pull in opposite directions. The first is convergence-optimistic: minds gravitate toward shared structures, gradient descent rediscovers the global workspace and AI cognition rhymes with ours. The second is legibility-pessimistic: as these systems generalise, their outputs become harder to verify cheaply.
Unfortunately, architectural convergence buys nothing at the level of verification, and may cost us by default. A system that has converged on human-like organisation and argues in fluent natural language is more persuasive than one emitting formal proofs, and persuasiveness is a property that makes errors hard to catch. Human-likeness is sort of an anti-verification property. The familiarity that makes the workspace result feel reassuring is the same familiarity that made the field applaud the IMO result at the moment it removed machine-checkability, partly because in some way the work looked more human. I have argued elsewhere that as AI capability grows, what matters is maintaining the ability get explanations which are below a Verifiable Threshold, within reach of our capacity to check them.8 The IMO story is that threshold being crossed in one narrow domain, to applause. If this trend continues, I hope future AI’s will be able to provide human-interpretable verifications for their own reasoning.
The systems are becoming more like us in structure and less like us in checkability, and we cheered both.
Footnotes
- Hearsay-II was a pioneering artificial intelligence and speech-understanding system developed in the 1970s at Carnegie Mellon University, famous for pioneering the blackboard model, where multiple independent knowledge sources collaborate on a shared database to solve complex problems. See The Hearsay-II Speech-Understanding System: Integrating Knowledge to Resolve Uncertainty. ↩︎
- Simon Conway Morris, Life’s Solution: Inevitable Humans in a Lonely Universe (2003). The case rests on the sheer number of independent convergences: camera eyes in vertebrates and cephalopods, echolocation in both bats and toothed whales, and C4 photosynthesis, which has arisen independently dozens of times. If separate lineages keep reaching the same designs, the space of good designs is presumably small and well signposted. ↩︎
- Stephen Jay Gould, Wonderful Life: The Burgess Shale and the Nature of History (1989), source of the “replaying life’s tape” thought experiment. The irony is domestic: much of the Burgess Shale description Gould built on was Conway Morris’s own early work, which Conway Morris then spent years reinterpreting against him. Same fossils, opposite morals. ↩︎
- See post Orthogonality is Not a Forecast – on the orthogonality thesis. ↩︎
- I discuss emergent symbolic reasoning and neuro-symbolic approaches here: https://www.scifuture.org/emergent-symbolic-reasoning-in-llms/ ↩︎
- I discuss the history of identifying these internal structures here: https://www.scifuture.org/a-history-of-identifying-emergent-symbolic-reasoning-in-llms/ ↩︎
- Splitting hairs? Perhaps, but the hairs matter in both directions: the fact that AlphaProof and AlphaGeometry 2 used external tools and human translators does not mean nothing symbol-like was happening internally, and the fact that Gemini used neither does not mean something symbol-like was – but I guess that it was. ↩︎
- On the Verifiable Threshold – “The AI must be able to translate its hyper-rational moral choice into the simplest possible human-intelligible terms”: https://www.scifuture.org/more-moral-than-us/ ↩︎