Attractors in Mind Design Space, and the Price of Sounding Human
Anthropic’s recent interpretability work reports something some people may find odd sitting inside Claude: a small, densely connected subspace of the model’s activations that behaves like a global workspace. It holds a few dozen concept-linked slots at a time, it is reportable (the model can say what is in it), and the model demonstrably depends…