Radical Interpretability
When Interpretability Becomes Mind-Reading Seatbelt Interpretability Most AI interpretability research asks a near-term engineering question: can we understand enough of a model’s internal processing to explain, predict or control important behaviour? Why did it refuse this request? Which internal features mattered? Is it representing a hidden objective? Can we tell whether a safety mechanism is…