For years, the honest answer to 'how does this model work?' was 'we don't know.' That is beginning to change — and the answers are stranger than expected.

Interpretability research — the project of understanding what happens inside neural networks — has produced a string of genuinely surprising findings in the past 18 months. Features that were expected to be diffuse turn out to be localized. Representations that seemed inscrutable turn out to have geometric structure that humans can recognize.

None of this means we understand these systems. But it means we are beginning to develop methods for asking better questions about them.