papersSEP 10 04:00 UTC
arXiv study proposes directly reading and writing transformer internals
A newly posted preprint investigates how many components inside a transformer actually determine a given token prediction, measuring the signed contributions of individual units and channels to the final logits. It reports that a single output can depend on anywhere from thousands to hundreds of thousands of components, whose effects partly offset one another. The work, titled Through the Looking Glass, presents methods for directly reading out and editing transformer internals.