papersSEP 10 04:00 UTC
Sparse Autoencoder Phase Diagram Shows Dominant Diffuse Phase
A new arXiv preprint maps out a phase diagram for sparse autoencoders, the tools widely used to pull interpretable features out of neural network activations. The work reports that a diffuse phase dominates the diagram, which helps explain why distinct features can be absorbed or merged when feature co-occurrence is systematic. It also engages with the MAIS-O43 open problem on controlling such feature merging.