SPICE Method Uses Clustering to Interpret Polysemantic Neurons
A new arXiv paper introduces SPICE, a technique that applies clustering to explain polysemantic features in neural networks, where individual neurons respond to multiple unrelated concepts. The approach aims to make functional interpretation of such neurons clearer for interpretability research. The abstract describes the work as a simple method for clustering-based explanation of these overlapping activations.