papersSEP 10 04:00 UTC
Study challenges the narrow-wide-narrow FFN convention in Transformer language models
An arXiv research paper questions why dense Transformers almost universally place most of their non-embedding parameters in narrow-wide-narrow feed-forward networks. Drawing on theoretical and empirical evidence, the authors explore an alternative wide-narrow-wide (hourglass) residual design for these blocks. The work is cross-listed across the cs.AI, cs.CL, and cs.LG categories on arXiv.