papersTODAY 04:00 UTC
URCHIN: A Horizontal Spiking Language Model for Data-Constrained Pretraining
Researchers introduce URCHIN, a spiking neural language model designed for pretraining on small, developmentally plausible text corpora, as targeted by the BabyLM challenge. The work argues that most existing language models ignore the biological properties of the neural circuitry that underlies human language acquisition. It tests how much language a model can learn from child-scale data instead of internet-scale datasets.