arXiv paper proposes brain-network priors for multimodal language models
A preprint introduces the "Platonic brain bridge hypothesis," arguing that large models handling video, audio and text together tend to converge on representations resembling those found in the human brain, with the relationship working in both directions. The authors suggest that brain-like alignment could shift from being only a way to measure models toward an actual design principle for building them. The revised version mentions the hypothesis as an architectural prior, while the earlier cross-listed version frames it around so-called omni models.