papersSEP 10 04:00 UTC
Study traces word-by-word embedding trajectories of Vietnamese legal headlines
A cs.CL paper on arXiv examines how dense retrieval models build question representations incrementally by encoding 2,144 held-out headlines from the Thu Vien Phap Luat Vietnamese legal library one word at a time. The authors track these embedding trajectories using Nemotron-3-Embed (8B/1B) and Qwen3-Embedding (8B/0.6B) models, offering insight into how vector representations evolve as words arrive. The v2 listing replaces the earlier version of the paper.