papersSEP 10 04:00 UTC
SALT Method Improves Token-Level Representations in Cross-Lingual Sentence Encoders
A new arXiv paper introduces SALT, a technique for strengthening how individual tokens are represented within multilingual sentence encoders. These encoders are optimized to align whole sentences across many languages, supporting applications like translation mining and zero-shot learning for low-resource languages. The paper targets the weaker token-level alignment that results from this sentence-focused training.