papersSEP 11 04:00 UTC
Paper Proposes Output Embedding Centering to Curb LLM Pretraining Instability
A new arXiv preprint introduces a method called output embedding centering aimed at reducing output logit divergence, a form of training instability that tends to appear late in large language model pretraining. The authors position it against commonly used mitigations such as z-loss. The work is a research contribution rather than a released model or product.