papersTODAY 04:00 UTC
Study Scales JugnuLM Language Models From 53M to 110M Parameters
A new arXiv paper examines how a fixed sub-150M pretraining recipe behaves when model size grows from 53.5M to 109.7M parameters. Both models use a Qwen3-style decoder with grouped-query attention, RoPE, SwiGLU, RMSNorm, QK-Norm and a z-loss, trained on FineWeb-Edu data, so only scale and depth differ. The work compares the 53M and 110M variants to isolate the effects of added capacity in this small-model regime.