papersSEP 12 04:00 UTC
T1: 122B Mixture-of-Experts Model Trained with RL for Terminal Agent Tasks
Researchers released T1, a 122-billion-parameter Mixture-of-Experts model trained via reinforcement learning to act as an agent in terminal environments. The work targets long-horizon workloads such as software development and scientific research, where sustained multi-step command-line use matters. It is presented as part of a broader shift in agent design away from short, single-turn interactions.