T1 Terminal Agent RL
Covered by
arxiv.org
Timeline
- T1 trains a 122B Mixture-of-Experts terminal agent with RL across 300-turn tasks
T1 is a rare full training recipe, not just a model card, for reinforcement learning that survives hundreds of shell turns, and its disjoint from Terminal-Bench 2.1 is the test of whether such gains are real capability or benchmark memorization.
1 source · excerpt-only, thin-sourcing
All sources (1)
- arXiv cs.LG2026-09-11