Technical Reports

XiaomiMiMo Releases Agentic RL Training Environment and Dataset

Overview of XiaomiMiMo/MiMo-V2.6-RL-oss and the verl fork repository for LLM agent reinforcement learning, detailing ...

Technical Reports

SmolDataEnvs: RL Tasks for Small Model Optimization

Explore SmolDataEnvs, a dataset with 5.5K+ RL tasks for hill-climbing optimization of small models in code and data s ...