Following MiMo-V2.6’s live RL run. Useful to see learning curves alongside sampling stats, queues, costs, and restarts!
The2B tokens per step is a substantial workload. What I’d like to understand is the compute allocation: at a fixed total budget, what produces the most improvement on unseen tasks?
A few questions this run brings up:
- More prompts, more rollouts, or better feedback?
The dashboard shows ~1,568 prompts × 16 rollouts per batch. Covering more tasks and exploring more trajectories per task serve different purposes. Grader compute adds another choice: when is it worth generating fewer trajectories and spending more on evaluating them? The best allocation may shift as the policy improves.
- What does agentic in-group credit assignment actually do?
Test cases and rubrics provide different kinds of feedback. I’m curious how the grader compares attempts within a group, and whether it assigns credit to whole trajectories, stages, or individual actions. The ablation I’d like to see is whether additional grading cost pays for itself through more efficient learning.
- Critic/RLOO choices also affect scheduling.
RLOO uses sibling rollout returns as a baseline, a learned value function predicts return from the current state.
If advantage computation requires the full group, slow trajectories can delay it even in an asynchronous pipeline. A value-based estimator that permits independent trajectory processing could reduce that dependency, but adds critic inference, training, and estimation costs. Retaining an RLOO component may retain the group dependency too.
This is a design question, not a claim about MiMo’s implementation. The dashboard’s critic/* metric names alone don’t establish that it uses a learned critic.
- Throughput needs to be read alongside policy lag.
More trajectories per hour don’t necessarily mean faster learning. By the time a long-running task finishes, its behavior policy may be several versions behind. How the system handles that gap - and avoids systematically favoring short tasks - matters alongside raw throughput.
The comparison I’d most like to see: time and total cost to reach the same held-out performance, accounting for rollouts, grading, environments, and any critic.
Appreciate the MiMo team making the run visible. There’s a lot to learn from the intermediate behavior of a training system.
https://mimo.xiaomi.com/rl/
Le Critique:
https://arxiv.org/abs/2608.16739