Five physical tasks, not a production result

Researchers from RAI Institute, Technical University of Munich and ETH Zurich have posted a method that uses sample-based model predictive control in simulation to create successful trajectories, then uses them to start sparse-reward reinforcement learning. They report transfer to an arm-equipped Boston Dynamics Spot quadruped and a Unitree G1 humanoid across five physical tasks.[1,2]

The physical work is specific: Spot reach, box pushing, tire uprighting and tire rolling, plus G1 box pushing. Those are useful whole-body contact tasks, and the project page supplies four clips of runs. They do not establish factory readiness, fleet operation or a general-purpose policy; the work remains an author-reported research result under submission.[1,2]

The published scorecard is still simulated

The paper's numerical gains stop before the hardware tests. Near-perfect success, more than 50 percent faster completion on some tasks and 11 to 45 percent lower duration variance refer to simulated training and evaluation. Neither the paper nor its project page gives a physical trial count or hardware task-success rate. Video confirms that a run occurred; it cannot measure the reliability of repeated operation.[1,2]

There is a second boundary. The authors say their real-world runs rely on state-based information, and that unstructured or outdoor use would require vision-based training or distillation. Unitree's G1 documentation lists a depth camera and three-dimensional lidar, but the paper does not say those sensors supplied the tested policy. This is sim-to-real transfer across two bodies, not a demonstrated perception-to-action system.[1,3]

The claimed workflow gain sits upstream of the robot. The authors say the simulated controller can be tuned in minutes, then generate roughly one million samples an hour on a single RTX 5090; they estimate about four graphics-processor hours to seed a task. That can reduce reward-design iteration for labs with an accurate simulator, but it does not yet measure physical commissioning time, operator time or maintenance cost.[1,2]

The decision change is narrower but real: simulation-generated expert data may remove some manual reward design for well-defined contact tasks, while the physical evidence still cannot price the intervention, sensing or failure burden. The next useful result is a preregistered physical scorecard that counts attempts, resets, failures and sensing conditions across new objects and sites.[1,2,3]