Physical intelligence introduced a new model π*0.6
π*0.6 can more than double throughput over a base model trained without RL, and can perform real-world tasks: making espresso drinks, folding diverse laundry, and assembling boxes.
Team trained a general-purpose value function on all of own data, which tells the π*0.6 VLA which actions are good or bad. By asking π*0.6 to produce only good actions, researchers get better performance. Team call this method Recap.
π*0.6 can then collect more autonomous data, which can be used to further train the value function and further improve π*0.6.
During autonomous data collection, a teleoperator can also intervene and provide corrections for significant mistakes, coaching π*0.6 further.
Quantitatively, training π*0.6 with RL can more than double throughput (number of successful task executions per hour) on the hardest tasks and cut the number of failures by as much as a factor of two.
Post #3796
780
- 🔥 5
- 🥰 3
- 👏 3