Post #4415
833
Chinese instagram Rednote dropped a 280B model and a new RL training algorithm for long-horizon rollouts based on test-time-scaled value estimation with macro-step policy optimization - TEMPO
- ❤ 4
AL All about AI, Web 3.0, BCI @alwebbci · 3.89K subscribers