Hugging Face turned Claude Code, Codex, Hermes, Pi, and other coding harnesses into RL environments
No changes to the harnesses, no changes to the training code.
Any open model, any task set, fully open source
Same model, same weights: 62% under Mini-SWE-Agent, 33% under Claude Code.
But training inside a real harness normally means reimplementing it as an environment, so most models get trained in a scaffold nobody actually ships.
Post #4518
399