Transformers struggle to generalize to tasks they were not explicitly trained on.
Instead, researchers propose in 2026 that it is the job of the harness to generalize through composition.
Researchers find that well-designed harnesses form a quotient set over task trajectories, meaning their individual LLM calls can see structurally “similar” tasks as near-identical, token-for-token.
Harnesses can effectively generalize for the Transformer during training, without relying on any intrinsic generalization capability from the model.
Post #4368
650