Post #835 200 Nov 19, 2024, 17:00 UTC curious approach, something akin to an in-context LoRA model finetuning (plus a lot of other tricks)https://arxiv.org/abs/2411.07279 arXiv.org The Surprising Effectiveness of Test-Time Training for Few-Shot Learning Language models (LMs) have shown impressive performance on tasks within their training distribution, but often struggle with structurally novel tasks even when given a small number of in-context...