Post #2269 107 Jan 25, 2026, 14:03 UTC #llm #chess #arc_agi_3 #arc_agi #reasoning #agi https://openreview.net/forum?id=65R1Dbfwzk openreview.net LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs... We introduce LLM CHESS, an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extended agentic...