Post #871 184 Mar 17, 2025, 19:56 UTC ^ that was the original idea that led me to write a custom benchmarkhttps://t.me/notatky/1292 Telegram χаотичні нотатки ok chat i've made a reasoning LLM benchmark that can't be saturated (inspired by AidanBench), what models should I test? currently I test on 200 easiest tasks solvable with pen and paper in seconds but the problem is NP complete and the number of tasks is…