ok chat i've made a reasoning LLM benchmark that can't be saturated (inspired by AidanBench), what models should I test?
currently I test on 200 easiest tasks solvable with pen and paper in seconds but the problem is NP complete and the number of tasks is infinite
Post #1292
445

- 🐳 1