Post #2181
258
Подход к оценке агентов через анализ трасс выполнения https://arxiv.org/abs/2605.08545
arXiv.org Log analysis is necessary for credible evaluation of AI agents Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated or deflated by shortcuts and benchmark...