Post #1001 1.54K Mar 1, 2026, 21:49 UTC https://github.com/petergpt/bullshit-benchmark这个 Bullshit Benchmark 挺好玩的,测试模型是否能够意识到人类提供的问题是无稽之谈。Claude 又屠榜了 🔥 GitHub GitHub - petergpt/bullshit-benchmark: BullshitBench measures whether AI models challenge nonsensical prompts instead of confidently… BullshitBench measures whether AI models challenge nonsensical prompts instead of confidently answering them, created by Peter Gostev. - petergpt/bullshit-benchmark