๐ LLM Faithfulness Journal Club
โ
This Week's Presentation:
๐น Title: Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
๐ธ Presenter: [Farzan Rahmani](https://www.linkedin.com/in/farzan-rahmani-51128b201)
๐ Abstract:
When large language models explain their decisions, their explanations may sound convincingโbut do they faithfully reflect the model's true reasoning? This paper presents a comprehensive counterfactual analysis of self-explanation faithfulness across 75 models from 13 model families. It investigates the tradeoff between concise and comprehensive explanations, introduces two new evaluation metrics (phi-CCT and F-AUROC), and studies how explanation verbosity influences faithfulness measurements. The results reveal a clear scaling trend: larger and more capable language models consistently produce more faithful self-explanations, providing valuable insights into the relationship between model scale, explanation quality, and AI safety.
๐ Paper: [Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations](https://arxiv.org/abs/2503.13445)
Session Details:
* ๐
Date: Wednesday (ฺูุงุฑุดูุจู)
* ๐ Time: 2:00 - 3:00 PM
* ๐ Location: Online at http://vc.sharif.edu/ch/rohban
We look forward to your participation! โ๏ธ
Post #257
4.49K