TGViewer
RIML Lab RIML Lab @rimllab ยท 3.25K subscribers
Post #257 4.49K
๐Ÿ” LLM Faithfulness Journal Club

โœ… This Week's Presentation:

๐Ÿ”น Title: Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

๐Ÿ”ธ Presenter: [Farzan Rahmani](https://www.linkedin.com/in/farzan-rahmani-51128b201)

๐ŸŒ€ Abstract:
When large language models explain their decisions, their explanations may sound convincingโ€”but do they faithfully reflect the model's true reasoning? This paper presents a comprehensive counterfactual analysis of self-explanation faithfulness across 75 models from 13 model families. It investigates the tradeoff between concise and comprehensive explanations, introduces two new evaluation metrics (phi-CCT and F-AUROC), and studies how explanation verbosity influences faithfulness measurements. The results reveal a clear scaling trend: larger and more capable language models consistently produce more faithful self-explanations, providing valuable insights into the relationship between model scale, explanation quality, and AI safety.

๐Ÿ“„ Paper: [Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations](https://arxiv.org/abs/2503.13445)

Session Details:

* ๐Ÿ“… Date: Wednesday (ฺ†ู‡ุงุฑุดู†ุจู‡)
* ๐Ÿ•‘ Time: 2:00 - 3:00 PM
* ๐ŸŒ Location: Online at http://vc.sharif.edu/ch/rohban

We look forward to your participation! โœŒ๏ธ
More from @rimllab
  1. Sep 27, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Exploiting LLM Quantizatโ€ฆ
  2. Sep 20, 2026we are looking for Teaching Assistants to join the Multi-Agent Reinforcement Learning (MARโ€ฆ
  3. Sep 20, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Robust Machine Unlearninโ€ฆ
  4. Sep 15, 2026๐Ÿ”˜ Open Research Position: Machine Unlearning ร— Model Quantization We are looking for motiโ€ฆ
  5. Sep 13, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Catastrophic Failure ofโ€ฆ
  6. Aug 30, 2026๐Ÿ“ข Join the IABI TA Team! ๐Ÿฉป The Intelligent Analysis of Biomedical Images (IABI) course iโ€ฆ
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook โ†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 โ†’