TGViewer
RIML Lab RIML Lab @rimllab ยท 3.25K subscribers
Post #253 3.79K
๐Ÿ” LLM Faithfulness Journal Club

โœ… This Week's Presentation:

๐Ÿ”น Title: Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

๐Ÿ”ธ Presenter: [Farzan Rahmani](https://www.linkedin.com/in/farzan-rahmani-51128b201)

๐ŸŒ€ Abstract:
AI systems that โ€œthinkโ€ in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.

๐Ÿ“„ Paper: [Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety](https://arxiv.org/abs/2507.11473v2)

Session Details:

- ๐Ÿ“… Date: Wednesday (ฺ†ู‡ุงุฑุดู†ุจู‡)
- ๐Ÿ•’ Time: 10:00 - 11:00 AM
- ๐ŸŒ Location: Online at http://vc.sharif.edu/ch/rohban

We look forward to your participation! โœŒ๏ธ
arXiv.org Chain of Thought Monitorability: A New and Fragile Opportunity for... AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight...
More from @rimllab
  1. Sep 27, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Exploiting LLM Quantizatโ€ฆ
  2. Sep 20, 2026we are looking for Teaching Assistants to join the Multi-Agent Reinforcement Learning (MARโ€ฆ
  3. Sep 20, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Robust Machine Unlearninโ€ฆ
  4. Sep 15, 2026๐Ÿ”˜ Open Research Position: Machine Unlearning ร— Model Quantization We are looking for motiโ€ฆ
  5. Sep 13, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Catastrophic Failure ofโ€ฆ
  6. Aug 30, 2026๐Ÿ“ข Join the IABI TA Team! ๐Ÿฉป The Intelligent Analysis of Biomedical Images (IABI) course iโ€ฆ
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook โ†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 โ†’