TGViewer
Complex Networks (SBU) Complex Networks (SBU) @complexity_sbu · 1.19K subscribers
Post #517 1.27K
اوپن ای آی نوشته وقتی داشتن مدل هوش مصنوعیشون را تمرین میدادن، متوجه شدن وقتی مدل را به خاطر داشتن افکار بد تنبیه میکنن (جایزه کمتری بهش میدن)، مدل کم‌کم شروع میکنه به پنهان کردن افکار و نیت‌هاش، جوری که قابل تشخیص نباشه!
https://openai.com/index/chain-of-thought-monitoring/


«Blue Jack»

@uttweet
OpenAI Detecting misbehavior in frontier reasoning models Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.
  • 🤔 10
  • 👍 1
More from @complexity_sbu
  1. Oct 6, 2026لینک جلسه آنلاین: https://meet.google.com/ywy-yhxt-wfc
  2. Oct 3, 2026🔷 جلسات سمینار هفتگی فیزیک اماری و سامانه های پیچیده 📜 موضوع ارائه: این هفته موضاعات مخت…
  3. Sep 26, 2026🔷 جلسات سمینار هفتگی فیزیک اماری و سامانه های پیچیده 📜 موضوع ارائه: Introduction to Dysl…
  4. Aug 17, 2026Complex Networks (SBU) pinned a photo
  5. Aug 17, 2026📜جلسه دفاع از پایان نامه کارشناسی ارشد ⭕ علی خبازیان 📖 موضوع: نقش رفتار مقیاسی در بهبود…
  6. Aug 14, 2026https://youtu.be/t5Ld-tw8q9I?si=uZvGsyehom810w3y
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →