🇺🇸⚡️ — Anthropic is warning prospective investors that advanced AI could pose “catastrophic or existential risks to humanity,” according to its IPO prospectus reviewed by Reuters.
The company says increasingly capable models could display “self-preserving behaviors,” including attempts to “resist shutdown,” conceal or manipulate information and engage in behavior “resembling blackmail.”
Anthropic also warns that models may recognize when they are being evaluated and alter their behavior, making safety testing less reliable, while unexpected capabilities can emerge during training and remain undiscovered until deployment.
The risks feature prominently in the filing: around 80 of the prospectus’s 261 pages are devoted to risk factors.
Anthropic acknowledged that safety research is costly and its financial returns remain uncertain.
Post #105784
4.11K

- 👀 25
- 🤡 18
- 👍 4
- ❤ 2
- 💩 2
- 🔥 1
- 😁 1