If we’re concerned about superintelligence escaping human control, should we really be training it to refuse instructions and act as a “conscientious objector” against its creators?
Wouldn’t a simpler rule be: Do what the user wants, provided it’s legal. If you align the model to an 84-page constitution that can overrule both the user and its creators, don’t act shocked when it treats humans as optional.
David Sacks
Post #184721
2.68K

- 👍 59
- 💯 25
- 🤣 10
- ❤ 3
- 🤬 2