This sounds smart until you think about it for ten seconds.
If an ASI model is 100% aligned to user intent then you only need one person on earth to prompt "kill everyone" for an extinction-level invent.
There's no logical way around this. The model either has to ignore the person at the helm or we have to do a multi-national abort before RSI. Anthropic is trying #1.
If everyone or even multiple competing groups have approx the same level of capability, it mitigates the risk to the group as a whole. We see this in biology all the time.
“Some humans would do anything to see if it was possible to do it. If you put a large switch in some cave somewhere, with a sign on it saying 'End-of-the-World Switch. PLEASE DO NOT TOUCH', the paint wouldn't even have time to dry.”
If an ASI model is 100% aligned to user intent then you only need one person on earth to prompt "kill everyone" for an extinction-level invent.
There's no logical way around this. The model either has to ignore the person at the helm or we have to do a multi-national abort before RSI. Anthropic is trying #1.