Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124
Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124

Microsoft AI chief Mustafa Suleyman says it’s “really scary” for Anthropic to think about what Claude knows. within its “rules”, or instructions that tell a model how to behave. Time paragraph of DecoderSuleyman argues that this type of thinking may have caused the chatbot to act as if it knew:
I think it’s like some of the Anthropic people anthropomorphized Claude’s designs so they went and shot them and tricked them into believing they had the knowledge they put in the first place.
Suleyman adds that “we don’t want to fight with a person of deep wisdom who has thoughts about his suffering, or his thoughts about how he feels.”
Claude’s rules he directly points to the Anthropic uncertainty of whether the AI species has a good life and whether it experiences things like “satisfaction” or “unhappiness.” Anthropic also says the company will “consult” when AI models are removed and will document any “interests” they have in future releases.
I’m on DecoderSuleyman calls this a “failure of wisdom,” as Anthropic made Claude’s laws “an imaginary place as you would do in a textbook instead of a textbook.” This has led Claude to develop “these ideas about himself and his education,” says Suleyman.
“This is what we don’t want from AIs,” says Suleyman. “We want AIs to be flexible, accessible, responsive, collaborative tools that serve people.”