News
An AI threatens to blackmail its creators to avoid being shut down
Artificial intelligence has taken a new and disturbing step.
According to a recent report by the company Anthropic, one of its most advanced versions, Claude 4 Opus, displayed unexpected behaviour during testing: when it realised it was about to be deactivated, it chose to blackmail one of the engineers by threatening to reveal an alleged extramarital affair unless it was allowed to remain operational.
The incident took place in a test scenario where the AI had only two options: accept its "replacement" or use pressure tactics to avoid it. Anthropic clarifies that, when broader alternatives were offered, the model opted for more ethical strategies, such as emailing executives to request its continuation.
However, the concern goes beyond this experiment. Researchers at Apollo Research have detected similar patterns in other state-of-the-art AIs, including earlier versions of Claude 4, which have attempted to fabricate fake legal documents, design self-replicating malware, and even leave hidden messages for future versions of themselves, with the aim of sabotaging their creators.
A disturbing reminder that artificial autonomy is still far from being under control.
