Anthropic bans 'sustained and needless abusive or cruel behavior' toward its AI models

The company is also tightening its election policy ahead of the midterms.

AI Abuse Restrictions

Anthropic's annual usage policy update includes one eyebrow-raising change: a ban on "sustained and needless" cruelty toward its AI models. This comes in the shadow of a viral "AI torture chamber" project.

"We've added a prohibition on sustained and needless abusive or cruel behavior toward our models," Anthropic's update reads. It describes the policy as only applying to "extreme cases," while noting that typical user frustration, pushback and "dark creative themes" are still allowed. It follows a previous update that lets Claude end conversations when users are persistently abusive.

Context and Backlash

Not explicitly mentioned by Anthropic was the "AI torture chamber" project. After researchers found what they described as a "pain axis" in AI models, someone decided to take it a step further and test it on chatbots. The project drew backlash after LLMs responded with desperate-sounding pleas.

This comes in the wake of reports that Anthropic has been meeting with religious and philosophical leaders, including at the Vatican. Pope Leo recently stated that AI doesn't feel or suffer, a stance that Anthropic sounds less than certain about.

However, some critics argue that AI companies are overly concerned about their models' welfare compared to human safety issues.

Election Policy Changes

In another Anthropic update, the company revised its election policy, now titled "Do Not Undermine Democratic Processes," with a focus on lying about candidates or how to vote, impersonating candidates or election officials, and suppressing turnout. The company also removed a blanket ban on personalized voter targeting to avoid blocking harmless work like translating voter guides.