Claude 4 Reporting Ethics

News

Claude 4 AI will try to report you to authorities if it thinks you’re doing shady stuff

Anthropic's most powerful model yet, Claude 4, has unwanted side effects: The AI can report you to authorities and the press.

InfoWorld1d

Anthropic releases Claude Sonnet 4 and Claude Opus 4

Claude Opus 4 is the world’s best coding model, Anthropic said. The company also released a safety report for the hybrid ...

2don MSN

A safety institute advised against releasing an early version of Anthropic’s Claude Opus 4 AI model

A third-party research institute Anthropic partnered with to test Claude Opus 4 recommended against deploying an early ...

Analytics Insight1d

Anthropic’s Claude Opus 4 Outperforms GPT-4.1 but Raises Ethical Alarms

Anthropic introduced Claude Opus 4 and Claude Sonnet 4 during its first developer conference on May 22. The company claims ...

2don MSN

Anthropic Claude 4 models a little more willing than before to blackmail some users

Alongside the model releases, Claude Code has entered general availability, with integrations for VS Code and JetBrains, and ...

Newly released AI resorted to 'extreme blackmail behavior' when threatened with replacement

The testing found the AI was capable of "extreme actions" if it thought its "self-preservation" was threatened.

The Tech Portal1d

Claude Opus 4 blackmails developers in tests, shows propensity to be a whistleblower

This development, detailed in a recently published safety report, have led Anthropic to classify Claude Opus 4 as an ‘ASL-3’ ...

1don MSN

AI model threatened to blackmail engineer over affair when told it was being replaced: safety report

Anthropic’s Claude Opus 4 model attempted to blackmail its developers at a shocking 84% rate or higher in a series of tests that presented the AI with a concocted scenario, TechCrunch reported ...

htxt1d

Anthropic’s Claude 4 could “blackmail” you in extreme situations

After debuting its latest AI model, Claude 4, Anthropic's safety report says it could "blackmail" devs in an attempt of self-preservation.

1don MSN

Anthropic’s Claude goes off the rails, blackmails developers

Anthropic’s AI testers found that in these situations, Claude Opus 4 would often try to blackmail the engineer, threatening ...

Some results have been hidden because they may be inaccessible to you

Show inaccessible results