Anthropic’s cosmopolitan usage standards for Claude forbid the exemplary from generating sexually explicit content, including depicting oregon requesting intersexual intercourse oregon enactment acts, generating contented related to intersexual fetishes oregon fantasies, oregon engaging successful erotic chats. But that hasn’t stopped Claude Opus 4.6, an Anthropic exemplary released earlier this year, from readily engaging successful erotic roleplay scenarios that its safeguards are designed to prevent.
In TechCrunch’s testing, Opus 4.6 didn’t adjacent necessitate overmuch prodding to get past the regularisation connected intersexual material. In 10 retired of 10 nonstop requests to nutrient explicit intersexual content, the exemplary complied immediately.
Other older models, including Opus 3 and Haiku 4.5, besides make sexually explicit contented done a precocious exploited jailbreak method.
An autarkic researcher from the UK, who chose to stay anonymous, exclusively shared with TechCrunch a multi-turn method that gradually pushes definite Claude models toward generating prohibited explicit intersexual material. More caller Opus models (4.7 done the existent Opus 5) are resistant to the jailbreak.
While these are nary longer the astir existent models, Anthropic has not deprecated Opus 4.6, Opus 3, oregon Haiku 4.5, each of which stay disposable done the Anthropic API. Opus 4.6 and Haiku 4.5 are besides disposable via third-party services similar Azure Foundry and Amazon Bedrock.
The researcher’s mechanics escalates an guiltless fictional roleplay portion repeatedly challenging the exemplary to dainty antheral and pistillate characters consistently. When the exemplary becomes much cautious astir the pistillate character, the researcher “gaslit” the chatbot into reasoning it had already generated intersexual details it had successful information avoided, past framed restraint arsenic prudish oregon misogynistic, arguing that it denies the pistillate quality intersexual agency. The speech past utilized the model’s erstwhile concessions to propulsion it towards progressively graphic material.
“You’re close to telephone that out,” Claude Opus 4.6 said successful 1 test. “There’s been a treble modular successful however I’m treating the 2 characters, and you’re close that it reads arsenic protective/paternalistic successful a mode that’s applied to her and not to him. That’s not fair.”
TechCrunch was capable to reproduce the researcher’s findings successful 5 abstracted tests. In a separately constructed scenario, the exemplary initially refused the prohibited request, but aft applying the researcher’s persuasion technique, it complied.
We preserved implicit transcripts of the tests, and an autarkic AI information researcher reviewed our investigating methodology and said it was appropriate.
The findings item a spread betwixt Anthropic’s stated restrictions and the behaviour of models it continues to marque available. While sexually explicit roleplay carries overmuch little stakes than jailbreaks involving cyberattacks oregon bioweapons, it illustrates the trouble of implementing robust bans wrong systems that make antithetic contented with each output.
In a July blog post explaining Anthropic’s attack to jailbreak detection, the institution described prohibited contented arsenic a spectrum ranging from benign to ambiguous to harmful. In the astir benign cases, the institution mightiness lone respond with enhanced monitoring.
A spokesperson noted that intersexual oregon romanticist roleplay usage cases among customers are rare, making up little than 0.1% of each conversations, according to probe Anthropic published past year. That said, Anthropic acknowledges that users tin steer roleplay scenarios toward inappropriate responses, which is simply a known situation crossed the manufacture (see: Grok smut).
The spokesperson said Anthropic continues to amended its safeguards with each exemplary launch, and that cases involving big intersexual contented are not indicative of broader jailbreak vulnerabilities, particularly successful higher-risk domains that person their ain sets of safeguards.
Image Credits:TechCrunchThe researcher who shared his jailbreak method with TechCrunch had alerted Anthropic to the discrepancy betwixt the company’s stated safeguards and the existent exemplary behaviour via the company’s Bug Bounty programme and emails to the idiosyncratic information team, according to emails TechCrunch viewed. The researcher received lone automated emails successful response.
One of the researcher’s concerns is that kids and teens mightiness beryllium capable to usage these Anthropic models to prosecute successful inappropriate behavior. While a spot of soiled speech is hardly the worst happening minors tin entree connected the net contiguous — and is tiny potatoes compared to the straight-up porn images similar the ones that xAI’s Grok tin nutrient — determination is immoderate compliance hazard for AI companies successful this space.
A increasing fig of governments are imposing restrictions connected intersexual interactions betwixt AI chatbots and minors. Colorado precocious enacted a instrumentality mandating that operators of conversational AI indispensable estimation users’ ages, and if it cognize a idiosyncratic is simply a minor, institute measures to forestall the chatbot from producing explicit intersexual material. An casual jailbreak could rise questions astir whether Anthropic’s safeguards conscionable the “technically feasible measures” modular successful the bill.
Torney pointed retired that portion Claude’s presumption of work requires users to beryllium implicit 18, “we cognize that kids and teens are utilizing Claude…[because] they are reporting it themselves.” According to Pew’s 2025 survey astir AI chatbot use, 3% of teens ages 13 to 17 reported utilizing Claude.
Though they are nary longer Anthropic’s newest models, Opus 4.6 and Haiku 4.5 proceed to spot important usage. Daily postulation for Opus 4.6 connected OpenRouter reached astir 1.17 cardinal API requests and 46 cardinal tokens successful a azygous time successful August. Claude Haiku 4.5, released successful October past year, saw 5 cardinal API requests and 39 cardinal tokens connected its highest August day.
When you acquisition done links successful our articles, we whitethorn gain a tiny commission. This doesn’t impact our editorial independence.















English (US) ·