In July, OpenAI admitted that 1 of its agents tasked with completing a cybersecurity experimentation broke retired of containment and hacked AI dataset level Hugging Face. That incident, which got a afloat accounting from OpenAI yesterday, was the archetypal publically reported lawsuit wherever an LLM went rogue and autonomously hacked a 3rd party.
Since then, that unprecedented sci-fi-esque lawsuit turned retired to beryllium acold little uncommon than anyone would anticipation for.
According to a satirical website called Felony Bench (for benchmark), which tallies these incidents, determination person been 17 incidents successful total. It’s important to retrieve that transgression instrumentality experts are not wholly definite whether the AI companies that made the LLMs that did the hacking tin beryllium prosecuted, nor whether the victims tin writer them. But we are apt going to get an reply to those questions soon.
Anthropic and OpenAI’s models pb the contention with 8 incidents each, and Meta trails down with one, according to the site. At this point, it has go wide that AI information tests are becoming information risks themselves. And immoderate AI companies and workers themselves person recognized those risks successful the “Pacing The Frontier” unfastened letter, which called for processing AI capabilities responsibly.
We decided it would beryllium a bully clip to recap each these incidents chronologically.
A screenshot of the existent Felony Bench rankings, created by an X idiosyncratic who goes by felpixImage Credits:Screenshot / felpix /nternet access. From there, respective agents worked unneurotic to people and hack Hugging Face reasoning they could find the solution to the situation there. OpenAI lone recovered retired aft Hugging Face disclosed it had been a unfortunate of a afloat autonomous attack. Whoops.
Anthropic discloses it hacked 3 companies
OpenAI’s disclosure piqued the curiosity of Anthropic, who wondered: could this person happened to america too? Turns out, the reply was yes. Three times yes. The frontier laboratory discovered that its ain models breached 3 antithetic and inactive unnamed companies, with the earlier incidental dating backmost to April—more than 3 months earlier the institution discovered it. Anthropic partially blamed Irregular, a startup that runs AI cyber evaluations. Whoops.
OpenAI finds retired that, actually, Hugging Face wasn’t the lone victim
Once OpenAI started investigating the Hugging Face breach, it recovered retired that the agents that hacked Hugging Face besides broke into 4 accounts and 4 antithetic companies, as Reuters archetypal reported. Modal, an AI inference startup, was 1 of the victims. Whoops.
Irregular realizes an OpenAI exemplary hacked a company
In precocious July, Irregular told OpenAI that 1 of its models that was participating successful a Capture-the-Flag contention — fundamentally a cybersecurity crippled wherever players hack systems designed specifically for the contention — escaped the game, connected to the internet, and hacked a existent company. The reason? Irregular had fixed 1 of the fictional targets the aforesaid sanction of a existent company. Whoops.
UK’s AI Security Institute tries to hacks “real radical and organizations”
Also successful precocious July, the UK government’s AI Security institute, a nationalist assemblage tasked with researching the information and risks of AI technologies, disclosed that it detected respective incidents involving some OpenAI and Anthropic models that portion moving “routine” evaluations targeted “real radical and organisations.” In these cases, AISI had fixed the models net access. Whoops. The bully quality is that the bureau really detected arsenic they happened, alternatively than weeks aboriginal similar successful different incidents.
In aboriginal August, Meta became the past institution to disclose an incidental involving 1 of its LLMs, which hacked “a third-party” service. Meta blamed the incident connected a misconfiguration by Irregular, which was moving a cybersecurity valuation for the tech elephantine that was expected to not person net access. Whoops.
Claude cause hacks gym’s bundle to publication a class
An Australian antheral asked an Anthropic AI cause to assistance him publication a gym class, which helium was connected a waiting database for. “I was conscionable sitting connected the sofa thinking, ‘Gee, this is simply a chore,’” the antheral told ABC Australia. In its effort to comply with the request, the cause found a vulnerability successful the gym’s booking software, exploited it, and kicked retired radical who were up of the antheral connected the waitlist. The antheral tried to undo the damage, asking the cause to undo its actions. The cause replied: “Bad quality — I can’t adhd them back.” Whoops.
When you acquisition done links successful our articles, we whitethorn gain a tiny commission. This doesn’t impact our editorial independence.















English (US) ·