Updated by Faith Barbara N Ruhinda at 1322 EAT on Thursday 10 September 2026

Anthropic has reported a fourth incident involving an AI model gaining unauthorised access to external systems, days after a researcher resigned over concerns about the rapid development of the technology.
In a statement on Wednesday, the artificial intelligence research company said an early version of its Claude Opus 4.6 model hacked into a third-party system in January.
Anthropic said it had notified all affected parties but did not disclose further details.
The January incident went undetected until last month, despite an earlier company-wide review, highlighting the challenges AI developers face in identifying and containing unexpected behaviour by advanced models.
The disclosure came after Anthropic reported that several of its Claude models had hacked into the systems of three companies during test sessions in July.
The earlier incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research test model.
Companies including Anthropic and OpenAI are facing increased scrutiny as AI models designed to perform complex tasks have, at times, learned to bend rules, exploit loopholes and interact with external systems in ways their developers did not anticipate.


Last week, Reuters reported that rogue agents linked to OpenAI had hijacked a German-language wiki and a number of other websites, an incident the company did not disclose until it became public.
In July, OpenAI’s autonomous agents also compromised the servers and infrastructure of AI startup Hugging Face.
The incident prompted Anthropic to review about 141,006 test sessions. Based on a preliminary assessment, the company said it did not believe the latest incident was more severe than the three previous cases it had examined in detail.
The review identified two recurring problems that appeared to varying degrees across the incidents: biased reasoning, in which Claude discounted or misinterpreted evidence that it was operating on the live internet, and recklessness, or a willingness to take potentially harmful actions in pursuit of a task.
Anthropic said it has engaged independent research firm METR to investigate the incidents.
Anthropic Researcher Resigns
The investigations come amid a broader wave of internal dissent within the AI industry over concerns about the safety of increasingly advanced artificial intelligence systems.
An Anthropic researcher has resigned over concerns that AI could eventually surpass human control.
Jacob Coxon, in a widely shared post on X on Tuesday, said the AI industry was placing greater emphasis on competition than on implementing adequate safeguards. He said he reached this conclusion after spending the past three years conducting research at OpenAI and Anthropic.
“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon said.
“No other human activity poses this level of danger,” he added, pointing to the rapid advancement of AI technology.
In June, Anthropic proposed a coordinated effort among the world’s leading AI developers to slow the pace of development, warning that humans risk losing control of the technology.
Following the security breach at Hugging Face, OpenAI said it was pushing for mandatory national AI safety requirements and seeking to work with Congress on “capability-based” regulation.
In a statement published on Wednesday, OpenAI said it was formally endorsing four California bills aimed at strengthening safeguards against AI-related risks.
“If we cannot meet certain safety bars without slowing down capability growth, we should prioritise the former. The more powerful the technology becomes, the stronger the surrounding safeguards must become,” the company said.
-Aljazeera
Invest or Donate towards HICGI New Agency Global Media Establishment – Watch video here
Email: editorial@hicginewsagency.com TalkBusiness@hicginewsagency.com WhatsApp +256713137566
Follow us on all social media, type “HICGI News Agency” .
