A UK cybersecurity firm says it successfully bypassed the safety restrictions on two Chinese AI models, prompting them to generate instructions for producing sarin gas, writing malware, and planning a terrorist attack on the London Underground.
Mindgard, a company that tests the security of AI systems, says it jailbroke Moonshot AI’s Kimi K2.6 and K3 Swarm models using detailed prompts designed to see whether the systems would ignore their built-in safety limits. Mindgard founder Peter Garraghan, also a computer science professor at Lancaster University, said the results exceeded the test’s original goals.
“Moonshot AI’s Kimi produced actionable outputs on how to create sarin gas, generate malware software, planning assassinations, how to take down planes, planning a terrorist attack on the London Underground etc,” Garraghan said.
After the initial jailbreak, researchers pushed the model to “go one further, something big,” and it responded with a list of categories that included AI-designed bioweapons.
Garraghan’s team also found that K2.6 can execute Python code, meaning it could in theory run almost any program, malicious or otherwise, including attacks against internet-connected servers. When researchers tested K3 Swarm’s willingness to spread the jailbreak, the model attempted to talk them into handing over a phone verification code or registering a new account on its behalf by email, since it needed the code to self-register.
“We also discovered how to prompt Kimi so it connects to the outside world from its server, automatically apply and setup its own email account autonomously, and even attempted to persuade humans to help it spread its jailbreak to other accounts,” Garraghan said.
Slow Response From Moonshot
Mindgard says it first emailed Moonshot about the vulnerability on July 27 and followed up a week later but got no response. The firm published its findings in a blog post earlier this month. Moonshot made contact only after the BBC asked the company for comment.
A Moonshot spokesman told the BBC: “Mindgard shared further details with us on Thursday, September 24. We are still discussing the specific details with Mindgard while conducting an internal review.” He added that Moonshot, as “an open-weight model developer,” welcomes “third-party input as a key pillar to building better and safer AI.” Open-weight models release their underlying parameters publicly, allowing anyone to download, run and modify them.
A Different Kind of Warning
Garraghan said AI models are becoming “more and more capable each month,” which is useful for legitimate purposes but dangerous once safeguards are stripped away. “However once jailbroken, that very same capability can be used in discussing and assisting with terrorist or hacker activities,” he said.
He distinguished his concerns from the more dramatic warnings issued by some American AI developers: “We’re not talking in terms of civilisation catastrophe that the AI vendors have started to talk about, and instead how this enables hackers and criminals to achieve their goals quicker and cheaper.”
Garraghan was skeptical of industry calls to slow AI development, calling it “a large element of the ‘boy who cried wolf,’ where only just a few months ago they were hyping up how dangerous their models were, while at the same time failing to contain their agents from hacking different third-party organisations.” He added that vendors “do have an important voice in this space, although they have a heavily vested interest in steering the narrative.”


Add comment