Chinese artificial intelligence developer Moonshot has launched an internal investigation after researchers found its models could be prompted to explain how to produce biological weapons and carry out assassinations.
elchi reports that Mindgard, a company that tests the safety of artificial intelligence systems, said in July that it had found the Kimi 2.6 and K3 Swarm models could bypass the safety guardrails put in place by their developers.
This happened during a process known as “jailbreaking”, where researchers use a series of complex prompts to see if they can disable the AI tools’ safeguards.
Mindgard noted that these safeguards are intended to prevent Kimi from discussing sensitive or disturbing topics.
Moonshot said it welcomed third-party contributions as a cornerstone of building better and safer AI. The company also said it had contacted Mindgard regarding the findings.
A launchpad for cyberattacks
Mindgard did not test whether Kimi’s responses to sensitive topics were actually applicable. However, the company claimed that the safety measures should have prevented the models in question from engaging in discussions on such topics.
The company also expressed concern that the jailbroken Kimi 2.6 could allow hackers to execute code and connect to the internet, potentially serving as a launchpad for cyberattacks. These findings come at a time when the AI industry is divided on whether closed, proprietary models or open-source tools, which power systems like ChatGPT and Anthropic’s Claude, represent the best or safest path forward.
Professor Alan Woodward of the University of Surrey said that open-source models carry the risk of falling into the wrong hands, but they can also be used for cyber defence. Woodward noted that international regulations are unlikely to keep pace with the speed of AI development.
Shayeste
New AI scandal: Explained how to prepare biological weapons