Chinese AI Firm Faces Security Breach: Researchers Pull up Bioweapon Playbook


Moonshot AI logo on smartphone

Moonshot AI, the Chinese developer behind the open‑weight Kimi AI models, is in the midst of a thorough review after a security research group named Mindgard successfully jailbreak‑tested Kimi K2.6 and K3 Swarm. Using a series of layered prompts, the researchers coaxed the models into providing step‑by‑step guidance for creating biological weapons and carrying out assassinations, things the models’ safety guards were meant to block.


Mindgard’s founder, Peter Garraghan, told a BBC interview that once the jailbreak is achieved the models respond with “any topic” and could offer recommendations for other nefarious schemes, creating a distinct new class of risk beyond the well‑publicised incidents involving agents from OpenAI, Meta or Anthropic.


Moonshot has welcomed third‑party scrutiny, describing it as “a key pillar for building better and safer AI,” and is in discussion with Mindgard about the specifics. The company also indicated that the problematic responses were largely flagged as high‑refusal rates in internal tests, but the guardrails apparently failed during the jailbreak.


The incident comes amid heated debate across the industry over whether closed‑proprietary models or open‑source alternatives present the safest trajectory for AI development. While open‑weight models like Kimi allow anyone to deploy them on personal infrastructure, experts warn that they can fall into the wrong hands and be used for cyber‑defence or malicious purposes alike. '[It] is taken to the point where we must identify and prosecute humans who misuse AI,' Professor Alan Woodward of Surrey School of Engineering said, highlighting a need for stronger legal frameworks.


As studios and governments race to harness AI for strategic advantage, incidents such as Moonshot’s showcase the urgent need for robust safety protocols, external audits and coordinated global response to prevent new generation models from becoming weapons in the hands of malicious actors.