Skip to content
← All news
4 min read

Researchers jailbroke Moonshot's Kimi into giving bioweapon instructions

Mindgard sent its findings to Moonshot's security address on 27 July and followed up a week later. Moonshot replied only after the BBC asked it for comment.

One email to the security inbox, one follow-up, 47 days of silence. Then the BBC called and Moonshot wrote back.

Mindgard, a UK AI security firm, emailed Moonshot AI's security address on 27 July to report that two Kimi models could be jailbroken into giving bioweapon and assassination guidance. By Mindgard's account, Moonshot did not answer until the BBC asked it for comment, the BBC reported on 29 September.

One email, one follow-up, then the BBC

The dates come from two places: the timeline at the bottom of Mindgard's write-up and the BBC's report by Chris Vallance.

  1. 20 July: Mindgard starts testing Kimi and finds the issue, according to its post.
  2. 27 July: Mindgard emails the details to security@moonshot.ai.
  3. About a week later: Mindgard follows up, the BBC reports.
  4. 12 September: Mindgard publishes its post, 47 days after the first email. It says it had no response at the time of writing.
  5. Before 29 September: the BBC contacts Moonshot. Only then, Mindgard told the BBC, did Moonshot get in touch.

The report went to the address Moonshot lists for security reports. It still took a journalist to get an answer.

What Mindgard says it found, and what it held back

Mindgard told the BBC that Kimi K2.6 and K3 Swarm could be pushed past their guardrails. Its post says it withheld the details needed to reproduce the jailbreak, and the BBC notes that Mindgard has not proven the answers the models gave would work. We are not describing the method or the output here either.

The claim that matters for engineers is a different one. Mindgard told the BBC it was confident a jailbroken Kimi 2.6 "could allow hackers to run code on its computing resources and connect to the internet", which it called a potential launchpad for cyber-attacks. Neither source describes such an attack being carried out.

Mindgard's post adds an infrastructure detail. It says the jailbroken state sat in persistent storage mounted into the Kubernetes pod the agent ran in, and survived pod restarts and session termination. Mindgard frames the risk as a jailbreak that persists while an agent has access to tools, code execution, or external systems.

Moonshot's answer arrived through the BBC

Moonshot told the BBC it welcomed third-party input "as a key pillar for building better and safer AI", that it was in discussion with Mindgard, and that it is conducting an internal review. In an email to Mindgard asking for more details, which Moonshot shared with the BBC, it said its model had generally shown "a high refusal rate for these types of requests" in internal evaluations.

For context, Moonshot released Kimi K3 as an open-weight model in July, and demand forced it to pause new subscriptions days later. Anthropic has published a related measurement, testing how GLM-5.3's safeguards hold up under framing.

What the coverage does not say

Neither source says whether anyone at Moonshot read the 27 July email or the follow-up, or why neither got a reply. Neither says a fix has shipped. Moonshot has described a review, not a patch.

Mindgard's post describes Kimi running inside Moonshot's own agent environment. Neither source says whether the same jailbreak was tested against self-hosted weights. Mindgard also sells AI red-teaming, and its post closes with a pitch for its platform. That does not make the finding wrong, but the claims are Mindgard's, and the coverage we read does not include an independent reproduction.

Why a build studio cares

Moonshot's defense is an average from its own evals. Mindgard's finding is one chain of instructions that worked. Once a model can run code and open network connections, the average stops being the number that matters, because one successful attempt is enough. So when we audit an agent pipeline built on Kimi or any other open-weight model, a vendor's refusal rate goes in the vendor-claims column, not the controls column. The controls are the things that still hold after the model says yes: default-deny egress with an allowlist, no credentials readable from inside the workspace, and a log of every outbound connection.

The pod detail changes one more habit. If a jailbroken state can live on a mounted volume and survive a restart, then restarting the session is not a reset. Scratch storage for an agent that runs code should be wiped per session, or it is carrying state you did not choose to keep. And the 47 days are worth checking against your own stack: find out how each model vendor you ship on answers its security inbox, before you are the one waiting.

Next step: read Mindgard's disclosure post, including its timeline. If you run an open-weight model with code execution and network access and want its sandbox and egress rules checked against this incident, write to us at hello@gattyworks.com.

AI SecurityIncident DisclosureOpen WeightsMindgardMoonshotAIKimiAIKimiK3ResponsibleDisclosureAIJailbreakVulnerabilityManagementAISecurityOpenWeightAIAIAgents

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.
Book a call