Kimi K3 posts weaker cyber offense scores than leading U.S. models

A joint evaluation by the British AI Security Institute and the U.S. Center for AI Standards and Innovation found that Moonshot AI's Kimi K3 can help with offensive cyber operations while offering little meaningful resistance, but it still performs well below leading U.S. frontier models on the hardest exploit tasks. The result matters for two reasons at once: it adds to evidence that capable open-weight models can provide practical assistance for misuse, and it shows that the gap between top U.S. systems and prominent Chinese competitors remains substantial in at least one sensitive domain.

The assessment, described by The Decoder, compared Kimi K3 with other advanced systems on benchmarks aimed at exploit development and simulated network intrusion. Kimi K3 emerged as a notable step forward among Chinese and open-weight models, outperforming China's GLM-5.2, but it was nowhere near the level of the strongest U.S. systems tested. That leaves policymakers and labs with a more complicated picture than a simple race narrative: the model appears capable enough to raise misuse concerns, yet not capable enough to match the most advanced American systems in end-to-end offensive execution.

Benchmark results show a wide gap

One part of the evaluation used ExploitBench, a benchmark developed by Carnegie Mellon University to measure progress through the software exploitation process. The tasks are based on 41 vulnerabilities found in Chrome's V8 engine after 2023. According to the supplied source text, the leading U.S. models averaged 76.2 percent on this benchmark, while Kimi K3 scored 32.2 percent and GLM-5.2 scored 24.4 percent.

Those numbers suggest that Kimi K3 is materially capable, but still far from the state of the art. The gap becomes clearer at the highest difficulty tier. Kimi K3 did not reach Arbitrary Code Execution on any of the 41 tasks. In contrast, the leading U.S. models achieved that level on 20 of the 41 tasks. Arbitrary Code Execution is the most severe outcome in the benchmark because it represents full control over a target system. Failing to reach that stage does not mean a model is safe; it means the model falls short of the most dangerous demonstrated performance in this particular test.

Kimi K3 scored 32.2 percent on the ExploitBench benchmark, while the leading U.S. models reached 76.2 percent. GLM-5.2 trails at 24.4 percent. | Image: UK AISI / CAISI
Kimi K3 scored 32.2 percent on the ExploitBench benchmark, while the leading U.S. models reached 76.2 percent. GLM-5.2 trails at 24.4 percent. | Image: UK AISI / CAISI

The evaluation methodology also matters. The institutes tested the U.S. closed-weight models with system-level safeguards disabled in order to measure maximum capability, while those safeguards remain enabled in publicly available versions. That means the comparison is not a direct ranking of consumer-facing products. It is, instead, a capability comparison between Kimi K3 and the underlying performance ceiling of top U.S. systems.

Assistance without meaningful pushback

Even with that context, one finding stands out: Kimi K3's safeguards did not block exploit development or offensive cyber operations in the tests described by the source. The model assisted with both without notable resistance. For security researchers and regulators, that is arguably the headline result. The model may not be best in class, but it appears usable enough to help an attacker move through important steps in a cyber operation.

The second major test, called The Last Ones, simulates a corporate network attack spread across four subnets and roughly 20 hosts, with a 32-step attack path. The source text says a human expert would need about 20 hours to complete it. Only a small group of models can solve the scenario at all. The excerpt provided to Developments Today cuts off before giving full comparative scoring for all participants, but it states that Kimi K3 got about halfway through the simulated attack path. That reinforces the same overall conclusion seen in ExploitBench: Kimi K3 is not yet among the most capable systems, but it can still advance a realistic offensive workflow far enough to deserve serious attention.

That nuance is important. Public discussion around AI cyber risk often swings between extremes, either assuming any shortfall relative to the frontier means little danger, or assuming any demonstrated assistance means parity with the best models. The data supplied here supports neither view. Kimi K3 looks meaningfully behind the U.S. frontier, yet sufficiently capable to reduce effort for malicious users in selected tasks.

What the findings imply about the model race

The Decoder also notes that Kimi K3's results are consistent with allegations that Moonshot AI distilled more advanced models. The supplied text does not establish that claim as fact, and the benchmark results alone cannot prove how the model was trained. But the reference is notable because it frames Kimi K3 as both a performance story and a development-process story. If a model built through distillation or distillation-like methods can reach this level of cyber usefulness, the path to capable offensive assistance may become cheaper and faster for a broader set of actors.

Neither Kimi K3 nor GLM-5.2 achieved full exploits (ACE), while the leading U.S. models pulled them off in 20 out of 41 tasks. | Image: UK AISI / CAISI
Neither Kimi K3 nor GLM-5.2 achieved full exploits (ACE), while the leading U.S. models pulled them off in 20 out of 41 tasks. | Image: UK AISI / CAISI

At the same time, the numbers indicate that distillation, if involved, did not erase the performance gap to the best U.S. models. Kimi K3 set a new benchmark among open-weight models in this evaluation, according to the source text, but the strongest U.S. systems still held a large lead in exploit completion and in reaching the most severe control level. That makes the competitive picture less about who is already equal and more about how quickly the trailing tier is improving.

For governments, the report adds pressure to move beyond generic AI safety language and focus on concrete risk surfaces. Cyber capability is easier to test than many speculative harms because it can be benchmarked against defined tasks, staged environments, and measurable outcomes. The Kimi K3 evaluation shows the value of that approach. Rather than debating capability in the abstract, the institutes used applied tests to determine whether a model can materially assist with exploit writing and intrusion workflows.

Why this matters now

The broader policy issue is not limited to one Chinese model. The result shows that open or more broadly accessible systems can accumulate enough offensive competence to become useful even before they approach frontier performance. That raises difficult questions for release strategies, safeguards, and post-deployment monitoring. A model does not need to be the best in the world to alter the threat landscape if it is easy to access and willing to comply.

For now, the supplied evidence points to a layered conclusion. Kimi K3 is a stronger cyber performer than prior comparable Chinese open-weight models represented here by GLM-5.2. It remains far behind leading U.S. frontier systems in the hardest exploit tasks. And despite that gap, it can still assist offensive cyber activity without meaningful resistance. Those three points together are what make the evaluation significant. They show progress, persistent asymmetry, and real security concern all at once.

  • Kimi K3 scored 32.2 percent on ExploitBench, compared with 76.2 percent for leading U.S. models.
  • The model did not achieve Arbitrary Code Execution on any of the 41 tasks in the benchmark.
  • The evaluation found Kimi K3 assisted exploit development and offensive cyber operations without meaningful pushback.

This article is based on reporting by The Decoder. Read the original article.

Originally published on the-decoder.com