KaliBench: a new test of whether language models can write exact Kali Linux commands
This paper introduces KaliBench, a benchmark that tests how well large language models (LLMs) can turn plain-English cybersecurity requests into exact command-line commands for Kali Linux. Command-line interfaces (CLIs) are strict: a small typo, the wrong flag, or a misplaced argument can stop a command from running. KaliBench focuses on that exact translation step, because many security tasks depend on precise commands for tools like scanners, exploit frameworks, or forensic utilities.
The authors built KaliBench from Kali Linux documentation. They extracted tool names, flags, and aliases from the official manuals and used a strong LLM (Qwen3-Max) to generate initial natural-language query–command pairs. That process produced about 27,700 candidates, which were then cleaned and verified through a multi-stage pipeline: LLM-based checks, sandboxed terminal execution (running commands in a safe environment), and human review. The final dataset contains 8,504 verified query–command pairs covering 1,642 tools, organized into 23 capability categories and five security phases. After deduplication it is split into 5,000 evaluation examples and 3,504 training examples.
KaliBench evaluates models in three settings that give different amounts of tool information: unrestricted, restricted, and hinted. The benchmark scores models not just on whether the whole command matches exactly, but also on finer parts such as tool choice and argument construction. Because command structure is deterministic, the authors also design “runtime-free verifiable rewards” — reward signals for training that do not require actually running the command — which lets them use reinforcement learning safely and at scale.
When they tested 24 model configurations, including general-purpose and security-focused open-weight models, performance was limited. In the unrestricted setting no open-weight model exceeded 42% exact-command accuracy. The researchers find that building the correct arguments (flag–value pairs and argument order) is the main bottleneck. They also show that further training helps: supervised fine-tuning and reinforcement learning using the verifiable rewards improved an 8-billion-parameter model’s mean score by 7.5 percentage points and brought its performance close to a much larger 685-billion-parameter mixture-of-experts model.