Editorial illustration for Open Model apex-flash-1 Solves 40 of 60 Security Research Bug Tasks
Open Model Solves 40 of 60 Security Bug Tasks
Cantina Security put a number on something most security teams argue about over drinks: can an open-weights model actually hunt bugs, or is that job stuck behind a Claude or GPT paywall. Working with Yeta Labs, the company built apex-flash-1, a reinforcement-learning fine-tune of Z.ai's GLM-5.3-Flash, and released it on Hugging Face under the MIT license. The base model is a Mixture-of-Experts system with 321.3 billion total parameters and 18 billion active, trained further with GRPO using a rank-256 LoRA plus selective full-parameter updates.
The training set wasn't generic code. Cantina built 150 tasks from 50 real vulnerability cases, each run three ways, guided whitebox, focused whitebox, and focused blackbox, with authorization and identity flaws making up 72% of the cases. Rollouts happened inside the Codex agent harness against production-like software and protocols, not toy sandboxes.
Then Cantina tested the result against 20 held-out cases it never trained on, comparing apex-flash-1 to its own base model and to Anthropic's Claude Opus 5 High. The gap in results, and in price per solved bug, is where this gets interesting.
Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license.
Why this matters
The gap between apex-flash-1 and Claude Opus 5 High is three tasks. The gap in cost is about 31x. For teams running bug bounty triage or internal vuln scans at volume, that math is hard to ignore: $0.06 per solved task versus $1.74 changes what "run it on everything" actually means.
We'd treat the 66.7% pass@1 figure as a floor, not a ceiling, since this is a reinforcement learning fine-tune of an already-capable base model (GLM-5.3-Flash jumped from 60% to 66.7% after training), which suggests the technique generalizes to other open bases. The catch is deployment cost: 640 GB of GPU memory in BF16 isn't something most startups have sitting around, MIT license or not. That's a real barrier even as the per-query economics look good.
Watch whether Cantina or others quantize this down to something that runs on commodity hardware, and whether the 60-task held-out set holds up against messier, real-world codebases rather than curated benchmarks. Security research is exactly the kind of high-stakes, verifiable-output task where open models closing the gap with frontier labs matters more than in chat or writing use cases.
Common Questions Answered
What is apex-flash-1 and how was it developed?
apex-flash-1 is an open-weights model created by Cantina Security and Yeta Labs specifically for vulnerability research and bug detection. It is a reinforcement learning fine-tune of Z.ai's GLM-5.3-Flash base model, which is a Mixture-of-Experts system with 321.3 billion total parameters and 18 billion active parameters, and it was released on Hugging Face under the MIT license.
How does apex-flash-1 perform on security research bug tasks compared to Claude Opus 5 High?
apex-flash-1 solves 40 out of 60 held-out bug tasks with a 66.7% pass@1 figure, achieving only a three-task gap behind Claude Opus 5 High. This performance is particularly significant given that apex-flash-1 costs approximately $0.06 per solved task compared to Claude Opus 5 High's $1.74 per task, representing a 31x cost difference.
What are the cost implications of using apex-flash-1 for security teams running bug bounty triage?
For teams conducting bug bounty triage or internal vulnerability scans at scale, apex-flash-1 offers substantial cost savings at $0.06 per solved task versus $1.74 for Claude Opus 5 High. This significant price difference makes running security analysis on large volumes of tasks economically feasible with the open-weights model, fundamentally changing the cost-benefit calculation for vulnerability research workflows.
Why is apex-flash-1 released as an open-weights model under the MIT license?
Cantina Security released apex-flash-1 as an open-weights model under the MIT license to address the debate about whether open models can perform security research tasks that were previously locked behind proprietary APIs like Claude and GPT. This open release democratizes access to vulnerability research capabilities and allows security teams to run the model without relying on paid API services or vendor lock-in.
What does the 66.7% pass@1 figure represent for apex-flash-1's capabilities?
The 66.7% pass@1 figure represents the percentage of bug tasks that apex-flash-1 successfully solves on the first attempt, and this metric should be treated as a floor rather than a ceiling for the model's capabilities. Since apex-flash-1 is a reinforcement learning fine-tune of an already-capable base model (GLM-5.3-Flash), there is potential for further improvement and optimization of the model's security research performance.
Further Reading
- Cantina releases an open-weights model trained for vulnerability research - RuntimeWire
- Cantina releases open-weight security model trained on 50 vulnerability cases - RuntimeWire
- Cantina Releases Apex Flash-1, a 321B Open-Weight Security Research Model - AI.info
- cantina-security/apex-flash-1 at main - Hugging Face
- Cantina | AI-Native Security, Backed by Human Expertise - Cantina Security