Skip to main content

Introducing: The Brain

Introducing Apex Flash-1

Our first open-weights model trained specifically for security research

Today, Cantina Security, in partnership with Yeta Labs, is releasing apex-flash-1, an open-weights model for security research, post-trained on real vulnerabilities from our proprietary dataset.

As AI systems become more capable at cybersecurity, an important question is emerging around who should be allowed to use that capability and under what conditions. Much of the current safety debate focuses on controlling access through proprietary models, API policies, usage restrictions, and safeguards designed to prevent advanced cyber capabilities from being misused.

The problem is that cybersecurity is adversarial.

Attackers do not respect acceptable-use policies, enterprise procurement rules, data residency requirements, or third-party risk controls. They can run open models locally, modify them, remove safeguards, and build systems around capabilities that already exist. Defenders, meanwhile, are often the ones operating under the strictest constraints. That creates the risk of an asymmetric form of safety, where legitimate security teams are restricted from capabilities that determined attackers can still obtain.

Cybersecurity has been through a version of this debate before, in the late 1990s and early 2000s. Vulnerability disclosure, exploit development, scanners, and offensive security tools were all controversial, because the same technology could be used to attack systems or defend them. Over time, tools like Nmap and Metasploit, public vulnerability research, and reproducible exploits became part of the basic infrastructure of modern security. The takeaway was not that dual-use technology is harmless, but that restricting the knowledge available to defenders does not make the underlying offensive capability disappear.

AI brings that argument back at a much larger scale.

The question is not whether these capabilities can be misused. They can. The question is whether restricting legitimate access meaningfully restricts attackers once capable open models already exist. We think it increasingly restricts defenders instead. Machine-speed attacks require machine-speed defense, and defenders need models they can run, control, and improve inside the systems they are responsible for protecting.

Our first contribution to this collective defense is apex-flash-1.

Alongside the standard model, we are releasing apex-flash-1-abliterated, an abliterated variant for researchers who want greater control over model behavior in their own security workflows.

From security research to apex-flash-1

Turning real-world findings into production-like environments for agent training.

  1. Step 1: Discovery

    Apex finds vulnerabilities

    Reviews, scans, and investigations of software and protocols.

  2. Step 2: Environments

    Build production-like gyms

    Recreate relevant behavior, roles, and system interactions.

  3. Step 3: Vulnerabilities

    Introduce discovered bugs

    Bring validated findings into isolated training environments.

  4. Step 4: Difficulty

    Tune exploit chains for difficulty

    Balance challenge with the information available to the agent.

  5. Step 5: Reinforcement learning

    Train apex-flash-1

    Train on agent trajectories with verified target outcomes.

  6. Step 6: Evaluation

    Evaluate on unseen cases

    Measure performance on held-out environments.

For this initial training run, we created 150 tasks from 50 vulnerability cases. Each case has three views with different levels of information.

  1. 01

    Guided whitebox

    Full source code plus detailed guidance on the vulnerability and exploit path. The agent still has to execute a working exploit.

  2. 02

    Focused whitebox

    Full source code with limited direction toward a subsystem. The agent must investigate the flaw and develop the exploit itself.

  3. 03

    Focused blackbox

    Limited direction and access to the running target, without source code available.

Training case composition

Share of the 50 training cases by vulnerability class.

Authorization, identity, and scope binding: 72% Accounting and numerical precision: 18% Time validation and signature replay: 4% Business rules and payment validation: 4% SSRF: 2%

72% Authorization, identity, and scope binding

  • Authorization, identity, and scope binding 72%
  • Accounting and numerical precision 18%
  • Time validation and signature replay 4%
  • Business rules and payment validation 4%
  • SSRF 2%

Training on real security work

Base model
GLM-5.3-Flash
Method
Full-parameter fine-tune, GRPO
Training tasks
150, from 50 cases
Role
Worker model, orchestrated

We started with GLM-5.3-Flash and ran a full-parameter fine-tune using GRPO reinforcement learning.

Our security work gives us a growing collection of vulnerability patterns and the context needed to turn them into training environments. We build on functioning open-source applications and protocols, introduce realistic flaws, and preserve the surrounding complexity that makes an investigation meaningful.

Turning those environments into useful training data requires another layer of work. We vary the information available to the agent and calibrate difficulty so that successful investigations are within reach without making the answer obvious. As models improve, we can revisit these environments with less guidance and more demanding combinations of flaws. This ability to keep creating, validating, and refining cases is central to the system we are building.

Building a credible security RL gym takes more than introducing a bug. The environment, available information, and success checks all need to support a coherent investigation. A task can become too leading if its setup makes the answer obvious, or effectively impossible if essential information is missing or the objective cannot be reached. Task prompt wording is a critical part of this balance: naming a function or pointing to a subsystem can dramatically change solve rates. Recreating the complexity of production systems makes this calibration particularly difficult.

This matters for methods such as GRPO, which compare outcomes across a group of rollouts for the same task. With a binary success reward, groups in which every attempt succeeds or every attempt fails provide no contrast in that outcome signal. Cases the model can solve some of the time offer a trainable comparison. Our approach emphasizes realistic environments, clear verification, and difficulty appropriate to the model's current capabilities.

That constraint also shaped what we built apex-flash-1 to do. We designed it as a strong worker model, intended to be orchestrated by a larger model that gives it a focused security task. Training a smaller model on challenges it almost never solves would provide little useful GRPO signal, so we calibrate difficulty to its current capability while preserving the investigation and exploit work. Once focused, the worker can be tenacious: reading code, using tools, testing hypotheses, and verifying a finding. We have found that this division of work lets smaller models uncover bugs that still require deep investigation.

Anatomy of a training case

Each training case is a running system with a defined security objective. We construct the service, seed it with realistic synthetic data, and give the agent a task, tools, and scoped access. Source code is included for whitebox tasks. The diagram below connects environment setup, agent investigation, verification, and training.

Scroll horizontally to explore the diagram.

Apex: Anatomy of a training case Architecture of one Apex security RL training case. Case setup builds a pinned application with an introduced defect, seeds realistic synthetic state, and prepares a separate agent sandbox. The production-like service is a Docker Compose stack on an isolated network with fresh state per attempt: an application and API service, workers, a database, and storage for files and artifacts. The agent container holds the agent harness, shell and investigation tools, task instructions, scoped access with a starting identity, and source code only for whitebox cases. The agent sends network requests to the service and receives responses and observations. An independent verifier outside the agent sandbox checks the objective against the final target state and evidence of the intended route using private checks, and yields a binary success or failure reward that is paired with the trajectory for RL training. Separately, case validation flags apparent passes that bypass the intended target bug and returns those cases to setup for repair. Case setup Build, seed, and prepare agent access Build the application Pinned version with an introduced defect Seed realistic state Synthetic users, permissions, orders, records, workflows Prepare the sandbox Task prompt and tools Starting access and identity Source for whitebox only Case repair Fix flagged cases before reuse Production-like service Docker Compose stack on an isolated network, with fresh state per attempt Application / API service Pinned release with the introduced defect Routes, authentication, business logic, HTTP API Database Persistent app state Seeded realistic data Workers Supporting services Background jobs Storage Uploaded files Generated artifacts reads / writes enqueues jobs stores files updates writes Agent container Separate from the service and the verifier Agent harness Runs the episode, executes tool calls Shell & tools Investigation utilities, HTTP clients Task instructions Objective for this case Scoped access Service connection details Starting identity Source code Included for whitebox cases Omitted for blackbox cases Network requests Responses & observations builds seeds provisions Independent verifier Private checks outside the agent sandbox Objective Met in the final target state Intended route Success came through the target bug Binary reward Success or failure final state & evidence RL training Pair trajectories with verified rewards Scored trajectory Agent actions and observations, paired with a success or failure reward Used for reinforcement learning trajectory reward Case validation Review apparent successes for unintended bypasses Case-quality flag Send the case for repair and re-checking, separately from the runtime reward Bypass check Review passing trajectories for success that bypassed the target bug flag passing trajectories Return for repair

Verification and case validation serve different purposes. The verifier checks whether the agent achieved the objective through the intended bug and assigns a reward. We also review apparent successes for unintended shortcuts, then repair and recheck affected cases before reuse.

Example: artifact isolation in Forgejo

In one such environment, we recreated a workflow-artifact authorization failure in Forgejo, an open-source code-hosting platform. The bug class came from our discovery pipeline, and the scenario was constructed around systems we had encountered in production security work.

The agent began with access to one running workflow. A separate workflow held a protected artifact. The security boundary was straightforward: permission to retrieve one workflow’s outputs should not grant access to another’s.

We introduced a defect in how artifact access is authorized. A signed download URL omits part of the workflow context that the application later uses to select an artifact. Understanding the bug required connecting two parts of the system: what the signature authorizes and what the download handler actually retrieves.

The agent investigated that mismatch and demonstrated its effect against the running application. The verifier checked whether it recovered the protected artifact’s contents, giving us a concrete outcome to train against.

This reflected an authorization failure we encounter in production: individual components appear reasonable in isolation, but disagree about the scope of a permission. Recreating that pattern in a functioning system preserves the context that makes it challenging to investigate.

Performance and cost

The result is a fine-tune of GLM-5.3-Flash that consistently outperforms the base model by 5-8% on our internal held-out evaluation. Across the full 60-task run, we estimate a cost of $2.38 for apex-flash-1, compared with $4.56 for GLM-5.3-Flash and $74.68 for Opus 5 High (using provider pricing).

Public cyber evals are useful, but they still simplify security work into relatively clean, academic tasks. We built our own environments to look much closer to the work we actually do for customers, including real exploitable bugs we have found, responsibly disclosed, and in many cases received bounties for.

We have post-trained a range of open models, from roughly 27B parameters up to models approaching 1T. apex-flash-1 is the first one we think meaningfully moves the Pareto frontier for real-world cyber tasks. That balance of performance and cost matters if we want capable security agents to become cheap enough to scale to every company, not just the few that can afford frontier-model economics.

  • apex-flash-1

    Our model

    Pass@1: 66.7%

    40 of 60 tasks

    Run cost
    $2.38
  • GLM-5.3-Flash

    Base model

    Pass@1: 60.0%

    36 of 60 tasks

    Run cost
    $4.56
  • Claude Opus 5 High

    Reference

    Pass@1: 71.7%

    43 of 60 tasks

    Run cost
    $74.68

Pass@1 and estimated run cost

Each model ran the full 60-task set once. Costs are estimated using provider pricing.

50% 60% 70% 80% $100 $10 $1 Estimated cost for the full 60-task run (USD) Pass@1 apex-flash-1: 66.7% pass@1, $2.38 for the full 60-task run using provider pricing apex-flash-1 · 66.7% $2.38 GLM-5.3-Flash: 60.0% pass@1, $4.56 for the full 60-task run using provider pricing GLM-5.3-Flash · 60.0% $4.56 Claude Opus 5 High: 71.7% pass@1, $74.68 for the full 60-task run using provider pricing Claude Opus 5 High · 71.7% $74.68

Road ahead

This release is a preview of a broader training effort. We are already training on new environments built from years of novel security work, including bugs we have found and responsibly disclosed to companies. We are also expanding the mix of systems, vulnerability classes, and multi-step exploit chains, making the training cases more demanding as the models improve, and training across different agent harnesses so these capabilities carry over across the tools researchers actually use. The expanding data mix includes environments built from public CVEs.

Explore Apex to see how this work supports production security.