Clear your external findings backlog

Novee Exploitability Validation

Clear your external findings backlog

Novee Exploitability Validation

Teaching Offensive AI New Movesets

How Novee post-trains offensive reasoning inside the Novee Gym, and turns overlooked small models into capable exploit agents.

Dan Padnos, Head of AI
Or Dancig, AI Research Team Lead

5 mins

Explore Article +

Training Offensive AI In the Novee Gym

The Novee Gym is our continuous training and optimization environment, where our full offensive AI system runs against thousands of real applications. Every exploit gets scored, and the results feed back in so the system gets sharper each cycle. Much of that loop is about routing every task to the best-performing model and tuning the harness around it.

Below, we drop a high-level overview of how we teach our own proprietary AI models new movesets, via the NoveeGym-RL, our post-training pipeline. This is the specific part of the Gym that improves Novee’s proprietary offensive model through reinforcement learning.

Learning on the Field

Reasoning about a vulnerability and exploiting one are different problems. A base model has fragments of offensive capability scattered through its weights, but it’s never been drilled on the full sequence: the long back-and-forth of probing a system, forming a hypothesis, and revising an exploit until it works. Pentesting is a long-horizon task, i.e. hundreds of steps per target, and post-training is how you reinforce the behaviors that carry a model from “can reason about the bug” to “reliably lands the exploit.”

We do it in two stages. First we teach form, by having the model imitate attempts that already worked. Then we let it learn from experience, running full attempts against live targets and keeping what works while throwing out what doesn’t. We also train the model inside the exact harness it’ll be deployed in, with the same tools and skills.

Scored on what lands, not what sounds good

A quick way to validate results would be to have one model generate the report, and another one grade it, but we don’t do that; it rewards persuasive prose over working exploits, and a language-model judge is exactly the kind of soft signal a model learns to game. 

A deterministic validator plays referee instead, taking the exploit the model wrote, running it against the live application, and confirming whether the attack truly fired. 

This is reinforcement learning from verifiable rewards, and because the signal is mechanical rather than human, there’s no ceiling on how far the model can climb.

A self-perpetuating loop

A model trained this way finds fresh zero-days in the wild, and those become training data for the next round, so the moveset compounds.

The hard part, though, was never the hacking. The trickiest challenges came from everything around the model: the machinery that has to work in concert for training to run at all. Every attempt needs its own fresh, isolated copy of a real web application to train against, and at peak there are thousands running at once. In one campaign, those live environments cost more than the GPUs, somewhere between 50 and 70 percent of the bill. Good training data isn’t free, and building the Gym turns out to be most of the work.

Finding potential anywhere

Frontier models make strong offensive drivers, but they’re slow and costly to run, and the price of pointing them at every application, continuously, climbs fast. We wanted to know how far a small, open-weight model could go instead.

Out of the box, not far. In our testing, the strongest small models landed around 55% of exploits on our benchmark, while the smallest, a 4-billion-parameter model, came in close to useless at about 2%. On paper, not a prospect worth signing.

Then we put them through the Gym. With our own training recipe, the numbers moved sharply, and the Qwen 4B model went from 2% to 51%, the 9B from 10% to 56%, and a 27B model from 55% to 82%. Training even changed their temperament. Where the untrained versions gave up on hard targets, the trained ones pushed through and kept swinging, cracking vulnerabilities the base models never could. A few of those were real zero-days in live open-source software.

The real payoff is in the tradeoff between capability and cost. A fine-tuned 27B model reached roughly 80% of a frontier model’s exploit rate for about 15% of the cost, small enough to run efficiently at scale, including on-premises for teams with strict data-privacy requirements, and specialized enough to do serious offensive work. That’s what “finding potential anywhere” comes down to. The Gym develops capability wherever it’s hiding, including in a model most people would pass over.

Finding the balance between performance and scale

The goal behind the research is straightforward. Continuous, attacker-level testing across an entire environment only works if the economics work, and specialized models that deliver frontier-grade offensive capability at a fraction of the cost are what make deep coverage viable everywhere.

The Gym never closes

An offensive AI is only as good as its current moveset, and one that stops growing starts to fall behind. The point of the Novee Gym is that it never stops. Every cycle brings new opponents and new techniques to learn from, all of it scored on the only thing that counts, which is what actually lands against a live target.

That’s how you build an AI hacker that keeps getting better, and an AI defender that closes the loop right behind it.


Dive deep into how the Novee Gym works: 

And see the results for yourself in your own environment.

Stay updated

Get the latest insights on AI, cybersecurity, and continuous pentesting delivered to your inbox