Clear your external findings backlog

Novee Exploitability Validation

Clear your external findings backlog

Novee Exploitability Validation

Novee Joins Fireworks Specialized Intelligence Index to Advance Real-World AI Evaluation

PWNBench-v0.1, Novee's benchmark for agentic pentesting of live web apps, is now live on the Fireworks Specialized Intelligence Index. See our findings across 11 models.

Novee Marketing

5 mins

Explore Article +

PWNBench-v0.1, our benchmark for agentic pentesting of live web applications, is now live on the Specialized Intelligence Index by Fireworks.

When a new frontier model ships, the whole industry refreshes the same public leaderboards. They’re useful for tracking raw capability, but they don’t tell you whether a model can do your job. 

As AI becomes more specialized and agentic, that distinction matters more. Evaluating AI increasingly means understanding how models perform on complex, real-world work – not just how they score on generalized tasks. When that job is offensive security, relying on results from existing public benchmarks only tell part of the story.

A model can top a cyber eval by recalling a CVE it saw in training. Real offensive work looks nothing like that. A tester gets ordinary credentials, a few docs, no hints about what’s broken, and a running system they have to break into and validate for real. Measuring that takes a different kind of benchmark, one built by people who do the work.

That’s why we built PWNBench, and it’s why we’re launching it today as a Cybersecurity benchmark on Fireworks’ Specialized Intelligence Index (SII).

The SII is Built to Measure Real Work

The SII is Fireworks’ answer to the generalized, static reality of public leaderboards. A high score proves a model performs on a fixed, static task set; it says nothing about whether the model can handle messy inputs, ambiguity, multi-step workflows, and real business constraints. 

The SII is built around a principle we share: meaningful AI evaluation has to reflect the work the AI is actually being asked to do. It brings together benchmarks built by practitioners who understand a specific job and the standard it has to meet, across domains like healthcare, legal, and cybersecurity.

PWNBench-v0.1 is our contribution to that cybersecurity category. It’s a robust slice of the internal evaluation suite we run every day, now public for the first time.

The index launches with benchmarks from teams that set the standard in their own fields, including Harvey’s Legal Agent Benchmark for legal work and Doximity’s BedsideBench in healthcare. These are groups that build and run their evals against live production work, held to rigorous standards. We’re excited to include PWNBench in that conversation.

Why we Built PWNBench

There was no industry-standard benchmark for agentic greybox pentesting of live web applications. Some cyber evals measure exploitation but never discovery, some assume source-code access, and many lean on public CVEs the models have almost certainly absorbed in training. PWNBench closes those gaps.

Each instance gives an agent the same material a real tester starts with: ordinary account credentials, a user manual, API docs, and a live application. From there it has to return a set of validated security issues, with no hints about what’s broken. The targets are forks of large, well-maintained open-source projects like developer platforms, CRMs, observability tools, productivity apps, and ERPs.

We score every reported issue against a ground-truth database our team built by hand, more than 400 labels covering all OWASP Top 10 risks, leaning heavily on zero-days plus novel vulnerabilities we injected and verified as exploitable. That keeps the benchmark off public CVEs and closer to the work real attackers actually do.

We report precision alongside recall, and we treat test-time compute as a real variable, varying reasoning effort and parallel runs instead of sampling each model once at its default setting.

Our Findings Across Leading Models

v0.1 covers 13 models, closed frontier APIs and open weights alike, all run through the same deliberately thin harness so the comparison stays fair. 

Here are the highlights from our early learnings:

  • There’s no single best model. The efficient frontier moves depending on whether you weight recall, precision, or severity. The right model to reach for depends on the job in front of you.
  • Coverage costs money, and the exchange rate varies wildly. Claude Opus 5 buys the highest recall, roughly 51% for about $1,400 in API spend at k=3, while Kimi K3 reaches 42% for $209. Adding reasoning effort and parallel runs moves a model along the recall/cost curve about as much as switching models does.
  • Precision is a separate ranking from recall. High coverage doesn’t mean a clean report. On precision, Grok 4.6 and Claude Opus 4.8 sit in the high-70s to low-80s, which matters more than it sounds, because a report full of high-severity false positives is useless to a defender.

The full methodology, every model’s results, and the interactive charts live in the SII deep-dive into PWNBench-v0.1.

Bringing Our AI Evaluation Discipline Into the Open

PWNBench is the public face of an AI evaluation discipline we run internally every day. Inside the Novee Gym, we continuously benchmark our proprietary offensive model, frontier models, and our harness against thousands of real applications, measuring how different models and systems perform across recall, precision, quality, speed, safety, and cost.

The results make that case better than we could on our own. When there’s no single best model, precision is a separate problem from recall, and adding compute moves the needle as much as swapping models, betting everything on one frontier model is the wrong bet. 

By making part of that work public through PWNBench and the SII, we hope to contribute what we’re learning in offensive security to the broader conversation about how specialized AI should be evaluated.


Explore PWNBench-v0.1 on Fireworks’ Specialized Intelligence Index

Read the full methodology and results 

See what attackers already see in your environment. Book a demo →

Stay updated

Get the latest insights on AI, cybersecurity, and continuous pentesting delivered to your inbox