:novee-gym: Elite AI hackers aren’t born. They’re trained.

Join the Novee Gym

:novee-gym: Elite AI hackers aren’t born. They’re trained.

Join the Novee Gym

Field Debrief: The Black Hat USA 2026 Trends CISOs Can’t Afford to Miss

Offensive AI dominated Black Hat this year. Here’s our breakdown, filtered through the noise, and written for the people defending real environments.

Novee Team

6 mins

Explore Article +

Offensive AI dominated Black Hat this year. Here’s our breakdown, filtered through the noise, and written for the people defending real environments.


Offensive security is no longer a niche corner of the show, reserved for hackers alone. One analysis of the floor counted 26 penetration testing and offensive security vendors among 470 exhibitors. Once you fold in the AI SOC crowd (close to 10%) and the AI model security and red teaming vendors, roughly 20% of Black Hat was building AI for offensive operations (the full breakdown is here). 

A year ago, this was a handful of startups. Now it has its own crowded aisle.

Takeaways from Black Hat 2026 for Security Leaders

Just as noisy as Black Hat is the post-show recap scene; analyst reports and influence breakdowns aplenty. Here’s our 2-min recounting of the offensive security AI landscape, as told by the biggest hacker conference of the year:

  • Finding is table stakes; proving and fixing is the real test. Most platforms still hand exploitability and remediation back to you. Make a vendor prove a finding is genuinely exploitable and show you the retested fix, not just a longer list of issues.
  • Safety is top of mind. With the recent rash of AI sandbox escapes (Open AI’s Hugging Face Incident and Anthropic’s unauthorized hacks), guardrails for offensive AI were front of mind for security leaders at the event. Novee co-founder and CEO Ido Geffen breaks down the importance of building guardrails not only into the models, but into the harnesses around them. See the full interview here.
  • Ask whether the model is purpose-trained or rented. General-purpose frontier models don’t automatically win at specialized offensive work, and purpose-trained smaller models often beat them (our benchmark). The question that cuts through the AI marketing is whether a vendor runs a purpose-trained offensive model or orchestrates a frontier LLM.
  • Push for consolidation. The clear appetite this year was fewer tools doing more. Count the handoffs in your current workflow across scanning, exposure management, pentesting, and remediation, then favor platforms that keep testing, validation, and fixing in one loop.
  • Put your AI agents on the attack-surface map. Agent trust and the AI supply chain drove much of the week, including research showing RCE-class flaws across widely used coding agents. Inventory where AI coding agents run in your pipelines and treat the trust boundaries between them as testable attack surfaces.
  • Security can’t hide behind obscurity anymore. AI is hypercharging speed and breadth, plumbing hidden corners of code that human attackers haven’t bothered to check. This is revealing, at scale, critical security issues in the “boring” plumbing behind modern software. Our own threat research into pulling off RCE in Enterprise Java shows one example of how middleware is as viable an attack surface as the edge.

If you’re curious how these trends map to a practical offensive security solution, we’ve collected the eight questions worth asking any AI pentesting vendor in the Novee Buyer’s Guide, the definitive resource for making an informed call on the category. Get your copy here.

Inside the Novee Gym

Elite AI hackers aren’t born, they’re trained.

Training offensive AI is the whole job here, so at Black Hat we put that work on display with live demos of attacker-level AI running against real systems, surfacing validated exploit paths and the fixes that close them, with no theoretical findings and no scanner output. 

Our AI hacker clears PRs in the Novee Gym every day. At Black Hat, we got to give some humans the opportunity to do the same.

On the Black Hat stage

Novee’s founding research team gave two briefings on the types of attacks that are not only possible thanks to AI, but prevalent.

Lidor Ben Shitrit and Assaf Levkovich presented Pre-Auth RCE in Enterprise Java: When Middleware Becomes the Exploit. Twelve vulnerabilities and multiple pre-auth remote code execution chains in enterprise Java middleware that sits under thousands of applications and rarely gets scrutinized. Full write-up here.

Elad Meged presented Trusted Enough to Run: Breaking AI Agents in Official Workflows, timely given the week’s focus on AI safety. He showed how trust breaks down between the stages of official AI agent workflows across Claude Code, Gemini CLI, and Codex, where one component marks an attacker-influenced state as safe and the next consumes it with more authority, ending in remote code execution and supply chain compromise. If you run these coding agents in your own pipelines, the same pattern likely applies to you. The details, including the CVE, are here.

Mobile coverage across every side of the application attack surface

We came with research, and we came with new updates to our platform.

Novee now tests mobile applications, making us the first AI pentesting platform to cover the whole application attack surface across web, mobile, and the APIs underneath both. 

Attackers chain a mobile weakness into an API and the API into your data, so leaving mobile untested leaves that path open. Because Novee works from inside the running app, the encrypted, pinned traffic that other tools stop at becomes a fully testable backend. Upload an Android package and proven findings land in hours, mapped to OWASP MASVS, alongside your web results. 

The first research pass across twenty Android apps is up now. You can read more about how we run our mobile application tests here.

What’s next

None of this slows after Vegas. The coverage map keeps growing, with mobile one step in a plan to test every surface an attacker can reach. We’re still pushing hard on the part that sets us apart, proving what’s genuinely exploitable and pairing each finding with a stack-specific, double-validated fix. 

The Novee Gym is closed to the public (for now), though the training never really stops. If we missed you in Las Vegas, or the briefings and the mobile launch left you wondering what Novee would surface in your environment, the fastest way to find out is to watch it run. Book a demo, and step into the Gym.

Stay updated

Get the latest insights on AI, cybersecurity, and continuous pentesting delivered to your inbox