Clear your external findings backlog
Closing the gap to frontier model accuracy at a fraction of the cost by fine-tuning small, open-weights language models specifically for offensive security.
Training LLMs for pentesting is harder than the hacking itself. Real lessons on prefix breaks, weight sync, silent bugs, and environment costs in RL runs.
Presenting the architecture and design choices underlying Novee's LLM post-training infrastructure for agentic pentesting.
How we turned hundreds of broken open-source apps into deterministic training environments
In live-browser exploit benchmarks, Novee’s 4B-parameter model achieved up to 90% accuracy, outperforming Claude 4 Sonnet and other frontier LLMs by over ~55%.
Modern web apps layer defenses so thick that XSS should be impossible. Attackers still find ways through. We wondered: could we teach AI to probe, adapt, and reason its way…
Get the latest insights on AI, cybersecurity, and continuous pentesting delivered to your inbox