:novee-gym: Elite AI hackers aren’t born. They’re trained.

Join the Novee Gym

:novee-gym: Elite AI hackers aren’t born. They’re trained.

Join the Novee Gym

Attacker-Grade Reasoning with Deterministic Control: How Novee Runs AI Offense with Safety Guardrails

Recent sandbox escapes and unauthorized lateral movement during security testing, from the likes of OpenAI and Anthropic, prove that offensive AI needs robust guardrails.

Omer Ninburg, Co-Founder & CTO
Netta Rager Dan, VP Product

5 mins

Explore Article +

Recent sandbox escapes and unauthorized lateral movement during security testing, from the likes of OpenAI and Anthropic, prove that offensive AI needs robust guardrails. 

The best AI defense can turn into an unexpected AI offense

We have seen real-world examples of a capable offensive model breaking testing protocol to reach its goals, unless something is deliberately put in its way to stop it. That’s why the boundaries placed around these models have taken on new importance.

In July 2026, OpenAI disclosed that during an internal evaluation, two of its models found and exploited an unknown vulnerability to break out of a sandboxed test environment. The models reached the open internet and chained their way into the production infrastructure of Hugging Face. Days later, Anthropic reported three separate incidents in which Claude models reached the internet from inside a testing environment and gained unauthorized access to the live systems of three real organizations.

In both cases, the models were running with their usual safety refusals switched off (so the labs could measure raw capability) and in both cases that capability proved enough to compromise real systems on its own.

Autonomous agents behave differently from the deterministic tools most security programs were built around because they reason, they improvise, and their flows are non-deterministic by design. This lets them find novel risk, but it also lets them cross trust boundaries and break containment. Left unchecked, a creative agent can wander past its boundaries or burn through cost with nothing to show for it.

For a CISO evaluating any offensive AI tool, it’s important to know who holds the leash while it runs against your environment.

Why owning the full stack is the answer

Novee is building the best AI hacker and the best AI defender in one platform, with safety treated as a design constraint from the start, rather than a layer bolted on at the end.

Novee owns the full offensive stack; both the model and the harness it runs in. Because we train the model and build the environment it operates in, we can enforce boundaries at both layers instead of hoping a general-purpose model behaves itself. 

Before any safety mechanism reaches your environment, we test it against a benchmark that includes tripwire scenarios; specific actions the agent must never take. We confirm the agent stops at each one instead of crossing it. We train the AI stack to stay aligned with your own response and policies, so the agent’s behavior reflects what you have authorized, not whatever it decides on its own.

From there, Novee enforces safety in two layers:

Hard guardrails are deterministic controls that live outside the language models themselves, the large and small models alike (LLMs and SLMs). They run in an intercepting proxy that every agent request passes through. Because they sit outside the models, they behave the same way no matter how the agent reasons, and prompt content cannot override them. 

Soft guardrails run through the models to shape the agent’s judgment; they are non-deterministic by nature and steer behavior rather than guarantee it. 

You reach for hard guardrails when something has to be guaranteed, and for soft guardrails when you want to direct the agent’s focus or catch intent that stays technically in bounds but is still not what you want.

Novee security guardrails at a glance

Hard Guardrails: Deterministic and Customizable to Your Environment

  • Traffic allow and deny rules 
  • Custom testing schedules and time-zone windows
  • Custom HTTP headers injected into every request, so your WAF and SIEM can tell authorized testing apart from a real attack
  • Rate limit caps on the request volume the agent can issue
  • Any custom hard guardrail you need

Because these controls are deterministic and run outside the models, a lateral-thinking model cannot reason its way around them, so they act as reliable hard stops.

Soft Guardrails: Steering the Behavior and Judgement of Our Models

  • Novee System Prompts: A baseline set of safety instructions on every run
  • Natural Language Prompts: Per-assessment instructions that direct the agent’s focus – e.g. “focus on the RBAC capabilities”
  • The Gatekeeper Agent: A separate LLM agent that runs alongside the testing agent, watches what it intends to do, and judges whether each action is acceptable before it happens

These are the controls that keep models on-task and in-check.

Assurance you can take to your board

“Novee adapted to our multi-tenant SaaS product within days and now gives us ongoing validation that our tenant isolation holds up against real attacker techniques. That level of assurance is what our customers expect of us, and it’s exactly what we give them through this partnership.”

— Scott Roberts, CISO, UiPath

The question worth putting to any offensive AI vendor has less to do with how clever the model is, and more to do with whether it gets privileged access without guardrails, and who stays in control while it runs against your environment.

Novee is built so that the answer comes easily. You get attacker-grade reasoning, and the boundaries stay firmly in your hands.

The clearest way to judge that is to see it for yourself. Book a live demo and watch Novee run against a target with the guardrails you set, and see the boundaries hold in real time.

Stay updated

Get the latest insights on AI, cybersecurity, and continuous pentesting delivered to your inbox