The gains are real, and they are not the whole story.
There is a version of the AI story in offensive security that has become almost reflexive in how it gets told: the tools got faster, the coverage got broader, and the teams that once spent weeks on manual testing cycles now move in days. That account is accurate as far as it goes, and most practitioners would recognise it, but it tends to stop well short of the parts of the job that have proven resistant to the kind of transformation the headline numbers suggest.
The observable gains are worth taking seriously. AI has expanded the discovery surface considerably, accelerating the identification of vulnerabilities and giving practitioners access to a scale of coverage that pure manual effort could never sustain. More significantly, AI has begun to participate in testing decisions themselves, not just assisting with discrete tasks but operating with a degree of autonomy within defined parameters and given sufficient context. The role of the tool has shifted, and anyone working in this space has felt that shift in practical terms.
The instability nobody put in the brochure
What has been harder to articulate, partly because it sits awkwardly against the dominant enthusiasm for AI-assisted security, is that a significant number of experienced teams have been quietly moving back toward building their own environments in-house. The reason tends not to be scepticism about AI’s capabilities in principle. The reason is operational, specific, and for the teams affected, fairly costly.
LLMs shift their guardrails without announcing it. They remove teams from security programmes without prior warning and introduce instability into workflows that were designed around the assumption of consistency. The disruption that follows is difficult to anticipate and often harder still to explain to the stakeholders who approved the original setup. Teams rebuilding in-house are not expressing nostalgia for older ways of working; instead, they are making a considered judgement that control over their own environment, and the reliability that comes with it, is worth considerably more than the convenience of not having to construct it.
The problems that were always going to survive automation
The speed-and-scale narrative, for all its accuracy on its own terms, tends to leave unexamined the category of problems that AI was never going to touch, because they are not problems that more processing power or broader coverage can solve.
The complexity of working across multiple stakeholder groups and business units has not simplified under AI. Workflow friction in any given organisation is structural, embedded in the way decisions get made and ownership gets distributed, and it is particular enough to each context that generalised automation has very little purchase on it. Increasing the velocity of a testing cycle does not do anything about the internal negotiation required to act on what the testing surfaces, and in most organisations, that negotiation is where the real work is.
The burnout question sits in a similar place. There was an implicit promise running through a lot of the early enthusiasm for AI in security work: that automation would absorb the toil, reduce the cognitive overhead, and give practitioners more space for the judgement-heavy work that actually interests most of them. What has emerged instead is considerably more complicated. Managing tools has its own demands. Prompting well, evaluating outputs critically, and identifying the cases where a model has drifted or produced a confident-sounding mistake all require sustained attention, and for some teams that overhead has settled on top of the existing workload rather than replacing any part of it.
The mindset question
The deepest issue is evaluative, and it runs underneath the more visible debates about capability and reliability. Using AI effectively in offensive security requires a kind of critical fluency that has become more demanding as the tools have become more capable, because the surface area for misjudgement has expanded alongside everything else. A practitioner needs to be able to assess whether the AI is doing a good job in the particular context it is operating in, which is a question that requires enough domain knowledge to evaluate outputs with genuine scepticism rather than deference. They need to distinguish between the cases where autonomous capabilities are solving a problem that actually needed solving and the cases where they are being introduced because they are available and adoption feels like progress. And they need to hold that evaluative capacity across different tools, different contexts, and a landscape that keeps moving.
This is not something AI can supply. It has to precede the AI to be worth anything, and it has to be actively exercised, including through the experience of watching AI get things wrong in ways that are not flagged and not always obvious.
What actually separates good security work from the rest
The teams that navigate this period well are unlikely to be distinguished by the breadth of their AI integration. They will be the ones who stayed clear about what the work actually requires: environments they can control, workflows that do not depend on third-party models behaving consistently, and a critical posture that makes it possible to tell the difference between a tool that is performing well and one that is performing convincingly. That posture is not new to offensive security. It is, in most respects, the same orientation that has always separated work that holds up under scrutiny from work that only appears to until something goes wrong.
Daniel Bechenea
Daniel Bechenea is Product Security Manager at Pentest Tools, where he coordinates, builds, and maintains highly accurate detection and validation capabilities for high-impact vulnerabilities such as Log4Shell and SessionReaper, with a focus on results that hold up in real environments. He is also an OSCP and CRTP-certified penetration tester and he brings that attacker mindset into how he designs scanning logic, evidence, and validation workflows.


