Artificial intelligence is currently transforming many areas of cybersecurity. AI models can now analyze code, identify vulnerabilities, and even operate security tools. In pentesting, the discussion increasingly focuses on new attack capabilities, autonomous security assessments, and significant productivity gains.
Together with our Pentesting team, we take a closer look at current developments and examine which changes are already noticeable in practice, where AI delivers real value, and where reality still falls short of expectations.
Recommendations for mid-sized organizations
AI makes pentesters more productive
Many tasks that previously had to be performed manually can now be supported or accelerated, at least to some extent.
- Analyzing large code bases
- Searching for potential vulnerabilities
- Creating small test scripts
- Researching known attack techniques
- Documenting test results
This delivers noticeable efficiency gains, particularly for recurring or time-consuming activities. In the day-to-day work of pentesters, this is especially evident in so-called "throwaway scripts" – small utilities created for a specific testing scenario. Many of these tasks can now be completed much faster. The actual security assessment, however, remains firmly in human hands.
This partly contrasts with current discussions that often portray AI as an immediate productivity breakthrough for the entire security industry. In reality, it is primarily structured and repetitive tasks that currently benefit from automation. The most challenging aspects of a penetration test can only be delegated to AI to a limited extent.
AI handles routine tasks – experience remains the bottleneck
This development is particularly visible in vulnerability analysis and exploit development. With modern large language models, it is now possible in some cases to generate proof-of-concept exploits or attack tools based on known vulnerabilities. The term "vibecoding" has emerged to describe the practice of creating code through collaboration with an AI system.
For organizations, this means one thing above all: attackers may need less time and potentially less specialized knowledge for certain tasks than they did in the past.
At the same time, practical experience shows that the truly critical questions still require expertise:
- Is the vulnerability actually exploitable?
- What would be the impact of a successful attack?
- Which systems are affected?
- How realistic is a specific attack path?
Today, no AI system can reliably provide this level of assessment.
Today, AI can easily write small utility scripts or explain attack techniques. That saves time. But the most interesting parts of a penetration test happen where experience, creativity, and contextual knowledge are required. Those aspects are difficult to automate.
Autonomous Pentesting: significant potential, but no revolution yet
Another topic receiving considerable attention is autonomous pentesting. In this approach, systems independently attempt to analyze environments and develop attack paths. Modern agent-based systems can already coordinate tools, evaluate results, and derive new actions from their findings. Examples include commercial platforms such as Horizon3.ai NodeZero and research projects such as PentAGI and RedAmon. We are currently analyzing and benchmarking these approaches to assess their practical effectiveness. The results will be presented in a future article.
Public discussions sometimes create the impression that traditional penetration testing could become largely automated in the near future. However, our experience from real-world engagements paints a much more cautious picture. While current solutions already achieve impressive results in clearly defined scenarios, we are still a long way from fully autonomous security assessments.
In practice, these approaches continue to face limitations whenever complex environments, custom business applications, or creative attack scenarios come into play. Situations that require contextual knowledge, creativity, or an understanding of business processes remain particularly challenging - precisely the areas that often determine the value and relevance of pentest results.
Many discussions about AI in cybersecurity focus on supposedly massive productivity gains. My view is more nuanced. Yes, some tasks can be completed faster. But the more important question is: How trustworthy are the results? That is exactly where many current approaches still reach their limits.
This aspect is frequently underestimated in the current debate. An autonomous system may be able to identify potential vulnerabilities. For organizations, however, the crucial question is whether those findings are traceable, reproducible, and reliable - especially in regulated environments. Security measures must often be justified to customers, auditors, regulatory authorities, or certification bodies. That is why documentation, transparency, and methodological rigor are often just as important as the vulnerability itself.
AI is becoming part of established pentesting tools
The trend is not limited to specialized AI vendors. Established security tools are increasingly integrating AI-powered capabilities as well. Platforms used to analyze web applications, source code, or networks already assist pentesters in identifying potential vulnerabilities and evaluating large volumes of data. One example is Burp Suite, one of the industry standards for web application penetration testing, which now includes AI-assisted features.
This is noticeably changing the way pentesters work. Less time is spent on routine tasks, while more time becomes available for risk assessment, developing realistic attack scenarios, and analyzing business impact. However, integrating AI does not automatically result in better security assessments. In many cases, existing workflows simply become more efficient. Whether this ultimately leads to more reliable results still depends heavily on methodology, experience, and expert judgment.
AI itself is becoming a test target
At the same time, a new field of work is emerging alongside traditional pentesting processes. Organizations are increasingly integrating AI capabilities into applications, customer portals, support processes, and internal systems. This creates new attack surfaces that cannot be fully addressed through conventional security assessments. Key risks include prompt injection, data leakage through AI systems, and manipulation of AI-generated outputs.
Real-world incidents and current research demonstrate that these risks are far from theoretical. Prompt injection attacks are now considered one of the most significant security threats to AI applications and are ranked as the number one risk for LLM-based systems by OWASP. At the same time, studies and security assessments show that attackers can, under certain conditions, extract personal or confidential information from AI-powered applications.
As a result, specialized security assessments for AI and LLM systems are becoming increasingly common. Organizations such as Germany's Federal Office for Information Security (BSI), the Alliance for Cyber Security, and OWASP are actively developing testing methodologies and security requirements for these technologies.
Learn More:
- Allianz für Cyber-Sicherheit: Leitfaden für Penetrationstests von Large Language Modellen
- OWASP: Top 10 for LLM Applications
- BSI: KI-Sicherheit und sichere KI-Anwendungen
For organizations, this means: if AI is used productively, the security of those AI-powered and AI-assisted systems should also be assessed on a regular basis.
The real challenge begins after the pentest
A much more important question than how vulnerabilities are discovered is often how quickly organizations respond to them. AI accelerates vulnerability discovery and therefore increases the likelihood that security flaws will be identified and exploited more quickly by attackers. AI-assisted attacks and automated assessments place significant pressure on existing vulnerability management processes.
This leads to a clear conclusion for mid-sized organizations: effective and efficient Vulnerability Management and Patch Management are becoming even more critical. Not because AI creates entirely new risks, but because the time between vulnerability disclosure and potential exploitation is becoming increasingly short.
The future of pentesting: Hybrid
Current developments make one thing clear: neither purely manual nor fully autonomous security assessments represent the most effective approach today. The greatest benefits arise when AI and experienced pentesters combine their respective strengths.
A hybrid model is emerging:
- AI handles analysis and routine tasks.
- Security tools are becoming increasingly automated.
- Pentesters can assess larger environments more efficiently.
- New testing methodologies for AI systems are emerging.
- Documentation and reporting can be partially accelerated with AI assistance.
At the same time, the most important tasks remain uniquely human:
- Assessing risks
- Recognizing relationships and dependencies
- Evaluating business impact
- Translating pentest findings into actionable insights for management, business units, and decision-makers
The often-discussed vision of fully autonomous security assessments remains exactly that for now: a vision. Organizations should avoid both AI hype and blanket skepticism. The greatest value currently comes from using automation to support human expertise – not replace it.
SUMMED UP
4 Recommendations for Mid-Sized Organizations
The discussion around AI in pentesting is often driven by technological possibilities. For organizations, however, a much more practical question remains:
What does this mean for their cybersecurity strategy?
The following topics should be on the agenda now:
#1
1. Align pentesting with change
As vulnerabilities can be identified and potentially exploited more quickly, annual snapshots provide less value than they once did.
Recommendation: Do not rely exclusively on fixed testing intervals. New applications, major releases, cloud migrations, or significant infrastructure changes should always trigger additional security assessments.
#2
2. Use automated testing strategically – but don’t confuse it with pentesting
Automated security tools can accelerate vulnerability identification and effectively bridge the gap between traditional penetration tests. However, evaluating realistic attack paths, potential impact, and business risks still requires experience and contextual understanding.
Recommendation: Use automated assessments as a valuable complement, but leave the interpretation of results to experienced pentesters.
#3
Integrate pentesting more closely with Vulnerability and Patch Management
The true value of a penetration test is not the report itself, but the remediation of the vulnerabilities it identifies.
Recommendation: Systematically prioritize and track findings from pentests and integrate them into existing Vulnerability Management and Patch Management processes.
#4
4. Include AI applications in the pentesting scope
Many organizations regularly assess web applications, networks, or cloud services. AI applications, however, are often still excluded, even though they introduce entirely new attack vectors.
Recommendation: Include AI-powered applications, chatbots, copilots, and LLM-based systems as a standard component of penetration testing and risk assessments.