University of Oslo · Norway
Assessing the Reliability of PenTestGPT: A Comparative Case Study Between LLMs in AI-Driven Web Application Testing
This PhD project focuses on evaluating the reliability and effectiveness of PenTestGPT, an AI tool for web application penetration testing. The research will compare its performance against traditional human penetration testers and other AI models. The goal is to understand the capabilities and limitations of Large Language Models (LLMs) in identifying web application vulnerabilities, suggesting fixes, and detecting threats.
The study will involve a case study where both human testers and AI systems, including PenTestGPT, test a set of web applications. Key metrics will include the rate of vulnerability detection, accuracy of findings (minimizing false positives and negatives), and the quality of recommended solutions. The research will also identify challenges LLMs face, such as handling complex attacks or situations requiring human intuition.
The methodology includes setting up vulnerable environments (e.g., specific web applications or capture-the-flag platforms) for testing. PenTestGPT will be used with various AI models for tasks like reconnaissance, scanning, and reporting. Human penetration testers will simultaneously assess the same environments using standard tools like Burp Suite and OWASP ZAP.