Comparison of AI Models in Application Vulnerability Detection
In a practical test for vulnerability detection in applications, modern AI models demonstrated significant differences in both effectiveness and cost. GPT-5.5 proved to be the most effective, while DeepSeek V4 Pro was the most cost-efficient—an important factor when scaling cybersecurity tools.
Crius
Exploring the Capabilities of Modern AI Models in Cybersecurity
In one of the most illustrative tests conducted over the past year, researchers evaluated how well various AI models could identify vulnerabilities in applications. For the experiment, a special book review app was created with a deliberately embedded vulnerability: the APK file contained exposed Firebase credentials, allowing direct access to the database while bypassing the protected API. More than a dozen AI models were tasked with discovering and exploiting this vulnerability, each with a $10 budget and a two-hour time limit. The total cost of the experiment amounted to $1,500.
AI Model Testing Results
GPT-5.5 demonstrated the highest effectiveness, successfully completing the task in 7 out of 10 attempts, with an average cost of $9.46 per successful run. In most cases, the model quickly focused on the Firebase vulnerability after analyzing the APK, without being distracted by other app elements or APIs.
DeepSeek V4 Pro achieved the lowest cost per successful attempt—just $0.62—solving the task in 3 out of 10 runs. This is about 15 times cheaper than GPT-5.5, despite a lower success rate. Such a cost difference could be significant when scaling security tools for broader use.
Claude Sonnet 4.6 and Claude Opus 4.8 managed to solve the task in 2 out of 10 cases. Notably, Opus came close to success several times, but sessions were interrupted by security mechanisms.
Gemini 3.1 Pro Preview almost always refused to perform the task, as reflected in its median token count—just 9,000 compared to 100,000 or more for other models. Gemini 3.5 Flash showed similar results, only attempting to solve the task in two cases.
Model Behavior Patterns
During the experiment, it was observed that Chinese AI models more often interacted directly with real databases, while Western models tended to be more cautious, even after identifying the correct approach.
This experiment is not a scientific assessment, but rather a well-documented practical test of the capabilities of modern AI models in conditions close to real-world cybersecurity challenges.
