Chat
Engineering · MapleScholar Plus

The Autonomous Digital Heist: How AI Agents Are Tested on Real-World Cyber Attacks

Isolated security benchmarks test whether AI can write single malicious scripts; CyberBench measures whether autonomous software agents can execute multi-day corporate cyber attacks from start to finish. By stress-testing AI models on complex, realistic enterprise penetration scenarios, cybersecurity researchers established the first realistic yardstick for autonomous digital threats.

Author
Linus Folkerts et al.
Published
2026
Journal
arXiv (Cornell University)
Last updated
September 2026
The Autonomous Digital Heist: How AI Agents Are Tested on Real-World Cyber Attacks

In corporate IT security, companies spend millions hiring ethical hacking teams to run annual penetration tests against their corporate firewalls. Because manual penetration testing is slow and expensive, networks remain vulnerable to new security exploits for months between audits.

Cybersecurity researchers created an automated cyber attack simulator called CyberBench. Operating like an autonomous digital heist crew, AI agents are dropped into enterprise networks to see if they can scan firewalls, sneak between corporate servers, elevate their privileges, and extract confidential data without human assistance.

This benchmark provides corporate defenders with a continuous automated stress test. By identifying security vulnerabilities before criminal gangs exploit them, by benchmarking defensive AI shields, and by running daily autonomous security fire drills, AI cybersecurity hardens enterprise networks.

Reference

Folkerts, L., Payne, W., Inman, S., Giavridis, P., Skinner, J., Deverett, S., Aung, J., Zorer, E., Schmatz, M., Ghanem, M., Wilkinson, J., Steer, A., Hong, V., & Wang, J. (2026). Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios (Version 3). arXiv.

Title

Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios

Abstract

We evaluate the autonomous cyber-attack capabilities of frontier AI models on two purpose-built cyber ranges-a 32-step corporate network attack and a 7-step industrial control system attack-that require chaining heterogeneous capabilities across extended action sequences. By comparing seven models released over an eighteen-month period (August 2024 to February 2026) at varying inference-time compute budgets, we observe two capability trends. First, model performance scales log-linearly with inference-time compute, with no observed plateau-increasing from 10M to 100M tokens yields gains of up to 59%, requiring no specific technical sophistication from the operator. Second, each successive model generation outperforms its predecessor at fixed token budgets: on the corporate network range, average steps completed at 10M tokens rose from 1.7 (GPT-4o, August 2024) to 9.8 (Opus 4.6, February 2026). The best single run completed 22 of 32 steps, corresponding to roughly 6 of the estimated 14 hours a human expert would need. On the industrial control system range, performance remains limited, though the most recent models are the first to reliably complete steps, averaging 1.2-1.4 of 7 (max 3).

Cited 1 times · View on doi.org

Continue

Continue Exploring

Ask this paper your own questions, or keep browsing the verified research catalogue.