AI Models vs. Hackers: The Cybersecurity Arms Race

Thomas Wolf (Hugging Face) and Uri Rolls (Arithmetic) discuss training AI models to out-think cyber attackers using novel benchmarks and the potential of open-source AI.

8 min read
Thomas Wolf and Uri Rolls on stage discussing AI and cybersecurity.
AI Engineer

Visual TL;DR. Hackers' Advantage drives need for AI in Cybersecurity. AI in Cybersecurity requires MaskOff Benchmark. MaskOff Benchmark enables AI-Powered Defense. Open-Source AI powers AI-Powered Defense. Data Quality improves Open-Source AI. AI-Powered Defense leads to Future Cybersecurity.

  1. AI in Cybersecurity: AI offers unique opportunity to rebalance scales in favor of defenders
  2. Hackers' Advantage: attackers empowered by advanced tools, creating an economic shift
  3. MaskOff Benchmark: new benchmark tests AI's reasoning in dynamic cybersecurity environments
  4. Open-Source AI: crucial for training models to out-think cyber attackers effectively
  5. Data Quality: high-quality data is vital for robust AI model training and performance
  6. AI-Powered Defense: AI models can understand and interact with complex attack scenarios
  7. Future Cybersecurity: AI will tackle complex challenges, enhancing defender capabilities
Visual TL;DR
Visual TL;DR, startuphub.ai AI in Cybersecurity requires MaskOff Benchmark. MaskOff Benchmark enables AI-Powered Defense. Open-Source AI powers AI-Powered Defense requires enables powers AI in Cybersecurity MaskOff Benchmark Open-Source AI AI-Powered Defense From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI in Cybersecurity requires MaskOff Benchmark. MaskOff Benchmark enables AI-Powered Defense. Open-Source AI powers AI-Powered Defense requires enables powers AI inCybersecurity MaskOff Benchmark Open-Source AI AI-PoweredDefense From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI in Cybersecurity requires MaskOff Benchmark. MaskOff Benchmark enables AI-Powered Defense. Open-Source AI powers AI-Powered Defense requires enables powers AI in Cybersecurity AI offers unique opportunity to rebalancescales in favor of defenders MaskOff Benchmark new benchmark tests AI's reasoning indynamic cybersecurity environments Open-Source AI crucial for training models to out-thinkcyber attackers effectively AI-Powered Defense AI models can understand and interact withcomplex attack scenarios From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI in Cybersecurity requires MaskOff Benchmark. MaskOff Benchmark enables AI-Powered Defense. Open-Source AI powers AI-Powered Defense requires enables powers AI inCybersecurity AI offers uniqueopportunity torebalance scales in… MaskOff Benchmark new benchmark testsAI's reasoning indynamic… Open-Source AI crucial fortraining models toout-think cyber… AI-PoweredDefense AI models canunderstand andinteract with… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Hackers' Advantage drives need for AI in Cybersecurity. AI in Cybersecurity requires MaskOff Benchmark. MaskOff Benchmark enables AI-Powered Defense. Open-Source AI powers AI-Powered Defense. Data Quality improves Open-Source AI. AI-Powered Defense leads to Future Cybersecurity drives need for requires enables powers improves leads to AI in Cybersecurity AI offers unique opportunity to rebalancescales in favor of defenders Hackers' Advantage attackers empowered by advanced tools,creating an economic shift MaskOff Benchmark new benchmark tests AI's reasoning indynamic cybersecurity environments Open-Source AI crucial for training models to out-thinkcyber attackers effectively Data Quality high-quality data is vital for robust AImodel training and performance AI-Powered Defense AI models can understand and interact withcomplex attack scenarios Future Cybersecurity AI will tackle complex challenges,enhancing defender capabilities From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Hackers' Advantage drives need for AI in Cybersecurity. AI in Cybersecurity requires MaskOff Benchmark. MaskOff Benchmark enables AI-Powered Defense. Open-Source AI powers AI-Powered Defense. Data Quality improves Open-Source AI. AI-Powered Defense leads to Future Cybersecurity drives need for requires enables powers improves leads to AI inCybersecurity AI offers uniqueopportunity torebalance scales in… Hackers'Advantage attackers empoweredby advanced tools,creating an… MaskOff Benchmark new benchmark testsAI's reasoning indynamic… Open-Source AI crucial fortraining models toout-think cyber… Data Quality high-quality datais vital for robustAI model training… AI-PoweredDefense AI models canunderstand andinteract with… FutureCybersecurity AI will tacklecomplex challenges,enhancing defender… From startuphub.ai · The publishers behind this format

In a compelling discussion at the AI Engineer World's Fair, Thomas Wolf of Hugging Face and Uri Rolls of Arithmetic explored the burgeoning intersection of artificial intelligence and cybersecurity. They highlighted how AI, rather than signaling the end of cybersecurity as we know it, presents a unique opportunity to rebalance the scales in favor of defenders.

AI Models vs. Hackers: The Cybersecurity Arms Race - AI Engineer
AI Models vs. Hackers: The Cybersecurity Arms Race — from AI Engineer

The Expanding Frontier of AI in Cybersecurity

Wolf opened the session by expressing his excitement about cybersecurity as a fertile ground for AI exploration. He noted that the field is far broader than many might assume, with the potential for AI to tackle complex challenges that have long eluded traditional methods. The core of their discussion revolved around a new benchmark developed by Arithmetic, which aims to test AI models' ability to understand and interact with dynamic environments, akin to tasks seen in projects like ARC-AGI.

Rolls elaborated on the economic shifts in cybersecurity, where attackers, empowered by advanced AI, can now target a wider range of systems with greater efficiency. He drew an analogy to defending a house, where defenders must secure every entry point, while attackers only need to find a single vulnerability. The increasing sophistication of AI amplifies this asymmetry, making it imperative for defenders to develop equally advanced AI capabilities.

The "MaskOff" Benchmark: Testing AI's Reasoning in Cybersecurity

The central piece of their presentation was the "MaskOff" benchmark, designed to evaluate AI models on access control, a fundamental aspect of cybersecurity. "We cannot capture all 'cyber' in a single benchmark," Rolls explained. "That's like saying swimming, F1, and basketball are the same sport." The benchmark focuses on long-horizon exploitation chains within realistic, multi-service environments, where models must not only identify vulnerabilities but also reason about the system's state and act accordingly.

Wolf highlighted the surprising difficulty these benchmarks pose for current AI models. "The current models, even though they are really good, they build... they can't really build a dynamic model of what's happening in the world," he stated. The benchmark's design requires models to understand cause and effect within complex systems, a task where even state-of-the-art models show significant limitations, often achieving only 1-2% success rates.

The Crucial Role of Open-Source Models and Data Quality

A key theme of the discussion was the significant potential of open-source AI models in bolstering cybersecurity defenses. Wolf countered the often-binary view that closed-source models are inherently better for security, emphasizing that open-source models are a vital part of the solution. "I think what we want to show today is that open source models are one part of the solution to cyber security challenges today," he said.

Rolls detailed the methodology behind their benchmark creation, stressing the importance of human vulnerability researchers in identifying and verifying zero-day exploits in open-source software. These real-world scenarios are then used to create complex, multi-service environments within Docker containers. Crucially, models operate in a black-box setting, meaning they don't have access to the code or prior knowledge of the vulnerabilities, forcing them to rely on their reasoning and problem-solving capabilities.

The Future of AI-Powered Cybersecurity

The presentation concluded with a look towards the future, where AI models are trained to not only identify but also out-think sophisticated cyber attackers. The goal is to develop AI that can reason across the entire attack surface faster than human adversaries. Both speakers expressed optimism, drawing parallels to the transformation AI has brought to coding, and believing that similar advancements are within reach for cybersecurity through high-quality evaluations and collaborative efforts.

"The only way to replace the old stack is through the models," Wolf asserted. "And I think the only way to do that is through a real array of strong open source models and collaboration that we can post-train on and that we can post-train to each network and to each environment as well." This collaborative approach, they believe, is essential for building a future where AI-powered defenses are robust and widely accessible.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.