Today in AI: Finding Zero-Days, Data Quality's Multiplier Effect

Today, Ada and Sam dive into how AI is learning to find real software vulnerabilities, the crucial role of data quality as a compute multiplier, and why long-horizon AI agents need better verification. Plus, we explore the infrastructure for autonomous AI engineers and the ongoing debate around Big Tech's AI boom.

6 min read
Today in AI: Finding Zero-Days, Data Quality's Multiplier Effect
🎙️ Listen to this episode of Today in AI

Today, Ada and Sam dive into how AI is learning to find real software vulnerabilities, the crucial role of data quality as a compute multiplier, and why long-horizon AI agents need better verification. Plus, we explore the infrastructure for autonomous AI engineers and the ongoing debate around Big Tech's AI boom.

In this episode

Transcript

Ada: Welcome to Today in AI, I'm Ada.

Sam: And I'm Sam. Today, we're dissecting how AI is getting smarter at finding critical software bugs and why data quality is becoming a compute superpower.

Ada: First up, a fascinating development from David Brumley, detailing how AI models are now reliably discovering real software vulnerabilities, what we call 'zero-days.' He explains that using reinforcement learning sandboxes and deterministic graders allows these AI models to actually pinpoint these elusive flaws. This isn't just theoretical- it's about finding real, exploitable bugs.

Sam: That's a huge leap. Traditionally, finding zero-days is a highly skilled, time-intensive human endeavor. If AI can automate this, even partially, it changes the game for cybersecurity both offensively and defensively. It means we could see more robust software faster, but also potentially more sophisticated attacks. The mechanism Brumley describes- combining reinforcement learning with a 'deterministic grader'- is key to ensuring the AI isn't just guessing, but truly understanding and confirming the vulnerability.

Ada: Exactly. Moving on, Rayan Garg from Theta Software is highlighting a critical challenge for long-horizon AI agents: the need for better verifiers. He argues that current benchmarks for these agents often lack accurate environment design and robust final-state verifiers. This means we might be overestimating their true capabilities without a proper way to confirm their success over extended, complex tasks.

Sam: This really resonates with the broader conversation around AI safety and reliability. If we're building agents to perform complex, multi-step tasks, we absolutely need to be sure they're not just 'appearing' to succeed. Garg's point about environment design is crucial. If the test environment isn't a high-fidelity representation of the real world, an agent's success there might not translate. And without strong final-state verifiers, we can't definitively say if the agent truly achieved its goal, or just landed in a similar-looking state by chance. It's about true intelligence versus superficial performance.

Ada: Absolutely. And speaking of foundational elements, Ari Morcos is making a strong case for data quality as the ultimate compute multiplier. He explains that high-quality data dramatically lowers training and inference costs for AI models. Essentially, better data means you need less compute to achieve the same or even better results.

Sam: This is a concept that's gaining increasing traction, and for good reason. For so long, the focus has been on throwing more compute at the problem- bigger models, more GPUs. But Morcos is articulating that the quality of your input data can be a more efficient lever. If your data is clean, relevant, and well-labeled, the model learns faster and generalizes better, requiring fewer cycles and less energy. It's a fundamental shift in how we think about optimizing AI development, moving beyond just raw processing power.

Ada: It’s a powerful argument for investing more upfront in data pipelines. Next, the co-founders of Emulated, Joseph Wang and Sid Patlollu, are shedding light on what it takes to train truly autonomous AI software engineers. They emphasize the necessity of multi-node, real-cloud environments for this kind of advanced training. This isn't something you can do in a simulated sandbox.

Sam: This ties into the idea of long-horizon agents and realistic environments. To create an AI that can genuinely act as a software engineer- understanding requirements, writing code, debugging, and deploying- it needs to operate in an environment that mirrors the complexity of real-world development. That means interacting with actual cloud services, dealing with dependencies, and navigating the nuances of distributed systems. It's a significant infrastructure challenge, but one that's essential for moving beyond basic code generation to truly autonomous engineering.

Ada: Indeed. And as we scale these models, Ross and Chengxi Taylor of General Reasoning are discussing the infrastructure needed for scaling AI models to long horizons. They touched on topics like Galactica, RLHF, and KellyBench. Their work highlights the immense engineering effort required to push AI capabilities further into complex, multi-step reasoning.

Sam: Their insights are particularly relevant in the context of both Rayan Garg's point about verifiers and Emulated's work on autonomous engineers. Scaling to long horizons isn't just about making models bigger; it's about enabling them to maintain coherence and perform effectively over extended periods and numerous steps. The techniques they mention, like RLHF- reinforcement learning from human feedback- are crucial for aligning these powerful models with human intentions, especially as tasks become more open-ended and complex. It's all about building reliable, capable AI systems for the future.

Ada: Before we wrap up, we have to touch on Ed Zitron's recent warning about what he calls a 'scandalous lie' at the heart of the Big Tech AI boom. He suggests that much of the growth is built on circular financing, with cash-burning startups like OpenAI and Anthropic essentially being funded by the very hyperscalers who then report massive AI-driven cloud revenue.

Sam: It's a provocative take, and one that highlights the intricate, sometimes opaque, financial relationships within the AI ecosystem. Zitron's argument is that this dynamic could be fueling a $1.6 trillion data center bubble, where the demand for compute is artificially inflated by these interconnected investments. While the underlying technology and innovation are real, the financial architecture supporting the current boom merits scrutiny. It’s a reminder that even in the most exciting tech revolutions, economic fundamentals and potential market distortions need to be considered.

Ada: A lot to chew on today. That's all for this episode of Today in AI. For full stories and more in-depth analysis, visit startuphub.ai.

Sam: We'll see you tomorrow.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.