Modern software systems are a labyrinth of interconnected components, generating an overwhelming deluge of data. When an outage or performance degradation strikes, identifying the true root cause amidst this cacophony of signals becomes a monumental, often manual, task. This critical challenge, a persistent thorn in the side of enterprise operations, is precisely what Traversal AI, a startup founded by Anish Agarwal and Raaz Dwivedi, aims to tackle with a sophisticated blend of causal machine learning and reinforcement learning.
In a recent Latent Space podcast, co-hosted by Alessio Fanelli of Kernel Labs and Swyx of Smol AI, Anish and Raaz, both veterans of MIT and academia, detailed their journey from deep research to building a product addressing this complex problem. Their backgrounds, rooted in cutting-edge AI research, provided a unique lens through which to view the inefficiencies plaguing incident response.
Anish, whose PhD research at MIT focused on "how do you get these AI systems to pick up cause-and-effect relationships from data, and also reinforcement learning, which is to me fundamentally about how do you search large spaces effectively," saw a convergence point. This specialized expertise, honed over years, directly informed their approach to what co-founder Raaz, a Berkeley PhD who previously worked in observability, termed the "needle in a haystack" problem. Raaz humorously noted, "Correlation isn't causation, and I joke, well, when I say it, then I'm allowed to say it because I have a degree." Their combined academic rigor and practical experience positioned them uniquely to address a challenge that has long defied simple solutions.
The core of the problem, as they articulated, is not merely detecting an anomaly. When a critical system experiences a latency spike, thousands of other metrics might also show unusual behavior. Distinguishing between a genuine cause, a mere symptom, or a spurious correlation becomes incredibly difficult. Traversal AI seeks to provide clarity in these high-stakes scenarios.
Modern enterprise systems operate at an astounding scale. DigitalOcean, one of Traversal's early customers, manages over 1300 microservices, generating billions of logs and tens of billions of time-series data points daily. Traditional observability tools struggle to make sense of this volume, often leaving engineers to manually sift through disparate dashboards, logs, and code repositories. Furthermore, the expectation for incident resolution is incredibly high, with companies demanding root cause identification within minutes.
