Eric Allam, co-founder of Trigger.dev, recently discussed the evolution of AI agents and the critical need for durability in their execution. In his presentation, Allam highlighted the shift from agents simply using existing backend infrastructure to becoming a fundamental part of that infrastructure themselves. This evolution necessitates new approaches to ensure agents can perform long-running, meaningful work reliably.
The Evolution of Web Backends and the Rise of Agents
Allam traced the history of web backends, starting with CGI in the early 1990s, which operated on a simple, stateless model where each request spawned a new process. This evolved into the LAMP stack and later serverless architectures, all largely adhering to a "shared nothing" or stateless compute model. However, as applications became more complex, they began incorporating "side effects" like sending emails or processing payments, which required more sophisticated handling of state and execution flow.
The emergence of AI agents, particularly with the advent of large language models (LLMs) in 2023, has further accelerated this trend. Allam explained that agents are fundamentally changing the dynamic, with the LLM now orchestrating the code, rather than the other way around. This shift demands a new approach to building "durable agents" that can maintain their state and recover from failures.
Two Roads to Durability: Replay vs. Snapshot
Allam outlined two primary methods for achieving agent durability: replay and snapshotting. The replay model, common in traditional software, involves logging every step of an execution. While this provides an audit trail and allows for retries, it can become cumbersome and inefficient for long-running, complex agent tasks, leading to issues with log size and versioning.
The alternative, snapshotting, involves capturing the entire state of a machine or process at a given moment and storing it. This approach, while conceptually simpler, has historically faced challenges with efficiency and size. Allam noted that early methods like CRIU (Checkpoint/Restore In Userspace) in 2011 offered a way to checkpoint and restore processes, but struggled with compatibility and performance for certain applications like media processing or headless browsers. He also pointed out the significant size of memory snapshots (512MB per snapshot) and the associated storage and network costs.
