Quentin Anthony, Head of Model Training at Zyphra and an advisor at EleutherAI, recently joined Alessio Fanelli, Founder of Kernel Labs, on the Latent Space podcast to dissect the evolving landscape of AI hardware and developer productivity. The conversation offered a deep dive into Zyphra's bold move to AMD’s MI300X GPUs and Anthony’s unique perspective on leveraging AI tools for coding, providing sharp insights for founders, VCs, and AI professionals navigating the industry's complex technical frontiers.
Zyphra, a full-stack model company focused on building foundation models for edge deployment, has made a significant strategic pivot: migrating its entire training cluster to AMD. This decision stems from a conviction that the AMD ecosystem offers a "really compelling training cluster" that significantly reduces their bottom line. Anthony's journey with AMD began out of necessity during his PhD work on the Frontier supercomputer at Oak Ridge National Lab, an AMD MI250X-based system. This early exposure to AMD's hardware, even when its software stack lagged, provided invaluable experience in porting complex operations like Flash Attention.
The current generation MI300X GPUs, Anthony asserts, represent a pivotal shift. With 192GB of VRAM and superior memory bandwidth, these accelerators can outperform NVIDIA H100s on specific workloads. For tasks not bottlenecked by FP8 dense compute, such as certain Mixture of Experts (MoE) models, AMD's offerings provide a distinct advantage. This ample VRAM and bandwidth minimize the need for intricate parallelism strategies, streamlining development and boosting efficiency.
While AMD's hardware has undeniably caught up, the software ecosystem has been a historical hurdle. Anthony notes, however, that the software stack for the MI300X has also matured significantly. "Not a lot of people have sort of discovered that they've caught up on software and we're kind of capitalizing on that," he states, highlighting a strategic arbitrage opportunity. This parity in hardware and improving software support enables companies like Zyphra to extract substantial value, especially when coupled with a deep understanding of GPU architecture.
Anthony's approach to kernel development is rooted in a "bottom-up" philosophy. Rather than relying on high-level frameworks like Triton or Torch Compile, which he describes as being "beholden to the compiler," he often dives directly into ROCm or even GPU assembly. This granular control allows him to dictate precisely where tensors are materialized, optimizing for specific hardware properties. "The hardware of MI300X has these properties, and I want my algorithm to pull these properties out," he explains, underscoring the importance of tailoring code to exploit the underlying silicon.