LLMs Fail to Write Fast Multi-GPU Kernels
Simran Arora from Together AI discusses the challenges of multi-GPU kernel development and why current LLMs struggle to optimize them, despite ongoing research.
6 min read

Visual TL;DR
communication across multiple GPUs is a critical hurdle for AI performance
From the article 6 mentionsArora explained that while significant advancements have been made in optimizing single-GPU kernels and memory-efficient architectures, the bottleneck has shifted to multi-GPU communication.
From the articleSimran Arora, a principal scientist at Together AI and an incoming professor at Caltech, recently shed light on the challenges and opportunities in developing efficient multi-GPU AI kernels during a presentation titled 'Can LLMs Write Fast Multi-GPU Kernels?'.
current LLMs fail to write fast multi-GPU kernels for optimization
From the article 4 mentionsArora highlighted that while LLMs can often generate compilable CUDA code, they struggle with the nuanced reasoning required for optimizing multi-GPU kernels.
LLM limitations prevent efficient orchestration of multi-GPU communication
From the articleShe concluded by emphasizing the ongoing challenges and the potential for future research to unlock more sophisticated AI-driven kernel optimization.
From the article 3 mentionsIn the ever-evolving landscape of AI hardware, optimizing performance for large language models and other complex AI workloads is paramount.
From the articleSimran Arora, a principal scientist at Together AI and an incoming professor at Caltech, recently shed light on the challenges and opportunities in developing efficient multi-GPU AI kernels during a presentation titled 'Can LLMs Write Fast Multi-GPU Kernels?'.
advancements in single-GPU kernels and memory-efficient architectures
From the article 2 mentionsThe pace of innovation in GPU compute and memory has outstripped the progress in communication hardware.
From the articleThe pace of innovation in GPU compute and memory has outstripped the progress in communication hardware.
communication across multiple GPUs is a critical hurdle for AI performance
From the article 6 mentionsArora explained that while significant advancements have been made in optimizing single-GPU kernels and memory-efficient architectures, the bottleneck has shifted to multi-GPU communication.
current LLMs fail to write fast multi-GPU kernels for optimization
From the article 4 mentionsArora highlighted that while LLMs can often generate compilable CUDA code, they struggle with the nuanced reasoning required for optimizing multi-GPU kernels.
LLM limitations prevent efficient orchestration of multi-GPU communication
From the articleShe concluded by emphasizing the ongoing challenges and the potential for future research to unlock more sophisticated AI-driven kernel optimization.
Contents(4)
© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.

