Modal CTO on the 100,000 Sandbox Problem

Modal CTO Akshat Bubna discusses the "100,000 Sandbox Problem" and Modal's approach to scalable, flexible LLM inference infrastructure.

4 min read
Akshat Bubna, CTO of Modal, speaking into a microphone.
Akshat Bubna, CTO of Modal, discusses LLM inference challenges.· Latent Space
Visual TL;DR
100,000 Sandbox ProblemDriver
diverse, unpredictable LLM workloads needing flexible infrastructure
Flexible HardwareContext
running models across various hardware configurations
From the articleBubna highlighted a key aspect of their platform: the ability to run models on specific hardware and in specific regions, providing users with fine-grained control over their inference setups.
Geographic DistributionContext
serving models in different geographic locations
LLM Inference ComplexityDriver
difficulty providing consistent, performant inference experience
User Needs VaryContext
specific GPUs, regions, latency requirements for models
From the article 2 mentionsBubna explained that the core issue lies in the difficulty of providing a consistent and performant inference experience for users who may have vastly different needs.
Modal's ApproachCore
From the article 4 mentionsHe elaborated on Modal's approach to tackling this problem, emphasizing their focus on providing a platform that allows users to "own their inference." This means giving users the control and flexibility to manage their models, from data preparation and training to deployment and scaling.
Scalable InfrastructureEffect
enabling flexible and efficient LLM serving
From the article 2 mentionsModal's infrastructure is built to accommodate this, offering features like the ability to run workloads across multiple cloud providers and to dynamically scale resources based on demand.
Optimized LLM ServingOutcome
handling diverse and unpredictable inference demands

In a recent discussion, Akshat Bubna, CTO of Modal, delved into the complexities of serving large language models (LLMs) at scale, highlighting what he terms the "100,000 Sandbox Problem." This challenge stems from the diverse and often unpredictable nature of LLM workloads, which require flexible and efficient infrastructure to run optimally across various hardware configurations and geographic locations.

Modal CTO on the 100,000 Sandbox Problem - Latent Space
Modal CTO on the 100,000 Sandbox Problem — from Latent Space

Bubna explained that the core issue lies in the difficulty of providing a consistent and performant inference experience for users who may have vastly different needs. "We see this all the time where customers want to run models that are very specific, maybe they want to run them on different GPUs, or they want to run them in different regions, or they have very specific latency requirements," Bubna stated.

He elaborated on Modal's approach to tackling this problem, emphasizing their focus on providing a platform that allows users to "own their inference." This means giving users the control and flexibility to manage their models, from data preparation and training to deployment and scaling. Modal's infrastructure is built to accommodate this, offering features like the ability to run workloads across multiple cloud providers and to dynamically scale resources based on demand.

Bubna highlighted a key aspect of their platform: the ability to run models on specific hardware and in specific regions, providing users with fine-grained control over their inference setups. He also touched upon the importance of observability, noting that "performance on a benchmark is not enough. Performance in production needs to be observable." Modal provides metrics and tools to help users understand and optimize their model performance in real-world scenarios.

The conversation also touched upon the evolving landscape of AI development, where teams are increasingly looking for ways to streamline the process of deploying and scaling models. Bubna suggested that Modal's platform offers a solution by abstracting away much of the underlying infrastructure complexity, allowing developers to focus on building and iterating on their models. He pointed to the company's commitment to open-source principles and providing transparent tooling as key differentiators.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.