In a recent TWIML AI Podcast episode, Philip Kiely, Head of AI Education at Baseten, joined host Sam Charrington to discuss the intricacies of AI inference. Kiely, who has spent over four years in the AI space, shared his insights on the challenges and opportunities in making AI models efficient and accessible for real-world applications. The conversation highlighted the critical differences between AI training and inference, emphasizing the need for specialized approaches to optimize the latter.
Companies working on this
Profiles of the companies named in this story, with founding year, headquarters, and a short description from our database.
AI inference platform for deploying models in production with high performance and availability.
- Founded
- 2020
- Location
- San Francisco, United States
- Valuation
- $5.0B
Kiely noted that while the AI field has seen tremendous progress in model training, the subsequent step of deploying these models for inference, where they are used to make predictions on new data, is often overlooked. This phase presents unique challenges related to cost, latency, and scalability, particularly as AI models become larger and more complex.
The full discussion can be found on TWIML's YouTube channel.
The Inference Challenge
The core thesis of the discussion revolved around the unique difficulties associated with AI inference. "If you think about medicine, for example, it can take decades for research to reach a pharmacy," Kiely explained. "Even within AI, if you want to train a model off of a new technique, it can still take weeks or months to find the exact right way to express that technique. But with inference, the timeline is often hours." He elaborated that inference, unlike training, needs to happen in real-time or near real-time to be useful for many applications.
This demand for speed and efficiency means that companies need to carefully consider how their models are deployed. "Inference is often the bottleneck," Kiely stated. "It's where the rubber meets the road. If your inference is too slow or too expensive, your product simply won't be viable." He contrasted this with the training phase, which, while computationally intensive, can often be done asynchronously and with more tolerance for longer processing times.
