AI Inference: 10x Faster Models & Self-Optimization
Philip Kiely and Ali Taha of Baseten discuss AI inference, LLM optimization, speculative decoding, and the engineering behind cutting-edge AI models.

Visual TL;DR
optimizing LLMs for faster response and higher efficiency in real-world applications
From the article 7 mentionsIn a recent episode of the Latent Space podcast, hosts and guests delved into the intricate world of AI inference, exploring how large language models (LLMs) are being optimized for speed and efficiency.
navigating complex, extended user inputs and multi-step tool calling processes
LLMs like GLM-5.2 write and optimize their own GPU kernels for inference engines
From the article 2 mentionsThe conversation highlighted the remarkable capability of LLMs, such as GLM-5.2, to not only generate text but also to write optimized GPU kernels.
a key technique to accelerate LLM output generation by predicting future tokens
From the article 2 mentionsSpeculative decoding emerged as a key technique for accelerating LLM inference.
Baseten's Ali Taha and Philip Kiely discuss practical implementation and support
From the article 9+ mentionsPhilip Kiely, author of "Inference Engineering," and Ali Taha, Head of Model Performance at Baseten, shared insights into the techniques and challenges involved in getting AI models to perform at their peak.
From the articleThis process involves a loop where the model analyzes profiling traces, identifies bottlenecks in its inference engine, writes new kernels, and then re-evaluates.
achieving significant speed improvements in LLM inference performance and throughput
From the article 2 mentionsThis method involves attaching a smaller, faster model that predicts several tokens ahead.
getting AI models to perform at their highest potential beyond standard benchmarks
From the article 4 mentionsPhilip Kiely, author of "Inference Engineering," and Ali Taha, Head of Model Performance at Baseten, shared insights into the techniques and challenges involved in getting AI models to perform at their peak.
Contents(4)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.