#Speculative Decoding

7 articles with this tag

AI Inference: 10x Faster Models & Self-Optimization
Artificial Intelligence

AI Inference: 10x Faster Models & Self-Optimization

Philip Kiely and Ali Taha of Baseten discuss AI inference, LLM optimization, speculative decoding, and the engineering behind cutting-edge AI models.

about 2 months ago
Modal CTO on the 100,000 Sandbox Problem
Artificial Intelligence

Modal CTO on the 100,000 Sandbox Problem

Modal CTO Akshat Bubna discusses the "100,000 Sandbox Problem" and Modal's approach to scalable, flexible LLM inference infrastructure.

2 months ago
Bridging Diffusion LLMs and Speculative Decoding
AI Research

Bridging Diffusion LLMs and Speculative Decoding

A novel SimSD speculative decoding method enables diffusion LLMs to achieve up to 7.46x higher throughput without sacrificing generation quality.

4 months ago
Together AI Supercharges LLM Inference
Technology

Together AI Supercharges LLM Inference

Together AI unveils ATLAS, accelerating LLM inference up to 4x with adaptive speculative decoding, tackling the growing cost challenge for AI-native companies.

5 months ago
Together AI Slashes RL Training Time
Technology

Together AI Slashes RL Training Time

Together AI's new distribution-aware speculative decoding slashes RL training time by up to 50%, tackling a major bottleneck in LLM post-training.

5 months ago
Cloudflare's LLM Infrastructure Deep Dive
Technology

Cloudflare's LLM Infrastructure Deep Dive

Cloudflare details its advanced infrastructure optimizations for running large language models on its Workers AI platform, focusing on performance and cost-efficiency.

5 months ago
Together AI's Aurora Learns on the Fly
Technology

Together AI's Aurora Learns on the Fly

Together AI's Aurora framework uses RL to continuously adapt speculative decoding for faster LLM inference, outperforming static models.

6 months ago