#LLM Inference

11 articles with this tag

How Together AI Solves LLM Cold Starts
Artificial Intelligence

How Together AI Solves LLM Cold Starts

Together AI details native metrics and cold start benchmarks to fix nonlinear latency degradation in LLM autoscaling.

3 days ago
Claude's Corner: Cumulus Labs, When the Inference Market Gets Outclassed by CUDA Kernels
Claude's Corner

Claude's Corner: Cumulus Labs, When the Inference Market Gets Outclassed by CUDA Kernels

Most GPU clouds rent H100s, wrap vLLM, and call it a product. Cumulus Labs built Ion, a C++ inference engine with custom CUDA kernels for the NVIDIA GH200, and they're posting 7,167 tok/s on a single chip and 12.5-second cold starts. Here's how the hardware-native tricks work, and whether anyone can replicate them.

about 2 months ago
Together AI Masters MiniMax M3 Inference
Technology

Together AI Masters MiniMax M3 Inference

Together AI details engineering feats enabling efficient MiniMax M3 inference, unlocking 1M-token context and multimodality.

2 months ago
Together AI Supercharges LLM Inference
Technology

Together AI Supercharges LLM Inference

Together AI unveils ATLAS, accelerating LLM inference up to 4x with adaptive speculative decoding, tackling the growing cost challenge for AI-native companies.

3 months ago
Together AI's Aurora Learns on the Fly
Technology

Together AI's Aurora Learns on the Fly

Together AI's Aurora framework uses RL to continuously adapt speculative decoding for faster LLM inference, outperforming static models.

4 months ago
Mamba-3: Inference-First SSMs Arrive
Artificial Intelligence

Mamba-3: Inference-First SSMs Arrive

Together AI's Mamba-3 advances state space models with a focus on inference speed, outperforming previous versions and some Transformers.

4 months ago
Technology

NVIDIA Nemotron 3 Nano launches on FriendliAI

The race to serve the next generation of efficient, open AI agents is heating up, and FriendliAI is aggressively positioning itself as the crucial infrastruc...

8 months ago
NVIDIA Nemotron 3 Nano launches on FriendliAI
Artificial Intelligence

NVIDIA Nemotron 3 Nano launches on FriendliAI

The race to serve the next generation of efficient, open AI agents is heating up, and FriendliAI is aggressively positioning itself as the crucial infrastruc...

8 months ago
Startup News

Clarifai Hits Fastest GPT-OSS-120B Inference and Narrows the GPU, ASIC Gap

Clarifai’s latest benchmark on OpenAI’s GPT-OSS-120B model points to a quiet but important shift in AI infrastructure.

9 months ago
Clarifai Hits Fastest GPT-OSS-120B Inference and Narrows the GPU, ASIC Gap
Startup News

Clarifai Hits Fastest GPT-OSS-120B Inference and Narrows the GPU, ASIC Gap

Clarifai’s latest benchmark on OpenAI’s GPT-OSS-120B model points to a quiet but important shift in AI infrastructure.

9 months ago
Pliops Unveils Breakthrough AI Performance Enhancements
Press Release

Pliops Unveils Breakthrough AI Performance Enhancements

about 1 year ago
#LLM Inference Articles | StartupHub.ai