#AI Inference

36 articles with this tag

Crusoe Cloud Offers Dedicated AI Inference
AI

Crusoe Cloud Offers Dedicated AI Inference

Crusoe Cloud launches Self-Serve Deployments, offering dedicated AI inference capacity for growing applications, moving beyond shared serverless pools.

1 day ago
CoreWeave Dominates AI Inference Benchmarks
AI

CoreWeave Dominates AI Inference Benchmarks

CoreWeave leads AI inference benchmarks for Moonshot AI's Kimi K2.6 model, showcasing superior speed and cost-efficiency through full-stack optimization.

1 day ago
OpenAI GPT-5.6: Smarter, Cheaper
Artificial Intelligence

OpenAI GPT-5.6: Smarter, Cheaper

OpenAI's new GPT-5.6 family cuts costs and boosts performance through significant inference and agentic harness optimizations.

6 days ago
AMD, Cerebras Target Low-Latency AI
Semiconductors

AMD, Cerebras Target Low-Latency AI

AMD and Cerebras announce a new partnership to deliver ultra-low latency AI inference solutions by combining AMD Helios and Cerebras Wafer-Scale Engine.

12 days ago
AMD and Cerebras Team on AI Inference
Technology

AMD and Cerebras Team on AI Inference

AMD and Cerebras unveil a new AI inference solution combining AMD Helios GPUs with Cerebras' Wafer-Scale Engine for ultra-low latency and high throughput.

12 days ago
AMD Helios Powers Azure AI Inference
Semiconductors

AMD Helios Powers Azure AI Inference

Microsoft is deploying AMD's Helios AI inference platform on Azure, integrating new EPYC CPUs and Instinct GPUs to power frontier models and AI services.

15 days ago
99.9% Uptime: What It Really Means for AI Inference
Technology

99.9% Uptime: What It Really Means for AI Inference

Achieving 99.9% uptime for AI inference means surviving data center failures, demanding active multi-facility traffic and direct infrastructure control.

19 days ago
Together AI adds Inkling multimodal model
Technology

Together AI adds Inkling multimodal model

Together AI integrates Inkling, a new multimodal AI model from Thinking Machines Lab, offering text, image, and audio processing with controllable reasoning.

20 days ago
Together AI Offers Predictable Inference
Technology

Together AI Offers Predictable Inference

Together AI introduces Provisioned Throughput, offering reserved inference capacity for open models with token-based pricing and a 99% uptime SLA.

27 days ago
Mozilla.ai Unveils transcribe.cpp
Technology

Mozilla.ai Unveils transcribe.cpp

Mozilla.ai launches transcribe.cpp, an open-source C/C++ library for fast, GPU-accelerated speech-to-text inference with broad model support.

about 1 month ago
OpenAI's Custom AI Chip with Broadcom
Semiconductors

OpenAI's Custom AI Chip with Broadcom

OpenAI unveils its first custom AI chip, 'Jalapeno', co-developed with Broadcom, aiming for 50% cost savings and enhanced AI inference performance.

about 1 month ago
Claude's Corner: Cumulus Labs, When the Inference Market Gets Outclassed by CUDA Kernels
Claude's Corner

Claude's Corner: Cumulus Labs, When the Inference Market Gets Outclassed by CUDA Kernels

Most GPU clouds rent H100s, wrap vLLM, and call it a product. Cumulus Labs built Ion, a C++ inference engine with custom CUDA kernels for the NVIDIA GH200, and they're posting 7,167 tok/s on a single chip and 12.5-second cold starts. Here's how the hardware-native tricks work, and whether anyone can replicate them.

about 2 months ago
Coding Agent Inference Benchmark Revealed
Technology

Coding Agent Inference Benchmark Revealed

Together AI unveils a new benchmark for coding agent inference, highlighting performance under real-world load and significant cost advantages.

3 months ago
Together AI Taps Blockchain for Cheaper AI
Technology

Together AI Taps Blockchain for Cheaper AI

Together AI and Pearl Research Labs are integrating blockchain to cut AI inference costs, offering discounted model access subsidized by cryptocurrency mining.

3 months ago
Together AI: Deploy Any Hugging Face Model Instantly
Technology

Together AI: Deploy Any Hugging Face Model Instantly

Together AI's Dedicated Container Inference lets developers deploy any Hugging Face model instantly, bypassing complex setups and accelerating AI experimentation.

3 months ago
Superlinked's Filip Makraduli on Small Model Inference Infrastructure
Artificial Intelligence

Superlinked's Filip Makraduli on Small Model Inference Infrastructure

Filip Makraduli of Superlinked discusses the critical need for robust small model inference infrastructure, highlighting Superlinked's open-source solution.

3 months ago
Intel's AI Chip Demand: A Boon for Semiconductor Stocks?
Semiconductors

Intel's AI Chip Demand: A Boon for Semiconductor Stocks?

Intel CEO Pat Gelsinger discusses the surging demand for Intel CPUs in AI inference and the company's strategy to leverage its integrated hardware offerings and partnerships for growth.

3 months ago
Orbital aims for space AI data centers
Funding Round

Orbital aims for space AI data centers

Orbital plans its first test mission for space-based AI data centers in 2027, aiming to overcome Earth's power constraints.

4 months ago
llm-d Enters CNCF Sandbox
Artificial Intelligence

llm-d Enters CNCF Sandbox

The llm-d project's entry into the CNCF Sandbox marks a pivotal moment for cloud-native AI inference and open infrastructure.

4 months ago
OpenAI Cerebras Deal Targets Real Time AI Speed
AI Research

OpenAI Cerebras Deal Targets Real Time AI Speed

OpenAI's Cerebras partnership prioritizes reducing AI inference latency, aiming for real-time interactions to drive deeper user engagement with deployed models.

7 months ago
Google TPU Ironwood: Inference Powerhouse Arrives
AI Research

Google TPU Ironwood: Inference Powerhouse Arrives

8 months ago
Google Cloud’s AI Storage Strategy: Optimizing Performance and Cost
AI Video

Google Cloud’s AI Storage Strategy: Optimizing Performance and Cost

9 months ago
vLLM Solves the AI Model Serving Conundrum at Scale
AI Video

vLLM Solves the AI Model Serving Conundrum at Scale

9 months ago
Google Cloud Unveils Blueprint for Reliable, Scalable AI Inference
AI Video

Google Cloud Unveils Blueprint for Reliable, Scalable AI Inference

9 months ago
NVIDIA Dynamo AI Inference Scales Data Center AI
AI Research

NVIDIA Dynamo AI Inference Scales Data Center AI

9 months ago
Impala AI Targets LLM Inference Costs with $11M Seed
Funding Round

Impala AI Targets LLM Inference Costs with $11M Seed

9 months ago
Fireworks AI raises $250M to advance its AI inference platform
Funding Round

Fireworks AI raises $250M to advance its AI inference platform

9 months ago
Tensormesh exits stealth with $4.5M to slash AI inference caching costs
AI Research

Tensormesh exits stealth with $4.5M to slash AI inference caching costs

The generative AI gold rush has an expensive secret: running the models costs a fortune.

9 months ago
Tensormesh exits stealth with $4.5M to slash AI inference caching costs
AI Research

Tensormesh exits stealth with $4.5M to slash AI inference caching costs

The generative AI gold rush has an expensive secret: running the models costs a fortune.

9 months ago
Qualcomm’s Bold AI Inference Play Challenges NVIDIA Dominance
AI Video

Qualcomm’s Bold AI Inference Play Challenges NVIDIA Dominance

9 months ago
AI Research

Blackwell AI Inference: NVIDIA's Extreme-Scale Bet

11 months ago
Groq Secures $750M Investment to Expand the American AI Stack
Funding Round

Groq Secures $750M Investment to Expand the American AI Stack

11 months ago
NVIDIA Details SMART Framework for AI Inference at Scale
AI Research

NVIDIA Details SMART Framework for AI Inference at Scale

NVIDIA has outlined its comprehensive strategy for optimizing AI inference performance at scale, introducing the "Think SMART" framework as a guide for enterprises building and operating "AI factories."

12 months ago
NVIDIA Dynamo Redefines AI Inference Economics
AI Video

NVIDIA Dynamo Redefines AI Inference Economics

about 1 year ago
Chalk Secures $50M Series A to Revolutionize AI Inference
Funding Round

Chalk Secures $50M Series A to Revolutionize AI Inference

about 1 year ago
Making Machine Learning Inference Meet Real-World Performance Demands
Interview

Making Machine Learning Inference Meet Real-World Performance Demands

FPGAs offer the configurability needed for real-time machine learning inference, with the flexibility to adapt to future workloads. Making these advantages accessible to data-scientists and developers calls for tools that are both comprehensive and easy to use. Daniel Eaton, Sr Manager, Strategic Marketing Development, Xilinx

over 7 years ago