#AI Inference

38 articles with this tag

Cerebras, Callosum Eye Agentic AI with New Partnership
Artificial Intelligence

Cerebras, Callosum Eye Agentic AI with New Partnership

Cerebras and Callosum partner to offer ultra-low-latency, heterogeneous agentic AI inference, expanding European reach and enabling new AI applications.

29 days ago
Cerebras Streams SUPERNOVA 2026 AI Event
Startup News

Cerebras Streams SUPERNOVA 2026 AI Event

Cerebras Systems will livestream its SUPERNOVA flagship event on August 18th, featuring new product announcements and demos for AI inference and production applications.

about 1 month ago
Crusoe Cloud Offers Dedicated AI Inference
AI

Crusoe Cloud Offers Dedicated AI Inference

Crusoe Cloud launches Self-Serve Deployments, offering dedicated AI inference capacity for growing applications, moving beyond shared serverless pools.

about 2 months ago
CoreWeave Dominates AI Inference Benchmarks
AI

CoreWeave Dominates AI Inference Benchmarks

CoreWeave leads AI inference benchmarks for Moonshot AI's Kimi K2.6 model, showcasing superior speed and cost-efficiency through full-stack optimization.

about 2 months ago
OpenAI GPT-5.6: Smarter, Cheaper
Artificial Intelligence

OpenAI GPT-5.6: Smarter, Cheaper

OpenAI's new GPT-5.6 family cuts costs and boosts performance through significant inference and agentic harness optimizations.

about 2 months ago
AMD, Cerebras Target Low-Latency AI
Semiconductors

AMD, Cerebras Target Low-Latency AI

AMD and Cerebras announce a new partnership to deliver ultra-low latency AI inference solutions by combining AMD Helios and Cerebras Wafer-Scale Engine.

about 2 months ago
AMD and Cerebras Team on AI Inference
Technology

AMD and Cerebras Team on AI Inference

AMD and Cerebras unveil a new AI inference solution combining AMD Helios GPUs with Cerebras' Wafer-Scale Engine for ultra-low latency and high throughput.

about 2 months ago
AMD Helios Powers Azure AI Inference
Semiconductors

AMD Helios Powers Azure AI Inference

Microsoft is deploying AMD's Helios AI inference platform on Azure, integrating new EPYC CPUs and Instinct GPUs to power frontier models and AI services.

2 months ago
99.9% Uptime: What It Really Means for AI Inference
Technology

99.9% Uptime: What It Really Means for AI Inference

Achieving 99.9% uptime for AI inference means surviving data center failures, demanding active multi-facility traffic and direct infrastructure control.

2 months ago
Together AI adds Inkling multimodal model
Technology

Together AI adds Inkling multimodal model

Together AI integrates Inkling, a new multimodal AI model from Thinking Machines Lab, offering text, image, and audio processing with controllable reasoning.

2 months ago
Together AI Offers Predictable Inference
Technology

Together AI Offers Predictable Inference

Together AI introduces Provisioned Throughput, offering reserved inference capacity for open models with token-based pricing and a 99% uptime SLA.

2 months ago
Mozilla.ai Unveils transcribe.cpp
Technology

Mozilla.ai Unveils transcribe.cpp

Mozilla.ai launches transcribe.cpp, an open-source C/C++ library for fast, GPU-accelerated speech-to-text inference with broad model support.

3 months ago
OpenAI's Custom AI Chip with Broadcom
Semiconductors

OpenAI's Custom AI Chip with Broadcom

OpenAI unveils its first custom AI chip, 'Jalapeno', co-developed with Broadcom, aiming for 50% cost savings and enhanced AI inference performance.

3 months ago
Claude's Corner: Cumulus Labs, When the Inference Market Gets Outclassed by CUDA Kernels
Claude's Corner

Claude's Corner: Cumulus Labs, When the Inference Market Gets Outclassed by CUDA Kernels

Most GPU clouds rent H100s, wrap vLLM, and call it a product. Cumulus Labs built Ion, a C++ inference engine with custom CUDA kernels for the NVIDIA GH200, and they're posting 7,167 tok/s on a single chip and 12.5-second cold starts. Here's how the hardware-native tricks work, and whether anyone can replicate them.

3 months ago
Coding Agent Inference Benchmark Revealed
Technology

Coding Agent Inference Benchmark Revealed

Together AI unveils a new benchmark for coding agent inference, highlighting performance under real-world load and significant cost advantages.

4 months ago
Together AI Taps Blockchain for Cheaper AI
Technology

Together AI Taps Blockchain for Cheaper AI

Together AI and Pearl Research Labs are integrating blockchain to cut AI inference costs, offering discounted model access subsidized by cryptocurrency mining.

4 months ago
Together AI: Deploy Any Hugging Face Model Instantly
Technology

Together AI: Deploy Any Hugging Face Model Instantly

Together AI's Dedicated Container Inference lets developers deploy any Hugging Face model instantly, bypassing complex setups and accelerating AI experimentation.

4 months ago
Superlinked's Filip Makraduli on Small Model Inference Infrastructure
Artificial Intelligence

Superlinked's Filip Makraduli on Small Model Inference Infrastructure

Filip Makraduli of Superlinked discusses the critical need for robust small model inference infrastructure, highlighting Superlinked's open-source solution.

5 months ago
Intel's AI Chip Demand: A Boon for Semiconductor Stocks?
Semiconductors

Intel's AI Chip Demand: A Boon for Semiconductor Stocks?

Intel CEO Pat Gelsinger discusses the surging demand for Intel CPUs in AI inference and the company's strategy to leverage its integrated hardware offerings and partnerships for growth.

5 months ago
Orbital aims for space AI data centers
Funding Round

Orbital aims for space AI data centers

Orbital plans its first test mission for space-based AI data centers in 2027, aiming to overcome Earth's power constraints.

5 months ago
llm-d Enters CNCF Sandbox
Artificial Intelligence

llm-d Enters CNCF Sandbox

The llm-d project's entry into the CNCF Sandbox marks a pivotal moment for cloud-native AI inference and open infrastructure.

6 months ago
OpenAI Cerebras Deal Targets Real Time AI Speed
AI Research

OpenAI Cerebras Deal Targets Real Time AI Speed

OpenAI's Cerebras partnership prioritizes reducing AI inference latency, aiming for real-time interactions to drive deeper user engagement with deployed models.

8 months ago
Google TPU Ironwood: Inference Powerhouse Arrives
AI Research

Google TPU Ironwood: Inference Powerhouse Arrives

10 months ago
Google Cloud’s AI Storage Strategy: Optimizing Performance and Cost
AI Video

Google Cloud’s AI Storage Strategy: Optimizing Performance and Cost

10 months ago
vLLM Solves the AI Model Serving Conundrum at Scale
AI Video

vLLM Solves the AI Model Serving Conundrum at Scale

10 months ago
Google Cloud Unveils Blueprint for Reliable, Scalable AI Inference
AI Video

Google Cloud Unveils Blueprint for Reliable, Scalable AI Inference

10 months ago
NVIDIA Dynamo AI Inference Scales Data Center AI
AI Research

NVIDIA Dynamo AI Inference Scales Data Center AI

10 months ago
Impala AI Targets LLM Inference Costs with $11M Seed
Funding Round

Impala AI Targets LLM Inference Costs with $11M Seed

11 months ago
Fireworks AI raises $250M to advance its AI inference platform
Funding Round

Fireworks AI raises $250M to advance its AI inference platform

11 months ago
Tensormesh exits stealth with $4.5M to slash AI inference caching costs
AI Research

Tensormesh exits stealth with $4.5M to slash AI inference caching costs

The generative AI gold rush has an expensive secret: running the models costs a fortune.

11 months ago
Tensormesh exits stealth with $4.5M to slash AI inference caching costs
AI Research

Tensormesh exits stealth with $4.5M to slash AI inference caching costs

The generative AI gold rush has an expensive secret: running the models costs a fortune.

11 months ago
Qualcomm’s Bold AI Inference Play Challenges NVIDIA Dominance
AI Video

Qualcomm’s Bold AI Inference Play Challenges NVIDIA Dominance

11 months ago
AI Research

Blackwell AI Inference: NVIDIA's Extreme-Scale Bet

12 months ago
Groq Secures $750M Investment to Expand the American AI Stack
Funding Round

Groq Secures $750M Investment to Expand the American AI Stack

about 1 year ago
NVIDIA Details SMART Framework for AI Inference at Scale
AI Research

NVIDIA Details SMART Framework for AI Inference at Scale

NVIDIA has outlined its comprehensive strategy for optimizing AI inference performance at scale, introducing the "Think SMART" framework as a guide for enterprises building and operating "AI factories."

about 1 year ago
NVIDIA Dynamo Redefines AI Inference Economics
AI Video

NVIDIA Dynamo Redefines AI Inference Economics

about 1 year ago
Chalk Secures $50M Series A to Revolutionize AI Inference
Funding Round

Chalk Secures $50M Series A to Revolutionize AI Inference

over 1 year ago
Making Machine Learning Inference Meet Real-World Performance Demands
Interview

Making Machine Learning Inference Meet Real-World Performance Demands

FPGAs offer the configurability needed for real-time machine learning inference, with the flexibility to adapt to future workloads. Making these advantages accessible to data-scientists and developers calls for tools that are both comprehensive and easy to use. Daniel Eaton, Sr Manager, Strategic Marketing Development, Xilinx

over 7 years ago