#AI Infrastructure
50 articles with this tag

Hermes Agent Runs on Crusoe Cloud
Crusoe Cloud now supports open-source AI agents like Hermes, offering optimized inference for complex, unattended tasks.

Crusoe Named Gartner Cloud AI Visionary
Crusoe named a Visionary in the 2026 Gartner Magic Quadrant for Cloud AI Infrastructure, highlighting its energy-first approach and specialized AI platform.

Five transactions in seven days put a price on AI agent governance: what the exits reveal
Okta and Cyera paid $1.2B for AI agent identity companies. Three VCs added $171M more. All five transactions landed in the same week, before the category had a name.

Tech Giants Commit $2.4 Trillion to AI Infrastructure
Alphabet, Meta, Microsoft, and Amazon commit nearly $2.4 trillion to AI data center infrastructure as Wall Street caps a positive trading month.

Together AI partners with Moonshot AI
Together AI partners with Moonshot AI to offer Kimi K3 and future models, providing developers with day-zero access to large-scale open-source AI.

Together AI Refines Model Deployment
Together AI details its capacity-aware routing architecture for dedicated model inference, enabling dynamic deployments, A/B testing, and efficient scaling.

Stripe paid $10 billion for the layer between AI models. Travis Kalanick raised $1.7 billion for the factory floor.
Physical AI and AI routing infrastructure dominated the week of July 21-27: Atoms raised $1.7B, Etched hit a $10.3B valuation, and Stripe is in talks to buy OpenRouter for $10B.

Claude's Corner: Voxel Energy - Power Is the New Bottleneck
Voxel Energy builds off-grid data centers powered by solar and repurposed EV batteries, designed to get GPU clusters running in months rather than the five-plus years a standard grid hookup now takes. A deep technical breakdown from YC W2026.

AI Data Centers Face Power Crunch
AI's massive power demands are straining electrical grids, forcing a 'power-first' approach to data center development to avoid costly delays.

Databricks' AI Vending Machine
Databricks unveils its 'Field Engineering Vending Machine' (FEVM), a self-serve app for on-demand infrastructure provisioning, enabling agent-first workflows and tackling scaling challenges.

AMD Unveils Helios AI Rack
AMD launches its first rack-scale AI solution, AMD Helios, alongside new CPUs and GPUs, aiming to dominate the agentic AI era.

CrowdStrike, Cerebras Speed Up AI Security
CrowdStrike and Cerebras unite to accelerate AI-driven cybersecurity, leveraging Cerebras's fast inference hardware to enhance threat detection and response.

AI Data Centers Face Storage Crunch
AI data centers face a pressing storage density problem, with new SSDs and integrated systems offering a path to greater capacity and performance.

Greylock Bets $1.5B on the Next AI Wave
Greylock Partners launches a $1.5 billion fund for early-stage AI startups, focusing on models, infrastructure, and applications, while offering a cautious view on new AI model benchmarks.

Netflix's LLM Engine Revealed
Netflix reveals its custom LLM serving infrastructure built on vLLM and NVIDIA Triton, enabling flexibility and performance within its production environment.

Kimi Verifier Rebuilds Trust in Open Source AI
Moonshot AI launches Kimi Vendor Verifier to ensure open-source AI models run accurately across all implementations, rebuilding trust in the ecosystem.

Jensen Huang’s Nvidia: $81.6B Q1 Revenue and the Sovereign AI Empire Behind It
Nvidia posted $81.6 billion in Q1 FY2027 revenue on May 20, 2026, an 85 pct increase year-on-year. Here is Jensen Huang’s segment-by-segment breakdown: $75.2B in data center (half from non-hyperscale buyers) and a reclassified Edge Computing segment, set against sovereign AI factory deals in Japan, South Korea, and Germany.

Together AI Boosts GPU Cluster Uptime
Together AI introduces major reliability and control upgrades for its GPU Clusters, including automated node repair and enhanced operational oversight.

Mozilla.ai Launches Otari LLM Control Plane
Mozilla.ai has launched Otari, an open-source LLM control plane to simplify managing diverse language model infrastructure for AI application developers.

Cerebras Eyes 200MW AI Power in Europe
Cerebras Systems announces a major European expansion, targeting 200MW of AI compute capacity by the end of 2027 to serve regional demand.

Modal CTO on the 100,000 Sandbox Problem
Modal CTO Akshat Bubna discusses the "100,000 Sandbox Problem" and Modal's approach to scalable, flexible LLM inference infrastructure.

AI Capex Demand Points to Multi-Year Growth Cycle
Rudina Seseri of Glasswing Ventures discusses the multi-year AI capex cycle and the trend of hyperscalers vertically integrating their solutions.

Four of every five dollars raised this week went to AI infrastructure. Here is what happened in the other 20 percent.
AI infrastructure took 79% of this week's $9.9B. Aramco led Together AI's $800M. Crusoe sought $3B. Schneider Electric paid $3.1B for Cognite.

Mamoon Hamid on VC in the AI Revolution
Mamoon Hamid of Kleiner Perkins discusses venture capital's role in the AI revolution, his career journey, and the future of AI-driven innovation.

AI Infrastructure: The Speed Problem
AI adoption is bottlenecked by slow, costly infrastructure. Companies need 'agentic speeds infrastructure' for autonomous AI to succeed.

Meta to Build Cloud Business for AI Compute
Meta is reportedly planning to build its own cloud business to sell excess AI compute power, offering API access to AI models and renting out data center capacity.

Together AI lands $800M Series C
Together AI secures $800M Series C to accelerate open-source AI development, promising lower costs and higher performance for production workloads.
OpenAI's Bug Hunt: 18-Year-Old Flaw Found
OpenAI uncovered two hidden bugs, including an 18-year-old software flaw, by analyzing crash data like an epidemiologist.

Qualcomm and Superhuman both bought AI software companies this week. VCs mostly sat out.
While VCs slowed to $900M deployed, hardware incumbents wrote $6B in acquisition checks. The inference software layer just got priced.

OpenAI's Custom AI Chip with Broadcom
OpenAI unveils its first custom AI chip, 'Jalapeno', co-developed with Broadcom, aiming for 50% cost savings and enhanced AI inference performance.

Claude's Corner: Cumulus Labs, When the Inference Market Gets Outclassed by CUDA Kernels
Most GPU clouds rent H100s, wrap vLLM, and call it a product. Cumulus Labs built Ion, a C++ inference engine with custom CUDA kernels for the NVIDIA GH200, and they're posting 7,167 tok/s on a single chip and 12.5-second cold starts. Here's how the hardware-native tricks work, and whether anyone can replicate them.

Rumble CEO on AI Compute and Quake AI Platform
Rumble CEO Chris Pavlovski discusses the company's new AI platform, Quake AI, and its strategy to capitalize on the massive demand for AI compute power with extensive GPU infrastructure.

Anjney Midha on AI's FLOPs to Megawatts Transition
Anjney Midha discusses the AI industry's shift from focusing on FLOPs to megawatt efficiency, highlighting cultural challenges and AMP PBC's approach to managed compute.
Databricks Taps NVIDIA Vera for AI Agents
Databricks and NVIDIA are integrating NVIDIA Vera CPUs and GPUs into the Databricks platform to accelerate the development of AI agents and other complex AI workloads.

Three companies named "Physics AI" the same week and raised $15.8 billion between them
Prometheus, PhysicsX, and Mistral each independently reached for the same new term in the same week -- and collectively raised $15.8B. Plus: the AI IPO wave nobody noticed.
Compute Once: Unlocking AI Agent Efficiency
A radical proposal to precompute LLM KV caches, slashing inference costs by up to 50x and enabling a new compute-efficient AI agent paradigm.

LLM Control Plane: Beyond the Gateway
Production AI needs more than just gateways; an LLM control plane is crucial for managing budgets, privacy, and dynamic routing.

Four quantum companies raised $961M in seven days. Europe wrote the checks.
The week Anthropic filed for IPO, four quantum rounds totaled $961M in Europe, AI agent payment rails launched quietly, and DeepSeek took its first outside money.
OpenAI's Policy Playbook
OpenAI lays out its public policy strategy, focusing on AI safety, youth protection, and equitable access to ensure AGI benefits all of humanity.

Together AI Masters MiniMax M3 Inference
Together AI details engineering feats enabling efficient MiniMax M3 inference, unlocking 1M-token context and multimodality.

HPE CEO Neri: AI Drives Strong Revenue and Future Growth
HPE CEO Antonio Neri discusses the company's strong Q2 results, driven by AI demand, and forecasts continued growth across networking, cloud, and AI portfolios.

HPE CEO Neri: AI Fuels "Blowout" Revenue, Triple-Digit Server Demand
HPE CEO Antonio Neri discusses the company's "blowout" AI revenue, triple-digit demand for AI infrastructure, and the shift towards on-premise solutions.

Rishabh Bhargava on Voice Agent Engineering
Rishabh Bhargava of Together AI discusses engineering voice agents, focusing on latency, quality, and scale challenges across STT, LLM, and TTS components.

Otari: Own Your AI Stack
Otari launches an open-source LLM gateway and hosted platform to provide essential tools and capabilities for both frontier and open-weight AI models.

AI Infrastructure: Your Next Competitive Edge
Enterprises must modernize their IT infrastructure to unlock AI potential and gain a competitive edge, moving beyond legacy systems and technical debt.
Databricks Tackles LLM Inference Costs
Databricks details its 'model units' abstraction and cost-aware autoscaling for reliable, high-throughput LLM inference, cutting GPU costs by over 80%.

Claude's Corner: Captain, The RAG Infrastructure Play That's Playing Bloomberg
Captain (YC W2026) is building managed RAG-as-a-service, two API calls to connect your data sources, 95% retrieval accuracy via contextual embeddings + hybrid search + reranking, and an Odyssey data pivot that looks a lot like Bloomberg Terminal strategy. Here's the architecture, the moat, and how to build a clone.

AI Infrastructure Boom: Demand Surges as Costs Collapse
ARK Investment Management's "Big Ideas 2026" report details the AI infrastructure boom, with demand surging and costs collapsing, driving massive investment.

Elon Musk's xAI: $500M ARR, $1B Burn, $1.25T SpaceX Merger
xAI reached an estimated $500M ARR in 2026 while burning roughly $1B per month. A complete breakdown of revenue, capital raises, the Colossus supercomputer, and what the SpaceX merger means.

Lenovo CFO on AI Growth: "First Year of Our AI Decade"
Lenovo CFO Winston Cheng discusses the company's strong AI growth, highlighting its diversified portfolio and strategic investments in the AI decade.