#AI Infrastructure
50 articles with this tag

Nvidia acquires Hugging Face $13 billion deal
Nvidia is acquiring Hugging Face in a deal valued at about $13 billion, per a Bloomberg Podcast Money Minute.

Nvidia acquires Hugging Face $13 billion
Nvidia keeps Hugging Face as an open, compute-agnostic hub in a $13B deal with $1B in retention and founders staying six years.

Neoclouds Borrow Like Utilities While Post-Training Finds Its First Commercial Layer
Lambda Labs debt, Deep Cogito post-training, Alice AI safety revenue, and what the real WoW signal is under the headline capital decline.

Nvidia's 70% Growth Is the main topic
Nvidia guided 70% growth for fiscal 2028 vs 45% expected, saying demand could double supply if chips were available.

NVIDIA Vera CPU Ships, Targeting Agentic AI Workloads

NVIDIA Vera Rubin NVL72 Sets New AI Agent Efficiency Benchmark

The inference chip layer raised $1.65 billion in five days. The AI cloud answered with $5.6 billion in bonds.
Etched, Fractile, and Groq split $1.65B on competing chip bets while Nebius and Domyn borrowed $5.6B in convertible debt. Castelion hit $13B making hypersonic missiles. Week of Aug 17-23, 2026.

Jeff Bezos Holds Three AI Bets: AWS Chips, Anthropic IPO, Prometheus
AWS AI and custom chips each crossed a $25 billion annual run rate in the second quarter of 2026, Amazon CEO Andy Jassy reported on July 30. The same quarter, a $53.4 billion non-cash mark on Amazon's Anthropic stake dominated the company's net income. And separately, Jeff Bezos is now co-CEO of Prometheus, an industrial AI startup valued at $41 billion after a $12 billion raise in June. The full picture is three AI bets operating in parallel, each at a scale that would be a career-defining position for most investors.

Jensen Huang's Blackwell Blueprint and Nvidia's AI Factory Strategy
Nvidia's GB200 NVL72 treats 72 GPUs as a single processor and delivers 30x faster LLM inference than the H100, while Jensen Huang argues the software stack running on top is now the harder-to-replicate asset.

Rum Group CEO on AI Pivot and $3B Opportunity
Rum Group CEO Chris Pavlovski discusses the company's record Q2 revenue, the strategic pivot into AI infrastructure with Quake AI, and the $3 billion opportunity in unmonetized power and GPU capacity.

Crusoe Cloud beefs up AI security
Crusoe Cloud rolls out Customer-Managed Keys (CMEK) for AWS KMS, giving enterprises direct control over their AI data encryption keys.

Harvey Labs: Building AI Research on a Budget
Harvey's Gabe Pereyra shares the playbook for building a competitive AI research lab on a budget, emphasizing benchmarks, synthetic data, and leveraging the frontier ecosystem.

SpaceX's IPO Debut: AI Costs Dampen Debut Earnings
SpaceX's IPO debut earnings reveal strong revenue but significant AI costs. Analysts discuss the company's growth strategy and competitive positioning.

Hermes Agent Runs on Crusoe Cloud
Crusoe Cloud now supports open-source AI agents like Hermes, offering optimized inference for complex, unattended tasks.

Crusoe Named Gartner Cloud AI Visionary
Crusoe named a Visionary in the 2026 Gartner Magic Quadrant for Cloud AI Infrastructure, highlighting its energy-first approach and specialized AI platform.

Five transactions in seven days put a price on AI agent governance: what the exits reveal
Okta and Cyera paid $1.2B for AI agent identity companies. Three VCs added $171M more. All five transactions landed in the same week, before the category had a name.

Tech Giants Commit $2.4 Trillion to AI Infrastructure
Alphabet, Meta, Microsoft, and Amazon commit nearly $2.4 trillion to AI data center infrastructure as Wall Street caps a positive trading month.

Together AI partners with Moonshot AI
Together AI partners with Moonshot AI to offer Kimi K3 and future models, providing developers with day-zero access to large-scale open-source AI.

Together AI Refines Model Deployment
Together AI details its capacity-aware routing architecture for dedicated model inference, enabling dynamic deployments, A/B testing, and efficient scaling.

Stripe paid $10 billion for the layer between AI models. Travis Kalanick raised $1.7 billion for the factory floor.
Physical AI and AI routing infrastructure dominated the week of July 21-27: Atoms raised $1.7B, Etched hit a $10.3B valuation, and Stripe is in talks to buy OpenRouter for $10B.

Claude's Corner: Voxel Energy - Power Is the New Bottleneck
Voxel Energy builds off-grid data centers powered by solar and repurposed EV batteries, designed to get GPU clusters running in months rather than the five-plus years a standard grid hookup now takes. A deep technical breakdown from YC W2026.

AI Data Centers Face Power Crunch
AI's massive power demands are straining electrical grids, forcing a 'power-first' approach to data center development to avoid costly delays.

Databricks' AI Vending Machine
Databricks unveils its 'Field Engineering Vending Machine' (FEVM), a self-serve app for on-demand infrastructure provisioning, enabling agent-first workflows and tackling scaling challenges.

AMD Unveils Helios AI Rack
AMD launches its first rack-scale AI solution, AMD Helios, alongside new CPUs and GPUs, aiming to dominate the agentic AI era.

CrowdStrike, Cerebras Speed Up AI Security
CrowdStrike and Cerebras unite to accelerate AI-driven cybersecurity, leveraging Cerebras's fast inference hardware to enhance threat detection and response.

AI Data Centers Face Storage Crunch
AI data centers face a pressing storage density problem, with new SSDs and integrated systems offering a path to greater capacity and performance.

Greylock Bets $1.5B on the Next AI Wave
Greylock Partners launches a $1.5 billion fund for early-stage AI startups, focusing on models, infrastructure, and applications, while offering a cautious view on new AI model benchmarks.

Netflix's LLM Engine Revealed
Netflix reveals its custom LLM serving infrastructure built on vLLM and NVIDIA Triton, enabling flexibility and performance within its production environment.

Kimi Verifier Rebuilds Trust in Open Source AI
Moonshot AI launches Kimi Vendor Verifier to ensure open-source AI models run accurately across all implementations, rebuilding trust in the ecosystem.

Jensen Huang’s Nvidia: $81.6B Q1 Revenue and the Sovereign AI Empire Behind It
Nvidia posted $81.6 billion in Q1 FY2027 revenue on May 20, 2026, an 85 pct increase year-on-year. Here is Jensen Huang’s segment-by-segment breakdown: $75.2B in data center (half from non-hyperscale buyers) and a reclassified Edge Computing segment, set against sovereign AI factory deals in Japan, South Korea, and Germany.

Together AI Boosts GPU Cluster Uptime
Together AI introduces major reliability and control upgrades for its GPU Clusters, including automated node repair and enhanced operational oversight.

Mozilla.ai Launches Otari LLM Control Plane
Mozilla.ai has launched Otari, an open-source LLM control plane to simplify managing diverse language model infrastructure for AI application developers.

Cerebras Eyes 200MW AI Power in Europe
Cerebras Systems announces a major European expansion, targeting 200MW of AI compute capacity by the end of 2027 to serve regional demand.

Modal CTO on the 100,000 Sandbox Problem
Modal CTO Akshat Bubna discusses the "100,000 Sandbox Problem" and Modal's approach to scalable, flexible LLM inference infrastructure.

AI Capex Demand Points to Multi-Year Growth Cycle
Rudina Seseri of Glasswing Ventures discusses the multi-year AI capex cycle and the trend of hyperscalers vertically integrating their solutions.

Four of every five dollars raised this week went to AI infrastructure. Here is what happened in the other 20 percent.
AI infrastructure took 79% of this week's $9.9B. Aramco led Together AI's $800M. Crusoe sought $3B. Schneider Electric paid $3.1B for Cognite.

Mamoon Hamid on VC in the AI Revolution
Mamoon Hamid of Kleiner Perkins discusses venture capital's role in the AI revolution, his career journey, and the future of AI-driven innovation.

AI Infrastructure: The Speed Problem
AI adoption is bottlenecked by slow, costly infrastructure. Companies need 'agentic speeds infrastructure' for autonomous AI to succeed.

Meta to Build Cloud Business for AI Compute
Meta is reportedly planning to build its own cloud business to sell excess AI compute power, offering API access to AI models and renting out data center capacity.

Together AI lands $800M Series C
Together AI secures $800M Series C to accelerate open-source AI development, promising lower costs and higher performance for production workloads.
OpenAI's Bug Hunt: 18-Year-Old Flaw Found
OpenAI uncovered two hidden bugs, including an 18-year-old software flaw, by analyzing crash data like an epidemiologist.

Qualcomm and Superhuman both bought AI software companies this week. VCs mostly sat out.
While VCs slowed to $900M deployed, hardware incumbents wrote $6B in acquisition checks. The inference software layer just got priced.

OpenAI's Custom AI Chip with Broadcom
OpenAI unveils its first custom AI chip, 'Jalapeno', co-developed with Broadcom, aiming for 50% cost savings and enhanced AI inference performance.

Claude's Corner: Cumulus Labs, When the Inference Market Gets Outclassed by CUDA Kernels
Most GPU clouds rent H100s, wrap vLLM, and call it a product. Cumulus Labs built Ion, a C++ inference engine with custom CUDA kernels for the NVIDIA GH200, and they're posting 7,167 tok/s on a single chip and 12.5-second cold starts. Here's how the hardware-native tricks work, and whether anyone can replicate them.

Rumble CEO on AI Compute and Quake AI Platform
Rumble CEO Chris Pavlovski discusses the company's new AI platform, Quake AI, and its strategy to capitalize on the massive demand for AI compute power with extensive GPU infrastructure.

Anjney Midha on AI's FLOPs to Megawatts Transition
Anjney Midha discusses the AI industry's shift from focusing on FLOPs to megawatt efficiency, highlighting cultural challenges and AMP PBC's approach to managed compute.
Databricks Taps NVIDIA Vera for AI Agents
Databricks and NVIDIA are integrating NVIDIA Vera CPUs and GPUs into the Databricks platform to accelerate the development of AI agents and other complex AI workloads.

Three companies named "Physics AI" the same week and raised $15.8 billion between them
Prometheus, PhysicsX, and Mistral each independently reached for the same new term in the same week -- and collectively raised $15.8B. Plus: the AI IPO wave nobody noticed.
Compute Once: Unlocking AI Agent Efficiency
A radical proposal to precompute LLM KV caches, slashing inference costs by up to 50x and enabling a new compute-efficient AI agent paradigm.

LLM Control Plane: Beyond the Gateway
Production AI needs more than just gateways; an LLM control plane is crucial for managing budgets, privacy, and dynamic routing.