#Kubernetes
37 articles with this tag

AI Agents Need Budgets, Not Just Tokens
Anthropic's Sachin Malhotra argues that AI agents in production need budgets, not just broad tokens, proposing primitives like asymmetric verbs, rate limits, and tripwires.

Netflix's Autoscaler Evolution
Netflix is consolidating its two Apache Flink autoscalers into one open-source solution, achieving significant cost savings and improved stability.

Open Source AI: Beyond Virtue to Ownership
Open source AI infrastructure is an ownership strategy, not charity. Controlling the AI model orchestration layer is key for long-term stability and auditability.

Krea.ai Details K2 Training & Serving Infrastructure
Gabriel from Krea.ai discusses the infrastructure behind K2, detailing challenges in large-scale GPU training and innovative serving solutions.

Crusoe Cloud Cuts AI Image Pulls
Crusoe Cloud integrates Spegel into its Managed Kubernetes to drastically cut AI/ML container image pull times, improving GPU utilization.

CoreWeave Launches AI Sandboxes
CoreWeave Sandboxes offers secure, isolated environments for AI reinforcement learning, agent tool use, and model evaluation, accessible on-cluster or serverless.

CoreWeave SUNK Adds Self-Service, Anywhere
CoreWeave enhances its SUNK system with self-service deployment and multi-cloud capabilities to speed up AI cluster setup.

Hugging Face Scales Infrastructure to Serve 3 Million Models
Arek Borucki from Hugging Face discusses how the platform scaled its infrastructure to serve millions of AI models and users, detailing architectural decisions and database optimizations.

Claude's Corner: Chamber (YC W2026), Your GPU Infrastructure Shouldn't Need a Babysitter
Between 30-60 pct of enterprise GPU capacity sits idle. Chamber's AI agents autonomously monitor, diagnose, and fix GPU cluster failures, and chat in Slack. Here's how it works and what it would take to build.

NVIDIA's AI Factory: Software, Validation, and Internal Use Cases
NVIDIA execs discuss Enterprise Validated Designs, the company's internal AI Factory scale, and the future of secure, autonomous AI agents.

Together AI Boosts GPU Cluster Uptime
Together AI introduces major reliability and control upgrades for its GPU Clusters, including automated node repair and enhanced operational oversight.

AI Agent Fleets: What Broke and How to Fix It
Kyle Jaejun Lee from KRAFTON shares the five key failures encountered when running a fleet of AI agents across multiple machines, and outlines solutions for building a more scalable and reliable system.
Databricks AI Serving Adapts to Any Model
Databricks unveils an AI serving platform that dynamically adapts to any model and traffic, slashing costs and boosting performance.
Nextdoor engineers build faster with Codex
Nextdoor engineers are using OpenAI's Codex to accelerate development, enabling end-to-end feature building and faster debugging.

Uber's Hybrid Core Allocation
Uber Engineering's hybrid core allocation system blends dedicated and shared CPUs for better efficiency and reliability.

Sally Ann O'Malley on OpenClaw in Containers
Sally Ann O'Malley from Red Hat discusses how OpenClaw agents can be containerized for reproducible, secure, and portable AI development from local machines to Kubernetes.
Kubernetes Security Goes Deep
LinkedIn enhances Kubernetes security with a new framework automating workload identity and credential management, ensuring trust across its massive infrastructure.

Uber Tackles AI Agent Identity
Uber is enhancing AI security with a new identity system for autonomous agents, ensuring accountability and traceability in complex workflows.

Scaling AI Agents on Kubernetes with OpenClaw
Onur Solmaz from OpenClaw discusses scaling AI agents on Kubernetes, highlighting ACP, acpx, and the future of agent orchestration.

Superlinked's Filip Makraduli on Small Model Inference Infrastructure
Filip Makraduli of Superlinked discusses the critical need for robust small model inference infrastructure, highlighting Superlinked's open-source solution.
OpenAI's Voice AI: Breaking Latency Barriers
OpenAI details its WebRTC rearchitecture for low-latency, high-scale voice AI, using a split relay and transceiver model.

Red Hat's Clyburn on Podman's AI Potential
Red Hat's Cedric Clyburn discusses Podman, highlighting its features for AI development, including Systemd integration and bootable containers.

Shared GPUs, Zero Conflict
Together AI's multi-tenant GPU clusters offer a path to cost-effective, scalable AI compute without sacrificing team isolation.

llm-d Enters CNCF Sandbox
The llm-d project's entry into the CNCF Sandbox marks a pivotal moment for cloud-native AI inference and open infrastructure.

Cedric Clyburn on Models as a Service
Red Hat's Cedric Clyburn discusses the evolution of AI from code assistants to Models as a Service (MaaS), highlighting on-premise and hybrid deployments with Kubernetes and OpenShift.
Databricks Tackles Kubernetes Load Balancing
Databricks engineers are presenting innovations in Kubernetes load balancing and AI-powered debugging at SRECon 2026.

AI Deployment Lags App Delivery
Most organizations struggle with AI deployment velocity due to lacking mature delivery infrastructure, a problem solvable by adapting cloud-native practices.

AI Coding Tests Flawed by Infrastructure Noise
The infrastructure powering AI coding tests can significantly inflate or deflate model scores, potentially masking true capabilities and misleading deployment decisions.

TAHO funding nets $3.5M to challenge Kubernetes in the AI compute war
The cost of training and running large AI models is rapidly becoming the single biggest bottleneck in the tech industry.
TAHO funding nets $3.5M to challenge Kubernetes in the AI compute war
The cost of training and running large AI models is rapidly becoming the single biggest bottleneck in the tech industry.

NVIDIA Dynamo AI Inference Scales Data Center AI
MetalBear Raises $12.5M Seed Round to Address Cloud Development Bottlenecks
MetalBear, the company behind the open-source Kubernetes development tool mirrord, today announced it has raised $12.5 million in Seed funding.

MetalBear Raises $12.5M Seed Round to Address Cloud Development Bottlenecks
MetalBear, the company behind the open-source Kubernetes development tool mirrord, today announced it has raised $12.5 million in Seed funding.

Llama Stack: Kubernetes for Generative AI Applications
Operant Launches Woodpecker: Open-Source Automated Red Teaming Engine for Kubernetes, APIs, and AI

ARMO Selected by Orange Business to Secure Managed Kubernetes Services
