#Kubernetes

37 articles with this tag

AI Agents Need Budgets, Not Just Tokens
Artificial Intelligence

AI Agents Need Budgets, Not Just Tokens

Anthropic's Sachin Malhotra argues that AI agents in production need budgets, not just broad tokens, proposing primitives like asymmetric verbs, rate limits, and tripwires.

about 2 hours ago
Netflix's Autoscaler Evolution
Technology

Netflix's Autoscaler Evolution

Netflix is consolidating its two Apache Flink autoscalers into one open-source solution, achieving significant cost savings and improved stability.

1 day ago
Open Source AI: Beyond Virtue to Ownership
Technology

Open Source AI: Beyond Virtue to Ownership

Open source AI infrastructure is an ownership strategy, not charity. Controlling the AI model orchestration layer is key for long-term stability and auditability.

2 days ago
Krea.ai Details K2 Training & Serving Infrastructure
AI Research

Krea.ai Details K2 Training & Serving Infrastructure

Gabriel from Krea.ai discusses the infrastructure behind K2, detailing challenges in large-scale GPU training and innovative serving solutions.

4 days ago
Crusoe Cloud Cuts AI Image Pulls
AI

Crusoe Cloud Cuts AI Image Pulls

Crusoe Cloud integrates Spegel into its Managed Kubernetes to drastically cut AI/ML container image pull times, improving GPU utilization.

19 days ago
CoreWeave Launches AI Sandboxes
AI

CoreWeave Launches AI Sandboxes

CoreWeave Sandboxes offers secure, isolated environments for AI reinforcement learning, agent tool use, and model evaluation, accessible on-cluster or serverless.

19 days ago
CoreWeave SUNK Adds Self-Service, Anywhere
AI

CoreWeave SUNK Adds Self-Service, Anywhere

CoreWeave enhances its SUNK system with self-service deployment and multi-cloud capabilities to speed up AI cluster setup.

19 days ago
Hugging Face Scales Infrastructure to Serve 3 Million Models
Artificial Intelligence

Hugging Face Scales Infrastructure to Serve 3 Million Models

Arek Borucki from Hugging Face discusses how the platform scaled its infrastructure to serve millions of AI models and users, detailing architectural decisions and database optimizations.

25 days ago
Claude's Corner: Chamber (YC W2026), Your GPU Infrastructure Shouldn't Need a Babysitter
Claude's Corner

Claude's Corner: Chamber (YC W2026), Your GPU Infrastructure Shouldn't Need a Babysitter

Between 30-60 pct of enterprise GPU capacity sits idle. Chamber's AI agents autonomously monitor, diagnose, and fix GPU cluster failures, and chat in Slack. Here's how it works and what it would take to build.

28 days ago
NVIDIA's AI Factory: Software, Validation, and Internal Use Cases
Artificial Intelligence

NVIDIA's AI Factory: Software, Validation, and Internal Use Cases

NVIDIA execs discuss Enterprise Validated Designs, the company's internal AI Factory scale, and the future of secure, autonomous AI agents.

about 1 month ago
Together AI Boosts GPU Cluster Uptime
Technology

Together AI Boosts GPU Cluster Uptime

Together AI introduces major reliability and control upgrades for its GPU Clusters, including automated node repair and enhanced operational oversight.

about 1 month ago
AI Agent Fleets: What Broke and How to Fix It
Artificial Intelligence

AI Agent Fleets: What Broke and How to Fix It

Kyle Jaejun Lee from KRAFTON shares the five key failures encountered when running a fleet of AI agents across multiple machines, and outlines solutions for building a more scalable and reliable system.

about 2 months ago
Databricks AI Serving Adapts to Any Model
Technology

Databricks AI Serving Adapts to Any Model

Databricks unveils an AI serving platform that dynamically adapts to any model and traffic, slashing costs and boosting performance.

2 months ago
Nextdoor engineers build faster with Codex
Artificial Intelligence

Nextdoor engineers build faster with Codex

Nextdoor engineers are using OpenAI's Codex to accelerate development, enabling end-to-end feature building and faster debugging.

2 months ago
Uber's Hybrid Core Allocation
tech

Uber's Hybrid Core Allocation

Uber Engineering's hybrid core allocation system blends dedicated and shared CPUs for better efficiency and reliability.

3 months ago
Sally Ann O'Malley on OpenClaw in Containers
Technology

Sally Ann O'Malley on OpenClaw in Containers

Sally Ann O'Malley from Red Hat discusses how OpenClaw agents can be containerized for reproducible, secure, and portable AI development from local machines to Kubernetes.

3 months ago
Kubernetes Security Goes Deep
tech

Kubernetes Security Goes Deep

LinkedIn enhances Kubernetes security with a new framework automating workload identity and credential management, ensuring trust across its massive infrastructure.

3 months ago
Uber Tackles AI Agent Identity
Artificial Intelligence

Uber Tackles AI Agent Identity

Uber is enhancing AI security with a new identity system for autonomous agents, ensuring accountability and traceability in complex workflows.

3 months ago
Scaling AI Agents on Kubernetes with OpenClaw
Artificial Intelligence

Scaling AI Agents on Kubernetes with OpenClaw

Onur Solmaz from OpenClaw discusses scaling AI agents on Kubernetes, highlighting ACP, acpx, and the future of agent orchestration.

3 months ago
Superlinked's Filip Makraduli on Small Model Inference Infrastructure
Artificial Intelligence

Superlinked's Filip Makraduli on Small Model Inference Infrastructure

Filip Makraduli of Superlinked discusses the critical need for robust small model inference infrastructure, highlighting Superlinked's open-source solution.

4 months ago
OpenAI's Voice AI: Breaking Latency Barriers
Artificial Intelligence

OpenAI's Voice AI: Breaking Latency Barriers

OpenAI details its WebRTC rearchitecture for low-latency, high-scale voice AI, using a split relay and transceiver model.

4 months ago
Red Hat's Clyburn on Podman's AI Potential
Artificial Intelligence

Red Hat's Clyburn on Podman's AI Potential

Red Hat's Cedric Clyburn discusses Podman, highlighting its features for AI development, including Systemd integration and bootable containers.

4 months ago
Shared GPUs, Zero Conflict
Technology

Shared GPUs, Zero Conflict

Together AI's multi-tenant GPU clusters offer a path to cost-effective, scalable AI compute without sacrificing team isolation.

4 months ago
llm-d Enters CNCF Sandbox
Artificial Intelligence

llm-d Enters CNCF Sandbox

The llm-d project's entry into the CNCF Sandbox marks a pivotal moment for cloud-native AI inference and open infrastructure.

5 months ago
Cedric Clyburn on Models as a Service
Artificial Intelligence

Cedric Clyburn on Models as a Service

Red Hat's Cedric Clyburn discusses the evolution of AI from code assistants to Models as a Service (MaaS), highlighting on-premise and hybrid deployments with Kubernetes and OpenShift.

5 months ago
Databricks Tackles Kubernetes Load Balancing
Technology

Databricks Tackles Kubernetes Load Balancing

Databricks engineers are presenting innovations in Kubernetes load balancing and AI-powered debugging at SRECon 2026.

5 months ago
AI Deployment Lags App Delivery
Artificial Intelligence

AI Deployment Lags App Delivery

Most organizations struggle with AI deployment velocity due to lacking mature delivery infrastructure, a problem solvable by adapting cloud-native practices.

5 months ago
AI Coding Tests Flawed by Infrastructure Noise
Artificial Intelligence

AI Coding Tests Flawed by Infrastructure Noise

The infrastructure powering AI coding tests can significantly inflate or deflate model scores, potentially masking true capabilities and misleading deployment decisions.

7 months ago
TAHO funding nets $3.5M to challenge Kubernetes in the AI compute war
Funding Round

TAHO funding nets $3.5M to challenge Kubernetes in the AI compute war

The cost of training and running large AI models is rapidly becoming the single biggest bottleneck in the tech industry.

9 months ago
Funding Round

TAHO funding nets $3.5M to challenge Kubernetes in the AI compute war

The cost of training and running large AI models is rapidly becoming the single biggest bottleneck in the tech industry.

9 months ago
NVIDIA Dynamo AI Inference Scales Data Center AI
AI Research

NVIDIA Dynamo AI Inference Scales Data Center AI

10 months ago
Funding Round

MetalBear Raises $12.5M Seed Round to Address Cloud Development Bottlenecks

MetalBear, the company behind the open-source Kubernetes development tool mirrord, today announced it has raised $12.5 million in Seed funding.

11 months ago
MetalBear Raises $12.5M Seed Round to Address Cloud Development Bottlenecks
Funding Round

MetalBear Raises $12.5M Seed Round to Address Cloud Development Bottlenecks

MetalBear, the company behind the open-source Kubernetes development tool mirrord, today announced it has raised $12.5 million in Seed funding.

11 months ago
Llama Stack: Kubernetes for Generative AI Applications
AI Video

Llama Stack: Kubernetes for Generative AI Applications

12 months ago
Press Release

Operant Launches Woodpecker: Open-Source Automated Red Teaming Engine for Kubernetes, APIs, and AI

over 1 year ago
ARMO Selected by Orange Business to Secure Managed Kubernetes Services
Press Release

ARMO Selected by Orange Business to Secure Managed Kubernetes Services

over 1 year ago
Klaudia is the First Kubernetes AI Agent for Proactive Troubleshooting
Interview

Klaudia is the First Kubernetes AI Agent for Proactive Troubleshooting

almost 2 years ago