Together AI Boosts GPU Cluster Uptime

Together AI introduces major reliability and control upgrades for its GPU Clusters, including automated node repair and enhanced operational oversight.

Abstract representation of interconnected GPUs and data flow within a server rack.
Together AI's latest updates focus on improving the reliability and control of their GPU Clusters for demanding AI workloads.
Visual TL;DR
Common Failure PointsDriver
From the articleThese enhancements address common failure points in distributed systems, such as hardware malfunctions and scheduler issues, which can derail lengthy training jobs.
GPU Cluster UptimeContext
improving reliability and control for large-scale AI training and inference environments
From the article 4 mentionsTogether AI is bolstering its Together GPU Clusters with a suite of updates designed to tackle the realities of large-scale AI training and inference.
Passive Health ChecksCore
From the article 2 mentionsThe company has introduced passive health checks, which continuously monitor nodes by observing real workloads, logs, and metrics.
Auto Node RepairCore
system automatically detects and addresses issues like GPUs falling off PCIe bus
From the articleComplementing these checks is an auto node repair system.
Enhanced Operational ControlEffect
providing more granular oversight for production environments and AI development
From the articleThe focus is squarely on improving platform health and providing more granular operational control for production environments.
Robust AI InfrastructureOutcome
offering a more manageable and reliable platform for AI development and deployment
From the article 3 mentionsThe goal is to offer a more robust and manageable infrastructure for AI development and deployment.
Contents(4)

Together AI is bolstering its Together GPU Clusters with a suite of updates designed to tackle the realities of large-scale AI training and inference. The focus is squarely on improving platform health and providing more granular operational control for production environments.

Companies working on this

StartupHub profiles of the companies this article names, with funding and a one-liner from our database.

Together AI
$8.3B
AI acceleration cloud platform for open source and enterprise.
OpenAI
$852.0B
OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity.
Cisco
$240.0B
Technology infrastructure and networking company providing enterprise solutions
OpenAI
$852.0B
An AI research and deployment company building safe and beneficial artificial general intelligence.

These enhancements address common failure points in distributed systems, such as hardware malfunctions and scheduler issues, which can derail lengthy training jobs. The goal is to offer a more robust and manageable infrastructure for AI development and deployment.

Platform Health Improvements

The company has introduced passive health checks, which continuously monitor nodes by observing real workloads, logs, and metrics. This system aims to detect degradation, like GPUs falling off the PCIe bus or thermal throttling, as it happens, providing an early warning system distinct from traditional active checks used at provisioning time.

Complementing these checks is an auto node repair system. When a node issue is detected, the system suggests remediation actions, Reboot, Reprovision, Failover, or Remove, for an operator to approve. This human-in-the-loop approach balances automation with safety, ensuring critical workloads remain uninterrupted.

These passive health checks and repair mechanisms promise to significantly reduce the time operators spend on troubleshooting, shifting from hours of support tickets to minutes of in-product workflow.

A Rebuilt Slurm Stack

Together AI has also rebuilt its Slurm-on-Kubernetes stack from the ground up. This overhaul addresses issues like crashing daemons, zombie processes, and scheduler drift. The new stack features self-healing worker daemons and ensures reliable cleanup of orphaned processes.

Job accounting is now stored on durable storage, preventing data loss from pod restarts. The system also accurately tracks GPU state after reschedules, ensuring the schedulable pool always matches the available hardware. DCGM metrics are now exposed in Grafana dashboards for detailed GPU utilization visibility.

This upgraded stack is now the default for newly provisioned Slurm clusters, with options for existing managed clusters to migrate. This represents a significant step towards more reliable Together GPU Clusters reliability.

Enhanced Operational Control

Beyond reliability, Together AI is enhancing control for growing teams. A redesigned cluster details view provides at-a-glance information on node health, live usage metrics, and an event timeline, consolidating critical operational data.

External OIDC support has been added for Kubernetes RBAC. This allows teams to integrate with existing identity providers for per-user authentication, authorization, and audit trails, moving away from shared admin kubeconfig files.

Startup scripts offer another layer of customization. These scripts can be configured to run at specific lifecycle events on nodes, enabling self-serve setup for internal packages, scratch space preparation, or custom notifications without manual intervention or support tickets. This level of control is crucial for effective AI infrastructure management.

These updates aim to provide the visibility, access, and customization needed to manage production GPU clusters effectively as organizations scale. The focus on resilience and control signals a maturing approach to AI infrastructure management, essential for demanding production GPU clusters.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer