# Together AI Boosts GPU Cluster Uptime _Together AI introduces major reliability and control upgrades for its GPU Clusters, including automated node repair and enhanced operational oversight._ **Published:** 2026-07-15 **Source:** https://www.startuphub.ai/ai-news/technology/2026/together-ai-boosts-gpu-cluster-uptime --- Together AI is bolstering its [Together GPU Clusters](https://www.together.ai/blog/new-in-together-gpu-clusters-reliability-and-control-for-production-gpu-clusters) with a suite of updates designed to tackle the realities of large-scale AI training and inference. The focus is squarely on improving platform health and providing more granular operational control for production environments. Common Failure PointsDriver From the articleThese enhancements address common failure points in distributed systems, such as hardware malfunctions and scheduler issues, which can derail lengthy training jobs.impactsGPU Cluster UptimeContextimproving reliability and control for large-scale AI training and inference environmentsFrom the article 4 mentionsTogether AI is bolstering its Together GPU Clusters with a suite of updates designed to tackle the realities of large-scale AI training and inference.achieved byPassive Health ChecksCoreFrom the article 2 mentionsThe company has introduced passive health checks, which continuously monitor nodes by observing real workloads, logs, and metrics.enablesAuto Node RepairCoresystem automatically detects and addresses issues like GPUs falling off PCIe busFrom the articleComplementing these checks is an auto node repair system.leads toEnhanced Operational ControlEffectproviding more granular oversight for production environments and AI developmentFrom the articleThe focus is squarely on improving platform health and providing more granular operational control for production environments.results inRobust AI InfrastructureOutcomeoffering a more manageable and reliable platform for AI development and deploymentFrom the article 3 mentionsThe goal is to offer a more robust and manageable infrastructure for AI development and deployment. These enhancements address common failure points in distributed systems, such as hardware malfunctions and scheduler issues, which can derail lengthy training jobs. The goal is to offer a more robust and manageable infrastructure for AI development and deployment. ## Platform Health Improvements The company has introduced passive health checks, which continuously monitor nodes by observing real workloads, logs, and metrics. This system aims to detect degradation, like GPUs falling off the PCIe bus or thermal throttling, as it happens, providing an early warning system distinct from traditional active checks used at provisioning time. Complementing these checks is an auto node repair system. When a node issue is detected, the system suggests remediation actions, Reboot, Reprovision, Failover, or Remove, for an operator to approve. This human-in-the-loop approach balances automation with safety, ensuring critical workloads remain uninterrupted. These [passive health checks](/ai-news/technology/2026/databricks-tackles-gpu-woes) and repair mechanisms promise to significantly reduce the time operators spend on troubleshooting, shifting from hours of support tickets to minutes of in-product workflow. ## A Rebuilt Slurm Stack Together AI has also rebuilt its Slurm-on-Kubernetes stack from the ground up. This overhaul addresses issues like crashing daemons, zombie processes, and scheduler drift. The new stack features self-healing worker daemons and ensures reliable cleanup of orphaned processes. Job accounting is now stored on durable storage, preventing data loss from pod restarts. The system also accurately tracks GPU state after reschedules, ensuring the schedulable pool always matches the available hardware. DCGM metrics are now exposed in Grafana dashboards for detailed GPU utilization visibility. This upgraded stack is now the default for newly provisioned Slurm clusters, with options for existing managed clusters to migrate. This represents a significant step towards more reliable [Together GPU Clusters reliability](/ai-news/technology/2026/shared-gpus-zero-conflict). ## Enhanced Operational Control Beyond reliability, Together AI is enhancing control for growing teams. A redesigned cluster details view provides at-a-glance information on node health, live usage metrics, and an event timeline, consolidating critical operational data. External OIDC support has been added for Kubernetes RBAC. This allows teams to integrate with existing identity providers for per-user authentication, authorization, and audit trails, moving away from shared admin kubeconfig files. Startup scripts offer another layer of customization. These scripts can be configured to run at specific lifecycle events on nodes, enabling self-serve setup for internal packages, scratch space preparation, or custom notifications without manual intervention or support tickets. This level of control is crucial for effective [AI infrastructure management](/ai-news/artificial-intelligence/2026/us-builds-ai-safety-framework). These updates aim to provide the visibility, access, and customization needed to manage [production GPU clusters](/ai-news/ai-news/2026/ai-infrastructure-week-july-6-2026) effectively as organizations scale. The focus on resilience and control signals a maturing approach to [AI infrastructure management](/ai-news/funding-round/2026/risk-ledger-secures-32m-series-b), essential for demanding [production GPU clusters](/ai-news/ai-news/2026/ai-chip-inference-wars-july-8-2026). --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.