CoreWeave Shatters MLPerf Records

CoreWeave sets new MLPerf training records with DeepSeek-V3 671B in just over 2 minutes on 8,192 NVIDIA GB300 GPUs.

8 min read
CoreWeave data center with rows of GPU servers
CoreWeave Newsroom

Visual TL;DR. AI Training Speed Race drives CoreWeave Sets Record. CoreWeave Sets Record achieved with 8,192 GB300 GPUs. 8,192 GB300 GPUs is Largest GB300 Cluster. CoreWeave Sets Record enables Faster Iteration. 8,192 GB300 GPUs requires Beyond Raw Hardware. Beyond Raw Hardware leads to Production-Ready Infra.

  1. AI Training Speed Race: rapidly accelerating AI development makes training speed a critical bottleneck
  2. CoreWeave Sets Record: shatters MLPerf records with DeepSeek-V3 671B in just over 2 minutes
  3. 8,192 GB300 GPUs: accomplished using a colossal cluster of NVIDIA GB300 NVL72 GPUs
  4. Largest GB300 Cluster: reportedly the largest GB300 cluster ever submitted to the MLPerf benchmark
  5. Beyond Raw Hardware: true differentiator lies in interplay of networking, orchestration, scheduling, storage
  6. Faster Iteration: ability for AI teams to iterate quickly directly impacts their competitive edge
  7. Production-Ready Infra: CoreWeave's infrastructure is built for demanding, real-world AI workloads
Visual TL;DR
Visual TL;DR, startuphub.ai AI Training Speed Race drives CoreWeave Sets Record. CoreWeave Sets Record achieved with 8,192 GB300 GPUs. CoreWeave Sets Record enables Faster Iteration drives achieved with enables AI Training Speed Race CoreWeave Sets Record 8,192 GB300 GPUs Faster Iteration From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Training Speed Race drives CoreWeave Sets Record. CoreWeave Sets Record achieved with 8,192 GB300 GPUs. CoreWeave Sets Record enables Faster Iteration drives achieved with enables AI Training SpeedRace CoreWeave SetsRecord 8,192 GB300 GPUs Faster Iteration From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Training Speed Race drives CoreWeave Sets Record. CoreWeave Sets Record achieved with 8,192 GB300 GPUs. CoreWeave Sets Record enables Faster Iteration drives achieved with enables AI Training Speed Race rapidly accelerating AI development makestraining speed a critical bottleneck CoreWeave Sets Record shatters MLPerf records with DeepSeek-V3671B in just over 2 minutes 8,192 GB300 GPUs accomplished using a colossal cluster ofNVIDIA GB300 NVL72 GPUs Faster Iteration ability for AI teams to iterate quicklydirectly impacts their competitive edge From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Training Speed Race drives CoreWeave Sets Record. CoreWeave Sets Record achieved with 8,192 GB300 GPUs. CoreWeave Sets Record enables Faster Iteration drives achieved with enables AI Training SpeedRace rapidlyaccelerating AIdevelopment makes… CoreWeave SetsRecord shatters MLPerfrecords withDeepSeek-V3 671B in… 8,192 GB300 GPUs accomplished usinga colossal clusterof NVIDIA GB300… Faster Iteration ability for AIteams to iteratequickly directly… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Training Speed Race drives CoreWeave Sets Record. CoreWeave Sets Record achieved with 8,192 GB300 GPUs. 8,192 GB300 GPUs is Largest GB300 Cluster. CoreWeave Sets Record enables Faster Iteration. 8,192 GB300 GPUs requires Beyond Raw Hardware. Beyond Raw Hardware leads to Production-Ready Infra drives achieved with is enables requires leads to AI Training Speed Race rapidly accelerating AI development makestraining speed a critical bottleneck CoreWeave Sets Record shatters MLPerf records with DeepSeek-V3671B in just over 2 minutes 8,192 GB300 GPUs accomplished using a colossal cluster ofNVIDIA GB300 NVL72 GPUs Largest GB300 Cluster reportedly the largest GB300 cluster eversubmitted to the MLPerf benchmark Beyond Raw Hardware true differentiator lies in interplay ofnetworking, orchestration, scheduling,storage Faster Iteration ability for AI teams to iterate quicklydirectly impacts their competitive edge Production-Ready Infra CoreWeave's infrastructure is built fordemanding, real-world AI workloads From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Training Speed Race drives CoreWeave Sets Record. CoreWeave Sets Record achieved with 8,192 GB300 GPUs. 8,192 GB300 GPUs is Largest GB300 Cluster. CoreWeave Sets Record enables Faster Iteration. 8,192 GB300 GPUs requires Beyond Raw Hardware. Beyond Raw Hardware leads to Production-Ready Infra drives achieved with is enables requires leads to AI Training SpeedRace rapidlyaccelerating AIdevelopment makes… CoreWeave SetsRecord shatters MLPerfrecords withDeepSeek-V3 671B in… 8,192 GB300 GPUs accomplished usinga colossal clusterof NVIDIA GB300… Largest GB300Cluster reportedly thelargest GB300cluster ever… Beyond RawHardware true differentiatorlies in interplayof networking,… Faster Iteration ability for AIteams to iteratequickly directly… Production-ReadyInfra CoreWeave'sinfrastructure isbuilt for… From startuphub.ai · The publishers behind this format

CoreWeave has once again pushed the boundaries of AI training performance, announcing new record-breaking results in the MLPerf Training v6.0 benchmark. The company managed to train the computationally intensive DeepSeek-V3 671B model in just 2.02 minutes. This remarkable feat was accomplished using a colossal cluster of 8,192 NVIDIA GB300 NVL72 GPUs, reportedly the largest GB300 cluster ever submitted to the benchmark.

The Race for Training Speed

In the rapidly accelerating world of AI development, the speed at which models can be trained is a critical bottleneck. As frontier models swell to trillion-parameter scales and agentic workloads become the norm, the ability for AI teams to iterate quickly directly impacts their competitive edge. CoreWeave's latest benchmark performance, detailed on their newsroom, highlights that raw hardware power is only one piece of the puzzle. The true differentiator lies in the intricate interplay of networking, orchestration, scheduling, storage, and software working in concert. This announcement underscores CoreWeave's sustained investment in full-stack optimization, a strategy that has consistently turned cutting-edge hardware into reliable, production-ready performance at scale.

Scaling the Summit with GB300

CoreWeave submitted three configurations for the DeepSeek-V3 671B workload, all achieving top marks among closed/available cloud submissions. The 8,192-GPU cluster hit target quality in just over two minutes. Even scaling down to 4,096 GPUs completed the training in 3.09 minutes, and 2,048 GPUs in 5.54 minutes. The near-linear scaling efficiency observed across these configurations is a testament to CoreWeave's platform-wide optimizations. Notably, CoreWeave was the sole participant in the v6.0 round to scale a GB300 platform beyond 2,048 GPUs for this demanding workload, proving that their full-stack approach yields more usable performance per GPU than sheer scale alone. For AI teams operating under strict compute budgets, this translates directly into faster development cycles and a quicker path to production.

Consistent Performance Across Scales

The company's engineering prowess isn't limited to the extreme scale of frontier models. CoreWeave also demonstrated impressive performance on smaller, yet still significant, deployments. On a 4,096-GPU NVIDIA GB300 NVL72 cluster, they reached the Llama-3.1-405B reference quality target in 9.77 minutes. This performance was achieved using 20% fewer GPUs compared to larger GB200 deployments, while delivering near-parity results. The technical underpinnings include NVIDIA NeMo Framework Release 26.04, full CUDA graphs, and tailored tensor/pipeline/context-parallel sharding. For more compact deployments, a 64-GPU NVIDIA HGX B200 cluster using InfiniBand delivered competitive results for GPT-OSS-20B and Llama-3.1-8B training, rivaling larger, newer-generation systems. This breadth of performance validates that CoreWeave's advantages benefit customers across all deployment sizes.

The Engine Room: Mission Control and More

Behind these record-breaking numbers is CoreWeave's meticulously engineered infrastructure. CoreWeave Mission Control™ plays a vital role, continuously monitoring hardware and firmware health across systems like the GB300 to ensure a consistent, high-performance baseline for training jobs. Their SUNK scheduler is topology-aware, intelligently placing workloads to maximize data locality and minimize inter-rack communication for complex models. Furthermore, a rail-aware networking strategy balances traffic efficiently, preventing bottlenecks even at multi-thousand-GPU scale. Brendan Burke, Research Director at Futurum Research, commented on the significance, noting that CoreWeave's ability to translate benchmark performance into real-world gains, especially as new hardware emerges, is a critical advantage for AI researchers racing to stay ahead. StartupHub.ai data shows CoreWeave with a score of 67/100, placing it among the competitive AI infrastructure providers, though behind leaders like Nebius (85/100) and Applied Digital (70/100). CoreWeave recently secured verified financials with a $900M junk-bond sale in 2026.

Production-Ready Infrastructure

Crucially, CoreWeave emphasizes that these MLPerf v6.0 results were achieved on the very same production infrastructure available to their customers today. The networking, scheduler, storage architecture, and Mission Control orchestration platform are not benchmark-specific environments but the core systems powering real-world AI workloads. This commitment to production-ready performance is further validated by CoreWeave's consistent top Platinum ranking in SemiAnalysis ClusterMAX™ assessments and strong performance in independent inference benchmarking, such as for Moonshot AI's Kimi K2.6. The company, publicly listed as CoreWeave, Inc. (NASDAQ:CRWV), is solidifying its position as a key player in the AI cloud space.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.