1 articles with this tag
Nvidia's GB200 NVL72 treats 72 GPUs as a single processor and delivers 30x faster LLM inference than the H100, while Jensen Huang argues the software stack running on top is now the harder-to-replicate asset.