DeepSeek-V3, the model that turned heads in early 2025, cost roughly $5.6 million to train on H800 GPUs, according to the company's own technical report — a number that helped reframe global assumptions about the compute cost of frontier AI. Its successor, DeepSeek-V4-Pro, released on April 24, 2026, scales to 1.6 trillion total parameters while keeping API input pricing at $1.74 per million tokens, around 13 times cheaper than comparable US models at launch.
How DeepSeek built frontier AI for under $6 million
The headline number from DeepSeek's V3 technical report, published in December 2024 on arXiv, is 2.788 million H800 GPU hours for the full training run, which at roughly $2 per GPU-hour comes to approximately $5.576 million. The model has 671 billion total parameters but uses a Mixture-of-Experts architecture that activates only 37 billion per token, reducing inference and training compute relative to a dense model of the same nominal size. Pre-training covered 14.8 trillion tokens in under two months on a 2,048-GPU cluster.