Research and technical references

DeepSeek-V3 Technical Report (27 December 2024)

27 December 2024

Claim

DeepSeek-AI team, “DeepSeek-V3 Technical Report,” arXiv:2412.19437 (initial 27 December 2024, updated 18 February 2025). DeepSeek-V3 is a “Mixture-of-Experts” model — one that has 671 billion parameters in total but only switches on 37 billion at a time — trained on 14.8 trillion tokens of text. The reported training cost: 2.664 million H800 GPU-hours for the main training, 119,000 for extending its context length, and 5,000 for final tuning, for a total of 2.788 million GPU-hours; at $2 per H800 GPU-hour, that comes to about $5.576 million. That figure leaves out earlier research, trial-and-error experiments, and the cost of the underlying hardware spread over time. Its technical innovations include FP8 mixed-precision training, a way of balancing the load without an extra penalty term, and predicting several tokens at once.

Sources

Referenced in

← Back to the Evidence Base