Research and technical references
DeepSeek-V3 Technical Report (27 December 2024)
Claim
DeepSeek-AI team, “DeepSeek-V3 Technical Report,” arXiv:2412.19437 (initial 27 December 2024, updated 18 February 2025). DeepSeek-V3 is a “Mixture-of-Experts” model — one that has 671 billion parameters in total but only switches on 37 billion at a time — trained on 14.8 trillion tokens of text. The reported training cost: 2.664 million H800 GPU-hours for the main training, 119,000 for extending its context length, and 5,000 for final tuning, for a total of 2.788 million GPU-hours; at $2 per H800 GPU-hour, that comes to about $5.576 million. That figure leaves out earlier research, trial-and-error experiments, and the cost of the underlying hardware spread over time. Its technical innovations include FP8 mixed-precision training, a way of balancing the load without an extra penalty term, and predicting several tokens at once.
Sources
- arXiv:2412.19437 ↗
Referenced in
- Framework §2
- AI-SAF-N Models pillar
- Refining §3.2