DeepSeek-V2 released — sparks a "price war" in China's AI industry with rock-bottom API pricing
DeepSeek has released its new Mixture-of-Experts model, “DeepSeek-V2.” Beyond its technical advances, the model drew major attention for triggering a “price war” that swept through China’s entire AI industry, thanks to its remarkably low API pricing.
Details
- Model architecture: A Mixture-of-Experts model with 236 billion total parameters, of which 21 billion are activated per token. Uses a custom architecture combining MLA (Multi-head Latent Attention) and DeepSeekMoE
- Efficiency: Cut training costs by 42.5% and reduced the KV cache needed for inference by 93.3% compared to the previous model
- Price disruption: Priced at about 1 yuan (roughly $0.14) per million input tokens — around 1/70th the cost of GPT-4 Turbo and about 1/7th that of Llama 3 70B
- Industry impact: Prompted major Chinese players like Alibaba and Baidu to cut their own model prices by more than 95%, triggering an industry-wide price war
- Lightweight version: A 16-billion-parameter “DeepSeek-V2-Lite” was also released on May 16
What happened next
The price war sparked by DeepSeek-V2 became the decisive moment that cemented DeepSeek’s reputation for “low cost, high performance.” This trajectory continued with DeepSeek-V3 that December, and shook the world on an even larger scale with DeepSeek-R1 in January 2025.