ZHEJIANG — DeepSeek, a Zhejiang-based artificial intelligence company, launched its V4 model on Friday, April 24, 2026. According to a research paper published by the company, the V4 model was trained using a technique called On-Policy Distillation, drawing on outputs from ten separate teacher models.
On-Policy Distillation allows a model to generate its own responses before consulting multiple teacher models for refinement, accelerating the learning cycle, according to the research paper. DeepSeek said the V4 model is based on its V3 design with a series of modifications, according to the research paper. In late January 2025, the company said it used knowledge distillation techniques to train its V3 model.
DeepSeek said that DeepSeek-V4-Pro-Max outperforms GPT-5.2 and Gemini-3.0-Pro on standard reasoning benchmarks through expanded reasoning tokens. In an announcement, the company said DeepSeek-V4-Pro outperforms all rival open-source models in math and coding tasks and ranks second only to Google's Gemini 3.1-Pro in world knowledge. DeepSeek also said the DeepSeek-V4-Flash model offers faster response times and highly cost-effective usage pricing while having similar reasoning abilities to the Pro version.
According to the research paper, DeepSeek V4's performance trails frontier models like GPT-5.4 and Gemini-3.1-Pro by approximately three to six months. OpenAI released GPT-5.2 in December 2025, and Google launched Gemini 3.0 Pro in November 2025.
DeepSeek released its DeepSeek-R1 model in January 2025 with capabilities comparable to ChatGPT and Gemini. The company said it spent less than $6 million on computing costs to train that model.
Australia, Taiwan, South Korea, Denmark and Italy introduced bans or restrictions on DeepSeek-R1, citing privacy and national security concerns. The R1 model's January 2025 release drew attention to DeepSeek's low reported training costs, as the company said computing expenses totaled less than $6 million.
forum Comments (0)
No comments yet. Be the first to comment.