Good Papers

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

DeepSeek-V3.2 improves efficiency via sparse attention, scaled reinforcement learning matching GPT-5, and agentic synthesis, with a special variant surpassing GPT-5 and reaching gold-medal IMO and IOI levels.

DeepSeek-AI, Aixin Liu, Mei, Aoxue, Lin, Bangcai, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bowen Wu, Bowei Zhang, Chaofan Lin, Dong, Chen, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Yang Dejian, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Dai, Fucong, Hao, Guangbo, Guan-Ting Chen, Guowei Li, Zhang, H., Hanwei Xu, Hao Li, Liang, Haofen, Haoran Wei, Haowei Zhang, Haowen Luo, Haozhe Ji, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Qu, Hui

Published Dec 2, 20258 citations▲ 274 on Hugging FacearXiv ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
DeepSeek-V3.2 delivers impressive RL scaling and agent synthesis that genuinely competes at the frontier, yet its sparse-attention gains, missing ablations, and vague "comparable to GPT-5" benchmarks leave critical questions of variance, diminishing returns, and synthetic-data collapse…

Abstract

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance in long-context scenarios. (2) Scalable Reinforcement Learning Framework: By implementing a robust reinforcement learning protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par with Gemini-3.0-Pro, achieving gold-medal performance in both the 2025 International Mathematical Olympiad (IMO) and the International Olympiad in Informatics (IOI). (3) Large-Scale Agentic Task Synthesis Pipeline: To integrate reasoning into tool-use scenarios, we developed a novel synthesis pipeline that systematically generates training data at scale. This methodology facilitates scalable agentic post-training, yielding substantial improvements in generalization and instruction-following robustness within complex, interactive environments.