DeepSeek Open-Sources DSpark, a Speculative Decoding Framework That Cuts Latency by Up to 85%
点击文中橙色高亮词查释义
On June 27, 2026, DeepSeek released DSpark, an for , the technique that lets a large language model generate text faster without losing quality. The code, the model weights, and a research paper were all published on GitHub the same day. DSpark is co-authored by the DeepSeek team and researchers at Peking University, with founder Liang Wenfeng listed among the authors.
DSpark uses a small to propose several tokens at once, in . The main model then checks those tokens in a single batch and accepts the ones it agrees with. Tokens that fail the check are discarded, and the model simply continues from the last accepted point. This pattern keeps output quality close to the original model while moving the bottleneck from generation speed to verification.
According to the team's measurements, running on real DeepSeek-V4 production traffic at the same as the previous system, DSpark cuts user-perceived generation time by 60 to 85 percent. The comparison is MTP-1, the multi- prediction system already serving live users. Both frameworks were tested under identical traffic, so the gain comes from the new approach itself, not from extra hardware.
For users, a 60 to 85 percent drop in generation time means long answers arrive in seconds rather than tens of seconds. For labs, the lesson is that inference efficiency is no longer only about better chips. With last week's OpenAI Jalapeño chip, custom silicon cuts the cost of running a model; with DSpark, smarter a lgorithms cut the cost of generating a . The two paths are now racing side by side.
/生词 · 点击查释义
/课后 5 题
1. When did DeepSeek release the DSpark framework?
2. What does DSpark use to propose tokens before the main model checks them?
3. Roughly how much faster is user-perceived generation in DeepSeek's production tests?
4. Which previous system did DeepSeek use as the comparison baseline for DSpark?
5. What wider point does the lesson make about DSpark and OpenAI's Jalapeño chip?