← ← 课程列表AI 研究中高级2026-07-06· 330 words

DeepSeek DSpark Adds DeepSpec Suffix Decay — A Second Look at Speculative Decoding at Scale

/听全文·· MP3 · 330 词

点击文中橙色高亮词查释义

On June 27, 2026, DeepSeek and a research group at Peking University released DSpark, an open-source framework that speeds up large language model inference through a refined form of . The release also included DeepSpec, a full training toolkit that lets developers build their own draft models for popular open models such as Qwen3 and Gemma. Both the paper and the code are publicly available, and the team has already deployed DSpark across the live DeepSeek-V4 serving system.

At the core of DSpark is a drafting layer. Traditional parallel drafts are fast, but their accuracy drops sharply near the end of each block, a problem the team calls suffix decay. DSpark keeps the parallel backbone for speed and adds a small sequential module on top, so that tokens inside one block still depend on each other. This single change raised the average accepted length per round by 16 to 31 percent compared with earlier open draft designs, without adding noticeable compute.

A confidence-scheduled verifier then decides how many draft tokens deserve a full check from the main model. Each token is paired with a , and a lightweight scheduler chooses the based on current load. Under light traffic the budget grows so a single user gets the fastest possible answer; under heavy traffic it shrinks so the system stays stable. The published numbers show single-user generation speed gains of 60 to 85 percent on DeepSeek-V4-Flash and 57 to 78 percent on DeepSeek-V4-Pro, with output that remains .

Beyond the model itself, the practical impact lies in the open release. With DeepSpec, a small team can train a for a new base model in a few hundred GPU-hours and plug it into the same serving loop. Industry watchers read DSpark as a signal that the next round of AI competition is moving from raw parameter counts to inference engineering, where and cost decide who can actually ship agent and voice products at scale.

/生词 · 点击查释义

/课后 5 题

  1. 1. When did DeepSeek and Peking University release DSpark?

  2. 2. What problem does DSpark's semi-autoregressive drafting layer mainly solve?

  3. 3. How much faster is single-user generation on DeepSeek-V4-Flash after DSpark?

  4. 4. Why is DSpark described as lossless?

  5. 5. What does the open release of DeepSpec mainly enable for small teams?

5 / 5