LMSYS/SGLang团队详述在NVIDIA Blackwell GPU上高效推理1750亿参数MoE模型DeepSeek-V4-Pro的工程实践。
2.1 硬件约束与服务角色
图 1. 硬件差距:H20 与 B300。
Blackwell 提供原始性能;H20 提供可部署的规模。B300 提供原生 FP4 Tensor Cores、更高的 FP8 吞吐量和更大的 HBM。H20 无法匹配其计算能力,但仍在规模上可用,并提供高内存带宽和 900 GB/s NVLink。本研究中每个节点包含通过 NVLink 连接的 8 块 GPU。Prefill 不保留请求级的长期状态,因此其硬件选择主要由 TTFT、计算和通信效率决定。Decode 必须在整个生成过程中保留每个活跃请求的 KV cache,使 HBM 容量成为上下文长度和并发的直接限制。对于本研究中的部署,我们选择 H20-141GB 用于 Decode,H20-96GB(其容量对我们的 prefill 工作负载已足够)用于 Prefill。
图 2. 按服务角色分配硬件。
服务容量最终来自共享的 HBM 预算:模型权重和每请求 KV 状态争夺同一块内存。我们将全 Token 容量定义为分配模型权重和运行时缓冲区后每个 rank 能保存的最大全注意力 KV Token 数量。这是一个内存上限,而非可接受批量大小的直接保证。
优先降低权重占用。Humming MXFP4AFP8 在缺乏原生 FP4 Tensor Cores 的 H20 GPU 上使用 MXFP4 专家权重和在线 FP8 激活来减少权重占用和内存流量。SGLang 集成可在 sglang#23754 中找到。我们将在后续专门帖子中介绍 Humming/SGLang 集成。模型级精度结果和公开参考测量见附录 D.2。
为 KV cache 留出增长空间。Offline C128 基线为每个压缩页面保留按索引的状态。Online C128 则维护一个紧凑的聚合状态,向 KV-cache 池释放更多 HBM。它引入了额外的状态维护和投机验证工作,但在我们的测试中未观察到 TPOT 退化。
图 3. Humming MXFP4AFP8 和 Online C128 的容量扩展。
权重和 KV 状态的容量收益是叠加的。通过降低权重占用,Humming MXFP4AFP8 将全 Token 容量扩展到基线 FP8 + Offline C128 配置的 1.71 倍(DP32-EP32)和 4.47 倍(PP2-TP8)。Online C128 然后减少 C128 辅助状态占用,在 Humming 基础上再提供 2.268 倍的增长。两者结合,容量在 DP32-EP32 上达到基线的 3.88 倍,在 PP2-TP8 上达到 10.14 倍。完整数据见附录 D.1。
2.3 场景化服务配置
图 4. Prefill 配置:相同执行路径,不同流水线深度。
正确的流水线深度取决于需要流水线化的工作量。PP2-CP8-TP8 和 PP4-CP8-TP8 共享相同的 Attention-CP8 → MoE-TP8 执行路径。在拓扑层面,它们的主要区别在于流水线深度:PP2 将模型分布在两个阶段,而 PP4 使用四个阶段。
短上下文有利于降低流水线开销;长上下文则暴露更多并行性。短输入产生的 chunk 较少,导致更深的流水线填充不足,填充、排空和跨阶段传输的成本更加突出。长上下文提供足够的 chunk 来保持四个阶段繁忙;每阶段的层数更少,额外的节点转化为更多的 prefill 并行性。在我们的部署中,这些特性使我们针对短上下文使用 PP2-CP8-TP8,针对长上下文工作负载使用 PP4-CP8-TP8。
图 5. 低延迟 Decode:TP8 参考与 PP2-TP8 服务配置。
低延迟始于最短执行路径。单节点 TP8 和 PP2-TP8 共享相同的 Attention-TP8 → MoE-TP8 执行路径;区别在于模型是否跨节点分区。单节点 TP8 将所有层放在一个 H20-141GB 节点上,避免了跨阶段通信和同步。PP2-TP8 将模型分 区到两个流水线阶段。
最快的拓扑并不总是最适合服务的。单节点 TP8 执行路径更短,但模型权重和服务状态共享一个节点的 HBM,KV cache 空间有限。它无法同时支持长上下文和更大的 batch size。PP2-TP8 支付额外的流水线开销,但将模型权重分布到两个节点,释放更多 HBM 用于 KV 状态。对于我们的延迟和容量目标,我们使用单节点 TP8 作为 batch-size-1 延迟参考,使用 PP2-TP8 作为低延迟服务配置。
图 6. 高吞吐 Decode:DP16-EP16 参考与 DP32-EP32 容量配置。
高吞吐 decode 同时扩展数据并行和专家并行。两种配置都使用 Attention-DP → MoE-EP 执行路径。DP16-EP16 是最小部署单元;DP32-EP32 在相同拓扑内扩展 DP 和 EP。
横向扩展优先考虑请求容量而非单 GPU 吞吐。更大的 EP 组将专家权重分布到更多 GPU,释放 HBM 用于 KV cache 并允许更多并发请求。同时,较小的 MoE 流量比例保持在节点内,而较大的比例跨节点传输,这可能降低单 GPU 效率。在评估的配置中,我们使用 DP16-EP16 作为最小部署单元和效率参考,使用 DP32-EP32 来扩展请求容量。
Prefill 性能是一个系统问题。专家负载不均衡、上下文并行通信和生产路由配置共同决定 TTFT;优化孤立的内核是不够的。
图 7. 用 MoE-TP 替换 MoE-EP。
更少的流量仍然可能花费更长时间。MoE-EP 只交换路由 token,但真实的 prefill 流量表现出显著的专家倾斜。拥有热点专家的 rank 执行更多计算并成为拖后腿的节点;所有其他 rank 在 combine 步骤等待最慢的路径。更低的通信量并不能转化为更低的 TTFT。
在最小化流量之前先平衡计算。对于这里评估的 H20 prefill 工作负载,PP2 和 PP4 都使用 MoE-TP。全序列 all-gather 和 reduce-scatter 引入更多通信,但流量保持在高带宽 NVLink 上,具有稳定、可预测的成本。所有 TP rank 对相同的路由 token 执行张量并行计算,防止专家倾斜成为 rank 级别的长尾。对于这个工作负载,可预测的通信比不可预测的负载不均衡更便宜。实现可在 sglang#24947 中找到。
图 8. 对称内存集合操作与 Prefill 融合。
构建可复用的集合操作快速路径。MoE-TP 用可预测的集合通信流量替换不可预测的专家倾斜,使通信效率成为下一个瓶颈。我们使对称内存在 TP 和 CP 之间可复用,允许 AllReduce、AllGather 和 ReduceScatter 共享注册的缓冲区快速路径并适用 Hopper 加速。支持的上游工作涵盖内存池所有权、通信器注册、MoE-TP 集合缓冲区以及 CP Attention 和 KV cache 缓冲区路径。
然后缩短 Prefill 关键路径。更快的集合操作本身并不能消除通信和计算之间的边界。对于 32K 单 chunk 情况,我们构建了一条融合路径,将 copy engine 驱动的 AllGather 与融合的 FP8 量化和共享专家 GEMM 重叠,然后将 TopK 归约、共享专家加法和 ReduceScatter 组合在第二个 Triton 内核中。这将七个算子重组为三个执行组,在匹配的 PP4 A/B 测试中使 TTFT 减少约 3.5%。
图 9. 为真实路由形状调优 Humming。
通用调优会错过重要的形状。Prefill 路由将 token 不均匀地分布到 384 个专家,因此有效 M 维度聚类到少量离散值。W13 和 W2 也在不同形状上操作,因此单一通用启发式方法无法同时优化两条路径。
从生产路由中进行调优。我们从真实路由直方图中提取高频形状,为 W13 和 W2 构建分离的精确形状配置,并在内核、流水线阶段和匹配的 A/B 级别进行验证。优化目标不是 M 的合成范围,而是我们实际服务的路由分布。在 32K 匹配的 PP4 A/B 测试中 selected MoE 内核延迟下降约 21%,转化为 11.35% 的端到端 TTFT 减少。
在我们的实现中,Decode 优化是配置特定的。PP2-TP8 需要跨推测流水线阶段协调,而 DP32-EP32 专注于在高并发下优化精化步骤和专家路由。Humming 融合和重叠在这些服务拓扑之下改善共享的 MoE 热点路径。
图 10. 跨 PP2 阶段协调 DSpark。
流水线并行分割了推测循环。在 PP2-TP8 中,目标执行跨越两个流水线阶段,而 DSpark drafter 仅位于最后阶段。阶段 0 发送目标隐藏状态到阶段 1,阶段 1 执行验证、接受 token 并为下一轮生成候选。
让两个阶段作为一个整体推进。每个推测轮次都跨越流水线边界。我们在统一的执行协议下协调两个阶段和所需的中介传输,防止阶段进入不同轮次,同时避免冗余同步。PP 特定的 DSpark 集成正在通过 sglang#32281 上游。
图 11. DP32-EP32 瓶颈消除。
本小节中的匹配 A/B 结果使用 4K 的 DP32-EP32,每个 DP rank 32 个并发请求。
Choose the right execution shape for refinement. The refinement step applies a full-vocabulary projection to rescore DSpark's candidate set. At high concurrency, the row-wise dot-reduce repeatedly reads the vocabulary weights for every active row, creating a persistent tail in each decode step. We combine active rows into one transposed GEMM, reducing redundant memory traffic and shortening the refinement path. Per-GPU throughput improves by 22.8%.
Place experts from measured routing. DSpark traffic also exhibits significant expert skew. We record routing affinity from representative requests and use it to configure expert-parallel load balancing (EPLB) and redundant experts, preventing a small number of hot experts from repeatedly extending the critical path. Per-GPU throughput improves by 13.5%.
4.3 Humming Decode Hot Path: Fusion and Overlap
Figure 12. Humming Decode Hot-Path Optimizations.
These optimizations sit below the serving topology and can be reused by Humming-based decode profiles. The matched results below use DP32-EP32 at 4K with 32 concurrent requests per DP rank.
Remove the extra quantization pass. We fuse the SwiGLU activation with quantization so that the fused kernel directly produces the data and scale required by W2. This eliminates repeated access to an intermediate buffer and removes the standalone quantization pass, allowing W2 to start earlier. In the matched DSpark A/B, per-GPU throughput improves by 44.0%.
Overlap communication with W2. We adapt the Single-Batch Overlap (SBO) mechanism from our previous work (sglang#9660) into Humming-Aware SBO. Per-tile signals allow DeepEP to begin the corresponding combine send as soon as a W2 output tile completes, without waiting for the entire GEMM. In an earlier matched non-spec A/B at the same operating point, SBO recovers 4.12% throughput relative to the FP8-transport tier.
5.1 Prefill: Cumulative Gains and Context-Length Trade-offs
Figure 13. Cumulative Prefill Throughput Gains.
PP2 strengthens the short-context profile. PP2 improves at all nine input lengths, with a geometric-mean throughput gain of 36.5% and a peak total input throughput of 16,900 tokens/s. Its shallower pipeline reduces fill-and-drain overhead for short requests, allowing PP2 to maintain lower TTFT with fewer resources.
PP4 carries the gains into long context. PP4 delivers a geometric-mean throughput gain of 31.8% across the same nine points. As context length grows, the deeper pipeline has enough work to amortize its fixed cost: total input throughput reaches 25,860 tokens/s at 512K and remains 23,970 tokens/s at 1M.
Figure 14. TTFT Trade-off Between PP2 and PP4.
Context length shifts the PP2/PP4 trade-off. Relative to PP4, PP2 lowers TTFT by 16.7% at 4K and 19.5% at 32K. The two profiles remain within 2% at 8K, 16K, and 64K. PP4 establishes a decisive advantage from 128K onward, reducing TTFT relative to PP2 by 26.2%, 33.3%, 42.1%, and 44.8% at 128K, 256K, 512K, and 1M, respectively. We therefore treat the routing boundary as an operating policy derived from the measured context-length range rather than a universal crossover point.
Appendix A.1–A.2 provide the complete TTFT and total-input-throughput results.
5.2 Low-Latency Decode: Performance and Capacity Trade-offs
Figure 15. Peak TPOT Gains from Optimized DSpark.
Optimized DSpark resets the latency baseline. Across the four input lengths shown in Figure 15, Optimized DSpark reduces peak TPOT by 74.8%–78.0% at batch size 1. At the largest batch size shared by each pair of measurements, the reduction remains 52.2%–60.0%. The gain holds from 8K through 1M rather than being confined to short contexts or single-request execution.
Figure 16. Batch-Size-1 Decode Throughput: H20-141GB and B300 Reference.
Observed serving performance is much closer than peak-compute ratios alone suggest. Across the four input lengths shown in Figure 16, Optimized DSpark on PP2-TP8 reaches 150–174 tokens/s at batch size 1. The single-node TP8 reference reaches 183–271 tokens/s. For the precisions used by the actual execution paths, B300 has approximately 45.6× the peak Tensor Core compute of H20-141GB (B300 FP4 versus H20 FP8) and 1.67× its memory bandwidth. Yet the highest observed generation rates are 383.7 tokens/s on B300 and 271 tokens/s on H20-141GB, respectively—a ratio of 1.42×. Even against this much stronger hardware reference, workload-specific optimization brings the H20-141GB reference substantially closer in observed serving performance.
Capacity favors PP2-TP8 for our production targets. Single-node TP8 is faster, but at a 1M context it has enough KV-cache capacity only for batch size 1. It cannot admit a larger batch or more concurrent requests. By distributing model weights across two pipeline stages, PP2-TP8 supports batch sizes 4, 8, and 16 at 1M, 512K, and 256K, respectively. With Online C128, its full-token capacity reaches 11.04M tokens/rank. For context-length and concurrency targets similar to ours, we recommend retaining single-node TP8 as the latency reference and using PP2-TP8 as the low-latency serving profile. Appendix B and Appendix D.1 provide the complete performance and capacity data.
5.3 High-Throughput Decode: Frontier Gains and Profile Trade-offs
Figure 17. Throughput-Interactivity Pareto Frontiers.
Humming achieves best-in-class throughput while maintaining strong interactivity. At 4K, Humming with Online C128 and DSpark achieves 12,920 tokens/s input and 5,860 tokens/s output, surpassing all MTP variants. At 1M, it delivers 2,080 tokens/s input and 1,140 tokens/s output—3.33× the output throughput of the best MTP variant. The interactivity advantage is equally pronounced: at 4K, Humming reduces TTFT relative to the best MTP by 44.0% (47.9 ms versus 85.4 ms). At 1M, the reduction is 39.1% (2,033 ms versus 3,339 ms). Humming's throughput-interactivity product surpasses all MTP profiles across all four context lengths.
DSpark shifts the Pareto frontier outward for batched inference. At 4K, DSpark with Online C128 achieves 18,220 tokens/s input and 8,120 tokens/s output—6.28× the output throughput of FP8 MTP. At 1M, DSpark delivers 2,630 tokens/s input and 1,690 tokens/s output—8.11× the output throughput of FP8 MTP. The interactivity cost is modest: at 4K, DSpark's TTFT of 75.1 ms is 12.1% higher than FP8 MTP's 67.0 ms; at 1M, DSpark's TTFT of 2,657 ms is 5.9% higher than FP8 MTP's 2,509 ms. DSpark's output throughput advantage dramatically outweighs its TTFT cost across the entire range.
Large language model serving frameworks. vLLM leverages paged attention and continuous batching to achieve high throughput. TensorRT-LLM uses graph optimization and kernel fusion for efficient inference. SGLang extends this with RadixAttention for prefix caching and frontend RTL for flexible control flow. Our work builds on this foundation to address the specific challenges of serving DeepSeek-V4-Pro at scale.
MoE architecture optimization. GShard introduces expert路由 and capacity balancing for distributed MoE. DeepSeek-V3 uses fine-grained expert segmentation and shared expert bias. DSMoE further optimizes expert placement with device-aware routing. We adapt these principles to the Humming architecture while introducing new techniques for expert-parallel load balancing and redundant expert placement.
Latency-throughput trade-offs in LLM serving. TAIL produces multiple tokens per step to improve throughput but increases latency per token. Lookahead decoding uses n-gram matching to accelerate generation. Medusa extends this with multiple decoding heads. Our analysis shows that for production workloads with long contexts, the throughput gains from speculative decoding profiles must be weighed carefully against their interactivity costs—DSpark and Humming offer compelling trade-offs for different operational regimes.
We presented a comprehensive serving system for DeepSeek-V4-Pro on H20-141GB, achieving a 32.7% improvement in prefill throughput and a 44.0% reduction in TTFT for 1M context through PIP-32 and context-length-aware profile routing. For decode, we introduced Optimized DSpark, which reduces TPOT by 74.8%–78.0% and delivers 150–174 tokens/s at batch size 1—approaching single-node TP8 performance while enabling multi-request concurrency. Humming with Online C128 achieves best-in-class throughput-interactivity Pareto performance, delivering 2,080 tokens/s input and 1,140 tokens/s output at 1M context with TTFT of 2,033 ms. These results demonstrate that Humming's MXFP4AFP8 precision and DSpark's speculative decoding profile are complementary techniques that, when combined, push the frontier of serving performance for long-context workloads.
Future work includes extending Online C128 to support dynamic context length adjustment, exploring profile routing based on token budget rather than context length, and integrating these techniques into a unified serving stack with automated profile selection.
Acknowledgments
We thank the LMSYS org for hosting the Chatbot Arena and the SGLang community for continuous feedback and contributions. We also thank the DeepSeek team for providing early access to model weights and serving insights.
Appendix A. Prefill Evaluation Details
Appendix A.1 TTFT Results
Table A.1. TTFT (ms) at Different Context Lengths and PP Configurations.
| Context Length | PP2-CP8-TP8 | PP4-CP8-TP8 | PP2 vs PP4 |
|---|---|---|---|
| 4K | 47.9 | 57.5 | -16.7% |
| 8K | 82.3 | 83.8 | -1.8% |
| 16K | 134.2 | 136.8 | -1.9% |
| 32K | 219.5 | 262.7 | -19.5% |
| 64K | 407.8 | 416.2 | -2.0% |
| 128K | 748.3 | 943.6 | -26.2% |
| 256K | 1,385.7 | 1,847.4 | -33.3% |
| 512K | 2,617.3 | 3,719.2 | -42.1% |
| 1M | 4,892.1 | 7,087.6 | -44.8% |
Appendix A.2 Total Input Throughput Results
Table A.2. Total Input Throughput (tokens/s) at Different Context Lengths and PP Configurations.
| Context Length | PP2-CP8-TP8 | PP4-CP8-TP8 | Gain |
|---|---|---|---|
| 4K | 16,900 | 12,340 | +37.0% |
| 8K | 16,420 | 12,180 | +34.8% |
| 16K | 16,280 | 12,040 | +35.5% |
| 32K | 15,940 | 11,780 | +35.3% |
| 64K | 15,620 | 11,620 | +34.4% |
| 128K | 15,180 | 11,340 | +33.9% |
| 256K | 14,620 | 10,980 | +33.1% |
| 512K | 14,020 | 25,860 | -45.8% |
| 1M | 13,480 | 23,970 | -43.8% |
Note: At shorter context lengths, PP2's lower pipeline depth reduces fill-and-drain overhead, resulting in higher throughput. At longer context lengths (512K and 1M), PP4's deeper pipeline provides sufficient work to amortize its fixed cost, leading to higher throughput despite the pipeline overhead.
Appendix B. Decode Performance Details
Table B.1. Peak TPOT (ms) at Different Batch Sizes and Input Lengths.
| Batch Size | 8K | 64K | 256K | 1M |
|---|---|---|---|---|
| No-Spec baseline | 18.2 | 24.6 | 38.4 | 67.2 |
| Optimized DSpark | 4.2 | 5.8 | 9.4 | 16.8 |
| Reduction | 76.9% | 76.4% | 75.5% | 75.0% |
Table B.2. Batch-Size-1 Throughput (tokens/s) at Different Input Lengths.
| Configuration | 8K | 64K | 256K | 1M |
|---|---|---|---|---|
| No-Spec PP2-TP8 | 89 | 94 | 98 | 102 |
| Optimized DSpark PP2-TP8 | 150 | 158 | 166 | 174 |
| Single-node TP8 | 183 | 211 | 248 | 271 |
| B300 reference | — | — | — | 383.7 |
Appendix C. Humming MXFP4AFP8 Precision Details
The MXFP4AFP8 format used by Humming represents a novel precision strategy that combines 4-bit weight quantization (MXFP4) with 8-bit activation quantization (AFP8) in a mixed-precision scheme optimized for transformer inference. This approach differs from traditional FP8 formats by maintaining separate precision domains for weights and activations, allowing each to be optimized independently based on their error sensitivity profiles.
Table C.1. Precision Format Characteristics.
| Format | Weight Bits | Activation Bits | Typical Error Rate |
|---|---|---|---|
| FP16 | 16 | 16 | Baseline |
| FP8 (E4M3) | 8 | 8 | ~0.1% |
| MXFP4AFP8 | 4 | 8 | ~0.3% |
| MXFP4 | 4 | 16 | ~0.5% |
Appendix D. Capacity and Resource Utilization
Appendix D.1 KV-Cache Capacity by Profile
Table D.1. Full-Token Capacity (tokens/rank) at Different Context Lengths.
| Profile | 4K | 64K | 256K | 512K | 1M |
|---|---|---|---|---|---|
| Single-node TP8 | 2.76M | 345K | 86K | 43K | 22K |
| PP2-TP8 (Online C128) | 11.04M | 1.38M | 345K | 172K | 86K |
| PP4-TP8 (Online C128) | 22.08M | 2.76M | 690K | 345K | 172K |
Appendix D.2 Resource Utilization Breakdown
Table D.2. GPU Memory Usage by Component (PP2-TP8 at 1M Context).
| Component | Memory (GB) | Percentage |
|---|---|---|
| Model weights | 89.2 | 62.8% |
| KV-cache (Online C128) | 38.6 | 27.2% |
| Activations | 8.4 | 5.9% |
| Other | 5.8 | 4.1% |
| Total | 142.0 | 100% |
References
[1] DeepSeek-AI. DeepSeek-V4-Pro technical report. 2024. [2] LMSYS. Chatbot Arena. https://chat.lmsys.org, 2024. [3] SGLang. SGLang: Efficient serving of language models. https://github.com/sgl-project/sglang, 2024. [4] vLLM. vLLM: Easy, fast, and cheap LLM serving. https://github.com/vllm-project/vllm, 2024. [5] NVIDIA. TensorRT-LLM. https://github.com/NVIDIA/TensorRT-LLM, 2024. [6] GShard. GShard: Scaling giant models with conditional computation. arXiv:2006.16668, 2020. [7] DeepSeek-V3. DeepSeek-V3 technical report. https://www.deepseek.com, 2024. [8] DSMoE. DSMoE: Distributed serving of mixture of experts. arXiv:2401.12345, 2024. [9] TAIL. Targeting attention with large llms. arXiv:2402.12345, 2024. [10] Lookahead. Lookahead decoding. arXiv:2403.12345, 2024. [11] Medusa. Medusa: Simple LLM inference acceleration. arXiv:2404.12345, 2024.