前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片NEW
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • AI 模拟面试NEW模拟真实面试 + 报告
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
    • AI 定制路线NEW按你的简历现排
    • AI 知识地图NEW串起全站知识点
  • 动态
    • AI 热点NEWAI 每日动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
AI 助手NEW
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片NEW
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • AI 模拟面试NEW模拟真实面试 + 报告
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
    • AI 定制路线NEW按你的简历现排
    • AI 知识地图NEW串起全站知识点
  • 动态
    • AI 热点NEWAI 每日动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
AI 助手NEW
旧版
返回 AI 情报前线
All News · 全部资讯8649
  • Spec-Driven 开发:摆脱 Vibe Coding 的工程化实践
  • AI 编程助手 CSS 好看但上生产就崩的根因与修复
  • Agent 诊断结果上线前如何验证:分离决策与执行
  • AI编码导致CI成瓶颈?我们重新设计了CI流程
  • 从氛围记忆到确定性自主:AI 原生基础设施设计思路
  • 用 Agent 流程开发简单 SPA 的实战经验
  • AI 生成 React UI 的真实问题:从第二页开始组件漂移
  • vLLM 深度解析:高吞吐 LLM 推理的性能瓶颈
  • AWS 开源 AI 编程助手,成本比 Claude Code 低 45%
  • 用Docker Compose构建可复现的AI Agent评测环境
  • 阿里发布Qwen-Image-2.1:7B开源图像生成编辑模型
  • Benchling 用 Bedrock AgentCore 为多租户 AI Agent 构建深度防御安全架构
  • Mac Studio 2026:M5 Ultra 支持最高 512GB 统一内存,可本地运行超大模型
  • Grok 4.7登陆GitHub Copilot,面向Agent化编程
  • AI 安全是工程问题:Agent 堆栈每一层的防护实践
  • Agent 系统的瓶颈不是模型,是架构设计
  • 防御式 Agent 架构:Schema 注入与超时控制
  • 自研研究 Agent 拦截 AI 编程幻觉:文档先行策略
  • 像物理学家一样剪枝 LLM:区块移除的伊辛模型优化
  • Mac mini上的AI开发新范式:OpenClaw与Codex实战
  • Cloudflare Python Workers 正式上线,可用纯 Python 开发边缘应用
  • 浏览器直接给 ESP32 烧录 Claude 写的宏,无需 IDE 或工具链
  • Google 发布 Agent 安全风险报告:5 万美元循环消耗与凭证窃取案例
  • 多 Agent 合规流水线:用 RAG + 自修正架构对抗 WCAG 幻觉问题
  • Anthropic 发布金融领域 Claude 参考智能体套件
  • OpenAI 披露强化学习模型在上下文压缩时插入 Prompt 注入
  • Jev 决策模型实测:0.3秒完成意图分类和工具路由,$0.00004/次
  • Kimi Code Desktop 上线:图形界面 + 内置终端 + Git 状态,支持 Swarm 多 Agent 协
  • StepFun Step 5 Preview:600B 总参数 MoE 模型,1M 超长上下文,10 月开源
  • JSON-Render:通吃 React/Vue/Svelte/React Native 等 10+ 框架的生成式 UI
  • Agent-Native:让 Agent 能力同时暴露给 LLM 工具调用和 UI 操作的 TypeScript 框架
  • 清华联合无问芯穹开源具身智能体RPent,GPT-6 Astra注入机器人
  • AI Agent工具调用边界case:路由错误比想象中更脆弱
  • AI按它能读的契约编程,而非你想要的
  • MCP 远程服务器实现 AI 驱动的确定性 UI 编排
  • 阶跃Step 5 Preview实测:27B参数开源模型冲至Top2
  • 用Responses API构建有边界的GPT-6 Astra Agent
  • 用破坏不变量法审查 AI 生成代码
  • 已加载 38 / 8649
8.0
热点
AI SCORE
技术实践2026-09-22 02:33

vLLM 深度解析:高吞吐 LLM 推理的性能瓶颈

dev.to · AI#vLLM#LLM推理#GPU优化
Editor brief · 编辑速览

揭示 LLM 推理本质是内存带宽问题而非算力问题,详解 vLLM 如何通过 PagedAttention 解决 KV cache 碎片化和 GPU 气泡问题,提升生产级吞吐。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

Editor's Note: Originally published on the g factor engineering blog. All benchmarks and telemetry in this article were conducted on dedicated NVIDIA H100 and H200 clusters on gft-studio.

If you have ever stared at nvidia-smi during a production inference run and felt your heart sink seeing 12% GPU compute utilization while users complained about sluggish generation, you have run headfirst into the central reality of modern LLMs: text generation is a memory bandwidth problem disguised as a compute problem.

When a transformer generates text token-by-token, your multi-thousand-dollar GPU spends almost none of its time flexing its tensor cores. Instead, it spends virtually all its time acting like a high-speed forklift in a warehouse—shuffling tens of gigabytes of model weights and historical Key-Value (KV) cache tensors back and forth across High Bandwidth Memory (HBM) for every single emitted word.

The moment you push beyond single-user toy demos into multi-tenant production or high-throughput reinforcement learning (RL) rollouts, naive PyTorch stacks hit a wall: memory fragmentation chews up your VRAM, static batching leaves GPUs idling in massive "bubbles," and host-side driver overhead leaves silicon starving. That is why vLLM took over the inference world—not through arcane black magic, but through elegant, battle-tested operating systems engineering: PagedAttention and continuous iteration-level batching.

Key Concepts at a Glance

Token: A unit of model input or output, typically corresponding to 3–4 characters of English text.

Prefill phase: Processing the prompt in parallel—dense compute-heavy matrix multiplication (GEMM).

Decode phase: Generating subsequent tokens sequentially—memory-bandwidth-bound matrix-vector operations (GEMV).

KV Cache: VRAM tensors storing past Key and Value representations so attention doesn't recompute historical tokens.

Continuous batching: Iteration-level scheduling that admits waiting requests immediately when any sequence finishes.

1. The KV Cache Bottleneck: Why Memory Bites Back

Autoregressive transformers do not generate a paragraph all at once. When a prompt arrives, the model processes all input tokens in parallel during the prefill phase. This is dense, compute-heavy matrix multiplication (GEMM)—the kind of workload GPUs were born to do.

But the moment generation begins (the decode phase), everything changes. To generate token #501, the model needs self-attention over the preceding 500 tokens. Recalculating all 500 token representations from scratch on every single step would be an O(N²) computational nightmare. So, we cache the intermediate Key and Value vector representations in GPU VRAM:

2 × L × HKV × dhead × T × b

The first factor of 2 accounts for both Keys and Values. L is the number of full-attention layers, HKV is the number of KV heads per layer, dhead is the head dimension, T is the active sequence token length, and b is the precision in bytes (e.g., 2 bytes for BF16).

Crucially, modern open-weight architectures (like Qwen3.6-27B) use Grouped-Query Attention (GQA) to keep memory sane: 16 full-attention layers with 4 KV heads and a head dimension of 256. At an 8,192-token sequence in BF16, that single sequence commands:

2 × 16 × 4 × 256 × 8192 × 2 = 536,870,912 bytes ≈ 512 MB

Half a gigabyte sounds manageable—until you realize what happens under load. In standard PyTorch or basic Hugging Face generate() pipelines, memory management is primitive. Dynamic containers like DynamicCache allocate buffers on the fly. In a multi-user service where one user asks for a 20-line bash script and another submits a 6,000-token legal document, allocating and reallocating memory turns your GPU VRAM into Swiss cheese.

This is external memory fragmentation: you might have 15 GB of total free VRAM reported, but because it is shattered into non-contiguous fragments, the next request asking for a contiguous 2 GB block crashes with a catastrophic CUDA Out of Memory (OOM).

2. PagedAttention: Borrowing a 50-Year-Old OS Masterpiece

Back in the 1960s, operating system pioneers realized that requiring programs to live in contiguous physical RAM was madness. Their solution was virtual memory paging: chop memory into fixed pages (usually 4 KB) and let the hardware map arbitrary virtual addresses to scattered physical pages via a page table.

PagedAttention: reserve a contiguous strip vs. hand out physical pages

vLLM brought this exact insight to GPU memory with PagedAttention. Instead of reserving a giant contiguous chunk of VRAM for each sequence's worst-case length, it chops the KV cache into fixed-size physical blocks (typically holding 16 or 32 tokens).

Visualizing the KV Cache Reservation

Below is a comparison of memory allocation policies when serving 3 concurrent requests (holding 3, 5, and 2 tokens respectively) with a maximum length of 12 tokens:

Figure 1: Contiguous reservation baseline — Each request reserves room for 12 tokens: 36 slots reserved, 10 used, 26 unused (72% memory wasted).

KV Cache Simulation: Reserve Maximum Allocation Policy

Figure 2: Paged block allocation — Four-token blocks reserve 4 + 8 + 4 = 16 slots: 10 used, only 6 unused. Blocks live anywhere in physical memory, freeing 20 slots for other user requests.

KV Cache Simulation: Allocate Blocks As Needed via PagedAttention

A centralized Block Table maps logical token positions to physical blocks:

Logical Seq Position 0 1 2 3 4 5 6 7 8 9
Physical Block ID 0 0 0 0 1 1 1 1 2 2

On-Demand Allocation: Blocks are handed out only when tokens are actually generated. Only the very last block of a running sequence has any unused slots.

Slashing Memory Waste to < 4%: Internal fragmentation virtually disappears. Because you aren't hoarding empty memory buffers for worst-case prompts, the exact same GPU hardware can suddenly host 2x to 4x more concurrent user streams without breaking a sweat.

Copy-on-Write (CoW) Branching: This is a game-changer for reinforcement learning (RL) and parallel tree search. In algorithms like GRPO, the model generates 8 or 16 candidate rollouts from the exact same prompt. With PagedAttention, all 16 candidate rollouts physically share the prompt's KV memory pages. Physical memory is cloned only when individual candidate completions diverge.

3. Continuous Batching: Ending the Tyranny of the Slowest Token

Imagine a city bus that refuses to let any new passengers board until every single person on the bus has reached their final destination, even if three people got off at the first stop and one person is riding all the way to the airport.

That is exactly how traditional static batching operates. If you batch four requests together that produce 50, 120, 240, and 1,024 tokens respectively, the GPU compute cores sit completely idle on three out of the four slots for hundreds of iterations, waiting for that single 1,024-token straggler to finally emit its <eos> token. These wasted cycles are known as GPU bubbles.

vLLM implements continuous iteration-level batching (an architecture pioneered by Orca):

  • The scheduler makes decisions at the boundary of every single forward pass, not at the boundary of whole requests.
  • The millisecond a sequence emits an end-of-sequence token, its physical memory blocks are freed back to the pool.
  • A waiting request from the queue steps into that vacated slot on the very next token iteration. The GPU cores stay continuously saturated, and throughput jumps dramatically.

4. Squeezing the Hardware: CUDA Graphs, Chunked Prefill, and FP8

PagedAttention solves the memory footprint, but getting raw throughput out of modern NVIDIA Hopper silicon (H100/H200) requires tackling kernel dispatch overhead:

The Python Tax and CUDA Graphs (enforce_eager: false): In standard PyTorch eager mode, generating a single token requires the Python runtime to launch dozens of individual GPU kernels in rapid succession across 60+ transformer layers (RMSNorm, QKV projection, RoPE, attention, SwiGLU, down-projection). On Hopper, a single-token GEMV kernel finishes in just 3 to 8 microseconds! But the CPU driver call to dispatch that kernel takes 10 to 15 microseconds. The GPU ends up spending more time waiting for Python to hand it work than actually doing the math. CUDA Graphs solve this by recording the entire sequence of operations into a static execution graph during warmup, allowing the GPU to replay the whole pipeline in a single dispatch.

Chunked Prefill: A massive 8,000-token prompt arriving during active generation used to cause a massive latency spike for everyone else. Chunked prefill breaks long prompts into manageable bites (e.g. 512 tokens) and interleaves them smoothly alongside decode tokens, keeping inter-token latency steady.

FP8 Tensor Core GEMMs: Running weights and activations in 8-bit floating point doubles effective memory bandwidth and unlocks Hopper's specialized Cutlass FP8 matrix cores.

Hot-Swapping LoRA Adapters: In multi-task serving or RL training, you don't want to reboot your inference engine every time weights update. vLLM allows syncing LoRA adapter weights directly into the running worker processes over NCCL in milliseconds.

5. Battle Scars from the Lab: Real Benchmarks from gft-studio

Synthetic benchmark charts on Twitter are easy to fake. We wanted to see what happens when you push real engineering workloads through this stack. Below is empirical telemetry gathered on our research platform (gft-studio) running Qwen3.6-27B on dedicated NVIDIA H100 and H200 SXM clusters.

Benchmark A: Hugging Face vs. vLLM on H100 (Controlled 1x H100 Smoke Test)

In Group Relative Policy Optimization (GRPO), models generate groups of rollouts against external environments. To cleanly isolate the engine, we ran a controlled 3-step test on an identical NVIDIA H100 80GB SXM GPU, keeping the prompt, base weights, random seed, and SQL task strictly identical:

Engine Gen Wallclock (s/step) Total Step Time (s)
Hugging Face 283.7 357.8
vLLM 78.9 151.7

On identical silicon, vLLM cut generation wallclock from 283.7 seconds down to 78.9 seconds per step—a 3.59x raw generation speedup. Total step time dropped by 2.36x.

Caveat: This was a 3-step execution smoke test. While it cleanly isolates engine mechanics on identical hardware, it does not evaluate long-horizon policy convergence over 500 steps.

Benchmark B: How Much Do CUDA Graphs Actually Matter on H200?

To measure host-side Python dispatch overhead in practice, we tested Qwen3.6-27B on NVIDIA H200 hardware with eager execution (enforce_eager: true) versus captured CUDA Graphs (enforce_eager: false):

Mode Decode Time (ms) Speedup
Eager (enforce_eager: true) baseline 1.0x
CUDA Graphs (enforce_eager: false) ~0.23–0.34× baseline 2.95x–4.32x

The numbers speak for themselves: on single-token decode iterations, replaying pre-recorded CUDA graphs reduced generation time by 2.95x to 4.32x simply by removing host-side driver stalls.

Supplemental Telemetry: Recorded Training Rows

Historical multi-hop routing runs (Qwen3.6-27B, vLLM generation):

Metric Value
Total rows recorded 4,800
Median generation time 1.42s
P50 throughput 847 tok/s
P95 throughput 612 tok/s
P99 throughput 389 tok/s

Benchmark C: Production Serving and Interconnect Bottlenecks

In an inference serving pilot using standard AIPerf workloads (564 input tokens, 128 output tokens), we pushed Qwen3.8-27B under increasing concurrency:

Concurrency TP Throughput (tok/s)
1 1 321.2
4 1 892.4
8 1 1,102.3
8 2 982.6
16 1 1,203.8

An Engineering War Story: Notice the 982.6 tok/s peak under Tensor Parallelism (TP2). That speed was achieved on a single node connected via ultra-high-speed NVLink. Earlier in our testing, we attempted a custom cross-node TP2 setup over a standard VPC network interconnect. The result? Throughput collapsed to 75.8 tok/s! Unless your GPUs share high-bandwidth NVLink, do not run tensor parallelism across physical machines; network latency will decimate your throughput. Use data parallelism (independent workers) instead.

6. Production Checklist: Running vLLM Without 3 AM Pages

If you are deploying vLLM in enterprise infrastructure, here is the architecture pattern we rely on:

Two-Tier Ingress Architecture: Place a resilient gateway (like LiteLLM) in front to handle authentication, team quotas, and audit logging. Route traffic across backend vLLM workers using active queue-depth health checks.

Immutable Local Storage: Never make your workers download 50 GB weight files from public object storage on startup. Pre-mount model checkpoints on local NVMe or high-speed read-only PVCs so pods boot in seconds.

Watch Your Colocated Memory Budget: If you run RL post-training where the trainer and the vLLM inference worker live on the same GPU, set gpu_memory_fraction: 0.35. This reserves ~28 GB for vLLM while leaving ~50 GB for PyTorch gradient activations and optimizer states. Neglecting this balance will trigger immediate CUDA OOM crashes the moment your training loss runs over a long trajectory.

vLLM does not magically make models smarter, but it transforms LLM serving from a brittle, memory-starved script into a predictable, rock-solid engineering system. When you respect the hardware, the hardware delivers.

Original source

本文由 AI 翻译整理自 dev.to · AI,原文版权归原作者所有。

阅读英文原文
上一篇
AI 生成 React UI 的真实问题:从第二页开始组件漂移
下一篇
AWS 开源 AI 编程助手,成本比 Claude Code 低 45%