前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 场景篇按分类整理的大前端场景考点
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 场景篇按分类整理的大前端场景考点
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
返回 AI 情报前线
All News · 全部资讯9275
  • 加向量数据库前,先用 NumPy 搞定 Embedding
  • 自实现令牌桶限流:比等 429 更优
  • 长时 AI 任务用后台任务队列:状态机比框架重要
  • asyncio 并发调用 40 个 LLM:代码少但坑多
  • AI 项目中 API 密钥的泄密风险与防护层级
  • 电话Agent工程实践:SIP/WebRTC音频路径全解析
  • pgvector索引调优:HNSW与IVFFlat成本精确分析
  • 多租户应用按客户计费:成本与可计费用量必须分开
  • PDF 解析难点深度剖析:字符间距阈值是万恶之源
  • Hugging Face开源OlmoEarth文本嵌入
  • 阿里开源 Qwen3.8:2.4T MoE、激活 95B、256K 上下文
  • MiniMax H3:单个 Transformer 替代视频生成完整流水线
  • DeepSeek V4 Pro 正式版发布:多项测试接近 Fable 5 水平
  • AI编程助手Lovable完成新一轮4亿美元融资,估值达133亿美元
  • MiniMax Music 3.0:开放权重生产级音乐生成模型
  • Microsoft 发布 MindTopo:VLMs 空间推理能力新基准
  • Grok 4.6 发布:剑指 GPT-5.6 Sol,主打长时间 Agent 任务
  • Grok 4.6 中文详解:训练数据、Agent 能力边界与定价
  • 市场份额报告:Google Gemini 份额从 12% 跌至 1.9%
  • 用Embeddings+Reranking+LLM构建可靠的内容分类流水线
  • 推理冷启动从10分钟降至秒级:容器镜像瘦身实战
  • Anthropic研究揭示:Claude Code旧版权限提示97%被机械通过
  • DeepSeek-V4-Pro-0813 悄然上线,支持思考与非思考模式
  • AI 生图 Prompt 审核实战:Node.js 调用 Chat JSON Schema 方案
  • 企业AI分析的隐藏陷阱:语义漂移问题深度剖析
  • AI 正在消除软件工程中层:代码看不懂、没人负责的团队困境
  • Qwen3.8-2.4T-A95B 模型发布
  • TraceMotive:本地优先的AI代理执行追踪调试工具
  • AI编程工具正在离开IDE:终端原生Agent工作流崛起
  • 个人开发者用AI编程的项目架构经验
  • FastAPI五个安全漏洞发现与修复全过程
  • 长文档AI审核的审计设计:Map-Reduce优于检索增强
  • AI Agent辅助发现SharePoint RCE漏洞链(CVSS 9.1)
  • 我用Claude Code将API的P99延迟降低一半
  • AI Agent读了你的secrets并删了生产数据库——PocketOS事故详解
  • 用聊天模型做金融内容审核:结构化输出设计实践
  • Prompt注入攻击原理与防御实践指南
  • 大规模漏洞扫描活动泛滥,攻击者冒充ClaudeBot等AI爬虫
  • 谷歌DeepMind发布手语转文本模型SL2T
  • LFM2.5-VL-3B:边缘设备高性能视觉语言模型发布
  • 通义千问3.8-27B发布,刷新开源大模型参数效率
  • 多Agent协作陷阱:个体测试全过,团队输出仍错误
  • Agent Plugins:Vercel/OpenAI/Microsoft 等联合推出 Agent 技能打包新标准
  • 企业实战:50+ AI Agent在UK主权云上的部署架构
  • ZeroGPU Router:让 AI Agent 用小模型处理例行任务
  • Solv Labs在AWS Bedrock上构建可审计的Agent支付系统
  • CodeBurn:让AI编程投入产出可见化
  • AWS SageMaker HyperPod分层KV缓存:LLM推理新范式
  • 2026年AI Agent安全开发指南
  • 从专有LLM API窃取推理痕迹研究
  • AI代码审查门控:让AI生成的补丁先自证再合入主分支
  • 已加载 51 / 9275
8.0
热点
AI SCORE
技术实践2026-08-12 23:42

推理冷启动从10分钟降至秒级:容器镜像瘦身实战

dev.to · AI#ML推理#容器优化#DevOps
Editor brief · 编辑速览

通过将容器镜像从 25GB 压缩到 6.7GB(移除冗余依赖和重复文件),将冷启动时间从 600 秒降至 60 秒级别,提供了容器化和模型加载的优化路径。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

Co-authored with Netanel Kadosh

You've built a great inference service. Auto-scaling is configured. The model performs beautifully in testing. You ship it to production feeling good.

Then a traffic spike hits at 2 AM.

New pods start spinning up. Users start waiting. Your on-call phone lights up. You watch the dashboards— nodes are healthy, no errors, but the pods just... aren't ready. Eight minutes pass. Then nine. Then ten. Sometimes fourteen.

Fourteen minutes for a service that responds in under 200 milliseconds once it's warm.

So you start digging into the cold start timeline to understand where all that time is actually going. This is exactly what we did and when we profiled the startup sequence, the breakdown was humbling:

The numbers made the problem clear: 90% of our cold start time was spent simply moving massive files around.

Our container image had bloated to a massive 25 GB. Pulling, unpacking and running an image that large was consuming 70% of our startup time. Inside, we found three different inference backends such as PyTorch, TensorRT-LLM and ONNX Runtime—bundled alongside gigabytes of static weight files. After that, we spent another two full minutes just transferring those weights into the GPU.

We didn't need to compromise on a smaller, less capable model nor did we need to pay for more powerful nodes. We just needed to rethink how we packaged our environment, how we pulled it over the network and how we streamed the weights into memory.

We had to make three engineering shifts that helped us cut our cold start time from minutes to seconds without changing a single line of application code.

Phase 1: Shrinking the Image

If you show a DevOps engineer a slow cold start, their immediate knee-jerk reaction is almost always: "The image is too big, shrink it." We are no different. Naturally, our very first instinct was to put this container on a strict diet.

At 25 GB, there was simply too much bloat inside. We did a deep dive into our image layers using dive and identified several indirect dependency sinks. Deleting a few temp files wasn't going to fix this. We had to ask: What does the inference service actually need at runtime?

We started with a heavy nvidia/cuda base image. In an EKS environment running Bottlerocket nodes, the NVIDIA device plugin handles driver exposure to workloads at runtime. Packaging full CUDA dev toolkits into the application image was redundant. We switched to python:3.13-slim, shedding several gigabytes immediately.

Eliminating the Double Tax

We discovered TensorRT components were being installed twice: the C++ development package via apt and the Python runtime via pip. The C++ dev libraries weren't required for runtime inference, so removing them eliminated another 3 GB.

Trimming PyTorch Ecosystem Bloat

Default PyTorch GPU wheels pull in massive libraries like NCCL, cuSPARSE and Triton. While crucial for distributed training, these added roughly 4.5 GB of unused weight for single-node inference.

We considered mounting Python's site-packages externally via an S3 CSI driver to make the image smaller. However, Python imports touch thousands of small files during boot. Turning local disk reads into thousands of network requests replaced a slow image pull with an even slower import phase. We rejected this approach.

The Strategic Shift: CTranslate2 & ONNX

For mixed-model workloads (like Silero VAD paired with LLM generation), tensorrt_llm forced heavy PyTorch dependencies. Even after our initial cleanup, the image was still stubbornly sitting around 15 GB.

We refactored the runtime stack: we converted Silero VAD to ONNX (~50 MB) and adopted CTranslate2 (~100 MB) for LLM generation. This gave us native tensor parallelism without PyTorch or MPI overhead.

Combined, these changes stripped a massive 15 GB of bloat from our base image, bringing our final footprint down to 10 GB.

We patted ourselves on the back, deployed the leaner container and checked the metrics. Reality hit us fast: a 10 GB image is still incredibly heavy and pulling it was still agonizingly slow.

The classic "just shrink the image" DevOps reflex had taken us as far as it could. To go faster, we couldn't just change the size of the data— we had to change how the data was downloaded.

Phase 2: If You Can't Shrink It, Parallelize It (with SOCI)

When we looked under the hood to see why the pull was still taking so long, the real culprit became clear: the default container runtime I/O. Runtimes fetch image layers sequentially over a single network stream and decompress them serially on a single CPU core. On modern cloud nodes with high-bandwidth interfaces and dozens of CPU cores, this approach leaves most host resources completely idle.

To fix this, we enabled Seekable OCI (SOCI) Parallel Pull Mode on our EKS nodes.

SOCI parallelizes both network downloads and extraction. It splits layers into smaller chunks, executes concurrent HTTP range requests to saturate network bandwidth and distributes decompression across all available CPU cores.

You'd expect a fix this powerful to be complicated, but it was actually the easiest part of the project. All we had to do was pass this tiny configuration block into our Bottlerocket node user-data:

[settings.container-runtime]
snapshotter = "soci"

[settings.container-runtime-plugins.soci-snapshotter]
pull-mode = "parallel-pull-unpack"

[settings.container-runtime-plugins.soci-snapshotter.parallel-pull-unpack]
max-concurrent-downloads-per-image = 20
concurrent-download-chunk-size = "16mb"
max-concurrent-unpacks-per-image = 10
discard-unpacked-layers = true

Hitting a New Wall: Storage I/O

When you parallelize network fetches and CPU decompression, disk I/O becomes your new bottleneck. Because SOCI writes decompressed data directly to disk to maintain predictable memory usage, nodes must be backed by fast storage. We ensured our EKS nodes were backed by high-performance NVMe instance store disks or EBS volumes provisioned for at least 400 MiB/s throughput.

Result: Image pull times for our base containers dropped from nearly 5 minutes down to 55 seconds— an 87% reduction in pull latency.

Phase 3: Streaming Weights Directly from S3 to VRAM

Even with a smaller image and parallel pulls, packaging model weights inside a container image creates a fundamental architectural flaw: it treats model data as code. Every time a model updated, we had to rebuild, push and pull a massive new container image.

The obvious solution was to decouple the model weights from the container image entirely and store them externally in Amazon S3.

But getting those external weights into the GPU introduced a new problem. To get the endless flexibility of S3 without a massive latency penalty, we adopted the open-source NVIDIA Run:ai Model Streamer.

Direct-to-VRAM Parallel Streaming

Traditional weight loading from remote storage is agonizingly slow because it forces a sequential hop. The Run:ai streamer fixes this by executing concurrent HTTP range requests directly against S3 and bypassing the disk entirely.

Here is how the workflow changes:

The Traditional load (Sequential):

S3 ➔ [Local Disk] ➔ [CPU RAM] ➔ [GPU VRAM] (Each step waits for the massive download to finish before moving to the next)

The Streamer Architecture (Concurrent):

[ S3 Bucket ] 
     │
     │  (Parallel Network Pulls)
     ▼
[ CPU RAM ] (Acts only as a transit buffer, no disk staging)
     │
     │  (Direct PCIe DMA Transfer)
     ▼
[ GPU VRAM ]

Instead of downloading a 15GB file to disk and moving it step-by-step, the streamer does everything at once. It uses your system RAM simply as a transit pipeline.

As parallel requests pull data from S3, those chunks are instantly injected into the GPU using Direct Memory Access (DMA) over PCIe. Local disk staging is bypassed entirely and network downloads overlap perfectly with GPU ingestion.

Integrating this into vLLM required zero application code changes. With S3 read permissions configured, we simply updated the vLLM execution command:

vllm serve s3://my-bucket-name/my-weights \
  --load-format runai_streamer \
  --model-loader-extra-config '{"concurrency": 32}'

By setting concurrency to 32, the streamer fully saturated our node's network bandwidth.

Result: Loading 15 GB of model weights dropped from 120 seconds down to 4.9 seconds— a 96% reduction in load time.

By addressing each step, we transformed our start time from a 10+ minute liability into a responsive, scalable infrastructure:

These optimizations build directly on top of each other. We stripped unnecessary dependencies from the image, parallelized the remaining layer pulls with SOCI and decoupled model weights entirely by streaming them from S3.

When your GPU services take minutes to start, don't immediately assume you need pre-warmed fleets, smaller models or to change your business logic. First look at where data is moving and eliminate sequential I/O, you might just find that you can fix it all without changing a single line of application code.

Original source

本文由 AI 翻译整理自 dev.to · AI,原文版权归原作者所有。

阅读英文原文
上一篇
用Embeddings+Reranking+LLM构建可靠的内容分类流水线
下一篇
Anthropic研究揭示:Claude Code旧版权限提示97%被机械通过