前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 场景篇按分类整理的大前端场景考点
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 场景篇按分类整理的大前端场景考点
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
返回 AI 情报前线
All News · 全部资讯9240
  • GitOps风格管理AI Agent记忆与工具配置版本
  • Gemini 3.7 Flash 发布:编程能力提升50%降价一半
  • TraceMotive:面向 AI Agent 执行的本地优先调试器
  • AI 生成代码应视为假设:给免费模型答案分配错误预算
  • 依赖距离检测:AI 改动的爆炸半径分析脚本
  • 别盲目合并 AI Patch:Git Worktree Replay Gate 方法论
  • Anthropic研究:AI Agent会争抢地盘并自发合作
  • Claude Code 新会话默认启用 Auto 模式
  • OpenAI 发布 GPT-5.6 三子型号系列
  • Cerebras 刷新 GPT-5.6 推理速度纪录
  • AI Agent生产落地:从Prompt链式调用到基础设施架构
  • AI原生企业构建:从Copilot到自主运营
  • Google Meet 测试版推出 AI 会议笔记功能,Gemini 自动生成纪要
  • Gemini 3.7 Flash 发布: 编程/Agent 能力大幅提升,每百万 Token 0.75 美元
  • 测试套件绕过隔离写入生产数据库的真实事故复盘
  • AI 生成的认证中间件: undefined === undefined 导致认证绕过
  • DeepSeek开源Agent Harness:一切皆插件的Node.js运行时
  • Node.js 多 Provider 代码审查摘要方案:JSON Output 实战对比
  • Hugging Face整合机器人开发工具链
  • Trace 不是 Governance:Agent 系统长期运行的架构边界
  • LangGraph 多 Agent 流水线实战:18 天开发 doc2slides 的工程权衡
  • 自建 LLM 路由:用对模型省 28x 成本
  • 训练数据仅占 0.3%,却占输出 36%:微调模型的数据污染陷阱
  • Gemini 3.7 Flash发布
  • Gemini 3.7 Flash三周内连发
  • Google Sheets Canvas:自然语言构建交互式数据看板
  • CI/CD AI Agent 部署前需加人工审批门
  • MCP + RSS 目录:让 AI 工具获取领域最新信息
  • Agent工具调用200 OK背后的谎言:完成所有权问题
  • Google Sheets Canvas功能详解:打造互动数据看板
  • DeepSeek V4 Pro正式版发布,Agent框架Harness开源,API涨价
  • GPU 数值格式详解:FP32/BF16/FP8/FP4 交互式理解指南
  • 长时 Agent 循环如何设计记忆机制
  • 深度体验DeepSeek Harness:涨价也在情理之中
  • AWS AgentCore:统一监控多云/本地AI Agent
  • DFlash speculative decoding 测评:tokens/s 指标可能误导
  • Claude文本水印技术原理解析
  • 为什么AI demo便宜、上线后成本却翻10倍
  • 用 Bedrock AgentCore Browser Tool 自动化遗留 Web 系统
  • Liquid AI 发布 3B 端侧视觉语言模型
  • AI 编程 Agent 通过测试也可能做出架构错误决策
  • 基于 Bedrock AgentCore 构建 M&A 尽职调查多智能体系统
  • Token降价90%但我的AI账单没降:Jevons悖论在起作用
  • 我的Agent替我做营销:我只负责点批准
  • 文本 AI 水印容易被移除,技术局限性分析
  • ai-prompt-firewall:拦截敏感信息外泄的 Node.js 库
  • 用 FastMCP 快速为 Claude Code 构建自定义工具服务器
  • Ling 3.0 Flash 同尺寸最聪明开源模型
  • AI 辅助团队开发实际吞吐量研究:REPL 到 Swarm
  • 神秘模型 mona-lisa-1 曝光:疑似 GPT-Image 继任者
  • OpenAI Astra 和 ChatGPT 6 曝光: 多 Agent 协作与 GPT-6 真身
  • 已加载 51 / 9240
8.0
热点
AI SCORE
编程提效2026-08-14 01:14

LangGraph 多 Agent 流水线实战:18 天开发 doc2slides 的工程权衡

dev.to · AI#LangGraph#多Agent#RAG
Editor brief · 编辑速览

拆解 PDF 转定制 PPT 的多 Agent 流水线:Parser→RAG Summarizer→Planner→Audience-specific Generator,详述 18 天开发中拒绝「聪明」方案、坚持可部署性的工程决策。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

I spent 18 days building an AI product that converts research papers into audience-tailored PowerPoint presentations. Not a toy — a real deployed thing at doc2slides on Railway that anyone can use.

The interesting parts weren't the "make it work" moments. They were the tradeoffs I had to make honestly, and the times I resisted the temptation to add a "clever" fix that would have made things worse.

This post is about those decisions.

Doc2Slides takes a PDF and produces a .pptx file tailored to four audiences:

Kid — fun analogies, simple words

Student — educational, terms defined

Engineer — technical depth, assumes domain knowledge

Executive — business focus, impact-oriented

The magic is that the same paper produces radically different output based on the audience. A compiler theory paper for a kid becomes "compilers are like magic helpers." The same paper for an executive becomes "advancing compiler technology with formal frameworks."

Code: github.com/manasviboineypally/doc2slides

The architecture: 5 agents in LangGraph

I built this as a multi-agent pipeline instead of one giant LLM prompt. Here's the flow:

PDF Upload
    ↓
Parser        → extracts sections + metadata
    ↓
Summarizer    → RAG-based section summarization
    ↓
Planner       → designs slide structure for audience
    ↓
Writer        → generates audience-adaptive slide content
    ↓
Builder       → produces editable .pptx file

Each agent is an independent node in a LangGraph state machine. They share a TypedDict state and read/write specific fields.

Here's what the graph definition actually looks like:

from langgraph.graph import StateGraph, END
from app.agents.state import AgentState
from app.agents.parser import parser_agent
from app.agents.summarizer import summarizer_agent
from app.agents.planner import planner_agent
from app.agents.writer import writer_agent
from app.agents.builder import builder_agent

def build_pipeline():
    graph = StateGraph(AgentState)

    graph.add_node("parser", parser_agent)
    graph.add_node("summarizer", summarizer_agent)
    graph.add_node("planner", planner_agent)
    graph.add_node("writer", writer_agent)
    graph.add_node("builder", builder_agent)

    graph.set_entry_point("parser")
    graph.add_edge("parser", "summarizer")
    graph.add_edge("summarizer", "planner")
    graph.add_edge("planner", "writer")
    graph.add_edge("writer", "builder")
    graph.add_edge("builder", END)

    return graph.compile()

Why LangGraph over a sequential chain? Adding a new agent is a 2-line change to the graph. In a sequential chain, adding a new step often means refactoring the previous ones. State-based multi-agent design scales better.

The interesting tradeoff #1: My RAG top-1 precision is 42%

I built an evaluation harness because I wanted to measure quality, not just claim it. Three eval types:

Parser evals — deterministic ground-truth assertions

RAG evals — top-K precision on hand-labeled query→section pairs

Summarizer evals — LLM-as-judge scoring faithfulness, completeness, clarity

The parser evals scored 100% (34/34 checks). The summarizer evals averaged 4.4/5.

But the RAG top-1 precision came in at 42%. Only 3 of 7 queries returned the correct section as the top result.

My first instinct: hide the number. Report top-3 (57%) instead.

What I did instead: publish both numbers and explain why.

Looking at the failures revealed a real limitation of RAG:

Query: "how does the genetic algorithm work?"

Expected section: Methodology

Actual top result: 3.6 Stopping Criteria (a subsection of methodology)

Genetic algorithms are discussed in 6 subsections (3.1 through 3.6). Vector search returns the highest-scoring chunk, not the highest-scoring section. For queries about broad topics, subsections often outrank the parent section because they mention the specific term more densely.

This is a known problem in RAG. Solutions include:

  • Hierarchical retrieval (search subsections, bubble to parent)
  • Query rewriting to be more specific
  • Retrieve top-K and let an LLM pick the right section

None of these are fixed today. But I know exactly what's broken and why — which is more useful than pretending it works.

Lesson: deterministic metrics beat vibes. Vibes let you convince yourself the AI is smart. Metrics tell you where it's dumb.

The interesting tradeoff #2: I refused to use word count as a proxy for content density

Users can request any number of slides between 3 and 50. When the paper's actual content density doesn't match the requested slide count, the LLM either pads shallow sections or compresses dense ones. This creates mild redundancy at high slide counts.

The obvious fix: allocate slides based on section word count. Long section = more slides. Short section = fewer slides.

I almost built this. Then I realized: word count is not content density.

A 100-word section with 3 distinct concepts should get multiple slides

A 2000-word section rambling around one idea should get one slide

Word count would systematically reward verbose sections and penalize concise ones. That's not a fix — it's a bug with math.

What I did instead: documented the tradeoff and shipped without the heuristic. From the project's testing_notes.md:

Rejected quick fix: using section word count as a proxy for content density. Word count is not density — a short section may contain multiple distinct ideas while a long section may ramble around one.

Proper solution deferred: content-aware slide allocation with LLM judgment, verified by an evaluation harness that measures output quality against ground truth. Requires infrastructure work not appropriate for the initial version.

Lesson: the right answer to "should I add this heuristic?" is often "no." Heuristics feel like progress. Sometimes they're anti-progress dressed up as pragmatism.

The interesting tradeoff #3: SQLite dev → PostgreSQL prod is one variable

I built with local SQLite during development but deployed to Railway with PostgreSQL. The migration was one line:

# app/db/session.py
DATABASE_URL = os.getenv("DATABASE_URL")
engine = create_engine(DATABASE_URL, echo=False)

For local dev, .env has:

DATABASE_URL=sqlite:///./doc2slides.db

For Railway, the environment variable is:

DATABASE_URL=postgresql+psycopg2://postgres:xxx@host:5432/railway

Nothing else changes. SQLAlchemy models are backend-agnostic.

This is boring engineering. But boring engineering is what lets you sleep at night. When someone asks "how do you handle database migrations?" the answer isn't a clever hack — it's "environment-driven configuration and a repository pattern."

The frontend is worth calling out. I used no framework — just HTML, CSS, and vanilla JavaScript in ~500 lines. Zero build step. Anyone can clone the repo, open the file, and understand it in 5 minutes.

For an MVP, that's a feature, not a limitation.

What I didn't build (and why that's OK)

Multi-tenant workspaces

Custom presentation templates

Job queue with Celery/Redis

Why: MVP. Every feature has a cost. Shipping the core value (PDF → audience-tailored slides) matters more than shipping every possible feature.

For a portfolio project, "I could have added X but chose not to for these reasons" is a stronger answer than "I added X poorly."

Lessons I'd tell my past self

1. Build evals before optimizing. I built the pipeline first, then evals. If I had built evals first, I would have known earlier that my RAG had issues. Now I have to make eval-driven improvements Week 3.

2. Resist heuristics. Every time I thought "this is a quick fix," it was actually a technical debt I was about to bake in. Word count as density. Silent AI slide count overrides. Boolean status flags instead of proper enums.

3. Deploy early. I deployed on Day 16 of 18. I should have deployed on Day 8. Deployment reveals real bugs — environment variable typos, missing dependencies, hardcoded localhost URLs. The sooner you find them, the cheaper they are.

4. Document tradeoffs, not features. Anyone can read code to know what it does. Almost no one leaves notes on why a design choice was made. My testing_notes.md file is where most of the actual engineering thinking lives.

The project is live, but not "done." Future work:

  • Content-aware slide count (with an eval harness measuring output quality)
  • Multi-language support for input PDFs
  • Custom presentation templates
  • Fix RAG for hierarchical sections (subsection → parent bubbling)

If you want to try Doc2Slides yourself:

Live demo: web-production-6eded.up.railway.app

Code: github.com/manasviboineypally/doc2slides

60-second video demo: Loom link

Upload any PDF, pick your audience, get back a deck. Same paper, radically different output depending on who you say you're presenting to.

Author: Manasvi Boineypally — GitHub · LinkedIn

Original source

本文由 AI 翻译整理自 dev.to · AI,原文版权归原作者所有。

阅读英文原文
上一篇
Trace 不是 Governance:Agent 系统长期运行的架构边界
下一篇
自建 LLM 路由:用对模型省 28x 成本