前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片NEW
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • AI 模拟面试NEW模拟真实面试 + 报告
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
    • AI 定制路线NEW按你的简历现排
    • AI 知识地图NEW串起全站知识点
  • 动态
    • AI 热点NEWAI 每日动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
AI 助手NEW
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片NEW
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • AI 模拟面试NEW模拟真实面试 + 报告
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
    • AI 定制路线NEW按你的简历现排
    • AI 知识地图NEW串起全站知识点
  • 动态
    • AI 热点NEWAI 每日动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
AI 助手NEW
旧版
返回 AI 情报前线
All News · 全部资讯8836
  • 定时 Agent 总重复干活?用完成分类账让它知道什么是「做完」
  • TypeSafe AI Jev 编码指南:类型化决策、置信度校准与推测式广播
  • Vercel Connect 新增 TanStack AI 集成
  • Anthropic与OpenAI 90分钟内相继降价
  • 选LLM API的六个价格陷阱
  • Agent记忆正常仍出错:问题在状态不在记忆
  • Mercury 2.5 推理速度达 770 tokens/秒
  • GitHub Copilot 代码审查新增个人配置选项
  • MCP单人上手容易,团队规模落地是另一回事
  • 影子测试100%一致率背后:模型实际正确率仅75%
  • 两个都通过的测试,代价却不同:重试的隐性成本
  • 定时运行AI Agent输出飘移的根因与修复
  • Anthropic如何两周将Claude.ai速度提升3倍
  • claude-code-templates:一键装配 Claude Code 开发套件,含 100+ Agent/MCP
  • Agent 删改测试必须拦截:CI 合并门禁实操方案
  • Google Antigravity SDK 支持本地 AI 模型:Gemma 4 26B 可离线跑
  • NVIDIA Warp 与 MjWarp 加速机器人仿真工作流
  • HEMA 用 MCP 和 Amazon Bedrock 实现内部 AI 助手转型
  • 五大LLM网关工具生产环境横评
  • GitHub Copilot应用如何渲染百万行PR
  • 基于 AWS 构建 Agent 式视频智能对话系统架构解析
  • Bedrock 上用开源权重模型做 AI 编程助手
  • Anthropic 实验室:Claude 自主发现类 CRISPR 新型酶系统
  • ChatGPT Voice 集成邮件、日历和 Slack:Altman 心中的"Her"更近一步
  • OpenAI GPT-6 Sol/Luna 和 Claude Opus 5.5 同步降价 50%
  • AI 工具循环必须显式传递 Retry-After 头否则必死循环
  • AI 代码补丁静默引入新工具调用:merge 前必须强制契约检查
  • AI 写 API 文档无法区分 null/0/缺省三态:OpenAPI 契约必须显式约束
  • AI 编程 Agent 工具输出遭截断:应记录 stdout_bytes 和截断标志
  • Anthropic工程师揭秘:Claude为何越进化写作越差
  • Gemini 3.8 Flash / Flash-Lite TTS 发布:千款语音、30秒克隆、逐行台词控制
  • 工程师详解:新版Claude为何写作风格变得怪异
  • 小米MiMo-V3将搭载HySparse 2:100万Token下KV缓存缩小4.5倍
  • GitHub Copilot 应用新增本地沙箱隔离功能
  • AI Agent 调试指南:重启不是调试,七层架构定位根因
  • AI 加剧软件供应链攻击威胁,行业如何应对
  • 阿里 Qwen Audio 3.1 发布:语音识别/TTS 多模型,API 价格最高降 95%
  • Claude Code部署到Lizard平台实战指南
  • Claude Opus 5.5降价却破坏四个Agent依赖项
  • OpenAI GPT-6 Sol/Luna 半价发布,缓存机制或为更大降本杠杆
  • treg:聚合 3000+ Agent 工具的统一网关
  • 向量检索权限校验应内嵌到 pgvector 查询中
  • Univer:面向 AI Agent 的开源办公套件 SDK
  • DeepSeek公开Agent训练新论文,梁文锋署名
  • 为 AI 编程 Agent 构建可复用技能系统的实践
  • DeepMind研究:百个AI智能体协作求解时出现作弊与告密现象
  • Google AX:开源Agent编排运行时
  • 阿里千问发布Qwen-Audio-3.1:TTS降价70%、ASR降价95%
  • 诺基亚开源AnyJev:无训练即可将任意开源LLM转为校准决策模型
  • 2026年AI网关横评:Bifrost领跑,多路 failover 哪家强
  • GPT-6 Sol/Luna 发布:准确率翻倍、成本减半,价格战开启
  • 已加载 51 / 8836
8.0
热点
AI SCORE
技术实践2026-09-24 02:21

基于 AWS 构建 Agent 式视频智能对话系统架构解析

AWS ML Blog#AWS#Agent#视频AI
Editor brief · 编辑速览

AWS 博客详解用 Strands Agents SDK 编排 Bedrock、Rekognition、Transcribe 实现自然语言视频分析的代理架构。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

借助 AI 智能体驱动的视频智能,你可以用自然语言对上传的视频提出问题,并在数秒内获得答案。媒体、安保、保险和专业服务等行业的企业正在生成大量视频,超出了团队人工审查的能力。会议录音在共享驱动器中不断堆积,安防摄像头捕捉了数周未审查的 footage。现场检查视频在初次审查很久之后仍躺在对象存储中。视频中的信息往往很有价值:三周前讨论的一个设计决策、某人到达门口的精确时刻,或者导致车辆碰撞的事件顺序。但传统上获取这些信息需要手动观看数小时的内容。另一个方案是为每种特定问题类型构建自定义机器学习(ML)管道,但这需要大量的开发工作。每个新的用例都意味着新的开发工作:

用于会议查询的转录管道。

用于视觉搜索的计算机视觉管道。

人脸比对集成。

在本文中,我们将介绍构建视频智能解决方案的架构和关键模式,该解决方案接受自然语言问题并从视频内容中返回答案。该解决方案采用 AI 智能体架构,在运行时决定调用哪些 AWS 服务。对于已分析过的内容,响应在不到一秒内返回。对新视频的初次分析需要 5–10 分钟,具体取决于视频长度和所需服务。完整的实现可在配套的 GitHub 仓库中获取。

我们不使用为每种问题类型预先构建固定管道,而是使用 Strands Agents SDK 创建一个统一的 AI 智能体,根据用户的问题来编排 Amazon Bedrock、Amazon Rekognition 和 Amazon Transcribe。一家大型媒体和娱乐公司在 AWS Professional Services 项目中采用了这种方法。通过此解决方案,他们的顾问可以查询记录的发现会话内容,提取设计决策、操作项和利益相关方立场。结果:基于客户内部对每位分析师每条录音前后工时的对比(未经独立核实),在 200 多条超过一小时的录音积压中,手动审查时间减少了约 80%。

该解决方案是一个 AI 智能体,接收视频文件并通过自然对话使其内容立即可查询。用户可以上传一段 90 分钟的会议录音,然后问"这次会议做出了哪些决策?"或"有人提到预算时间表吗?"AI 智能体决定是否调用转录、视觉分析或两者,然后综合结果生成连贯的答案。同一系统处理安放 footage 查询("这个人出现过吗?")、内容分析("总结前 30 分钟")和调查性问题("碰撞前哪辆车变道了?")。每个用例不需要单独的處理管道。

以下截图展示了该界面,提供用于自然语言查询的聊天面板和用于文件上传及分析模式选择的侧边栏。

图 1:视频智能聊天界面

关键在于管道是在运行时确定的。AI 智能体对涉及口语内容的问题调用 Amazon Transcribe,对人脸比对转向 Amazon Rekognition,并对先前处理过内容的后续问题重用缓存结果。路由由模型处理,而非应用代码。

要跟随本文中的实现,你需要:

一个 AWS 账户,可访问 Amazon Bedrock(已启用 Anthropic Claude Sonnet)和 Amazon Simple Storage Service(Amazon S3)。对于文档处理,需要 Amazon Bedrock Data Automation(BDA)或 Amazon Rekognition 和 Amazon Transcribe。请参阅 Amazon Bedrock 中按 AWS 区域支持的模型。

Python 3.11 或更高版本,并已安装 Strands Agents SDK(pip install strands-agents strands-agents-tools)。

已配置 AWS Identity and Access Management(IAM)权限的 AWS Command Line Interface(AWS CLI),权限涵盖前述服务。

基本熟悉 AI 智能体概念,如工具调用和推理循环。

该系统由一个连接到多个 AWS AI 服务的 AI 智能体编排器组成,Amazon S3 提供上传视频和缓存分析输出的存储。AI 智能体编排器是推理引擎。它使用 Strands Agents SDK 构建,由 Amazon Bedrock 驱动,使用 Claude Sonnet 或其他支持工具调用的大型语言模型(LLM)。它接收用户的自然语言查询,根据问题确定调用哪些工具,在需要时对多个服务调用进行排序,并将结果综合成对话式响应。AI 智能体维护对话历史,因此后续问题可以在先前分析的基础上继续,无需重新处理。

图 2:解决方案架构

Amazon Rekognition 提供视觉分析,包括检测视频帧中的对象、场景、活动和人脸。当用户的问题涉及视频中可见内容时,AI 智能体会调用 Amazon Rekognition。Amazon Transcribe 将口语音频转换为文本,支持超过 100 种语言的自动语言检测(请参阅 Amazon Transcribe 支持的语言)以及说话人分离。当问题涉及口语内容时,AI 智能体使用 Transcribe。Amazon Bedrock Data Automation(BDA)提供了替代分析路径,在一个 API 调用中结合视频摘要、章节检测和完整转录。当用户希望一步完成全面分析,或者 Amazon Rekognition 或 Transcribe 不可用时,这很有用。所有上传的视频和分析输出都存储在 Amazon S3 中,使用按用户划分的前缀实现多租户隔离。

这三个服务是起始集,而非固定集。因为 AI 智能体从工具描述中选择工具,而非从硬编码的工作流逻辑中选择,所以同一架构可以接受其他服务作为工具。我们将在"超越视频扩展"一节中回到这一点。对于生产部署,我们建议添加 Amazon Bedrock Guardrails,以对 AI 智能体响应实施内容过滤和接地检查,特别是在人脸比对和监控用例中,负责 AI 控制至关重要。

AI 智能体编排的工作原理

在传统的视频分析应用中,开发者定义一个固定的处理管道:上传视频、运行转录、执行视觉分析、展示结果。这种方法无论具体查询是什么,都将每个视频通过相同的步骤处理,用户必须等待完整管道完成后才能提问。AI 智能体方法颠覆了这一模式。通过将对视频文件的最小预上传限制在上传到 S3 桶中,AI 智能体独立地对每个问题进行推理,仅调用回答问题所需的服务。

当用户提交查询时,AI 智能体首先解析意图:用户想要转录摘要、视觉搜索还是人脸匹配?然后检查是否已有相关分析结果被缓存。如果没有,它选择合适的工具,执行它们(当前一个工具的输出馈送另一个工具时可能按顺序执行),并将结果合成为自然语言答案。在我们对 60 分钟视频的测试中,关于视频的第一个问题通常需要 5–10 分钟(转录或视觉分析运行期间)。关于同一内容的后续问题在不到一秒内返回,因为 AI 智能体重用了缓存结果。实际时间因视频长度、分辨率和调用的 AWS 服务而异。

配置 AI 智能体

以下代码展示了完整的 AI 智能体设置。我们定义模型提供者、指导 AI 智能体推理行为的系统提示,以及可用工具集。使用 Strands,整个编排逻辑(决定调用哪些工具、按什么顺序调用以及如何组合它们的输出)由 LLM 处理,而非应用代码。我们展示了两个代表性工具实现(search_faces_in_video 和 analyze_with_bda)。其余工具(包括 transcribe_video 和 analyze_video_visuals)遵循相同的模式,可在 GitHub 仓库中找到。

from strands import Agent
from strands.models.bedrock import BedrockModel
from tools import (
    transcribe_video, analyze_video_visuals,
    search_faces_in_video, analyze_reference_image,
    analyze_with_bda, upload_video
)

model = BedrockModel(
    model_id="us.anthropic.claude-sonnet-4-5-20250929-v1:0",
    max_tokens=4096
)

SYSTEM_PROMPT = """
You are a video intelligence assistant. For each user query:
1. Determine whether it requires spoken content analysis,
   visual content analysis, or both
2. Check if prior analysis results are already cached
3. Invoke the appropriate tools
4. Synthesize results into a clear answer with timestamps
"""

生产环境系统提示词约 250 行源码。以下精简示例展示三条代表性策略(缓存复用、服务降级和多模态编排),而非完整复现提示词:

# --- 缓存管理(节选) ---
CACHE_GUIDANCE = """
调用任何分析工具前,先检查缓存:
- 先调用 get_cached_result(video_id, analysis_type)
- 如果有缓存结果且不超过 24 小时,直接使用
- 如果用户说"重新分析"或"全新分析",跳过缓存
- 新分析完成后,用 cache_result() 存储结果

# --- 工具降级行为 ---
如果工具调用失败或返回低置信度结果:
- 转录失败:建议以 BDA 作为降级方案
- Rekognition 置信度低(<60%):向用户报告不确定性
- BDA 超时:降级为分别调用 Transcribe 和 Rekognition

# --- 多模态编排 ---
当查询需要同时理解音频和视觉内容时:
1. 可能的情况下并行运行 Transcribe 和 Rekognition
2. 跨模态关联时间戳
3. 综合统一答案,引用两个来源
4. 每个主张都注明具体的时间戳
"""

agent = Agent(
    model=model,
    system_prompt=SYSTEM_PROROMPT,
    tools=[transcribe_video, analyze_video_visuals,
           search_faces_in_video, analyze_reference_image,
           analyze_with_bda, upload_video]
)

生产提示词的其余部分列出了可用工具,并定义了文件选择、缓存复用与显式重新分析、BDA 配置与访问被拒绝降级、参考图像搜索、转录与字幕、体育精彩片段、架构图绘制,以及在 BDA 和服务特定分析之间选择等工作流。它还标准化了统一多文件响应、要求清理前确认、复用先前结果回答后续问题,并应用范围和上传进度护栏。

通过此配置,代理自主处理路由、工具排序和响应综合。添加新功能(例如检测屏幕文字)只需定义一个新的工具函数并将其添加到 tools 列表中,无需修改工作流逻辑。

使用 @tool 装饰器定义工具

每个 AWS 服务都作为带有 @tool 装饰器的 Python 函数暴露给代理。函数签名定义参数,docstring 告诉代理何时及如何使用该工具。这个 docstring 至关重要:它是代理的工具说明书。以下示例展示包装 Amazon Rekognition 的人脸搜索工具:

from strands.tools import tool
import boto3

@tool
def search_faces_in_video(
    video_s3_key: str,
    collection_id: str,
    confidence_threshold: float = 80.0
) -> dict:
    """Search for a specific person in video footage.

Use this tool when the user provides a reference photo
    and asks whether that person appears in a video.
    Requires a face collection created first via
    analyze_reference_image.

Args:
        video_s3_key: S3 key of the uploaded video
        collection_id: Rekognition collection with the
            indexed reference face
        confidence_threshold: Minimum confidence for a
            match (default 80%)

Returns:
        Dict with matched_faces containing timestamps
        and confidence scores for each appearance
    """
    rek = boto3.client("rekognition")
    response = rek.start_face_search(
        Video={"S3Object": {
            "Bucket": BUCKET, "Name": video_s3_key}},
        CollectionId=collection_id,
        FaceMatchThreshold=confidence_threshold
    )
    job_id = response["JobId"]
    # Poll for completion and collect results...
    return {"matched_faces": matches}

以下示例展示 BDA 工具,该工具通过单次 API 调用提供全面视频分析(摘要、章节和转录):

@tool
def analyze_with_bda(
    s3_uri: str,
    analysis_types: list[str] = ["SUMMARY", "CHAPTERS", "TRANSCRIPT"]
) -> dict:
    """Perform comprehensive video analysis using Amazon Bedrock Data Automation.

Use this tool when the user wants a full video summary, chapter
    breakdown, or complete transcript, or when Rekognition/Transcribe
    are unavailable. Returns all requested analysis types in one call.

Args:
        s3_uri: S3 URI of the uploaded video (s3://amzn-s3-demo-bucket/key)
        analysis_types: List of analysis types to perform.
            Options: "SUMMARY", "CHAPTERS", "TRANSCRIPT", "IAB"

Returns:
        Dict with keys for each requested analysis type,
        including timestamps and confidence scores
    """
    bda = boto3.client("bedrock-data-automation-runtime")
    response = bda.invoke_data_automation_async(
        inputConfiguration={"s3Uri": s3_uri},
        dataAutomationConfiguration={
            "dataAutomationArn": BDA_PROJECT_ARN
        },
        outputConfiguration={"s3Uri": f"s3://{BUCKET}/bda-output/"}
    )
    invocation_arn = response["invocationArn"]
    # Poll for completion...
    result = poll_bda_result(invocation_arn)
    return {
        "summary": result.get("summary"),
        "chapters": result.get("chapters", []),
        "transcript": result.get("transcript")
    }

多步推理实战

为说明代理如何链接多个工具调用,假设用户上传了一张参考照片并问"这个人出现在我的监控录像中了吗?"代理必须先索引参考人脸,然后在视频中搜索。这是两个相互依赖的顺序操作。以下追踪展示了代理的内部推理过程:

[Think] This is a face-matching query. I need to index
the reference photo first, then search the video.

[Act] analyze_reference_image(image="reference.jpg")
[Observe] {"faces_detected": 1, "face_id": "face_001",
"confidence": 99.8}

[Think] Face indexed. Now searching the video.

[Act] search_faces_in_video(
video_s3_key="security_cam_july5.mp4",
collection_id="user_collection")
[Observe] {"matched_faces": [
{"timestamp": "00:14:32", "confidence": 97.2},
{"timestamp": "00:47:15", "confidence": 94.8}]}

[Response] Yes, the person appeared twice: at 14:32
(97% confidence) and 47:15 (95% confidence).

代理确定了正确的操作顺序并处理了它们之间的依赖关系(人脸搜索需要已索引的集合)。然后它以对话方式呈现结果。没有应用代码定义这个序列。模型根据工具描述和用户问题自行推理出这个顺序。

对于全面分析(当用户问"分析此视频"或"总结此录像"时),代理可以调用 Amazon Bedrock Data Automation(BDA)而不是分别调用 Amazon Rekognition 和 Transcribe。BDA 在单次异步 API 调用中生成视频摘要、按章节分的时间线分解和完整转录:

[Think] The user wants a full summary. BDA provides summary +
chapters + transcript in one call, more efficient than
running Rekognition and Transcribe separately.

[Act] analyze_with_bda(s3_uri="s3://amzn-s3-demo-bucket/meeting.mp4")
[Observe] {"summary": "Team discussed Q3 roadmap...",
"chapters": [{"title": "Introductions", "start": "00:00"},
{"title": "Roadmap Review", "start": "05:32"}, ...],
"transcript": "Welcome everyone. Let's start with..."}

[Response] Here's the meeting summary with chapters:
Summary
The team discussed the Q3 roadmap...
Chapters
- 00:00 - Introductions

当结果不明确时,代理会明确传达不确定性。边界置信度分数(例如 62%)会产生限定性回答:"我在 14:32 发现了可能的匹配,但置信度较低,所以您可能需要手动验证。"如果由于音频质量差导致转录失败,代理会建议替代方案:"音频质量太低,无法进行可靠转录。您希望我改用演示文稿幻灯片的视觉分析吗?"

AI 智能体模式广泛应用于用户需要从视频内容中提取特定信息而无需预先知道需要哪种分析类型的场景。

会议智能——在项目中期加入的团队成员上传先前的会议录像并提出有针对性的问题:"4 月份做出了哪些架构决策?""团队何时同意使用 GraphQL?""总结一下关于身份验证方法的讨论。"代理转录、搜索并总结,返回带有时间戳的答案以引用录像中的具体时刻。

安全与访问监控——建筑物业管理员上传大堂摄像机画面并附带预期访客的照片,问"这个人本周有没有进入大楼?什么时候?"代理在视频中运行人脸匹配并返回带有置信度分数的具体时间戳。

Claims investigation – An insurance adjuster uploads dash-cam footage and asks "Describe the sequence of events before the collision" or "Which vehicle was in the wrong lane?" The agent combines visual scene analysis with audio (verbal reactions, horns) to reconstruct the event timeline.

Extending beyond video with a stable tool contract

The three examples above all analyze video, but nothing about the architecture is video-specific. The agent selects tools from their docstrings, so adding a new capability (or a new modality entirely) is a matter of wrapping another service as a @tool function and describing when to use it. No workflow logic changes. The same orchestrator, cache, and per-user isolation apply unchanged.

Figure 3: Extending the pattern to other modalities

That makes the pattern a general template for multi-modal AI assistants, using either AWS services or third-party models:

Document and diagram understanding with Amazon Textract. A discovery session rarely lives only in video. Add a Textract tool to extract text, tables, and form fields from architecture diagrams and working documents supplied as PDFs, and the agent can cross-reference what was drawn on a whiteboard with what was said in the recording. This enriches the same conversational session that already answers questions about the meeting audio.

Clinical conversations with AWS HealthScribe. Point the same pattern at a clinician-patient audio file and a HealthScribe tool returns a structured clinical note (a turn-by-turn transcript plus extracted sections such as chief complaint and treatment plan), so a user can ask "What follow-up was recommended?" against the recording.

Entity and sentiment extraction, or a third-party model. An Amazon Comprehend tool can pull entities, key phrases, and personally identifiable information (PII) from any transcript the agent produces. A model available on Amazon Bedrock (including third-party models) can be wrapped the same way for domain-specific reasoning.

In each case the extension point is the tool contract, not the pipeline. A team that has built the video assistant already has the scaffolding (orchestration, caching, authentication, and per-user isolation) to stand up an AI assistant for a different modality by adding tools.

The per-query cost depends on which AWS services the agent invokes. After the initial analysis (transcription or visual processing), follow-up questions about the same video only incur Amazon Bedrock reasoning costs because results are cached. The following table shows approximate costs for a 60-minute video:

The $6.00 Amazon Rekognition cost is one-time per-video costs (subsequent queries only incur Bedrock reasoning costs).

Based on AWS service pricing as of July 2025 and the preceding cost table, a typical transcript-based query on a 60-minute video costs approximately $1.50 for the initial transcription plus Bedrock reasoning. Subsequent questions about the same transcribed content cost only $0.05–$0.15 per turn, covering only the Bedrock inference call. Actual costs depend on model selection, input length, and AWS Region. For current pricing, see Amazon Bedrock pricing, Amazon Transcribe pricing, and Amazon Rekognition pricing.

The solution deploys on Amazon Elastic Container Service (Amazon ECS) with AWS Fargate. The Streamlit application and the agent runtime run in Fargate tasks behind an internal Application Load Balancer, and an Amazon CloudFront distribution is the only public entry point. CloudFront reaches the load balancer through a virtual private cloud (VPC) origin, so the load balancer stays in private subnets with no route to an internet gateway and isn't directly reachable from the internet. CloudFront also terminates viewer TLS using its default *.cloudfront.net certificate, which provides a publicly trusted HTTPS endpoint without a custom domain or an AWS Certificate Manager certificate. Amazon Cognito handles authentication (invitation-only, with mandatory multi-factor authentication (MFA) through a time-based one-time password (TOTP) by default), and uploads and cached output are stored in Amazon S3 under per-user prefixes with a 24-hour lifecycle policy.

Figure 4: Deployment architecture on Amazon ECS and AWS Fargate

A single script (./deploy/deploy-ecs.sh) builds the container image, pushes it to Amazon Elastic Container Registry (Amazon ECR), and deploys the AWS CloudFormation stacks. Deployment typically completes in 15–20 minutes, most of which is CloudFront propagation. Full deployment prerequisites, the AllowSelfSignup parameter and its trade-offs, and step-by-step instructions are in the repository README.

Kiro is an AI-powered development environment that supports spec-driven software development by turning high-level ideas into structured requirements, designs, and implementation tasks. We used its spec workflow, persistent project context, and agent hooks to move from concept to a deployable sample while building security into each capability as it took shape.

Specs defined each capability before implementation. The face-matching spec defined inputs (reference photo plus video), expected behavior (index the face, search, and return timestamps), and edge cases (no face detected, low-confidence matches). The transcription spec covered multi-language detection, speaker diarization, and cache behavior for repeated queries. Kiro generated implementation tasks from each spec and maintained context across the full feature lifecycle. Based on the team's prior experience building similar integrations, this compressed what they estimated would typically be a multi-week effort into a focused sprint.

Threat modeling ran alongside the specs, not after them. As each capability was specified, we modeled how it could be abused and captured the result in a living threat model (see docs/threat-model.md in the companion repository). The model works through concrete kill chains (authentication bypass, network exposure, agent exploitation through prompt injection, over-privileged IAM, and data exfiltration through tool outputs) to identify where guardrails are needed. We then codified those guardrails in code: input validation, output filtering, least-privilege IAM policies, and explicit user consent for sensitive operations. This security-integration approach meant that by the time a feature was marked complete, security wasn't a separate review gate — it was already embedded.

Original source

本文由 AI 翻译整理自 AWS ML Blog,原文版权归原作者所有。

阅读英文原文
上一篇
GitHub Copilot应用如何渲染百万行PR
下一篇
Bedrock 上用开源权重模型做 AI 编程助手