教程演示如何配置布局分析、DocTR OCR、表格提取,并实现自定义实体识别服务,生成结构化 JSONL 数据用于 RAG 工作流。
在本教程中,我们使用 deepDoctection 1.2.x 实现了一个文档智能pipeline,将布局检测、表格结构识别、OCR、阅读顺序重建、标注链接和结构化导出整合在单个工作流中。我们明确配置了基于 DocLayNet 的布局检测、Table Transformer 结构识别和 DocTR OCR 分析器,然后检查生成的 Page 对象,以理解 deepDoctection 如何表示文本、图表、表格、关系、来源和阅读顺序。我们还通过注册自定义对象类型并实现自己的 PipelineComponent 来提取金额和日期实体、同时根据表格特征对文档进行分类来扩展框架。最后,我们使用 ServiceFactory 手动组装自定义pipeline,探索过滤和服务回滚,对处理后的页面进行序列化,并将文档标注转换为适合下游 RAG 和检索系统的有序 JSONL 块。
!pip install -q "deepdoctection" "transformers>=5.2.0" "timm" "python-doctr" "pdfplumber" "networkx" "lxml"
import os
os.environ["DD_USE_TORCH"] = "True"
os.environ["DPI"] = "200"
os.environ["LOG_LEVEL"] = "INFO"
os.environ["ENABLE_DYNAMIC_OBJECT_TYPES"] = "False"
import json, re, textwrap
from pathlib import Path
from collections import Counter
import numpy as np
import matplotlib.pyplot as plt
from IPython.display import HTML, display
import deepdoctection as dd
print("deepdoctection:", dd.__version__)
import transformers.integrations.peft as _hf_peft
if _hf_peft.is_peft_available():
_hf_peft.is_peft_available = lambda: False
print("patched: PEFT adapter lookup disabled for from_pretrained")
!mkdir -p /content/docs /content/imgs
!wget -q -O /content/docs/paper.pdf \
Click to access 2312.13560.pdf
!wget -q -O /content/imgs/finance.png \
https://raw.githubusercontent.com/deepdoctection/notebooks/main/sample/finance/1bcac3899c9cb1c0b0f650b1431d3d52_7.png
PDF = Path("/content/docs/paper.pdf")
PNG = Path("/content/imgs/finance.png")
OUT = Path("/content/out"); OUT.mkdir(exist_ok=True)
def show(img, w=16):
if img is None: return
plt.figure(figsize=(w, w * 1.3)); plt.axis("off"); plt.imshow(img); plt.show()
def analyze_any(pipe, path, **kw):
"""
Dispatch correctly for a directory, a PDF, or a single image file.
DoctectionPipe can stream a directory or a PDF from disk, but a *single*
image has no reader — path= only supplies the file name / provenance, and
the pixels must be handed in via bytes=. Without this you get:
ValueError: When passing a path to a single image, bytes of the image
must be passed
"""
path = Path(path)
if path.is_dir():
kw.setdefault("file_type", [".jpg", ".png", ".jpeg", ".tif"])
return pipe.analyze(path=path, **kw)
if path.suffix.lower() == ".pdf":
return pipe.analyze(path=path, **kw)
if path.suffix.lower() in (".png", ".jpg", ".jpeg", ".tif"):
return pipe.analyze(path=path, bytes=path.read_bytes(), **kw)
raise ValueError(f"unsupported input: {path}")
我们安装所需的 deepDoctection 依赖项,配置其运行时环境,并应用 Transformers 和 PEFT 的兼容性补丁。我们下载整个教程中使用的示例 PDF 和图像文件,并准备输出目录。我们还定义了辅助函数来可视化图像并一致地分析目录、PDF 和单个图像文件。
dd.print_model_infos(add_description=False, add_config=False, add_categories=False)
profile = dd.ModelCatalog.get_profile("Aryn/deformable-detr-DocLayNet/model.safetensors")
print("\nlayout model categories:", profile.categories)
print("is registered:", dd.ModelCatalog.is_registered("Aryn/deformable-detr-DocLayNet/model.safetensors"))
config_overwrite = [
"USE_ROTATOR=False",
"USE_LAYOUT=True",
"USE_LAYOUT_NMS=True",
"USE_TABLE_SEGMENTATION=True",
"USE_TABLE_REFINEMENT=False",
"USE_PDF_MINER=False",
"USE_OCR=True",
"USE_LAYOUT_LINK=True",
"LAYOUT.WEIGHTS=Aryn/deformable-detr-DocLayNet/model.safetensors",
"ITEM.WEIGHTS=deepdoctection/tatr_tab_struct_v2/model.safetensors",
"ITEM.FILTER=['table']",
"OCR.USE_DOCTR=True",
"OCR.USE_TESSERACT=False",
"OCR.USE_TEXTRACT=False",
"OCR.WEIGHTS.DOCTR_WORD=doctr/db_resnet50/db_resnet50-ac60cadc.pt",
"OCR.WEIGHTS.DOCTR_RECOGNITION=doctr/crnn_vgg16_bn/crnn_vgg16_bn-0417f351.pt",
"SEGMENTATION.THRESHOLD_ROWS=0.4",
"SEGMENTATION.THRESHOLD_COLS=0.4",
"SEGMENTATION.FULL_TABLE_TILING=True",
"WORD_MATCHING.RULE=ioa",
"WORD_MATCHING.THRESHOLD=0.3",
"WORD_MATCHING.MAX_PARENT_ONLY=True",
"TEXT_ORDERING.INCLUDE_RESIDUAL_TEXT_CONTAINER=True",
"TEXT_ORDERING.PARAGRAPH_BREAK=0.035",
"TEXT_ORDERING.BROKEN_LINE_TOLERANCE=0.003",
"LAYOUT_LINK.PARENTAL_CATEGORIES=['figure','table']",
"LAYOUT_LINK.CHILD_CATEGORIES=['caption']",
]
analyzer = dd.get_dd_analyzer(config_overwrite=config_overwrite)
print("\n--- pipeline ---")
for sid, name in analyzer.get_pipeline_info().items():
print(f"{sid} {name}")
print("\n--- what this pipeline produces ---")
print(analyzer.get_meta_annotation())
我们检查 deepDoctection 的模型注册表,验证布局模型及其支持的文档类别。我们明确配置分析器以组合布局检测、表格分割、DocTR OCR、单词匹配、阅读顺序重建和布局链接。然后我们初始化分析器并检查其 pipeline 组件及生成的标注类型。
df = analyze_any(analyzer, PDF, session_id="tutorial01", max_datapoints=3)
df.reset_state()
pages = list(df)
print(f"\nparsed {len(pages)} pages")
page = pages[0]
show(page.viz(show_figures=True, show_residual_layouts=True, show_table_structure=True))
print("== narrative text ==")
print(textwrap.fill(page.text[:900], 110))
print("\n== layout blocks in reading order ==")
for doc_id, img_id, pno, ann_id, order, cat, txt in page.chunks[:12]:
print(f"[{order:>3}] {str(cat):<15} {txt[:70]!r}")
print("\n== category histogram ==")
print(Counter(a.category_name for a in page.get_annotation()))
for fig in page.figures:
linked = fig.get_relationship("layout_link")
print("figure", fig.annotation_id[:8], "-> caption ids:", [i[:8] for i in linked])
if page.words:
w = page.words[0]
print("\nword:", w.characters, "| service:", w.service_id,
"| model:", w.model_id, "| bbox:", [round(x) for x in w.bbox])
tbl_pages = [p for p in pages if p.tables]
if tbl_pages:
t = tbl_pages[0].tables[0]
print(f"table {t.number_of_rows}x{t.number_of_columns}, "
f"max_row_span={t.max_row_span}, max_col_span={t.max_col_span}")
display(HTML(t.html))
for row in t.csv[:5]:
print([c[:22] for c in row])
for c in t.cells[:5]:
print(f" r{c.row_number} c{c.column_number} "
f"(span {c.row_span}x{c.column_span}) {c.text[:40]!r}")
else:
print("no table on these pages — the finance.png sample below has one")
我们对示例 PDF 运行配置好的分析器,并从惰性数据流中具体化生成的页面。我们检查叙述文本、阅读顺序块、标注类别、图表-标题关系、单词来源和边界框。我们还通过 HTML、CSV 和单个单元格表示访问检测到的表格,以检查其结构化输出。
@dd.object_types_registry.register("CustomKey")
class CustomKey(dd.ObjectTypes):
"""Custom summary keys — must be registered to be serialisable."""
MONEY_MENTIONS = "money_mentions"
DATE_MENTIONS = "date_mentions"
DOC_FLAVOUR = "doc_flavour"
@dd.object_types_registry.register("FlavourLabel")
class FlavourLabel(dd.ObjectTypes):
TABULAR = "tabular"
NARRATIVE = "narrative"
MIXED = "mixed"
MONEY = re.compile(r"(?:[$€£]\s?\d[\d,.]*|\d[\d,.]*\s?(?:USD|EUR|GBP|million|bn))")
DATE = re.compile(r"\b(?:\d{1,2}[/-]\d{1,2}[/-]\d{2,4}|\d{4}-\d{2}-\d{2}|"
r"(?:Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec)\w*\s+\d{1,2},?\s+\d{4})\b")
class EntityAndFlavourService(dd.PipelineComponent):
def __init__(self, name="entity_flavour", tabular_ratio=0.25):
self.tabular_ratio = tabular_ratio
super().__init__(name)
def serve(self, dp: dd.Image) -> None:
page = dd.Page.from_image(dp, text_container=dd.LayoutLabel.WORD)
text = page.text_no_line_break
money = sorted(set(MONEY.findall(text)))
dates = sorted(set(DATE.findall(text)))
tables = page.tables
table_area = sum((b[2] - b[0]) * (b[3] - b[1]) for b in (t.bbox for t in tables))
ratio = table_area / float(page.width * page.height or 1)
flavor = (FlavourLabel.TABULAR if ratio > self.tabular_ratio
else FlavourLabel.NARRATIVE if not tables
else FlavourLabel.MIXED)
self.dp_manager.set_summary_annotation(
summary_key=CustomKey.MONEY_MENTIONS, summary_name=CustomKey.MONEY_MENTIONS,
summary_value=money)
self.dp_manager.set_summary_annotation(
summary_key=CustomKey.DATE_MENTIONS, summary_name=CustomKey.DATE_MENTIONS,
summary_value=dates)
self.dp_manager.set_summary_annotation(
summary_key=CustomKey.DOC_FLAVOUR, summary_name=flavour,
summary_score=round(ratio, 4))
def clone(self):
return self.__class__(self.name, self.tabular_ratio)
def get_meta_annotation(self) -> dd.MetaAnnotation:
return dd.MetaAnnotation(
image_annotations=(),
sub_categories={},
relationships={},
summaries=(CustomKey.MONEY_MENTIONS, CustomKey.DATE_MENTIONS, CustomKey.DOC_FLAVOUR),
)
for k in (CustomKey.MONEY_MENTIONS, CustomKey.DATE_MENTIONS, CustomKey.DOC_FLAVOUR):
dd.Page.add_attribute_name(k)
我们为提取的金额提及、日期提及和文档风格分类注册自定义对象类型。我们实现了一个自定义 deepDoctection pipeline 组件,分析页面文本和表格覆盖率以生成这些页面级摘要。然后我们将自定义摘要字段公开为 Page 属性,以便我们可以直接从处理的文档中访问它们。
from deepdoctection.analyzer import cfg, ServiceFactory
cfg.freeze(False)
cfg.USE_TABLE_SEGMENTATION = True
cfg.freeze(True)
components = []
layout_detector = ServiceFactory.build_layout_detector(cfg, mode="LAYOUT")
components.append(ServiceFactory.build_layout_service(cfg, detector=layout_detector, mode="LAYOUT"))
components.append(ServiceFactory.build_layout_nms_service(cfg))
item_detector = ServiceFactory.build_layout_detector(cfg, mode="ITEM")
components.append(ServiceFactory.build_sub_image_service(cfg, detector=item_detector, mode="ITEM"))
components.append(ServiceFactory.build_table_segmentation_service(cfg, detector=item_detector))
word_detector = ServiceFactory.build_doctr_word_detector(cfg)
components.append(ServiceFactory.build_doctr_word_detector_service(word_detector))
components.append(ServiceFactory.build_text_extraction_service(cfg, ServiceFactory.build_ocr_detector(cfg)))
components.append(ServiceFactory.build_word_matching_service(cfg))
components.append(ServiceFactory.build_text_order_service(cfg))
components.append(EntityAndFlavourService())
custom_pipe = dd.DoctectionPipe(pipeline_component_list=components)
print("\ncustom pipeline:", list(custom_pipe.get_pipeline_info().values()))
df2 = analyze_any(custom_pipe, PNG)
df2.reset_state()
fin_page = next(iter(df2))
print("flavour :", fin_page.doc_flavour)
print("money :", fin_page.money_mentions[:10])
print("dates :", fin_page.date_mentions[:10])
show(fin_page.viz(show_table_structure=True), w=13)
def skip_if_no_table(dp: dd.Image) -> bool:
return "table" not in {a.category_name for a in dp.get_annotation()}
components[-1].set_inbound_filter(skip_if_no_table)
det_sid = next(sid for sid, n in analyzer.get_pipeline_info().items()
if n.startswith("image_doctr"))
det_comp = analyzer.get_pipeline_component(service_id=det_sid)
df_undo = det_comp.undo(dd.DataFromList([p.base_image for p in pages]))
df_undo.reset_state()
undone = list(df_undo)
print("annotations before/after undo:",
len(pages[0].get_annotation()),
len(dd.Page.from_image(undone[0]).get_annotation()))
我们使用 ServiceFactory 手动组装 deepDoctection pipeline,结合布局分析、表格处理、OCR、文本排序和我们自定义的组件。我们在金融文档图像上执行此自定义 pipeline,并检查检测到的风格、金额、日期和表格结构。我们还应用入站过滤器并演示如何撤销选定 DocTR 服务生成的标注。
for i, p in enumerate(pages):
p.save(image_to_json=False, path=OUT / f"page_{i}.json")
restored = dd.Page.from_file(str(OUT / "page_0.json"))
print("round-trip:", len(restored.get_annotation()), "of",
len(pages[0].get_annotation()), "annotations restored")
records = []
for p in pages:
for doc_id, img_id, pno, ann_id, order, cat, txt in p.chunks:
if txt and txt.strip():
records.append({"document_id": doc_id, "page": pno, "order": order,
"category": str(cat), "annotation_id": ann_id, "text": txt})
for t in p.tables:
records.append({"document_id": p.document_id, "page": p.page_number,
"order": -1, "category": "table_html",
"annotation_id": t.annotation_id, "text": t.html})
(OUT / "chunks.jsonl").write_text("\n".join(json.dumps(r) for r in records))
print(f"\n{len(records)} chunks -> {OUT/'chunks.jsonl'}")
print(json.dumps(records[0], indent=2)[:400])
我们将每个处理的页面序列化为 JSON,同时保留其结构标注而不嵌入原始图像数据。我们重新加载保存的页面并比较标注数量,以验证结构信息在序列化后得以保留。最后,我们将叙述块和表格 HTML 转换为 JSONL 记录,我们可以直接在 RAG、检索和下游文档处理 pipeline 中使用。
总之,我们培养了对 deepDoctection 如何将多个文档分析模型和基于规则的服务编排成可配置处理 pipeline 的实用理解。我们不仅仅运行预定义的分析器,而是通过检查模型注册、控制各个服务、访问结构化页面级标注、提取表格、创建自定义摘要元数据以及组合我们自己的 pipeline 阶段来深入研究。我们还研究了服务过滤和撤销操作如何影响标注,使我们能够更好地控制复杂的文档处理工作流。最后,我们序列化了处理的文档结构,生成了 RAG 就绪的块,为构建文档搜索、知识提取、检索增强生成和其他面向生产的文档 AI 应用提供了可重用的基础。
点击这里查看完整代码。另外,请务必关注我们的 Twitter,不要忘记订阅我们的通讯。你用 telegram 吗?现在你也可以加入我们的 telegram 群了!
需要与我们合作推广您的 GitHub 仓库、Hugging Face 页面、产品发布或网络研讨会吗?联系我们
Sana Hassan,Marktechpost 的咨询实习生,也是 IIT Madras 的双学位学生,热衷于将技术和 AI 应用于解决现实世界的挑战。凭借解决实际问题的浓厚兴趣,他为 AI 与现实生活解决方案的交汇带来了新鲜的视角。