You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何不使用COM提取DOCX文档的批注内容及关联文本?

无需COM提取DOCX批注及关联文本的解决方案

完全可以不依赖Word COM组件实现需求,DOCX本质是基于OOXML的压缩文件,我们可以用python-docx库直接解析文档结构,无需安装Microsoft Word,且跨平台可用。

实现步骤

  1. 先安装依赖库:
pip install python-docx
  1. 提取批注及关联文本的代码实现:
from docx import Document

def extract_comments(doc_path):
    comments = []
    doc = Document(doc_path)
    # 遍历文档中的所有批注
    for comment in doc.comments:
        # 过滤掉回复类批注(对应原COM方案中的c.Ancestor is None)
        if not comment.parent:
            # 获取批注关联的文本范围
            scope_text = ""
            if comment.scope:
                # 遍历范围中的段落和文本块提取内容
                for element in comment.scope._element:
                    if element.tag.endswith('p'):
                        scope_text += element.text + '\n'
                    elif element.tag.endswith('r'):
                        scope_text += element.text
            # 收集批注关联文本和批注内容
            comments.append([scope_text.strip(), comment.text])
    return comments

# 调用示例
report_path = "你的文档路径.docx"
comments = extract_comments(report_path)
for scope, content in comments:
    print(f"关联文本:{scope}\n批注内容:{content}\n---")

代码说明

  • comment.parent用于判断是否为顶级批注(无祖先批注),和原COM方案的c.Ancestor is None逻辑一致
  • comment.scope对应COM中的c.Scope,通过解析其XML元素提取关联文本
  • comment.text直接获取批注的内容,对应COM中的c.Range.Text

这个方案无需依赖Word环境,运行更轻量化,且支持Windows、Linux、macOS多平台。

内容的提问来源于stack exchange,提问作者Finntech

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 00:18:22