You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用python-docx读取内部超链接指向的文档章节内容

解决python-docx解析文档内部章节超链接的问题

docx里的内部章节超链接本质是指向书签(Bookmark),而非直接关联章节。要获取超链接指向的章节内容,需按以下步骤操作:

1. 提取超链接的目标书签名称

遍历文档段落中的超链接,通过关系ID(rId)获取链接目标,去掉前缀#得到书签名称:

from docx import Document

doc = Document("你的文档路径.docx")

# 遍历段落查找超链接
for para in doc.paragraphs:
    for run in para.runs:
        if run.hyperlink is not None:
            rel_id = run.hyperlink.rId
            rel = doc.part.rels[rel_id]
            # 内部链接的target以#开头,截取书签名称
            if rel.target.startswith("#"):
                bookmark_name = rel.target[1:]
                print(f"超链接目标书签:{bookmark_name}")

2. 根据书签定位章节起始位置

python-docx没有直接的书签定位API,需解析底层XML查找书签所在的段落(包括表格内的段落):

from docx.oxml.text.paragraph import CT_P
from docx.oxml.table import CT_Tbl

def find_bookmark_paragraph(doc, bookmark_name):
    # 遍历文档正文的所有元素
    for element in doc.element.body:
        # 检查段落
        if isinstance(element, CT_P):
            if element.xpath(f".//w:bookmarkStart[@w:name='{bookmark_name}']"):
                return doc._body._element_to_object(element)
        # 检查表格内的段落
        elif isinstance(element, CT_Tbl):
            for cell in element.xpath(".//w:tc"):
                for p in cell.xpath(".//w:p"):
                    if p.xpath(f".//w:bookmarkStart[@w:name='{bookmark_name}']"):
                        return doc._body._element_to_object(p)
    return None

# 调用函数定位书签段落
target_para = find_bookmark_paragraph(doc, bookmark_name)
if target_para:
    print(f"书签所在章节标题:{target_para.text}")
else:
    print(f"未找到书签 {bookmark_name}")

3. 提取章节的完整内容

从书签所在段落开始,收集后续的段落、图片、表格内容,直到遇到下一个章节标题(假设标题使用默认的Heading样式):

def extract_section_content(start_para):
    content = []
    current_para = start_para
    while current_para:
        # 添加段落文本
        content.append(current_para.text)
        # 处理段落内的图片
        for shape in current_para.inline_shapes:
            if shape.type == 3:  # 图片类型
                content.append(f"[图片:{shape._inline.graphic.graphicData.pic.nvPicPr.cNvPr.name}]")
        # 处理后续的表格
        next_elem = current_para._element.getnext()
        while next_elem and isinstance(next_elem, CT_Tbl):
            table = doc._body._element_to_object(next_elem)
            table_rows = []
            for row in table.rows:
                table_rows.append("\t".join([cell.text for cell in row.cells]))
            content.append("\n".join(table_rows))
            next_elem = next_elem.getnext()
        # 遇到下一个章节标题则停止收集
        if current_para != start_para and current_para.style.name.startswith("Heading"):
            break
        # 切换到下一段落
        current_para = current_para.next_paragraph
    return "\n\n".join(content)

# 提取并打印章节内容
if target_para:
    section_content = extract_section_content(target_para)
    print("章节内容:")
    print(section_content)

注意事项

  • 若文档使用自定义标题样式,需修改startswith("Heading")的判断逻辑,匹配实际样式名称。
  • 图片处理部分可扩展为保存图片到本地,需额外解析图片的二进制数据。
  • 表格解析仅提取文本内容,如需保留格式需进一步处理XML结构。

内容的提问来源于stack exchange,提问作者user4571139

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 02:35:27