You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用lxml获取Word文档XML中bookmarkStart与bookmarkEnd间的元素?

问题描述

我提取了Word文档的XML内容,想要查找其中所有书签,但书签是由<w:bookmarkStart>和<w:bookmarkEnd>标签配对组成(通过相同的w:id属性关联),而非闭合标签结构。目前用lxml只能找到<w:bookmarkStart>标签,无法获取这两个标签之间的元素,请问该如何实现?

XML示例

<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<w:document xmlns:wpc="http://schemas.microsoft.com/office/word/2010/wordprocessingCanvas" xmlns:cx="http://schemas.microsoft.com/office/drawing/2014/chartex" xmlns:cx1="http://schemas.microsoft.com/office/drawing/2015/9/8/chartex" xmlns:cx2="http://schemas.microsoft.com/office/drawing/2015/10/21/chartex" xmlns:cx3="http://schemas.microsoft.com/office/drawing/2016/5/9/chartex" xmlns:cx4="http://schemas.microsoft.com/office/drawing/2016/5/10/chartex" xmlns:cx5="http://schemas.microsoft.com/office/drawing/2016/5/11/chartex" xmlns:cx6="http://schemas.microsoft.com/office/drawing/2016/5/12/chartex" xmlns:cx7="http://schemas.microsoft.com/office/drawing/2016/5/13/chartex" xmlns:cx8="http://schemas.microsoft.com/office/drawing/2016/5/14/chartex" xmlns:mc="http://schemas.openxmlformats.org/markup-compatibility/2006" xmlns:aink="http://schemas.microsoft.com/office/drawing/2016/ink" xmlns:am3d="http://schemas.microsoft.com/office/drawing/2017/model3d" xmlns:o="urn:schemas-microsoft-com:office:office" xmlns:oel="http://schemas.microsoft.com/office/2019/extlst" xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships" xmlns:m="http://schemas.openxmlformats.org/officeDocument/2006/math" xmlns:v="urn:schemas-microsoft-com:vml" xmlns:wp14="http://schemas.microsoft.com/office/word/2010/wordprocessingDrawing" xmlns:wp="http://schemas.openxmlformats.org/drawingml/2006/wordprocessingDrawing" xmlns:w10="urn:schemas-microsoft-com:office:word" xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main" xmlns:w14="http://schemas.microsoft.com/office/word/2010/wordml" xmlns:w15="http://schemas.microsoft.com/office/word/2012/wordml" xmlns:w16cex="http://schemas.microsoft.com/office/word/2018/wordml/cex" xmlns:w16cid="http://schemas.microsoft.com/office/word/2016/wordml/cid" xmlns:w16="http://schemas.microsoft.com/office/word/2018/wordml" xmlns:w16sdtdh="http://schemas.microsoft.com/office/word/2020/wordml/sdtdatahash" xmlns:w16se="http://schemas.microsoft.com/office/word/2015/wordml/symex" xmlns:wpg="http://schemas.microsoft.com/office/word/2010/wordprocessingGroup" xmlns:wpi="http://schemas.microsoft.com/office/word/2010/wordprocessingInk" xmlns:wne="http://schemas.microsoft.com/office/word/2006/wordml" xmlns:wps="http://schemas.microsoft.com/office/word/2010/wordprocessingShape" mc:Ignorable="w14 w15 w16se w16cid w16 w16cex w16sdtdh wp14">
    <w:body>
        <w:p w14:paraId="2DDA6990" w14:textId="44789F6F" w:rsidR="0067078D" w:rsidRDefault="003F5B0A">
            <w:bookmarkStart w:id="0" w:name="testmark"/>
            <w:proofErr w:type="spellStart"/>
            <w:r>
                <w:t>sometext</w:t>
            </w:r>
            <w:bookmarkEnd w:id="0"/>
            <w:proofErr w:type="spellEnd"/>
        </w:p>
        <w:sectPr w:rsidR="0067078D">
            <w:pgSz w:w="11906" w:h="16838"/>
            <w:pgMar w:top="1417" w:right="1417" w:bottom="1134" w:left="1417" w:header="708" w:footer="708" w:gutter="0"/>
            <w:cols w:space="708"/>
            <w:docGrid w:linePitch="360"/>
        </w:sectPr>
    </w:body>
</w:document>

当前代码

from lxml import etree as ET

ns = {'w': 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'}
ns2 = '{http://schemas.openxmlformats.org/wordprocessingml/2006/main}'

with open('document.xml', 'r', encoding='utf-8') as xml_file:
    tree_word = ET.parse(xml_file)

findall_param = 'w:bookmarkStart'
find_param = 'w:t'

root_word = tree_word.getroot()
field_content = tree_word.findall('.//'+findall_param, ns)

for bookmark in field_content:
    textmarker = bookmark.attrib[f"{ns2}name"]
    print(ET.tostring(bookmark))
    t = bookmark.find('.//w:t', ns)

解决方法

Word的OpenXML书签通过bookmarkStart和bookmarkEnd的w:id属性配对,我们可以通过以下步骤提取两个标签之间的元素:

1. 建立书签ID映射

先收集所有bookmarkEnd标签,用其w:id作为键,元素本身作为值,建立快速查找的映射:

bookmark_end_map = {}
for end in root_word.findall('.//w:bookmarkEnd', ns):
    end_id = end.attrib[f"{ns2}id"]
    bookmark_end_map[end_id] = end

2. 遍历bookmarkStart提取中间元素

对每个bookmarkStart,找到对应的bookmarkEnd,然后遍历它们的共同父节点下的子元素,收集两个标签之间的内容:

for start in field_content:
    start_id = start.attrib[f"{ns2}id"]
    bookmark_name = start.attrib[f"{ns2}name"]
    end = bookmark_end_map.get(start_id)
    
    if not end:
        print(f"警告:书签「{bookmark_name}」没有对应的bookmarkEnd标签")
        continue
    
    parent = start.getparent()
    in_bookmark = False
    bookmark_elements = []
    
    # 遍历父节点子元素,筛选书签范围内的内容
    for elem in parent.iterchildren():
        if elem is start:
            in_bookmark = True
            continue
        if elem is end:
            in_bookmark = False
            break
        if in_bookmark:
            bookmark_elements.append(elem)
    
    # 提取并打印书签内的文本
    print(f"\n书签「{bookmark_name}」的内容:")
    total_text = ''
    for elem in bookmark_elements:
        text_parts = elem.findall('.//w:t', ns)
        if text_parts:
            total_text += ''.join([t.text or '' for t in text_parts])
    print(total_text)

完整代码

整合后的完整代码如下:

from lxml import etree as ET

ns = {'w': 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'}
ns2 = '{http://schemas.openxmlformats.org/wordprocessingml/2006/main}'

with open('document.xml', 'r', encoding='utf-8') as xml_file:
    tree_word = ET.parse(xml_file)

root_word = tree_word.getroot()
bookmark_starts = root_word.findall('.//w:bookmarkStart', ns)
bookmark_end_map = {end.attrib[f"{ns2}id"]: end for end in root_word.findall('.//w:bookmarkEnd', ns)}

for start in bookmark_starts:
    start_id = start.attrib[f"{ns2}id"]
    bookmark_name = start.attrib[f"{ns2}name"]
    end = bookmark_end_map.get(start_id)
    
    if not end:
        print(f"警告:书签「{bookmark_name}」没有对应的bookmarkEnd标签")
        continue
    
    parent = start.getparent()
    in_bookmark = False
    bookmark_elements = []
    
    for elem in parent.iterchildren():
        if elem is start:
            in_bookmark = True
            continue
        if elem is end:
            in_bookmark = False
            break
        if in_bookmark:
            bookmark_elements.append(elem)
    
    print(f"\n书签「{bookmark_name}」的内容:")
    total_text = ''
    for elem in bookmark_elements:
        text_parts = elem.findall('.//w:t', ns)
        if text_parts:
            total_text += ''.join([t.text or '' for t in text_parts])
    print(total_text)

    # 若需查看完整元素XML,取消以下注释
    # print("\n完整元素XML:")
    # for elem in bookmark_elements:
    #     print(ET.tostring(elem, encoding='utf-8').decode())

注意事项

  • 常规情况下,bookmarkStart和bookmarkEnd会处于同一个父节点下;若存在跨父节点的极端情况,需修改遍历逻辑,改为遍历整个文档元素直到找到对应bookmarkEnd。
  • 处理w:t标签时,需考虑文本为空的情况,用or ''避免报错。

内容的提问来源于stack exchange,提问作者user1701820

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 22:10:50