如何用lxml获取Word文档XML中bookmarkStart与bookmarkEnd间的元素?
问题描述
我提取了Word文档的XML内容,想要查找其中所有书签,但书签是由<w:bookmarkStart>和<w:bookmarkEnd>标签配对组成(通过相同的w:id属性关联),而非闭合标签结构。目前用lxml只能找到<w:bookmarkStart>标签,无法获取这两个标签之间的元素,请问该如何实现?
XML示例
<?xml version="1.0" encoding="UTF-8" standalone="yes"?> <w:document xmlns:wpc="http://schemas.microsoft.com/office/word/2010/wordprocessingCanvas" xmlns:cx="http://schemas.microsoft.com/office/drawing/2014/chartex" xmlns:cx1="http://schemas.microsoft.com/office/drawing/2015/9/8/chartex" xmlns:cx2="http://schemas.microsoft.com/office/drawing/2015/10/21/chartex" xmlns:cx3="http://schemas.microsoft.com/office/drawing/2016/5/9/chartex" xmlns:cx4="http://schemas.microsoft.com/office/drawing/2016/5/10/chartex" xmlns:cx5="http://schemas.microsoft.com/office/drawing/2016/5/11/chartex" xmlns:cx6="http://schemas.microsoft.com/office/drawing/2016/5/12/chartex" xmlns:cx7="http://schemas.microsoft.com/office/drawing/2016/5/13/chartex" xmlns:cx8="http://schemas.microsoft.com/office/drawing/2016/5/14/chartex" xmlns:mc="http://schemas.openxmlformats.org/markup-compatibility/2006" xmlns:aink="http://schemas.microsoft.com/office/drawing/2016/ink" xmlns:am3d="http://schemas.microsoft.com/office/drawing/2017/model3d" xmlns:o="urn:schemas-microsoft-com:office:office" xmlns:oel="http://schemas.microsoft.com/office/2019/extlst" xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships" xmlns:m="http://schemas.openxmlformats.org/officeDocument/2006/math" xmlns:v="urn:schemas-microsoft-com:vml" xmlns:wp14="http://schemas.microsoft.com/office/word/2010/wordprocessingDrawing" xmlns:wp="http://schemas.openxmlformats.org/drawingml/2006/wordprocessingDrawing" xmlns:w10="urn:schemas-microsoft-com:office:word" xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main" xmlns:w14="http://schemas.microsoft.com/office/word/2010/wordml" xmlns:w15="http://schemas.microsoft.com/office/word/2012/wordml" xmlns:w16cex="http://schemas.microsoft.com/office/word/2018/wordml/cex" xmlns:w16cid="http://schemas.microsoft.com/office/word/2016/wordml/cid" xmlns:w16="http://schemas.microsoft.com/office/word/2018/wordml" xmlns:w16sdtdh="http://schemas.microsoft.com/office/word/2020/wordml/sdtdatahash" xmlns:w16se="http://schemas.microsoft.com/office/word/2015/wordml/symex" xmlns:wpg="http://schemas.microsoft.com/office/word/2010/wordprocessingGroup" xmlns:wpi="http://schemas.microsoft.com/office/word/2010/wordprocessingInk" xmlns:wne="http://schemas.microsoft.com/office/word/2006/wordml" xmlns:wps="http://schemas.microsoft.com/office/word/2010/wordprocessingShape" mc:Ignorable="w14 w15 w16se w16cid w16 w16cex w16sdtdh wp14"> <w:body> <w:p w14:paraId="2DDA6990" w14:textId="44789F6F" w:rsidR="0067078D" w:rsidRDefault="003F5B0A"> <w:bookmarkStart w:id="0" w:name="testmark"/> <w:proofErr w:type="spellStart"/> <w:r> <w:t>sometext</w:t> </w:r> <w:bookmarkEnd w:id="0"/> <w:proofErr w:type="spellEnd"/> </w:p> <w:sectPr w:rsidR="0067078D"> <w:pgSz w:w="11906" w:h="16838"/> <w:pgMar w:top="1417" w:right="1417" w:bottom="1134" w:left="1417" w:header="708" w:footer="708" w:gutter="0"/> <w:cols w:space="708"/> <w:docGrid w:linePitch="360"/> </w:sectPr> </w:body> </w:document>
当前代码
from lxml import etree as ET ns = {'w': 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'} ns2 = '{http://schemas.openxmlformats.org/wordprocessingml/2006/main}' with open('document.xml', 'r', encoding='utf-8') as xml_file: tree_word = ET.parse(xml_file) findall_param = 'w:bookmarkStart' find_param = 'w:t' root_word = tree_word.getroot() field_content = tree_word.findall('.//'+findall_param, ns) for bookmark in field_content: textmarker = bookmark.attrib[f"{ns2}name"] print(ET.tostring(bookmark)) t = bookmark.find('.//w:t', ns)
解决方法
Word的OpenXML书签通过bookmarkStart和bookmarkEnd的w:id属性配对,我们可以通过以下步骤提取两个标签之间的元素:
1. 建立书签ID映射
先收集所有bookmarkEnd标签,用其w:id作为键,元素本身作为值,建立快速查找的映射:
bookmark_end_map = {} for end in root_word.findall('.//w:bookmarkEnd', ns): end_id = end.attrib[f"{ns2}id"] bookmark_end_map[end_id] = end
2. 遍历bookmarkStart提取中间元素
对每个bookmarkStart,找到对应的bookmarkEnd,然后遍历它们的共同父节点下的子元素,收集两个标签之间的内容:
for start in field_content: start_id = start.attrib[f"{ns2}id"] bookmark_name = start.attrib[f"{ns2}name"] end = bookmark_end_map.get(start_id) if not end: print(f"警告:书签「{bookmark_name}」没有对应的bookmarkEnd标签") continue parent = start.getparent() in_bookmark = False bookmark_elements = [] # 遍历父节点子元素,筛选书签范围内的内容 for elem in parent.iterchildren(): if elem is start: in_bookmark = True continue if elem is end: in_bookmark = False break if in_bookmark: bookmark_elements.append(elem) # 提取并打印书签内的文本 print(f"\n书签「{bookmark_name}」的内容:") total_text = '' for elem in bookmark_elements: text_parts = elem.findall('.//w:t', ns) if text_parts: total_text += ''.join([t.text or '' for t in text_parts]) print(total_text)
完整代码
整合后的完整代码如下:
from lxml import etree as ET ns = {'w': 'http://schemas.openxmlformats.org/wordprocessingml/2006/main'} ns2 = '{http://schemas.openxmlformats.org/wordprocessingml/2006/main}' with open('document.xml', 'r', encoding='utf-8') as xml_file: tree_word = ET.parse(xml_file) root_word = tree_word.getroot() bookmark_starts = root_word.findall('.//w:bookmarkStart', ns) bookmark_end_map = {end.attrib[f"{ns2}id"]: end for end in root_word.findall('.//w:bookmarkEnd', ns)} for start in bookmark_starts: start_id = start.attrib[f"{ns2}id"] bookmark_name = start.attrib[f"{ns2}name"] end = bookmark_end_map.get(start_id) if not end: print(f"警告:书签「{bookmark_name}」没有对应的bookmarkEnd标签") continue parent = start.getparent() in_bookmark = False bookmark_elements = [] for elem in parent.iterchildren(): if elem is start: in_bookmark = True continue if elem is end: in_bookmark = False break if in_bookmark: bookmark_elements.append(elem) print(f"\n书签「{bookmark_name}」的内容:") total_text = '' for elem in bookmark_elements: text_parts = elem.findall('.//w:t', ns) if text_parts: total_text += ''.join([t.text or '' for t in text_parts]) print(total_text) # 若需查看完整元素XML,取消以下注释 # print("\n完整元素XML:") # for elem in bookmark_elements: # print(ET.tostring(elem, encoding='utf-8').decode())
注意事项
- 常规情况下,
bookmarkStart和bookmarkEnd会处于同一个父节点下;若存在跨父节点的极端情况,需修改遍历逻辑,改为遍历整个文档元素直到找到对应bookmarkEnd。 - 处理
w:t标签时,需考虑文本为空的情况,用or ''避免报错。
内容的提问来源于stack exchange,提问作者user1701820
相关产品推荐
相关产品推荐

