如何从docx XML文件跟踪更改状态及评论resolved状态?构建相关应用
从DOCX的XML文件中获取评论的解决状态
要跟踪Word文档评论的resolved(已解决)/未解决状态,核心是解析DOCX压缩包内的两个关键XML文件:word/comments.xml和word/commentsExtended.xml。以下是具体实现思路:
1. 先获取DOCX内的XML文件
DOCX本质是ZIP压缩包,直接用解压工具(或代码里的ZIP处理库)打开,找到word目录下的:
comments.xml:存储所有评论的基础信息(作者、内容、创建时间等)commentsExtended.xml:存储评论的扩展状态(仅当有评论被标记为已解决时才会生成)
2. 关联评论ID与解决状态
每个评论都有唯一的w:id属性,用来关联两个XML文件中的对应条目:
- 在
comments.xml中,每个评论对应<w:comment>元素,w:id是它的唯一标识:<w:comment w:id="1" w:author="Jane Smith" w:date="2024-05-20T09:15:00Z"> <w:p> <w:r> <w:t>建议调整这段文字的语序</w:t> </w:r> </w:p> </w:comment> - 在
commentsExtended.xml中,<w:commentEx>元素的w:id与上述ID一一对应,其中<w:resolved>元素的w:val属性就是状态值:<w:commentEx w:id="1"> <w:resolved w:val="true"/> <!-- true=已解决,false=未解决 --> <w:dateResolved w:date="2024-05-21T11:30:00Z"/> </w:commentEx>
3. 处理边界情况
- 如果不存在
commentsExtended.xml,说明所有评论都处于未解决状态。 - 如果某个评论的ID在
commentsExtended.xml中没有对应条目,同样默认是未解决状态。
4. 解析逻辑示例(伪代码)
# 伪代码示例,实际需用XML解析库(如lxml、xml.etree) comments = {} # 解析comments.xml,存储评论ID和基础信息 with open("word/comments.xml", "r") as f: tree = ET.parse(f) ns = {"w": "http://schemas.openxmlformats.org/wordprocessingml/2006/main"} for comment in tree.findall(".//w:comment", namespaces=ns): comment_id = comment.get(f"{{{ns['w']}}}id") content = comment.find(".//w:t", namespaces=ns).text comments[comment_id] = {"content": content, "resolved": False} # 解析commentsExtended.xml,更新解决状态 try: with open("word/commentsExtended.xml", "r") as f: tree = ET.parse(f) ns = {"w": "http://schemas.openxmlformats.org/wordprocessingml/2006/main"} for comment_ex in tree.findall(".//w:commentEx", namespaces=ns): comment_id = comment_ex.get(f"{{{ns['w']}}}id") resolved_val = comment_ex.find(".//w:resolved", namespaces=ns).get(f"{{{ns['w']}}}val") if comment_id in comments: comments[comment_id]["resolved"] = (resolved_val == "true") except FileNotFoundError: # 没有扩展文件,所有评论保持未解决状态 pass
内容的提问来源于stack exchange,提问作者Purushotham Parthy
相关产品推荐
相关产品推荐

