如何用Python解析pptx中的XML文件?pptx转h5p项目技术求助
解决PPTX XML解析(ElementTree)及转JSON问题
核心问题分析
你遇到的解析失败大概率是因为PPTX的XML文件带有命名空间,直接使用p:sld这类带前缀的节点名无法匹配元素,ElementTree需要显式绑定命名空间才能正确查找节点。
步骤与代码实现
1. 导入依赖模块
import xml.etree.ElementTree as ET import json
2. 加载XML并绑定命名空间
PPTX的XML中常用两个命名空间:
p:演示文稿主内容命名空间a:绘图内容命名空间
# 加载PPTX解压后的slide XML文件(例如ppt/slides/slide1.xml) tree = ET.parse("slide1.xml") root = tree.getroot() # 定义命名空间字典,用于ElementTree查找节点 namespaces = { "p": "http://schemas.openxmlformats.org/presentationml/2006/main", "a": "http://schemas.openxmlformats.org/drawingml/2006/main" }
3. 提取目标节点文本并转JSON
根据你给出的节点路径,我们需要遍历形状(p:sp)、文本框(p:txBody)、段落(a:p)、文本段(a:r),最终提取文本节点(a:t)的值:
extracted_content = [] # 遍历所有形状节点 for shape in root.findall(".//p:sp", namespaces=namespaces): # 获取当前形状的文本框 text_body = shape.find("p:txBody", namespaces=namespaces) if not text_body: continue # 递归查找所有段落节点(包含嵌套的a:p) for paragraph in text_body.findall(".//a:p", namespaces=namespaces): # 遍历段落内的所有文本段 for run in paragraph.findall("a:r", namespaces=namespaces): text_node = run.find("a:t", namespaces=namespaces) # 仅提取非空文本 if text_node is not None and text_node.text: extracted_content.append(text_node.text) # 将提取结果写入JSON文件 with open("slide_texts.json", "w", encoding="utf-8") as f: json.dump({"extracted_texts": extracted_content}, f, ensure_ascii=False, indent=2)
关键说明
- 使用
.//前缀可以递归查找所有层级的目标节点,适配你路径中嵌套的a:p节点 - 命名空间是PPTX XML解析的核心,必须通过
namespaces参数传递给find/findall方法 - 如果需要保留更完整的结构(比如每个形状对应一组文本),可以调整代码构建嵌套字典
内容的提问来源于stack exchange,提问作者Tuấn Anh Trần
相关产品推荐
相关产品推荐

