You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python解析pptx中的XML文件?pptx转h5p项目技术求助

解决PPTX XML解析(ElementTree)及转JSON问题

核心问题分析

你遇到的解析失败大概率是因为PPTX的XML文件带有命名空间,直接使用p:sld这类带前缀的节点名无法匹配元素,ElementTree需要显式绑定命名空间才能正确查找节点。

步骤与代码实现

1. 导入依赖模块

import xml.etree.ElementTree as ET
import json

2. 加载XML并绑定命名空间

PPTX的XML中常用两个命名空间:

  • p:演示文稿主内容命名空间
  • a:绘图内容命名空间
# 加载PPTX解压后的slide XML文件(例如ppt/slides/slide1.xml)
tree = ET.parse("slide1.xml")
root = tree.getroot()

# 定义命名空间字典,用于ElementTree查找节点
namespaces = {
    "p": "http://schemas.openxmlformats.org/presentationml/2006/main",
    "a": "http://schemas.openxmlformats.org/drawingml/2006/main"
}

3. 提取目标节点文本并转JSON

根据你给出的节点路径,我们需要遍历形状(p:sp)、文本框(p:txBody)、段落(a:p)、文本段(a:r),最终提取文本节点(a:t)的值:

extracted_content = []

# 遍历所有形状节点
for shape in root.findall(".//p:sp", namespaces=namespaces):
    # 获取当前形状的文本框
    text_body = shape.find("p:txBody", namespaces=namespaces)
    if not text_body:
        continue
    
    # 递归查找所有段落节点(包含嵌套的a:p)
    for paragraph in text_body.findall(".//a:p", namespaces=namespaces):
        # 遍历段落内的所有文本段
        for run in paragraph.findall("a:r", namespaces=namespaces):
            text_node = run.find("a:t", namespaces=namespaces)
            # 仅提取非空文本
            if text_node is not None and text_node.text:
                extracted_content.append(text_node.text)

# 将提取结果写入JSON文件
with open("slide_texts.json", "w", encoding="utf-8") as f:
    json.dump({"extracted_texts": extracted_content}, f, ensure_ascii=False, indent=2)

关键说明

  • 使用.//前缀可以递归查找所有层级的目标节点,适配你路径中嵌套的a:p节点
  • 命名空间是PPTX XML解析的核心,必须通过namespaces参数传递给find/findall方法
  • 如果需要保留更完整的结构(比如每个形状对应一组文本),可以调整代码构建嵌套字典

内容的提问来源于stack exchange,提问作者Tuấn Anh Trần

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 13:45:53