You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取PPT图片?含组合图提取及顺序问题求解

问题解决思路与代码优化

核心问题分析

  • 无法提取组合形式的图片:当前代码仅判断isinstance(shape, Picture),但PPT中的组合图属于GroupShape类型,不会被识别
  • 图片提取顺序混乱:slide.shapes的遍历顺序是PPT内部的存储顺序,并非视觉上的从上到下、从左到右
  • PPT转PDF的可行性:转PDF后组合图会被渲染为单张图片,提取时无需处理组合,但PDF图片提取同样需要处理顺序,且会丢失原图片的矢量属性(若原图片为矢量图)

优化方案

1. 处理组合形状

如果需要提取组合内的单张图片,可递归遍历GroupShape的子形状;如果需要将整个组合作为一张图片导出,可借助Office COM组件(Windows环境)实现:

递归提取组合内图片

修改遍历逻辑,增加对组合形状的递归处理:

from pptx.shapes.group import GroupShape

def traverse_shapes(shapes, doc, index):
    # 按视觉顺序排序:先按垂直位置top,再按水平位置left
    sorted_shapes = sorted(shapes, key=lambda s: (s.top, s.left))
    for shape in sorted_shapes:
        if shape.has_text_frame:
            text_frame = shape.text_frame
            for paragraph in text_frame.paragraphs:
                cleaned_text = clean_text_for_xml(paragraph.text)
                doc.add_paragraph(cleaned_text)
        elif shape.has_table:
            process_table(shape.table, doc)
        elif isinstance(shape, Picture):
            # 原图片处理逻辑
            temp_img_path = f'temp_img_{index}.jpg'
            with open(temp_img_path, 'wb') as f:
                f.write(shape.image.blob)
            doc.add_picture(temp_img_path, width=Inches(3.5))
            index += 1
            print(f'Image saved as: {temp_img_path}')
            os.remove(temp_img_path)
        elif isinstance(shape, GroupShape):
            # 递归遍历组合内的子形状
            index = traverse_shapes(shape.shapes, doc, index)
    return index

导出组合形状为单张图片(Windows环境)

若需将整个组合作为一张图片导出,可调用PowerPoint的原生导出功能:

import win32com.client as win32

def export_group_as_image(group_shape, slide, save_path):
    slide.Shapes(group_shape.name).Select()
    win32.gencache.EnsureDispatch('PowerPoint.Application').ActiveWindow.Selection.Export(
        save_path, 'JPG', 1920, 1080
    )

2. 修正图片提取顺序

通过形状的top(垂直坐标)和left(水平坐标)属性排序,实现视觉上的从上到下、从左到右顺序,如上述traverse_shapes函数中的排序逻辑。

3. PPT转PDF方案评估

  • 优势:组合图会被自动渲染为单张图片,无需处理组合逻辑
  • 劣势:
    • 转PDF会丢失原图片的可编辑性,矢量图会转为位图
    • PDF提取图片仍需处理顺序问题,且可能存在质量损耗
    • 额外依赖PDF处理库(如PyPDF2、pdfplumber)

完整优化后代码

from pptx import Presentation
from docx import Document
from pptx.shapes.picture import Picture
from pptx.shapes.group import GroupShape
from docx.shared import Inches
import re
import os

def ppt_to_docx(ppt_path, docx_path):
    ppt = Presentation(ppt_path)
    doc = Document()

    index = 1
    for slide in ppt.slides:
        # 按视觉顺序遍历幻灯片形状
        index = traverse_shapes(slide.shapes, doc, index)

    doc.save(docx_path)

def traverse_shapes(shapes, doc, index):
    # 按从上到下、从左到右排序形状
    sorted_shapes = sorted(shapes, key=lambda s: (s.top, s.left))
    for shape in sorted_shapes:
        if shape.has_text_frame:
            text_frame = shape.text_frame
            for paragraph in text_frame.paragraphs:
                cleaned_text = clean_text_for_xml(paragraph.text)
                doc.add_paragraph(cleaned_text)
        elif shape.has_table:
            process_table(shape.table, doc)
        elif isinstance(shape, Picture):
            temp_img_path = f'temp_img_{index}.jpg'
            with open(temp_img_path, 'wb') as f:
                f.write(shape.image.blob)
            doc.add_picture(temp_img_path, width=Inches(3.5))
            index += 1
            print(f'Image saved as: {temp_img_path}')
            os.remove(temp_img_path)
        elif isinstance(shape, GroupShape):
            # 递归处理组合内的形状
            index = traverse_shapes(shape.shapes, doc, index)
    return index

def clean_text_for_xml(text):
    texts = re.sub(u"[\\x00-\\x08\\x0b\\x0e-\\x1f\\x7f]", "", text)
    texts = re.sub("\f", "", texts)
    return texts

def process_table(table, doc):
    doc_table = doc.add_table(rows=len(table.rows), cols=len(table.columns))
    doc_table.style = 'Light Grid'

    for ppt_row, row in enumerate(table.rows):
        for ppt_col, cell in enumerate(row.cells):
            cell_text = clean_text_for_xml(cell.text)
            doc_table.cell(ppt_row, ppt_col).text = cell_text

ppt_path = 'static/测试报告.pptx'
docx_path = 'static/demo1.docx'
ppt_to_docx(ppt_path, docx_path)

内容的提问来源于stack exchange,提问作者user23640279

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 15:13:13