如何用Python提取PPT图片?含组合图提取及顺序问题求解
问题解决思路与代码优化
核心问题分析
- 无法提取组合形式的图片:当前代码仅判断
isinstance(shape, Picture),但PPT中的组合图属于GroupShape类型,不会被识别 - 图片提取顺序混乱:
slide.shapes的遍历顺序是PPT内部的存储顺序,并非视觉上的从上到下、从左到右 - PPT转PDF的可行性:转PDF后组合图会被渲染为单张图片,提取时无需处理组合,但PDF图片提取同样需要处理顺序,且会丢失原图片的矢量属性(若原图片为矢量图)
优化方案
1. 处理组合形状
如果需要提取组合内的单张图片,可递归遍历GroupShape的子形状;如果需要将整个组合作为一张图片导出,可借助Office COM组件(Windows环境)实现:
递归提取组合内图片
修改遍历逻辑,增加对组合形状的递归处理:
from pptx.shapes.group import GroupShape def traverse_shapes(shapes, doc, index): # 按视觉顺序排序:先按垂直位置top,再按水平位置left sorted_shapes = sorted(shapes, key=lambda s: (s.top, s.left)) for shape in sorted_shapes: if shape.has_text_frame: text_frame = shape.text_frame for paragraph in text_frame.paragraphs: cleaned_text = clean_text_for_xml(paragraph.text) doc.add_paragraph(cleaned_text) elif shape.has_table: process_table(shape.table, doc) elif isinstance(shape, Picture): # 原图片处理逻辑 temp_img_path = f'temp_img_{index}.jpg' with open(temp_img_path, 'wb') as f: f.write(shape.image.blob) doc.add_picture(temp_img_path, width=Inches(3.5)) index += 1 print(f'Image saved as: {temp_img_path}') os.remove(temp_img_path) elif isinstance(shape, GroupShape): # 递归遍历组合内的子形状 index = traverse_shapes(shape.shapes, doc, index) return index
导出组合形状为单张图片(Windows环境)
若需将整个组合作为一张图片导出,可调用PowerPoint的原生导出功能:
import win32com.client as win32 def export_group_as_image(group_shape, slide, save_path): slide.Shapes(group_shape.name).Select() win32.gencache.EnsureDispatch('PowerPoint.Application').ActiveWindow.Selection.Export( save_path, 'JPG', 1920, 1080 )
2. 修正图片提取顺序
通过形状的top(垂直坐标)和left(水平坐标)属性排序,实现视觉上的从上到下、从左到右顺序,如上述traverse_shapes函数中的排序逻辑。
3. PPT转PDF方案评估
- 优势:组合图会被自动渲染为单张图片,无需处理组合逻辑
- 劣势:
- 转PDF会丢失原图片的可编辑性,矢量图会转为位图
- PDF提取图片仍需处理顺序问题,且可能存在质量损耗
- 额外依赖PDF处理库(如
PyPDF2、pdfplumber)
完整优化后代码
from pptx import Presentation from docx import Document from pptx.shapes.picture import Picture from pptx.shapes.group import GroupShape from docx.shared import Inches import re import os def ppt_to_docx(ppt_path, docx_path): ppt = Presentation(ppt_path) doc = Document() index = 1 for slide in ppt.slides: # 按视觉顺序遍历幻灯片形状 index = traverse_shapes(slide.shapes, doc, index) doc.save(docx_path) def traverse_shapes(shapes, doc, index): # 按从上到下、从左到右排序形状 sorted_shapes = sorted(shapes, key=lambda s: (s.top, s.left)) for shape in sorted_shapes: if shape.has_text_frame: text_frame = shape.text_frame for paragraph in text_frame.paragraphs: cleaned_text = clean_text_for_xml(paragraph.text) doc.add_paragraph(cleaned_text) elif shape.has_table: process_table(shape.table, doc) elif isinstance(shape, Picture): temp_img_path = f'temp_img_{index}.jpg' with open(temp_img_path, 'wb') as f: f.write(shape.image.blob) doc.add_picture(temp_img_path, width=Inches(3.5)) index += 1 print(f'Image saved as: {temp_img_path}') os.remove(temp_img_path) elif isinstance(shape, GroupShape): # 递归处理组合内的形状 index = traverse_shapes(shape.shapes, doc, index) return index def clean_text_for_xml(text): texts = re.sub(u"[\\x00-\\x08\\x0b\\x0e-\\x1f\\x7f]", "", text) texts = re.sub("\f", "", texts) return texts def process_table(table, doc): doc_table = doc.add_table(rows=len(table.rows), cols=len(table.columns)) doc_table.style = 'Light Grid' for ppt_row, row in enumerate(table.rows): for ppt_col, cell in enumerate(row.cells): cell_text = clean_text_for_xml(cell.text) doc_table.cell(ppt_row, ppt_col).text = cell_text ppt_path = 'static/测试报告.pptx' docx_path = 'static/demo1.docx' ppt_to_docx(ppt_path, docx_path)
内容的提问来源于stack exchange,提问作者user23640279
相关产品推荐
相关产品推荐

