You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取Word中的形状、Microsoft图表等非文本图形为图片?

提取Word文档中形状、图表等非文本内容的解决方案

目前你用docx2python只能提取普通图片,无法获取图表、形状这类元素,转HTML的方案也没效果,试试以下两种解决方法:

方法一:Windows环境下用Win32COM操控Word

直接调用Word原生API,精准识别并导出所有形状、图表和图片:

  1. 安装依赖包:
pip install pywin32
  1. 运行提取代码:
import win32com.client as win32
import os

def extract_non_text_elements(doc_path, output_dir):
    # 创建输出目录(不存在则自动生成)
    os.makedirs(output_dir, exist_ok=True)
    
    # 启动Word后台进程,不显示界面
    word = win32.gencache.EnsureDispatch('Word.Application')
    word.Visible = False
    
    try:
        doc = word.Documents.Open(os.path.abspath(doc_path))
        file_count = 1
        
        # 导出所有形状(包括自选图形、文本框等)
        for shape in doc.Shapes:
            shape.Export(
                os.path.join(output_dir, f"shape_{file_count}.png"),
                Filter=win32.constants.wdExportFormatPNG
            )
            file_count += 1
        
        # 导出所有内嵌图表
        for inline_shape in doc.InlineShapes:
            if inline_shape.HasChart:
                inline_shape.Chart.Export(
                    os.path.join(output_dir, f"chart_{file_count}.png"),
                    Filter="PNG"
                )
                file_count += 1
        
        # 导出嵌入式图片
        for inline_shape in doc.InlineShapes:
            if inline_shape.Type == win32.constants.wdInlineShapePicture:
                inline_shape.Export(
                    os.path.join(output_dir, f"image_{file_count}.png"),
                    Filter=win32.constants.wdExportFormatPNG
                )
                file_count += 1
                
        print(f"共提取{file_count-1}个非文本元素,保存至{output_dir}")
    except Exception as e:
        print(f"提取失败:{str(e)}")
    finally:
        # 关闭文档和Word进程,不保存修改
        doc.Close(SaveChanges=False)
        word.Quit()

# 替换为你的Word文档路径和输出目录
extract_non_text_elements("someWord.docx", r"c:\filename\")

该方法能分别区分形状、图表和图片,导出格式可自行调整(如JPG、BMP)。

方法二:跨平台用LibreOffice命令行

适用于Windows、macOS、Linux环境,无需依赖Word:

  1. 先安装LibreOffice
  2. 执行命令行转换:
# Windows系统
soffice --headless --convert-to html:"HTML (StarWriter)" --outdir c:\filename\ someWord.docx

# macOS/Linux系统
libreoffice --headless --convert-to html:"HTML (StarWriter)" --outdir ./filename/ someWord.docx

转换完成后,输出目录会生成一个HTML文件和同名资源文件夹,里面包含所有导出的非文本元素对应的图片文件。


内容的提问来源于stack exchange,提问作者ahmad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 10:26:03