You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用unstructured库转换Docx为JSON时无法识别图片问题排查

问题

尝试用unstructured库把Word文档转成JSON,结果元素列表里本该有的Image类型没被识别出来,程序没报错就是不返回图片元素。相关代码、项目结构如下:

代码

from unstructured.partition.docx import partition_docx
import os
# 原代码遗漏json模块导入
import json

# 设置环境变量
os.environ['UNSTRUCTURED_API_KEY'] = "your unstructured.io api key"
os.environ['UNSTRUCTURED_API_URL'] = "https://api.unstructuredapp.io/general/v0/general"

elements = partition_docx(filename="input/test.docx")

with open("input/test.docx", "rb") as f:
    elements = partition_docx(file=f)
    elements = [element.to_dict() for element in elements]
    # 保存为JSON
    with open("output/test.json", "w") as f_json:
        json.dump(elements, f_json, indent=2)

项目结构

├── root
│   ├── input
│   └── output

测试文档里包含一段文本、一张图片、另一段文本,但图片未被检测到。

解决方法
  • API模式需显式开启图片提取:你设置了API相关环境变量,此时unstructured调用云端API处理文档。默认API不会提取图片,需要在调用partition_docx时加上strategy="hi_res"参数:

    elements = partition_docx(file=f, strategy="hi_res")
    
  • 本地处理需补全依赖:如果不想用云端API,删掉环境变量配置改用本地处理:

    1. 安装系统依赖poppler:Ubuntu用sudo apt-get install poppler-utils,Mac用brew install poppler,Windows需手动下载配置环境。
    2. 安装unstructured的docx扩展依赖:
      pip install "unstructured[docx]"
      
      完成后本地处理即可自动识别嵌入的图片元素。
  • 修复代码遗漏:原代码未导入json模块,必须补上,否则会抛出NameError。

  • 确认图片类型:确保Word文档里的图片是直接嵌入到文档内的,而非外部链接形式,unstructured无法提取链接型图片。

内容的提问来源于stack exchange,提问作者ThaNoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 00:22:39