You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用unstructured分割PDF时遇NameError及ocr_languages错误的解决

用unstructured分割含文本、表格、图片的PDF问题及解决方法

我尝试用Python的unstructured、unstructured[pdf]库分割包含文本、表格、图片的PDF文件,遇到了无法解决的错误。

操作代码

from unstructured.partition.pdf import partition_pdf

path = '/content/'
file_name = 'ABCABC.pdf'

raw_pdf_elements = partition_pdf(
    filename=path + file_name,
    extract_images_in_pdf=True,
    infer_table_structure=True,
    chunking_strategy="by_title",
    max_characters=4000,
    new_after_n_chars=3800,
    combine_text_under_n_chars=2000,
    image_output_dir_path=path
)

首次报错信息

NameError: name 'sort_page_elements' is not defined

环境信息

  • Python:3.9.19
  • unstructured:0.15.1(也曾尝试0.7.12和0.12.2版本)
  • 操作系统:Ubuntu 20.04

无效尝试及新错误

尝试过安装numpy(1.26.4)和opencv-python(4.10.0.84)的解决方案,但无效。重新激活环境后出现新错误:

get_model() got an unexpected keyword argument 'ocr_languages'

最终解决步骤

  • 为partition_pdf函数添加额外参数,例如:languages=["vie"]
  • 下载对应语言的Tesseract训练数据
  • 设置环境变量:export TESSDATA_PREFIX=你的tessdata文件夹路径(即训练数据所在的文件夹)

内容的提问来源于stack exchange,提问作者happy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 20:26:05