使用unstructured分割PDF时遇NameError及ocr_languages错误的解决
用unstructured分割含文本、表格、图片的PDF问题及解决方法
我尝试用Python的unstructured、unstructured[pdf]库分割包含文本、表格、图片的PDF文件,遇到了无法解决的错误。
操作代码
from unstructured.partition.pdf import partition_pdf path = '/content/' file_name = 'ABCABC.pdf' raw_pdf_elements = partition_pdf( filename=path + file_name, extract_images_in_pdf=True, infer_table_structure=True, chunking_strategy="by_title", max_characters=4000, new_after_n_chars=3800, combine_text_under_n_chars=2000, image_output_dir_path=path )
首次报错信息
NameError: name 'sort_page_elements' is not defined
环境信息
- Python:3.9.19
- unstructured:0.15.1(也曾尝试0.7.12和0.12.2版本)
- 操作系统:Ubuntu 20.04
无效尝试及新错误
尝试过安装numpy(1.26.4)和opencv-python(4.10.0.84)的解决方案,但无效。重新激活环境后出现新错误:
get_model() got an unexpected keyword argument 'ocr_languages'
最终解决步骤
- 为
partition_pdf函数添加额外参数,例如:languages=["vie"] - 下载对应语言的Tesseract训练数据
- 设置环境变量:
export TESSDATA_PREFIX=你的tessdata文件夹路径(即训练数据所在的文件夹)
内容的提问来源于stack exchange,提问作者happy
相关产品推荐
相关产品推荐

