You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python multiprocessing结合pytesseract时触发TypeError错误求助

解决multiprocessing结合pytesseract/pdf2image的pickle序列化错误

我来帮你分析一下这个问题,你遇到的TypeError: __init__() takes 1 positional argument but 2 were given错误本质是多进程间传递PIL Image对象时的pickle序列化/反序列化失败,再加上代码里有个小拼写错误,咱们一步步解决:

错误根源解析

这个错误出现在多进程池处理结果的阶段(日志里的_handle_results),因为你直接把pdf2image.convert_from_path返回的PIL Image对象传给了子进程。这类复杂对象在跨进程传递时,pickle(Python的序列化工具)无法正确完成反序列化,就会抛出这个看似奇怪的初始化参数错误。另外你的代码里还有个拼写错误:['orientaton']少了一个i,应该是['orientation'],这个也会导致后续提取方向信息失败。

修复步骤

1. 先修正拼写错误

把extract函数里的键名修正:

return pytesseract.image_to_osd(img, output_type=pytesseract.Output.DICT)['orientation']

2. 避免跨进程传递PIL Image对象

更可靠的做法是:利用你已经设置的output_folder保存图像文件,然后在子进程中通过文件路径加载图像,而不是直接传递Image对象。修改后的完整代码如下:

import pytesseract
import pdf2image
import multiprocessing
from PIL import Image
import os

def extract(img_path, page_num):
    print(page_num)
    # 子进程中独立加载图像
    img = Image.open(img_path)
    return pytesseract.image_to_osd(img, output_type=pytesseract.Output.DICT)['orientation']

if __name__ == "__main__":
    pdf_path = r"C:/Users/erik7/Documents/Late Scans for Testing/scans_template2.pdf"
    output_fmt = 'jpeg'
    img_dpi = 300
    pop_path = r"C:\Users\erik7\Downloads\poppler-0.90.1\bin"
    pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract.exe"
    converted_path = r"C:\Users\erik7\Downloads\converted_images"
    
    # 将PDF转换为图像文件,自动保存到converted_path
    pdf2image.convert_from_path(
        pdf_path=pdf_path, 
        fmt=output_fmt, 
        dpi=img_dpi, 
        poppler_path=pop_path, 
        output_folder=converted_path, 
        grayscale=True, 
        thread_count=2
    )
    
    # 整理图像文件路径,按页码排序
    img_files = sorted(
        [os.path.join(converted_path, f) for f in os.listdir(converted_path) 
         if f.lower().endswith(output_fmt)]
    )
    # 生成迭代对象:(图像路径, 页码)
    iterable = [[img_path, page_num] for page_num, img_path in enumerate(img_files)]
    
    results = []
    # 使用with语句管理进程池,自动处理关闭/回收
    with multiprocessing.Pool() as p:
        r = p.starmap(extract, iterable)
        results.append(r)
    
    print("\n**PROCESS COMPLETED SUCCESSFULLY")
    print("页面方向检测结果:", results)

3. 额外优化建议

  • 使用with multiprocessing.Pool()替代手动调用close(),这样进程池会自动在代码块结束后关闭并回收资源,更安全简洁
  • 如果不需要保留转换后的临时图像,可以在处理完成后添加代码删除这些文件:
    for img_path in img_files:
        os.remove(img_path)
    

内容的提问来源于stack exchange,提问作者erik7970

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 23:33:13