使用Python multiprocessing结合pytesseract时触发TypeError错误求助
解决multiprocessing结合pytesseract/pdf2image的pickle序列化错误
我来帮你分析一下这个问题,你遇到的TypeError: __init__() takes 1 positional argument but 2 were given错误本质是多进程间传递PIL Image对象时的pickle序列化/反序列化失败,再加上代码里有个小拼写错误,咱们一步步解决:
错误根源解析
这个错误出现在多进程池处理结果的阶段(日志里的_handle_results),因为你直接把pdf2image.convert_from_path返回的PIL Image对象传给了子进程。这类复杂对象在跨进程传递时,pickle(Python的序列化工具)无法正确完成反序列化,就会抛出这个看似奇怪的初始化参数错误。另外你的代码里还有个拼写错误:['orientaton']少了一个i,应该是['orientation'],这个也会导致后续提取方向信息失败。
修复步骤
1. 先修正拼写错误
把extract函数里的键名修正:
return pytesseract.image_to_osd(img, output_type=pytesseract.Output.DICT)['orientation']
2. 避免跨进程传递PIL Image对象
更可靠的做法是:利用你已经设置的output_folder保存图像文件,然后在子进程中通过文件路径加载图像,而不是直接传递Image对象。修改后的完整代码如下:
import pytesseract import pdf2image import multiprocessing from PIL import Image import os def extract(img_path, page_num): print(page_num) # 子进程中独立加载图像 img = Image.open(img_path) return pytesseract.image_to_osd(img, output_type=pytesseract.Output.DICT)['orientation'] if __name__ == "__main__": pdf_path = r"C:/Users/erik7/Documents/Late Scans for Testing/scans_template2.pdf" output_fmt = 'jpeg' img_dpi = 300 pop_path = r"C:\Users\erik7\Downloads\poppler-0.90.1\bin" pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract.exe" converted_path = r"C:\Users\erik7\Downloads\converted_images" # 将PDF转换为图像文件,自动保存到converted_path pdf2image.convert_from_path( pdf_path=pdf_path, fmt=output_fmt, dpi=img_dpi, poppler_path=pop_path, output_folder=converted_path, grayscale=True, thread_count=2 ) # 整理图像文件路径,按页码排序 img_files = sorted( [os.path.join(converted_path, f) for f in os.listdir(converted_path) if f.lower().endswith(output_fmt)] ) # 生成迭代对象:(图像路径, 页码) iterable = [[img_path, page_num] for page_num, img_path in enumerate(img_files)] results = [] # 使用with语句管理进程池,自动处理关闭/回收 with multiprocessing.Pool() as p: r = p.starmap(extract, iterable) results.append(r) print("\n**PROCESS COMPLETED SUCCESSFULLY") print("页面方向检测结果:", results)
3. 额外优化建议
- 使用
with multiprocessing.Pool()替代手动调用close(),这样进程池会自动在代码块结束后关闭并回收资源,更安全简洁 - 如果不需要保留转换后的临时图像,可以在处理完成后添加代码删除这些文件:
for img_path in img_files: os.remove(img_path)
内容的提问来源于stack exchange,提问作者erik7970
相关产品推荐
相关产品推荐

