You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ProcessPoolExecutor批量处理PDF OCR无响应问题求助

问题诊断与修复方案

核心问题

你的代码存在两个关键错误:

  1. ppe.map() 与 as_completed() 不兼容:map() 返回的是任务结果的迭代器,而非 Future 对象,as_completed() 无法处理这种类型,导致后续循环完全不执行,自然看不到完成提示。
  2. 多进程输出缓冲:子进程的 print 内容会被系统缓冲,无法实时同步到主进程控制台,这就是你看不到处理进度的原因。

修复后的完整代码

# scrap text from pdf's and store content in files for nlp analysis
# tried to use both camelot and tabular and both packages could not scrap the required table contents
# this script implements ocr using tesseract  

from glob import glob 
import pytesseract
from concurrent.futures import ProcessPoolExecutor
from concurrent.futures import as_completed
from pdf2image import convert_from_path as pdf2img
import pathlib as pl
import multiprocessing as mpc

def ProcessPDF(par_FilePath):
    lstImages = pdf2img(par_FilePath)
    intImgs = len(lstImages)
    strOCRd = ''
    file_name = pl.Path(par_FilePath).name
    for it, im in enumerate(lstImages):
        npg = '='*50+f'Pg:{it+1}'+'='*50+'\n' #end each page
        pgText = pytesseract.image_to_string(im) #perform ocr
        strOCRd += pgText + '\n' + npg # add to string
        # 强制刷新输出,避免子进程缓冲
        print(f'Processing: {file_name} : {int(it/intImgs*100)}%', flush=True)
    fStem  = pl.Path(par_FilePath).stem
    fDir =  str(pl.Path(par_FilePath).parent)+'/'

    with open(fDir + fStem + '.txt', 'w') as fobj: #save file
        fobj.write(strOCRd)
    return f'Completed: {file_name}'

if __name__ == '__main__':
    strFolderPDF = r'/home/*****/proj/rfp_model/pdfFiles/'
    lstFiles = glob(strFolderPDF+'*.pdf')
    numFiles = len(lstFiles)
    numCPUs = mpc.cpu_count()
    print(f'Starting pool executor, processing {numFiles} files with {numCPUs} workers.')
    with ProcessPoolExecutor(max_workers=numCPUs) as ppe:
        # 改用submit提交任务,生成Future对象列表
        futures = [ppe.submit(ProcessPDF, file_path) for file_path in lstFiles]
    
        # 遍历完成的Future对象,打印结果
        for future in as_completed(futures):
            print(future.result())

    #this works 
    #for ipath in lstFiles:
    #    ProcessPDF(ipath)

修复细节说明

  1. 替换 map() 为 submit():submit() 会为每个任务返回独立的 Future 对象,as_completed() 可以正确识别并处理这些对象,任务完成后立即输出结果。
  2. 添加 flush=True 到print语句:强制子进程的输出直接刷新到控制台,解决多进程环境下的输出缓冲问题,让进度提示实时显示。
  3. 提前提取文件名:避免在循环中重复调用路径解析方法,小幅优化性能。

额外建议

  • 确保系统已安装 poppler-utils(pdf2image 的依赖),Ubuntu下可通过 sudo apt install poppler-utils 安装。
  • 若处理大体积PDF,可给 pdf2img 添加 dpi=150 参数降低图片分辨率,减少内存占用:pdf2img(par_FilePath, dpi=150)。

内容的提问来源于stack exchange,提问作者rrrrrrs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 23:27:03