使用Paddle OCR实现并行处理失败问题求助
PaddleOCR多线程并行处理报错,改用多进程解决
问题现象
- 单线程(
num_threads=1)调用predictx_parallel方法时,OCR识别正常运行; - 多线程(
num_threads>1)调用时,触发两类框架层面错误:
错误信息
UnimplementedError: Currently, only can set dims from DenseTensor or SelectedRows. (at /paddle/paddle/fluid/framework/infershape_utils.cc:314) [operator < fused_conv2d > error]
NotFoundError: Variable Id 29797 is not registered. [Hint: Expected it != Instance().id_to_type_map_.end(), but received it == Instance().id_to_type_map_.end().] (at /paddle/paddle/fluid/framework/var_type_traits.cc:103) [operator < fused_conv2d > error]
原实现代码
def predictx_parallel(input_images: List[Image], ocr_params: OcrParams, num_threads: int) -> Tuple[List[Dict], List[Image]]: def ocr_image(image): image_array = np.array(image) # type(image) : <class 'PIL.Image.Image'> # type(image_array) : <class 'numpy.ndarray'> results = ocr.ocr(image_array) # [[[[[381.0, 285.0], [537.0, 285.0], [537.0, 333.0], [381.0, 333.0]], ('PAGE1A', 0.9997838139533997)], [[[388.0, 371.0], [530.0, 371.0], [530.0, 419.0], [388.0, 419.0]], ('PAGE1B', 0.998117983341217)]]] return results ocr = get_ocr_obj(params=ocr_params) # <class 'paddleocr.paddleocr.PaddleOCR'> with ThreadPoolExecutor(max_workers=num_threads) as executor: results = list(executor.map(ocr_image, input_images))
使用环境
paddlepaddle==2.6.0paddleocr==2.7.0.3python==3.9.12
问题原因
单个PaddleOCR实例无法在多线程环境下安全复用,跨线程共享会导致Paddle内部张量、变量管理出现线程安全问题,触发框架报错。
解决方案
改用多进程实现并行处理,在每个进程内独立创建新的PaddleOCR实例,避免跨线程共享OCR对象。
修复后代码
pipeline.py
def predictx_parallel_processes(input_images, num_processes): with Pool(processes=num_processes) as pool: pool.map(ocr_image_x, input_images)
ocr_processing.py
def ocr_image_x(image): process_pid = os.getpid() logger.info(f"Process PID: {process_pid}") ocr = PaddleOCR() # 每个进程创建独立的OCR实例 image_array = np.array(image) results = ocr.ocr(image_array) logger.info(results)
内容的提问来源于stack exchange,提问作者Soumya
相关产品推荐
相关产品推荐

