使用Pytesseract分类图片时部分文件出现PermissionError求助
解决Windows下pytesseract处理图片时的PermissionError(文件被占用)
问题背景
在Windows 10系统使用Tesseract 5.3.3 + pytesseract进行图片分类时,部分图片触发PermissionError:删除临时文件时提示文件被其他进程占用,错误信息如下:
PermissionError: [WinError 32] The process cannot access the file because it is being used by another process: 'C:\\Users\\V34FD~1.MED\\AppData\\Local\\Temp\\tmpzg2vxncf.png'
已尝试提前删除临时文件夹,但问题未解决。
核心原因
- pytesseract处理numpy格式图片时,会自动生成临时文件供Tesseract引擎读取,若Tesseract子进程未完全退出或文件句柄未及时释放,会导致后续删除操作失败。
- Windows系统的文件锁机制较严格,进程未释放句柄时无法删除文件。
解决方案
1. 给临时文件删除操作添加重试逻辑
在删除临时文件的代码块中,增加重试机制,多次尝试删除,给系统足够时间释放文件句柄:
import time # 替换原删除临时文件的代码 temp_image_path = "你的临时文件路径" max_retries = 3 retry_delay = 0.5 for attempt in range(max_retries): try: if os.path.exists(temp_image_path): os.remove(temp_image_path) break except PermissionError: if attempt < max_retries - 1: time.sleep(retry_delay) else: print(f"Failed to delete {temp_image_path} after {max_retries} attempts")
2. 修改pytesseract调用方式,减少临时文件残留
当处理numpy数组格式的图片时,pytesseract会自动创建临时文件。可以尝试直接传入图片文件路径(如果原始图片未经过预处理),或者手动控制临时文件的生成与清理:
# 示例:直接传入图片路径而非numpy数组(适用于无需预处理的场景) extracted_text = pytesseract.image_to_string(image_path) # 替代传入img_org
如果必须使用预处理后的numpy数组,可以手动指定临时文件路径,并在处理后立即清理:
import tempfile def classify_axial_img(classification_dict, includes_dict, platform, analysis_case, plate_number, image_path, manual_classification=None): try: img_org = cv2.imread(image_path) # 尝试原图片 extracted_text = pytesseract.image_to_string(img_org) check = False for col in includes_dict.keys(): if includes_dict[col]['keyword'] in extracted_text: classification_dict[platform][analysis_case][col][0].append(plate_number) classification_dict[platform][analysis_case][col][1].append(image_path) return classification_dict # 预处理尝试 if not check: for preprocessing in preprocessings: preprocessed_img = apply_preprocessing(preprocessing, img_org) # 手动创建临时文件 with tempfile.NamedTemporaryFile(suffix='.png', delete=False) as temp_file: cv2.imwrite(temp_file.name, preprocessed_img) temp_path = temp_file.name try: extracted_text = pytesseract.image_to_string(temp_path) for col in includes_dict.keys(): if includes_dict[col]['keyword'] in extracted_text: classification_dict[platform][analysis_case][col][0].append(plate_number) classification_dict[platform][analysis_case][col][1].append(image_path) return classification_dict finally: # 立即清理临时文件 max_retries = 3 retry_delay = 0.5 for attempt in range(max_retries): try: if os.path.exists(temp_path): os.remove(temp_path) break except PermissionError: if attempt < max_retries - 1: time.sleep(retry_delay) return classification_dict except Exception as e: print(e) print('Could not classify image {}.'.format(image_path)) return classification_dict
3. 确保Tesseract子进程完全退出
强制pytesseract使用同步方式执行,设置超时避免进程挂起,确保子进程结束后再继续后续操作:
# 指定Tesseract路径(根据实际安装路径调整) pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' # 设置超时时间,确保进程及时退出 extracted_text = pytesseract.image_to_string(img_org, timeout=10)
4. 排查第三方进程干扰
临时文件被占用也可能是杀毒软件、文件预览工具等后台进程导致的,可以临时关闭这类软件测试,或者将临时文件目录加入杀毒软件白名单。
内容的提问来源于stack exchange,提问作者user24023817
相关产品推荐
相关产品推荐

