Python OCR图片转文本后逐元素检测重复问题求助
Python OCR图片转文本重复检测问题排查与修复
问题场景
正在进行Python OCR图片转文本开发,需要通过逐元素检测识别重复文本并精准定位重复项,但当前代码运行时listOfElems始终为空,无法实现重复检测功能。
需求示例
例如列表listA = [1,2,3,4,4,5,6],在添加元素时能提示4是重复项。
原代码
from PIL import Image import pytesseract import cv2 import numpy as np from os import listdir from os.path import isfile, join mypath = "/home/DC_ton/desktop/test_11_8/output02" onlyfiles = [f for f in listdir(mypath) if isfile(join(mypath, f))] print(onlyfiles) i = 1 listOfElems = [] Number_of_onlyfiles = len(onlyfiles) while i < Number_of_onlyfiles : each_file_path = '/home/DC_ton/desktop/test_11_8/output02/'+ onlyfiles[i] image = Image.open(each_file_path) text = pytesseract.image_to_string(image, lang='eng') print(text) for text in listOfElems: if text not in listOfElems: listOfElems.append(text) else: print("here get duplicate") i +=1 print(listOfElems) newlist = [] duplist = [] def checkIfDuplicates_1(listOfElems): ''' Check if given list contains any duplicates ''' if len(listOfElems) == len(set(listOfElems)): return False else: return True result = checkIfDuplicates_1(listOfElems) if result: print('Yes, list contains duplicates') else: print('No duplicates found in list') for k in listOfElems: if k not in newlist: newlist.append(k) else: duplist.append(k) print("List of duplicates", duplist)
原运行输出
['final_output_11.png', 'final_output_6.png', 'final_output_17.png', 'final_output_8.png', 'final_output_15.png', 'final_output_14.png', 'final_output_2.png', 'final_output_12.png', 'final_output_21.png', 'final_output_3.png', 'final_output_24.png', 'final_output_18.png', 'final_output_19.png', 'final_output_10.png', 'final_output_29.png', 'final_output_9.png', 'final_output_20.png', 'final_output_7.png', 'final_output_31.png', 'final_output_30.png', 'final_output_25.png', 'final_output_1.png', 'final_output_16.png', 'final_output_5.png', 'final_output_27.png', 'final_output_13.png', 'final_output_28.png', 'final_output_4.png', 'final_output_23.png', 'final_output_26.png', 'final_output_22.png'] CA7T4B2 CAT7T4BF CAT4B8 CAT4BE CAT4C4 CAT4C1 CAT4B7 CA7T4CB CAT4cs CAT4B4 CAT4BA CAT7T4BC CA74B9 CAT4BD (CAT4AF CAT4CA [] No duplicates found in list List of duplicates []
问题分析与修复方案
核心问题点
- 循环起始索引错误:
i = 1会跳过第一个文件,且初始listOfElems为空时,后续的for text in listOfElems循环根本不会执行,导致没有元素被添加到列表。 - 重复检测逻辑完全错误:遍历
listOfElems来处理当前识别的text逻辑颠倒,应该直接判断当前识别的文本是否已存在于列表中。 - OCR文本未清理:识别出的文本包含大量换行、空白字符,会导致判断重复时出现误差。
修复后的代码
from PIL import Image import pytesseract from os import listdir from os.path import isfile, join mypath = "/home/DC_ton/desktop/test_11_8/output02" onlyfiles = [f for f in listdir(mypath) if isfile(join(mypath, f))] print(onlyfiles) listOfElems = [] # 循环从第一个文件开始,遍历所有文件 for file_name in onlyfiles: each_file_path = join(mypath, file_name) image = Image.open(each_file_path) text = pytesseract.image_to_string(image, lang='eng').strip() # 清理空白字符 print(text) # 跳过空文本,避免无效数据 if not text: continue # 核心重复检测逻辑 if text in listOfElems: print(f"检测到重复项:{text}") else: listOfElems.append(text) print("\n所有识别到的文本:", listOfElems) # 统计重复项 seen = set() duplist = [] for item in listOfElems: if item in seen: duplist.append(item) else: seen.add(item) if duplist: print('列表包含重复项') print("重复项列表:", duplist) else: print('列表中未发现重复项')
关键修改说明
- 把
while循环改为更简洁的for循环遍历文件列表,避免索引错误。 - 用
strip()清理OCR识别文本中的换行、空格等空白字符,确保文本匹配准确。 - 直接判断当前识别的文本是否在
listOfElems中,逻辑清晰,符合需求。 - 添加空文本过滤,避免无效的空字符串加入列表。
- 简化重复项统计逻辑,提高效率。
内容的提问来源于stack exchange,提问作者user20396381
相关产品推荐
相关产品推荐

