You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python OCR图片转文本后逐元素检测重复问题求助

Python OCR图片转文本重复检测问题排查与修复

问题场景

正在进行Python OCR图片转文本开发,需要通过逐元素检测识别重复文本并精准定位重复项,但当前代码运行时listOfElems始终为空,无法实现重复检测功能。

需求示例

例如列表listA = [1,2,3,4,4,5,6],在添加元素时能提示4是重复项。

原代码

from PIL import Image
import pytesseract
import cv2 
import numpy as np
from os import listdir
from os.path import isfile, join

mypath = "/home/DC_ton/desktop/test_11_8/output02"
onlyfiles = [f for f in listdir(mypath) if isfile(join(mypath, f))]
print(onlyfiles)

i = 1
listOfElems = []
Number_of_onlyfiles = len(onlyfiles)
while i < Number_of_onlyfiles :
    each_file_path = '/home/DC_ton/desktop/test_11_8/output02/'+ onlyfiles[i]
    image = Image.open(each_file_path)
    text = pytesseract.image_to_string(image, lang='eng')
    print(text)
       
    for text in listOfElems:
        if text not in listOfElems:
            listOfElems.append(text)
        else:
            print("here get duplicate")
     
    i +=1
    
print(listOfElems)  

newlist = [] 
duplist = []

def checkIfDuplicates_1(listOfElems):
    ''' Check if given list contains any duplicates '''
    if len(listOfElems) == len(set(listOfElems)):
        return False
    else:
        return True
    
result = checkIfDuplicates_1(listOfElems)
if result:
    print('Yes, list contains duplicates')
else:
    print('No duplicates found in list')  


for k in listOfElems:
    if k not in newlist:
        newlist.append(k)
    else:
        duplist.append(k) 
print("List of duplicates", duplist)

原运行输出

['final_output_11.png', 'final_output_6.png', 'final_output_17.png', 'final_output_8.png', 'final_output_15.png', 'final_output_14.png', 'final_output_2.png', 'final_output_12.png', 'final_output_21.png', 'final_output_3.png', 'final_output_24.png', 'final_output_18.png', 'final_output_19.png', 'final_output_10.png', 'final_output_29.png', 'final_output_9.png', 'final_output_20.png', 'final_output_7.png', 'final_output_31.png', 'final_output_30.png', 'final_output_25.png', 'final_output_1.png', 'final_output_16.png', 'final_output_5.png', 'final_output_27.png', 'final_output_13.png', 'final_output_28.png', 'final_output_4.png', 'final_output_23.png', 'final_output_26.png', 'final_output_22.png']
CA7T4B2


CAT7T4BF


CAT4B8


CAT4BE


CAT4C4





CAT4C1


CAT4B7


CA7T4CB

 

CAT4cs


CAT4B4


CAT4BA


CAT7T4BC


CA74B9


CAT4BD


(CAT4AF


CAT4CA


[]
No duplicates found in list
List of duplicates []

问题分析与修复方案

核心问题点

  1. 循环起始索引错误:i = 1会跳过第一个文件,且初始listOfElems为空时,后续的for text in listOfElems循环根本不会执行,导致没有元素被添加到列表。
  2. 重复检测逻辑完全错误:遍历listOfElems来处理当前识别的text逻辑颠倒,应该直接判断当前识别的文本是否已存在于列表中。
  3. OCR文本未清理:识别出的文本包含大量换行、空白字符,会导致判断重复时出现误差。

修复后的代码

from PIL import Image
import pytesseract
from os import listdir
from os.path import isfile, join

mypath = "/home/DC_ton/desktop/test_11_8/output02"
onlyfiles = [f for f in listdir(mypath) if isfile(join(mypath, f))]
print(onlyfiles)

listOfElems = []
# 循环从第一个文件开始,遍历所有文件
for file_name in onlyfiles:
    each_file_path = join(mypath, file_name)
    image = Image.open(each_file_path)
    text = pytesseract.image_to_string(image, lang='eng').strip()  # 清理空白字符
    print(text)
    
    # 跳过空文本,避免无效数据
    if not text:
        continue
    
    # 核心重复检测逻辑
    if text in listOfElems:
        print(f"检测到重复项:{text}")
    else:
        listOfElems.append(text)

print("\n所有识别到的文本:", listOfElems)  

# 统计重复项
seen = set()
duplist = []
for item in listOfElems:
    if item in seen:
        duplist.append(item)
    else:
        seen.add(item)

if duplist:
    print('列表包含重复项')
    print("重复项列表:", duplist)
else:
    print('列表中未发现重复项')

关键修改说明

  • 把while循环改为更简洁的for循环遍历文件列表,避免索引错误。
  • 用strip()清理OCR识别文本中的换行、空格等空白字符,确保文本匹配准确。
  • 直接判断当前识别的文本是否在listOfElems中,逻辑清晰,符合需求。
  • 添加空文本过滤,避免无效的空字符串加入列表。
  • 简化重复项统计逻辑,提高效率。

内容的提问来源于stack exchange,提问作者user20396381

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 22:20:27