You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用glob()加载图片目录至标签检测工具时遇属性错误求助

解决你的图片加载与标签检测问题

Hey there! Let's figure out why your code is throwing that Attribute Error and get you set up with an easier way to load images recursively. I'll break this down simply since you're new to glob and programming.

问题出在哪?

Looking at your code, the main issue is that you're storing all loaded images in a list called image, but then trying to use functions like imutils.resize() and pytesseract.image_to_string() directly on this list. These functions only work on single image objects (not lists of images), which is why you're getting an Attribute Error.

Also, your current glob pattern only grabs .JPG files in one folder—if you have subfolders with images, it won't find them. Let's fix both issues.

修复后的完整代码

Here's a revised version of your code that handles multiple images correctly, plus a simpler recursive image loading method:

import re
import cv2
import glob
import imutils
import pytesseract
import pandas as pd

# --------------------------
# 简便的递归加载图片方法
# 用glob的recursive参数查找所有子文件夹里的JPG/PNG(可扩展格式)
# --------------------------
# 匹配所有JPG/JPEG/PNG文件,包括子文件夹
image_paths = glob.glob('Image Samples/CE Test Pics/**/*.JPG', recursive=True)
# 可以加上其他格式,比如:
# image_paths += glob.glob('Image Samples/CE Test Pics/**/*.jpg', recursive=True)
# image_paths += glob.glob('Image Samples/CE Test Pics/**/*.png', recursive=True)

# 准备存储结果的列表
results = []

for img_path in image_paths:
    # 加载单张图片
    image = cv2.imread(img_path)
    if image is None:
        print(f"无法加载图片: {img_path}")
        continue  # 跳过加载失败的图片
    
    # 调整单张图片大小
    resized_image = imutils.resize(image, width=500)
    
    # 识别文字
    text = pytesseract.image_to_string(resized_image, lang='eng', config='--psm 11')
    
    # 提取合同编号
    contract_number = None
    for line in text.split('\n'):
        line = line.strip()  # 去掉空格和换行符
        if re.match(r'^\d{2}-\w{5} \w{2}$', line):
            contract_number = line
            break  # 找到第一个匹配的就停止
        elif re.match(r'^\d{2}-\w{4}', line):
            contract_number = line
            break
        elif re.match(r'^\d{2}-\w{5}', line):
            contract_number = line
            break
        elif re.match(r'^\d{2}-\w{6}', line):
            contract_number = line
            break
        elif re.match(r'^\d{2}-\d{4}\w{1}$', line):
            contract_number = line
            break
        elif re.match(r'^\w{2}-\d{4}\w{2}$', line):
            contract_number = line
            break
    
    # 将结果添加到列表(存储图片路径而不是图片对象,因为Excel存不了图片)
    results.append({
        '图片路径': img_path,
        '识别到的合同编号': contract_number if contract_number else '未找到匹配编号'
    })

# 转换为DataFrame并导出到Excel
df = pd.DataFrame(results)
df.to_excel('./ContractNumberTest.xlsx', sheet_name='Air Testing', index=False)
print("Excel文件已成功导出!")

关键改进点说明

  • 递归加载图片: 使用glob.glob(..., recursive=True),**表示匹配任意层级的子文件夹,这样不用手动遍历每个子目录。
  • 单张图片处理: 循环遍历每个图片路径,逐个加载、调整大小、识别文字——这样就不会对列表调用单图函数了。
  • 结果存储: 存图片路径而不是图片对象(因为Excel无法直接存储OpenCV的图片对象),这样你可以在Excel里直接查看原始图片位置。
  • 容错处理: 加入了if image is None的判断,避免加载失败的图片导致程序崩溃。

更适合新手的递归加载方法(可选)

如果你觉得glob还是有点绕,可以试试Python 3.4+自带的pathlib模块,语法更直观:

from pathlib import Path

# 定义根目录
root_dir = Path('Image Samples/CE Test Pics')

# 递归查找所有图片文件(支持多种格式)
image_paths = list(root_dir.rglob('*.JPG')) + list(root_dir.rglob('*.jpg')) + list(root_dir.rglob('*.png'))

# 转换成字符串路径(因为cv2.imread需要字符串)
image_paths = [str(path) for path in image_paths]

这样写更像自然语言,新手更容易理解~

内容的提问来源于stack exchange,提问作者RR3327

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 19:03:13