循环中仅首张图像存入DataFrame的原因排查求助
问题分析与修复方案
核心问题
你代码里的内层os.walk循环重复使用了外层循环的变量名root、dirnames、filenames,这会直接覆盖外层循环的迭代状态,导致外层循环只执行第一次迭代就终止,自然只会处理第一张图像。
另外还有几个可优化的点:
- 没必要每次处理图像都重新遍历整个标签目录,提前把所有标签文件名和路径映射好,能大幅提升效率
- 硬编码的
2000缩放值改成用图像实际尺寸计算,避免尺寸不匹配问题 cv2.rectangle不需要重新赋值给mask,因为它是直接在原数组上修改的
修复后的代码
import os import re import cv2 import numpy as np import pandas as pd mask_images = [] # 灰度掩码数据集 RGB_images = [] # RGB图像数据集 # 提前构建标签文件路径映射:文件名 -> 文件路径 label_path_map = {} for root, _, filenames in os.walk("dataset/test/detection_labels"): for txtname in filenames: label_path_map[txtname] = os.path.join(root, txtname) # 遍历处理所有图像 for root, _, filenames in os.walk("dataset/test/images"): for imgname in filenames: if re.search(r"\.(jpg|jpeg|png|bmp|tiff)$", imgname, flags=re.IGNORECASE): imgpath = os.path.join(root, imgname) # 生成对应的标签文件名 txt_name = os.path.splitext(imgname)[0] + ".txt" # 读取并转换图像 img = cv2.imread(imgpath) if img is None: print(f"无法读取图像: {imgpath}") continue rgb_image = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) RGB_images.append(rgb_image) # 初始化掩码 mask = np.zeros((img.shape[0], img.shape[1]), dtype=np.uint8) # 查找对应标签文件并处理 if txt_name in label_path_map: txtpath = label_path_map[txt_name] with open(txtpath, 'r') as f: lines = f.readlines() for line in lines: coords = [float(num) for num in re.findall(r'-?\d+\.?\d*', line)] # 用图像实际尺寸计算坐标(替换原硬编码的2000) img_h, img_w = img.shape[:2] x = int(coords[1] * img_w) y = int(coords[2] * img_h) w = int(coords[3] * img_w) h = int(coords[4] * img_h) # 绘制填充矩形到掩码 cv2.rectangle(mask, (x, y), (x + w, y + h), color=255, thickness=-1) mask_images.append(mask) # 将数据存入DataFrame df = pd.DataFrame({'rgb_image': RGB_images, 'mask_image': mask_images})
关键修改说明
- 提前遍历一次标签目录,构建文件名到路径的字典,避免重复遍历,提升效率
- 内层循环不再使用外层的变量名,避免覆盖迭代状态
- 用图像实际尺寸替换硬编码的2000,适配不同大小的图像
- 增加图像读取失败的判断,避免程序崩溃
- 简化
cv2.rectangle的调用,不需要重新赋值给mask
内容的提问来源于stack exchange,提问作者RF_07
相关产品推荐
相关产品推荐

