You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速pd.DataFrame()将图像数据集转换为CSV的过程?

优化图像转CSV的速度与内存占用问题

你的问题核心在于一次性将所有图像数据存入列表后再转换为DataFrame,这会导致内存占用飙升,且DataFrame初始化时的类型转换、内存分配操作会消耗大量时间。下面给出两种高效的解决方案,帮你大幅提升处理速度:

方案一:直接用CSV模块逐行写入(最推荐)

完全跳过创建大型DataFrame的步骤,处理一张图像就写入一行CSV,内存占用极低,速度最快。

import os
import cv2
import numpy as np
import csv

root = 'test_case_images/Test_1/'
width = 224
height = 224

# 以追加模式打开CSV,newline=''避免自动添加空行
with open('data.csv', 'a', newline='', encoding='utf-8') as csvfile:
    writer = csv.writer(csvfile)
    
    for folder in os.listdir(root):
        folder_path = os.path.join(root, folder)
        # 过滤非文件夹的条目(避免目录下的杂文件)
        if not os.path.isdir(folder_path):
            continue
            
        for filename in os.listdir(folder_path):
            img_path = os.path.join(folder_path, filename)
            # 读取图像,处理读取失败的情况
            img = cv2.imread(img_path)
            if img is None:
                print(f"警告:无法读取图像 {img_path},已跳过")
                continue
                
            # 图像预处理:转灰度→调整尺寸→扁平化
            img_gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
            img_resized = cv2.resize(img_gray, (width, height))
            pixel_list = img_resized.flatten().tolist()
            
            # 拼接标签与像素值,写入一行
            writer.writerow([folder] + pixel_list)

为什么这个方法更快?

  • 不需要在内存中存储所有图像数据,处理一个写一个,内存占用几乎可以忽略。
  • 避免了DataFrame初始化时的大量内部操作(比如类型推断、内存对齐等),这些操作在数据量大时会非常耗时。

方案二:批量写入(保留Pandas习惯)

如果你更习惯用Pandas,可以通过批量处理+分批写入的方式,避免一次性创建超大DataFrame,分散内存压力和写入耗时。

import os
import cv2
import numpy as np
import pandas as pd

root = 'test_case_images/Test_1/'
width = 224
height = 224
batch_size = 1000  # 每处理1000张图像就写入一次
data_batch = []
# 判断文件是否已存在,决定是否需要写表头
file_exists = os.path.isfile('data.csv')

for folder in os.listdir(root):
    folder_path = os.path.join(root, folder)
    if not os.path.isdir(folder_path):
        continue
        
    for filename in os.listdir(folder_path):
        img_path = os.path.join(folder_path, filename)
        img = cv2.imread(img_path)
        if img is None:
            print(f"警告:无法读取图像 {img_path},已跳过")
            continue
            
        img_gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
        img_resized = cv2.resize(img_gray, (width, height))
        pixel_array = img_resized.flatten()
        # 拼接标签与像素
        row_data = np.hstack((folder, pixel_array))
        data_batch.append(row_data)
        
        # 达到批量大小就写入CSV
        if len(data_batch) >= batch_size:
            df = pd.DataFrame(data_batch)
            df.to_csv('data.csv', mode='a', header=not file_exists, index=False)
            # 第一次写入后,后续批次不再写表头
            file_exists = True
            data_batch.clear()

# 处理剩余的不足一个批次的图像
if data_batch:
    df = pd.DataFrame(data_batch)
    df.to_csv('data.csv', mode='a', header=not file_exists, index=False)
    data_batch.clear()

额外优化建议

  1. 并行处理:如果图像数量极多,可以用multiprocessing库并行处理图像读取与预处理(注意写入CSV时要加锁,避免多进程写入冲突)。
  2. 图像读取优化:可以尝试用Pillow代替OpenCV读取图像,某些场景下速度会略有提升。
  3. 路径缓存:提前拼接好所有图像的完整路径,避免循环中重复调用os.path.join,减少小开销的累积。

内容的提问来源于stack exchange,提问作者xafak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:05:25