Google Colab中cv2.imread读取4000张图片过慢及代码复用问题求助
解决Colab读取4000张图片慢+重复运行问题的方案
一、加速图片读取与处理
1. 先把Drive图片拷贝到Colab本地
Google Drive是网络挂载,读取速度远不如Colab本地磁盘。先批量复制图片到/content/目录,能大幅降低读取延迟:
import shutil import os # 替换成你Drive里的图片文件夹路径 drive_img_dir = '/content/drive/MyDrive/your_image_folder' # Colab本地存放路径 local_img_dir = '/content/local_images' # 创建本地文件夹 os.makedirs(local_img_dir, exist_ok=True) # 批量复制所有图片 for img_name in os.listdir(drive_img_dir): src_path = os.path.join(drive_img_dir, img_name) dst_path = os.path.join(local_img_dir, img_name) shutil.copy(src_path, dst_path)
2. 用多进程并行处理
单循环一张一张读图是串行执行,效率极低。用多进程同时处理多张图片,直接把时间砍到原来的1/4甚至1/8:
import cv2 import numpy as np from concurrent.futures import ProcessPoolExecutor # 定义单张图片的处理函数 def process_single_img(img_path): img = cv2.imread(img_path) if img is not None: # 跳过损坏的图片 img = cv2.resize(img, (64, 64)) return img return None # 获取本地所有图片的完整路径 all_img_paths = [os.path.join(local_img_dir, name) for name in os.listdir(local_img_dir)] # 开启多进程,max_workers设为Colab给的CPU核心数(一般4-8) with ProcessPoolExecutor(max_workers=4) as executor: processed_imgs = list(executor.map(process_single_img, all_img_paths)) # 过滤掉损坏的图片,转成numpy数组 final_array = np.array([img for img in processed_imgs if img is not None])
3. 换更快的读取库试试
有些场景下PIL.Image比cv2.imread读取速度更快,你可以替换试试:
from PIL import Image def process_with_pil(img_path): try: with Image.open(img_path) as img: img_resized = img.resize((64, 64)) return np.array(img_resized) except: # 捕获打开失败的情况 return None # 同样用多进程处理 with ProcessPoolExecutor(max_workers=4) as executor: processed_imgs = list(executor.map(process_with_pil, all_img_paths)) final_array = np.array([img for img in processed_imgs if img is not None])
二、避免重复运行:直接保存处理好的数据集
每次重新登录都重跑太浪费时间,处理完直接把numpy数组存到Drive,下次打开直接加载就行:
# 保存到Drive,替换成你想存的路径 save_path = '/content/drive/MyDrive/processed_64x64_images.npy' np.save(save_path, final_array) # 下次打开Colab,直接加载这个文件就行,不用再读图处理 final_array = np.load(save_path)
如果需要同时保存图片和标签,用np.savez打包:
# 假设你有对应的标签数组labels_array np.savez('/content/drive/MyDrive/image_data.npz', images=final_array, labels=labels_array) # 加载时 data = np.load('/content/drive/MyDrive/image_data.npz') loaded_images = data['images'] loaded_labels = data['labels']
额外提醒
- Colab闲置太久会自动断开会话,处理时尽量保持页面活跃,或者用Colab Pro的长期会话功能。
- 提前检查图片是否有损坏,损坏的图片会拖慢读取速度,处理时记得加异常捕获跳过。
内容的提问来源于stack exchange,提问作者Ejaz
相关产品推荐
相关产品推荐

