使用CV2读取Google Drive中图像速度过慢,求优化方法
优化Google Colab中从Google Drive读取图像速度的方案
核心瓶颈分析
当前单张图像读取耗时0.3-0.5秒,主要瓶颈是Google Drive的网络IO延迟,其次是单线程串行读取的低效性。以下是针对性优化方案:
1. 将Drive图像文件夹复制到Colab本地磁盘
Colab本地磁盘的IO速度远高于Drive的网络传输速度,先批量复制文件到本地,再读取本地文件能大幅降低读取耗时。
# 复制Drive中的目标图像文件夹到Colab本地(替换为你的实际路径) !cp -r "/content/drive/MyDrive/your_mri_images_folder" "/content/local_mri_images"
之后读取时直接使用本地路径/content/local_mri_images,而非Drive路径。
2. 使用多线程并行读取图像
图像读取属于IO密集型任务,单线程串行处理会浪费CPU空闲时间,用多线程并行读取可同时处理多个文件的IO操作,提升整体速度。
import concurrent.futures import cv2 import numpy as np import random import os # 配置参数 local_jpg_folder = "/content/local_mri_images" n_samples = 10000 width = 112 height = 112 # 打乱文件列表并取前n_samples个 jpg_files = os.listdir(local_jpg_folder) random.shuffle(jpg_files) target_files = jpg_files[:n_samples] # 预分配数组 mri_images = np.empty((n_samples, height, width), dtype=np.float32) # 定义单张图像读取函数 def load_single_image(idx, file_name): file_path = os.path.join(local_jpg_folder, file_name) img = cv2.imread(file_path, 0) return idx, img # 启动线程池(max_workers可根据Colab资源调整,建议8-16) with concurrent.futures.ThreadPoolExecutor(max_workers=12) as executor: # 提交所有读取任务 task_futures = [executor.submit(load_single_image, i, fname) for i, fname in enumerate(target_files)] # 收集结果并写入数组 for future in concurrent.futures.as_completed(task_futures): idx, img = future.result() mri_images[idx] = img
3. 预转换为Numpy二进制格式(.npy)
如果需要重复读取图像,可将所有图像一次性转换并保存为Numpy的.npy二进制文件,后续读取时直接加载整个数组,速度比逐个解码JPG快一个数量级。
# 首次运行:读取所有图像并保存为npy all_images = [] for fname in target_files: img_path = os.path.join(local_jpg_folder, fname) img = cv2.imread(img_path, 0) all_images.append(img) mri_images_np = np.array(all_images, dtype=np.float32) np.save("/content/local_mri_images/all_mri_samples.npy", mri_images_np) # 后续直接加载 mri_images = np.load("/content/local_mri_images/all_mri_samples.npy")
4. 小细节优化
- 使用
os.path.join拼接路径,避免字符串拼接的潜在错误,同时提升代码兼容性; - 提前过滤非JPG文件,避免读取时的无效IO操作;
- 若图像尺寸统一,可跳过尺寸检查,减少额外计算。
内容的提问来源于stack exchange,提问作者Baki
相关产品推荐
相关产品推荐

