You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CV2读取Google Drive中图像速度过慢,求优化方法

优化Google Colab中从Google Drive读取图像速度的方案

核心瓶颈分析

当前单张图像读取耗时0.3-0.5秒,主要瓶颈是Google Drive的网络IO延迟,其次是单线程串行读取的低效性。以下是针对性优化方案:


1. 将Drive图像文件夹复制到Colab本地磁盘

Colab本地磁盘的IO速度远高于Drive的网络传输速度,先批量复制文件到本地,再读取本地文件能大幅降低读取耗时。

# 复制Drive中的目标图像文件夹到Colab本地(替换为你的实际路径)
!cp -r "/content/drive/MyDrive/your_mri_images_folder" "/content/local_mri_images"

之后读取时直接使用本地路径/content/local_mri_images,而非Drive路径。


2. 使用多线程并行读取图像

图像读取属于IO密集型任务,单线程串行处理会浪费CPU空闲时间,用多线程并行读取可同时处理多个文件的IO操作,提升整体速度。

import concurrent.futures
import cv2
import numpy as np
import random
import os

# 配置参数
local_jpg_folder = "/content/local_mri_images"
n_samples = 10000
width = 112
height = 112

# 打乱文件列表并取前n_samples个
jpg_files = os.listdir(local_jpg_folder)
random.shuffle(jpg_files)
target_files = jpg_files[:n_samples]

# 预分配数组
mri_images = np.empty((n_samples, height, width), dtype=np.float32)

# 定义单张图像读取函数
def load_single_image(idx, file_name):
    file_path = os.path.join(local_jpg_folder, file_name)
    img = cv2.imread(file_path, 0)
    return idx, img

# 启动线程池(max_workers可根据Colab资源调整,建议8-16)
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as executor:
    # 提交所有读取任务
    task_futures = [executor.submit(load_single_image, i, fname) for i, fname in enumerate(target_files)]
    
    # 收集结果并写入数组
    for future in concurrent.futures.as_completed(task_futures):
        idx, img = future.result()
        mri_images[idx] = img

3. 预转换为Numpy二进制格式(.npy)

如果需要重复读取图像,可将所有图像一次性转换并保存为Numpy的.npy二进制文件,后续读取时直接加载整个数组,速度比逐个解码JPG快一个数量级。

# 首次运行:读取所有图像并保存为npy
all_images = []
for fname in target_files:
    img_path = os.path.join(local_jpg_folder, fname)
    img = cv2.imread(img_path, 0)
    all_images.append(img)

mri_images_np = np.array(all_images, dtype=np.float32)
np.save("/content/local_mri_images/all_mri_samples.npy", mri_images_np)

# 后续直接加载
mri_images = np.load("/content/local_mri_images/all_mri_samples.npy")

4. 小细节优化

  • 使用os.path.join拼接路径,避免字符串拼接的潜在错误,同时提升代码兼容性;
  • 提前过滤非JPG文件,避免读取时的无效IO操作;
  • 若图像尺寸统一,可跳过尺寸检查,减少额外计算。

内容的提问来源于stack exchange,提问作者Baki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 14:40:24