You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Google Drive文件夹连接到Google Colab并加载所有狗品种图像

在Google Colab中加载Google Drive中的狗品种图像

1. 挂载Google Drive到Colab

首先需要将你的Google Drive挂载到Colab环境中,执行以下代码并按照提示完成授权:

from google.colab import drive
drive.mount('/content/drive')

2. 定义图像根路径

根据你的Drive路径,Colab中对应的完整路径为:

root_dir = '/content/drive/MyDrive/Data Science/Deep_learning/dogImages'

3. 加载所有图像到dog_breed_image变量

方法一:手动遍历加载(自定义性强)

适合需要对图像做个性化预处理的场景,使用PIL和os模块遍历所有文件夹:

import os
from PIL import Image

dog_breed_image = []
# 遍历train、test、valid三个子文件夹
for split_dir in ['train', 'test', 'valid']:
    split_path = os.path.join(root_dir, split_dir)
    # 遍历每个品种文件夹
    for breed_dir in os.listdir(split_path):
        breed_path = os.path.join(split_path, breed_dir)
        # 跳过非文件夹项
        if not os.path.isdir(breed_path):
            continue
        # 遍历文件夹内的所有图像
        for img_name in os.listdir(breed_path):
            img_path = os.path.join(breed_path, img_name)
            try:
                # 打开图像并添加到列表
                img = Image.open(img_path)
                dog_breed_image.append(img)
            except Exception as e:
                print(f"无法加载图像 {img_path}: {e}")

dog_breed_image会存储所有加载成功的PIL.Image对象,你可以后续统一调整尺寸、转换为数组等。

方法二:用Keras数据生成器加载(适合深度学习)

如果是为深度学习任务准备数据,推荐使用tensorflow.keras的ImageDataGenerator,可以同时完成加载和预处理:

import tensorflow as tf
from tensorflow.keras.preprocessing.image import ImageDataGenerator

# 初始化数据生成器(可添加预处理参数,比如归一化)
datagen = ImageDataGenerator(rescale=1./255)

# 从目录加载数据,指定图像尺寸
generator = datagen.flow_from_directory(
    root_dir,
    target_size=(224, 224),  # 根据需求调整图像大小
    batch_size=32,
    class_mode='categorical',
    shuffle=False
)

# 将所有数据加载到变量中(注意:数据量过大时可能占用大量内存)
images, labels = next(generator)
for i in range(len(generator)-1):
    batch_img, batch_label = next(generator)
    images = tf.concat([images, batch_img], axis=0)
    labels = tf.concat([labels, batch_label], axis=0)

# 最终images即为所有图像的numpy数组形式,可赋值给dog_breed_image
dog_breed_image = images

这种方式会将图像转换为统一尺寸的numpy数组,同时返回对应的品种标签,适合直接输入到深度学习模型中。

注意事项

  • 如果图像数量较多,方法二一次性加载所有数据可能导致内存不足,此时建议保留生成器形式分批处理,而非全部存入变量。
  • 手动加载时注意过滤非图像文件(比如隐藏文件),避免加载报错。

内容的提问来源于stack exchange,提问作者Surendra kumawat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 13:24:28