加载Google Drive数据集遇NotFoundError,求有效读取方案
解决Google Drive表情数据集加载问题
错误原因分析
报错路径显示为/content/drive/MyDrive/Dataset/images/images/train/angry671.jpg,明显是将angry目录名与目录内的图片文件名671.jpg错误拼接,大概率是路径处理逻辑或加载函数的问题,同时需先确认Google Drive是否正确挂载。
一、先完成Google Drive的正确挂载
在Colab环境中,必须先执行挂载授权才能访问Drive文件:
from google.colab import drive drive.mount('/content/drive')
执行后按照提示完成授权操作,确保Drive目录能在Colab左侧文件管理器中正常访问。
二、检查路径有效性
手动验证路径:打开Colab左侧文件管理器,导航至/content/drive/MyDrive/Dataset/images/images/train/,确认存在angry、disgust等子目录,且每个目录下有对应的图片文件。注意Linux系统路径区分大小写,确保目录名拼写完全一致。
三、推荐高效加载方案(TensorFlow内置方法)
使用tf.keras.utils.image_dataset_from_directory可以自动按子目录分类加载数据集,无需手动逐个类别处理,更稳定高效:
import tensorflow as tf # 训练集根目录路径 train_dir = '/content/drive/MyDrive/Dataset/images/images/train' # 加载数据集 train_ds = tf.keras.utils.image_dataset_from_directory( train_dir, image_size=(48, 48), # 表情数据集常用48x48尺寸,根据实际情况调整 batch_size=32, label_mode='categorical' # 多分类任务用categorical,二分类用binary ) # 查看数据集信息 class_names = train_ds.class_names print(f"数据集类别:{class_names}") # 查看单批次数据形状 for images, labels in train_ds.take(1): print(f"单批次图片形状:{images.shape}") print(f"单批次标签形状:{labels.shape}")
四、修复原有代码方案(若坚持自定义加载)
先实现正确的图片加载函数,同时修正代码中的语法错误:
import numpy as np from PIL import Image import os # 实现正确的批量图片加载函数 def load_images(dir_path): images = [] # 遍历目录下所有图片文件 for filename in os.listdir(dir_path): if filename.lower().endswith(('.jpg', '.png', '.jpeg')): img_path = os.path.join(dir_path, filename) # 按需转灰度图(表情数据集常用灰度图)或保留RGB img = Image.open(img_path).convert('L') # 'L'表示灰度图,转RGB用'RGB' img = np.array(img) # 若为灰度图,添加通道维度,统一形状为(样本数, 高, 宽, 通道数) if len(img.shape) == 2: img = np.expand_dims(img, axis=-1) images.append(img) return np.array(images) # 基础路径 path = '/content/drive/MyDrive/Dataset/images/images' # 逐个加载类别 angry = load_images(os.path.join(path, 'train', 'angry')) print('Loaded angry dataset: ', angry.shape) disgust = load_images(os.path.join(path, 'train', 'disgust')) print('Loaded disgust dataset: ', disgust.shape) fear = load_images(os.path.join(path, 'train', 'fear')) print('Loaded fear dataset: ', fear.shape) happy = load_images(os.path.join(path, 'train', 'happy')) print('Loaded happy dataset: ', happy.shape) neutral = load_images(os.path.join(path, 'train', 'neutral')) print('Loaded neutral dataset: ', neutral.shape) sad = load_images(os.path.join(path, 'train', 'sad')) print('Loaded sad dataset: ', sad.shape) surprise = load_images(os.path.join(path, 'train', 'surprise')) print('Loaded surprise dataset: ', surprise.shape) # 合并所有数据集(numpy的vstack仅接受一个数组序列参数) dataA = np.vstack([angry, disgust, fear, happy, neutral, sad, surprise]) print('Combined dataset shape: ', dataA.shape)
内容的提问来源于stack exchange,提问作者Jacob
相关产品推荐
相关产品推荐

