使用Keras训练神经网络时遇NotFoundError:文件不存在问题求助
问题诊断与解决方案
核心问题定位
你的报错源于两个关键问题:预处理代码的逻辑错误,以及数据集路径与实际存储不匹配。
1. 预处理代码的致命错误
在「Batch and resize after batch」部分,你错误地将训练数据集train_ds用验证集val_ds的映射结果覆盖:
train_ds = val_ds.map(lambda imgs, lbls: (tf.image.resize(imgs, (IMG_SIZE, IMG_SIZE)), lbls), num_parallel_calls=AUTOTUNE)
这会导致训练集与验证集完全重复,同时之前的文件读取逻辑也存在冗余(定义了map_fn却未使用)。
2. 文件路径不匹配
报错提示train/60377.jpg不存在,需确认:
train.file中的路径是相对路径还是绝对路径?如果是相对路径,需保证代码运行目录与数据集目录层级匹配。- FairFace数据集是否完整解压,有没有遗漏文件?Linux/macOS系统需注意文件名大小写敏感问题。
修正后的完整预处理代码
import os import tensorflow as tf IMG_SIZE = 224 AUTOTUNE = tf.data.AUTOTUNE BATCH_SIZE = 224 NUM_CLASSES = len(labels_map) # 先过滤不存在的文件,避免后续报错 train = train[train.file.apply(lambda x: os.path.exists(x))] val = val[val.file.apply(lambda x: os.path.exists(x))] # 标签编码 y_train = tf.keras.utils.to_categorical(train.race, num_classes=NUM_CLASSES, dtype='float32') y_val = tf.keras.utils.to_categorical(val.race, num_classes=NUM_CLASSES, dtype='float32') # 创建数据集 train_ds = tf.data.Dataset.from_tensor_slices((train.file, y_train)).shuffle(len(y_train)) val_ds = tf.data.Dataset.from_tensor_slices((val.file, y_val)) # 验证数据集长度 assert len(train_ds) == len(train.file) == len(train.race) assert len(val_ds) == len(val.file) == len(val.race) # 统一预处理函数 def preprocess(path, label): # 处理文件读取异常 try: img = tf.io.read_file(path) img = tf.io.decode_jpeg(img) except tf.errors.NotFoundError: img = tf.zeros((IMG_SIZE, IMG_SIZE, 3), dtype=tf.uint8) img = tf.image.resize(img, (IMG_SIZE, IMG_SIZE)) img = tf.cast(img, tf.float32) / 255.0 # 归一化 return img, label # 应用预处理并优化性能 train_ds = train_ds.map(preprocess, num_parallel_calls=AUTOTUNE) val_ds = val_ds.map(preprocess, num_parallel_calls=AUTOTUNE) train_ds = train_ds.batch(BATCH_SIZE).prefetch(AUTOTUNE) val_ds = val_ds.batch(BATCH_SIZE).prefetch(AUTOTUNE)
额外验证步骤
- 手动检查缺失文件:在终端执行
ls train/60377.jpg(Linux/macOS)或dir train\60377.jpg(Windows),确认文件是否真的存在。 - 若路径是相对路径,可在代码中统一拼接绝对路径:
# 替换为你的FairFace数据集根目录 DATA_ROOT = "/path/to/your/fairface/dataset" train['file'] = train['file'].apply(lambda x: os.path.join(DATA_ROOT, x)) val['file'] = val['file'].apply(lambda x: os.path.join(DATA_ROOT, x))
内容的提问来源于stack exchange,提问作者oceanspray22
相关产品推荐
相关产品推荐

