You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Keras训练神经网络时遇NotFoundError:文件不存在问题求助

问题诊断与解决方案

核心问题定位

你的报错源于两个关键问题:预处理代码的逻辑错误,以及数据集路径与实际存储不匹配。


1. 预处理代码的致命错误

在「Batch and resize after batch」部分,你错误地将训练数据集train_ds用验证集val_ds的映射结果覆盖:

train_ds = val_ds.map(lambda imgs, lbls: (tf.image.resize(imgs, (IMG_SIZE, IMG_SIZE)), lbls), num_parallel_calls=AUTOTUNE)

这会导致训练集与验证集完全重复,同时之前的文件读取逻辑也存在冗余(定义了map_fn却未使用)。

2. 文件路径不匹配

报错提示train/60377.jpg不存在,需确认:

  • train.file中的路径是相对路径还是绝对路径?如果是相对路径,需保证代码运行目录与数据集目录层级匹配。
  • FairFace数据集是否完整解压,有没有遗漏文件?Linux/macOS系统需注意文件名大小写敏感问题。

修正后的完整预处理代码

import os
import tensorflow as tf

IMG_SIZE = 224
AUTOTUNE = tf.data.AUTOTUNE
BATCH_SIZE = 224
NUM_CLASSES = len(labels_map)

# 先过滤不存在的文件,避免后续报错
train = train[train.file.apply(lambda x: os.path.exists(x))]
val = val[val.file.apply(lambda x: os.path.exists(x))]

# 标签编码
y_train = tf.keras.utils.to_categorical(train.race, num_classes=NUM_CLASSES, dtype='float32')
y_val = tf.keras.utils.to_categorical(val.race, num_classes=NUM_CLASSES, dtype='float32')

# 创建数据集
train_ds = tf.data.Dataset.from_tensor_slices((train.file, y_train)).shuffle(len(y_train))
val_ds = tf.data.Dataset.from_tensor_slices((val.file, y_val))

# 验证数据集长度
assert len(train_ds) == len(train.file) == len(train.race)
assert len(val_ds) == len(val.file) == len(val.race)

# 统一预处理函数
def preprocess(path, label):
    # 处理文件读取异常
    try:
        img = tf.io.read_file(path)
        img = tf.io.decode_jpeg(img)
    except tf.errors.NotFoundError:
        img = tf.zeros((IMG_SIZE, IMG_SIZE, 3), dtype=tf.uint8)
    
    img = tf.image.resize(img, (IMG_SIZE, IMG_SIZE))
    img = tf.cast(img, tf.float32) / 255.0  # 归一化
    return img, label

# 应用预处理并优化性能
train_ds = train_ds.map(preprocess, num_parallel_calls=AUTOTUNE)
val_ds = val_ds.map(preprocess, num_parallel_calls=AUTOTUNE)

train_ds = train_ds.batch(BATCH_SIZE).prefetch(AUTOTUNE)
val_ds = val_ds.batch(BATCH_SIZE).prefetch(AUTOTUNE)

额外验证步骤

  1. 手动检查缺失文件:在终端执行ls train/60377.jpg(Linux/macOS)或dir train\60377.jpg(Windows),确认文件是否真的存在。
  2. 若路径是相对路径,可在代码中统一拼接绝对路径:
# 替换为你的FairFace数据集根目录
DATA_ROOT = "/path/to/your/fairface/dataset"
train['file'] = train['file'].apply(lambda x: os.path.join(DATA_ROOT, x))
val['file'] = val['file'].apply(lambda x: os.path.join(DATA_ROOT, x))

内容的提问来源于stack exchange,提问作者oceanspray22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:20:32