You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

搭建CNN训练用数据生成器遇梯度缺失问题求助

问题:CNN回归任务中数据生成器导致"No gradients provided"错误

任务目标:搭建CNN模型,输入图片后输出25个连续值(回归任务)。数据存储在带自目录的文件夹中,每个子目录包含.jpg图片和对应存储25个值的.csv文件。尝试用自定义数据生成器实现,但训练时返回错误:

ValueError: No gradients provided for any variable:

图片已确认能正确加载,但标签加载存在问题。以下是用户提供的代码:


自定义数据生成器代码

import os, shutil, csv
import tensorflow as tf
from tensorflow.keras.preprocessing.image import ImageDataGenerator

def data_generator(data_dir, batch_size):
    # Set up the data generator
    datagen = ImageDataGenerator(rescale=1./255)

    # Read in the data from the specified directory
    generator = datagen.flow_from_directory(
        data_dir,
        target_size=image_size,
        batch_size=batch_size,
        class_mode=None,  # since the labels are stored in separate .csv files
        shuffle=True
    )
    
    # Iterate over the data
    for data_batch, _ in generator:
        # Read in the corresponding .csv file for each image
        labels_batch = []
        for i, filename in enumerate(data_batch):
            # Get the file name without the extension
            file_base = os.path.splitext(filename)[0]
    
            # Read in the labels from the .csv file
            with open(file_base + '.csv', 'r') as csv_file:
                reader = csv.reader(csv_file)
                labels_batch.append(next(reader))
        yield data_batch, labels_batch

train_generator = data_generator(os.path.join(data_dir, 'train'), batch_size)

模型编译与训练代码

# Compile the model
model.compile(optimizer='adam', loss='mean_squared_error', metrics=['accuracy'])

# Train the model using the data generators
history = model.fit_generator(
train_generator,
steps_per_epoch=len(train_generator),
epochs=10,
validation_data=val_generator,
validation_steps=len(val_generator)
)

# Evaluate the model on the testing data
test_loss, test_acc = model.evaluate_generator(test_generator, steps=len(test_generator))

问题分析与解决方法

核心错误原因

  1. 标签格式不兼容:生成器返回的labels_batch是Python字符串列表,不是TensorFlow能计算梯度的数值型张量(numpy数组或tf.Tensor)。
  2. 文件名获取错误:enumerate(data_batch)遍历的是图片张量的维度(如像素),而非图片文件名,导致无法匹配对应的.csv标签文件。
  3. 指标误用:回归任务使用分类指标accuracy,会导致计算逻辑异常,干扰梯度传递。

修正后的实现方案

1. 修复数据生成器

import os, csv
import numpy as np
import tensorflow as tf
from tensorflow.keras.preprocessing.image import ImageDataGenerator

def data_generator(data_dir, batch_size, image_size):
    datagen = ImageDataGenerator(rescale=1./255)
    
    # 配置生成器以保留文件名
    generator = datagen.flow_from_directory(
        data_dir,
        target_size=image_size,
        batch_size=batch_size,
        class_mode=None,
        shuffle=True,
        save_format='jpg'
    )
    
    while True:
        data_batch = next(generator)
        # 获取当前batch对应的文件名
        current_idx = generator.batch_index
        if current_idx >= batch_size:
            filenames = generator.filenames[current_idx - batch_size : current_idx]
        else:
            # 处理循环到数据集末尾的情况
            filenames = generator.filenames[-batch_size:]
        
        labels_batch = []
        for filename in filenames:
            # 拼接完整的csv文件路径
            img_full_path = os.path.join(data_dir, filename)
            csv_full_path = os.path.splitext(img_full_path)[0] + '.csv'
            
            # 读取csv并转为float数组
            with open(csv_full_path, 'r') as f:
                reader = csv.reader(f)
                label = np.array(next(reader), dtype=np.float32)
                labels_batch.append(label)
        
        # 转为numpy数组,确保形状为(batch_size, 25)
        labels_batch = np.array(labels_batch)
        yield data_batch, labels_batch

# 调用示例
train_generator = data_generator(os.path.join(data_dir, 'train'), 32, (224,224))

2. 修正模型编译

回归任务替换为适合的指标(如平均绝对误差mae):

model.compile(optimizer='adam', loss='mean_squared_error', metrics=['mae'])

3. 修正训练代码

TensorFlow 2.10+已弃用fit_generator,改用fit,并手动计算steps_per_epoch:

# 统计训练集总样本数
train_total = 0
for root, dirs, files in os.walk(os.path.join(data_dir, 'train')):
    train_total += sum(1 for f in files if f.endswith('.jpg'))
steps_per_epoch = train_total // batch_size

# 训练模型
history = model.fit(
    train_generator,
    steps_per_epoch=steps_per_epoch,
    epochs=10,
    validation_data=val_generator,
    validation_steps=val_total // batch_size
)

# 评估模型
test_loss, test_mae = model.evaluate(test_generator, steps=test_total // batch_size)

可选:更高效的tf.data.Dataset实现(推荐)

若允许调整数据存储结构(将图片和csv放在同一目录,无嵌套子目录),可使用TensorFlow原生数据集API:

def load_image_and_label(img_path):
    # 加载并预处理图片
    img = tf.io.read_file(img_path)
    img = tf.image.decode_jpeg(img, channels=3)
    img = tf.image.resize(img, (224,224))
    img = img / 255.0
    
    # 加载并解析标签
    csv_path = tf.strings.regex_replace(img_path, '.jpg', '.csv')
    csv_raw = tf.io.read_file(csv_path)
    label = tf.io.decode_csv(csv_raw, record_defaults=[tf.float32]*25)
    label = tf.stack(label)
    
    return img, label

# 构建训练数据集
train_dataset = tf.data.Dataset.list_files(os.path.join(data_dir, 'train', '*.jpg'))
train_dataset = train_dataset.map(load_image_and_label, num_parallel_calls=tf.data.AUTOTUNE)
train_dataset = train_dataset.batch(32).shuffle(1000).prefetch(tf.data.AUTOTUNE)

# 训练
history = model.fit(train_dataset, epochs=10, validation_data=val_dataset)

内容的提问来源于stack exchange,提问作者sbudoff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 03:40:45