You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何正确将图像集传入Keras模型进行回归训练?

解决方法

你的问题出在x_train的格式上:用apply得到的是一个Series,每个元素是带**额外batch维度(shape=(1, height, width, channels))**的numpy数组,这种嵌套数组结构无法被Keras直接解析成训练用的张量。下面是两种可行的解决方案,适配回归任务的需求:

方法1:将图像数据转为统一的NumPy数组(适合小数据集)

首先修正图像加载函数,去掉不必要的tf.expand_dims,同时强制统一图像尺寸(避免因输入图像大小不一致导致的错误):

import numpy as np
import tensorflow as tf

def load_image(img_path, target_size=(224, 224)):
    # 加载图像并统一尺寸
    img = tf.keras.preprocessing.image.load_img(img_path, target_size=target_size)
    # 转为数组并归一化(可选但推荐,加速模型收敛)
    img_array = tf.keras.preprocessing.image.img_to_array(img) / 255.0
    return img_array

# 加载所有图像并转为NumPy数组
x_train = np.array([load_image(path) for path in df['filename']])
# y_train需确保是NumPy数组(回归任务的连续值)
y_train = df['target_column'].values  # 替换成你的回归目标列名

之后就可以正常调用model.fit:

history = model.fit(x_train, y_train, epochs=10, batch_size=32)

方法2:用tf.data.Dataset构建数据集(适合大数据集,内存友好)

这种方法不需要一次性加载所有图像到内存,适合数据量较大的场景:

# 从路径和标签创建数据集
dataset = tf.data.Dataset.from_tensor_slices((df['filename'].values, df['target_column'].values))

def process_data(img_path, label):
    # 加载图像
    img = tf.io.read_file(img_path)
    img = tf.image.decode_jpeg(img, channels=3)  # 如果是PNG用decode_png
    # 统一尺寸并归一化
    img = tf.image.resize(img, (224, 224)) / 255.0
    return img, label

# 映射处理函数,设置批量和预取
train_dataset = dataset.map(process_data, num_parallel_calls=tf.data.AUTOTUNE)\
                     .batch(32)\
                     .prefetch(tf.data.AUTOTUNE)

# 训练模型
history = model.fit(train_dataset, epochs=10)

关键注意事项

  • 所有输入图像必须尺寸统一,否则模型无法接收不一致的输入形状。
  • 图像归一化(除以255)是常规操作,能让模型训练更稳定。
  • 回归任务的y_train必须是连续数值类型(如float32),确保数据类型正确。

内容的提问来源于stack exchange,提问作者Luca Sagoleo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 14:55:31