You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何结合TensorFlow Dataset与ImageDataGenerator构建图像识别流水线(解决TypeError报错)

解决TensorFlow Datasets结合ImageDataGenerator的适配问题

错误原因分析

你遇到的TypeError: All values in column x_col=image must be strings本质是API设计的匹配问题:flow_from_dataframe的x_col参数要求传入的是图片文件的路径字符串(它会根据路径去磁盘加载图片),但tfds.as_dataframe()转换后,image列存储的是已经解码完成的numpy数组(不是文件路径),自然不符合API的输入要求。

下面提供两种解决方案,优先满足你想用ImageDataGenerator的需求,同时也给出更贴合TF2生态的原生方案。


方案1:适配ImageDataGenerator的实现(保存图片到本地)

既然flow_from_dataframe需要文件路径,我们可以先把TFDS中的图片保存到本地目录(和你之前用的Keras推荐目录结构一致),再生成包含路径和标签的DataFrame供API调用。

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import os
import tensorflow as tf
from tensorflow.keras.optimizers import Adam
from tensorflow.keras.preprocessing.image import ImageDataGenerator
import tensorflow_datasets as tfds

# 加载数据集
(train_ds, test_ds), ds_info = tfds.load(
    'beans', 
    split=['train', 'test'], 
    as_supervised=True, 
    download=True, 
    with_info=True
)

# 获取数据集的类别名称
class_names = ds_info.features['label'].names

# 创建本地存储目录,按训练/测试+类别划分
base_dir = './beans_local'
train_dir = os.path.join(base_dir, 'train')
test_dir = os.path.join(base_dir, 'test')

# 递归创建目录
for split_dir in [train_dir, test_dir]:
    for cls in class_names:
        os.makedirs(os.path.join(split_dir, cls), exist_ok=True)

# 定义函数:将数据集图片保存到本地,并生成带路径的DataFrame
def save_images(dataset, save_dir, class_names):
    img_paths = []
    labels = []
    for idx, (img, label) in enumerate(dataset):
        cls_name = class_names[label.numpy()]
        # 构建图片保存路径
        img_path = os.path.join(save_dir, cls_name, f'img_{idx}.png')
        # 将Tensor格式的图片转成numpy数组并保存
        tf.keras.preprocessing.image.save_img(img_path, img.numpy())
        img_paths.append(img_path)
        labels.append(cls_name)
    return pd.DataFrame({'image': img_paths, 'label': labels})

# 生成训练/测试集的DataFrame
df_train = save_images(train_ds, train_dir, class_names)
df_test = save_images(test_ds, test_dir, class_names)

# 初始化数据生成器
train_datagen = ImageDataGenerator(
    rescale=1./255,
    rotation_range=40,
    width_shift_range=0.2,
    height_shift_range=0.2,
    shear_range=0.2,
    zoom_range=0.2,
    horizontal_flip=True,
    fill_mode='nearest'
)

test_datagen = ImageDataGenerator(rescale=1./255)

# 现在可以正常使用flow_from_dataframe了
train_generator = train_datagen.flow_from_dataframe(
    df_train,
    x_col="image",
    y_col="label",
    target_size=(500, 500),
    batch_size=20,
    class_mode='categorical'
)

test_generator = test_datagen.flow_from_dataframe(
    df_test,
    x_col="image",
    y_col="label",
    target_size=(500, 500),
    batch_size=20,
    class_mode='categorical'
)

方案2:TF2原生方案(tf.data + 数据增强)

如果不局限于ImageDataGenerator,更推荐用TF2的tf.data.Dataset结合原生数据增强方式,这种方式不需要额外磁盘存储,效率更高,也更贴合TFDS的使用场景。

方式A:用tf.image函数实现增强

import tensorflow as tf
import tensorflow_datasets as tfds

# 加载数据集
(train_ds, test_ds), ds_info = tfds.load(
    'beans', 
    split=['train', 'test'], 
    as_supervised=True, 
    download=True, 
    with_info=True
)

class_names = ds_info.features['label'].names
num_classes = len(class_names)
img_size = (500, 500)
batch_size = 20

# 定义训练集预处理(含数据增强)
def preprocess_train(image, label):
    # 归一化到[0,1]
    image = tf.cast(image, tf.float32) / 255.0
    # 调整图片尺寸
    image = tf.image.resize(image, img_size)
    # 数据增强操作
    image = tf.image.random_flip_left_right(image)
    image = tf.image.random_brightness(image, max_delta=0.2)
    image = tf.image.random_contrast(image, lower=0.8, upper=1.2)
    image = tf.image.random_zoom(image, zoom_range=(0.8, 1.2))
    # 标签转one-hot编码
    label = tf.one_hot(label, depth=num_classes)
    return image, label

# 定义测试集预处理(仅归一化和 resize)
def preprocess_test(image, label):
    image = tf.cast(image, tf.float32) / 255.0
    image = tf.image.resize(image, img_size)
    label = tf.one_hot(label, depth=num_classes)
    return image, label

# 处理数据集,提升性能
train_ds = train_ds.map(preprocess_train, num_parallel_calls=tf.data.AUTOTUNE)
train_ds = train_ds.shuffle(1000).batch(batch_size).prefetch(tf.data.AUTOTUNE)

test_ds = test_ds.map(preprocess_test, num_parallel_calls=tf.data.AUTOTUNE)
test_ds = test_ds.batch(batch_size).prefetch(tf.data.AUTOTUNE)

# 后续直接传入model.fit()即可
# model.fit(train_ds, epochs=10, validation_data=test_ds)

方式B:用Keras预处理层(可集成到模型,方便部署)

这种方式可以把数据增强层直接集成到模型中,后续部署时不需要单独处理数据:

import tensorflow as tf
import tensorflow_datasets as tfds
from tensorflow.keras import Sequential, layers

# 加载数据集
(train_ds, test_ds), ds_info = tfds.load(
    'beans', 
    split=['train', 'test'], 
    as_supervised=True, 
    download=True, 
    with_info=True
)

class_names = ds_info.features['label'].names
num_classes = len(class_names)
img_size = (500, 500)
batch_size = 20

# 定义数据增强层(仅训练时生效)
data_augmentation = Sequential([
    layers.Resizing(*img_size),
    layers.Rescaling(1./255),
    layers.RandomFlip("horizontal"),
    layers.RandomRotation(0.2),
    layers.RandomZoom(0.2),
])

# 训练集预处理
def preprocess_train(image, label):
    image = data_augmentation(image)
    label = tf.one_hot(label, depth=num_classes)
    return image, label

# 测试集预处理
def preprocess_test(image, label):
    image = layers.Resizing(*img_size)(image)
    image = layers.Rescaling(1./255)(image)
    label = tf.one_hot(label, depth=num_classes)
    return image, label

# 处理数据集
train_ds = train_ds.map(preprocess_train, num_parallel_calls=tf.data.AUTOTUNE)
train_ds = train_ds.shuffle(1000).batch(batch_size).prefetch(tf.data.AUTOTUNE)

test_ds = test_ds.map(preprocess_test, num_parallel_calls=tf.data.AUTOTUNE)
test_ds = test_ds.batch(batch_size).prefetch(tf.data.AUTOTUNE)

内容的提问来源于stack exchange,提问作者Enrico Anderlini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 23:27:35