You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NumPy数组转Tensor失败:文本转图像模型训练报错求助

解决NumPy数组转Tensor失败的问题

错误根源

你遇到的ValueError: Failed to convert a NumPy array to a Tensor (Unsupported object type list),本质是texts数组的类型问题:

  • tf.keras.preprocessing.text.one_hot返回的是长度不一的整数列表,当你用np.array()把这些列表转成numpy数组时,因为每个元素长度不同,numpy只能生成object类型的数组(数组里存的是列表对象),而TensorFlow无法直接将这种object数组转换为张量。

具体解决方案

方案1:统一文本序列长度(最常用)

对one-hot后的文本序列做填充/截断,让所有序列长度一致,生成二维numpy数组,TensorFlow可直接转换。

修改texts的生成代码:

data = pd.read_csv('data/captions.csv', delimiter=',')
captions = data['caption'].values
image_paths = "data/images/" + data['filename'].values

def load_image(image_path):
    return np.array(Image.open(image_path).resize((IMAGE_SIZE, IMAGE_SIZE)))

images = np.array([load_image(path) for path in image_paths])

# 步骤1:生成one-hot序列
one_hot_texts = [tf.keras.preprocessing.text.one_hot(caption, TEXT_EMBEDDING_DIM) for caption in captions]
# 步骤2:确定统一的序列长度(可选:取最长序列长度,或指定固定值)
MAX_SEQ_LEN = max(len(seq) for seq in one_hot_texts)
# 步骤3:填充/截断序列,统一长度
texts = tf.keras.preprocessing.sequence.pad_sequences(
    one_hot_texts,
    maxlen=MAX_SEQ_LEN,
    padding='post',  # 不足长度的在末尾补0
    truncating='post'  # 超过长度的截断末尾
)

修改后texts是形状为(样本数, MAX_SEQ_LEN)的二维整数数组,可正常传入模型训练。

方案2:用tf.data.Dataset处理可变长度序列

如果不想统一序列长度,可以用TensorFlow的数据集管道处理可变长度输入:

修改训练代码:

# 构建训练数据集
train_dataset = tf.data.Dataset.from_tensor_slices(
    (
        [np.random.randn(len(image_train), LATENT_DIM), text_train],
        image_train
    )
).batch(BATCH_SIZE)

# 构建验证数据集
val_dataset = tf.data.Dataset.from_tensor_slices(
    (
        [np.random.randn(len(image_val), LATENT_DIM), text_val],
        image_val
    )
).batch(BATCH_SIZE)

# 训练模型
generator_model.fit(train_dataset, epochs=EPOCHS, validation_data=val_dataset)

注意:模型的文本输入层需要设置为接受可变长度张量,比如Input(shape=(None,))。

额外优化:改用Tokenizer更规范

one-hot编码功能有限,推荐用Tokenizer处理文本,更灵活且便于后续嵌入层使用:

tokenizer = tf.keras.preprocessing.text.Tokenizer(num_words=TEXT_EMBEDDING_DIM)
tokenizer.fit_on_texts(captions)
# 转成序列
text_sequences = tokenizer.texts_to_sequences(captions)
# 统一长度
MAX_SEQ_LEN = max(len(seq) for seq in text_sequences)
texts = tf.keras.preprocessing.sequence.pad_sequences(
    text_sequences,
    maxlen=MAX_SEQ_LEN,
    padding='post',
    truncating='post'
)

内容的提问来源于stack exchange,提问作者Sheker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 01:32:46