You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow报错:Roberta-Base模型期望2个输入却仅收到1个

问题排查与解决方案

核心问题定位

报错明确说明模型需要2个输入,但数据集仅输出1个张量。你通过map_dataset转换为字典形式但仍失败,大概率是数据集输出结构与模型输入要求不匹配,或输入类型、键名存在不一致。

具体排查与修复步骤

1. 核对模型输入层定义

Roberta-Base在TensorFlow中默认需要input_ids和attention_mask两个输入,确认你的模型输入层是否正确定义:

# 标准输入层定义(需与数据集输出键名一致)
input_ids = tf.keras.layers.Input(shape=(128,), dtype=tf.int32, name="input_ids")
attention_mask = tf.keras.layers.Input(shape=(128,), dtype=tf.int32, name="attention_mask")

# 加载预训练Roberta并关联输入
roberta_model = TFAutoModel.from_pretrained("roberta-base")
x = roberta_model([input_ids, attention_mask])[0]
# 后续添加分类/回归输出层
output_layer = tf.keras.layers.Dense(1, activation="sigmoid")(x[:, 0, :])

# 绑定输入输出,明确模型需要2个输入
model = tf.keras.Model(inputs=[input_ids, attention_mask], outputs=output_layer)

2. 修正map_dataset函数输出

确保map函数返回**(输入字典/列表, 标签)**的结构,而非单一张量:

# 错误示例:仅返回单个输入张量
def map_dataset(example):
    return example["input_ids"]  # 缺少attention_mask和标签

# 正确示例1:返回字典(键名需与模型输入层name完全匹配)
def map_dataset(example):
    return {
        "input_ids": tf.cast(example["input_ids"], tf.int32),  # 转int32,匹配模型输入类型
        "attention_mask": tf.cast(example["attention_mask"], tf.int32)
    }, example["label"]

# 正确示例2:返回输入列表+标签
def map_dataset(example):
    return [tf.cast(example["input_ids"], tf.int32), tf.cast(example["attention_mask"], tf.int32)], example["label"]

注意:你的报错中输入是float64类型,而Roberta要求int32,必须通过tf.cast转换类型。

3. 验证数据集输出结构

训练前打印数据集的一个批次,确认输入格式是否符合要求:

for inputs, labels in train_dataset.take(1):
    print("输入结构:", inputs)
    print("输入数量:", len(inputs) if isinstance(inputs, (list, dict)) else 1)
    print("输入类型:", inputs["input_ids"].dtype if isinstance(inputs, dict) else inputs[0].dtype)

如果输出的输入是单一张量,说明map函数未正确处理;如果是字典,需确认键名与模型输入层name完全一致(大小写、拼写均不能错)。

完整可运行代码示例

import tensorflow as tf
from transformers import TFAutoModel, AutoTokenizer

# 1. 加载tokenizer与构建模型
tokenizer = AutoTokenizer.from_pretrained("roberta-base")
def build_roberta_model():
    input_ids = tf.keras.layers.Input(shape=(128,), dtype=tf.int32, name="input_ids")
    attention_mask = tf.keras.layers.Input(shape=(128,), dtype=tf.int32, name="attention_mask")
    
    roberta = TFAutoModel.from_pretrained("roberta-base")
    roberta_output = roberta([input_ids, attention_mask])[0]
    cls_token_output = roberta_output[:, 0, :]  # 取序列的<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]>token输出
    final_output = tf.keras.layers.Dense(2, activation="softmax")(cls_token_output)
    
    model = tf.keras.Model(inputs=[input_ids, attention_mask], outputs=final_output)
    model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=2e-5),
                  loss="sparse_categorical_crossentropy",
                  metrics=["accuracy"])
    return model

# 2. 构建并预处理数据集
# 假设texts是你的文本数据,labels是对应的标签
texts = ["sample text 1", "sample text 2", ...]
labels = [0, 1, ...]

# 文本tokenization
tokenized_data = tokenizer(texts, max_length=128, padding="max_length", truncation=True, return_tensors="tf")

# 转为tf.data.Dataset
dataset = tf.data.Dataset.from_tensor_slices({
    "input_ids": tokenized_data["input_ids"],
    "attention_mask": tokenized_data["attention_mask"],
    "label": labels
})

# 预处理函数
def preprocess(example):
    return {
        "input_ids": example["input_ids"],
        "attention_mask": example["attention_mask"]
    }, example["label"]

train_dataset = dataset.map(preprocess).batch(16).shuffle(1000)

# 3. 训练模型
model = build_roberta_model()
model.fit(train_dataset, epochs=3)

Transformer新手必注意事项

  • 输入类型严格匹配:Roberta的input_ids和attention_mask必须是int32类型,禁止使用float64(你的报错中已出现此问题)
  • 输入键名完全一致:模型输入层的name必须与数据集返回的字典键名完全匹配,TensorFlow会通过键名映射输入
  • 数据集输出格式规范:多输入模型的数据集必须返回(输入集合, 标签)的结构,不能只返回输入或标签
  • 预训练模型输入方式:调用预训练Roberta时,需将多个输入以列表形式传入,不能单独传单个张量
  • 训练前验证数据:务必打印数据集的一个批次,确认输入数量、类型、结构均符合模型要求

内容的提问来源于stack exchange,提问作者DeadSec

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 20:55:23