You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tensorflow RNN聊天机器人训练报错:形状不兼容问题求助

基于TensorFlow的RNN聊天机器人训练报错排查

问题背景

我用TensorFlow搭建基于循环神经网络(RNN)的通用聊天机器人,数据集为问答对格式,训练时出现形状不兼容的ValueError,请求排查问题。

数据集格式

questions = (("How are you?"), ("Who are you?"))
answers = (("I'm fine, thank you"), ("I'm Mr. Nobody"))

数据集划分代码

self.train_x, self.test_x, self.train_y, self.test_y = train_test_split(
            self.questions, self.answers, test_size=0.2, random_state=42)

数据预处理代码

def preprocess_data(self):
    # Tokenize texts
    t = Tokenizer()
    t.fit_on_texts(self.train_x)

    # Convert text to sequences of integers
    train_questions = t.texts_to_sequences(self.train_x)
    test_questions = t.texts_to_sequences(self.test_x)

    train_answers = t.texts_to_sequences(self.train_y)
    test_answers = t.texts_to_sequences(self.test_y)
        
    # Pad sequences
    self.train_x = pad_sequences(train_questions, maxlen=self.max_sequences_length)
    self.test_x = pad_sequences(test_questions, maxlen=self.max_sequences_length)

    self.train_y = pad_sequences(train_answers, maxlen=self.max_sequences_length)
    self.test_y = pad_sequences(test_answers, maxlen=self.max_sequences_length)

RNN模型结构代码

def build_model(self):
    model = keras.models.Sequential()
    model.add(keras.layers.Embedding(10000, 128, input_length=self.max_sequences_length))
    model.add(keras.layers.Dropout(0.2))
    model.add(keras.layers.SimpleRNN(64, return_sequences=True))
    model.add(keras.layers.Dense(len(self.vocabulary), activation="sigmoid"))

    return model

训练函数代码

def train_model():
     self.preprocess_data()
     model = self.build_model()

     model.compile(optimizer="rmsprop", loss="categorical_crossentropy", metrics=["accuracy"])
     model.fit(self.train_x, self.train_y, epochs=5, batch_size=32)

报错信息

Traceback (most recent call last):
  File "/mnt/c/Users/matte/Coding/projects/ai-projects/general-chatbot/chatbot.py", line 183, in <module>
    main()
  File "/mnt/c/Users/matte/Coding/projects/ai-projects/general-chatbot/chatbot.py", line 138, in main
    neural_network.train_model()
  File "/mnt/c/Users/matte/Coding/projects/ai-projects/general-chatbot/chatbot.py", line 115, in train_model
    model.fit(self.train_x, self.train_y, epochs=5, batch_size=32)
  File "/home/breadpit/.local/lib/python3.10/site-packages/keras/utils/traceback_utils.py", line 70, in error_handler
    raise e.with_traceback(filtered_tb) from None
  File "/tmp/__autograph_generated_file6oeizk3o.py", line 15, in tf__train_function
    retval_ = ag__.converted_call(ag__.ld(step_function), (ag__.ld(self), ag__.ld(iterator)), None, fscope)
ValueError: in user code:

    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1284, in train_function  *
        return step_function(self, iterator)
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1268, in step_function  **
        outputs = model.distribute_strategy.run(run_step, args=(data,))
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1249, in run_step  **
        outputs = model.train_step(data)
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1051, in train_step
        loss = self.compute_loss(x, y, y_pred, sample_weight)
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1109, in compute_loss
        return self.compiled_loss(
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/compile_utils.py", line 265, in __call__
        loss_value = loss_obj(y_t, y_p, sample_weight=sw)
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/losses.py", line 142, in __call__
        losses = call_fn(y_true, y_pred)
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/losses.py", line 268, in call  **
        return ag_fn(y_true, y_pred, **self._fn_kwargs)
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/losses.py", line 1984, in categorical_crossentropy
        return backend.categorical_crossentropy(
    File "/home/breadpit/.local/lib/python3.10/site-packages/keras/backend.py", line 5559, in categorical_crossentropy
        target.shape.assert_is_compatible_with(output.shape)

    ValueError: Shapes (None, 64) and (None, 64, 2365) are incompatible

数据集形状

train_x shape: (2980, 64)
train_y shape: (2980, 64)
test_x shape: (745, 64)
test_y shape: (745, 64)

问题原因与解决方法

核心问题

  1. 输出形状不匹配:模型因SimpleRNN(return_sequences=True)输出三维张量(None, 64, 2365),但标签train_y是二维整数序列(None, 64),且categorical_crossentropy要求标签为one-hot编码的三维张量,两者形状不兼容。
  2. Tokenizer范围缺失:仅在训练集问题上拟合Tokenizer,未包含答案文本,会导致答案中部分词汇无法被转换为整数序列。
  3. 损失函数与标签类型不匹配:当前标签是整数序列,不是one-hot编码,使用categorical_crossentropy不合适。

具体修复步骤

1. 修正Tokenizer拟合范围

将训练集的问题和答案文本都加入Tokenizer拟合,确保覆盖所有词汇:

def preprocess_data(self):
    # 同时拟合问题和答案的训练数据
    t = Tokenizer()
    t.fit_on_texts(self.train_x + self.train_y)
    self.tokenizer = t
    self.vocabulary_size = len(t.word_index) + 1  # 词汇表索引从1开始,+1补全0位

    # 转换文本为整数序列
    train_questions = t.texts_to_sequences(self.train_x)
    test_questions = t.texts_to_sequences(self.test_x)

    train_answers = t.texts_to_sequences(self.train_y)
    test_answers = t.texts_to_sequences(self.test_y)
        
    # 填充序列到统一长度
    self.train_x = pad_sequences(train_questions, maxlen=self.max_sequences_length)
    self.test_x = pad_sequences(test_questions, maxlen=self.max_sequences_length)

    self.train_y = pad_sequences(train_answers, maxlen=self.max_sequences_length)
    self.test_y = pad_sequences(test_answers, maxlen=self.max_sequences_length)

2. 转换标签为one-hot编码

将整数序列标签转换为三维one-hot数组,匹配模型输出形状:

from tensorflow.keras.utils import to_categorical

def preprocess_data(self):
    # ... 保留上述预处理代码 ...

    # 将标签转换为one-hot编码
    self.train_y = to_categorical(self.train_y, num_classes=self.vocabulary_size)
    self.test_y = to_categorical(self.test_y, num_classes=self.vocabulary_size)

转换后train_y形状变为(2980, 64, vocabulary_size),与模型输出形状一致。

3. 调整模型激活函数

最后一层激活函数改为softmax,适配多分类任务的概率分布要求:

def build_model(self):
    model = keras.models.Sequential()
    model.add(keras.layers.Embedding(self.vocabulary_size, 128, input_length=self.max_sequences_length))
    model.add(keras.layers.Dropout(0.2))
    model.add(keras.layers.SimpleRNN(64, return_sequences=True))
    model.add(keras.layers.Dense(self.vocabulary_size, activation="softmax"))  # 替换为softmax

    return model

4. 修正训练函数缩进

原训练函数缺少self参数,修正为类的方法格式:

def train_model(self):
     self.preprocess_data()
     model = self.build_model()

     model.compile(optimizer="rmsprop", loss="categorical_crossentropy", metrics=["accuracy"])
     model.fit(self.train_x, self.train_y, epochs=5, batch_size=32)

内容的提问来源于stack exchange,提问作者user17348736

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 09:14:59