Tensorflow RNN聊天机器人训练报错:形状不兼容问题求助
基于TensorFlow的RNN聊天机器人训练报错排查
问题背景
我用TensorFlow搭建基于循环神经网络(RNN)的通用聊天机器人,数据集为问答对格式,训练时出现形状不兼容的ValueError,请求排查问题。
数据集格式
questions = (("How are you?"), ("Who are you?")) answers = (("I'm fine, thank you"), ("I'm Mr. Nobody"))
数据集划分代码
self.train_x, self.test_x, self.train_y, self.test_y = train_test_split( self.questions, self.answers, test_size=0.2, random_state=42)
数据预处理代码
def preprocess_data(self): # Tokenize texts t = Tokenizer() t.fit_on_texts(self.train_x) # Convert text to sequences of integers train_questions = t.texts_to_sequences(self.train_x) test_questions = t.texts_to_sequences(self.test_x) train_answers = t.texts_to_sequences(self.train_y) test_answers = t.texts_to_sequences(self.test_y) # Pad sequences self.train_x = pad_sequences(train_questions, maxlen=self.max_sequences_length) self.test_x = pad_sequences(test_questions, maxlen=self.max_sequences_length) self.train_y = pad_sequences(train_answers, maxlen=self.max_sequences_length) self.test_y = pad_sequences(test_answers, maxlen=self.max_sequences_length)
RNN模型结构代码
def build_model(self): model = keras.models.Sequential() model.add(keras.layers.Embedding(10000, 128, input_length=self.max_sequences_length)) model.add(keras.layers.Dropout(0.2)) model.add(keras.layers.SimpleRNN(64, return_sequences=True)) model.add(keras.layers.Dense(len(self.vocabulary), activation="sigmoid")) return model
训练函数代码
def train_model(): self.preprocess_data() model = self.build_model() model.compile(optimizer="rmsprop", loss="categorical_crossentropy", metrics=["accuracy"]) model.fit(self.train_x, self.train_y, epochs=5, batch_size=32)
报错信息
Traceback (most recent call last): File "/mnt/c/Users/matte/Coding/projects/ai-projects/general-chatbot/chatbot.py", line 183, in <module> main() File "/mnt/c/Users/matte/Coding/projects/ai-projects/general-chatbot/chatbot.py", line 138, in main neural_network.train_model() File "/mnt/c/Users/matte/Coding/projects/ai-projects/general-chatbot/chatbot.py", line 115, in train_model model.fit(self.train_x, self.train_y, epochs=5, batch_size=32) File "/home/breadpit/.local/lib/python3.10/site-packages/keras/utils/traceback_utils.py", line 70, in error_handler raise e.with_traceback(filtered_tb) from None File "/tmp/__autograph_generated_file6oeizk3o.py", line 15, in tf__train_function retval_ = ag__.converted_call(ag__.ld(step_function), (ag__.ld(self), ag__.ld(iterator)), None, fscope) ValueError: in user code: File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1284, in train_function * return step_function(self, iterator) File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1268, in step_function ** outputs = model.distribute_strategy.run(run_step, args=(data,)) File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1249, in run_step ** outputs = model.train_step(data) File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1051, in train_step loss = self.compute_loss(x, y, y_pred, sample_weight) File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/training.py", line 1109, in compute_loss return self.compiled_loss( File "/home/breadpit/.local/lib/python3.10/site-packages/keras/engine/compile_utils.py", line 265, in __call__ loss_value = loss_obj(y_t, y_p, sample_weight=sw) File "/home/breadpit/.local/lib/python3.10/site-packages/keras/losses.py", line 142, in __call__ losses = call_fn(y_true, y_pred) File "/home/breadpit/.local/lib/python3.10/site-packages/keras/losses.py", line 268, in call ** return ag_fn(y_true, y_pred, **self._fn_kwargs) File "/home/breadpit/.local/lib/python3.10/site-packages/keras/losses.py", line 1984, in categorical_crossentropy return backend.categorical_crossentropy( File "/home/breadpit/.local/lib/python3.10/site-packages/keras/backend.py", line 5559, in categorical_crossentropy target.shape.assert_is_compatible_with(output.shape) ValueError: Shapes (None, 64) and (None, 64, 2365) are incompatible
数据集形状
train_x shape: (2980, 64) train_y shape: (2980, 64) test_x shape: (745, 64) test_y shape: (745, 64)
问题原因与解决方法
核心问题
- 输出形状不匹配:模型因
SimpleRNN(return_sequences=True)输出三维张量(None, 64, 2365),但标签train_y是二维整数序列(None, 64),且categorical_crossentropy要求标签为one-hot编码的三维张量,两者形状不兼容。 - Tokenizer范围缺失:仅在训练集问题上拟合Tokenizer,未包含答案文本,会导致答案中部分词汇无法被转换为整数序列。
- 损失函数与标签类型不匹配:当前标签是整数序列,不是one-hot编码,使用
categorical_crossentropy不合适。
具体修复步骤
1. 修正Tokenizer拟合范围
将训练集的问题和答案文本都加入Tokenizer拟合,确保覆盖所有词汇:
def preprocess_data(self): # 同时拟合问题和答案的训练数据 t = Tokenizer() t.fit_on_texts(self.train_x + self.train_y) self.tokenizer = t self.vocabulary_size = len(t.word_index) + 1 # 词汇表索引从1开始,+1补全0位 # 转换文本为整数序列 train_questions = t.texts_to_sequences(self.train_x) test_questions = t.texts_to_sequences(self.test_x) train_answers = t.texts_to_sequences(self.train_y) test_answers = t.texts_to_sequences(self.test_y) # 填充序列到统一长度 self.train_x = pad_sequences(train_questions, maxlen=self.max_sequences_length) self.test_x = pad_sequences(test_questions, maxlen=self.max_sequences_length) self.train_y = pad_sequences(train_answers, maxlen=self.max_sequences_length) self.test_y = pad_sequences(test_answers, maxlen=self.max_sequences_length)
2. 转换标签为one-hot编码
将整数序列标签转换为三维one-hot数组,匹配模型输出形状:
from tensorflow.keras.utils import to_categorical def preprocess_data(self): # ... 保留上述预处理代码 ... # 将标签转换为one-hot编码 self.train_y = to_categorical(self.train_y, num_classes=self.vocabulary_size) self.test_y = to_categorical(self.test_y, num_classes=self.vocabulary_size)
转换后train_y形状变为(2980, 64, vocabulary_size),与模型输出形状一致。
3. 调整模型激活函数
最后一层激活函数改为softmax,适配多分类任务的概率分布要求:
def build_model(self): model = keras.models.Sequential() model.add(keras.layers.Embedding(self.vocabulary_size, 128, input_length=self.max_sequences_length)) model.add(keras.layers.Dropout(0.2)) model.add(keras.layers.SimpleRNN(64, return_sequences=True)) model.add(keras.layers.Dense(self.vocabulary_size, activation="softmax")) # 替换为softmax return model
4. 修正训练函数缩进
原训练函数缺少self参数,修正为类的方法格式:
def train_model(self): self.preprocess_data() model = self.build_model() model.compile(optimizer="rmsprop", loss="categorical_crossentropy", metrics=["accuracy"]) model.fit(self.train_x, self.train_y, epochs=5, batch_size=32)
内容的提问来源于stack exchange,提问作者user17348736
相关产品推荐
相关产品推荐

