Keras命令匹配模型推理遇维度错误:Found array with dim 3
命令匹配模型运行错误分析与修复
错误信息
Started 2023-09-06 13:44:05.970618: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX AVX2 To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags. Enter an command: check connection 1/1 [==============================] - 0s 327ms/step 1/1 [==============================] - 0s 26ms/step 1/1 [==============================] - 0s 31ms/step 1/1 [==============================] - 0s 16ms/step 1/1 [==============================] - 0s 31ms/step 1/1 [==============================] - 0s 31ms/step 1/1 [==============================] - 0s 47ms/step Traceback (most recent call last): File "D:\Advanced Robot\commands\test.py", line 32, in <module> similarity_scores = cosine_similarity(user_input_embedding, command_embeddings) File "C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\metrics\pairwise.py", line 1393, in cosine_similarity X, Y = check_pairwise_arrays(X, Y) File "C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\metrics\pairwise.py", line 163, in check_pairwise_arrays Y = check_array( File "C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\utils\validation.py", line 915, in check_array raise ValueError( ValueError: Found array with dim 3. check_pairwise_arrays expected <= 2.
核心问题分析
错误存在两个关键根源:
- train.py模型设计偏离需求:原模型是二分类任务架构(
Dense(1, sigmoid)+binary_crossentropy损失),但实际需求是文本相似匹配,需要能生成文本特征向量的特征提取模型,而非分类模型。 - main.py两处致命错误:
- 未加载训练时的Tokenizer词汇表,导致文本转序列时无法映射训练过的单词
- 模型输出维度不匹配,
cosine_similarity要求输入为2维数组,但当前输出是3维,无法计算相似度
修复步骤
1. 修正train.py(重新训练特征提取模型)
调整模型结构为特征提取型,同时保存Tokenizer配置供main.py复用:
print("Started") import tensorflow as tf import numpy as np from tensorflow.keras.preprocessing.text import Tokenizer from tensorflow.keras.preprocessing.sequence import pad_sequences import json # 加载命令文本 with open('commands.txt', 'r') as command_file: text_data = command_file.read().splitlines() vocab_size = 1000 embedding_dim = 16 num_epochs = 50 batch_size = 32 max_sequence_length = 100 # 初始化Tokenizer并训练词汇表 tokenizer = Tokenizer(num_words=vocab_size) tokenizer.fit_on_texts(text_data) sequences = tokenizer.texts_to_sequences(text_data) paded_sequences = pad_sequences(sequences, maxlen=max_sequence_length, padding='post', truncating='post') # 保存Tokenizer配置,避免main.py重新训练词汇表 with open('tokenizer_config.json', 'w') as f: json.dump(tokenizer.to_json(), f) # 构建特征提取模型:去掉分类层,用LSTM直接输出64维特征向量 model = tf.keras.Sequential([ tf.keras.layers.Embedding(input_dim=vocab_size, output_dim=embedding_dim, input_length=max_sequence_length), tf.keras.layers.LSTM(units=64) ]) # 用自监督方式训练:让模型学习命令自身的特征表示 model.compile(loss='mse', optimizer='adam', metrics=['mae']) model.fit(paded_sequences, paded_sequences, epochs=num_epochs, batch_size=batch_size) print("Saving model and tokenizer....") model.save("commands_model.h5") print("Model saved")
2. 修正main.py
加载训练好的Tokenizer,调整输出维度适配相似度计算:
print("Started") import tensorflow as ten from tensorflow.keras.preprocessing.text import tokenizer_from_json from tensorflow.keras.preprocessing.sequence import pad_sequences from sklearn.metrics.pairwise import cosine_similarity import numpy as np import json # 加载训练时保存的Tokenizer配置 with open('tokenizer_config.json', 'r') as f: tokenizer = tokenizer_from_json(json.load(f)) loaded_model = ten.keras.models.load_model('commands_model.h5') max_sequence_length = 100 inp = input("Enter a command: ") user_input_sequence = tokenizer.texts_to_sequences([inp]) padded_user_input = pad_sequences(user_input_sequence, maxlen=max_sequence_length, padding='post', truncating='post') # 获取用户输入的特征向量,压缩多余维度至2维 user_input_embedding = loaded_model.predict(padded_user_input).squeeze() user_input_embedding = user_input_embedding.reshape(1, -1) command_embeddings = [] with open('commands.txt', 'r') as file: commands = file.read().splitlines() for command in commands: command_sequence = tokenizer.texts_to_sequences([command]) paded_command = pad_sequences(command_sequence, maxlen=max_sequence_length, padding='post', truncating='post') # 压缩每个命令特征的维度 command_embedding = loaded_model.predict(paded_command).squeeze() command_embeddings.append(command_embedding) command_embeddings = np.array(command_embeddings) similarity_scores = cosine_similarity(user_input_embedding, command_embeddings) most_similar_command_index = np.argmax(similarity_scores) most_similar_command = commands[most_similar_command_index] print("Most similar: ", most_similar_command)
额外说明
- 原train.py的
LabelEncoder和二分类损失完全不符合文本相似匹配需求,改为自监督训练(用命令自身作为目标)更贴合场景 - 必须保存并加载Tokenizer,否则main.py无法复用训练时的词汇映射关系,导致文本编码错误
squeeze()方法用于移除模型输出的多余维度,将(1,1,64)格式转换为(64),最终组合成(命令数量,64)的二维数组,满足cosine_similarity的输入要求
内容的提问来源于stack exchange,提问作者MaxMouse
相关产品推荐
相关产品推荐

