You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras命令匹配模型推理遇维度错误:Found array with dim 3

命令匹配模型运行错误分析与修复

错误信息

Started
2023-09-06 13:44:05.970618: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations:  AVX AVX2
To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags.
Enter an command: check connection
1/1 [==============================] - 0s 327ms/step
1/1 [==============================] - 0s 26ms/step
1/1 [==============================] - 0s 31ms/step
1/1 [==============================] - 0s 16ms/step
1/1 [==============================] - 0s 31ms/step
1/1 [==============================] - 0s 31ms/step
1/1 [==============================] - 0s 47ms/step
Traceback (most recent call last):
  File "D:\Advanced Robot\commands\test.py", line 32, in <module>
    similarity_scores = cosine_similarity(user_input_embedding, command_embeddings)
  File "C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\metrics\pairwise.py", line 1393, in cosine_similarity
    X, Y = check_pairwise_arrays(X, Y)
  File "C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\metrics\pairwise.py", line 163, in check_pairwise_arrays
    Y = check_array(
  File "C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\utils\validation.py", line 915, in check_array
    raise ValueError(
ValueError: Found array with dim 3. check_pairwise_arrays expected <= 2.

核心问题分析

错误存在两个关键根源:

  1. train.py模型设计偏离需求:原模型是二分类任务架构(Dense(1, sigmoid)+binary_crossentropy损失),但实际需求是文本相似匹配,需要能生成文本特征向量的特征提取模型,而非分类模型。
  2. main.py两处致命错误:
    • 未加载训练时的Tokenizer词汇表,导致文本转序列时无法映射训练过的单词
    • 模型输出维度不匹配,cosine_similarity要求输入为2维数组,但当前输出是3维,无法计算相似度

修复步骤

1. 修正train.py(重新训练特征提取模型)

调整模型结构为特征提取型,同时保存Tokenizer配置供main.py复用:

print("Started")

import tensorflow as tf
import numpy as np
from tensorflow.keras.preprocessing.text import Tokenizer
from tensorflow.keras.preprocessing.sequence import pad_sequences
import json

# 加载命令文本
with open('commands.txt', 'r') as command_file:
    text_data = command_file.read().splitlines()

vocab_size = 1000
embedding_dim = 16
num_epochs = 50
batch_size = 32
max_sequence_length = 100

# 初始化Tokenizer并训练词汇表
tokenizer = Tokenizer(num_words=vocab_size)
tokenizer.fit_on_texts(text_data)
sequences = tokenizer.texts_to_sequences(text_data)
paded_sequences = pad_sequences(sequences, maxlen=max_sequence_length, padding='post', truncating='post')

# 保存Tokenizer配置,避免main.py重新训练词汇表
with open('tokenizer_config.json', 'w') as f:
    json.dump(tokenizer.to_json(), f)

# 构建特征提取模型:去掉分类层,用LSTM直接输出64维特征向量
model = tf.keras.Sequential([
    tf.keras.layers.Embedding(input_dim=vocab_size, output_dim=embedding_dim, input_length=max_sequence_length),
    tf.keras.layers.LSTM(units=64)
])

# 用自监督方式训练:让模型学习命令自身的特征表示
model.compile(loss='mse', optimizer='adam', metrics=['mae'])
model.fit(paded_sequences, paded_sequences, epochs=num_epochs, batch_size=batch_size)

print("Saving model and tokenizer....")
model.save("commands_model.h5")
print("Model saved")

2. 修正main.py

加载训练好的Tokenizer,调整输出维度适配相似度计算:

print("Started")
import tensorflow as ten
from tensorflow.keras.preprocessing.text import tokenizer_from_json
from tensorflow.keras.preprocessing.sequence import pad_sequences
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np
import json

# 加载训练时保存的Tokenizer配置
with open('tokenizer_config.json', 'r') as f:
    tokenizer = tokenizer_from_json(json.load(f))

loaded_model = ten.keras.models.load_model('commands_model.h5')
max_sequence_length = 100

inp = input("Enter a command: ")
user_input_sequence = tokenizer.texts_to_sequences([inp])
padded_user_input = pad_sequences(user_input_sequence, maxlen=max_sequence_length, padding='post', truncating='post')

# 获取用户输入的特征向量,压缩多余维度至2维
user_input_embedding = loaded_model.predict(padded_user_input).squeeze()
user_input_embedding = user_input_embedding.reshape(1, -1)

command_embeddings = []
with open('commands.txt', 'r') as file:
    commands = file.read().splitlines()

for command in commands:
    command_sequence = tokenizer.texts_to_sequences([command])
    paded_command = pad_sequences(command_sequence, maxlen=max_sequence_length, padding='post', truncating='post')
    # 压缩每个命令特征的维度
    command_embedding = loaded_model.predict(paded_command).squeeze()
    command_embeddings.append(command_embedding)

command_embeddings = np.array(command_embeddings)
similarity_scores = cosine_similarity(user_input_embedding, command_embeddings)
most_similar_command_index = np.argmax(similarity_scores)
most_similar_command = commands[most_similar_command_index]

print("Most similar: ", most_similar_command)

额外说明

  • 原train.py的LabelEncoder和二分类损失完全不符合文本相似匹配需求,改为自监督训练(用命令自身作为目标)更贴合场景
  • 必须保存并加载Tokenizer,否则main.py无法复用训练时的词汇映射关系,导致文本编码错误
  • squeeze()方法用于移除模型输出的多余维度,将(1,1,64)格式转换为(64),最终组合成(命令数量,64)的二维数组,满足cosine_similarity的输入要求

内容的提问来源于stack exchange,提问作者MaxMouse

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 12:24:57