You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python中NLTK训练数组的ValueError(形状不均)

解决AI聊天机器人代码中的ValueError及其他问题

错误原因分析

你遇到的ValueError: setting an array element with a sequence错误,主要由以下几个问题导致:

  • 分类数据保存错误:代码中将words错误写入classes.pkl文件,而非classes列表,会导致后续数据匹配异常。
  • 数组转换与模型拟合语法错误:model.fit中使用了错误的数组转换语法np.array[train_x],正确写法应为np.array(train_x);同时直接将嵌套列表转为numpy数组时,若子元素长度不一致也会触发报错。
  • 潜在版本兼容问题:新版本TensorFlow/Keras中,SGD的参数lr已更名为learning_rate,旧写法会触发警告或错误。

修复步骤

1. 修正分类数据保存

将保存classes.pkl的代码行修改为:

pickle.dump(classes, open('classes.pkl', 'wb'))

2. 避免不均匀数组转换问题

拆分training列表时直接提取数据,跳过numpy数组转换步骤,避免形状不匹配报错:

# 替换原有的training数组转换和拆分代码
random.shuffle(training)
train_x = [item[0] for item in training]
train_y = [item[1] for item in training]

3. 修正模型拟合语法

将model.fit中的错误语法修正:

model.fit(np.array(train_x), np.array(train_y), epochs=200, batch_size=5, verbose=1)

4. 适配新版本Keras参数(可选)

若使用TensorFlow 2.6+版本,更新SGD参数写法:

sgd = SGD(learning_rate=0.01, decay=1e-6, momentum=0.9, nesterov=True)

修正后的完整代码

import random
import json
import pickle
import numpy as np
import nltk
import tensorflow
from nltk.stem import WordNetLemmatizer
from keras.models import Sequential
from keras.layers import Dense, Activation, Dropout
from keras.optimizers import SGD

lemmatizer = WordNetLemmatizer()
intents = json.loads(open("X:/Memory/intents.json").read())
words = []
classes = []
documents = []
ignore_letters = ["!", "?", ".", ","]

for intent in intents['intents']:
    for pattern in intent['patterns']:
        word_list = nltk.word_tokenize(pattern)
        words.extend(word_list)
        documents.append((word_list, intent['tag']))
        if intent['tag'] not in classes:
            classes.append(intent['tag'])

words = [lemmatizer.lemmatize(word) for word in words if word not in ignore_letters]
words = sorted(set(words))
pickle.dump(words, open('words.pkl', 'wb'))
pickle.dump(classes, open('classes.pkl', 'wb'))  # 修正分类数据保存

training = []
output_empty = [0] * len(classes)
for document in documents:
    bag = []
    word_patterns = document[0]
    word_patterns = [lemmatizer.lemmatize(word.lower()) for word in word_patterns]
    for word in words:
        bag.append(1) if word in word_patterns else bag.append(0)

    output_row = list(output_empty)
    output_row[classes.index(document[1])] = 1
    training.append([bag, output_row])

random.shuffle(training)
# 直接拆分训练数据
train_x = [item[0] for item in training]
train_y = [item[1] for item in training]

model = Sequential()
model.add(Dense(128, input_shape=(len(train_x[0]),), activation='relu'))
model.add(Dropout(0.5))
model.add(Dense(64, activation='relu'))
model.add(Dropout(0.5))
model.add(Dense(len(train_y[0]), activation='softmax'))

# 适配新版本SGD参数
sgd = SGD(learning_rate=0.01, decay=1e-6, momentum=0.9, nesterov=True)
model.compile(loss='categorical_crossentropy', optimizer=sgd, metrics=['accuracy'])
# 修正模型拟合语法
model.fit(np.array(train_x), np.array(train_y), epochs=200, batch_size=5, verbose=1)
model.save('X:/Memory/ModelLearning.mem')

额外注意事项

  • 确保intents.json文件路径正确,且格式符合要求(每个intent需包含patterns和tag字段)。
  • 若运行时提示nltk资源缺失,需先下载必要数据:
nltk.download('punkt')
nltk.download('wordnet')

内容的提问来源于stack exchange,提问作者Nixfee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 19:40:33