调用model.fit时遇NumPy转Tensor错误求助
问题分析与解决
核心错误原因
- 输入数据类型不匹配:
tokenizer.texts_to_sequences()返回的X是不等长的嵌套列表,TensorFlow无法直接将其转换为张量——这就是你看到Failed to convert a NumPy array to a Tensor (Unsupported object type list)报错的直接原因。你已经用pad_sequences(X)生成了统一长度的数组Y,但训练时却错误传入了未处理的X。 - 损失函数与标签不匹配:你使用
categorical_crossentropy作为损失函数,但传入的标签X是整数索引序列,该损失函数要求标签是one-hot编码格式。 - 训练逻辑偏差:文本生成任务的正确逻辑是用前k个词预测第k+1个词,而非直接将整个序列同时作为输入和标签。
修正后的代码
# Importing the dataset import pandas as pd import re import pickle from tensorflow.keras.preprocessing.text import Tokenizer from tensorflow.keras.preprocessing.sequence import pad_sequences from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Embedding, Bidirectional, LSTM, Dense from tensorflow.keras.regularizers import regularizers from tensorflow.keras.optimizers import Adam from tensorflow.keras.utils import to_categorical filename = "MoviePlots.csv" data = pd.read_csv(filename, encoding= 'unicode_escape') # Keeping only the necessary columns data = data[['Plot']] # Keep only rows where 'Plot' is a string data = data[data['Plot'].apply(lambda x: isinstance(x, str))] # Clean the data data['Plot'] = data['Plot'].apply(lambda x: x.lower()) data['Plot'] = data['Plot'].apply((lambda x: re.sub('[^a-zA-z0-9\s]', '', x))) # Create the tokenizer tokenizer = Tokenizer(num_words=5000, split=" ") tokenizer.fit_on_texts(data['Plot'].values) # Save the tokenizer with open('tokenizer.pickle', 'wb') as handle: pickle.dump(tokenizer, handle, protocol=pickle.HIGHEST_PROTOCOL) # Create the sequences and pad them to fixed length sequences = tokenizer.texts_to_sequences(data['Plot'].values) padded_sequences = pad_sequences(sequences, maxlen=None) # maxlen=None自动取最长序列长度 # 构建训练输入和标签:用前n-1个词预测第n个词 X_train = padded_sequences[:, :-1] y_train = padded_sequences[:, -1] # 对标签做one-hot编码,匹配categorical_crossentropy要求 y_train = to_categorical(y_train, num_classes=5000) # Create the model model = Sequential() # 注意input_length改为X_train的列数(即序列长度-1) model.add(Embedding(5000, 256, input_length=X_train.shape[1])) model.add(Bidirectional(LSTM(256, return_sequences=True, dropout=0.1, recurrent_dropout=0.1))) model.add(LSTM(256, return_sequences=True, dropout=0.1, recurrent_dropout=0.1)) model.add(LSTM(256, dropout=0.1, recurrent_dropout=0.1)) model.add(Dense(256, activation='relu', kernel_regularizer=regularizers.l2(0.01))) model.add(Dense(5000, activation='softmax')) # Compile the model model.compile(loss='categorical_crossentropy', optimizer=Adam(lr=0.01), metrics=['accuracy']) # Train the model model.fit(X_train, y_train, epochs=500, batch_size=256, verbose=1)
关键调整点说明
- 替换输入数据:用
padded_sequences(原代码中的Y)替代未padding的X,确保输入是统一长度的NumPy数组。 - 重构训练数据结构:将每个序列拆分为“输入序列(前n-1个词)”和“目标标签(第n个词)”,符合文本生成的任务逻辑。
- 标签编码转换:用
to_categorical()将整数标签转为one-hot格式,适配categorical_crossentropy损失函数。 - 修正模型输入长度:
Embedding层的input_length改为X_train.shape[1],与输入序列长度匹配。
内容的提问来源于stack exchange,提问作者SIDHANT YADAV
相关产品推荐
相关产品推荐

