调用model.fit()抛出ValueError张量转换错误,求排查原因
问题:调用model.fit()时出现ValueError的原因及解决方法
问题背景与代码重现
先通过BERT生成嵌入矩阵,再以此为初始权重构建LSTM模型,调用model.fit()时触发类型转换错误。
生成BERT嵌入矩阵的代码
# Load pre-trained model tokenizer and model tokenizer = BertTokenizer.from_pretrained('bert-base-multilingual-cased') model = BertModel.from_pretrained('bert-base-multilingual-cased') # Define batch size batch_size = 1 # Tokenize and encode input data in batches encoded_inputs = [] for i in range(0, len(labeled_data), batch_size): inputs = labeled_data[i:i+batch_size] encoded_inputs.append(tokenizer.batch_encode_plus(inputs, padding=True, truncation=True, return_tensors="pt")) # Generate embeddings for each batch embeddings_new = [] for encoded_input in tqdm(encoded_inputs): with torch.no_grad(): model_output = model(**encoded_input) batch_embeddings = model_output.last_hidden_state.mean(dim=1) embeddings_new.append(batch_embeddings) embeddings_new = tf.concat(embeddings_new, axis=0) embedding_matrix = model.embeddings.word_embeddings.weight embedding_matrix = embedding_matrix.cpu().detach().numpy() embed_tensor = tf.convert_to_tensor(embedding_matrix, dtype=tf.float32)
构建LSTM模型的代码
lstm_out1 = 150 embed_dim = 768 model = Sequential() model.add(Embedding(embedding_matrix.shape[0], embed_dim, weights=[embed_tensor], input_length=50, trainable=False)) model.add(LSTM(lstm_out1, dropout=0.2, recurrent_dropout=0.2)) model.add(Dense(64, activation='relu')) model.add(Dense(1, activation='sigmoid')) adam = Adam(lr=0.001, beta_1=0.9, beta_2=0.999, epsilon=1e-08, decay=0.0) model.compile(loss='binary_crossentropy', optimizer=adam, metrics=['accuracy']) model.summary()
触发错误的调用与错误信息
调用代码:
model.fit(tokenized_sentences, labels, batch_size=5, epochs=1, shuffle=True)
错误信息:
ValueError: Failed to convert a NumPy array to a Tensor (Unsupported object type list).
输入数据说明
tokenized_sentences是二维整数列表,由以下代码生成:
tokenized_sentences = [] for sentence in labeled_data: # Apply the tokenizer to each sentence to obtain its tokens tokens = tokenizer.encode(sentence, add_special_tokens=True) # Append the tokenized sentence to the list tokenized_sentences.append(tokens)
示例内容:[[101, 10406, 10161, ..., 102], ...]
labels是一维整数列表,示例内容:[1, 0, 1, 1, 0, ..., 0]
错误原因
- 序列长度不统一:
tokenized_sentences中每个子列表(单句token序列)的长度不一致,但LSTM模型的Embedding层设置了input_length=50,要求输入必须是固定长度的规则张量。TensorFlow无法将长度参差不齐的嵌套列表直接转换为符合要求的张量,因此抛出类型转换错误。 - 数据格式不符合要求:嵌套列表不属于TensorFlow默认支持的输入格式,需要先转换为固定长度的NumPy数组或TensorFlow张量。
解决方法
方法一:tokenize阶段直接生成固定长度序列
修改tokenize代码,利用BERT tokenizer的参数直接生成符合input_length要求的序列:
tokenized_sentences = [] max_len = 50 # 与Embedding层的input_length保持一致 for sentence in labeled_data: tokens = tokenizer.encode( sentence, add_special_tokens=True, max_length=max_len, padding='max_length', truncation=True ) tokenized_sentences.append(tokens) # 转换为NumPy数组 import numpy as np tokenized_sentences = np.array(tokenized_sentences) labels = np.array(labels)
方法二:手动补全/截断已有序列
如果已经生成了tokenized_sentences,可以手动调整为固定长度:
max_len = 50 processed_sentences = [] for tokens in tokenized_sentences: if len(tokens) < max_len: # 补零至指定长度 padded_tokens = tokens + [0]*(max_len - len(tokens)) else: # 截断至指定长度 padded_tokens = tokens[:max_len] processed_sentences.append(padded_tokens) # 转换为数组 tokenized_sentences = np.array(processed_sentences) labels = np.array(labels)
方法三:使用Keras内置工具处理序列
利用pad_sequences快速统一序列长度:
from tensorflow.keras.preprocessing.sequence import pad_sequences max_len = 50 tokenized_sentences = pad_sequences( tokenized_sentences, maxlen=max_len, padding='post', truncating='post' ) labels = np.array(labels)
完成上述处理后,再调用model.fit()即可正常运行。
内容的提问来源于stack exchange,提问作者Debbie
相关产品推荐
相关产品推荐

