基于SentenceTransformer句嵌入构建神经网络时的张量转换失败问题求助
Hey there, I totally get stuck on this kind of tensor conversion issue when feeding embeddings to Keras models—let’s break down what’s going wrong and fix it step by step.
What’s Causing the Error?
Your X_train right now is an object-type NumPy array filled with individual smaller NumPy arrays (run print(X_train.dtype) and you’ll see object). TensorFlow can’t handle this nested structure—it expects a single 2D NumPy array where each row is one sample’s embedding vector (shape like (number_of_samples, embedding_dimension)).
When you do df_embed['Embeddings'].values, you’re getting an array where each element is a separate embedding array, not a unified 2D matrix. That’s why the conversion fails.
How to Fix It
We just need to convert that object array into a proper 2D NumPy array. Here’s the adjusted code with key fixes:
title_list = df.Title.tolist() model = SentenceTransformer('paraphrase-distilroberta-base-v1') embeddings = model.encode(title_list) # Fixed the typo here (you had embeddings_ex before) # Avoid modifying the original DataFrame with .copy() df_embed = df.copy() # Directly store embeddings as a list of arrays in the DataFrame df_embed['Embeddings'] = list(embeddings) # Critical step: Convert the Embeddings column into a 2D NumPy array X = np.vstack(df_embed['Embeddings'].values) y = df_embed.Tags mlb = MultiLabelBinarizer(classes=top_tags) y_mlb = pd.DataFrame(mlb.fit_transform(y), columns=mlb.classes_, index=y.index) from sklearn.model_selection import train_test_split X_train, X_val, y_train, y_val = train_test_split(X, y_mlb, test_size=0.3, random_state=0) X_val, X_test, y_val, y_test = train_test_split(X_val, y_val, test_size=0.4, random_state=0) # Build the model with explicit input shape (helps avoid confusion) model = Sequential() model.add(Dense(100, activation="relu", input_shape=(X_train.shape[1],))) model.add(Dropout(0.3, noise_shape=None, seed=None)) model.add(Dense(50, activation="sigmoid")) model.compile(loss='binary_crossentropy', optimizer=Adam(0.01), metrics=['accuracy']) # Use validation_data instead of validation_split since we already split our data hist = model.fit(X_train, y_train, batch_size=8, epochs=10, validation_data=(X_val, y_val))
Quick Notes
- I fixed a typo in your code: you had
embeddings_exinstead ofembeddingswhen creatingembeddings_list—that would have caused an error too! - Adding
input_shapeto the first Dense layer makes the model’s input requirements explicit, which helps catch dimension mismatches early. - You were using both
validation_split=0.1and pre-splitting validation data—this would have led to double-splitting your data. Usingvalidation_datawith your pre-splitX_val/y_valis cleaner here.
After these changes, X_train will be a standard 2D NumPy array that TensorFlow can convert to a Tensor without issues.
内容的提问来源于stack exchange,提问作者Célia Bayet

