多输入特征神经网络训练:Keras中文本序列与概率值拼接问题
Hey there! Let's break down your problem and fix it the right way—because forcing text sequences and numeric values into a single array isn't just causing type issues, it's also not the best approach for your Keras multi-input model.
First, why are you getting all floats?
NumPy arrays require uniform data types. When you append a float (your Probability value) to an integer array, NumPy automatically upcasts all elements to float to maintain consistency. That's expected behavior, but you don't need to do this at all for your classification task.
The Correct Approach: Multi-Input Model
Your text sequences and numeric Probability are two distinct feature types—text is sequential (needs RNN/CNN processing) while the probability is a standalone numeric feature. The right way to handle this is to build a model with two separate input branches that merge later. Here's how to adapt your existing code:
Step 1: Prepare Your Inputs Properly
First, finish processing your text sequences and reshape the numeric feature:
# Pad sequences (even though your sentences are fixed-length, this keeps things standardized) max_length = 5 padded_sequences = pad_sequences(sequences, maxlen=max_length, padding='post') # Convert your Probability (stv) to a 2D numpy array (required for Keras input) numeric_input = np.array(stv).reshape(-1, 1) # Split into train/test sets target = df['Target'].values X_train_text, X_test_text, X_train_num, X_test_num, y_train, y_test = train_test_split( padded_sequences, numeric_input, target, test_size=0.2, random_state=seed )
Step 2: Build the Multi-Input Model
Use Keras' Model class to create two branches and merge them:
from keras.models import Model from keras.layers import Input, concatenate, Dense, LSTM, Embedding # Text input branch: handles the sequence of words text_input = Input(shape=(max_length,), name="text_sequence") # Embedding layer (replace with pre-trained Word2Vec if you have it) embedding = Embedding( input_dim=len(tokenizer_obj.word_index) + 1, output_dim=EMBEDDING_DIM, input_length=max_length )(text_input) lstm_out = LSTM(128)(embedding) # You can use GRU instead if preferred # Numeric input branch: handles the Probability value num_input = Input(shape=(1,), name="probability") num_dense = Dense(32, activation="relu")(num_input) # Merge the two branches merged = concatenate([lstm_out, num_dense]) # Final classification layer (binary classification uses sigmoid) output = Dense(1, activation="sigmoid")(merged) # Define and compile the model model = Model(inputs=[text_input, num_input], outputs=output) model.compile(optimizer="adam", loss="binary_crossentropy", metrics=["accuracy"])
Step 3: Train the Model
Feed both inputs to the model during training:
model.fit( [X_train_text, X_train_num], y_train, validation_data=([X_test_text, X_test_num], y_test), epochs=10, batch_size=32 )
If You Really Need a Mixed-Type List (Not for Model Input)
If you just want to create a list of lists with integers and floats (not for feeding into Keras), use a list comprehension instead of NumPy:
combined_features = [seq + [stv[i]] for i, seq in enumerate(sequences)]
This gives you exactly the structure you wanted: [[2, 77, 20, 17, 81, 0.456...], ...]—but remember, this won't work directly with Keras models since they require uniform-type arrays.
Why This Matters
By keeping your features separate, you let each branch process the data in a way that makes sense:
- The RNN/LSTM can focus on the sequential relationships between words in your sentence.
- The numeric branch can handle the Probability value as a standalone signal without forcing it to fit into a text sequence structure.
This approach will give you better model performance than trying to cram two different feature types into one array.
内容的提问来源于stack exchange,提问作者Nayantara Jeyaraj

