如何在TensorFlow中传入两个不同长度数组作为单个训练单元?
Hey there! Let’s walk through how to tackle your TensorFlow setup, from handling your custom data structure to building that 69-layer neural network you’re aiming for.
First, let’s get your myLargeArray into a format TensorFlow can work with. Your array is structured as [[[42-length-array], [7-length-array]], ...]—so we need to extract the two input components separately and clean up their shapes.
Here’s how to do it with NumPy:
import numpy as np # Extract all 42-dimensional inputs input_42 = np.array([sample[0] for sample in myLargeArray]).squeeze() # After squeezing, shape will be (number_of_samples, 42) # Extract all 7-dimensional inputs input_7 = np.array([sample[1] for sample in myLargeArray]).squeeze() # Shape becomes (number_of_samples, 7)
Don’t forget you’ll also need your labels (the 7-length arrays or single numbers you want the model to predict). Make sure those are in a NumPy array with shape (number_of_samples, 7) or (number_of_samples, 1) depending on your output type.
You have two options for feeding these inputs into your model:
Option A: Combine Inputs into One Vector
The simplest approach is to concatenate the 42 and 7-dimensional arrays into a single 49-dimensional input:
combined_input = np.concatenate([input_42, input_7], axis=1) # Shape: (number_of_samples, 49)
This works great if you don’t need to treat the two input types differently in your model.
Option B: Keep Inputs Separate (Multi-Input Model)
If you want to process the 42 and 7-dimensional inputs through distinct branches of your network (e.g., different layer sizes for each), go with a multi-input model. We’ll cover this in the model building step.
Let’s cover both input approaches, plus a critical note on handling so many hidden layers.
For Combined Input (Single Vector)
Use a Sequential model and loop to add your 69 hidden layers. We’ll add batch normalization to help stabilize training (more on why later):
import tensorflow as tf from tensorflow.keras import layers, models model = models.Sequential() # Input layer matching our combined 49-dimensional input model.add(layers.Input(shape=(49,))) # Add 69 hidden layers for _ in range(69): model.add(layers.Dense(64, activation='relu')) # Batch norm prevents vanishing/exploding gradients with deep layers model.add(layers.BatchNormalization()) # Output layer: choose based on your goal # For 7-length array output (classification or multi-output regression) model.add(layers.Dense(7, activation='softmax')) # Use 'softmax' for classification, 'linear' for regression # For single number output (regression) # model.add(layers.Dense(1, activation='linear')) # Compile the model model.compile(optimizer='adam', loss='categorical_crossentropy') # Use 'mse' for regression, 'categorical_crossentropy' for classification
For Multi-Input (Separate Branches)
Build two parallel branches for each input type, then combine them before the final output:
# Define input layers for each input type input_42_layer = layers.Input(shape=(42,)) input_7_layer = layers.Input(shape=(7,)) # Process 42-dimensional input (34 hidden layers here) x1 = layers.Dense(64, activation='relu')(input_42_layer) x1 = layers.BatchNormalization()(x1) for _ in range(33): # 33 more layers to make 34 total x1 = layers.Dense(64, activation='relu')(x1) x1 = layers.BatchNormalization()(x1) # Process 7-dimensional input (34 hidden layers here) x2 = layers.Dense(32, activation='relu')(input_7_layer) x2 = layers.BatchNormalization()(x2) for _ in range(33): # 33 more layers to make 34 total x2 = layers.Dense(32, activation='relu')(x2) x2 = layers.BatchNormalization()(x2) # Combine the two branches combined = layers.concatenate([x1, x2]) # Add 1 more hidden layer to hit 69 total (34+34+1) combined = layers.Dense(64, activation='relu')(combined) combined = layers.BatchNormalization()(combined) # Output layer output = layers.Dense(7, activation='softmax')(combined) # Or Dense(1, activation='linear') # Build the multi-input model model = models.Model(inputs=[input_42_layer, input_7_layer], outputs=output) model.compile(optimizer='adam', loss='categorical_crossentropy')
Training is straightforward once your data is prepped:
For Combined Input
# Replace 'labels' with your actual label array model.fit(combined_input, labels, epochs=10, batch_size=32, validation_split=0.2) # Uses 20% of data for validation
For Multi-Input
model.fit([input_42, input_7], labels, epochs=10, batch_size=32, validation_split=0.2)
69 layers is extremely deep—you’ll likely run into vanishing or exploding gradients (where the model can’t learn because gradient signals get lost or blown up through the layers). To fix this:
- Add residual connections (skip connections) that add the input of a layer to its output. Here’s how to modify a Sequential layer for this:
for _ in range(69): # Capture the current layer output before adding new layers prev_output = model.layers[-1].output x = layers.Dense(64, activation='relu')(prev_output) x = layers.BatchNormalization()(x) # Add residual connection if shapes match if x.shape[-1] == prev_output.shape[-1]: x = layers.add([x, prev_output]) model.add(x) - Stick with small layer sizes (like 32 or 64 units) to keep computation manageable and gradients stable.
内容的提问来源于stack exchange,提问作者Ben Gubler

