Keras多输出回归模型不收敛问题及模型配置咨询
Hey Eric! Let's walk through each of your questions and tweak your model to better fit your resource allocation prediction task. First, quick clarification: your task is technically a multi-output regression with a sum-to-1 constraint, which makes it more aligned with predicting probability distributions—this is key for optimizing your model.
1. Does the current model configuration fit my goal?
The core feed-forward structure is a solid starting point, but it’s missing a critical piece to account for your Y values' sum-to-1 constraint. Right now, using MSE loss doesn’t enforce that your model’s outputs add up to 1, which might lead to predictions that don’t make sense for resource allocation. The layer sizes (256 neurons) are reasonable, but you can adjust them later if you see underfitting/overfitting during training.
2. Do I need to change the activation functions?
The relu activation in your hidden layers is totally fine—it’s a reliable choice that avoids gradient vanishing issues and works well for most tabular tasks like yours. If you later notice slow convergence or gradient saturation, you could experiment with swish or leaky_relu, but start with relu first since it’s proven here.
3. Are the input shape (308,) and output shape 28 set correctly?
Perfect! Your X_train has shape (46038, 308), so input_shape=(308,) matches exactly. And your output layer with Dense(28) aligns perfectly with your Y_train’s (46038, 28) shape—no issues here.
4. Should I add an activation function to the output layer? If so, which one is best?
Absolutely yes, and softmax is the ideal choice here! The softmax activation ensures your 28 output values sum to exactly 1, which is exactly the constraint your resource allocation task requires. Without it, your model could output arbitrary values that don’t add up to 1, making the predictions useless for real-world allocation.
Along with softmax, I’d recommend switching your loss function from mean_squared_error to categorical_crossentropy. MSE is great for unconstrained regression, but cross-entropy is designed to measure the difference between two probability distributions (your true Y ratios and predicted ratios), which fits your task far better.
Here’s your adjusted model code (cleaned up a bit too—you don’t need to specify input_shape for the second Dense layer; Keras infers it automatically):
model = Sequential() model.add(Dense(256, input_shape=(308,) )) model.add(Activation('relu')) model.add(Dropout(0.5)) model.add(Dense(256)) model.add(Activation('relu')) model.add(Dropout(0.5)) model.add(Dense(28, activation='softmax')) # Integrated softmax here model.compile(loss='categorical_crossentropy', optimizer='adam')
One quick note: make sure your Y_train data is stored as float32 to match TensorFlow’s default dtype—this avoids any unexpected type mismatches during training.
内容的提问来源于stack exchange,提问作者Eric Broda

