Keras(TensorFlow)前馈神经网络销量预测:变异性优但幅值预测不佳求因
Hey there, I’ve run into this exact issue before when building regression models for sales forecasting—super frustrating when your model gets the trends right but misses the magnitude! Let’s break down the likely culprits and fixes, starting with your hunch about data normalization/bias since that’s a common offender.
Normalization/Scaling Missteps
Your suspicion here is spot-on. Mismatched scaling between input features and your target (sales volume) is one of the top reasons for this kind of underprediction. For example:- If you normalized input features with
StandardScalerorMinMaxScalerbut left the target variable untouched, the model’s output range might get constrained, pulling predictions toward the mean and undercutting high-volume values. Try scaling your target variable too (using the same scaler or a separate one), then inverse-scaling the predictions after inference to get back to real-world values. - Outliers in sales data can also throw off scaling. If high-sales outliers are compressed during normalization, the model never learns to predict their magnitude. Swap in
RobustScalerinstead—it’s designed to ignore extreme values and preserve the distribution of most data points.
- If you normalized input features with
Training Data Bias
The fact that underprediction persists across subsets tells me your training data might have inherent biases:- Check if high-sales samples are underrepresented in your training set. If the model sees too few instances of peak sales, it won’t learn the patterns that drive those high values. Compare histograms of sales distributions in your training and test sets—if they don’t align, try oversampling high-sales cases or using a weighted loss function (in Keras, you can customize loss to assign higher weights to high-volume samples, e.g.,
loss = lambda y_true, y_pred: K.mean(K.square(y_true - y_pred) * sales_weight_vector, axis=-1)). - Missing key features could also be the issue. If you’re leaving out impactful variables like promotions, holidays, or competitor activity, the model can only capture partial trends, not the full magnitude of sales swings.
- Check if high-sales samples are underrepresented in your training set. If the model sees too few instances of peak sales, it won’t learn the patterns that drive those high values. Compare histograms of sales distributions in your training and test sets—if they don’t align, try oversampling high-sales cases or using a weighted loss function (in Keras, you can customize loss to assign higher weights to high-volume samples, e.g.,
Model Architecture Limitations
Your feedforward network might not have the capacity or right activation functions to learn magnitude:- Double-check your output layer: if you used
sigmoidortanh, those functions clamp output to a narrow range, making it impossible to predict high sales volumes. For regression tasks, always use a linear activation for the output layer. - Try boosting the model’s capacity: add more neurons to hidden layers or an extra hidden layer entirely. Just be sure to add dropout or L2 regularization to avoid overfitting as you increase complexity.
- Double-check your output layer: if you used
Loss Function Choice
Mean Squared Error (MSE), the default for many regression tasks, prioritizes minimizing small errors, which can make the model ignore large errors from high-sales samples. Swap in Mean Absolute Error (MAE) or Huber Loss—both are more robust to outliers and force the model to pay attention to magnitude. In Keras, you can usetf.keras.losses.Huber()directly for this.
Start by auditing your scaling approach and training data distribution first—those are the quickest fixes and often the root cause. If that doesn’t resolve the issue, tweak your model’s architecture or loss function. You’ve already got the trend right, so you’re halfway there!
内容的提问来源于stack exchange,提问作者Joshua McPhie

