基于Python的机票价格回归模型:训练集高R²但测试集准确率偏低
Hey there! You’ve got a solid foundation with that 80%+ R² on your training data, but it’s frustrating when the test set’s accuracy (measured by predictions within 5% of actual prices) doesn’t match up. Let’s walk through some common causes and fixes tailored to your setup:
1. Metric Mismatch: R² vs. 5% Accuracy
First, it’s important to note that R² and your custom accuracy metric measure very different things. R² quantifies how well your model explains variance in the target variable, but it doesn’t care about individual prediction errors falling within a fixed percentage range.
For example:
- Your model might do great at capturing overall price trends, but struggle with extreme values (super cheap or ultra-expensive tickets). A 5% deviation on a $1000 ticket is $50, which is a bigger absolute gap than 5% on a $200 ticket—these high-value misses could be dragging down your accuracy rate.
- If you only log-transformed your inputs (not the target price), the model’s error in the linear space might not translate evenly to percentage errors in the original price scale.
Action: Break down your test set error by price ranges and days-to-departure buckets. This will show you exactly which segments the model is failing on.
2. Overfitting: The Classic Culprit
A high training R² paired with poor test performance is a textbook sign of overfitting, even with 100k+ data points. If your model is too complex (e.g., a deep decision tree without constraints, or a linear model with too many high-degree terms), it’s memorizing noise in the training data instead of learning generalizable patterns.
Fixes:
- Add regularization: For linear regression, swap plain
LinearRegressionforRidgeorLassoto penalize large coefficients. For tree-based models, set limits on max depth, min samples per leaf, or use ensemble methods like Random Forests which naturally reduce overfitting. - Validate properly: Replace your single 80-20 split with 5-fold or 10-fold cross-validation. This gives you a more reliable view of how the model performs across different data subsets.
- Check for data leakage: Did you calculate mean/median values using the entire dataset before splitting into train/test? That’s a common leakage issue—you should only compute these stats using the training set to avoid giving the model an unfair advantage on training data.
3. Log-Transform Nuances
You mentioned log-transforming your input variables (mean, median, days-to-departure), but have you considered transforming the target price as well?
If you’re predicting raw prices using log-transformed inputs, errors in the model’s predictions can scale unevenly with price. For example, a small error in the log-input space might turn into a large percentage error for high-priced flights.
Suggestion: Try log-transforming the target flight price, train the model on log prices, then convert predictions back to the original scale using exponentiation. Just remember to correct for bias from the log-normal distribution—add half the variance of the model’s residuals to the predicted log value before exponentiating.
4. Missing Key Features
Flight prices are driven by way more than just mean/median prices and days-to-departure. Factors like airline, route, day of week, time of day, seat class, holiday seasons, and even fuel prices can have massive impacts on pricing. If your model lacks these features, it can’t learn the true patterns that drive price changes, leading to poor generalization.
Action: Audit your crawled data to see what other variables you can extract. Even simple features like is_weekend_departure or is_holiday_season can make a big difference. If you don’t have these data points, consider adding them to your crawler pipeline.
5. Reassess Your Accuracy Metric
A 5% deviation threshold is pretty strict for flight prices, which are inherently volatile. Your model might be performing reasonably well, but this binary metric is masking that.
Try these metrics instead:
- MAPE (Mean Absolute Percentage Error): Gives you the average percentage error across all predictions, which is more informative than a pass/fail count.
- MAE (Mean Absolute Error): Measures the average absolute gap between predictions and actual prices, helping you understand the overall error magnitude.
内容的提问来源于stack exchange,提问作者Ananth

