Kaggle流量预测竞赛提交代码报错:ValueError维度不匹配问题
Got it, let's dig into why that ValueError is hitting you and get your submission code working. That error is screaming that the data you're passing (predictions) doesn't line up with the index and columns you're trying to use for your submission DataFrame. Here's the breakdown of what went wrong and how to fix it:
What's Causing the Error?
- Wrong Column Definition: Your final line uses
columns=pages_and_id['Visits']—but that's trying to set every value in theVisitscolumn as a separate column name for your submission. The Kaggle submission only needs two fixed columns:IdandVisits. - Unflattened Predictions: Your
predictionsare shaped(36266, 182)—that's 36266 pages, each with 182 days of forecasts. But the submission expects one forecast per row (one perIdinkey_1), so you need to turn that 2D array into a 1D array with length36266 * 182 = 17407560(matching the number of rows inkey_1). - Redundant Data Merging: Combining
sample_subandkey_1is unnecessary—key_1already has all theIdandPagepairs you need, andsample_subjust shows the required submission format.
Corrected Submission Code
Here's the cleaned-up code that fixes all these issues:
# Read required submission files key_1 = pd.read_csv('../input/web-traffic-time-series-forecasting/key_1.csv.zip') sample_sub = pd.read_csv('../input/web-traffic-time-series-forecasting/sample_submission_1.csv.zip') # Flatten your 2D predictions into a 1D array # Make sure the order matches: each page's 182 forecasts are consecutive, matching key_1's page order flattened_predictions = predictions.flatten(order='C') # 'C' order flattens row-wise (page first, then days) # Create the submission DataFrame with the correct structure submission = pd.DataFrame({ 'Id': key_1['Id'], 'Visits': flattened_predictions }) # Optional: Align columns to match sample_sub's exact order submission = submission[['Id', 'Visits']] # Verify the shape (should be (17407560, 2)) print(submission.shape)
Critical Check: Ensure Prediction Order Matches
Double-check that the order of pages in your predictions matches the unique page order in key_1. If they're out of sync, your forecasts will be assigned to the wrong pages. You can validate this with:
# Get unique pages from your training data (adjust to match your training data variable name) training_pages = your_train_data['Page'].unique() # Get unique pages from key_1 key_pages = key_1['Page'].unique() # Confirm they're identical in order print(np.array_equal(training_pages, key_pages)) # Should print True
If this returns False, you'll need to reorder your predictions to match key_pages before flattening.
内容的提问来源于stack exchange,提问作者HTTP 418

