NumPy数组追加至Pandas DataFrame及索引匹配、列删除问题
The core issue here is that your prediction DataFrame uses a default 0-based index, while your actual data has random, non-sequential indices—so pandas can’t align rows correctly when you try to concatenate or assign directly. Let’s fix this step by step:
Step 1: Get the index mapping for predictions
You mentioned prediction index 2 corresponds to first_100_train index 480, so you need a list/array that maps each prediction in first_100 to its matching row index in first_100_train. Let’s call this actual_indices:
# Example: Replace with your actual index mapping actual_indices = [66, 201, 480, ...] # Length 100, matches first_100's size
If you don’t have this list yet, retrieve it from how you generated first_100. For example, if you sampled rows from first_100_train to get your predictions, use:
# If you used something like first_100_train_subset = first_100_train.sample(100) actual_indices = first_100_train_subset.index.values
Step 2: Create a properly indexed prediction DataFrame
Instead of using the default 0-based index, set the prediction DataFrame’s index to match actual_indices:
import pandas as pd # Convert NumPy array to DataFrame with aligned indices pred_df = pd.DataFrame(first_100, columns=['Prediction'], index=actual_indices)
Now pred_df shares the same indices as its corresponding rows in first_100_train, so pandas can align them perfectly.
Step 3: Merge predictions with actual data
Choose one of these options based on your needs:
- Keep all rows in
first_100_train(with NaNs for rows without predictions):combined_df = first_100_train.join(pred_df) - Only keep rows with predictions (the 100 matching rows):
combined_df = first_100_train.loc[actual_indices].join(pred_df)
Step 4: Delete the last column
Once merged, remove the final column using either of these methods:
- Create a new DataFrame excluding the last column:
combined_df = combined_df.iloc[:, :-1] - Modify the DataFrame in place:
combined_df.drop(columns=combined_df.columns[-1], inplace=True)
Example Test Case
Let’s verify with a small sample:
# Sample actual data first_100_train = pd.DataFrame({'Actual': [10,20,30,40,50]}, index=[66,201,480,32,15]) # Sample predictions first_100 = np.array([11,22,49]) # Index mapping actual_indices = [66,201,480] # Execute steps 2-4 pred_df = pd.DataFrame(first_100, columns=['Prediction'], index=actual_indices) combined_df = first_100_train.join(pred_df) combined_df = combined_df.iloc[:, :-1]
The final combined_df will have correctly aligned rows (no mismatched NaNs) and the last column removed.
Content of the question comes from Stack Exchange, asked by Raph117

