You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NumPy数组追加至Pandas DataFrame及索引匹配、列删除问题

Solution to Align Predictions with Actual Data and Merge DataFrames

The core issue here is that your prediction DataFrame uses a default 0-based index, while your actual data has random, non-sequential indices—so pandas can’t align rows correctly when you try to concatenate or assign directly. Let’s fix this step by step:

Step 1: Get the index mapping for predictions

You mentioned prediction index 2 corresponds to first_100_train index 480, so you need a list/array that maps each prediction in first_100 to its matching row index in first_100_train. Let’s call this actual_indices:

# Example: Replace with your actual index mapping
actual_indices = [66, 201, 480, ...]  # Length 100, matches first_100's size

If you don’t have this list yet, retrieve it from how you generated first_100. For example, if you sampled rows from first_100_train to get your predictions, use:

# If you used something like first_100_train_subset = first_100_train.sample(100)
actual_indices = first_100_train_subset.index.values

Step 2: Create a properly indexed prediction DataFrame

Instead of using the default 0-based index, set the prediction DataFrame’s index to match actual_indices:

import pandas as pd

# Convert NumPy array to DataFrame with aligned indices
pred_df = pd.DataFrame(first_100, columns=['Prediction'], index=actual_indices)

Now pred_df shares the same indices as its corresponding rows in first_100_train, so pandas can align them perfectly.

Step 3: Merge predictions with actual data

Choose one of these options based on your needs:

  • Keep all rows in first_100_train (with NaNs for rows without predictions):
    combined_df = first_100_train.join(pred_df)
    
  • Only keep rows with predictions (the 100 matching rows):
    combined_df = first_100_train.loc[actual_indices].join(pred_df)
    

Step 4: Delete the last column

Once merged, remove the final column using either of these methods:

  • Create a new DataFrame excluding the last column:
    combined_df = combined_df.iloc[:, :-1]
    
  • Modify the DataFrame in place:
    combined_df.drop(columns=combined_df.columns[-1], inplace=True)
    

Example Test Case

Let’s verify with a small sample:

# Sample actual data
first_100_train = pd.DataFrame({'Actual': [10,20,30,40,50]}, index=[66,201,480,32,15])
# Sample predictions
first_100 = np.array([11,22,49])
# Index mapping
actual_indices = [66,201,480]

# Execute steps 2-4
pred_df = pd.DataFrame(first_100, columns=['Prediction'], index=actual_indices)
combined_df = first_100_train.join(pred_df)
combined_df = combined_df.iloc[:, :-1]

The final combined_df will have correctly aligned rows (no mismatched NaNs) and the last column removed.

Content of the question comes from Stack Exchange, asked by Raph117

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:36:26