为何编写Age列空值填充代码后,运行数据仍无变化?
Ah, I see the issue here! Let's break down why your code isn't updating the Age column in your DataFrame, and how to fix it.
Key Problems with Your Current Code
You're modifying a copy, not the original DataFrame
When you useiterrows(), it returns a copy of each row as a Series object. So when you doi['Age'] = ..., you're only updating that temporary copy—not the actual row in your originaldf. It's like editing a photocopy of a document and wondering why the original stays unchanged.Potential NaN propagation
If a specific combination ofPclassandSibSphas no non-nullAgevalues, calculatingmean()will returnNaN. Even if you fixed the copy issue, this would leave those rows empty instead of filling them.
Fixes That Actually Work
1. Use Pandas' Vectorized Grouped Fill (Recommended)
Pandas is designed for vectorized operations—they're way faster than looping, especially with large datasets. This method directly updates the original DataFrame:
# Fill missing Age values with the mean of their Pclass + SibSp group df['Age'] = df.groupby(['Pclass', 'SibSp'])['Age'].transform( lambda x: x.fillna(x.mean()) ) # Optional: Add a fallback for groups with all NaN values df['Age'] = df['Age'].fillna(df['Age'].mean())
The transform method ensures the calculated mean gets mapped back to the correct rows in your original DataFrame, no copies involved.
2. Fix the Loop (Not Recommended, But For Learning)
If you really want to use a loop, you need to modify the original DataFrame directly using loc to target the row index:
for index, row in df[df['Age'].isna()].iterrows(): pc = row['Pclass'] ss = row['SibSp'] # Use df.loc to update the original row by index group_mean = df[(df['Pclass'] == pc) & (df['SibSp'] == ss)]['Age'].mean() df.loc[index, 'Age'] = group_mean if not pd.isna(group_mean) else df['Age'].mean()
This bypasses the copy issue by directly writing to the original DataFrame's index position.
内容的提问来源于stack exchange,提问作者sanket srivastava

