如何为数据集填充缺失值?已替换问号为NaN但填充代码无效
com数据集缺失值填充无效的问题 Hey there! Let's dig into why your missing value filling code isn't getting rid of those NaNs in the com dataset. Let's walk through common pitfalls and fixes step by step:
1. 先确认问号转NaN是否真的完全生效
Sometimes, the question marks in your data might have extra spaces (like ' ? ' instead of '?'), which means a simple replace('?', np.nan) would miss them. Try using a regex-based replacement to catch all variants:
import pandas as pd import numpy as np # 读取数据后,替换所有带可能空格的问号为NaN com.replace(r'\s*\?\s*', np.nan, regex=True, inplace=True) # 验证替换结果:查看各列NaN数量 print("替换后的NaN分布:") print(com.isna().sum())
This will make sure every question mark (with or without surrounding spaces) gets turned into a proper NaN.
2. 检查你的填充代码是否有这些常见问题
问题A:没有修改原数据集
The fillna() method returns a new DataFrame by default—it doesn't change the original com unless you specify inplace=True, or assign the result back to com:
# 错误写法:原com不会被修改 com.fillna(com.mean()) # 正确写法二选一: com = com.fillna(com.mean()) # 赋值回原变量 # 或者 com.fillna(com.mean(), inplace=True) # 直接修改原数据
问题B:填充逻辑没覆盖所有列
If you're using fillna(com.mean()), this only works for numeric columns. String/categorical columns will keep their NaNs because you can't calculate a mean for text. For those columns, use the mode (most frequent value) instead:
# 遍历每一列,根据类型选择填充方式 for col in com.columns: if com[col].dtype in ['int64', 'float64']: # 数值列用均值填充 com[col] = com[col].fillna(com[col].mean()) else: # 非数值列用众数填充(mode()[0]取第一个众数) com[col] = com[col].fillna(com[col].mode()[0])
问题C:某些列全是NaN
If a column has nothing but NaNs, mean/mode won't work—you'll need to decide whether to drop that column (com.dropna(axis=1, how='all')) or assign a custom value (like com['col_name'] = com['col_name'].fillna('Unknown') for text columns).
3. 验证填充结果
After running your fix, double-check with:
print("填充后的NaN数量:") print(com.isna().sum())
This will show you if any NaNs are still lingering, and which columns need extra attention.
内容的提问来源于stack exchange,提问作者Jack

