如何用每行非NaN值的均值填充Pandas DataFrame中的NaN?
解决方案:按行均值填充DataFrame中的NaN
问题分析
你需要用每行非NaN数值的均值填充该行空缺值,但原代码未生效,核心问题有三点:
- 未排除非数值列(
rows),求和时包含字符串导致计算结果错误 fillna默认按列填充,无法匹配行级的均值数据- 硬编码数值列个数(
4)不够灵活,索引方式冗余
正确代码实现
方法一:先计算行均值再填充
import pandas as pd import numpy as np # 初始化示例DataFrame df = pd.DataFrame([['row 1', 1, 2, 3, 4], ['row 2', 5, np.nan, 7, 8], ['row 3', np.nan, np.nan, 11, 12]], index=['A', 'B', 'C'], columns=['rows', '1st', '2nd', '3th', '4th']) # 指定数值列(排除字符串列rows) numeric_cols = ['1st', '2nd', '3th', '4th'] # 计算每行数值列的均值(自动跳过NaN) row_means = df[numeric_cols].mean(axis=1) # 按行填充NaN:利用索引匹配,将每行均值填充到该行空缺位置 df[numeric_cols] = df[numeric_cols].fillna(row_means, axis=0) print(df)
方法二:用apply按行处理(更简洁)
numeric_cols = ['1st', '2nd', '3th', '4th'] df[numeric_cols] = df[numeric_cols].apply(lambda row: row.fillna(row.mean()), axis=1)
运行结果
处理后的DataFrame如下:
| rows | 1st | 2nd | 3th | 4th |
|---|---|---|---|---|
| row 1 | 1.0 | 2.0 | 3 | 4 |
| row 2 | 5.0 | 6.666667 | 7 | 8 |
| row 3 | 11.5 | 11.5 | 11 | 12 |
原代码修正说明
- 替换硬编码的
4为df[numeric_cols].count(axis=1),自动获取每行非NaN数值的个数 - 必须排除非数值列,避免字符串干扰求和计算
- 确保填充值的索引与DataFrame行索引匹配,或通过
apply按行处理实现精准填充
内容的提问来源于stack exchange,提问作者Maryam Mirfazaelian
相关产品推荐
相关产品推荐

