如何用Pandas计算DataFrame指定列的非空值占比?
计算Pandas DataFrame指定列的非空值占比并添加统计行
我来帮你搞定这个需求!要计算A、C、D列的非空值占比,并且把统计结果作为首行添加到原DataFrame中,我们可以用以下步骤实现:
1. 准备示例数据(还原你的原始DataFrame)
import pandas as pd import numpy as np # 构造你的原始DataFrame data = { 'id': [0, 1, 2, 3, 4, 5, 6, 7], 'A': [1.0, np.nan, 2.0, 55.0, 6.0, np.nan, -17.0, np.nan], 'B': ['one', 'one', 'two', 'three', 'two', 'two', 'one', 'three'], 'C': [4.0, 14.0, 3.0, np.nan, 8.0, 7.0, np.nan, 11.0], 'D': [np.nan, np.nan, -12.0, 12.0, 12.0, -12.0, np.nan, np.nan] } df = pd.DataFrame(data)
2. 计算目标列的非空值占比
# 计算A、C、D列的非空值占比,转成百分比格式 non_null_pct = df[['A', 'C', 'D']].notna().mean() * 100 non_null_pct = non_null_pct.apply(lambda x: f"{x}%")
这里的逻辑很直观:
notna()将非空值转为True(数值等价于1),空值转为False(数值等价于0)mean()计算列的平均值,也就是非空值占总行数的比例- 乘以100后用
apply格式化为带%的字符串,更符合阅读习惯
3. 构造统计行并合并到原DataFrame
# 创建统计行的DataFrame summary_row = pd.DataFrame({ 'id': ['not-nulls_pct'], 'A': [non_null_pct['A']], 'B': [np.nan], # B列不需要统计,设为NaN 'C': [non_null_pct['C']], 'D': [non_null_pct['D']] }) # 把统计行放在原DataFrame的最上方,重新生成连续索引 result_df = pd.concat([summary_row, df], ignore_index=True)
最终结果
运行完上面的代码后,result_df就会和你期望的结构完全一致:
id A B C D 0 not-nulls_pct 62.5% NaN 75.0% 50.0% 1 0 1.0 one 4.0 NaN 2 1 NaN one 14.0 NaN 3 2 2.0 two 3.0 -12.0 4 3 55.0 three NaN 12.0 5 4 6.0 two 8.0 12.0 6 5 NaN two 7.0 -12.0 7 6 -17.0 one NaN NaN 8 7 NaN three 11.0 NaN
内容的提问来源于stack exchange,提问作者ah bon
相关产品推荐
相关产品推荐

