You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas计算DataFrame指定列的非空值占比?

计算Pandas DataFrame指定列的非空值占比并添加统计行

我来帮你搞定这个需求!要计算A、C、D列的非空值占比,并且把统计结果作为首行添加到原DataFrame中,我们可以用以下步骤实现:

1. 准备示例数据(还原你的原始DataFrame)

import pandas as pd
import numpy as np

# 构造你的原始DataFrame
data = {
    'id': [0, 1, 2, 3, 4, 5, 6, 7],
    'A': [1.0, np.nan, 2.0, 55.0, 6.0, np.nan, -17.0, np.nan],
    'B': ['one', 'one', 'two', 'three', 'two', 'two', 'one', 'three'],
    'C': [4.0, 14.0, 3.0, np.nan, 8.0, 7.0, np.nan, 11.0],
    'D': [np.nan, np.nan, -12.0, 12.0, 12.0, -12.0, np.nan, np.nan]
}
df = pd.DataFrame(data)

2. 计算目标列的非空值占比

# 计算A、C、D列的非空值占比,转成百分比格式
non_null_pct = df[['A', 'C', 'D']].notna().mean() * 100
non_null_pct = non_null_pct.apply(lambda x: f"{x}%")

这里的逻辑很直观:

  • notna() 将非空值转为True(数值等价于1),空值转为False(数值等价于0)
  • mean() 计算列的平均值,也就是非空值占总行数的比例
  • 乘以100后用apply格式化为带%的字符串,更符合阅读习惯

3. 构造统计行并合并到原DataFrame

# 创建统计行的DataFrame
summary_row = pd.DataFrame({
    'id': ['not-nulls_pct'],
    'A': [non_null_pct['A']],
    'B': [np.nan],  # B列不需要统计,设为NaN
    'C': [non_null_pct['C']],
    'D': [non_null_pct['D']]
})

# 把统计行放在原DataFrame的最上方,重新生成连续索引
result_df = pd.concat([summary_row, df], ignore_index=True)

最终结果

运行完上面的代码后,result_df就会和你期望的结构完全一致:

id      A      B      C       D
0  not-nulls_pct  62.5%    NaN  75.0%   50.0%
1              0    1.0    one    4.0     NaN
2              1    NaN    one   14.0     NaN
3              2    2.0    two    3.0   -12.0
4              3   55.0  three    NaN    12.0
5              4    6.0    two    8.0    12.0
6              5    NaN    two    7.0   -12.0
7              6  -17.0    one    NaN     NaN
8              7    NaN  three   11.0     NaN

内容的提问来源于stack exchange,提问作者ah bon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:03:25