You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于分组变量填充数据中的NA值?

用Pandas填充同一用户的缺失体重数据(保留无有效值的NA)

这个场景在用户数据清洗里太常见了——我们需要把同一用户的缺失体重、体重日期用该用户已有的有效值补上,完全没有有效值的用户就保留NA不变。下面直接上解决方案:

先看原始数据

tableididnameweightweightdate
11david10001/01/2020
21david10001/01/2020
31davidNANA
42anneNANA
53peter15002/10/2020
63peter15002/10/2020

解决方案代码

用Pandas的groupby结合双向填充(bfill+ffill)是最简洁的方式,它会自动在同一用户分组内用有效值覆盖NA,全NA的分组则保持原样:

import pandas as pd

# 构造你的原始数据集
df = pd.DataFrame({
    'tableid': [1, 2, 3, 4, 5, 6],
    'id': [1, 1, 1, 2, 3, 3],
    'name': ['david', 'david', 'david', 'anne', 'peter', 'peter'],
    'weight': [100, 100, None, None, 150, 150],
    'weightdate': ['01/01/2020', '01/01/2020', None, None, '02/10/2020', '02/10/2020']
})

# 按用户id分组,对weight和weightdate字段做双向填充
df[['weight', 'weightdate']] = df.groupby('id')[['weight', 'weightdate']].transform(
    lambda x: x.bfill().ffill()
)

print(df)

如果你的数据集里同一用户的有效值可能分散在不同位置,双向填充能确保不管NA在分组的开头、中间还是结尾,都能被有效值补上。

运行后的结果

和你期望的完全一致:

tableididnameweightweightdate
11david10001/01/2020
21david10001/01/2020
31david10001/01/2020
42anneNaNNaN
53peter15002/10/2020
63peter15002/10/2020

补充说明

  • 这里用transform而不是apply,是因为它能保持原数据的索引和结构不变
  • 如果你确定同一用户的有效值都是相同的,也可以直接取分组内第一个非NA值来填充,效果是一样的:
def fill_with_first_valid(group):
    # 取分组内第一个非NA的体重值,没有的话留None
    valid_weight = group['weight'].dropna().iloc[0] if not group['weight'].dropna().empty else None
    valid_date = group['weightdate'].dropna().iloc[0] if not group['weightdate'].dropna().empty else None
    group['weight'] = group['weight'].fillna(valid_weight)
    group['weightdate'] = group['weightdate'].fillna(valid_date)
    return group

df = df.groupby('id').apply(fill_with_first_valid).reset_index(drop=True)

内容的提问来源于stack exchange,提问作者Andres Mora

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:44:48