You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并Pandas DataFrame中重复client_id的行并规整数据?

解决DataFrame重复client_id合并与NaN替换问题

核心思路

按client_id分组聚合,保留同一客户的有效数据,再将剩余NaN值统一替换为0.0。

具体实现代码

方式1:指定列手动聚合(适合列数较少的场景)

假设你的full_result包含client_id、收入列(如rzd_income/auto_income/air_income)和客户基础信息列(如client_name/client_age):

import pandas as pd

# 按client_id分组,收入列取有效值(自动忽略NaN),基础信息列取首个有效值
cleaned_result = full_result.groupby('client_id').agg(
    rzd_income=('rzd_income', 'max'),
    auto_income=('auto_income', 'max'),
    air_income=('air_income', 'max'),
    client_name=('client_name', 'first'),
    client_age=('client_age', 'first')
).reset_index()

# 将所有NaN替换为0.0
cleaned_result = cleaned_result.fillna(0.0)

方式2:自动适配列类型(适合列数较多的场景)

无需手动指定每一列,自动根据列类型选择聚合逻辑:

import pandas as pd

# 构建聚合规则:数值列取max(保留有效值),非数值列取first(取首个一致值)
agg_func = {}
for col in full_result.columns:
    if col == 'client_id':
        continue
    if pd.api.types.is_numeric_dtype(full_result[col]):
        agg_func[col] = 'max'
    else:
        agg_func[col] = 'first'

# 分组聚合+重置索引
cleaned_result = full_result.groupby('client_id').agg(agg_func).reset_index()

# 替换NaN为0.0
cleaned_result = cleaned_result.fillna(0.0)

注意事项

  • 如果同一client_id的同一收入列存在多个非NaN有效值(比如重复录入的多条收入记录),可将聚合函数max替换为sum,实现收入累加。
  • 客户基础信息列(如姓名、年龄)需确保同一client_id的记录一致,否则first会取第一条记录的值,若有不一致需先排查数据源头问题。

内容的提问来源于stack exchange,提问作者Follin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 12:00:24