You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Pandas groupby()忽略空字符串""?

处理Pandas分组中空值floatValue的合并问题

问题场景

导入CSV数据后,需要用Pandas的groupby()按idField和floatValue字段分组,要求floatValue为唯一浮点值。但存在同一idField对应floatValue为空字符串""的情况,需要将这类记录合并到同idField的有效floatValue分组中。

示例数据

import pandas as pd
import numpy as np

# 示例DataFrame
tempdf = pd.DataFrame({
    'idField': ['A1', 'A1', 'B2', "B2"],
    'floatValue': [1.0, "", 3.0, 4.0],
    'Term': ['this is 1st A', 'This is 2nd A', "This is 1st B", "This is 2nd B"]
})

解决方案

直接按原字段分组会把空字符串的floatValue单独分成一组,所以需要先将空字符串的floatValue填充为对应idField的有效浮点值,再进行分组聚合:

  • 为每个idField提取非空的floatValue(假设每个idField仅对应一个有效浮点值)
  • 将空字符串的floatValue替换为对应idField的有效浮点值
  • 按idField和floatValue分组,聚合Term字段

修改后的代码

# 为每个idField匹配有效floatValue
id_float_map = tempdf[tempdf['floatValue'] != ''].drop_duplicates('idField').set_index('idField')['floatValue']

# 填充空字符串的floatValue
tempdf['floatValue'] = tempdf.apply(
    lambda row: id_float_map[row['idField']] if row['floatValue'] == '' else row['floatValue'],
    axis=1
)

# 分组聚合处理Term字段
groupbyList = ['idField', 'floatValue']
tempdf = tempdf.groupby(groupbyList).agg(
    lambda x: ','.join([f"'{str(elem)}'" for elem in list(set(x))])
).replace(np.nan, "").reset_index()

print(tempdf)

输出结果

idFieldfloatValueTerm
A11.0'this is 1st A','This is 2nd A'
B23.0'This is 1st B'
B24.0'This is 2nd B'

内容的提问来源于stack exchange,提问作者rwjam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 16:17:20