You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组聚合后返回空数组:如何替换为NaN?

解决DataFrame中空数组替换为NaN的问题

问题背景

原始DataFrame定义及输出:

import numpy as np
import pandas as pd

data = {
  "Key": ["A1", "A2", np.nan, "A3", "A4"],
  "Name": ["Candy A", "Candy B", np.nan, "Candy C", "Candy D"],
  "Amout": [25, 50, np.nan, np.nan, 50],
  "Condition": ["Good", "Good", "Good", "Good", "Good"],
  "Packing": ["25 Nice", "49 Nice", "1 Damaged", "40 Nice", "50 Nice"],
  "Sunlight" : [np.nan, np.nan, np.nan, np.nan, "No Sunlight"]
}

df = pd.DataFrame(data)
print(df)

输出:

Key     Name  Amout Condition    Packing     Sunlight
0   A1  Candy A   25.0      Good    25 Nice          NaN
1   A2  Candy B   50.0      Good    49 Nice          NaN
2  NaN      NaN    NaN      Good  1 Damaged          NaN
3   A3  Candy C    NaN      Good    40 Nice          NaN
4   A4  Candy D   50.0      Good    50 Nice  No Sunlight

使用自定义聚合函数分组聚合后,Sunlight列出现了空数组(array([], dtype=object)),且replace和mask方法无法将其替换为NaN:

def custom_agg(s):
    if pd.api.types.is_numeric_dtype(s):
        return s.sum(min_count=1)
    s = s.dropna().drop_duplicates()
    if len(s) > 1:
        return ', '.join(s.astype(str))
    return s

df = df.groupby(df['Key'].notna().cumsum(), as_index=False).agg(custom_agg)
print(df)

聚合后输出:

Key     Name  Amout Condition             Packing     Sunlight
0  A1  Candy A   25.0      Good             25 Nice           []
1  A2  Candy B   50.0      Good  49 Nice, 1 Damaged           []
2  A3  Candy C    NaN      Good             40 Nice           []
3  A4  Candy D   50.0      Good             50 Nice  No Sunlight

解决办法

方法1:修改聚合函数,从源头避免空数组

问题出在当清洗后的序列为空时,直接返回了空的序列/数组。修改聚合函数,增加空序列判断,直接返回NaN:

def custom_agg(s):
    if pd.api.types.is_numeric_dtype(s):
        return s.sum(min_count=1)
    s_clean = s.dropna().drop_duplicates()
    # 空序列直接返回NaN
    if len(s_clean) == 0:
        return np.nan
    elif len(s_clean) > 1:
        return ', '.join(s_clean.astype(str))
    # 返回单个值而非序列/数组
    return s_clean.iloc[0]

重新执行聚合后,原本空数组的位置会直接显示NaN。

方法2:聚合后批量替换空数组

如果不想修改聚合函数,可在聚合后遍历列,判断并替换空数组:

# 针对全表处理
for col in df.columns:
    df[col] = df[col].apply(lambda x: np.nan if isinstance(x, np.ndarray) and x.size == 0 else x)

# 仅针对Sunlight列处理
df['Sunlight'] = df['Sunlight'].apply(lambda x: np.nan if isinstance(x, np.ndarray) and x.size == 0 else x)

执行后所有空数组都会被替换为NaN。

内容的提问来源于stack exchange,提问作者HizaCrenata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 13:32:22