如何从Pandas嵌套列表列取最小值?为何np.min()失效而np.mean()可用?
问题:Pandas处理嵌套列表时np.min()报错,np.mean()可正常运行
我在处理Pandas中嵌套列表类型的列时,发现np.mean()能正常计算邻居值的均值,但替换为np.min()时会抛出错误:ValueError: zero-size array to reduction operation minimum which has no identity。需要明确报错原因,以及实现取最小值的预期效果。
可正常运行的原代码(使用np.mean())
import pandas as pd import numpy as np def transformation(custom_df): dic = dict(zip(custom_df['customers'], custom_df['values'])) custom_df['values'] = np.where(custom_df['values'].isna() & (custom_df['valid_neighbors'] >= 1), custom_df['neighbors'].apply( lambda row: np.mean([dic[v] for v in row if dic.get(v)])), custom_df['values']) return custom_df customers = [1, 2, 3, 4, 5, 6] values = [np.nan, np.nan, 10, np.nan, 11, 12] neighbors = [[6], [3], [], [3, 5], [6], [5]] vn = [1, 1, 0, 2, 1, 1] df2 = pd.DataFrame({'customers': customers, 'values': values, 'neighbors': neighbors, 'valid_neighbors': vn}) print("原数据:") print(df2) df2 = transformation(df2) print("\n处理后结果:") print(df2)
原代码运行结果
原数据: customers values neighbors valid_neighbors 0 1 NaN [6] 1 1 2 NaN [3] 1 2 3 10.0 [] 0 3 4 NaN [3, 5] 2 4 5 11.0 [6] 1 5 6 12.0 [5] 1 处理后结果: customers values neighbors valid_neighbors 0 1 12.0 [6] 1 1 2 10.0 [3] 1 2 3 10.0 [] 0 3 4 10.5 [3, 5] 2 4 5 11.0 [6] 1 5 6 12.0 [5] 1
报错原因
np.mean()对空数组会返回np.nan,不会触发错误;但np.min()没有预设的空数组默认返回值,当传入空数组时,直接抛出ValueError。- 即使代码中判断了
valid_neighbors >=1,如果邻居对应的values全为NaN,[dic[v] for v in row if dic.get(v)]会生成空列表,此时调用np.min()就会报错。
解决方法及修改后代码
可以自定义一个处理函数,先检查列表是否为空,空则返回np.nan(或其他默认值),非空再计算最小值;也可以用pd.Series.min(),它对空序列返回np.nan,不会报错。
修改后代码(使用自定义函数)
import pandas as pd import numpy as np def get_min(arr): if not arr: return np.nan return np.min(arr) def transformation(custom_df): dic = dict(zip(custom_df['customers'], custom_df['values'])) custom_df['values'] = np.where(custom_df['values'].isna() & (custom_df['valid_neighbors'] >= 1), custom_df['neighbors'].apply( lambda row: get_min([dic[v] for v in row if dic.get(v)])), custom_df['values']) return custom_df customers = [1, 2, 3, 4, 5, 6] values = [np.nan, np.nan, 10, np.nan, 11, 12] neighbors = [[6], [3], [], [3, 5], [6], [5]] vn = [1, 1, 0, 2, 1, 1] df2 = pd.DataFrame({'customers': customers, 'values': values, 'neighbors': neighbors, 'valid_neighbors': vn}) df2 = transformation(df2) print(df2)
修改后运行结果(符合预期)
customers values neighbors valid_neighbors 0 1 12.0 [6] 1 1 2 10.0 [3] 1 2 3 10.0 [] 0 3 4 10.0 [3, 5] 2 4 5 11.0 [6] 1 5 6 12.0 [5] 1
内容的提问来源于stack exchange,提问作者enriicoo
相关产品推荐
相关产品推荐

