如何基于百分位区间为Pandas列赋值?遇Series真值歧义错误
问题
需要给Pandas DataFrame多列的每个值,根据所属百分位区间分配1-5的分数。编写get_percentiles函数后用apply调用时触发错误:Truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all(),期望得到包含对应分数的新DataFrame。
原代码如下:
import pandas as pd import numpy as np def get_percentiles(x, percentile_array): percentile_array = np.sort(np.array(percentile_array)) if x < x.quantile(percentile_array[0]) < 0: return 1 elif (x >= x.quantile(percentile_array[0]) & (x < x.quantile(percentile_array[1])): return 2 elif (x >= x.quantile(percentile_array[1]) & (x < x.quantile(percentile_array[2])): return 3 elif (x >= x.quantile(percentile_array[2]) & (x < x.quantile(percentile_array[3])): return 4 else: return 5 df = pd.DataFrame({'col1' : [1,10,5,9,15,4], 'col2' : [4,10,15,19,3,2], 'col3' : [10,5,6,9,1,24]}) percentile_array = [0.05, 0.25, 0.5, 0.75] df.apply(lambda x : get_percentiles(x, percentile_array), result_type = 'expand')
错误原因
- 参数类型误解:
apply传入get_percentiles的x是整列的Series,不是单个元素。直接用x < x.quantile(...)会生成布尔Series,if语句无法直接判断布尔Series的真假,导致歧义错误。 - 语法与逻辑错误:原代码的
elif括号不匹配,逻辑与&的使用缺少括号包裹单个条件;第一个if的x < x.quantile(...) < 0逻辑混乱,完全不符合需求。 - 函数逻辑错位:原函数试图对整列Series返回单个值,但我们需要给每个元素分配对应分数,逻辑方向错误。
解决方法
方法1:用pd.cut(推荐,高效简洁)
pd.cut是Pandas专门的分箱工具,直接按百分位划分区间并分配1-5的分数:
import pandas as pd import numpy as np df = pd.DataFrame({'col1' : [1,10,5,9,15,4], 'col2' : [4,10,15,19,3,2], 'col3' : [10,5,6,9,1,24]}) percentile_array = [0.05, 0.25, 0.5, 0.75] def assign_score(col): # 生成包含两端极值的分箱区间 bins = [-np.inf] + [col.quantile(p) for p in percentile_array] + [np.inf] # 对应区间的分数标签 labels = [1,2,3,4,5] # 分箱并分配标签,include_lowest确保最小值被包含在第一个区间 return pd.cut(col, bins=bins, labels=labels, include_lowest=True) # 对每列应用分箱逻辑 df_scores = df.apply(assign_score) print(df_scores)
方法2:修正自定义函数(用np.select处理Series)
如果要保留自定义函数的思路,用np.select替代if-elif,直接处理Series的条件判断:
import pandas as pd import numpy as np def get_percentiles(x, percentile_array): percentile_array = np.sort(np.array(percentile_array)) # 提前计算各百分位的阈值 p0, p1, p2, p3 = [x.quantile(p) for p in percentile_array] # 定义每个分数对应的条件 conditions = [ x < p0, (x >= p0) & (x < p1), (x >= p1) & (x < p2), (x >= p2) & (x < p3) ] # 对应条件的分数 values = [1,2,3,4] # 不满足前面条件的返回5 return np.select(conditions, values, default=5) df = pd.DataFrame({'col1' : [1,10,5,9,15,4], 'col2' : [4,10,15,19,3,2], 'col3' : [10,5,6,9,1,24]}) percentile_array = [0.05, 0.25, 0.5, 0.75] df_scores = df.apply(lambda x : get_percentiles(x, percentile_array)) print(df_scores)
两种方法都能输出预期的分数DataFrame,方法1更符合Pandas的惯用写法,效率更高。
内容的提问来源于stack exchange,提问作者Karthik S
相关产品推荐
相关产品推荐

