Pandas布尔列value_counts返回多计数异常问题咨询
问题原因与解决办法
异常原因
虽然df.dtypes显示列是bool类型,但你的数据里大概率混合了Python原生布尔值和numpy布尔值(比如numpy.bool_类型)。这两种类型在pandas的dtype检测中都会被识别为bool,但在分组、统计时会被当成不同取值,所以出现了重复的True条目,以及超过4种的组合。
你可以先验证这个猜测:
df['IsInTime'].apply(type).value_counts()
执行后会看到输出两种类型:<class 'bool'>和<class 'numpy.bool_'>。
解决步骤
方法1:用pd.to_numeric统一类型后转布尔
这种方法能彻底统一底层类型,适合大多数情况:
# 处理IsInTime列 df['IsInTime'] = pd.to_numeric(df['IsInTime'], downcast='integer').astype(bool) # 处理IsGoodKine列 df['IsGoodKine'] = pd.to_numeric(df['IsGoodKine'], downcast='integer').astype(bool)
方法2:直接映射为Python原生布尔值
用map函数强制把所有元素转成Python的bool类型:
df['IsInTime'] = df['IsInTime'].map(bool) df['IsGoodKine'] = df['IsGoodKine'].map(bool)
方法3:使用pandas nullable布尔类型
如果你的数据可能存在空值,或者需要更严格的类型控制,可以转成pandas的boolean类型(注意首字母大写):
df['IsInTime'] = df['IsInTime'].astype('boolean') df['IsGoodKine'] = df['IsGoodKine'].astype('boolean')
之后如果需要普通bool类型,再转回astype(bool)即可。
处理完成后,再执行df['IsInTime'].value_counts()和df.groupby(['IsGoodKine', 'IsInTime']).size(),就能得到正确的2种取值和4种逻辑组合的统计结果了。
内容的提问来源于stack exchange,提问作者Chandler Kenworthy
相关产品推荐
相关产品推荐

