You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于指定类别数组统计NumPy数组元素的出现次数?

问题描述

我目前使用np.bincount统计分类元素的出现次数,代码及输出如下:

print(df[attribute].cat.categories)
>>> Int64Index([0, 1, 2], dtype='int64')
print(df[attribute].to_numpy())
>>> [0 1 0 1 1]
partition = np.bincount(df[attribute].to_numpy())
print(partition)
>>> [2 3]

我希望基于categories数组作为统计区间,得到包含未出现类别(如类别2)的计数结果,即输出[2 3 0]。我的数据框中分类数据均为从0开始的整数编码,且希望避免使用df[attribute].value_counts(),因为性能分析显示它存在性能瓶颈。

解决方案

方法1:补全np.bincount结果到目标长度

由于你的分类是从0开始的连续整数,直接获取categories的长度,将bincount的结果补全到对应长度,缺失位置填充0即可:

categories = df[attribute].cat.categories
counts = np.bincount(df[attribute].to_numpy())
full_counts = np.zeros(len(categories), dtype=int)
full_counts[:len(counts)] = counts
print(full_counts)
>>> [2 3 0]

方法2:使用np.histogram指定统计区间

利用np.histogram,将bins设置为categories对应的整数边界,直接生成包含所有类别的计数:

categories = df[attribute].cat.categories
data = df[attribute].to_numpy()
counts, _ = np.histogram(data, bins=np.arange(len(categories)+1))
print(counts)
>>> [2 3 0]

方法3:通过索引映射对齐分类

如果后续分类出现非连续的情况(当前场景是连续从0开始,该方法同样适用),可以直接用categories作为索引从bincount结果中取值,不存在的类别会自动返回0:

categories = df[attribute].cat.categories
data = df[attribute].to_numpy()
counts = np.bincount(data)
full_counts = counts[categories]
print(full_counts)
>>> [2 3 0]

内容的提问来源于stack exchange,提问作者redrobinyum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 12:28:09