You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计含重复索引的DataFrame中各索引的出现频次?

统计DataFrame重复索引的出现频次

针对你带重复索引的DataFrame,要统计每个唯一索引的出现次数并得到目标数组,这里提供两种简洁的实现方法:

方法一:利用索引的value_counts方法

import pandas as pd
import numpy as np

# 构造示例DataFrame
df = pd.DataFrame(
    {'ages': [94.0, 94.0, 94.0, 30.0, 30.0, 64.0, 64.0, 64.0, 57.0, 57.0, 57.0]},
    index=[0, 0, 0, 1, 1, 2, 2, 2, 3, 3, 3]
)

# 统计索引频次、按索引排序后转成数组
index_counts = df.index.value_counts().sort_index().to_numpy()
print(index_counts)
# 输出:array([3, 2, 3, 3])

index.value_counts()会自动统计每个索引的出现次数,但默认结果是按频次降序排列,所以需要用sort_index()按索引本身的顺序重新排序,最后通过to_numpy()转换为目标数组。

方法二:按索引分组统计

index_counts = df.groupby(df.index).size().to_numpy()
print(index_counts)
# 输出:array([3, 2, 3, 3])

通过groupby(df.index)将数据按索引分组,size()会直接统计每组的行数,结果天然按索引升序排列,直接转成numpy数组就能得到目标结果。

内容的提问来源于stack exchange,提问作者Eisen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 20:09:52