You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组DataFrame取每组前50行后100行报错,求解决方案

问题:如何获取Pandas分组排序后的第51至150行?

我有一个名为euc_data的Pandas DataFrame,包含code1、code2、euclidean_distance列。我按code1分组、按euclidean_distance排序后,通过以下代码获取了每组前50行:

matrix_top_50 = euc_data.sort_values(['code1', 'euclidean_distance']).groupby('code1').head(50).reset_index(drop=True)

现在我想创建另一个矩阵,获取每组按euclidean_distance排序后的接下来100行,尝试使用代码:

start = 51
end = 151
next_matrix = euc_data.sort_values(['code1', 'euclidean_distance']).groupby('code1').iloc[start:end].reset_index(drop=True)

但报错:Cannot access callable attribute 'iloc' of 'DataFrameGroupBy' objects, try using the 'apply' method。请问如何实现该需求?


解决方案

你遇到的问题核心在于:groupby()返回的是DataFrameGroupBy对象,它本身不支持直接调用iloc,必须通过分组处理方法来操作每个组内的DataFrame。另外还要注意一个关键细节:Pandas的位置索引是从0开始的,前50行对应的是组内索引0-49,接下来的100行应该是索引50-149(对应第51到150行),你之前的start=51会跳过第51行,这点需要纠正。

下面给你两种可行的实现方式:

方法1:使用groupby.apply结合iloc

这种方式直观易懂,直接对每个分组应用iloc切片:

# 先排序,再分组处理每个组的切片
next_matrix = (
    euc_data.sort_values(['code1', 'euclidean_distance'])
    .groupby('code1')
    .apply(lambda group: group.iloc[50:150])  # 取组内第51到150行(索引50到149)
    .reset_index(drop=True)
)

方法2:使用cumcount标记组内行号(更高效)

如果你的数据集比较大,apply的循环操作效率可能稍低,推荐用cumcount()给每个组内的行分配序号,再通过筛选序号来获取目标行:

# 1. 先按要求排序
euc_sorted = euc_data.sort_values(['code1', 'euclidean_distance'])

# 2. 添加组内序号列,从0开始计数
euc_sorted['group_rank'] = euc_sorted.groupby('code1').cumcount()

# 3. 筛选序号在50到149之间的行,然后删除辅助列
next_matrix = (
    euc_sorted[(euc_sorted['group_rank'] >= 50) & (euc_sorted['group_rank'] < 150)]
    .drop('group_rank', axis=1)
    .reset_index(drop=True)
)

两种方法都能实现你的需求,其中方法2在处理大型数据集时性能更优,因为它避免了apply的循环开销。


内容的提问来源于stack exchange,提问作者Shubham R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:08:29