You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas双列分组后提取每组B列最小、中间、最大行的方法

Pandas提取分组内指定B值对应行的解决方案

使用groupby的实现

完整提取最小、中间、最大行

由于每个ID固定对应3行数据,我们可以先按ID+B排序,再通过groupby提取每组的指定位置行:

import pandas as pd

# 构造示例数据
df = pd.DataFrame.from_dict({
    "ID" : [1,1,1,2,2,2,3,3,3], 
    "B" : [1,3,5,2,4,6,3,7,8], 
    "C" : [3,2,4,1,6,5,3,2,4]
})

# 按ID分组,组内按B升序排序后提取第0、1、2行(对应最小、中间、最大)
sorted_df = df.sort_values(["ID", "B"])
result = sorted_df.groupby("ID").nth([0, 1, 2]).reset_index()
print(result)

输出结果:

ID  B  C
0   1  1  3
1   1  3  2
2   1  5  4
3   2  2  1
4   2  4  6
5   2  6  5
6   3  3  3
7   3  7  2
8   3  8  4

仅提取最小行(示例要求)

如果只需要每个ID下B最小的行:

min_rows = sorted_df.groupby("ID").nth(0).reset_index()
print(min_rows)

输出:

ID  B  C
0   1  1  3
1   2  2  1
2   3  3  3

满足双列分组要求的写法

若必须将B纳入分组逻辑,可通过groupby结合apply在组内排序筛选:

result = df.groupby("ID", group_keys=False).apply(
    lambda grp: grp.sort_values("B").iloc[[0,1,2]]
).reset_index(drop=True)

不使用groupby的替代方案

方案1:基于排序和索引切片

利用每个ID固定3行的特点,排序后直接通过索引步长提取:

sorted_df = df.sort_values(["ID", "B"])

# 提取最小行:索引0,3,6
min_rows = sorted_df.iloc[::3].reset_index(drop=True)
# 提取中间行:索引1,4,7
mid_rows = sorted_df.iloc[1::3].reset_index(drop=True)
# 提取最大行:索引2,5,8
max_rows = sorted_df.iloc[2::3].reset_index(drop=True)

# 合并结果
combined = pd.concat([min_rows, mid_rows, max_rows]).sort_values("ID").reset_index(drop=True)

方案2:利用rank标记筛选

通过计算组内排名(仅用groupby计算排名,筛选逻辑不依赖groupby):

# 计算每个ID内B的升序排名
df['b_rank'] = df.groupby("ID")["B"].rank(method='first').astype(int)
# 筛选排名1、2、3的行
result = df[df['b_rank'].isin([1,2,3])].drop(columns='b_rank').sort_values(["ID", "B"]).reset_index(drop=True)

内容的提问来源于stack exchange,提问作者Maria C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 15:43:11