You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按索引重复次数拆分pandas DataFrame为多个子DataFrame?

按索引出现次数拆分Pandas DataFrame

首先给出问题中的DataFrame构建代码:

import pandas as pd

a = [0.00000, 0.071928, 1.294, 2.592563, 0.000318, 2.575291, 0.439986, 2.232147, 6.091523, 2.075441, 0.96152]
b = [0.00000, 0.399791, 1.302446, 1.388957, 1.276451, 1.527568, 1.614107, 2.686325, 4.167600, 6.135689, 5.945807]

df = pd.DataFrame({'a' : a, 'b' : b})
df.index = [1,1,1,1,1,2,2,3,3,3,4]

需求:将每个索引值对应的行按出现顺序拆分,所有索引的第1次出现行归入df1,第2次归入df2,以此类推。


解决方案

核心思路是先给每个索引组内的行添加组内计数标记,再根据标记拆分。

步骤1:添加组内计数列

使用groupby(level=0).cumcount()对每个索引值的行从0开始计数(第1次出现为0,第2次为1,依此类推):

# 添加组内计数列,命名为occurrence
df['occurrence'] = df.groupby(level=0).cumcount()

步骤2:拆分DataFrame

方式1:自动生成所有拆分后的DataFrame(推荐)

适合出现次数较多的场景,用字典推导式按计数自动生成对应DataFrame:

# 按occurrence分组,生成字典,键为df1/df2...,值为对应数据(移除计数列)
dfs = {f'df{count+1}': group.drop('occurrence', axis=1) for count, group in df.groupby('occurrence')}

此时dfs['df1']对应所有索引第1次出现的行,dfs['df2']对应第2次,以此类推。

方式2:手动提取指定次数的行

如果只需要提取前几次的行,直接筛选即可:

# df1:所有索引第1次出现的行
df1 = df[df['occurrence'] == 0].drop('occurrence', axis=1)
# df2:所有索引第2次出现的行
df2 = df[df['occurrence'] == 1].drop('occurrence', axis=1)
# df3:所有索引第3次出现的行
df3 = df[df['occurrence'] == 2].drop('occurrence', axis=1)
# df4:所有索引第4次出现的行
df4 = df[df['occurrence'] == 3].drop('occurrence', axis=1)
# df5:所有索引第5次出现的行
df5 = df[df['occurrence'] == 4].drop('occurrence', axis=1)

结果验证

以df1为例,输出内容如下:

a         b
1  0.000000  0.000000
2  2.575291  1.527568
3  2.232147  2.686325
4  0.961520  5.945807

df2的内容:

a         b
1  0.071928  0.399791
2  0.439986  1.614107
3  6.091523  4.167600

完全符合需求。


原方法失效原因

df.duplicated(subset=['index'])是判断全局范围内索引值是否重复,返回的是全局重复标记,无法区分每个索引组内的出现次数(比如索引1的第2-5行都会被标记为重复,但无法区分是第2次还是第3次出现),因此无法按需求拆分。

内容的提问来源于stack exchange,提问作者Volti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 19:18:19