You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于条件高效合并Pandas DataFrame的相邻行

相邻簇合并实现(合并后簇内观测数≥1000)

问题说明

现有数据集包含3个字段:

  • clust:原始簇编号
  • obs:簇内观测样本量
  • x_max:簇内变量x的最大值

原始数据构造代码

import numpy as np
import pandas as pd

clust = np.arange(0, 10)
obs = np.array([1041, 544, 310, 1648, 1862, 2120, 2916, 5148, 12733, 1])
x_max = np.array([10, 20, 30, 40, 50, 60, 70, 80, 90, 100])
df = pd.DataFrame(np.c_[clust, obs, x_max], columns=['clust', 'obs', 'x_max'])

原始数据样例

clustobsx_max
0104110
154420
231030
3164840
4186250
5212060
6291670
7514880
81273390
91100

需求与现存问题

需求为合并相邻的原始簇行,保证合并后每个新簇的obs值不低于1000。
原有循环实现存在逻辑缺陷:遍历到最后一个仅含1条观测的簇时,始终无法满足obs≥1000的条件,会陷入无限死循环,错误代码如下:

condition = True
while (condition):
    condition = False
    for i in np.arange(0, len(df) + 1):
        if df.loc[i, 'x'] < 1000:
            df.loc[i, 'id'] = df.loc[i, 'id'] + 1
            df = df.groupby('id', as_index=False).agg({'x': 'sum', 'y': 'max'})
            condition = True
            break

类似SAS中if last.row的逐行边界处理逻辑,pandas没有原生同名语法,需要手动实现分组边界判断,避免末尾孤立小簇导致的死循环。

预期输出结果

clustobsx_max
0104110
1250240
2186250
3212060
4291670
5514880
612734100

高效实现方案

采用单次遍历生成分组标签的方式实现,时间复杂度O(n),无需反复重构DataFrame,从逻辑上避免末尾小簇导致的死循环:

  1. 逐行累加obs值,累加值达到1000阈值时切分新组
  2. 遍历到最后一行时不提前切分,将剩余不足1000的样本直接归入上一个相邻组
  3. 按生成的分组标签聚合,obs求和、x_max取组内最大值,最后重排新簇编号
    完整代码如下:
# 初始化分组参数
group_id = 0
current_obs_sum = 0
group_labels = []
total_rows = len(df)

for idx, o in enumerate(df['obs']):
    group_labels.append(group_id)
    current_obs_sum += o
    # 非最后一行且累加值达标时,切换到下一组
    if current_obs_sum >= 1000 and idx != total_rows - 1:
        group_id += 1
        current_obs_sum = 0

# 分组聚合得到结果
result = df.groupby(group_labels, as_index=False).agg(
    obs=('obs', 'sum'),
    x_max=('x_max', 'max')
).reset_index(names='clust')

运行代码后输出结果与预期完全一致。

内容的提问来源于stack exchange,提问作者Serge Kashlik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 20:01:22