You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在分组的Pandas DataFrame中识别并标记事件重叠

解决Pandas多组场景下的事件重叠标记问题

问题背景

需要在带分组的Pandas DataFrame中提取组内重叠事件:

  • 同一组内的重叠事件用overlap列标记为同一组
  • 生成time_start_overlap和time_end_overlap列,显示该重叠组的时间范围
  • 不同组的事件即使时间重叠,也不能被标记为同一重叠组

原代码在单组场景有效,但多组时会跨组标记重叠(例如不同组的事件被分配同一个overlap ID),需修复。

示例数据

import pandas as pd

df = pd.DataFrame({
    'id': [1, 2, 3, 4, 5, 6, 7, 8],
    'group': ['a', 'a', 'a', 'a', 'a', 'b', 'b', 'b'],
    'event': ['event1', 'event2', 'event3', 'event4', 'event5', 'event6', 'event7', 'event8'],
    'time_start': ['2000-01-01 08:00:00',
                   '2000-01-01 07:30:00',
                   '2000-01-01 11:00:00',
                   '2000-01-01 12:30:00',
                   '2000-01-01 13:00:00',
                   '2000-01-01 08:00:00',
                   '2000-01-01 07:30:00',
                   '2000-01-01 08:30:00'
                   ],
    'time_end': ['2000-01-01 09:00:00',
                 '2000-01-01 10:30:00',
                 '2000-01-01 12:00:00',
                 '2000-01-01 13:30:00',
                 '2000-01-01 14:00:00',
                 '2000-01-01 09:00:00',
                 '2000-01-01 09:30:00',
                 '2000-01-01 10:00:00'
                 ]
})

原代码问题分析

原代码的核心错误是全局创建IntervalIndex并判断重叠,没有按group分组处理。这导致不同组的事件即使属于不同分组,只要时间重叠就会被标记为同一个overlap ID,违反了"不同组事件不跨组标记"的要求。

修正后的代码

# 先将时间列转为datetime类型,确保时间计算准确
df['time_start'] = pd.to_datetime(df['time_start'])
df['time_end'] = pd.to_datetime(df['time_end'])

def process_group(g):
    # 组内创建IntervalIndex
    iix = pd.IntervalIndex.from_arrays(g['time_start'], g['time_end'], closed='both')
    # 初始化overlap列为组内索引
    g['overlap'] = g.index
    # 组内当前最大的overlap ID,从组长度+1开始
    current_overlap_id = len(g) + 1
    
    # 遍历组内每个区间,找到重叠的事件组
    processed = set()
    for idx in g.index:
        if idx in processed:
            continue
        # 获取当前区间的所有重叠项
        overlaps = iix.overlaps(iix[idx])
        # 如果重叠项数量大于1(即有重叠)
        if overlaps.sum() > 1:
            # 将这些重叠项标记为同一个overlap ID
            g.loc[overlaps, 'overlap'] = current_overlap_id
            # 标记已处理的索引
            processed.update(g[overlaps].index)
            current_overlap_id += 1
    return g

# 按group分组处理,合并结果
df_processed = df.groupby('group', group_keys=False).apply(process_group)

# 计算每个重叠组的时间范围
overlap_ranges = df_processed.groupby(['group', 'overlap']).agg(
    time_start_overlap=('time_start', 'min'),
    time_end_overlap=('time_end', 'max')
).reset_index()

# 合并回原数据
final_df = pd.merge(df_processed, overlap_ranges, on=['group', 'overlap'])

print(final_df)

代码说明

  1. 时间类型转换:先将time_start和time_end转为datetime类型,避免字符串比较导致的错误。
  2. 分组处理函数:对每个group单独处理,确保重叠判断仅在组内进行,不会跨组干扰。
  3. 重叠标记逻辑:组内遍历每个事件,找到所有重叠事件并分配同一个overlap ID,已处理的事件跳过,避免重复标记。
  4. 计算重叠范围:按group和overlap分组,取最小的time_start和最大的time_end作为该重叠组的时间范围。

最终效果

  • group 'a'内:event1和event2标记为同一overlap组,event4和event5标记为另一组,event3无重叠保持独立ID
  • group 'b'内:event6、event7、event8标记为同一overlap组
  • 不同组的overlap ID不会重复,完全隔离

内容的提问来源于stack exchange,提问作者Joysn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 02:05:58