You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何查找两个dataframe行区间的最大重叠并完成配对与排名?

DataFrame区间重叠匹配实现

需求描述

给定两个均包含start、end区间字段的DataFrame,需要完成两个功能:

  1. 为df_1的每一行匹配区间重叠度最高的df_2行,输出配对结果
  2. 支持对df_1每行的所有可能配对按重叠度降序排名,可自定义返回前n个重叠度最高的结果

示例数据

import pandas as pd

# 构建df_1
df_1 = pd.DataFrame(
    {'start': [1, 20, 70], 'end': [10, 50, 100]},
    index=['a', 'b', 'c']
)
print("df_1:")
print(df_1)

# 构建df_2
df_2 = pd.DataFrame(
    {'start': [5, 70, 20], 'end': [10, 120, 30]},
    index=['a', 'b', 'c']
)
print("\ndf_2:")
print(df_2)

核心计算逻辑

两个区间的重叠度计算规则:

  • 重叠长度 = max(0, min(df1_end, df2_end) - max(df1_start, df2_start))
  • 可选归一化处理:重叠长度 / df1区间总长度,可消除区间长度差异带来的排序影响

完整代码实现

def calculate_overlap(df1_row, df2_row):
    # 计算两个区间的重叠长度
    overlap_start = max(df1_row['start'], df2_row['start'])
    overlap_end = min(df1_row['end'], df2_row['end'])
    return max(0, overlap_end - overlap_start)

def match_top_overlap(df_1, df_2, top_n=1):
    result = []
    for df1_idx, df1_row in df_1.iterrows():
        # 计算当前df1行和所有df2行的重叠度
        overlap_list = []
        for df2_idx, df2_row in df_2.iterrows():
            overlap = calculate_overlap(df1_row, df2_row)
            overlap_list.append((df2_idx, overlap))
        # 按重叠度降序排序
        overlap_list.sort(key=lambda x: x[1], reverse=True)
        # 取前n个结果
        top_matches = overlap_list[:top_n]
        for rank, (df2_idx, overlap) in enumerate(top_matches, 1):
            result.append({
                'df_1_index': df1_idx,
                'df_2_index': df2_idx,
                'overlap_len': overlap,
                'rank': rank
            })
    res_df = pd.DataFrame(result)
    # 如果只取top1,返回简化的配对结果
    if top_n == 1:
        return res_df[['df_1_index', 'df_2_index']].rename(columns={'df_1_index':'df_1', 'df_2_index':'df_2'}).set_index('df_1')
    return res_df

# 测试1:取最高匹配(top1)
print("top1匹配结果:")
print(match_top_overlap(df_1, df_2, top_n=1))

# 测试2:取前2个匹配
print("\n前2个匹配结果:")
print(match_top_overlap(df_1, df_2, top_n=2))

输出示例

top1匹配输出

和需求给出的示例完全一致:

df_1    df_2
a       a
b       c
c       b

前2个匹配输出示例

df_1_index df_2_index  overlap_len  rank
0          a          a            5     1
1          a          c            0     2
2          b          c           10     1
3          b          a            0     2
4          c          b           30     1
5          c          c            0     2

内容的提问来源于stack exchange,提问作者user17005846

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 03:24:03