You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何匹配两个DataFrame中存在差异的同源文本片段?

匹配DataFrame片段与源句的高效方案

针对你遇到的DataFrame片段匹配问题,这里提供一个高效的解决思路,利用全文本拼接+位置映射的方式,既能处理跨句片段,又能适配大数据量场景:

核心思路

  1. 将DataFrame A的所有句子拼接成一个完整文本,同时记录每个源句在这个大文本中的起止位置和对应的索引。
  2. 对DataFrame B中的每个片段,在完整文本中定位其位置,再通过位置重叠关系反推对应的源句索引列表。

代码实现

import pandas as pd

# 初始化示例数据(实际使用时替换为你的DataFrame)
df_a = pd.DataFrame({
    'sentence': ["I was in the river.", "How was the water?", "Good, but cold"]
})
df_b = pd.DataFrame({
    'sentence': ["was in the river.", "was the water? Good", "Good, but co", "Good, but cold"]
})

# 构建A的全文本与位置映射
full_text = ""
sentence_positions = []
for idx_a, row in df_a.iterrows():
    start = len(full_text)
    full_text += row['sentence']
    end = len(full_text)
    sentence_positions.append({
        'idx_a': idx_a,
        'start': start,
        'end': end
    })
pos_df = pd.DataFrame(sentence_positions)

# 定义匹配函数
def get_matching_indices(b_sent, full_text, pos_df):
    # 定位片段在全文本中的起始位置
    start_idx = full_text.find(b_sent)
    end_idx = start_idx + len(b_sent)
    # 筛选所有与片段位置有重叠的源句
    matches = pos_df[
        (pos_df['start'] < end_idx) & (pos_df['end'] > start_idx)
    ]
    return matches['idx_a'].tolist()

# 生成匹配结果
df_b['idx_A'] = df_b['sentence'].apply(lambda x: get_matching_indices(x, full_text, pos_df))
result = df_b.reset_index().rename(columns={'index': 'idx_B'})[['idx_B', 'idx_A']]

print(result)

方案优势

  • 高效性:仅需一次文本拼接,后续每个片段仅需一次字符串查找,避免了逐句比对的O(n*m)复杂度,适合大数据量场景。
  • 兼容性:天然支持跨多个源句的片段匹配,比如示例中覆盖A第1、2句的片段能准确返回对应的索引列表。
  • 可靠性:题目明确B的片段均来自A的内容,因此无需处理匹配失败的异常情况。

内容的提问来源于stack exchange,提问作者lamjunioor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 09:22:08