You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效计算不同DataFrame短语行相似度并按阈值筛选?

计算两个DataFrame文本相似度的实现方案

核心需求

计算两个DataFrame的文本相似度:df1含150条特定短语,df2含25万条短语,用SequenceMatcher计算df1每行与df2每行的相似度,标记/筛选相似度高于阈值的df2行。

实现思路

  1. 生成df1与df2的笛卡尔积,确保每条df2短语都和df1的所有短语配对
  2. 对每一对短语调用SequenceMatcher计算相似度
  3. 整理结果格式,按需筛选阈值以上的行

高效优化建议

  • 去重df1短语:先对df1的Phrases列去重,减少重复计算次数(比如示例中df1有重复短语,去重后计算量减少25%)
  • 分块处理df2:25万条数据直接生成笛卡尔积会占用大量内存,分块加载计算可避免内存溢出
  • 并行计算:利用多进程/swifter库加速相似度计算,充分利用CPU资源
  • 替换高效算法:若对精度要求不高,可改用Jaccard相似度(基于词袋),计算速度远快于SequenceMatcher
  • 提前过滤短文本:对长度差异过大的短语直接判定相似度为0,跳过计算

代码示例

1. 基础实现(示例数据)

import pandas as pd
from difflib import SequenceMatcher

# 构造示例数据
df1 = pd.DataFrame({
    'Phrases': ['My cat is black', 'Dog is white', 'Peter is waiting', 'Dog is white']
})
df2 = pd.DataFrame({
    'Phrases': ['My cat is white', 'Dog is jumping', 'Marcos is waiting', 'Dog is white']
})

# 定义相似度计算函数
def calculate_similarity(text1, text2):
    return SequenceMatcher(None, text1, text2).ratio()

# 生成笛卡尔积并计算相似度
cross_df = df2.merge(df1, how='cross')
cross_df['Labels'] = cross_df.apply(lambda x: calculate_similarity(x['Phrases_x'], x['Phrases_y']), axis=1)

# 整理成预期输出格式
result_df = cross_df[['Phrases_x', 'Labels']].rename(columns={'Phrases_x': 'Phrases'})
# 保留两位小数匹配示例
result_df['Labels'] = result_df['Labels'].round(2)

print(result_df)

2. 大规模数据优化版(分块+去重)

import pandas as pd
from difflib import SequenceMatcher
import swifter

# df1去重,减少计算量
unique_df1 = df1.drop_duplicates(subset='Phrases').reset_index(drop=True)

# 定义相似度计算函数
def calculate_similarity(text1, text2):
    return SequenceMatcher(None, text1, text2).ratio()

# 分块处理df2(假设df2存储在csv文件中)
chunk_size = 10000
result_chunks = []

for chunk in pd.read_csv('df2.csv', chunksize=chunk_size):
    # 笛卡尔积
    cross_chunk = chunk.merge(unique_df1, how='cross')
    # 用swifter加速apply
    cross_chunk['Labels'] = cross_chunk.swifter.apply(lambda x: calculate_similarity(x['Phrases_x'], x['Phrases_y']), axis=1)
    # 整理格式
    result_chunk = cross_chunk[['Phrases_x', 'Labels']].rename(columns={'Phrases_x': 'Phrases'})
    result_chunk['Labels'] = result_chunk['Labels'].round(2)
    result_chunks.append(result_chunk)

# 合并所有分块结果
final_result = pd.concat(result_chunks, ignore_index=True)

# 筛选相似度≥0.5的行
filtered_result = final_result[final_result['Labels'] >= 0.5]

3. Jaccard相似度替代方案(更快)

def jaccard_similarity(text1, text2):
    set1 = set(text1.split())
    set2 = set(text2.split())
    intersection = len(set1 & set2)
    union = len(set1 | set2)
    return intersection / union if union != 0 else 0

# 替换calculate_similarity函数即可使用

结果说明

运行基础实现代码后,输出结果与预期一致:每条df2短语会对应df1所有短语的相似度值,例如My cat is white与df1四条短语的相似度分别为0.75、0.00、0.00、0.33。

内容的提问来源于stack exchange,提问作者D EA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 20:45:23