如何使用spacy计算DataFrame两列对应句子的相似度
解决方案
以下是可直接运行的实现代码:
import pandas as pd import spacy_sentence_bert # 加载模型,仅需加载一次,避免重复消耗资源 nlp = spacy_sentence_bert.load_model('en_roberta_large_nli_stsb_mean_tokens') def get_similarity(row): # 分别处理两列的句子计算相似度 sent1_doc = nlp(row['Sentence1']) sent2_doc = nlp(row['Sentence2']) return sent1_doc.similarity(sent2_doc) # 假设你的DataFrame变量名为df,直接应用函数生成新列 df['Similarity-Score-Sentence1-Sentence2'] = df.apply(get_similarity, axis=1)
优化建议
- 数据量较大时,可使用
swifter库替代原生apply,自动判断是否使用并行加速,减少运行耗时 - 若两列存在大量重复句子,可提前将所有唯一句子转换为向量缓存,计算相似度时直接调用缓存向量计算,避免重复处理相同文本
内容的提问来源于stack exchange,提问作者user1624562
相关产品推荐
相关产品推荐

