You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas中实现同列所有行的自定义规则相似度检测

问题

现有如下Pandas DataFrame:

Product
0 Checking savings account
1 Closing account
2 Debt collection
3 Credit reporting credit repair services personal consumer reports
4 Checking savings account

需要实现:

  • 遍历所有行两两组合计算相似度:索引0与其他4行逐一比较,索引1与包括0在内的其他4行逐一比较,以此类推
  • 自定义相似度规则:匹配的相同词数量 ÷ 较长句子的词数,结果以百分比呈现
  • 排除自身匹配(自匹配相似度100%无意义)

示例:

Checking savings account 与 Closing account 的相似度为33.3%——匹配词为"account"(数量1),较长句子词数为3,1/3=33.3%;
Checking savings account 与 Debt collection 的相似度为0%。

已尝试代码:

for i in df['Product']:
    compareItem = i.split()
    print(compareItem)
    for k in df['Product']:
        compareList = k.split()
        print(compareList)
    print('------')

注:此需求并非检测重复项,常规重复检测方案不适用,核心为上述自定义相似度规则。

解决方案

以下代码实现了自定义规则的相似度计算,同时排除自匹配:

import pandas as pd

# 构造示例DataFrame
data = {'Product': [
    'Checking savings account',
    'Closing account',
    'Debt collection',
    'Credit reporting credit repair services personal consumer reports',
    'Checking savings account'
]}
df = pd.DataFrame(data)

# 遍历所有索引对
for i in range(len(df)):
    sentence_i = df['Product'][i]
    words_i = sentence_i.split()
    len_i = len(words_i)
    
    for j in range(len(df)):
        # 跳过自身匹配
        if i == j:
            continue
        
        sentence_j = df['Product'][j]
        words_j = sentence_j.split()
        len_j = len(words_j)
        
        # 统计共同词数量(去重后)
        common_words = set(words_i) & set(words_j)
        match_count = len(common_words)
        
        # 取较长句子的词数
        max_sentence_len = max(len_i, len_j)
        
        # 计算相似度并转为百分比
        similarity = (match_count / max_sentence_len) * 100
        
        # 格式化输出结果
        print(f"索引{i}「{sentence_i}」与索引{j}「{sentence_j}」的相似度:{similarity:.1f}%")
    print('-' * 50)

关键说明

  1. 索引遍历:通过索引i和j遍历,方便直接判断并跳过i==j的自匹配场景
  2. 词数统计:保留原句子的词数(未去重),符合用户示例中"较长句子词数"的定义
  3. 共同词统计:利用集合交集快速获取两个句子的共同词(去重后),匹配数量符合示例逻辑
  4. 结果格式化:将相似度转换为百分比并保留1位小数,输出清晰直观

运行代码后,会输出所有有效两两组合的相似度结果,例如:

索引0「Checking savings account」与索引1「Closing account」的相似度:33.3%
索引0「Checking savings account」与索引2「Debt collection」的相似度:0.0%
索引0「Checking savings account」与索引3「Credit reporting credit repair services personal consumer reports」的相似度:0.0%
索引0「Checking savings account」与索引4「Checking savings account」的相似度:100.0%
--------------------------------------------------
...(后续索引的比较结果)

内容的提问来源于stack exchange,提问作者piseynir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 08:45:35