You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中支持重复词匹配的非集合字符串相似度计算方案

解决方案:统计重复词的字符串相似度计算

要解决重复词被去重的问题,我们需要基于词频统计来计算匹配数,而非依赖集合的交集操作。以下是符合需求的实现:

核心思路

  1. 将两个字符串拆分为词列表
  2. 统计目标字符串的词频(保留重复词的出现次数)
  3. 累加对比字符串中每个词在目标字符串中的出现次数,得到总匹配数
  4. 按照公式 (匹配词数 / 较长字符串的词数) * 100 计算相似度,保留两位小数

实现代码

首先导入必要的模块:

import pandas as pd
from collections import Counter

定义相似度计算函数:

def calculate_similarity(row1, row2):
    # 拆分字符串为词列表
    words1 = row1.split()
    words2 = row2.split()
    
    # 统计第二个字符串的词频
    word_counts = Counter(words2)
    
    # 计算总匹配词数:row1中每个词在row2中的出现次数之和
    match_count = sum(word_counts.get(word, 0) for word in words1)
    
    # 获取较长字符串的词数,避免除以0
    max_word_count = max(len(words1), len(words2))
    if max_word_count == 0:
        return 0.0
    
    # 计算相似度并保留两位小数
    similarity = (match_count / max_word_count) * 100
    return round(similarity, 2)

测试你的示例场景

  • 测试场景1:row1="Card",row2="Credit Card Debit Card"

    print(calculate_similarity("Card", "Credit Card Debit Card"))  # 输出 50.0
    

    匹配数为2,较长字符串词数为4,(2/4)*100=50.0,符合需求。

  • 测试场景2:row1="Credit Card",row2="Credit Card Debit Card"

    print(calculate_similarity("Credit Card", "Credit Card Debit Card"))  # 输出 75.0
    

    匹配数为3(Credit出现1次 + Card出现2次),较长字符串词数为4,(3/4)*100=75.0,符合需求。

结合你的DataFrame使用

加载你的示例数据并运行:

# 构建示例DataFrame
data = {
    'Product': ['Bank account or service', 'Credit card', 'Credit reporting', 'Credit reporting credit repair services or other personal consumer reports', 'Credit reporting', 'Mortgage', 'Debt collection', 'Mortgage', 'Mortgage', 'Credit reporting'],
    'Issue': ['Deposits and withdrawals', 'Billing disputes', 'Incorrect information on credit report', "Problem with a credit reporting company's investigation into an existing problem", 'Incorrect information on credit report', 'Applying for a mortgage or refinancing an existing mortgage', 'Disclosure verification of debt', 'Loan servicing payments escrow account', 'Loan servicing payments escrow account', 'Incorrect information on credit report'],
    'Company': ['CITIBANK NA', 'FIRST NATIONAL BANK OF OMAHA', 'EQUIFAX INC', 'Experian Information Solutions Inc', 'Experian Information Solutions Inc', 'BANK OF AMERICA NATIONAL ASSOCIATION', 'AllianceOne Recievables Management', 'SELECT PORTFOLIO SERVICING INC', 'OCWEN LOAN SERVICING LLC', 'Experian Information Solutions Inc'],
    'State': ['CA', 'WA', 'FL', 'UT', 'MI', 'CA', 'WA', 'IL', 'TX', 'CA'],
    'ZIP_code': ['92606', '98272', '329XX', '84321', '486XX', '94537', '984XX', '60473', '76247', '91401'],
    'Complaint_ID': [90452, 2334443, 1347696, 2914771, 1788024, 2871939, 1236424, 1619712, 2421373, 1803691]
}
dff = pd.DataFrame(data)

# 选择要测试的行索引
target_index = int(input("Which index should be tested:"))
target_product = dff['Product'].iloc[target_index]

# 计算所有行与目标行的相似度
dff['Similarity'] = dff['Product'].apply(lambda x: calculate_similarity(target_product, x))

# 打印结果
print(dff[['Product', 'Similarity']])

优化说明

对于大型CSV文件,使用apply比手动循环更高效。如果数据量极大,还可以考虑预先将所有Product列的词频统计好,进一步提升计算速度。

内容的提问来源于stack exchange,提问作者Berk Sunduri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 05:25:14