You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中遍历DataFrame两列并计算字符串相似度的实现方法

计算DataFrame每行字符串相似度得分的方法

你可以通过两种方式实现需求,推荐使用pandas的apply方法(更高效、符合pandas最佳实践),也可以手动循环遍历行:

方法一:使用apply(推荐)

直接对DataFrame的每行应用你的similar函数,无需手动遍历:

import pandas as pd
from difflib import SequenceMatcher

def similar(a, b):
    return SequenceMatcher(None, a, b).ratio()

# 假设你的DataFrame名为df
df['similarity_score'] = df.apply(lambda row: similar(row['Account_Name_HGI'], row['company_name_Ignite']), axis=1)

# 查看结果
print(df)

执行后会新增一列similarity_score,存储每行两个字符串的相似度得分。

方法二:手动循环遍历行

如果你确实需要手动循环,可以使用iterrows():

# 初始化空列表存储得分
scores = []

for index, row in df.iterrows():
    score = similar(row['Account_Name_HGI'], row['company_name_Ignite'])
    scores.append(score)

# 将得分添加到DataFrame
df['similarity_score'] = scores

注意:iterrows()的效率低于apply,处理大数据集时优先选择方法一。


内容的提问来源于stack exchange,提问作者Raul Gonzales

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 23:00:18