You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中筛选相似前缀但number列值不同的行?

筛选Pandas DataFrame中前缀相似但number值不同的行

首先还原你的测试DataFrame:

import pandas as pd

df_test = pd.DataFrame(data=None, columns=['file', 'number'])
df_test.file = ['washington_142', 'washington_287', 'chicago_453', 'chicago_221', 'chicago_345', 'seattle_976', 'seattle_977', 'boston_367', 'boston 098']
df_test.number = [20, 21, 33, 34, 33, 45, 45, 52, 52]

需求明确:找出file列字符串相似度≥50%,但对应number列值不同的行。

解决方案

使用difflib.SequenceMatcher计算字符串相似度,遍历所有行对筛选符合条件的行:

  1. 导入依赖库
from difflib import SequenceMatcher
  1. 定义相似度计算函数
def calculate_similarity(str1, str2):
    # 返回0-1之间的相似度值,1表示完全匹配
    return SequenceMatcher(None, str1, str2).ratio()
  1. 筛选符合条件的行索引
# 用集合存储符合条件的行索引,避免重复
target_indices = set()

# 遍历所有行对(i<j,避免重复对比)
for i in range(len(df_test)):
    for j in range(i + 1, len(df_test)):
        sim_score = calculate_similarity(df_test['file'].iloc[i], df_test['file'].iloc[j])
        # 满足相似度≥50%且number值不同
        if sim_score >= 0.5 and df_test['number'].iloc[i] != df_test['number'].iloc[j]:
            target_indices.add(i)
            target_indices.add(j)
  1. 生成结果DataFrame
result_df = df_test.iloc[list(target_indices)].sort_index()
print(result_df)

执行后输出结果:

file  number
0  washington_142      20
1  washington_287      21
2    chicago_453      33
3    chicago_221      34
4    chicago_345      33

说明

  • 相比difflib.get_close_matches,SequenceMatcher更适合逐行对比所有组合,能确保不遗漏任何符合条件的行对。
  • 用集合存储索引可以自动去重,避免同一行被多次添加。
  • 最后按原索引排序,保持结果的原始顺序。

内容的提问来源于stack exchange,提问作者Marcus K.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 07:52:50