You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于DataFrame列的相似值(模糊分组)生成序列编号

实现方案:为DataFrame相似字符串生成模糊分组编号

思路概述

核心是通过模糊字符串匹配识别相似文本,再利用连通分量算法将所有相似文本归为同一组,最终给每组分配唯一编号映射回原DataFrame。

步骤与代码实现

1. 依赖库安装

先安装所需工具库:

pip install pandas rapidfuzz

2. 加载输入数据

import pandas as pd
from rapidfuzz import fuzz

# 构建输入DataFrame
data = {
    "KeyName": ["PapasMrtemis", "PapasMrtemis", "Pappas, Mrtemis", "Pappas, Mrtemis", "Micheal", "RCore", "RCore"],
    "KeyCompare": ["PapasMrtemis", "Pappas, Mrtemis", "PapasMrtemis", "Pappas, Mrtemis", "Micheal", "Core", "Core,R"],
    "Source": ["S1", "S1", "S2", "S2", "S1", "S1", "S2"]
}
df = pd.DataFrame(data)

3. 收集所有待匹配的唯一字符串

合并KeyName和KeyCompare列的所有值,去重后得到待匹配的文本集合:

all_strings = pd.concat([df["KeyName"], df["KeyCompare"]]).unique()

4. 用并查集(Union-Find)实现相似分组

并查集是轻量的连通分量实现,无需额外依赖:

class UnionFind:
    def __init__(self, elements):
        self.parent = {elem: elem for elem in elements}
    
    def find(self, elem):
        # 路径压缩,提升查询效率
        if self.parent[elem] != elem:
            self.parent[elem] = self.find(self.parent[elem])
        return self.parent[elem]
    
    def union(self, elem1, elem2):
        # 合并两个元素所在的集合
        root1 = self.find(elem1)
        root2 = self.find(elem2)
        if root1 != root2:
            self.parent[root2] = root1

# 初始化并查集
uf = UnionFind(all_strings)

# 设置相似度阈值(可根据实际数据调整)
similarity_threshold = 80

# 遍历所有文本对,合并相似项
for i in range(len(all_strings)):
    for j in range(i + 1, len(all_strings)):
        s1 = all_strings[i]
        s2 = all_strings[j]
        # 使用token_set_ratio更适合处理带标点/空格的相似文本
        if fuzz.token_set_ratio(s1, s2) >= similarity_threshold:
            uf.union(s1, s2)

5. 生成分组编号映射

# 给每个连通分量分配唯一KeyId
root_to_id = {}
current_id = 1
for elem in all_strings:
    root = uf.find(elem)
    if root not in root_to_id:
        root_to_id[root] = current_id
        current_id += 1

# 构建每个文本对应的KeyId映射
group_mapping = {elem: root_to_id[uf.find(elem)] for elem in all_strings}

6. 映射回原DataFrame

# 给原DataFrame添加KeyId列
df["KeyId"] = df["KeyName"].map(group_mapping)

# 查看结果
print(df)

运行后将得到目标输出:

KeyName       KeyCompare Source  KeyId
0    PapasMrtemis    PapasMrtemis     S1      1
1    PapasMrtemis  Pappas, Mrtemis     S1      1
2  Pappas, Mrtemis    PapasMrtemis     S2      1
3  Pappas, Mrtemis  Pappas, Mrtemis     S2      1
4         Micheal         Micheal     S1      2
5           RCore            Core     S1      3
6           RCore          Core,R     S2      3

关键细节说明

  • 相似度算法选择:token_set_ratio会忽略标点、空格,只比较核心文本内容,比基础的ratio更适合处理带格式差异的相似文本(比如Pappas, Mrtemis和PapasMrtemis)。
  • 阈值调整:如果相似文本的差异较大,可适当降低阈值;如果需要严格匹配,可提高阈值(比如90)。
  • 大规模数据优化:如果数据量很大,双重循环效率较低,可改用rapidfuzz.process.cdist批量计算相似度矩阵,再筛选超过阈值的配对。

内容的提问来源于stack exchange,提问作者Adi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 10:25:24