You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

dedupe库使用报错求助:实现相似度≥80%的姓名匹配失败

问题:使用dedupe库匹配高相似度姓名时出现KeyError

我正在学习dedupe库,目标是匹配相似度超过80%的姓名,以下是我的实现代码及报错信息:

原代码

import dedupe
from Levenshtein import distance

def test():


    # 示例数据(替换为实际数据)
    data = [
        {'name': 'Alice Smith', 'address': '123 Main St', 'phone': '555-1212'},
        {'name': 'Alice SmIth', 'address': '123 Main Street', 'phone': '555-1213'},
        {'name': 'Bob Johnson', 'address': '456 Elm St', 'phone': '555-3434'},
        {'name': 'Charlie Brown', 'address': '789 Maple Ave', 'phone': '555-5656'},
    ]

    # 定义比较字段
    fields = [
        {'field': 'name', 'comparators': ['name_similarity']},
    ]

    # 自定义姓名相似度计算函数
    def name_similarity(s1, s2):
        distance1 = distance(s1, s2)
        similarity = 1 - (distance1 / max(len(s1), len(s2)))  # 归一化到0-1范围
        return similarity

    # 设置相似度阈值并初始化去重器
    deduper = dedupe.Dedupe(fields)
    deduper.threshold( threshold=0.8)

    # 执行去重
    deduped_data = deduper.dedupe(data)

    # 输出结果
    print("去重后的数据:")
    for cluster in deduped_data:
        print(cluster)


if __name__ == '__main__':
    test()

报错信息

C:\PythonProject\pythonProject\venv\Graph_POC\Scripts\python.exe C:\PythonProject\pythonProject\matching.py  
Traceback (most recent call last):
  File "C:\PythonProject\pythonProject\venv\Graph_POC\Lib\site-packages\dedupe\datamodel.py", line 152, in typify_variables
    variable_type = definition["type"]
                    ~~~~~~~~~~^^^^^^^^
KeyError: 'type'

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "C:\PythonProject\pythonProject\matching.py", line 45, in <module>
    test()
  File "C:\PythonProject\pythonProject\matching.py", line 32, in test
    deduper = dedupe.Dedupe(fields)
              ^^^^^^^^^^^^^^^^^^^^^
  File "C:\PythonProject\pythonProject\venv\Graph_POC\Lib\site-packages\dedupe\api.py", line 1155, in __init__
    self.data_model = datamodel.DataModel(variable_definition)
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\PythonProject\pythonProject\venv\Graph_POC\Lib\site-packages\dedupe\datamodel.py", line 42, in __init__
    self.primary_variables, all_variables = typify_variables(variable_definitions)
                                            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\PythonProject\pythonProject\venv\Graph_POC\Lib\site-packages\dedupe\datamodel.py", line 161, in typify_variables
    raise KeyError(
KeyError: "Missing variable type: variable specifications are dictionaries that must include a type definition, ex. {'field' : 'Phone', type: 'String'}"

Process finished with exit code 1

错误分析与解决方案

错误点

  1. 字段定义缺失type:dedupe要求每个字段定义必须包含type键,明确字段的数据类型,原代码的fields中未指定。
  2. 自定义比较器使用错误:原代码用字符串'name_similarity'指定比较器,dedupe无法识别,需直接传入函数对象。
  3. 未完成模型训练流程:dedupe是机器学习类库,初始化后必须先完成采样、标注、训练步骤,才能执行去重操作。

修正后的代码

import dedupe
from Levenshtein import distance

def name_similarity(s1, s2):
    # 优化相似度计算:忽略大小写,处理空值
    if not s1 or not s2:
        return 0.0
    distance_val = distance(s1.lower(), s2.lower())
    max_len = max(len(s1), len(s2))
    similarity = 1 - (distance_val / max_len)
    return similarity

def test():
    # 示例数据
    data = [
        {'name': 'Alice Smith', 'address': '123 Main St', 'phone': '555-1212'},
        {'name': 'Alice SmIth', 'address': '123 Main Street', 'phone': '555-1213'},
        {'name': 'Bob Johnson', 'address': '456 Elm St', 'phone': '555-3434'},
        {'name': 'Charlie Brown', 'address': '789 Maple Ave', 'phone': '555-5656'},
    ]

    # 字段定义:补充type,直接传入自定义比较器函数
    fields = [
        {'field': 'name', 'type': 'String', 'comparators': [name_similarity]},
    ]

    # 初始化Dedupe对象
    deduper = dedupe.Dedupe(fields)

    # 采样训练数据(可根据数据量调整样本数)
    sample = deduper.sample(data, 10)

    # 手动标注训练对(实际场景可使用dedupe.console_label(deduper, sample)交互式标注)
    labeled_pairs = [
        ((data[0], data[1]), True),
        ((data[0], data[2]), False),
        ((data[1], data[2]), False),
    ]
    deduper.mark_pairs(labeled_pairs)

    # 训练模型
    deduper.train()

    # 设置相似度阈值为0.8
    deduper.threshold = 0.8

    # 执行去重
    deduped_data = deduper.dedupe(data)

    # 输出结果
    print("去重后的数据集群:")
    for cluster in deduped_data:
        print(cluster)

if __name__ == '__main__':
    test()

说明

  • 补充了type: 'String'字段,解决KeyError问题。
  • 优化了姓名相似度计算逻辑,忽略大小写提升匹配准确性。
  • 完善了dedupe的训练流程,这是使用该类库的必要步骤。
  • 设置阈值为0.8,满足相似度超过80%的匹配需求。

内容的提问来源于stack exchange,提问作者pbh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 00:54:53