dedupe库使用报错求助:实现相似度≥80%的姓名匹配失败
问题:使用dedupe库匹配高相似度姓名时出现KeyError
我正在学习dedupe库,目标是匹配相似度超过80%的姓名,以下是我的实现代码及报错信息:
原代码
import dedupe from Levenshtein import distance def test(): # 示例数据(替换为实际数据) data = [ {'name': 'Alice Smith', 'address': '123 Main St', 'phone': '555-1212'}, {'name': 'Alice SmIth', 'address': '123 Main Street', 'phone': '555-1213'}, {'name': 'Bob Johnson', 'address': '456 Elm St', 'phone': '555-3434'}, {'name': 'Charlie Brown', 'address': '789 Maple Ave', 'phone': '555-5656'}, ] # 定义比较字段 fields = [ {'field': 'name', 'comparators': ['name_similarity']}, ] # 自定义姓名相似度计算函数 def name_similarity(s1, s2): distance1 = distance(s1, s2) similarity = 1 - (distance1 / max(len(s1), len(s2))) # 归一化到0-1范围 return similarity # 设置相似度阈值并初始化去重器 deduper = dedupe.Dedupe(fields) deduper.threshold( threshold=0.8) # 执行去重 deduped_data = deduper.dedupe(data) # 输出结果 print("去重后的数据:") for cluster in deduped_data: print(cluster) if __name__ == '__main__': test()
报错信息
C:\PythonProject\pythonProject\venv\Graph_POC\Scripts\python.exe C:\PythonProject\pythonProject\matching.py Traceback (most recent call last): File "C:\PythonProject\pythonProject\venv\Graph_POC\Lib\site-packages\dedupe\datamodel.py", line 152, in typify_variables variable_type = definition["type"] ~~~~~~~~~~^^^^^^^^ KeyError: 'type' During handling of the above exception, another exception occurred: Traceback (most recent call last): File "C:\PythonProject\pythonProject\matching.py", line 45, in <module> test() File "C:\PythonProject\pythonProject\matching.py", line 32, in test deduper = dedupe.Dedupe(fields) ^^^^^^^^^^^^^^^^^^^^^ File "C:\PythonProject\pythonProject\venv\Graph_POC\Lib\site-packages\dedupe\api.py", line 1155, in __init__ self.data_model = datamodel.DataModel(variable_definition) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\PythonProject\pythonProject\venv\Graph_POC\Lib\site-packages\dedupe\datamodel.py", line 42, in __init__ self.primary_variables, all_variables = typify_variables(variable_definitions) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\PythonProject\pythonProject\venv\Graph_POC\Lib\site-packages\dedupe\datamodel.py", line 161, in typify_variables raise KeyError( KeyError: "Missing variable type: variable specifications are dictionaries that must include a type definition, ex. {'field' : 'Phone', type: 'String'}" Process finished with exit code 1
错误分析与解决方案
错误点
- 字段定义缺失
type:dedupe要求每个字段定义必须包含type键,明确字段的数据类型,原代码的fields中未指定。 - 自定义比较器使用错误:原代码用字符串
'name_similarity'指定比较器,dedupe无法识别,需直接传入函数对象。 - 未完成模型训练流程:dedupe是机器学习类库,初始化后必须先完成采样、标注、训练步骤,才能执行去重操作。
修正后的代码
import dedupe from Levenshtein import distance def name_similarity(s1, s2): # 优化相似度计算:忽略大小写,处理空值 if not s1 or not s2: return 0.0 distance_val = distance(s1.lower(), s2.lower()) max_len = max(len(s1), len(s2)) similarity = 1 - (distance_val / max_len) return similarity def test(): # 示例数据 data = [ {'name': 'Alice Smith', 'address': '123 Main St', 'phone': '555-1212'}, {'name': 'Alice SmIth', 'address': '123 Main Street', 'phone': '555-1213'}, {'name': 'Bob Johnson', 'address': '456 Elm St', 'phone': '555-3434'}, {'name': 'Charlie Brown', 'address': '789 Maple Ave', 'phone': '555-5656'}, ] # 字段定义:补充type,直接传入自定义比较器函数 fields = [ {'field': 'name', 'type': 'String', 'comparators': [name_similarity]}, ] # 初始化Dedupe对象 deduper = dedupe.Dedupe(fields) # 采样训练数据(可根据数据量调整样本数) sample = deduper.sample(data, 10) # 手动标注训练对(实际场景可使用dedupe.console_label(deduper, sample)交互式标注) labeled_pairs = [ ((data[0], data[1]), True), ((data[0], data[2]), False), ((data[1], data[2]), False), ] deduper.mark_pairs(labeled_pairs) # 训练模型 deduper.train() # 设置相似度阈值为0.8 deduper.threshold = 0.8 # 执行去重 deduped_data = deduper.dedupe(data) # 输出结果 print("去重后的数据集群:") for cluster in deduped_data: print(cluster) if __name__ == '__main__': test()
说明
- 补充了
type: 'String'字段,解决KeyError问题。 - 优化了姓名相似度计算逻辑,忽略大小写提升匹配准确性。
- 完善了dedupe的训练流程,这是使用该类库的必要步骤。
- 设置阈值为0.8,满足相似度超过80%的匹配需求。
内容的提问来源于stack exchange,提问作者pbh
相关产品推荐
相关产品推荐

