使用Dedupe库数据去重遇TypeError错误,请求技术帮助
我正在学习Dedupe库,运行以下示例代码时,执行到deduper.prepare_training(data)步骤出现错误:
import dedupe from Levenshtein import distance # 定义相似度函数 - 根据匹配规则自定义 def name_similarity(s1, s2): # 实现姓名比较逻辑(比如Levenshtein距离) distance1 = distance(s1, s2) similarity = 1 - (distance1 / max(len(s1), len(s2))) # 将距离归一化为0-1的相似度 return similarity if __name__ == '__main__': # 示例数据(字典列表) data = {18709931: {'id': '18709931', 'name': 'TEST', 'ent_num': '8256364', 'ent_nm_txt': 'TST Corporation'}, 18484906: {'id': '18484906', 'name': 'VESTCOM', 'ent_num': '8256364', 'ent_nm_txt': 'TST Corporation'}, 18709961: {'id': '18709961', 'name': 'TESTMATERIALS', 'ent_num': '8256364', 'ent_nm_txt': 'TST Corporation'}, 19415694: {'id': '19415694', 'name': 'TEST', 'ent_num': '8256364', 'ent_nm_txt': 'TST Corporation'}} # 定义schema fields = [ {'field': 'name', 'type': 'Custom', 'comparator': name_similarity}, {'field': 'ent_num', 'type': 'Exact'}, ] # 初始化deduper deduper = dedupe.Dedupe(fields) # 准备训练数据(报错位置) deduper.prepare_training(data) # 交互式标注示例 dedupe.console_label(deduper) # 训练模型 deduper.train() # 保存训练好的模型到磁盘 with open('dedupe_model.pickle', 'wb') as f: dedupe.pickle.dump(deduper, f)
错误信息:
Traceback (most recent call last):
File "C:\Python_Projects\Python_extra_code\test_dedupe_code.py", line 30, in
deduper.prepare_training(data)
File "C:\Dev\Python3.11\Lib\site-packages\dedupe\api.py", line 1424, in prepare_training
self.active_learner = labeler.DedupeDisagreementLearner(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Dev\Python3.11\Lib\site-packages\dedupe\labeler.py", line 430, in init
self.mark(examples, labels)
File "C:\Dev\Python3.11\Lib\site-packages\dedupe\labeler.py", line 391, in mark
learner.fit(self.pairs, self.y)
File "C:\Dev\Python3.11\Lib\site-packages\dedupe\labeler.py", line 117, in fit
self.current_predicates = self.block_learner.learn(
^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Dev\Python3.11\Lib\site-packages\dedupe\training.py", line 58, in learn
coverable_dupes = frozenset.union(*match_cover.values())
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
TypeError: unbound method frozenset.union() needs an argumentProcess finished with exit code 1
问题原因
这个错误的核心是match_cover.values()为空集合,导致frozenset.union()没有可传入的参数。出现这种情况的主要原因是:
- 样本数据中所有条目的
ent_num完全相同,Dedupe基于Exact类型字段生成的blocker无法筛选出有效候选对,导致没有足够的训练样本对用于后续处理。 - 样本数据量过小,无法生成符合要求的训练候选对。
解决方案
添加更多多样化样本:
在data中加入一些ent_num不同的条目,让blocker能生成不同的候选组,比如:data = { # 原有数据 18709931: {'id': '18709931', 'name': 'TEST', 'ent_num': '8256364', 'ent_nm_txt': 'TST Corporation'}, 18484906: {'id': '18484906', 'name': 'VESTCOM', 'ent_num': '8256364', 'ent_nm_txt': 'TST Corporation'}, 18709961: {'id': '18709961', 'name': 'TESTMATERIALS', 'ent_num': '8256364', 'ent_nm_txt': 'TST Corporation'}, 19415694: {'id': '19415694', 'name': 'TEST', 'ent_num': '8256364', 'ent_nm_txt': 'TST Corporation'}, # 新增不同ent_num的条目 20000001: {'id': '20000001', 'name': 'ABC Corp', 'ent_num': '1234567', 'ent_nm_txt': 'ABC Corporation'}, 20000002: {'id': '20000002', 'name': 'ABC Company', 'ent_num': '1234567', 'ent_nm_txt': 'ABC Corporation'} }手动指定训练样本:
如果暂时无法添加更多数据,可以跳过自动生成训练对的步骤,手动提供标注好的训练样本,示例如下:# 初始化deduper后,手动定义训练样本 labeled_examples = { 'match': [ (data[18709931], data[19415694]), (data[18709931], data[18709961]) ], 'distinct': [ (data[18709931], data[18484906]), (data[18484906], data[18709961]) ] } deduper.prepare_training(data, labeled_examples=labeled_examples)调整Schema配置:
可以给ent_num字段添加has missing参数,或者调整blocker策略,避免因单一字段完全相同导致的候选对生成失败:fields = [ {'field': 'name', 'type': 'Custom', 'comparator': name_similarity}, {'field': 'ent_num', 'type': 'Exact', 'has missing': True}, ]
内容的提问来源于stack exchange,提问作者pbh

