Python dedupe报错TypeError:列表索引应为整数而非字符串求解决
问题背景
原本参考MySQL示例实现dedupe对接数据库,但旧版MSSQL无json_object函数,自行编写record_pairs函数后运行报错。
错误栈信息
Process Process-1:
Traceback (most recent call last):
File "/usr/lib/python3.9/multiprocessing/process.py", line 315, in _bootstrap
self.run()
File "/usr/lib/python3.9/multiprocessing/process.py", line 108, in run
self._target(*self._args, **self._kwargs)
File "/usr/local/lib/python3.9/dist-packages/dedupe/core.py", line 80, in call
self.fieldDistance(record_pairs)
File "/usr/local/lib/python3.9/dist-packages/dedupe/core.py", line 91, in fieldDistance
distances = self.data_model.distances(records)
File "/usr/local/lib/python3.9/dist-packages/dedupe/datamodel.py", line 98, in distances
if record_1[field] is not None and record_2[field] is not None:
TypeError: list indices must be integers or slices, not str
TypeError: list indices must be integers or slices, not strThe above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/root/client/msql2.py", line 286, in
clustered_dupes = deduper.cluster(deduper.score(record_pairs(read_cur)),
File "/usr/local/lib/python3.9/dist-packages/dedupe/api.py", line 109, in score
matches = core.scoreDuplicates(
File "/usr/local/lib/python3.9/dist-packages/dedupe/core.py", line 170, in scoreDuplicates
raise ChildProcessError from exc
ChildProcessError
当前生成的记录格式
(10430, [{'Rue': 'ROUTE DE Nantes X', 'nom': 'DOIGT', 'Email': None, 'TEL': None, 'SIRET': None}]) (15252, [{'Rue': 'RUE FAT X', 'nom': 'CLOPI', 'Email': None, 'TEL': '227868', 'SIRET': None}])
dedupe要求的正确格式
[ ((1, {'name' : 'Pat', 'address' : '123 Main'}), (2, {'name' : 'Pat', 'address' : '123 Main'})), ((1, {'name' : 'Pat', 'address' : '123 Main'}), (3, {'name' : 'Sam', 'address' : '123 Main'})) ]
问题分析与解决方案
核心问题
- 当前记录的字段数据被包裹在列表中(
[{'Rue': ...}]),但dedupe期望直接传入字典对象,导致代码尝试用字符串索引列表时触发TypeError。 record_pairs函数生成的是单条记录,而非dedupe要求的两两配对的记录元组。
修复步骤
- 调整记录结构:从数据库读取数据时,直接返回字典格式,避免用列表包裹。
- 生成合法配对:使用组合工具生成所有不重复的记录对,格式为
((id1, record_dict1), (id2, record_dict2))。
示例修复代码
def record_pairs(cursor): # 读取所有记录,整理为(id, 字典)格式 records = [] cursor.execute("SELECT id, Rue, nom, Email, TEL, SIRET FROM your_table") for row in cursor: record_id = row[0] record_dict = { 'Rue': row[1], 'nom': row[2], 'Email': row[3], 'TEL': row[4], 'SIRET': row[5] } records.append((record_id, record_dict)) # 生成两两配对(避免重复,用itertools.combinations) from itertools import combinations for pair in combinations(records, 2): yield pair
- 验证格式:修复后生成的记录对需与dedupe要求的格式一致,此时
deduper.score()和deduper.cluster()即可正常运行。
内容的提问来源于stack exchange,提问作者naylo

