You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python dedupe报错TypeError:列表索引应为整数而非字符串求解决

问题:使用dedupe库对接旧版MSSQL时的格式错误问题

问题背景

原本参考MySQL示例实现dedupe对接数据库,但旧版MSSQL无json_object函数,自行编写record_pairs函数后运行报错。

错误栈信息

Process Process-1:
Traceback (most recent call last):
File "/usr/lib/python3.9/multiprocessing/process.py", line 315, in _bootstrap
self.run()
File "/usr/lib/python3.9/multiprocessing/process.py", line 108, in run
self._target(*self._args, **self._kwargs)
File "/usr/local/lib/python3.9/dist-packages/dedupe/core.py", line 80, in call
self.fieldDistance(record_pairs)
File "/usr/local/lib/python3.9/dist-packages/dedupe/core.py", line 91, in fieldDistance
distances = self.data_model.distances(records)
File "/usr/local/lib/python3.9/dist-packages/dedupe/datamodel.py", line 98, in distances
if record_1[field] is not None and record_2[field] is not None:
TypeError: list indices must be integers or slices, not str
TypeError: list indices must be integers or slices, not str

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
File "/root/client/msql2.py", line 286, in
clustered_dupes = deduper.cluster(deduper.score(record_pairs(read_cur)),
File "/usr/local/lib/python3.9/dist-packages/dedupe/api.py", line 109, in score
matches = core.scoreDuplicates(
File "/usr/local/lib/python3.9/dist-packages/dedupe/core.py", line 170, in scoreDuplicates
raise ChildProcessError from exc
ChildProcessError

当前生成的记录格式

(10430, [{'Rue': 'ROUTE DE Nantes X', 'nom': 'DOIGT', 'Email': None, 'TEL': None, 'SIRET': None}])
(15252, [{'Rue': 'RUE FAT X', 'nom': 'CLOPI', 'Email': None, 'TEL': '227868', 'SIRET': None}])

dedupe要求的正确格式

[
    ((1, {'name' : 'Pat', 'address' : '123 Main'}), (2, {'name' : 'Pat', 'address' : '123 Main'})),
    ((1, {'name' : 'Pat', 'address' : '123 Main'}), (3, {'name' : 'Sam', 'address' : '123 Main'}))
]

问题分析与解决方案

核心问题

  1. 当前记录的字段数据被包裹在列表中([{'Rue': ...}]),但dedupe期望直接传入字典对象,导致代码尝试用字符串索引列表时触发TypeError。
  2. record_pairs函数生成的是单条记录,而非dedupe要求的两两配对的记录元组。

修复步骤

  1. 调整记录结构:从数据库读取数据时,直接返回字典格式,避免用列表包裹。
  2. 生成合法配对:使用组合工具生成所有不重复的记录对,格式为((id1, record_dict1), (id2, record_dict2))。

示例修复代码

def record_pairs(cursor):
    # 读取所有记录,整理为(id, 字典)格式
    records = []
    cursor.execute("SELECT id, Rue, nom, Email, TEL, SIRET FROM your_table")
    for row in cursor:
        record_id = row[0]
        record_dict = {
            'Rue': row[1],
            'nom': row[2],
            'Email': row[3],
            'TEL': row[4],
            'SIRET': row[5]
        }
        records.append((record_id, record_dict))
    
    # 生成两两配对(避免重复,用itertools.combinations)
    from itertools import combinations
    for pair in combinations(records, 2):
        yield pair
  1. 验证格式:修复后生成的记录对需与dedupe要求的格式一致,此时deduper.score()和deduper.cluster()即可正常运行。

内容的提问来源于stack exchange,提问作者naylo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 21:24:28