You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

嵌套循环匹配DataFrame列值对时触发KeyError问题求助

问题描述
  • 需求:计算DataFrame的dc_term列中,每一对不同单元格之间的匹配值数量
  • 单元格值格式示例:
    ['http://dbpedia.org/resource/Category:American_books,http://dbpedia.org/resource/Category:American_literature_by_medium,http://dbpedia.org/resource/Category:Autobiographies,http://dbpedia.org/resource/Category:Bertelsmann_subsidiaries']
    
  • 编写的代码:
    i = 0
    j = 0
    
    for i in range(len(book_dc.dc_term)):
        values_i = set(book_dc['dc_term'][i].split(','))
        for j in range(i+1, len(book_dc.dc_term)):
            values_j = set(book_dc['dc_term'][j].split(','))
            num_matching = len(values_i.intersection(values_j))
            print("i:", i, "j:", j, "num_matching:", num_matching)
            print('\n')
    
  • 报错信息:
    KeyError                                  Traceback (most recent call last) /usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3360             try: 3361                 return self._engine.get_loc(casted_key) 3362             except KeyError as err:
    
    5 frames pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.Int64HashTable.get_item()
    
    pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.Int64HashTable.get_item()
    
    KeyError: 1
    
    The above exception was the direct cause of the following exception:
    
    KeyError                                  Traceback (most recent call last) /usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3361                 return self._engine.get_loc(casted_key) 3362             except KeyError as err: 3363                 raise KeyError(key) from err 3364  3365         if is_scalar(key) and isna(key) and not self.hasnans:
    
    KeyError: 1
    
问题原因与解决方案

原因

range(len(book_dc.dc_term))生成的是连续整数,但你的DataFrame索引可能不是连续的整数(比如执行过删除行操作后,索引出现断层),此时用整数i去索引book_dc['dc_term'][i]会找不到对应的索引键,从而抛出KeyError。

解决方案

方案1:重置DataFrame索引

先将DataFrame的索引重置为连续的整数,再执行原代码:

# 重置索引,drop=True丢弃原索引列
book_dc = book_dc.reset_index(drop=True)

# 原循环代码
i = 0
j = 0

for i in range(len(book_dc.dc_term)):
    values_i = set(book_dc['dc_term'][i].split(','))
    for j in range(i+1, len(book_dc.dc_term)):
        values_j = set(book_dc['dc_term'][j].split(','))
        num_matching = len(values_i.intersection(values_j))
        print("i:", i, "j:", j, "num_matching:", num_matching)
        print('\n')

方案2:直接遍历列的元素列表

将dc_term列转为普通列表,遍历列表的索引即可避免索引不匹配问题:

# 将列转为列表
dc_terms = book_dc['dc_term'].tolist()

for i in range(len(dc_terms)):
    # 注意:如果单元格是列表格式(如示例中的['xxx,xxx']),需要先取列表第一个元素再split
    # values_i = set(dc_terms[i][0].split(','))
    values_i = set(dc_terms[i].split(','))
    for j in range(i+1, len(dc_terms)):
        # values_j = set(dc_terms[j][0].split(','))
        values_j = set(dc_terms[j].split(','))
        num_matching = len(values_i.intersection(values_j))
        print(f"i: {i}, j: {j}, num_matching: {num_matching}\n")

额外注意点

如果你的单元格值是列表格式(如示例中的['xxx,xxx']),直接调用split会报错,需要先提取列表中的字符串元素,再执行split操作(代码中已注释相关处理方式)。

内容的提问来源于stack exchange,提问作者Mohammed Esam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 02:50:52